Data PROMPT

Benchmark Claim Skeptic

July 26, 2026Optimized for: anyEvaluating model claims

Prompt

I am being sold on a model based on the benchmark claims below. Stress-test them.

For each claim:
1. Is the benchmark contamination-prone? Public repos, well-known problem sets and anything predating the model's training cutoff all are.
2. Was the number produced under conditions I can reproduce? Note any scaffolding, retries, best-of-N, or custom harness.
3. What does the benchmark actually measure, and is that the thing I care about?
4. What is the sibling benchmark this result should predict, and does it?
5. What is conspicuously absent from the claim set?

Then tell me: what would I have to test myself to know whether this model is better for MY workload, and what is the cheapest version of that test?

Claims:
[PASTE]
My workload:
[DESCRIBE]

Tags

Applies contamination and reproducibility scrutiny to vendor benchmark claims, then reduces it to the cheapest test you could run yourself.

Share This Prompt

Related Prompts

Have a Great Prompt to Share?

Submit your own AI prompts to the community. The best ones get featured on TokenCalculator - and credited to you.

Submit a Prompt

Ratings & Feedback

0.0 / 5 · 0 votes

Comments