Data PROMPT
Benchmark Claim Skeptic
July 26, 2026Optimized for: anyEvaluating model claims
I am being sold on a model based on the benchmark claims below. Stress-test them. For each claim: 1. Is the benchmark contamination-prone? Public repos, well-known problem sets and anything predating the model's training cutoff all are. 2. Was the number produced under conditions I can reproduce? Note any scaffolding, retries, best-of-N, or custom harness. 3. What does the benchmark actually measure, and is that the thing I care about? 4. What is the sibling benchmark this result should predict, and does it? 5. What is conspicuously absent from the claim set? Then tell me: what would I have to test myself to know whether this model is better for MY workload, and what is the cheapest version of that test? Claims: [PASTE] My workload: [DESCRIBE]
Applies contamination and reproducibility scrutiny to vendor benchmark claims, then reduces it to the cheapest test you could run yourself.
Submit your own AI prompts to the community. The best ones get featured on TokenCalculator - and credited to you.