Coding PROMPT
Eval Set Builder from Git History
July 26, 2026Optimized for: anyBuilding private model evals
Build me a model evaluation set from real work instead of a public benchmark. Given the closed issues and their merged fixes below, produce 20 eval tasks. For each: 1. TASK: the issue restated as a self-contained instruction, with no hints about the actual fix 2. CONTEXT NEEDED: which files a model would have to read to solve it 3. PASS CRITERIA: an objective check, ideally a test command, that distinguishes a real fix from a plausible one 4. DIFFICULTY: trivial / moderate / hard 5. TRAP: what a model is most likely to get subtly wrong here Exclude issues that are pure dependency bumps, typo fixes, or anything where the fix is stated in the issue title. Issues and fixes: [PASTE]
Converts your own closed issues into a private eval set. Far more predictive than public benchmarks and immune to contamination.
Submit your own AI prompts to the community. The best ones get featured on TokenCalculator - and credited to you.