Which Model Should You Actually Use? A July 2026 Decision Tree
Back to All Posts

Which Model Should You Actually Use? A July 2026 Decision Tree

Since mid June we have had Llama 5, Grok 4, DeepSeek V4.5, Mistral Large 3.1, GLM-5, Qwen 3.8, Claude Opus 5, Claude Sonnet 5, Kimi K3.7, Gemini 3.6 Pro, Gemini 3.6 Flash, GPT-5.6 and GPT-5.6 Sol. That is a lot of releases for eight weeks and the practical question has not changed: which one do you actually put in your application?

Here is a direct answer per situation. Prices are per million tokens, input then output.

By Situation

1. You want one model and you do not want to think about it

Claude Sonnet 5 at $3 / $15. Near-Opus quality, 1M window standard, sane defaults. It is the highest floor of any single choice right now.

2. You are running long autonomous coding agents

Claude Opus 5 at $5 / $25. The compaction layer is the differentiator, not the raw benchmark. Day-long runs actually hold together.

3. You need the absolute best output quality and cost is secondary

Claude Fable 5 at $10 / $50. Top of our leaderboard at 98.4 TC Score.

4. You are solving hard maths, science or research problems

GPT-5.6 Sol at $12 / $72. Best published GPQA and MATH scores. Expensive per completed task because of reasoning token spend. Worth it for the right problems, wasteful for anything else.

5. You are serving high volume and quality tolerance is moderate

Gemini 3.6 Flash at $0.50 / $3.00. Best throughput per dollar near the frontier, 1M window, multimodal input.

6. You are cost constrained and running general workloads

DeepSeek V4.5 at $0.45 / $1.80. Open weights if the hosted region is a problem. Roughly an order of magnitude cheaper than Western equivalents.

7. You are doing agentic coding at volume and want open weights

Kimi K3.7 at $0.60 / $2.50. 66.7% SWE-bench Verified with downloadable weights is currently unmatched at this price.

8. You need documents, video or very long corpora

Gemini 3.6 Pro at $2 / $12. 2M window on every paid tier, native video at 1 fps, and use context caching at $0.20 per million or the cost will surprise you.

9. You serve users outside the major Western languages

Qwen 3.8 at $0.90 / $3.60. 119 languages with tested structured-output quality. Nothing else is close on breadth.

10. You are in the EU with strict data residency requirements

Mistral Large 3.1 at $2 / $6. Residency on by default. The capability gap is real, and for extraction and function calling it does not matter much.

11. You need current, real-world information

Grok 4 at $4 / $20, or Grok 4 Mini at $0.40 / $1.60. The only models with genuinely live grounding at those prices.

12. It has to run on your own hardware

Llama 5 Scout. 1M window and vision on a single 24 GB GPU. Check your quantisation fit on the RAM calculator first.

The Meta-Answer: Route, Do Not Choose

The single most valuable architectural change most teams can make is to stop picking one model. Classify each request as easy or hard, send easy to a cheap tier and hard to an expensive one, and measure the split.

In most production workloads the easy bucket is seventy to ninety percent of traffic. Routing that to Gemini 3.6 Flash or GPT-5.6 Mini instead of a frontier model typically cuts total spend by sixty to eighty percent with no measurable quality change on those requests.

The classifier itself can be a cheap model. It does not need to be clever, only calibrated, and biasing it toward escalation is the safe failure mode.

Check the Numbers Yourself

All prices above are live in our model list and price comparison tool. Model your own traffic on the cost calculator before committing to any of this, because the right answer depends far more on your token mix than on any benchmark.

Try Our Token Calculator

Want to optimize your LLM tokens? Try our free Token Calculator tool to accurately measure token counts for various models.

Go to Token Calculator