Benchmark Data — July 2026

LLM Benchmark Scores - MMLU, HumanEval, MATH, GPQA and More

Standardized benchmark results for 20 leading AI models across 8 widely recognised evaluation suites. Scores are sourced from official model cards, research papers and reproducible third-party evaluations, last refreshed July 2026.

MMLU
57 academic subjects from elementary to graduate level.
Scale: 0–100%
HumanEval
164 hand-written Python programming problems.
Scale: 0–100%
MATH
12,500 competition math problems across 5 difficulty levels.
Scale: 0–100%
GPQA
448 expert-level questions in biology, chemistry, and physics.
Scale: 0–100%
GSM8K
8,500 linguistically diverse grade-school math problems.
Scale: 0–100%
SWE-bench
300 real software engineering issues from open-source repos.
Scale: 0–100%
MT-Bench
80 multi-turn questions judged by GPT-4 on a 1–10 scale.
Scale: 1–10
MMMU
11,500 university-level multimodal questions across 30 subjects.
Scale: 0–100%

All Benchmark Results

ModelProvider MMLU HumanEval MATH GPQA GSM8K SWE-bench MT-Bench MMMU
Claude Fable 5 Anthropic 94.6 96.2 95.1 84.7 98.9 79.4 9.7 86.2
Claude Opus 5 Anthropic 94.1 95.8 94.2 83.1 98.7 78.1 9.6 85.3
GPT-5.6 Sol OpenAI 93.8 94.7 97.8 89.6 99.3 74.6 9.4 83.1
GPT-5.6 OpenAI 93.1 95.4 93.4 80.2 98.6 75.8 9.5 82.6
Gemini 3.6 Pro Google 92.4 92.8 92.1 78.4 97.9 70.3 9.4 89.7
Claude Sonnet 5 Anthropic 92.2 94.1 90.8 76.3 97.4 72.6 9.4 82.1
Claude Opus 4.8 Anthropic 92.8 93.6 92.3 77.1 97.6 71.2 9.4 83.4
Gemini 3.6 Flash Google 87.6 88.2 84.7 64.8 93.8 56.4 8.9 84.2
Grok 4 xAI 90.2 90.6 89.4 72.1 96.2 62.8 9.1 79.6
Llama 5 Behemoth Meta 90.8 89.4 88.1 71.4 95.8 60.2 8.9 81.3
Kimi K3.7 Moonshot AI 89.1 92.3 88.6 69.8 95.4 66.7 8.8 74.2
GPT-5.6 Mini OpenAI 87.4 89.6 84.2 63.7 93.1 55.3 8.8 76.4
DeepSeek V4.5 DeepSeek 88.7 90.8 87.3 68.2 95.1 58.6 8.9 72.8
DeepSeek R2 DeepSeek 88.2 87.1 96.4 82.6 99.0 52.1 8.6 68.4
Qwen 3.8 Alibaba 87.9 88.4 85.7 66.1 94.3 54.8 8.7 74.6
Llama 5 Maverick Meta 86.4 85.7 82.1 62.8 93.2 50.4 8.6 79.1
Claude Haiku 4.5 Anthropic 80.4 81.3 74.2 52.1 86.4 38.7 8.3 72.1
GLM-5 Zhipu AI 85.8 86.9 82.6 61.4 92.7 49.7 8.5 70.2
Mistral Large 3.1 Mistral 83.1 82.4 76.8 55.2 89.4 44.1 8.3 68.7
Llama 5 Scout Meta 80.2 79.8 72.4 50.6 86.1 39.8 8.1 73.4
High score Mid score Lower score Best in benchmark

Top Performers per Benchmark

MMLU
94.60–100%
HumanEval
96.20–100%
MATH
97.80–100%
GPQA
89.60–100%
GSM8K
99.30–100%
SWE-bench
79.40–100%
MT-Bench
9.71–10
MMMU
89.70–100%

About Each Benchmark

MMLU — Massive Multitask Language Understanding
57 academic subjects from elementary to graduate level. Scale: 0–100%

Tests breadth of knowledge. High scores indicate strong general academic understanding across STEM, humanities, and professional domains.

#1
Claude Fable 594.6
#2
Claude Opus 594.1
#3
GPT-5.6 Sol93.8
HumanEval — Code Generation from Docstrings
164 hand-written Python programming problems. Scale: 0–100%

Evaluates ability to write correct code from docstrings. Directly relevant for coding assistants and software development use cases.

#1
Claude Fable 596.2
#2
Claude Opus 595.8
#3
GPT-5.695.4
MATH — Competition Mathematics
12,500 competition math problems across 5 difficulty levels. Scale: 0–100%

Measures mathematical reasoning from AMC to AIME difficulty. Correlates with broader reasoning capability.

#1
GPT-5.6 Sol97.8
#2
DeepSeek R296.4
#3
Claude Fable 595.1
GPQA — Graduate-Level Science Q&A
448 expert-level questions in biology, chemistry, and physics. Scale: 0–100%

Expert-level science questions that stump even PhD students. A tough frontier test of genuine understanding vs. pattern matching.

#1
GPT-5.6 Sol89.6
#2
Claude Fable 584.7
#3
Claude Opus 583.1
GSM8K — Grade School Math Word Problems
8,500 linguistically diverse grade-school math problems. Scale: 0–100%

Tests multi-step arithmetic reasoning in plain language. Near-saturated for frontier models - now mainly a baseline check.

#1
GPT-5.6 Sol99.3
#2
DeepSeek R299.0
#3
Claude Fable 598.9
SWE-bench — Real GitHub Issue Resolution
300 real software engineering issues from open-source repos. Scale: 0–100%

Measures real-world software engineering. The gold standard for agentic coding: resolve GitHub issues end-to-end.

#1
Claude Fable 579.4
#2
Claude Opus 578.1
#3
GPT-5.675.8
MT-Bench — Multi-Turn Instruction Following
80 multi-turn questions judged by GPT-4 on a 1–10 scale. Scale: 1–10

LLM-judged multi-turn conversations. Reflects instruction-following quality in realistic chat and assistant scenarios.

#1
Claude Fable 59.7
#2
Claude Opus 59.6
#3
GPT-5.69.5
MMMU — Multimodal University Questions
11,500 university-level multimodal questions across 30 subjects. Scale: 0–100%

The leading multimodal benchmark. Tests chart understanding, scientific diagrams, and cross-modal reasoning simultaneously.

#1
Gemini 3.6 Pro89.7
#2
Claude Fable 586.2
#3
Claude Opus 585.3