Blog

Insights, tutorials, and news about LLMs, tokens, and AI tools

Kimi K3.7 Scores 66.7% on SWE-bench with Open Weights. Do the Math on That

Kimi K3.7 Scores 66.7% on SWE-bench with Open Weights. Do the Math on That

Moonshot released a model at $0.60 per million input that beats every Western frontier model from six months ago on agentic coding, and gave away the weights. The implications are not subtle.

Read More
Claude Opus 5 and Sonnet 5 Landed Together. The Sonnet Tier Is the Real Story

Claude Opus 5 and Sonnet 5 Landed Together. The Sonnet Tier Is the Real Story

Opus 5 got the benchmark headlines at 78.1% on SWE-bench Verified. But Sonnet 5 at $3 per million with a 1M window now covers workloads that needed Opus six months ago, and that is where the money is.

Read More
SWE-bench Is Contaminated. Here Is How to Read Coding Benchmarks Now

SWE-bench Is Contaminated. Here Is How to Read Coding Benchmarks Now

Models scoring 79% on SWE-bench Verified land in the mid 50s on the rotating Live variant. Roughly 20 points of recent progress is memorisation. Here is what to trust instead.

Read More
Qwen 3.8 Speaks 119 Languages. Most Frontier Models Still Fumble Six of Them

Qwen 3.8 Speaks 119 Languages. Most Frontier Models Still Fumble Six of Them

Alibaba's Qwen 3.8 pushed tool-calling reliability from 89% to 97% and extended language coverage further than anything else on the market. If you serve users outside the big five languages, this matters.

Read More
GLM-5 Is the Open-Weight Model Nobody Outside China Is Talking About

GLM-5 Is the Open-Weight Model Nobody Outside China Is Talking About

Zhipu's GLM-5 posts 49.7% on SWE-bench Verified with open weights at $0.70 per million input, and fixed the tool-calling reliability that made GLM-4.5 frustrating. It deserves more attention than it is getting.

Read More