Stop Buying Bigger Context Windows. Compaction Beats Them
Agents that summarise and discard their own history outperform agents given a window four times larger. The research is clear, the pricing implications are large, and most teams are still optimising the wrong variable.
Kimi K3.7 Scores 66.7% on SWE-bench with Open Weights. Do the Math on That
Moonshot released a model at $0.60 per million input that beats every Western frontier model from six months ago on agentic coding, and gave away the weights. The implications are not subtle.
SWE-bench Is Contaminated. Here Is How to Read Coding Benchmarks Now
Models scoring 79% on SWE-bench Verified land in the mid 50s on the rotating Live variant. Roughly 20 points of recent progress is memorisation. Here is what to trust instead.
GLM-5 Is the Open-Weight Model Nobody Outside China Is Talking About
Zhipu's GLM-5 posts 49.7% on SWE-bench Verified with open weights at $0.70 per million input, and fixed the tool-calling reliability that made GLM-4.5 frustrating. It deserves more attention than it is getting.
Grok 4 Quietly Closed the Coding Gap, and Nobody Noticed
Grok 4 jumped from 84.3 to 90.6 on HumanEval and shipped a 1M context window. It is still the only model with genuinely live data, and that combination is now hard to ignore.