July 2026 Roundup

AI News and Developments

A curated overview of the most significant AI announcements, model releases and research breakthroughs of mid 2026. The Claude 5 family landed, OpenAI answered with GPT-5.6, and open weights got genuinely competitive. Updated regularly by the TokenCalculator team.

Featured

Claude Opus 5 and Sonnet 5 Ship Together

Anthropic shipped two Claude 5 models in a single release window: Opus 5 at the familiar $5 per million input tokens, and Sonnet 5 with a 1M context window as standard. Opus 5 posts 78.1% on SWE-bench Verified and holds context across day-long agent runs thanks to a reworked compaction layer. Haiku 4.5 remains the small tier for now. The bigger story is pricing. Frontier quality did not get more expensive this cycle, it got cheaper per unit of work, and the Sonnet tier now covers workloads that needed Opus six months ago.

AnthropicModel LaunchAgentic AI Jul 8, 2026
OpenAI Jul 21, 2026

OpenAI Ships GPT-5.6 and the Sol Long-Thinking Tier

OpenAI answered the Claude 5 launch two weeks later with GPT-5.6, held at GPT-5.5 pricing, plus a new premium tier called Sol that spends far more reasoning tokens per request by default. Sol tops OpenAI internal math and agentic evaluations and posts 89.6% on GPQA, the best public score on that benchmark. The catch is cost. Sol runs at $12 per million input and $72 per million output, so it is a scalpel rather than a daily driver.

GPT-5.6SolReasoningPricing
Google Jul 14, 2026

Gemini 3.6 Pro Brings the 2M Context Window to Every Paid Tier

Google stopped gating the 2 million token window behind enterprise agreements. Gemini 3.6 Pro now exposes it on every paid tier, with context caching that makes repeated queries over the same corpus roughly ten times cheaper than reprocessing. Native video understanding runs at 1 fps and grounding with Google Search cites live sources inline. Gemini 3.6 Flash keeps the 1M window at $0.50 per million input, which remains the best throughput-per-dollar of any hosted frontier-adjacent model.

Gemini 3.6Long ContextContext Caching
Moonshot AI Jul 9, 2026

Moonshot Kimi K3.7 Posts 66.7% on SWE-bench with Open Weights

Moonshot released Kimi K3.7 with a 512K context window and, unusually for a model at this capability level, full open weights. It scores 66.7% on SWE-bench Verified, which puts it ahead of every Western model from six months ago, at $0.60 per million input tokens. A distilled Mini variant runs on a single high-memory GPU. The release reset expectations for what open weights can do on long agentic coding runs.

Kimi K3.7Open WeightAgentic CodingChina
Alibaba Jul 2, 2026

Alibaba Qwen 3.8 Extends Multilingual Lead to 119 Languages

Qwen 3.8 ships with a 256K context window, sparse mixture-of-experts routing and first-class support for 119 languages, including several that no Western frontier model handles cleanly. Tool-calling reliability jumped from roughly 89% to 97% on Alibaba internal harnesses. A separate Qwen 3.8 Coder variant carries a 1M window and targets repository-scale work at $0.50 per million input tokens.

Qwen 3.8MultilingualOpen WeightTool Calling
Anthropic Jun 24, 2026

Anthropic Introduces the Mythos Class with Fable 5 and Mythos 5

Anthropic split its top of the range into two: Mythos 5, distributed under strict access controls, and Fable 5, the first publicly available Mythos-class model. Fable 5 shares the underlying intelligence and adds safety classifiers that fall back to Opus 4.8 on sensitive prompts. It holds the top spot on our TC Score at 98.4 and posts 79.4% on SWE-bench Verified. Adaptive thinking replaces fixed thinking budgets, so the model decides how long to reason rather than being told.

Fable 5Mythos 5FrontierSafety
Meta Jun 17, 2026

Meta Releases the Llama 5 Family Including a 2M Context Behemoth

Meta shipped three Llama 5 models under a revised community license: Scout for local and edge work on a single 24 GB GPU, Maverick as the general purpose mixture-of-experts model, and Behemoth, the largest open-weight model released to date. Behemoth scores 90.8 on MMLU and is served through partners rather than being practical to self-host. Scout is the more consequential release for most developers, since it puts a 1M window and vision input on consumer hardware.

Llama 5Open WeightEdge AIVision
xAI Jun 11, 2026

xAI Grok 4 Adds a 1M Window and Closes the Coding Gap

Grok 4 arrived with a 1M token context window, live X and web grounding retained from Grok 3.5, and a jump from 84.3 to 90.6 on HumanEval. The persistent cross-conversation memory introduced last year now syncs between the X app and the API. A Grok 4 Mini tier at $0.40 per million input keeps live grounding, which no other cheap model offers.

Grok 4Live GroundingCodingxAI
DeepSeek Jun 4, 2026

DeepSeek V4.5 Doubles Context and Keeps the Price Floor

DeepSeek pushed a mid-cycle refresh of V4 with a 256K window, better tool calling and noticeably improved multilingual output. Pricing stayed at $0.45 per million input and $1.80 per million output, which remains the cheapest way to run a model at this capability level. Weights are open. Together with R2, the reasoning sibling that exposes full chain of thought, DeepSeek now anchors the bottom of the price curve across both general and reasoning workloads.

DeepSeek V4.5Open WeightPricingLong Context
Mistral AI May 26, 2026

Mistral Large 3.1 Makes EU Data Residency the Default

Mistral shipped Large 3.1 with a 256K window, a stricter JSON mode that fails loudly instead of silently repairing malformed output, and EU data residency enabled by default on La Plateforme rather than as an add-on. For regulated European teams that single default change matters more than the benchmark deltas. A refreshed Mistral Embeddings v3 launched alongside with better cross-lingual retrieval.

MistralGDPRStructured OutputEurope
Zhipu AI May 19, 2026

Zhipu GLM-5 Lands as a Serious Bilingual Open-Weight Option

Zhipu released GLM-5, a 256K context mixture-of-experts model with strong performance in both Chinese and English and a 49.7% SWE-bench Verified score. At $0.70 per million input it sits between the Qwen and DeepSeek tiers. The notable detail is agentic tool use, where GLM-5 handles multi-step tool chains far more reliably than GLM-4.5 did, which had been the model family main weakness.

GLM-5Open WeightBilingualAgentic
Policy May 12, 2026

EU AI Act Codes of Practice Bite for General-Purpose Models

The general-purpose AI provisions of the EU AI Act moved from published guidance to active enforcement. Providers above the systemic-risk compute threshold must now file model documentation, training-data summaries and adversarial evaluation results. OpenAI, Anthropic, Google, Mistral and Meta have all published compliance packets. The practical effect for developers is that EU-facing terms of service changed again, so it is worth rereading the usage sections before shipping.

EU AI ActRegulationCompliancePolicy
Research May 5, 2026

SWE-bench Verified Gets a Contamination-Resistant Successor

With top models now clearing 75% on SWE-bench Verified, the benchmark authors published SWE-bench Live, which draws issues from repositories after each model training cutoff and rotates monthly. Early results are sobering: the same models that score in the high 70s on Verified land in the mid 50s on Live. The gap is the clearest evidence yet that a meaningful chunk of recent coding-benchmark progress is contamination rather than capability.

SWE-benchBenchmarksContaminationEvaluation
Anthropic Apr 3, 2026

Claude Opus 4.6 Tops Multiple Benchmarks and Cements Its Lead in Agentic Tasks

Anthropic Claude Opus 4.6 emerged as the highest-rated model on the LMSYS Chatbot Arena, surpassing GPT-5.4 and Gemini 3.1 Pro in head-to-head human preference evaluations. The model achieved record scores on SWE-bench Verified at 65.3%, reflecting a step change in agentic software engineering. Anthropic attributed the improvement to a hybrid architecture combining standard transformer layers with a sparse mixture-of-experts component for routing reasoning-heavy tokens.

Claude Opus 4.6BenchmarksAgentic AI
OpenAI Apr 5, 2026

OpenAI GPT-5.4 Receives Major Update with Improved Instruction Following

OpenAI pushed a significant update to GPT-5.4 that reduced refusals on benign edge-case requests by 40% while maintaining safety. The update also extended context window handling and improved performance on multi-document analysis tasks. GPT-5.4 Mini received a parallel update bringing its coding performance closer to the full model at a fraction of the cost.

GPT-5.4Model UpdateInstruction Following
Google Apr 1, 2026

Google Gemini 3.1 Pro Now Available in Vertex AI with 2M Context

Google made Gemini 3.1 Pro generally available on Vertex AI, bringing enterprise-grade 2 million token context to production workloads. New features included document-level caching for entire books or codebases, native video understanding at 1 fps, and improved grounding with Google Search integration that cites live web sources in responses.

GeminiLong ContextVertex AI
Meta Mar 28, 2026

Meta Releases Llama 4 Scout: A Compact VLM for Edge Deployment

Meta open-sourced Llama 4 Scout, a 17 billion parameter vision-language model optimized for edge devices. Scout achieved competitive vision benchmark scores while running at full speed on a single consumer GPU with 24 GB VRAM or an Apple M4 Pro. The model supports image, video frame and PDF inputs and shipped under the Llama 4 Community License.

Llama 4Open WeightVisionEdge AI
DeepSeek Mar 25, 2026

DeepSeek R2 Achieves State-of-the-Art on AIME 2025 and Competitive Math

DeepSeek released R2, their reasoning model, which achieved 92.7% on AIME 2025 and 89.4% on MATH-500, numbers that rivalled the best Western reasoning endpoints. R2 shipped via API at roughly 70% lower cost than comparable models, reigniting the debate about AI cost efficiency East versus West. An open-weight distilled 32B version was released alongside it.

DeepSeekReasoningMathOpen Weight
Anthropic Mar 20, 2026

Anthropic Launches Claude Code Agent: A Full Agentic Coding Environment

Anthropic launched Claude Code as a standalone product, a terminal-native AI agent that can clone repositories, write and run tests, fix failing CI pipelines and open pull requests autonomously. Claude Code integrates with GitHub, GitLab and Jira out of the box and supports sandboxed execution for safe code running.

Claude CodeAgentic AICodingDevTools
xAI Mar 18, 2026

xAI Grok 3 Adds Real-time Image Generation and Persistent Memory

xAI updated Grok 3 with two headline features: real-time image generation using a proprietary diffusion model integrated directly in the chat UI, and Grok Memory, a persistent cross-conversation context that remembers user preferences, past projects and key facts. The underlying API became available to enterprise customers.

Grok 3Image GenerationMemoryxAI
Policy Mar 15, 2026

EU AI Act Enters Full Enforcement: What It Means for LLM Providers

The EU AI Act entered full enforcement in March 2026, requiring all AI systems deployed in the EU to meet transparency, safety and risk classification requirements. High-risk applications must maintain detailed logs and pass conformity assessments. Major providers published their general-purpose AI compliance documentation.

EU AI ActRegulationCompliancePolicy
Mistral AI Mar 10, 2026

Mistral Large 3 Released with Improved Function Calling and EU Data Residency

Mistral AI released Large 3 with significant improvements to structured output generation, function calling accuracy and JSON mode reliability. Mistral began offering EU data residency through La Plateforme, making it a top choice for European enterprises with strict GDPR data locality requirements. Mistral Embeddings v2 launched alongside.

MistralFunction CallingGDPREurope
Research Mar 5, 2026

HuggingFace Open LLM Leaderboard v3 Launches with New Benchmarks

HuggingFace revamped the Open LLM Leaderboard with benchmarks designed to be more contamination resistant and practically meaningful: IFEval-Hard, MATH-Verify, LiveCodeBench-2026 and a new multi-turn conversation benchmark. The refresh reshuffled rankings significantly among open-weight models. Leaderboard data refreshes weekly.

HuggingFaceBenchmarksOpen SourceLeaderboard
OpenAI Feb 28, 2026

OpenAI o3-mini Becomes the Default Reasoning Model for ChatGPT Plus

OpenAI replaced o1-mini with o3-mini as the default reasoning model for ChatGPT Plus subscribers, citing roughly 3x faster responses at equivalent or better quality on math and science tasks. OpenAI also launched Flex compute, a pricing tier offering a 30% discount during off-peak hours.

o3ChatGPTReasoningPricing

Research Highlights (Q2 2026)

Context Compaction Beats Bigger Windows

A widely cited paper shows that agents using learned compaction, where the model summarises and discards its own history, outperform agents given a raw window four times larger on day-long tasks. Attention over a huge unfiltered window turns out to be worse than attention over a curated small one.

Benchmark Contamination Quantified at Scale

By comparing model scores on issues published before and after each training cutoff, researchers estimate that 15 to 25 points of recent SWE-bench progress is memorisation rather than generalisation. The finding drove the launch of rotating, contamination-resistant benchmark suites.

Sparse Attention Cuts Long-Context Cost by 70%

A joint academic and industry paper demonstrates that learned sparse attention patterns retain over 99% of dense-attention quality at million-token lengths while cutting inference FLOPs by roughly 70%, which is a large part of why 1M windows became standard pricing this year.

Interpretability Reaches Production Safety Tooling

Feature-level probes derived from sparse autoencoders moved out of research and into deployed safety classifiers, letting providers detect specific unsafe behaviours by inspecting internal activations rather than only by filtering output text.

2026 Model Releases

GPT-5.6 + Sol
OpenAI · Jul 2026
Gemini 3.6 Pro / Flash
Google · Jul 2026
Kimi K3.7
Moonshot AI · Jul 2026
Claude Opus 5 + Sonnet 5
Anthropic · Jul 2026
Qwen 3.8
Alibaba · Jul 2026
Claude Fable 5 + Mythos 5
Anthropic · Jun 2026
Llama 5 Scout / Maverick / Behemoth
Meta · Jun 2026
Grok 4
xAI · Jun 2026
DeepSeek V4.5
DeepSeek · Jun 2026
Gemini 3.5 Flash
Google · May 2026
Claude Opus 4.8
Anthropic · May 2026
Mistral Large 3.1
Mistral · May 2026
GLM-5
Zhipu AI · May 2026
GPT-5.5 + GPT-5.5 Pro
OpenAI · Apr 2026
Claude Opus 4.7
Anthropic · Mar 2026
DeepSeek R2
DeepSeek · Mar 2026

AI in Numbers (2026)

Models tracked here130+
Largest context window2M tokens
Cheapest frontier-class input$0.45/M
Fastest model (tok/s)~1400
SWE-bench Verified top score79.4%