AI Model Release Timeline

Every model we track, in the order it shipped, with launch pricing, context window and the two capabilities that defined it. 142 releases across 20 providers, most recent Jul 2026.

Jump to
July 2026 14
Anthropic

Anthropic Claude Opus 5 (200k)

200K$5/M in / $25/M out

Anthropic's flagship Opus-class model for July 2026. Opus 5 pairs Mythos-class reasoning with the Opus price point, adds adaptive effort control and a hardened long-horizon agent loop that holds context across day-long tasks.

Adaptive effort control (low to max)Best-in-class agentic coding
Anthropic

Anthropic Claude Opus 5 (1M context)

1M$6/M in / $30/M out

The 1M-token context tier of Claude Opus 5. Same model, long-context premium pricing applied to requests above 200K input tokens. Built for whole-repository and whole-corpus work.

1M token context windowWhole-repository reasoning
Anthropic

Anthropic Claude Sonnet 5 (1M)

1M$3/M in / $15/M out

The workhorse of the Claude 5 family. Sonnet 5 lands close to Opus 4.8 on coding and agentic benchmarks while staying at the long-standing Sonnet price, with a 1M token context window available by default.

1M token context windowNear-Opus coding quality
OpenAI

OpenAI GPT-5.6 Sol (1M)

1M$12/M in / $72/M out

OpenAI's Sol tier for GPT-5.6: a long-thinking configuration that spends far more reasoning tokens per request and tops OpenAI's own agentic and math evaluations. Priced as a premium reasoning endpoint.

Extended reasoning by default~1M token context window
OpenAI

OpenAI GPT-5.6 (1M)

1M$5/M in / $30/M out

The July 2026 refresh of the GPT-5 series. GPT-5.6 keeps GPT-5.5 pricing while improving instruction following, tool-call accuracy and refusal calibration, and it ships with a faster streaming stack.

~1M token context windowImproved instruction following
OpenAI

OpenAI GPT-5.6 Mini (1M)

1M$0.6/M in / $3.6/M out

The cost tier of GPT-5.6. Mini keeps the same tokenizer and tool interface as the full model, trading depth for roughly an eighth of the price and noticeably lower latency.

~1M token context windowVery low cost per million tokens
OpenAI

OpenAI GPT-5.6 Codex (1M)

1M$5/M in / $30/M out

The Codex-tuned build of GPT-5.6 for the Codex CLI, IDE extension and cloud agents. Optimised for long autonomous coding sessions, patch generation and test repair rather than open-ended chat.

Tuned for agentic codingLong autonomous sessions
Google

Google Gemini 3.6 Pro (2M)

2M$2/M in / $12/M out

Google's flagship Gemini for mid-2026. Gemini 3.6 Pro extends the 2M token window to all paid tiers, improves grounded answers with live Search, and adds native 1 fps video understanding.

2M token context windowNative video understanding at 1 fps
Google

Google Gemini 3.6 Flash (1M)

1M$0.5/M in / $3/M out

The fast, cheap tier of Gemini 3.6. Flash keeps a 1M token window and multimodal input while running several times faster than Pro, which makes it the default choice for high-volume production traffic.

1M token context windowVery high throughput
Moonshot AI

Moonshot Kimi K3.7 (512k)

512K$0.6/M in / $2.5/M out

Moonshot's flagship Kimi K3.7, an open-weight mixture-of-experts model with a 512K token window. K3.7 closes most of the gap to Western frontier models on agentic coding while staying dramatically cheaper.

512K token context windowOpen-weight MoE architecture
Moonshot AI

Moonshot Kimi K3.7 Turbo (512k)

512K$1.2/M in / $5/M out

The low-latency hosted tier of Kimi K3.7. Same weights, dedicated high-throughput serving, roughly double the price and a large jump in tokens per second.

Turbo serving tier512K token context window
Moonshot AI

Moonshot Kimi K3.7 Mini (256k)

256K$0.15/M in / $0.8/M out

A distilled K3.7 aimed at classification, extraction and sub-agent duty. Runs comfortably on a single high-memory GPU when self-hosted.

Distilled from K3.7256K token context window
Alibaba

Alibaba Qwen 3.8 (256k)

256K$0.9/M in / $3.6/M out

Alibaba's Qwen 3.8 flagship. A sparse MoE model with 256K context, first-class support for more than 100 languages and a notable jump in tool-calling reliability over Qwen 3.

256K token context window100+ language support
Alibaba

Alibaba Qwen 3.8 Coder (1M)

1M$0.5/M in / $2/M out

The coding-specialised Qwen 3.8 with a 1M token window. Tuned for repository-scale understanding, patch generation and terminal agents.

1M token context windowRepository-scale code understanding
June 2026 8
Anthropic

Anthropic Claude Mythos 5 (1M)

1M$12/M in / $60/M out

Anthropic's most powerful frontier model - the Mythos-class system with state-of-the-art agentic and cybersecurity capability. Distributed under strict access controls; Fable 5 is its safeguarded public sibling.

Mythos-class frontier intelligenceState-of-the-art long-horizon agents
Anthropic

Anthropic Claude Fable 5 (1M)

1M$10/M in / $50/M out

The first public Mythos-class model - Anthropic's most powerful generally available model, sharing Mythos 5's underlying intelligence with added safeguards that fall back to Opus 4.8 on sensitive prompts.

Mythos-class intelligence for everyoneAdaptive thinking (no fixed budget)
Meta

Meta Llama 5 Maverick (2M)

2M$0.35/M in / $1.1/M out

Meta's flagship open-weight model for 2026. Llama 5 Maverick is a large mixture-of-experts system with a 2M token window and native multimodal input, released under the Llama 5 Community License.

2M token context windowMixture-of-experts architecture
Meta

Meta Llama 5 Scout (1M)

1M$0.15/M in / $0.5/M out

The compact member of the Llama 5 family. Scout runs on a single 24 GB consumer GPU at full speed while keeping a 1M token window and vision input, which makes it the practical choice for local and edge work.

1M token context windowRuns on a single 24 GB GPU
Meta

Meta Llama 5 Behemoth (2M)

2M$1.2/M in / $4/M out

Meta's largest released model, positioned against closed frontier systems. Behemoth is served through partners rather than being practical to self-host, and it leads open-weight benchmarks on reasoning and math.

Largest open-weight model to date2M token context window
DeepSeek

DeepSeek V4.5 (256k)

256K$0.45/M in / $1.8/M out

A mid-cycle refresh of DeepSeek V4 with a doubled context window, better tool calling and a large improvement in multilingual output quality. Still open weights.

256K token context windowImproved tool calling
xAI

xAI Grok 4 (1M)

1M$4/M in / $20/M out

xAI's Grok 4 with a 1M token window, live X and web grounding, and a large jump in coding benchmarks over Grok 3.5. Available through the xAI API and X Premium+.

1M token context windowLive X and web grounding
xAI

xAI Grok 4 Mini (1M)

1M$0.4/M in / $1.6/M out

The cost tier of Grok 4. Keeps live grounding and the 1M window but targets high-volume traffic at roughly a tenth of the flagship price.

1M token context windowLive grounding retained
May 2026 4
Anthropic

Anthropic Claude Opus 4.8 (1M)

1M$5/M in / $25/M out

Anthropic's most capable Opus-tier model - highly autonomous, state-of-the-art on long-horizon agentic work, knowledge work and memory, with clearer, warmer writing. 1M context at standard pricing.

Adaptive thinking + effort (low→max)State-of-the-art long-horizon agents
Google

Google Gemini 3.5 Flash (1M)

1M$1.5/M in / $9/M out

Google's newest fast model (Google I/O 2026), built for agentic tasks and coding - beats Gemini 3.1 Pro on coding and agentic benchmarks at a fraction of the cost.

Fast, low-cost agentic model1M token context window
Mistral AI

Mistral Large 3.1 (256k)

256K$2/M in / $6/M out

Mistral's refreshed flagship with a 256K window, stricter JSON mode and EU data residency through La Plateforme. The default pick for European teams with GDPR locality requirements.

256K token context windowEU data residency
Zhipu AI

Zhipu GLM-5 (256k)

256K$0.7/M in / $2.4/M out

Zhipu's GLM-5, an open-weight MoE model with strong bilingual Chinese and English performance and a competitive agentic coding score at a fraction of frontier pricing.

256K token context windowOpen-weight MoE
April 2026 2
OpenAI

OpenAI GPT-5.5 (1M)

1M$5/M in / $30/M out

OpenAI's flagship GPT-5.5, released April 2026. A ~1.05M token context window with strong terminal/agentic benchmarks; available in the Responses and Chat Completions APIs.

~1.05M token context window128K max output tokens
OpenAI

OpenAI GPT-5.5 Pro (1M)

1M$30/M in / $180/M out

A higher-accuracy GPT-5.5 variant for the most demanding tasks, available via the API at premium pricing.

Maximum accuracy GPT-5.5 tier~1.05M token context window
March 2026 6
OpenAI

GPT-5.4

256K$25/M in / $100/M out

OpenAI's most advanced model with enhanced reasoning, longer context, and multimodal capabilities. Top of the line for complex tasks.

256K context windowAdvanced reasoning
OpenAI

GPT-5.4 Mini

128K$3/M in / $12/M out

Affordable version of GPT-5.4 with strong performance for everyday tasks at a fraction of the cost.

128K contextFast inference
Google

Gemini 3.1 Flash

1M$0.35/M in / $1.4/M out

Fast and affordable Gemini 3.1 Flash optimized for high-throughput applications at minimal cost.

1M context windowUltra-fast inference
xAI

xAI Grok 3.5 (256k)

256K$5/M in / $25/M out

xAI's latest flagship model with expanded context, improved reasoning, and deeper integration with real-time data from X platform.

256K token context windowAdvanced reasoning and coding
Anthropic

Anthropic Claude Opus 4.7 (1M)

1M$5/M in / $25/M out

Previous-generation Opus - highly autonomous with excellent long-horizon agentic work, vision and memory. Adaptive thinking only; high-resolution vision support.

Adaptive thinking with xhigh effortHigh-resolution vision (2576px)
DeepSeek

DeepSeek R2 (256k)

256K$0.4/M in / $1.6/M out

DeepSeek's second-generation reasoning model. R2 exposes its full chain of thought, scores near the top of competitive math benchmarks and is priced far below comparable Western reasoning endpoints.

Open reasoning traces256K token context window
February 2026 5
Anthropic

Anthropic Claude Opus 4.6 (1M)

1M$18/M in / $90/M out

Anthropic's Claude Opus 4.6 with extended 1M token context window for processing entire codebases, books, and massive datasets in a single prompt.

1M token context windowExtended thinking for complex reasoning
Anthropic

Anthropic Claude Opus 4.6 (200k)

200K$15/M in / $75/M out

Anthropic's latest flagship model with best-in-class reasoning, coding, and agentic capabilities. Supports extended thinking for complex multi-step problems.

Extended thinking for complex reasoning200K token context window
Anthropic

Anthropic Claude Sonnet 4.6 (200k)

200K$3/M in / $15/M out

Anthropic's balanced mid-tier model offering strong intelligence, speed, and cost-effectiveness. Excellent for everyday coding and enterprise tasks.

Strong balance of intelligence and speed200K token context window
Google

Google Gemini 3.1 Pro (2M)

2M$2/M in / $12/M out

Google's most capable Gemini Pro model with an industry-leading 2M token context window - ideal for large-document analysis and long multi-turn reasoning.

Industry-leading 2M context windowDeep multimodal reasoning
DeepSeek

DeepSeek V4 (128k)

128K$0.5/M in / $2/M out

DeepSeek's fourth-generation open-weights model with state-of-the-art reasoning at remarkably low cost. Strong performance on coding and math benchmarks.

128K token context windowOpen-weights model
January 2026 4
OpenAI

OpenAI o3 Deep Research (200k)

200K$20/M in / $80/M out

An o3-family model optimized for deep research tasks. Autonomously browses the web, synthesizes information, and produces comprehensive research reports.

Autonomous web browsing and research200K token context window
Mistral AI

Mistral Large 3 (128k)

128K$2/M in / $6/M out

Mistral's latest flagship model with multilingual excellence, strong coding, and enterprise-grade function calling. Open-weight model.

128K token context windowMultilingual excellence
Alibaba

Alibaba Qwen 3 (128k)

128K$0.8/M in / $3.2/M out

Alibaba's latest Qwen 3 flagship model with strong multilingual capabilities and improved reasoning. Competitive with frontier models at lower cost.

128K context windowStrong multilingual support (100+ languages)
Midjourney

Midjourney v7

N/A

Professional AI image generation with photorealistic quality and artistic control. Subscription-based: Basic $10/mo, Standard $30/mo, Pro $60/mo

Photorealistic outputsAdvanced style control
December 2025 3
OpenAI

OpenAI GPT-5.2 Pro (400k)

400K$21/M in / $168/M out

Higher-compute variant of GPT-5.2 for harder problems (Responses API only).

Responses API onlySupports higher reasoning effort
OpenAI

OpenAI GPT-5.1-Codex-Max (400k)

400K$1.25/M in / $10/M out

Most intelligent Codex model optimized for long-horizon agentic coding (Responses API only).

Long-horizon agentic codingResponses API only
Google

Google Gemini 3 Flash Preview (1M)

1M$0.5/M in / $3/M out

Gemini 3 Flash Preview on Vertex AI.

1M input contextMultimodal
November 2025 5
OpenAI

OpenAI GPT-5.2 (400k)

400K$1.75/M in / $14/M out

OpenAI flagship model for coding and agentic tasks across industries.

Text & image input, text output400K context window
OpenAI

OpenAI GPT-5.1 (400k)

400K$1.75/M in / $14/M out

Flagship GPT model with configurable reasoning effort; predecessor to GPT-5.2.

Configurable reasoning/non-reasoning effort400K context window
OpenAI

OpenAI GPT-5.1-Codex (400k)

400K$1.25/M in / $10/M out

GPT-5.1 variant optimized for agentic coding in Codex (Responses API only).

Agentic coding specializationResponses API only
Anthropic

Anthropic Claude Opus 4.5 (200k)

200K$5/M in / $25/M out

Claude 4.5 flagship Opus model for long-horizon coding and agentic workflows.

Extended thinking supportStrong long-horizon coding/agents
Google

Google Gemini 3 Pro Preview (1M)

1M$2/M in / $12/M out

Gemini 3 Pro Preview on Vertex AI.

1M context windowPremium long-context billing beyond 200K
October 2025 3
Anthropic

Anthropic Claude Haiku 4.5 (200k)

200K$1/M in / $5/M out

Anthropic's fastest 4.5-series model-cost-effective for high-volume workloads with extended thinking and strong tool use.

Extended thinking (reasoning) mode200K token context window
Cognition / Windsurf

Cognition SWE-1.5 (Windsurf in-house)

256K$0.5/M in / $2/M out

Windsurf/Cognition in-house frontier model for agentic coding. Consumed via Windsurf prompt credits (not USD per-token).

Agentic coding-optimizedNear SOTA performance at very high speed
Anthropic

Anthropic Claude Haiku 4.5 (200k)

200K$0.8/M in / $4/M out

Anthropic's fastest and most cost-effective 4.5-series model, optimized for high-volume workloads with strong tool use capabilities.

Fast and cost-effective200K token context window
September 2025 3
OpenAI

OpenAI GPT-5-Codex (400k)

400K$1.25/M in / $10/M out

A version of GPT-5 optimized for agentic coding in Codex. Default in Codex cloud & reviews; also usable via API key. Priced the same as GPT-5.

Optimized for agentic software engineeringStronger code review & large-refactor ability
Anthropic

Anthropic Claude Sonnet 4.5 (200k/1M beta)

1M$3/M in / $15/M out

Anthropic's most intelligent model for agents and coding, with extended thinking and state-of-the-art performance on SWE-bench Verified.

Extended thinking for long-horizon tasksBest-in-class coding and agent performance
Google

Google Gemini 3.0 (2M)

2M

Next-generation Gemini with advanced multimodal reasoning and long context. Rates pending official pricing page.

Advanced multimodal (text, image, audio, video)2M token context window
August 2025 10
Anthropic

Anthropic Claude Opus 4.1 (200k)

200K$15/M in / $75/M out

Opus 4.1 (200k) - higher-cost flagship Opus tier. Newer Opus 4.5 provides a more accessible price point.

Improved long-term memory supportSignificantly better performance on SWE-bench Verified
OpenAI

OpenAI GPT-5 Pro (400k)

400K$15/M in / $120/M out

Version of GPT-5 that produces smarter and more precise responses. Responses API only; higher max output than standard GPT-5.

Highest reasoning depth in GPT-5 familyResponses API only
OpenAI

OpenAI GPT-5 (400k)

400K$1.25/M in / $10/M out

Previous flagship GPT model for coding, reasoning, and agentic tasks. OpenAI recommends GPT-5.1/5.2 for newest improvements.

Text & vision400K token context window (128K max output)
OpenAI

OpenAI GPT-5 Mini (400k)

400K$0.25/M in / $2/M out

A faster, lower-cost GPT-5 for well-defined tasks. Text & vision with long context at a fraction of the price.

Cost-efficient reasoningText & vision
OpenAI

OpenAI GPT-5 Nano (400k)

400K$0.05/M in / $0.4/M out

Fastest, most cost-efficient GPT-5 variant.

Very low latencyGreat for classification/summarization
OpenAI

OpenAI GPT-OSS 120B (Open-Weight)

128K$0.1/M in / $0.5/M out

Open-weight model entry as listed by OpenAI (see models page). Token costs depend on where you run it.

Open-weightRuns locally/on your infra
OpenAI

OpenAI GPT-OSS 20B (Open-Weight)

128K$0.05/M in / $0.2/M out

Open-weight model entry as listed by OpenAI (see models page). Token costs depend on where you run it.

Open-weightLower-latency footprint
OpenAI

OpenAI GPT-5 Low (1M)

1M$0.25/M in / $2/M out

Alias tier aligned with GPT-5 Mini pricing; use when your workflow targets the "low" cost tier.

Cost-efficient tierGood for high-volume tasks
OpenAI

OpenAI GPT-5 Medium (1M)

1M

Mid-tier GPT-5 variant balancing performance and cost for general-purpose workloads.

Balanced price/performanceStrong reasoning and coding
OpenAI

OpenAI GPT-5 High (1M)

1M

High-tier GPT-5 with enhanced reasoning, reliability and coding for mission-critical workloads.

Advanced reasoningImproved factuality
July 2025 5
Moonshot AI

Moonshot Kimi K2 Base (128k)

128K$0.15/M in / $2.5/M out

A 1T parameter open-weight Mixture-of-Experts (MoE) model with 32B active parameters. This is the unaligned, pre-trained base model, suitable for further fine-tuning.

1T total parameters (32B active)Mixture-of-Experts (MoE) architecture
Moonshot AI

Moonshot Kimi K2 Instruct (128k)

128K$0.15/M in / $2.5/M out

The instruction-tuned version of Kimi K2, optimized for chat, agentic tasks, and tool use. Aligned with RLHF for helpful and safe responses.

Instruction-tuned for chat and agentic tasksOptimized for tool use
Alibaba

Alibaba Qwen3 Coder Flash (1M)

1M$0.3/M in / $1.5/M out

A 30B parameter model from the Qwen3 series, excelling in coding and agentic tasks with a 1M token context length.

30 billion parameters1M token context length
Zhipu AI

Zhipu GLM-4.5 (200k)

200K$0.6/M in / $2.2/M out

GLM-4.5 agentic foundation model. Official pricing is published in RMB on BigModel; USD prices vary by provider.

Expanded context window up to 200KStrong coding + agentic capabilities
Alibaba

Alibaba Qwen3 Coder (1M)

1M

Qwen3 coding-specialized model with long-context capabilities and strong tool-use.

Long-context code understandingTool calling
June 2025 1
OpenAI

OpenAI o3-pro (200k)

200K$20/M in / $80/M out

Version of o3 with more compute for better responses (Responses API only).

Higher-compute o3Responses API only
May 2025 8
Anthropic

Anthropic Claude Opus 4 (200k)

200K$15/M in / $75/M out

Anthropic's most powerful model, excelling in coding, advanced reasoning, and AI agent workflows. Handles complex, long-running tasks.

Extended thinking with tool use (beta)Parallel tool use
Anthropic

Anthropic Claude Sonnet 4 (200k)

200K$3/M in / $15/M out

Anthropic's highly capable and versatile model, offering a strong balance of intelligence, speed, and cost-effectiveness for enterprise applications.

Improved coding and reasoning over Claude 3.5 SonnetEnhanced vision capabilities
Google

Google Gemini 2.5 Pro (1M)

1M$1.25/M in / $10/M out

Google's most advanced reasoning Gemini model, capable of solving complex problems. Supports text, code, image, audio, and video inputs. Features a 1M token context window (up to 2M in some versions).

Advanced reasoning ('thinking model')Native multimodal (text, code, image, audio, video)
Mistral AI

Mistral Medium 3 (128k)

128K$0.4/M in / $2/M out

Mistral AI's frontier-class multimodal model balancing SOTA performance, lower cost, and simpler deployability for enterprise usage. Excels in coding and multimodal understanding.

Multimodal capabilities128K token context window
Mistral AI

Mistral Devstral Small (128k)

128K$0.1/M in / $0.3/M out

A 24B open-source text model from Mistral AI that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents. Apache 2.0 license.

24 billion parametersExcels at tool use for codebases
Google

Google Gemini 2.5 Flash (1M)

1M$0.15/M in / $0.6/M out

Google's best model for price and performance (as of May 2025), featuring hybrid reasoning capabilities. Supports text, code, image, audio, and video inputs. 1M token context window.

Hybrid reasoning with configurable thinking budgetMultimodal (text, code, image, audio, video)
OpenAI

OpenAI Codex Mini (200k)

200K$1.5/M in / $6/M out

OpenAI's fast coding model designed for the Codex coding agent. Optimized for rapid code generation, editing, and review within development workflows.

Optimized for fast code generation200K token context window
Pika Labs

Pika 2

N/A

Fast and creative video generation with emphasis on artistic styles and special effects. Free tier available with paid options.

Fast generationCreative effects
April 2025 11
OpenAI

OpenAI o3 (200k)

200K$2/M in / $8/M out

OpenAI reasoning model for complex tasks (text+image input, text output). Succeeded by GPT-5.x for many agentic workloads.

Advanced reasoningText and image processing
OpenAI

OpenAI o4-mini (200k)

200K$1.1/M in / $4.4/M out

A faster, cost-efficient reasoning model, successor to o3-mini, released in April 2025. Offers strong performance on math, coding, and vision. Can process text and images, and features autonomous tool use.

Faster and cost-efficient reasoningVision capabilities ('thinks with images')
Alibaba

Alibaba Qwen3 235B MoE (128k)

128K$0.22/M in / $0.88/M out

Alibaba's flagship Qwen3 Mixture-of-Experts model with 235B total parameters (22B active). Features hybrid reasoning and supports 119 languages. (Note: Not publicly available at release).

235 billion total parameters (22B active)Mixture-of-Experts (MoE) architecture
Alibaba

Alibaba Qwen3 30B MoE (128k)

128K$0.08/M in / $0.29/M out

Alibaba's Qwen3 Mixture-of-Experts model with 30B total parameters (3B active). Features hybrid reasoning and supports 119 languages. Apache 2.0 license.

30 billion total parameters (3B active)Mixture-of-Experts (MoE) architecture
Alibaba

Alibaba Qwen3 32B Dense (128k)

128K$0.1/M in / $0.3/M out

Alibaba's largest dense model in the Qwen3 family with 32B parameters. Features hybrid reasoning and supports 119 languages. Apache 2.0 license.

32 billion parameters (dense model)Hybrid reasoning (thinking/non-thinking modes)
OpenAI

OpenAI GPT-4.1 (400k)

400K$2/M in / $8/M out

Smartest non-reasoning GPT model (text+image in, text out).

Strong instruction followingHigh coding quality
OpenAI

OpenAI GPT-4.1 mini (Unknown context)

Unknown$0.4/M in / $1.6/M out

Smaller, faster GPT-4.1 tier with low cost.

Low costGood speed/quality tradeoff
OpenAI

OpenAI GPT-4.1 nano (Unknown context)

Unknown$0.1/M in / $0.4/M out

Fastest, cheapest GPT-4.1 tier.

Very low costText & image input
Google

Google Gemini 2.5 Flash Preview (1M)

1M$0.15/M in / $0.6/M out

Preview of Google's Gemini 2.5 Flash with hybrid reasoning capabilities. Ultra-fast and cost-efficient for high-volume applications.

Hybrid reasoning with configurable thinking budget1M token context window
Meta

Meta Llama 4 Maverick (1M)

1M$0.2/M in / $0.6/M out

Meta's Llama 4 Maverick with mixture-of-experts architecture, 1M context window, and strong multilingual support. Open-source model.

1M token context windowMixture-of-experts (128 experts)
Meta

Meta Llama 4 Scout (10M)

10M$0.1/M in / $0.3/M out

Meta's Llama 4 Scout with an industry-leading 10M token context window and 16 experts MoE architecture. Optimized for efficiency.

10M token context window16-expert MoE architecture
March 2025 5
Mistral AI

Mistral Small 3.1 (128k)

128K$0.1/M in / $0.3/M out

A new leader in the small models category by Mistral AI, with image understanding capabilities and an extended 128k context length. Apache 2.0 license.

Image understanding capabilities128K token context window
OpenAI

OpenAI GPT-4o with Search (128k)

128K$2.5/M in / $10/M out

GPT-4o variant with built-in web search grounding. Provides up-to-date, cited answers by searching the web in real time.

Built-in web search groundingReal-time information retrieval
Google

Google Gemini 2.5 Pro Preview (1M)

1M$1.25/M in / $10/M out

Early preview of Google's Gemini 2.5 Pro thinking model. Excels at reasoning, coding, and multimodal tasks with a 1M token context window.

Thinking model with advanced reasoning1M token context window
Cohere

Cohere Command A (256k)

256K$2.5/M in / $10/M out

Cohere's next-generation enterprise model with an expanded 256K context window. Optimized for agentic RAG, tool use, and structured outputs.

256K token context windowOptimized for agentic RAG and tool use
Google

Google Gemma 3 27B (128k)

128KOpen weights

Google's open-source 27B parameter model from the Gemma 3 family. Natively multimodal with strong performance on text, image, and video tasks. Free to use under open license.

27 billion parameters128K token context window
February 2025 5
Anthropic

Anthropic Claude 3.7 Sonnet (Deprecated, 200k)

200K$3/M in / $15/M out

Hybrid reasoning Claude 3.x model (extended thinking). Deprecated; recommended replacement is Claude Sonnet 4.5.

Enhanced reasoning over Claude 3.5Improved coding and mathematical capabilities
OpenAI

OpenAI o3-mini (200k)

200K$1.1/M in / $4.4/M out

A faster, more cost-effective version of o3, released in January 2025. Offers strong reasoning, coding, and vision capabilities. Optimized for math and coding tasks.

Cost-effective reasoningText and limited image processing
Mistral AI

Mistral Saba (32k)

32K$0.2/M in / $0.6/M out

A powerful and efficient model from Mistral AI for languages from the Middle East and South Asia.

Optimized for Middle Eastern and South Asian languages32K token context window
xAI

xAI Grok 3 (128k)

128K$3/M in / $15/M out

xAI's flagship large language model with strong reasoning, coding, and math capabilities. Trained on the Colossus supercluster.

Strong reasoning and coding capabilities128K token context window
xAI

xAI Grok 3 Mini (128k)

128K$0.3/M in / $0.5/M out

xAI's lightweight reasoning model with think mode. Faster and more cost-efficient than Grok 3 while maintaining strong reasoning capabilities.

Lightweight reasoning modelThink mode for step-by-step reasoning
January 2025 3
DeepSeek

DeepSeek 3.1 (128k)

128K$0.12/M in / $0.24/M out

DeepSeek's enhanced model with improved reasoning capabilities, expanded context window, and even more competitive pricing.

Enhanced reasoning capabilitiesImproved multilingual performance
Mistral AI

Mistral Codestral 2 (256k)

256K$0.3/M in / $0.9/M out

Mistral AI's cutting-edge language model for coding (second version). Specializes in low-latency, high-frequency tasks like fill-in-the-middle (FIM), code correction, and test generation.

Specialized for code generation256K token context window
DeepSeek

DeepSeek R1 (128k)

128K$0.55/M in / $2.19/M out

DeepSeek's reasoning model trained with reinforcement learning. Excels at math, coding, and complex reasoning tasks with transparent chain-of-thought.

Reinforcement learning-based reasoningTransparent chain-of-thought process
December 2024 9
Google

Google Gemini 2.0 Flash (1M)

1M$0.08/M in / $0.3/M out

Google's latest experimental model with breakthrough multimodal capabilities and enhanced reasoning at an extremely competitive price point.

Next-generation multimodal capabilities1M token context window
Meta

Meta Llama 3.3 70B Instruct

128K$0.6/M in / $0.6/M out

Meta's latest 70B parameter model with improved performance and capabilities, offering state-of-the-art results for its size.

70 billion parameters128K context length
DeepSeek

DeepSeek V3 (64k)

64K$0.14/M in / $0.28/M out

DeepSeek's latest model with strong performance across reasoning, coding, and general tasks at competitive pricing.

Strong reasoning capabilitiesExcellent coding performance
OpenAI

OpenAI o1 (200k)

200K$15/M in / $60/M out

Previous full o-series reasoning model (text+image in, text out).

200K context windowHigh reasoning depth
OpenAI

OpenAI o1-pro (200k)

200K$150/M in / $600/M out

Higher-compute variant of o1 for better responses.

Very high reasoning computePremium pricing
Microsoft

Microsoft Phi-4 (16k)

16KOpen weights

Microsoft's small language model with 14B parameters that punches well above its weight. Excels at STEM reasoning and coding despite its compact size. Open source under MIT license.

14 billion parameters16K token context window
Amazon

Amazon Nova Pro (300k)

300K$0.8/M in / $3.2/M out

Amazon's highly capable multimodal model balancing accuracy, speed, and cost. Processes text, images, and video inputs for a wide range of enterprise tasks.

300K token context windowMultimodal (text, image, video input)
Amazon

Amazon Nova Lite (300k)

300K$0.06/M in / $0.24/M out

Amazon's very low-cost multimodal model for high-volume tasks. Processes text, images, and video at extremely competitive pricing via Amazon Bedrock.

300K token context windowMultimodal (text, image, video input)
OpenAI

OpenAI Sora Turbo

N/A

Advanced text-to-video generation with realistic motion and scene understanding. Pricing per video generation varies by length and resolution.

Realistic motion physicsMulti-shot sequences
October 2024 2
Anthropic

Anthropic Claude 3.5 Sonnet (200k)

200K$3/M in / $15/M out

Anthropic's most advanced model, significantly improved over Claude 3 Sonnet with enhanced reasoning, coding, and vision capabilities.

Significantly improved reasoning and codingEnhanced vision capabilities
Anthropic

Anthropic Claude 3.5 Haiku (Deprecated, 200k)

200K$0.8/M in / $4/M out

Claude 3.5 Haiku snapshot. Deprecated; recommended replacement is Haiku 4.5.

Fast and cost-efficient200K context window
September 2024 6
OpenAI

OpenAI o1-preview (Deprecated, 128k)

128K$15/M in / $60/M out

Deprecated preview snapshot of OpenAI's first o-series reasoning model. Kept for backwards compatibility; prefer o1 for production.

Reasoning-first o-series model (preview snapshot)128K token context window (historical preview)
OpenAI

OpenAI o1-mini (128k)

128K$1.1/M in / $4.4/M out

Smaller, faster o-series reasoning model. Deprecated in favor of newer reasoning models, but still supported in legacy workflows.

Strong reasoning at lower costOptimized for STEM and coding
Meta

Meta Llama 3.2 90B Vision Instruct

128K$1.2/M in / $1.2/M out

Meta's multimodal model combining text and vision capabilities with strong performance across various tasks.

90 billion parametersVision and text capabilities
Meta

Meta Llama 3.2 11B Vision Instruct

128K$0.18/M in / $0.18/M out

A smaller, efficient multimodal model from Meta with vision capabilities, suitable for edge deployment and cost-sensitive applications.

11 billion parametersVision and text capabilities
Mistral AI

Mistral Small (32k)

32K$0.2/M in / $0.6/M out

Mistral AI's cost-effective model for straightforward tasks, offering good performance and efficiency.

Optimized for latency and costGood for simple tasks
Alibaba

Qwen 2.5 72B Instruct

32K$0.56/M in / $0.56/M out

Alibaba's large language model with strong performance in reasoning, coding, and multilingual tasks.

72 billion parametersStrong multilingual capabilities
August 2024 1
AI21 Labs

AI21 Jamba 1.5 Large (256k)

256K$2/M in / $8/M out

AI21's large hybrid SSM-Transformer model with a 256K context window. Uses a novel Jamba architecture combining Mamba SSM layers with Transformer attention for efficient long-context processing.

Hybrid SSM-Transformer (Jamba) architecture256K token context window
July 2024 5
OpenAI

OpenAI GPT-4o mini (128k)

128K$0.15/M in / $0.6/M out

OpenAI's most affordable and fastest model in the GPT-4o family, designed for high-volume, low-latency tasks.

Cost-effectiveHigh speed
Meta

Meta Llama 3.1 405B Instruct

128K$2.7/M in / $2.7/M out

Meta's largest and most capable Llama 3.1 model, designed for complex reasoning, coding, and nuanced instruction following.

405 billion parameters128K context length
Meta

Meta Llama 3.1 70B Instruct

128K$0.6/M in / $0.6/M out

A large instruction-tuned model from Meta's Llama 3.1 series, offering a strong balance of performance and efficiency for a wide range of tasks.

70 billion parameters128K context length
Meta

Meta Llama 3.1 8B Instruct

128K$0.06/M in / $0.06/M out

A highly efficient instruction-tuned model from Meta's Llama 3.1 series, suitable for fast, on-device, or edge applications.

8 billion parameters128K context length
Mistral AI

Mistral Large 2 (128k)

128K$2/M in / $6/M out

Mistral AI's flagship model with enhanced reasoning, coding, and multilingual capabilities.

Top-tier reasoning and codingImproved multilingual performance
June 2024 1
Runway

Runway Gen-3 Alpha

N/A

Professional video generation and editing AI with motion control and style consistency. Subscription-based pricing.

Text and image to videoMotion brush control
May 2024 3
OpenAI

OpenAI GPT-4o (128k)

128K$2.5/M in / $10/M out

OpenAI's flagship multimodal model, natively processing text, audio, and images for faster, more capable interactions.

Native multimodal (text, audio, image)128K token context window
Google

Google Gemini 1.5 Flash (1M)

1M$0.08/M in / $0.3/M out

Google's faster and lower-cost version of Gemini 1.5 Pro, optimized for high-volume, high-frequency tasks while retaining a large context window and multimodal capabilities.

1M token context windowMultimodal capabilities
Mistral AI

Mistral Codestral (22B)

32K$0.2/M in / $0.6/M out

Mistral AI's open-weight generative model specialized for code generation, supporting 80+ languages.

Specialized for code (80+ languages)22 billion parameters
April 2024 2
OpenAI

OpenAI GPT-4 Turbo (128k)

128K$10/M in / $30/M out

OpenAI's powerful model prior to GPT-4o, with a large context window and strong performance on complex tasks. Supports vision.

128K token context windowVision capabilities (image inputs)
Cohere

Cohere Command R+ (128k)

128K$2.5/M in / $10/M out

Cohere's most powerful model optimized for enterprise RAG and tool use. Excels at grounded generation with citations and multi-step tool workflows.

Optimized for enterprise RAG with citations128K token context window
March 2024 3
Anthropic

Anthropic Claude 3 Opus (200k)

200K$15/M in / $75/M out

Claude 3 Opus (deprecated) - previous highest-intelligence Claude 3 model.

Top-level performance on complex tasksStrong math & coding skills
Anthropic

Anthropic Claude 3 Sonnet (200k)

200K$3/M in / $15/M out

A balanced model from Anthropic, offering a blend of intelligence and speed, ideal for enterprise workloads and scaled AI deployments.

Ideal balance of intelligence and speed2x faster than Claude 2 and 2.1 for most workloads
Anthropic

Anthropic Claude 3 Haiku (200k)

200K$0.25/M in / $1.25/M out

Anthropic's fastest and most compact model, designed for near-instant responsiveness and high throughput tasks.

Fastest model in its intelligence categoryCost-effective
February 2024 1
Google

Google Gemini 1.5 Pro (2M)

2M$1.25/M in / $5/M out

Google's highly capable multimodal model with a breakthrough long context window of up to 2 million tokens. Excels at complex reasoning, problem-solving, and understanding long-form content.

2M token context windowNative multimodal (text, code, image, audio, video)
November 2023 1
Leonardo.AI

Leonardo AI

N/A

Game asset and creative image generation with consistent character and style creation. Credit-based pricing system.

Game asset generationConsistent characters
October 2023 1
OpenAI

OpenAI DALL-E 3

N/A

State-of-the-art text-to-image generation with improved prompt following and image quality. Pricing per image: HD 1024×1024 $0.040, Standard $0.020

Superior prompt adherenceHigh-quality image generation
July 2023 1
Stability AI

Stable Diffusion XL

N/A

Open-source image generation model with fine-tuning capabilities. API pricing varies by provider, self-hosting available.

Open-sourceFine-tunable
March 2023 1
OpenAI

OpenAI GPT-3.5 Turbo (16k)

16K$0.5/M in / $1.5/M out

OpenAI's fast, cost-effective model optimized for chat and simple tasks.

Fast and cost-effective16K token context window
142
Total Models
33
Release Periods
20
Providers
Jul 2026
Latest Update