Leaderboard

LLM Leaderboard 2026

60 models compared on price per million tokens, context window, public ELO where available, and what each one is actually best at. Includes the July 2026 wave — GPT-5.6 Sol/Terra/Luna, Claude Opus 5, Claude Sonnet 5, Gemini 3.6 Flash and open-weight Kimi K3 — which are listed without ELO until public arena scores stabilize.

Search models

60 models — ranked by ELO, newest first

Ranked list shows models with public ELO data. New models without enough Arena votes are listed separately below.

#ModelDeveloperELOMMLUContextPrice (Input)Model IDOfficialType
1Claude Opus 4.8Anthropic1512-1M$5.00claude-opus-4.8Claude models docs ->Closed
2GPT-5.5 ProOpenAI1510-256K$5.00gpt-5-5-proOpenAI models docs ->Closed
3GPT-5.5OpenAI1506-1M (API) / 400K (Codex)$5.00gpt-5.5OpenAI models docs ->Closed
4Claude Opus 4.7Anthropic1505-1M$5.00claude-opus-4.7Claude models docs ->Closed
5Gemini 3.1 ProGoogle1505-1M$2.00gemini-3.1-proGemini models docs ->Closed
6Grok 4.3xAI1498-N/Agrok-4.3xAI API docs ->Closed
7Grok 4.20xAI1496-256K$3.00grok-4-20xAI API docs ->Closed
8GPT-5.4OpenAI1495-1M$2.50gpt-5.4OpenAI models docs ->Closed
9Claude Opus 4.6Anthropic149091.11M$5.00claude-opus-4.6Claude models docs ->Closed
10Gemini 3 ProGoogle148691.81Mgemini-3-proGemini models docs ->Closed
11Claude Opus 4.5Anthropic146790.8200K$5.00claude-opus-4.5Claude models docs ->Closed
12Claude Sonnet 4.6Anthropic146789.3200K$3.00claude-sonnet-4.6Claude models docs ->Closed
13DeepSeek V4 ProDeepSeek1467-128Kdeepseek-v4-proDeepSeek API docs ->Closed
14GLM 5.1Zhipu AI1467-128K$1.00glm-5-1Find official docs ->Closed
15Kimi K2.6Moonshot AI1466-256K$1.50kimi-k2-6Moonshot API docs ->Closed
16DeepSeek V3.2DeepSeek1455-128K$0.27deepseek-v3-2DeepSeek API docs ->Open
17Kimi K2.5Moonshot AI1452-262Kkimi-k2.5Moonshot API docs ->Closed
18Gemini 2.5 ProGoogle1450-1M$1.25gemini-2.5-proGemini models docs ->Closed
19GLM 5Zhipu AI1450-128K$0.80glm-5Find official docs ->Closed
20GPT-4.5 PreviewOpenAI1444-128K$75.00gpt-4.5-previewOpenAI models docs ->Closed
21GPT-4oOpenAI144288.7128K$2.50gpt-4oOpenAI models docs ->Closed
22GPT-5.2OpenAI143789.6400K$1.75gpt-5.2OpenAI models docs ->Closed
23o3OpenAI1433-200K$10.00o3OpenAI models docs ->Closed
24Gemini 3.1 Flash-LiteGoogle1421-1M$0.10gemini-3-1-flash-liteGemini models docs ->Closed
25DeepSeek R1DeepSeek141890.864K$0.55deepseek-r1DeepSeek API docs ->Open
26Claude Opus 4Anthropic1414-200K$5.00claude-opus-4Claude models docs ->Closed
27Mistral Large 3Mistral AI1414-128K$2.00mistral-large-3Mistral models docs ->Open
28Grok-3xAI141192.7131K$3.00grok-3xAI API docs ->Closed
29Gemini 2.5 FlashGoogle1410-1M$0.30gemini-2.5-flashGemini models docs ->Closed
30DeepSeek V4 FlashDeepSeek1410-128Kdeepseek-v4-flashDeepSeek API docs ->Closed
31Claude Haiku 4.5Anthropic1404-200K$1.00claude-haiku-4.5Claude models docs ->Closed
32o1OpenAI140290.8200K$15.00o1OpenAI models docs ->Closed

New Models (Awaiting Public ELO)

ModelDeveloperReleasedContextPrice (Input)Official
GPT-5.6 SolOpenAI2026-071.05M$5.00OpenAI models docs ->
GPT-5.6 TerraOpenAI2026-071.05M$2.00OpenAI models docs ->
GPT-5.6 LunaOpenAI2026-071.05M$0.20OpenAI models docs ->
Gemini 3.6 FlashGoogle2026-071M$1.50Gemini models docs ->
Kimi K3Moonshot AI2026-071M$3.00Moonshot API docs ->
Claude Opus 5Anthropic2026-071M$5.00Claude models docs ->
Claude Sonnet 5Anthropic2026-061M$3.00Claude models docs ->
Claude Fable 5Anthropic2026-061M$10.00Claude models docs ->
GPT-5.5 InstantOpenAI2026-05128K$2.50OpenAI models docs ->
GPT-5.4 MiniOpenAI2026-03400K$0.75OpenAI models docs ->
GPT-5.4 NanoOpenAI2026-03400K$0.20OpenAI models docs ->
Claude Sonnet 4Anthropic2025-05200K$3.00Claude models docs ->
Llama 4 ScoutMeta2025-04N/A-Llama model hub ->
Llama 4 MaverickMeta2025-04N/A-Llama model hub ->
Gemini 2.0 FlashGoogle2025-021M$0.10Gemini models docs ->
o3-miniOpenAI2025-01200K$1.10OpenAI models docs ->
DeepSeek V3DeepSeek2024-12128K$0.14DeepSeek API docs ->
Claude 3.5 Sonnet (Oct 2024)Anthropic2024-10200K$3.00Claude models docs ->
Qwen 2.5 72B InstructAlibaba2024-09128K$0.30Qwen docs ->
Grok-2xAI2024-08128K$2.00xAI API docs ->
GPT-4o MiniOpenAI2024-07128K$0.15OpenAI models docs ->
Llama 3.1 405BMeta2024-07128K$0.80Llama model hub ->
Llama 3.1 70BMeta2024-07128K$0.35Llama model hub ->
Llama 3.1 8BMeta2024-07128K$0.05Llama model hub ->
Mistral Large 2Mistral AI2024-07128K$2.00Mistral models docs ->
Gemini 1.5 ProGoogle2024-052M$1.25Gemini models docs ->
Claude 3 OpusAnthropic2024-03200K$15.00Claude models docs ->
Claude 3 HaikuAnthropic2024-03200K$0.25Claude models docs ->

Model Profiles

#1Proprietary

Claude Opus 4.8

Anthropic

Anthropic's May 2026 Opus upgrade. 4x less likely to overlook code flaws than Opus 4.7. First model to complete every case on Super-Agent benchmark. Highest score ever on Legal Agent Benchmark at launch. Fast mode available at $10/$50 per million tokens.

1512

ELO

1M

Context

$5.00

per 1M tokens

#2Proprietary

GPT-5.5 Pro

OpenAI

OpenAI's most capable model as of May 2026. Leads the Arena leaderboard. 52.5% fewer hallucinations than GPT-5.4. Enhanced personalization with conversation memory and Gmail integration.

1510

ELO

256K

Context

$5.00

per 1M tokens

#3Proprietary

GPT-5.5

OpenAI

OpenAI's April 2026 flagship model for real-world coding and professional workflows, with stronger agentic performance and a 1M API context window.

1506

ELO

1M (API) / 400K (Codex)

Context

$5.00

per 1M tokens

#4Proprietary

Claude Opus 4.7

Anthropic

Anthropic's most capable generally available model, launched in April 2026 with stronger long-horizon agentic performance and the same $5/$25 MTok pricing as Opus 4.6.

1505

ELO

1M

Context

$5.00

per 1M tokens

#5Proprietary

Gemini 3.1 Pro

Google

Google's newer Pro generation surfaced at Cloud Next 2026 as its most capable model for complex workflows, with 1M context and Gemini 3-class multimodal support.

1505

ELO

1M

Context

$2.00

per 1M tokens

#6Proprietary

Grok 4.3

xAI

xAI's newer flagship family with top-ranked LMArena results for both thinking and non-thinking modes.

1498

ELO

N/A

Context

per 1M tokens

#7Proprietary

Grok 4.20

xAI

xAI's latest model with improved reasoning and coding capabilities.

1496

ELO

256K

Context

$3.00

per 1M tokens

#8Proprietary

GPT-5.4

OpenAI

OpenAI's March 2026 frontier model for professional reasoning and agentic coding, released across ChatGPT, API, and Codex.

1495

ELO

1M

Context

$2.50

per 1M tokens

#9Proprietary

Claude Opus 4.6

Anthropic

Current #1 on Chatbot Arena (1496 ELO), with 99.8% AIME 2025 and 80.8% SWE-bench, leading in coding and hard prompts.

1490

ELO

1M

Context

$5.00

per 1M tokens

#10Proprietary

Gemini 3 Pro

Google

Google's late-2025 Pro model with 94.3% GPQA and 100% AIME 2025, a 1M context window, and strong multimodal reasoning for complex workflows.

1486

ELO

1M

Context

per 1M tokens

#11Proprietary

Claude Opus 4.5

Anthropic

Major upgrade with 87% GPQA and 80.9% SWE-bench, the highest-rated Anthropic model before the 4.6 generation.

1467

ELO

200K

Context

$5.00

per 1M tokens

#12Proprietary

Claude Sonnet 4.6

Anthropic

Latest Sonnet with adaptive reasoning, 89.9% GPQA and 79.6% SWE-bench, excellent balance of speed and intelligence.

1467

ELO

200K

Context

$3.00

per 1M tokens

#13Proprietary

DeepSeek V4 Pro

DeepSeek

DeepSeek's higher-capability V4-generation API model introduced in April 2026 for deeper reasoning and agentic workloads.

1467

ELO

128K

Context

per 1M tokens

#14Proprietary

GLM 5.1

Zhipu AI

Chinese frontier model from Zhipu AI. Competitive with Claude Sonnet at a fraction of the cost.

1467

ELO

128K

Context

$1.00

per 1M tokens

#15Proprietary

Kimi K2.6

Moonshot AI

Moonshot AI's latest model with strong multilingual and long-context capabilities. $20B valuation.

1466

ELO

256K

Context

$1.50

per 1M tokens

#16Open Source

DeepSeek V3.2

DeepSeek

Updated open-source model with near-frontier capability at 1/20th the cost of GPT-5.5.

1455

ELO

128K

Context

$0.27

per 1M tokens

#17Proprietary

Kimi K2.5

Moonshot AI

Chinese model with the highest HumanEval score ever recorded (99.0%), excelling at code generation and reasoning.

1452

ELO

262K

Context

per 1M tokens

#18Proprietary

Gemini 2.5 Pro

Google

Google's hybrid thinking model combining fast responses with deep reasoning, top performer on coding and math benchmarks.

1450

ELO

1M

Context

$1.25

per 1M tokens

#19Proprietary

GLM 5

Zhipu AI

Previous generation Zhipu model, still competitive with Western mid-tier models.

1450

ELO

128K

Context

$0.80

per 1M tokens

#20Proprietary

GPT-4.5 Preview

OpenAI

OpenAI's largest and most knowledgeable non-reasoning model with broad world knowledge and reduced hallucinations.

1444

ELO

128K

Context

$75.00

per 1M tokens

#21Proprietary

GPT-4o

OpenAI

OpenAI's flagship multimodal model with native text, vision, and audio capabilities, offering strong all-around performance.

1442

ELO

128K

Context

$2.50

per 1M tokens

#22Proprietary

GPT-5.2

OpenAI

OpenAI's current-gen flagship with 400K context, 92.4% GPQA and 100% AIME 2025, strong reasoning at reduced cost.

1437

ELO

400K

Context

$1.75

per 1M tokens

#23Proprietary

o3

OpenAI

Advanced reasoning model succeeding o1, with significantly improved math and coding performance at reduced pricing.

1433

ELO

200K

Context

$10.00

per 1M tokens

#24Proprietary

Gemini 3.1 Flash-Lite

Google

Ultra-low latency model designed for sub-second responses at scale. Google's cheapest frontier-adjacent model.

1421

ELO

1M

Context

$0.10

per 1M tokens

#25Open Source

DeepSeek R1

DeepSeek

Open-source reasoning model matching o1 performance with 97.3% MATH-500, disrupted the AI industry with its efficiency.

1418

ELO

64K

Context

$0.55

per 1M tokens

#26Proprietary

Claude Opus 4

Anthropic

Anthropic's first Opus 4 generation model with extended thinking capabilities and strong agentic coding performance.

1414

ELO

200K

Context

$5.00

per 1M tokens

#27Open Source

Mistral Large 3

Mistral AI

Open-weight Apache 2.0 MoE flagship from Mistral 3 generation with strong multilingual and multimodal performance.

1414

ELO

128K

Context

$2.00

per 1M tokens

#28Proprietary

Grok-3

xAI

Trained on xAI's Colossus supercluster, top-tier math reasoning with 93.3% AIME 2025 score.

1411

ELO

131K

Context

$3.00

per 1M tokens

#29Proprietary

Gemini 2.5 Flash

Google

Fast reasoning model with excellent cost efficiency, balancing speed and intelligence for high-throughput applications.

1410

ELO

1M

Context

$0.30

per 1M tokens

#30Proprietary

DeepSeek V4 Flash

DeepSeek

DeepSeek's April 2026 V4-generation API model focused on speed and lower-cost production inference.

1410

ELO

128K

Context

per 1M tokens

#31Proprietary

Claude Haiku 4.5

Anthropic

Anthropic's fastest model in the 4.5 generation, offering near-Sonnet quality at Haiku-tier speed and pricing.

1404

ELO

200K

Context

$1.00

per 1M tokens

#32Proprietary

o1

OpenAI

OpenAI's first reasoning model that uses chain-of-thought to solve complex math, science, and coding problems.

1402

ELO

200K

Context

$15.00

per 1M tokens

#33Proprietary

GPT-5.6 Sol

OpenAI

Flagship tier of the GPT-5.6 family (July 9, 2026). Scores 94.6% on GPQA Diamond, 90.4% on BrowseComp and 89% on FrontierMath Tier 1-3. 922K input / 128K output; requests above 272K input tokens are billed at 2x input and 1.5x output.

-

ELO

1.05M

Context

$5.00

per 1M tokens

#34Proprietary

GPT-5.6 Terra

OpenAI

Balanced tier of the GPT-5.6 family: most of Sol's capability at 40% of the price. OpenAI cut Terra pricing 20% on July 30, 2026. The default choice for production workloads that don't need frontier reasoning.

-

ELO

1.05M

Context

$2.00

per 1M tokens

#35Proprietary

GPT-5.6 Luna

OpenAI

Fastest and cheapest GPT-5.6 tier, cut 80% in price on July 30, 2026. At $0.20 per million input tokens it is priced for high-volume classification, routing and extraction rather than deep reasoning.

-

ELO

1.05M

Context

$0.20

per 1M tokens

#36Proprietary

Gemini 3.6 Flash

Google

Google's speed-and-value play against GPT-5.6 (July 21, 2026). 1M-token context, knowledge cutoff March 2026, and the strongest price-per-task ratio in the current flagship group for high-throughput work.

-

ELO

1M

Context

$1.50

per 1M tokens

#37Open Source

Kimi K3

Moonshot AI

The first open-weight model in the ~3T parameter class (July 16, 2026): a 2.8-trillion-parameter mixture-of-experts with native vision and a 1M-token context. Cached input drops to $0.30 per million, which makes repeated long-context work unusually cheap.

-

ELO

1M

Context

$3.00

per 1M tokens

#38Proprietary

Claude Opus 5

Anthropic

Anthropic's July 2026 Opus release. Stronger than Opus 4.8 on coding, knowledge work, long-horizon agentic tasks, and cost per completed task. Same $5/$25 per million token base pricing, with Fast mode available.

-

ELO

1M

Context

$5.00

per 1M tokens

#39Proprietary

Claude Sonnet 5

Anthropic

Anthropic's most agentic Sonnet yet (June 30, 2026), approaching Opus 4.8 quality at a fraction of the cost. Best for agentic coding, multi-file refactors, long-document analysis and computer use. Introductory pricing of $2/$10 runs through August 31, 2026.

-

ELO

1M

Context

$3.00

per 1M tokens

#40Proprietary

Claude Fable 5

Anthropic

Anthropic's first generally available Mythos-class model (June 9, 2026), restored globally after the June access pause. Built for long-horizon autonomous work, advanced coding, vision, memory, and professional tasks. Priced at $10/$50 per million tokens, with safety routing for high-risk requests.

-

ELO

1M

Context

$10.00

per 1M tokens

#41Proprietary

GPT-5.5 Instant

OpenAI

Fast default model powering ChatGPT for all users. Optimized for speed while maintaining GPT-5.5 quality.

-

ELO

128K

Context

$2.50

per 1M tokens

#42Proprietary

GPT-5.4 Mini

OpenAI

Faster and lower-cost GPT-5.4 variant for high-volume coding, subagents, and computer-use workloads.

-

ELO

400K

Context

$0.75

per 1M tokens

#43Proprietary

GPT-5.4 Nano

OpenAI

OpenAI's smallest GPT-5.4-class model, optimized for ultra-cheap classification, extraction, and lightweight agent sub-tasks.

-

ELO

400K

Context

$0.20

per 1M tokens

#44Proprietary

Claude Sonnet 4

Anthropic

Balanced mid-tier model in the Claude 4 generation, offering strong coding and reasoning at competitive pricing.

-

ELO

200K

Context

$3.00

per 1M tokens

#45Open Source

Llama 4 Scout

Meta

Open-weight, natively multimodal Llama 4 model designed for efficient deployment and long-context workloads.

-

ELO

N/A

Context

per 1M tokens

#46Open Source

Llama 4 Maverick

Meta

Higher-capability open-weight Llama 4 model in Meta's multimodal MoE generation, available for download via llama.com.

-

ELO

N/A

Context

per 1M tokens

#47Proprietary

Gemini 2.0 Flash

Google

Ultra-fast and affordable multimodal model with native tool use, image/audio generation, and 1M token context.

-

ELO

1M

Context

$0.10

per 1M tokens

#48Proprietary

o3-mini

OpenAI

Cost-efficient reasoning model with adjustable effort levels (low/medium/high), matching o1 at medium on STEM tasks.

-

ELO

200K

Context

$1.10

per 1M tokens

#49Open Source

DeepSeek V3

DeepSeek

Chinese open-source MoE model rivaling GPT-4o at a fraction of the cost, trained for under $6M causing industry shock.

-

ELO

128K

Context

$0.14

per 1M tokens

#50Proprietary

Claude 3.5 Sonnet (Oct 2024)

Anthropic

Updated Sonnet with computer use capability and improved coding (93.7% HumanEval), the most popular coding model of late 2024.

-

ELO

200K

Context

$3.00

per 1M tokens

#51Open Source

Qwen 2.5 72B Instruct

Alibaba

Alibaba's leading open-source model with strong multilingual and coding capabilities, competitive with Llama 3.1 70B.

-

ELO

128K

Context

$0.30

per 1M tokens

#52Proprietary

Grok-2

xAI

xAI's second-gen model with real-time X/Twitter data access, competitive with GPT-4o on standard benchmarks.

-

ELO

128K

Context

$2.00

per 1M tokens

#53Proprietary

GPT-4o Mini

OpenAI

Cost-efficient small model replacing GPT-3.5 Turbo, offering strong performance at a fraction of GPT-4o's cost.

-

ELO

128K

Context

$0.15

per 1M tokens

#54Open Source

Llama 3.1 405B

Meta

Largest open-source model at release, competitive with GPT-4o and Claude 3.5 Sonnet across most benchmarks.

-

ELO

128K

Context

$0.80

per 1M tokens

#55Open Source

Llama 3.1 70B

Meta

Strong mid-size open-source model offering excellent performance-to-cost ratio for self-hosted deployments.

-

ELO

128K

Context

$0.35

per 1M tokens

#56Open Source

Llama 3.1 8B

Meta

Compact open-source model suitable for on-device and edge deployments with surprisingly strong capabilities for its size.

-

ELO

128K

Context

$0.05

per 1M tokens

#57Proprietary

Mistral Large 2

Mistral AI

Mistral's flagship with 123B parameters, multilingual in 80+ languages, and strong code generation (92% HumanEval).

-

ELO

128K

Context

$2.00

per 1M tokens

#58Proprietary

Gemini 1.5 Pro

Google

Google's first million-token context model (up to 2M), excelling at long-document understanding and multimodal tasks.

-

ELO

2M

Context

$1.25

per 1M tokens

#59Proprietary

Claude 3 Opus

Anthropic

Anthropic's original flagship model, excelling at complex analysis and nuanced writing with strong safety alignment.

-

ELO

200K

Context

$15.00

per 1M tokens

#60Proprietary

Claude 3 Haiku

Anthropic

Anthropic's fastest and most affordable model, designed for near-instant responses on simple queries and classification.

-

ELO

200K

Context

$0.25

per 1M tokens