LLM Benchmark Board
Intelligence, coding and agentic scores side by side, with value per 1M tokens.
Data as of 2026-09-30103 models scoredArtificial Analysis
| Rank | ||||
|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic · Adaptive Reasoning, Max Effort, Default Fallback | 57.6 | 6.5 | $4.40 / $22.00 |
| 2 | Claude Sonnet 5.5 Anthropic · Adaptive Reasoning, Max Effort, Default Fallback | 56.0 | 14 | $2.00 / $10.00 |
| 3 | Claude Fable 5.1 Anthropic · Adaptive Reasoning, Max Effort, Default Fallback | 53.4 | 5.3 | $5.00 / $25.00 |
| 4 | Qwen3.8 Max qwen | 53.4 | — | — |
| 5 | GPT-6 Astra OpenAI · max | 52.7 | 0.4 | $60.00 / $300 |
| 6 | GPT-6.1 Sol OpenAI · max | 51.8 | 12 | $2.20 / $11.00 |
| 7 | Claude Opus 5 Anthropic · Adaptive Reasoning, Max Effort | 50.8 | 4.6 | $5.50 / $27.50 |
| 8 | Claude Fable 5 Anthropic · Adaptive Reasoning, Max Effort, Opus 4.8 Fallback | 49.6 | 2.5 | $10.00 / $50.00 |
| 9 | GPT-6 Sol OpenAI · max | 47.5 | 12 | $2.00 / $10.00 |
| 10 | GPT-5.6 Sol OpenAI · max | 47.0 | 5.9 | $4.00 / $20.00 |
| 11 | Grok 4.7 SpaceXAI · xhigh | 46.4 | 14 | $2.20 / $6.60 |
| 12 | MiMo-V2.6-Pro Xiaomi | 46.3 | 85 | $0.435 / $0.870 |
| 13 | Qwen3.8 Max Qwen · 0902 | 45.4 | 15 | $2.00 / $6.00 |
| 14 | GLM-5.3 Z.ai · max | 44.8 | 19 | $1.54 / $4.84 |
| 15 | Grok 4.6 SpaceXAI · high | 44.3 | 13 | $2.20 / $6.60 |
| 16 | Kimi K3 MoonshotAI · max | 43.6 | 7.8 | $2.80 / $14.00 |
| 17 | GPT-5.6 Terra OpenAI · max | 42.1 | 4.7 | $4.00 / $24.00 |
| 18 | Claude Opus 4.8 Anthropic · Adaptive Reasoning, Max Effort | 41.8 | 2.1 | $10.00 / $50.00 |
| 19 | GLM 5.3 Flash Z.ai | 41.8 | 176 | $0.150 / $0.500 |
| 20 | Gemini 3.8 Flash Google · high | 40.9 | 15 | $1.35 / $6.75 |
| 21 | Qwen3.8 2.4T A95B Qwen | 39.9 | 13 | $2.00 / $6.00 |
| 22 | Muse Spark 1.2 Meta · xhigh | 39.6 | 20 | $1.25 / $4.25 |
| 23 | DeepSeek V4.1 Flash DeepSeek · Reasoning, Max Effort | 39.5 | 50 | $0.450 / $1.80 |
| 24 | Gemini 3.7 Flash Google · high | 39.1 | 14 | $1.35 / $6.75 |
| 25 | Grok 4.5 SpaceXAI · high | 38.8 | 6.5 | $4.00 / $12.00 |
| 26 | GPT-5.5 OpenAI · xhigh | 38.4 | 6.8 | $2.50 / $15.00 |
| 27 | Claude Sonnet 5 Anthropic · Adaptive Reasoning, Max Effort | 38.2 | 8.7 | $2.20 / $11.00 |
| 28 | MiMo-V2.6-Flash Xiaomi | 37.9 | 217 | $0.140 / $0.280 |
| 29 | GPT-6 Luna OpenAI · max | 37.3 | 186 | $0.100 / $0.500 |
| 30 | GPT-5.6 Luna OpenAI · max | 37.3 | 41 | $0.400 / $2.40 |
| 31 | DeepSeek V4 Pro 0813 DeepSeek · Reasoning, Max Effort | 36.0 | 36 | $0.150 / $3.50 |
| 32 | DeepSeek V4 Flash 0731 DeepSeek · Reasoning, Max Effort | 34.3 | 196 | $0.140 / $0.280 |
| 33 | Gemini 3.6 Flash Google · high | 34.0 | 13 | $1.35 / $6.75 |
| 34 | GLM-5.2 Z.ai · max | 33.7 | 30 | $0.180 / $4.00 |
| 35 | Muse Spark 1.1 Meta · xhigh | 33.7 | 17 | $1.25 / $4.25 |
| 36 | Qwen3.8 27B Qwen · xhigh | 33.7 | 30 | $0.025 / $4.35 |
| 37 | Gemini 3.5 Flash Google · high | 32.6 | 5.4 | $2.70 / $16.20 |
| 38 | DeepSeek V4 Pro 0424 DeepSeek · Reasoning, Max Effort | 30.4 | 26 | $0.400 / $3.50 |
| 39 | Claude Sonnet 4.6 Anthropic · Adaptive Reasoning, Max Effort | 30.1 | 5.0 | $3.00 / $15.00 |
| 40 | Gemini 3.1 Pro Preview | 29.7 | 3.7 | $3.60 / $21.60 |
| 41 | Qwen3.7 Max Qwen | 29.5 | 13 | $1.48 / $4.42 |
| 42 | MiniMax-M3 MiniMax | 29.2 | 28 | $0.600 / $2.40 |
| 43 | Kimi K2.6 MoonshotAI | 27.0 | 18 | $0.855 / $3.60 |
| 44 | GLM-5.1 Z.ai · Reasoning | 26.1 | 13 | $1.33 / $4.18 |
| 45 | MiMo-V2.5-Pro Xiaomi | 26.0 | 40 | $0.522 / $1.04 |
| 46 | Kimi K2.7 Code MoonshotAI | 25.8 | 15 | $0.950 / $4.00 |
| 47 | Hy3 Tencent | 25.3 | 89 | $0.180 / $0.600 |
| 48 | Qwen3.7 Plus Qwen | 25.2 | 45 | $0.320 / $1.28 |
| 49 | Inkling Thinking Machines · xhigh | 25.0 | Free | $0 / $0 |
| 50 | Grok 4.3 SpaceXAI · high | 24.9 | 20 | $1.00 / $2.00 |
Value = score ÷ blended price per 1M tokens. Higher is better.
Intelligence, coding and agentic scores side by side, with value per 1M tokens.
Scores come from Artificial Analysis, served through OpenRouter, across three independent indices: intelligence for general reasoning and knowledge, coding for real programming tasks, and agentic for multi-step tool use and autonomous execution. They disagree often — the most intelligent model is not always the best coder, and strong agentic scores frequently come with a matching price tag. Putting all three next to the pricing column is what turns “should we switch?” into a decision with evidence behind it.
How to use
- 1Switch the benchmark (intelligence / coding / agentic) and the board re-ranks by that score.
- 2Sort by value to see how many points each dollar per 1M tokens buys.
- 3Read across to the price column to find the best score inside your budget.