- AI index
- Models
AI models, ranked.
45 model configurations on capability and on what they cost to get a right answer, plus 378 models on head-to-head human preference. Every number is read from a published board with its source attached.
Intelligence vs cost
LiveBench overall score against what each model spent to solve a task it got right, across the full 2026-06-25 release. Up and to the left is better. The line joins every model that nothing cheaper beats; the shaded corner holds the models above the median score at below the median cost.
- OpenAI
- Anthropic
- Other labs
DeepSeek V4 Flash scores 65.5 for $0.016 a task. GPT 5.6 Sol Max scores 82.4 for $0.59: 37× the cost for 16.9 more points.
Cost frontier: no cheaper model scores higher
Most capable
LiveBench overall score, 0 to 100, across reasoning, math, coding, data analysis, language and instruction following.
- OpenAI
- Anthropic
- Other labs
What a million tokens costs
List price per million input and output tokens for the 16 most capable configurations. Output is where the bill is: across all 45, it runs 2 to 8 times the input price.
- Input
- Output
Where each model is strong
LiveBench category scores for the 14 most capable configurations. Darker is higher.
| Model | Reasoning | Math | Coding | Agentic coding | Data analysis | Language | Instruction following |
|---|---|---|---|---|---|---|---|
| GPT 5.6 Sol Max | 92 | 96 | 84 | 66 | 80 | 88 | 72 |
| Claude Fable 5 Max Effort | 90 | 96 | 86 | 47 | 81 | 91 | 76 |
| Claude Opus 5 xHigh Effort | 90 | 95 | 83 | 61 | 78 | 87 | 68 |
| GPT 5.5 xHigh | 90 | 96 | 82 | 52 | 82 | 87 | 71 |
| GPT 5.6 Terra Max | 91 | 95 | 78 | 68 | 79 | 83 | 65 |
| GPT 5.6 Sol xHigh | 90 | 95 | 82 | 57 | 80 | 86 | 67 |
| Claude Fable 5 xHigh Effort | 88 | 96 | 83 | 51 | 79 | 89 | 72 |
| Claude Opus 5 Max Effort | 91 | 96 | 81 | 59 | 75 | 89 | 64 |
| Claude Opus 5 High Effort | 87 | 95 | 81 | 62 | 79 | 86 | 64 |
| Claude Opus 4.8 xHigh Effort | 90 | 95 | 79 | 56 | 78 | 81 | 72 |
| GPT 5.5 High | 90 | 95 | 80 | 47 | 80 | 88 | 71 |
| Kimi K3 | 91 | 84 | 81 | 58 | 79 | 86 | 71 |
| GPT 5.4 xHigh | 88 | 94 | 78 | 54 | 79 | 83 | 70 |
| Gemini 3.1 Pro Preview High | 84 | 91 | 76 | 45 | 79 | 85 | 79 |
Best model from each lab
Arena text rating, style-controlled, from head-to-head human votes. The whisker is the 95% interval.
Open weights vs proprietary
Every Arena model rated above 1300, by license. The best open model sits 38 points behind the best closed one.
These models recommend brands.
When a buyer asks ChatGPT, Claude or Gemini who to buy from, the answer names a few brands and skips the rest. We rank which ones it names. Want us to rank yours?
LiveBench release 2026-06-25. Arena snapshot Jul 21, 2026. LiveBench publishes a full release about twice a year; the page updates when a new one lands.