Skip to content
  1. AI index
  2. Models

AI models, ranked.

45 model configurations on capability and on what they cost to get a right answer, plus 378 models on head-to-head human preference. Every number is read from a published board with its source attached.

Most capable82.4GPT 5.6 Sol Max
Best value$0.016DeepSeek V4 Flash, per solved task
Within 5 points of the top$0.36GPT 5.6 Sol xHigh
Humans prefer1507claude-fable-5, Arena rating

Intelligence vs cost

LiveBench overall score against what each model spent to solve a task it got right, across the full 2026-06-25 release. Up and to the left is better. The line joins every model that nothing cheaper beats; the shaded corner holds the models above the median score at below the median cost.

  • OpenAI
  • Anthropic
  • Google
  • Other labs

DeepSeek V4 Flash scores 65.5 for $0.016 a task. GPT 5.6 Sol Max scores 82.4 for $0.59: 37× the cost for 16.9 more points.

Cost frontier: no cheaper model scores higher

Source: LiveBench official release 2026-06-25, cost file frozen to release prices. 45 model configurations.Board and method

Most capable

LiveBench overall score, 0 to 100, across reasoning, math, coding, data analysis, language and instruction following.

  • OpenAI
  • Anthropic
  • Google
  • Other labs
  1. 1GPT 5.6 Sol MaxOpenAI82.4
  2. 2Claude Fable 5 Max EffortAnthropic80.8
  3. 3Claude Opus 5 xHigh EffortAnthropic80.3
  4. 4GPT 5.5 xHighOpenAI79.9
  5. 5GPT 5.6 Terra MaxOpenAI79.8
  6. 6GPT 5.6 Sol xHighOpenAI79.7
  7. 7Claude Fable 5 xHigh EffortAnthropic79.5
  8. 8Claude Opus 5 Max EffortAnthropic79.2
  9. 9Claude Opus 5 High EffortAnthropic79.0
  10. 10Claude Opus 4.8 xHigh EffortAnthropic78.9
  11. 11GPT 5.5 HighOpenAI78.7
  12. 12Kimi K3Moonshot78.5
  13. 13GPT 5.4 xHighOpenAI78.0
  14. 14Gemini 3.1 Pro Preview HighGoogle77.1
  15. 15Claude Opus 4.7 xHigh EffortAnthropic76.5
Hover a bar for the full record

What a million tokens costs

List price per million input and output tokens for the 16 most capable configurations. Output is where the bill is: across all 45, it runs 2 to 8 times the input price.

  • Input
  • Output
USD per million tokens, log scale. Release pricing.

Where each model is strong

LiveBench category scores for the 14 most capable configurations. Darker is higher.

ModelReasoningMathCodingAgentic codingData analysisLanguageInstruction following
GPT 5.6 Sol Max92968466808872
Claude Fable 5 Max Effort90968647819176
Claude Opus 5 xHigh Effort90958361788768
GPT 5.5 xHigh90968252828771
GPT 5.6 Terra Max91957868798365
GPT 5.6 Sol xHigh90958257808667
Claude Fable 5 xHigh Effort88968351798972
Claude Opus 5 Max Effort91968159758964
Claude Opus 5 High Effort87958162798664
Claude Opus 4.8 xHigh Effort90957956788172
GPT 5.5 High90958047808871
Kimi K391848158798671
GPT 5.4 xHigh88947854798370
Gemini 3.1 Pro Preview High84917645798579
4596

Best model from each lab

Arena text rating, style-controlled, from head-to-head human votes. The whisker is the 95% interval.

  1. 1Anthropicclaude-fable-51507
  2. 2Metamuse-spark-1.11495
  3. 3Googlegemini-3.1-pro-preview1486
  4. 4Moonshotkimi-k31486
  5. 5OpenAIgpt-5.6-sol-xhigh1485
  6. 6Alibabaqwen3.7-max-preview1475
  7. 7xAIgrok-4.20-beta11474
  8. 8Z.aiglm-5.11470
  9. 9baiduernie-5.11468
  10. 10xiaomimimo-v2.5-pro1467
  11. 11DeepSeekdeepseek-v4-pro1457
  12. 12bytedancedola-seed-2.0-pro1455
Bars start at the lowest interval shown, not at zero.

Open weights vs proprietary

Every Arena model rated above 1300, by license. The best open model sits 38 points behind the best closed one.

Arena text leaderboard, 378 models.Board and method

These models recommend brands.

When a buyer asks ChatGPT, Claude or Gemini who to buy from, the answer names a few brands and skips the rest. We rank which ones it names. Want us to rank yours?

Get my brand ranked →

LiveBench release 2026-06-25. Arena snapshot Jul 21, 2026. LiveBench publishes a full release about twice a year; the page updates when a new one lands.