Live chart · updates with every published snapshot

Best AI model score on LiveBench

The highest overall score on LiveBench, an objective benchmark with contamination-resistant questions, held by the current leading model configuration.

LiveBench overall score82.4held by GPT 5.6 Sol Max
Behind the number

The full field

Open the board →

  • GPT 5.6 Sol Max82.4
  • Claude Fable 5 Max Effort80.8
  • Claude Opus 5 xHigh Effort80.3
  • GPT 5.5 xHigh79.9
  • GPT 5.6 Terra Max79.8
  • GPT 5.6 Sol xHigh79.7
  • Claude Fable 5 xHigh Effort79.5
  • Claude Opus 5 Max Effort79.2
  • Claude Opus 5 High Effort79.0
  • Claude Opus 4 8 xHigh Effort78.9
  • GPT 5.5 High78.7
  • Kimi K378.5

The board in rank order. The headline number is the top-ranked row, highlighted. Showing the top 12 of 45 ranked entries.

LiveBench overall score by entry
RankEntryLiveBench overall score
1GPT 5.6 Sol Max82.4
2Claude Fable 5 Max Effort80.8
3Claude Opus 5 xHigh Effort80.3
4GPT 5.5 xHigh79.9
5GPT 5.6 Terra Max79.8
6GPT 5.6 Sol xHigh79.7
7Claude Fable 5 xHigh Effort79.5
8Claude Opus 5 Max Effort79.2
9Claude Opus 5 High Effort79.0
10Claude Opus 4 8 xHigh Effort78.9
11GPT 5.5 High78.7
12Kimi K378.5

82.4 on every publication so far. A line appears here as soon as the number moves.

About this metric

LiveBench is an objective benchmark that scores exact model configurations across seven categories, including reasoning, coding, mathematics, and data analysis. Questions rotate to resist contamination, and grading is automatic rather than judged by other models. This chart tracks the single best overall score on the board, whichever configuration holds it.

The score is the unweighted mean of the seven LiveBench category averages for each official release, republished here per exact configuration rather than per marketing name. A model listed at high effort or maximum thinking budget is scored as that configuration only. The leader shown beside the number is read live from the current snapshot.

The frontier score is the cleanest single number for how capable publicly available models actually are. Vendor announcements lead with selective evaluations; a rotating, objective benchmark does not care about launch narratives. Watching the ceiling move, and who holds it, is the fastest way to track real capability shifts.

Related charts

All charts →