Official style-controlled text human-preference ratings with confidence bounds, variance, vote counts, and a frozen dataset revision.
LiveBench model capability
45 models ranked by LiveBench overall score.
GPT 5.6 Sol Max is out in front at 82.4%.
45 modelschecked Jun 25, 2026, 11:59 PM UTCdaily88 failed refresheshow we measure this →CSV / JSON
GPT 5.6 Sol Max
82.4%LiveBench overall score
- Ahead of second place
- 1.6%
- Held by the top three
- 7%
- Middle of the board
- 74.3%
MiniMax M3
Ranks exact model configurations by the mean of seven official LiveBench category averages.
Some explanation inputs are unavailable Missing: publicFrozenEvidence, comparablePriorRanking.
- Published rank
- #40
- Score
- 67.257
- Sample size
- 8
100% confidence · 67% component coverage · as of Jun 25, 2026, 11:59 PM UTC
Component ledger
| Component | Value | Weight | Contribution | Evidence | Status |
|---|---|---|---|---|---|
| livebench agentic coding score | — | 0% | 0 | 1 | available |
| livebench coding score | — | 0% | 0 | 1 | available |
| livebench data analysis score | — | 0% | 0 | 1 | available |
| livebench instruction following score | — | 0% | 0 | 1 | available |
| livebench language score | — | 0% | 0 | 1 | available |
| livebench mathematics score | — | 0% | 0 | 1 | available |
| livebench overall score | — | 100% | 67.257 | 1 | available |
| livebench reasoning score | — | 0% | 0 | 1 | available |
| metrics | Structured | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| rankingDirection | descending | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| release | 2026-06-25 | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| sourceModelId | minimax-m3 | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
new since prior snapshot
no_prior_published_snapshot
No component moved since the prior snapshot.
Source trail
LiveBench Agentic Coding scorelivebench agentic coding score · observed Jun 25, 2026, 11:59 PM UTC40.656 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench Coding scorelivebench coding score · observed Jun 25, 2026, 11:59 PM UTC68.203 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench Data Analysis scorelivebench data analysis score · observed Jun 25, 2026, 11:59 PM UTC76.166 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench IF scorelivebench instruction following score · observed Jun 25, 2026, 11:59 PM UTC57.508 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench Language scorelivebench language score · observed Jun 25, 2026, 11:59 PM UTC76.836 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench Mathematics scorelivebench mathematics score · observed Jun 25, 2026, 11:59 PM UTC76.948 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench overall scorelivebench overall score · observed Jun 25, 2026, 11:59 PM UTC67.257 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
LiveBench Reasoning scorelivebench reasoning score · observed Jun 25, 2026, 11:59 PM UTC74.481 percent100% confidence
- Source
- LiveBench official model benchmark
- Grade
- A
- Freshness
- fresh
- Locator
{"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
More in Model rank
Capability-first rankings across frontier and open models.
Exact model-plus-inference configurations ranked by UC Berkeley BFCL V4 overall accuracy, with tool-use breakdowns, cost, latency, and source evidence.
Hugging Face repository downloads and engagement for a fixed, reviewed model cohort.