LiveBench 2026-06-25 model capability
Ranks exact model configurations by the mean of seven official LiveBench category averages.
- Version
- livebench-2026-06-25-official-v1-livebench-model-capability-v1
- Published
- Jul 27, 2026
- Live boards
- 1
- Method ID
- livebench-model-capability
Where the facts come from
- Source
livebench-official
How the order is produced
unweighted mean of the seven category means
When the method reruns
- daily
Boards using this version
Only live, quality-gated boards appear. Each board opens on the exact published table and evidence path.
| Board | Primary measure | Rows | Confidence | Freshness | Board version | Cadence | Evidence |
|---|---|---|---|---|---|---|---|
| LiveBench Model Capabilitylivebench-model-capability | livebench-overall-score | 45 | 100% | current | livebench-2026-06-25-official-v1-livebench-model-capability-v1 | daily | Open board → |
Complete measurement ledger
These fields are stored with the methodology version. A changed contract must publish a new version before comparable movement can resume.
Source
livebench-official
License
Apache-2.0
Release
2026-06-25
Universe
All 45 exact model-and-inference configurations in the pinned official release.
Versioning
Release identifier, asset SHA-256 hashes, category manifest, formula, and exact inference configuration are frozen.
Limitations
- This is one benchmark family, not Rank.ai's absolute composite.
- Scores compare only configurations within this frozen release.
- Release pricing does not represent negotiated or cached pricing.
- Cost per successful task is specific to the LiveBench workload.
Missing Data
The entire release is suppressed unless every model has every required metric. Missing values are never imputed.
Rank Formula
unweighted mean of the seven category means
Tie Breakers
- canonical entity slug ascending
Primary Metric
livebench-overall-score
Cost Construction
- Basis
prices and token counts in the official release cost file
- Cost Per Question
sum task cost / sum evaluated questions
- Cost Per Successful Task
cost per question / (overall score / 100)
Ranking Direction
descending
Category Construction
- Categories
- Reasoning
- Coding
- Agentic Coding
- Mathematics
- Data Analysis
- Language
- IF
- Overall Score
unweighted mean of category scores
- Category Score
unweighted mean of task scores
Required Metrics Per Entity
- livebench-overall-score
- livebench-reasoning-score
- livebench-coding-score
- livebench-agentic-coding-score
- livebench-mathematics-score
- livebench-data-analysis-score
- livebench-language-score
- livebench-instruction-following-score