LiveBench 2026-06-25 model cost efficiency
Ranks exact model configurations by official cost per successful task for the complete frozen LiveBench workload.
- Version
- livebench-2026-06-25-official-v1-livebench-model-efficiency-v1
- Published
- Jul 27, 2026
- Live boards
- 1
- Method ID
- livebench-model-efficiency
Where the facts come from
- Source
livebench-official
How the order is produced
(sum task inference cost / sum evaluated questions) / (overall score / 100)
When the method reruns
- daily
Boards using this version
Only live, quality-gated boards appear. Each board opens on the exact published table and evidence path.
| Board | Primary measure | Rows | Confidence | Freshness | Board version | Cadence | Evidence |
|---|---|---|---|---|---|---|---|
| LiveBench Model Cost Efficiencylivebench-model-efficiency | livebench-cost-per-successful-task-usd | 45 | 100% | current | livebench-2026-06-25-official-v1-livebench-model-efficiency-v1 | daily | Open board → |
Complete measurement ledger
These fields are stored with the methodology version. A changed contract must publish a new version before comparable movement can resume.
Source
livebench-official
License
Apache-2.0
Release
2026-06-25
Universe
All 45 exact model-and-inference configurations in the pinned official release.
Versioning
Release identifier, asset SHA-256 hashes, category manifest, formula, and exact inference configuration are frozen.
Limitations
- This is one benchmark family, not Rank.ai's absolute composite.
- Scores compare only configurations within this frozen release.
- Release pricing does not represent negotiated or cached pricing.
- Cost per successful task is specific to the LiveBench workload.
Missing Data
The entire release is suppressed unless every model has every required metric. Missing values are never imputed.
Rank Formula
(sum task inference cost / sum evaluated questions) / (overall score / 100)
Tie Breakers
- livebench-overall-score descending
- canonical entity slug ascending
Primary Metric
livebench-cost-per-successful-task-usd
Cost Construction
- Basis
prices and token counts in the official release cost file
- Cost Per Question
sum task cost / sum evaluated questions
- Cost Per Successful Task
cost per question / (overall score / 100)
Ranking Direction
ascending
Category Construction
- Categories
- Reasoning
- Coding
- Agentic Coding
- Mathematics
- Data Analysis
- Language
- IF
- Overall Score
unweighted mean of category scores
- Category Score
unweighted mean of task scores
Required Metrics Per Entity
- livebench-cost-per-successful-task-usd
- livebench-overall-score
- livebench-cost-per-question-usd
- livebench-input-price-per-million-usd
- livebench-output-price-per-million-usd
- livebench-average-input-tokens
- livebench-average-output-tokens