Rank.ai
Menu
Live
RANK.AI / PUBLISHED METHOD

LiveBench 2026-06-25 model capability

Ranks exact model configurations by the mean of seven official LiveBench category averages.

Version
livebench-2026-06-25-official-v1-livebench-model-capability-v1
Published
Jul 27, 2026
Live boards
1
Method ID
livebench-model-capability

Where the facts come from

Source

livebench-official

How the order is produced

unweighted mean of the seven category means

When the method reruns

  • daily

Boards using this version

Only live, quality-gated boards appear. Each board opens on the exact published table and evidence path.

BoardPrimary measureRowsConfidenceFreshnessBoard versionCadenceEvidence
LiveBench Model Capabilitylivebench-model-capabilitylivebench-overall-score45100% currentlivebench-2026-06-25-official-v1-livebench-model-capability-v1dailyOpen board →

Complete measurement ledger

These fields are stored with the methodology version. A changed contract must publish a new version before comparable movement can resume.

Source

livebench-official

License

Apache-2.0

Release

2026-06-25

Universe

All 45 exact model-and-inference configurations in the pinned official release.

Versioning

Release identifier, asset SHA-256 hashes, category manifest, formula, and exact inference configuration are frozen.

Limitations

  • This is one benchmark family, not Rank.ai's absolute composite.
  • Scores compare only configurations within this frozen release.
  • Release pricing does not represent negotiated or cached pricing.
  • Cost per successful task is specific to the LiveBench workload.

Missing Data

The entire release is suppressed unless every model has every required metric. Missing values are never imputed.

Rank Formula

unweighted mean of the seven category means

Tie Breakers

  • canonical entity slug ascending

Primary Metric

livebench-overall-score

Cost Construction

Basis

prices and token counts in the official release cost file

Cost Per Question

sum task cost / sum evaluated questions

Cost Per Successful Task

cost per question / (overall score / 100)

Ranking Direction

descending

Category Construction

Categories
  • Reasoning
  • Coding
  • Agentic Coding
  • Mathematics
  • Data Analysis
  • Language
  • IF
Overall Score

unweighted mean of category scores

Category Score

unweighted mean of task scores

Required Metrics Per Entity

  • livebench-overall-score
  • livebench-reasoning-score
  • livebench-coding-score
  • livebench-agentic-coding-score
  • livebench-mathematics-score
  • livebench-data-analysis-score
  • livebench-language-score
  • livebench-instruction-following-score