RANKING UNIVERSE / LIVE TABLES ONLY

Models

Capability, human preference, tool use, efficiency, price, latency, and model adoption—kept separate so one score never hides the tradeoffs.

Categories
2
Live tables
10
Tracked rows
589
Observations
3,915
PUBLICATION CATALOG

Live models tables

Inspect data health →
CategoryTablePrimary measureEntitiesConfidenceFreshnessCadence
MModel rankAbsolute model rankModels ordered by published intelligence score, without a cost adjustment.Published model measure100Current6-hour refresh
MModel rankCoding model rankA task-specific view of published coding scores.Published model measure100Current6-hour refresh
MModel rankAgentic model rankModels ordered by published agentic scores.Published model measure100Current6-hour refresh
MModel rankModel adoption signalsHugging Face repository downloads and engagement for a fixed, reviewed model cohort.Hugging Face downloads (30 days)12100%currentdaily
MModel rankLiveBench model capabilityExact model configurations ranked by the official, objective LiveBench release, with every category and source artifact inspectable.LiveBench overall score45100%currentdaily
MModel rankArena text preferenceOfficial style-controlled text human-preference ratings with confidence bounds, variance, vote counts, and a frozen dataset revision.Arena text style-controlled rating378100%currentdaily
MModel rankBFCL tool-use configurationsExact model-plus-inference configurations ranked by UC Berkeley BFCL V4 overall accuracy, with tool-use breakdowns, cost, latency, and source evidence.BFCL overall accuracy109100%currentdaily
EModel efficiencyQuality per dollarIntelligence points per weighted inference dollar at public API prices.Published model measure100Current6-hour refresh
EModel efficiencyContext per dollarUsable context capacity relative to published input-token pricing.Published model measure100Current6-hour refresh
EModel efficiencyLiveBench cost efficiencyOfficial cost per successful LiveBench task beside capability, token use, and the exact prices frozen into the release.LiveBench cost per successful task45100%currentdaily
01 / M

Model rank

Capability-first rankings across frontier and open models.

7 live tables
02 / E

Model efficiency

Capability relative to token price, context, latency, and throughput.

3 live tables