AI models · live, read Oct 1, 2026, 04:57 UTC
LLM leaderboard: every major AI benchmark, blended into one ranking
Artificial Analysis, LiveBench, LMArena and Berkeley BFCL each rank AI models their own way, and they rarely agree. We blend them: on each leaderboard a model scores the share of models it beats, and the average of those is its consensus score out of 100. Claude Fable 5 leads, first on 1 of 3.
Best for
The best AI model for each job
Each answer names the leaderboard it comes from. Pick one to see the model's page.
- Best overallClaude Fable 585consensus of 3 leaderboards
- Best valueDeepSeek V4 Pro 0423$0.32per million tokens, consensus 58
- Best for codingClaude Opus 578.0Artificial Analysis Coding Index
- People preferClaude Fable 51507LMArena rating, blind votes
- Best at tool useClaude Opus 4.5 20251101 Thinking 64K High Effort77.47%Berkeley Function Calling Leaderboard
- Best for agentsClaude Opus 556.5Artificial Analysis Agentic Index
The leaderboard
Every model, ranked on the consensus of the leaderboards
Each source column shows the model's rank there and its score on that source's own scale. Sort by what matters to you, search, or narrow to one lab.
Ranked on the consensus of every leaderboard that measures the model.
| # | Model | Consensus | AA | LiveBench | Arena | BFCL | Price / 1M | Context |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 85 | #249.6 | #479.5 | #11507 | – | $20OpenRouter | 1M |
| 2 | Claude Opus 5Anthropic | 83 | #150.8 | #180.3 | – | – | $10OpenRouter | 1M |
| 3 | GPT-5.6 SolOpenAI | 83 | #347.0 | #379.7 | #81485 | – | $4.00OpenRouter | 1.1M |
| 4 | Kimi K3Moonshot AI | 80 | #443.6 | #678.5 | #71486 | – | $2.86OpenRouter | 1.0M |
| 5 | GPT-5.5OpenAI | 80 | #838.4 | #279.9 | #111482 | – | $11OpenRouter | 1.1M |
| 6 | gemini-3-proGoogle | 79 | – | – | #61486 | #372.51% | – | – |
| 7 | Claude Opus 4.8Anthropic | 79 | #641.8 | #578.9 | #101484 | – | $10OpenRouter | 1M |
| 8 | Claude Opus 4.7 xHigh EffortAnthropic | 74 | – | #976.5 | #31502 | – | $10Vercel AI Gateway | – |
| 9 | GPT 5.4 xHighOpenAI | 73 | – | #778.0 | #121478 | – | $5.63Vercel AI Gateway | – |
| 10 | Gemini 3.1 Pro PreviewGoogle | 72 | #1829.7 | #877.1 | #51486 | – | $4.50OpenRouter | 1.0M |
| 11 | Muse Spark 1.1Meta | 72 | #1433.7 | #1176.2 | #41495 | – | $2.00OpenRouter | 1.0M |
| 12 | Grok 4.5xAI | 72 | #738.8 | #1076.2 | #191468 | – | $3.00OpenRouter | 500K |
| 13 | Claude Sonnet 5Anthropic | 69 | #938.2 | #1274.8 | #201461 | – | $4.00OpenRouter | 1M |
| 14 | Claude Opus 4.6 Thinking High EffortAnthropic | 69 | – | #1574.5 | #21505 | – | $10Vercel AI Gateway | – |
| 15 | Claude Opus 4.5 20251101 Thinking 64K High EffortAnthropic | 68 | – | #2172.6 | #151473 | #177.47% | $10Vercel AI Gateway | – |
| 16 | Gemini 3.5 FlashGoogle | 67 | #1632.6 | #1374.6 | #131476 | – | $3.38OpenRouter | 1.0M |
| 17 | Gemini 3.6 FlashGoogle | 67 | #1334.0 | #1773.6 | #91485 | – | $1.50OpenRouter | 1.0M |
| 18 | GPT-5.6 TerraOpenAI | 66 | #542.1 | #1674.3 | – | – | $4.50OpenRouter | 1.1M |
| 19 | grok-4-1-fast-reasoningxAI | 65 | – | – | #361431 | #569.57% | – | – |
| 20 | Claude Sonnet 4.5Anthropic | 64 | #3020.7 | – | #251456 | #273.24% | $6.00OpenRouter | 1M |
Head to head
Compare any two AI models
Every leaderboard that measures both, side by side. The longer bar wins the row.
Claude Opus 5 leads on 2 of 2 leaderboards that measure both.
- 85Consensus0 to 10083
- #2 · 49.6Artificial Analysisshare of models beaten#1 · 50.8
- #4 · 79.5LiveBenchshare of models beaten#1 · 80.3
- #1 · 1507LMArenashare of models beatenNot measured
- Not measuredBerkeley BFCLshare of models beatenNot measured
- $20Priceblended, per 1M tokens$10
Capability against price
What each point of capability costs
The line joins the cheapest model at each step up in consensus. gpt-oss-20b is the cheapest on it, at $0.036 per million tokens; Claude Fable 5 tops it at $20.
Best value line
Gateways
The same model, priced by OpenRouter and Vercel AI Gateway
Blended list price per million tokens where the two gateways differ. Providers behind each gateway set most of these prices, so gaps come and go.
| Model | OpenRouter | Vercel AI Gateway |
|---|---|---|
| GPT-5.6 Sol | $4.00 | $8.00 |
| Kimi K3 | $2.86 | $6.00 |
| GLM 5.1 | $1.48 | $2.15 |
| GLM 5.2 | $2.15 | $1.24 |
| Qwen3.7 Max | $2.21 | $3.75 |
| Qwen3.7 Plus | $0.56 | $0.70 |
| DeepSeek V4 Pro 0423 | $0.32 | $0.99 |
| Kimi K2.6 | $1.34 | $1.71 |
| DeepSeek V4 Flash 0423 | $0.052 | $0.095 |
| MiniMax M2.7 | $0.37 | $0.52 |
| Kimi K2.7 Code | $1.34 | $1.71 |
| DeepSeek V3.1 Terminus | $0.47 | $0.45 |
Sources
Sources, method and how to cite this leaderboard
Free to use with attribution. Every score belongs to the leaderboard that measured it; the consensus and the value ranking are Rank.ai's.
Sources and method
- Artificial Analysis
Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations. Via Artificial Analysis indices on the OpenRouter Models API, read Oct 1, 2026.
- LiveBench
Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language. Via LiveBench official release, published as a Rank.ai board, read Jun 25, 2026.
- LMArena
Blind head-to-head human votes, style controlled. Via Arena official leaderboard dataset, published as a Rank.ai board, read Jul 21, 2026.
- Berkeley BFCL
Function calling and tool use, single turn to multi turn and web search. Via UC Berkeley BFCL V4 official leaderboard, published as a Rank.ai board, read Apr 13, 2026.
- OpenRouter prices
List prices per input and output token, read Oct 1, 2026.
- Vercel AI Gateway prices
List prices per input and output token, read Oct 1, 2026.
Models are matched across sources by name, with lab prefixes, dates and reasoning settings removed, so “GPT 5.6 xHigh” on LiveBench and “openai/gpt-5.6” on OpenRouter are one model. Each source keeps the model's best run.
On each source, a model scores the share of the other ranked models it beats, 0 to 100. The consensus averages those with one neutral score of 50 added, and needs at least two sources. Best value divides the consensus by the blended price, for models scoring 50 or more.
Download the data
- LLM leaderboard (CSV)
85 models: consensus, rank and score on each source, gateway prices, context window.
Last updated October 1, 2026
Questions about the LLM leaderboard
What is the best AI model right now?
What is the best AI model for coding?
What is the best value AI model?
Which AI model do people prefer?
How is the consensus score calculated?
Why do AI leaderboards disagree?
Where do the prices come from?
How often is the LLM leaderboard updated?
More on AI models: LiveBench capability against cost per solved task, what an H100 costs an hour and the whole Index.
These models recommend brands. See if they name yours.
Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.
- Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
- Who gets namedEvery competitor in the answers, most named first.
- The pages AI readsThe sources behind each answer.
- Three fixesWhat to fix first, with a brief for the first article.