Skip to content

AI models · live, read Oct 1, 2026, 04:57 UTC

LLM leaderboard: every major AI benchmark, blended into one ranking

Artificial Analysis, LiveBench, LMArena and Berkeley BFCL each rank AI models their own way, and they rarely agree. We blend them: on each leaderboard a model scores the share of models it beats, and the average of those is its consensus score out of 100. Claude Fable 5 leads, first on 1 of 3.

One thread per model, left to right through each leaderboard's ranking.The top three carry their rank at every stop.Where threads cross, the leaderboards disagree. Tap a thread to pin it.

The leaderboard

Every model, ranked on the consensus of the leaderboards

Each source column shows the model's rank there and its score on that source's own scale. Sort by what matters to you, search, or narrow to one lab.

Ranked on the consensus of every leaderboard that measures the model.

AI models ranked by most capable, with each leaderboard's rank and the cheapest list price
#ModelConsensusAALiveBenchArenaBFCLPrice / 1MContext
1Claude Fable 5Anthropic85#249.6#479.5#11507–$20OpenRouter1M
2Claude Opus 5Anthropic83#150.8#180.3––$10OpenRouter1M
3GPT-5.6 SolOpenAI83#347.0#379.7#81485–$4.00OpenRouter1.1M
4Kimi K3Moonshot AI80#443.6#678.5#71486–$2.86OpenRouter1.0M
5GPT-5.5OpenAI80#838.4#279.9#111482–$11OpenRouter1.1M
6gemini-3-proGoogle79––#61486#372.51%––
7Claude Opus 4.8Anthropic79#641.8#578.9#101484–$10OpenRouter1M
8Claude Opus 4.7 xHigh EffortAnthropic74–#976.5#31502–$10Vercel AI Gateway–
9GPT 5.4 xHighOpenAI73–#778.0#121478–$5.63Vercel AI Gateway–
10Gemini 3.1 Pro PreviewGoogle72#1829.7#877.1#51486–$4.50OpenRouter1.0M
11Muse Spark 1.1Meta72#1433.7#1176.2#41495–$2.00OpenRouter1.0M
12Grok 4.5xAI72#738.8#1076.2#191468–$3.00OpenRouter500K
13Claude Sonnet 5Anthropic69#938.2#1274.8#201461–$4.00OpenRouter1M
14Claude Opus 4.6 Thinking High EffortAnthropic69–#1574.5#21505–$10Vercel AI Gateway–
15Claude Opus 4.5 20251101 Thinking 64K High EffortAnthropic68–#2172.6#151473#177.47%$10Vercel AI Gateway–
16Gemini 3.5 FlashGoogle67#1632.6#1374.6#131476–$3.38OpenRouter1.0M
17Gemini 3.6 FlashGoogle67#1334.0#1773.6#91485–$1.50OpenRouter1.0M
18GPT-5.6 TerraOpenAI66#542.1#1674.3––$4.50OpenRouter1.1M
19grok-4-1-fast-reasoningxAI65––#361431#569.57%––
20Claude Sonnet 4.5Anthropic64#3020.7–#251456#273.24%$6.00OpenRouter1M

Head to head

Compare any two AI models

Every leaderboard that measures both, side by side. The longer bar wins the row.

Claude Opus 5 leads on 2 of 2 leaderboards that measure both.

  1. 85Consensus0 to 10083
  2. #2 · 49.6Artificial Analysisshare of models beaten#1 · 50.8
  3. #4 · 79.5LiveBenchshare of models beaten#1 · 80.3
  4. #1 · 1507LMArenashare of models beatenNot measured
  5. Not measuredBerkeley BFCLshare of models beatenNot measured
  6. $20Priceblended, per 1M tokens$10

Capability against price

What each point of capability costs

The line joins the cheapest model at each step up in consensus. gpt-oss-20b is the cheapest on it, at $0.036 per million tokens; Claude Fable 5 tops it at $20.

Best value line

Gateways

The same model, priced by OpenRouter and Vercel AI Gateway

Blended list price per million tokens where the two gateways differ. Providers behind each gateway set most of these prices, so gaps come and go.

Blended price per million tokens by gateway
ModelOpenRouterVercel AI Gateway
GPT-5.6 Sol$4.00$8.00
Kimi K3$2.86$6.00
GLM 5.1$1.48$2.15
GLM 5.2$2.15$1.24
Qwen3.7 Max$2.21$3.75
Qwen3.7 Plus$0.56$0.70
DeepSeek V4 Pro 0423$0.32$0.99
Kimi K2.6$1.34$1.71
DeepSeek V4 Flash 0423$0.052$0.095
MiniMax M2.7$0.37$0.52
Kimi K2.7 Code$1.34$1.71
DeepSeek V3.1 Terminus$0.47$0.45

Sources

Sources, method and how to cite this leaderboard

Free to use with attribution. Every score belongs to the leaderboard that measured it; the consensus and the value ranking are Rank.ai's.

Sources and method

  • Artificial Analysis

    Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations. Via Artificial Analysis indices on the OpenRouter Models API, read Oct 1, 2026.

  • LiveBench

    Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language. Via LiveBench official release, published as a Rank.ai board, read Jun 25, 2026.

  • LMArena

    Blind head-to-head human votes, style controlled. Via Arena official leaderboard dataset, published as a Rank.ai board, read Jul 21, 2026.

  • Berkeley BFCL

    Function calling and tool use, single turn to multi turn and web search. Via UC Berkeley BFCL V4 official leaderboard, published as a Rank.ai board, read Apr 13, 2026.

  • OpenRouter prices

    List prices per input and output token, read Oct 1, 2026.

  • Vercel AI Gateway prices

    List prices per input and output token, read Oct 1, 2026.

Models are matched across sources by name, with lab prefixes, dates and reasoning settings removed, so “GPT 5.6 xHigh” on LiveBench and “openai/gpt-5.6” on OpenRouter are one model. Each source keeps the model's best run.

On each source, a model scores the share of the other ranked models it beats, 0 to 100. The consensus averages those with one neutral score of 50 added, and needs at least two sources. Best value divides the consensus by the blended price, for models scoring 50 or more.

Read the full method

Download the data

  • LLM leaderboard (CSV)

    85 models: consensus, rank and score on each source, gateway prices, context window.

Last updated October 1, 2026

Questions about the LLM leaderboard

What is the best AI model right now?
Claude Fable 5 from Anthropic, with a consensus score of 85 out of 100 across 3 leaderboards, ahead of Claude Opus 5 at 83. The consensus combines Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, so a model has to do well on several independent tests to rank first.
What is the best AI model for coding?
Claude Opus 5, with 78.0 on the Artificial Analysis Coding Index, the highest of the models on this leaderboard. Sort the table by Coding to see the rest.
What is the best value AI model?
DeepSeek V4 Pro 0423. It scores 58 on the consensus and lists at $0.32 per million blended tokens on OpenRouter, the most capability per dollar among models in the top half of the board.
Which AI model do people prefer?
Claude Fable 5 has the highest LMArena rating of the models here, 1507, from blind votes where people compare two answers without knowing which model wrote them.
How is the consensus score calculated?
For each leaderboard, a model gets the share of the other ranked models it beats there, from 0 to 100. The consensus is the average of those shares, with one extra neutral score of 50 added, so a model measured by two leaderboards cannot outrank one measured by five on a lucky pair. A model needs at least two leaderboards to be ranked. Each source keeps the model's best run, for example its highest reasoning setting.
Why do AI leaderboards disagree?
They test different things. Artificial Analysis and LiveBench grade answers to exam-style questions, Vals AI grades enterprise tasks against expert answers, LMArena counts which answer people prefer, and BFCL checks whether tool calls are correct. A model tuned for chat can win votes and trail on maths, which is why this page shows every source's rank next to the consensus.
Where do the prices come from?
List prices per token from OpenRouter and Vercel AI Gateway, blended at three input tokens to one output token, the ratio Artificial Analysis uses. Where both gateways list a model, the cheaper one is shown. Prices exclude caching discounts and batch rates.
How often is the LLM leaderboard updated?
Model prices and the Artificial Analysis scores are re-read from OpenRouter and Vercel AI Gateway every 10 minutes, and the LiveBench, LMArena and BFCL boards every hour, so the blend moves as soon as a leaderboard publishes. When each source was last read is listed under Sources and method.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first article.