Artificial Analysis
Not measured by Artificial Analysis.
Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.
Anthropic
Where Artificial Analysis, LiveBench, LMArena and Berkeley BFCL place Claude Opus 4.5 20251101 Thinking 64K High Effort, what it costs through OpenRouter and Vercel AI Gateway, and the models either side of it on the consensus.
Read Oct 1, 2026
#15of 85
on the LLM leaderboard, with a consensus score of 68 across 3 leaderboards.
Source: Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, combined by Rank.ai.
By leaderboard
Artificial Analysis
Not measured by Artificial Analysis.
Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.
LiveBench
#21of 34
Score 72.6. Best run: Claude Opus 4.5 20251101 Thinking 64K High Effort.
Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language.
LMArena
#15of 79
Score 1473. Best run: claude-opus-4-5-20251101-thinking-32k.
Blind head-to-head human votes, style controlled.
Berkeley BFCL
#1of 37
Score 77.47%. Best run: Claude-Opus-4-5-20251101 (FC).
Function calling and tool use, single turn to multi turn and web search.
Price and context
$10
Vercel AI Gateway: $5.00 in, $25 out LiveBench cost per solved task $0.610.
Nearby
Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.