Free public boards

Models

How the frontier models compare on capability, human preference, tool use, speed and price, kept apart so one score never hides a tradeoff.

One number cannot tell you which model to ship.
Capability, price and speed pull against each other.

So they stay on separate boards, with the release each score came from.

Pick your board

Every live models board

How these numbers are made →

All 7 boards are free to read and link to the evidence behind every row. Refresh rates differ by board and are shown on each one.

Model rank

How the frontier models score when the same tests are run on all of them.

  1. 1claude-fable-51,507 arena score
  2. 2claude-opus-4-6-thinking1,505 arena score
  3. 3claude-opus-4-7-thinking1,502 arena score
The top three sit within 5.3 of each other.378 models
  1. 1Claude-Opus-4-5-20251101 (FC)77%
  2. 2Claude-Sonnet-4-5-20250929 (FC)73%
  3. 3Gemini-3-Pro-Preview (Prompt)73%
Claude-Opus-4-5-20251101 (FC) leads by 4.2 points.109 models
  1. 1all-MiniLM-L6-v2253.38M downloads
  2. 2Qwen3 0.6B28.51M downloads
  3. 3Qwen3 8B16.76M downloads
all-MiniLM-L6-v2 is 236615494 clear of third.12 models

Model efficiency

What that capability costs per dollar, per token and per second.

  1. 1DeepSeek V4 Flash$0.016
  2. 2Grok Build 0.1$0.024
  3. 3DeepSeek V4 Pro$0.050
The top three sit within 0 of each other.45 models