RANK.AI DATA / LATEST PUBLICATION

LiveBench model capability

45 models ranked by LiveBench overall score.
GPT 5.6 Sol Max is out in front at 82.4%.

45 modelschecked Jun 25, 2026, 11:59 PM UTCdaily88 failed refresheshow we measure this →CSV / JSON

Follow changes
In front today

GPT 5.6 Sol Max

82.4%LiveBench overall score

Ahead of second place
1.6%
Held by the top three
7%
Middle of the board
74.3%
Highest 82.4%Middle 74.3%
1. GPT 5.6 Sol Max — 82.4%2. Claude Fable 5 Max Effort — 80.8%3. Claude Opus 5 xHigh Effort — 80.3%4. GPT 5.5 xHigh — 79.9%5. GPT 5.6 Terra Max — 79.8%6. GPT 5.6 Sol xHigh — 79.7%7. Claude Fable 5 xHigh Effort — 79.5%8. Claude Opus 5 Max Effort — 79.2%9. Claude Opus 5 High Effort — 79.0%10. Claude Opus 4 8 xHigh Effort — 78.9%11. GPT 5.5 High — 78.7%12. Kimi K3 — 78.5%13. GPT 5.4 xHigh — 78.0%14. Gemini 3.1 Pro Preview High — 77.1%15. Claude Opus 4 7 xHigh Effort — 76.5%16. Grok 4.5 — 76.2%17. Muse Spark 1.1 xHigh — 76.2%18. Muse Spark 1.1 High — 75.1%19. Claude Sonnet 5 xHigh Effort — 74.8%20. Gemini 3.5 Flash High — 74.6%21. GPT 5.2 2025 12 11 High — 74.6%22. Claude Opus 4 6 Thinking Auto High Effort — 74.5%23. GPT 5.6 Luna Max — 74.3%24. GPT 5.6 Terra xHigh — 74.3%25. GPT 5.2 Codex — 74.0%26. Gemini 3.6 Flash High — 73.6%27. GLM 5.2 — 73.2%28. Qwen3.7 Max — 73.1%29. Claude Sonnet 4 6 Thinking Auto Medium Effort — 73.0%30. Claude Opus 4 5 20251101 Thinking 64K High Effort — 72.6%31. GPT 5.5 — 72.2%32. Inkling xHigh — 71.7%33. DeepSeek V4 Pro — 71.6%34. GPT 5.6 Luna xHigh — 71.0%35. Kimi K2.6 Thinking — 70.5%36. GPT 5.4 Nano xHigh — 69.6%37. Qwen3.6 Plus — 68.9%38. Kimi K2.7 Code — 68.4%39. Grok Build 0.1 — 67.8%40. MiniMax M3 — 67.3%41. GPT 5.4 Mini xHigh — 66.4%42. DeepSeek V4 Flash — 65.5%43. Qwen3.6 27b — 64.0%44. Gemini 3.5 Flash Lite High — 63.9%45. Grok 4.3 — 62.2%
Rank 1one bar per published rowRank 45
WHY THIS RANK / VERIFIED SNAPSHOT

GPT 5.6 Luna xHigh

Ranks exact model configurations by the mean of seven official LiveBench category averages.

Some explanation inputs are unavailable Missing: publicFrozenEvidence, comparablePriorRanking.

Published rank
#34
Score
71.007
Sample size
8

100% confidence · 67% component coverage · as of Jun 25, 2026, 11:59 PM UTC

SCORE CONSTRUCTION

Component ledger

8/12 evidenced
Components contributing to GPT 5.6 Luna xHigh's rank
ComponentValueWeightContributionEvidenceStatus
livebench agentic coding score0%01available
livebench coding score0%01available
livebench data analysis score0%01available
livebench instruction following score0%01available
livebench language score0%01available
livebench mathematics score0%01available
livebench overall score100%71.0071available
livebench reasoning score0%01available
metricsStructuredUnavailable0unavailableno_public_frozen_evidence
rankingDirectiondescendingUnavailable0unavailableno_public_frozen_evidence
release2026-06-25Unavailable0unavailableno_public_frozen_evidence
sourceModelIdgpt-5.6-luna-xhighUnavailable0unavailableno_public_frozen_evidence
WHY IT MOVED

new since prior snapshot

No rank delta

no_prior_published_snapshot

No component moved since the prior snapshot.

PRIMARY EVIDENCE

Source trail

8 public records
  • LiveBench Agentic Coding scorelivebench agentic coding score · observed Jun 25, 2026, 11:59 PM UTC48.838 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench Coding scorelivebench coding score · observed Jun 25, 2026, 11:59 PM UTC76.715 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench Data Analysis scorelivebench data analysis score · observed Jun 25, 2026, 11:59 PM UTC72.897 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench IF scorelivebench instruction following score · observed Jun 25, 2026, 11:59 PM UTC57.446 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench Language scorelivebench language score · observed Jun 25, 2026, 11:59 PM UTC70.221 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench Mathematics scorelivebench mathematics score · observed Jun 25, 2026, 11:59 PM UTC86.269 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench overall scorelivebench overall score · observed Jun 25, 2026, 11:59 PM UTC71.007 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
  • LiveBench Reasoning scorelivebench reasoning score · observed Jun 25, 2026, 11:59 PM UTC84.664 percent100% confidence
    Source
    LiveBench official model benchmark
    Grade
    A
    Freshness
    fresh
    Locator
    {"scoreAsset":"https://livebench.ai/table_2026_06_25.csv","aggregation":"mean_of_category_means","categoryAsset":"https://livebench.ai/categories_2026_06_25.json"}
    Open primary evidence ↗
REPRODUCIBILITYLiveBench 2026-06-25 model capability · vlivebench-2026-06-25-official-v1-livebench-model-capability-v1360 observations · 1 sources · passed quality
Snapshot
7b5efe94-c8bd-447f-90d9-4f4f2bd55271
Data hash
f2cdb2a65d1683993ff9ccba4cf399472ea8fa2a31d795f741c694957e64f219
Method hash
b04d138a8fea8028e2a5a4e9656eb410bd1d8cf6958a50b6e106d211946f733e

More in Model rank

Capability-first rankings across frontier and open models.

3 live tables
Arena text preference

Official style-controlled text human-preference ratings with confidence bounds, variance, vote counts, and a frozen dataset revision.

BFCL tool-use configurations

Exact model-plus-inference configurations ranked by UC Berkeley BFCL V4 overall accuracy, with tool-use breakdowns, cost, latency, and source evidence.