Skip to content

xAI

Grok Build 0.1 benchmarks, ranks and price

Where Artificial Analysis, LiveBench, LMArena and Berkeley BFCL place Grok Build 0.1, what it costs through OpenRouter and Vercel AI Gateway, and the models either side of it on the consensus.

Updated Oct 5, 2026Full leaderboardSources and method

Read Oct 5, 2026

#68of 126

on the LLM leaderboard, with a consensus score of 48 across 2 leaderboards.

Source: Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, combined by Rank.ai.

By leaderboard

Where each leaderboard places Grok Build 0.1

Artificial Analysis

#29of 115

Score 27.2. Best run: x-ai/grok-build-0.1.

Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.

LiveBench

#30of 36

Score 67.8. Best run: Grok Build 0.1.

Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language.

LMArena

Not measured by LMArena.

Blind head-to-head human votes, style controlled.

Berkeley BFCL

Not measured by Berkeley BFCL.

Function calling and tool use, single turn to multi turn and web search.

Price and context

$1.25

OpenRouter: $1.00 in, $2.00 out Context window 256K tokens. LiveBench cost per solved task $0.024.

Nearby

The models either side of Grok Build 0.1

  1. #65Mistral Medium 3.549
  2. #66o4 Mini48
  3. #67Gemini 3.5 Flash Lite48
  4. #68Grok Build 0.148
  5. #69Kimi K2.7 Code47
  6. #70DeepSeek V3.1 Terminus47
  7. #71R147

Questions about Grok Build 0.1

How good is Grok Build 0.1?
Grok Build 0.1 ranks #68 of 126 models on the Rank.ai LLM leaderboard, with a consensus score of 48 out of 100 across 2 leaderboards.
Is Grok Build 0.1 better than Gemini 3.5 Flash Lite?
Not on the consensus: Gemini 3.5 Flash Lite scores 48 to Grok Build 0.1's 48. Compare them source by source below, since the leaderboards test different things.
Is Grok Build 0.1 better than Kimi K2.7 Code?
On the consensus, yes: 48 to 47.
How much does Grok Build 0.1 cost?
$1.00 per million input tokens and $2.00 per million output tokens on OpenRouter, the cheaper listing we read. Blended at three input tokens to one output, that is $1.25 per million.
Is Grok Build 0.1 good for coding?
It scores 51.5 on the Artificial Analysis Coding Index, placing it #32 for coding among the models on this leaderboard.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first article.