Skip to content

Google

Gemma 3 12B benchmarks, ranks and price

Where Artificial Analysis, LiveBench, LMArena and Berkeley BFCL place Gemma 3 12B, what it costs through OpenRouter and Vercel AI Gateway, and the models either side of it on the consensus.

Updated Oct 1, 2026Full leaderboardSources and method

Read Oct 1, 2026

#75of 85

on the LLM leaderboard, with a consensus score of 25 across 3 leaderboards.

Source: Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, combined by Rank.ai.

By leaderboard

Where each leaderboard places Gemma 3 12B

Artificial Analysis

#51of 51

Score 3.8. Best run: google/gemma-3-12b-it.

Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.

LiveBench

Not measured by LiveBench.

Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language.

LMArena

#63of 79

Score 1342. Best run: gemma-3-12b-it.

Blind head-to-head human votes, style controlled.

Berkeley BFCL

#26of 37

Score 30.43%. Best run: Gemma-3-12b-it (Prompt).

Function calling and tool use, single turn to multi turn and web search.

Price and context

$0.075

OpenRouter: $0.050 in, $0.15 out Context window 131K tokens.

Nearby

The models either side of Gemma 3 12B

  1. #72phi-428
  2. #73Gemma 3 27B27
  3. #74Qwen3 30B A3B Thinking 250726
  4. #75Gemma 3 12B25
  5. #76llama-3.1-nemotron-ultra-253b-v124
  6. #77amazon-nova-pro-v1.024
  7. #78granite-3.1-8b-instruct24

Questions about Gemma 3 12B

How good is Gemma 3 12B?
Gemma 3 12B ranks #75 of 85 models on the Rank.ai LLM leaderboard, with a consensus score of 25 out of 100 across 3 leaderboards.
Is Gemma 3 12B better than Qwen3 30B A3B Thinking 2507?
Not on the consensus: Qwen3 30B A3B Thinking 2507 scores 26 to Gemma 3 12B's 25. Compare them source by source below, since the leaderboards test different things.
Is Gemma 3 12B better than llama-3.1-nemotron-ultra-253b-v1?
On the consensus, yes: 25 to 24.
How much does Gemma 3 12B cost?
$0.050 per million input tokens and $0.15 per million output tokens on OpenRouter, the cheaper listing we read. Blended at three input tokens to one output, that is $0.075 per million.
Is Gemma 3 12B good for coding?
It scores 5.8 on the Artificial Analysis Coding Index, placing it #51 for coding among the models on this leaderboard.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first article.