Skip to content

Microsoft

phi-4 benchmarks, ranks and price

Where Artificial Analysis, LiveBench, LMArena and Berkeley BFCL place phi-4, what it costs through OpenRouter and Vercel AI Gateway, and the models either side of it on the consensus.

Updated Oct 1, 2026Full leaderboardSources and method

Read Oct 1, 2026

#72of 85

on the LLM leaderboard, with a consensus score of 28 across 2 leaderboards.

Source: Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, combined by Rank.ai.

By leaderboard

Where each leaderboard places phi-4

Artificial Analysis

Not measured by Artificial Analysis.

Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.

LiveBench

Not measured by LiveBench.

Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language.

LMArena

#73of 79

Score 1256. Best run: phi-4.

Blind head-to-head human votes, style controlled.

Berkeley BFCL

#28of 37

Score 28.79%. Best run: Phi-4 (Prompt).

Function calling and tool use, single turn to multi turn and web search.

Price and context

–

No gateway lists it.

Nearby

The models either side of phi-4

  1. #69o3 Mini High30
  2. #70llama-4-scout-17b-16e-instruct30
  3. #71Mistral Large 3 251228
  4. #72phi-428
  5. #73Gemma 3 27B27
  6. #74Qwen3 30B A3B Thinking 250726
  7. #75Gemma 3 12B25

Questions about phi-4

How good is phi-4?
phi-4 ranks #72 of 85 models on the Rank.ai LLM leaderboard, with a consensus score of 28 out of 100 across 2 leaderboards.
Is phi-4 better than Mistral Large 3 2512?
Not on the consensus: Mistral Large 3 2512 scores 28 to phi-4's 28. Compare them source by source below, since the leaderboards test different things.
Is phi-4 better than Gemma 3 27B?
On the consensus, yes: 28 to 27.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first article.