Skip to content

Alibaba

qwen3-235b-a22b-instruct-2507 benchmarks, ranks and price

Where Artificial Analysis, LiveBench, LMArena and Berkeley BFCL place qwen3-235b-a22b-instruct-2507, what it costs through OpenRouter and Vercel AI Gateway, and the models either side of it on the consensus.

Updated Oct 1, 2026Full leaderboardSources and method

Read Oct 1, 2026

#36of 85

on the LLM leaderboard, with a consensus score of 52 across 2 leaderboards.

Source: Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, combined by Rank.ai.

By leaderboard

Where each leaderboard places qwen3-235b-a22b-instruct-2507

Artificial Analysis

Not measured by Artificial Analysis.

Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.

LiveBench

Not measured by LiveBench.

Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language.

LMArena

#41of 79

Score 1423. Best run: qwen3-235b-a22b-instruct-2507.

Blind head-to-head human votes, style controlled.

Berkeley BFCL

#16of 37

Score 52.15%. Best run: Qwen3-235B-A22B-Instruct-2507 (Prompt).

Function calling and tool use, single turn to multi turn and web search.

Price and context

–

No gateway lists it.

Nearby

The models either side of qwen3-235b-a22b-instruct-2507

  1. #33gemini-2.5-flash54
  2. #34Claude Haiku 4.553
  3. #35Kimi K2.653
  4. #36qwen3-235b-a22b-instruct-250752
  5. #37gpt-4.1-2025-04-1452
  6. #38Inkling xHigh51
  7. #39Command A50

Questions about qwen3-235b-a22b-instruct-2507

How good is qwen3-235b-a22b-instruct-2507?
qwen3-235b-a22b-instruct-2507 ranks #36 of 85 models on the Rank.ai LLM leaderboard, with a consensus score of 52 out of 100 across 2 leaderboards.
Is qwen3-235b-a22b-instruct-2507 better than Kimi K2.6?
Not on the consensus: Kimi K2.6 scores 53 to qwen3-235b-a22b-instruct-2507's 52. Compare them source by source below, since the leaderboards test different things.
Is qwen3-235b-a22b-instruct-2507 better than gpt-4.1-2025-04-14?
On the consensus, yes: 52 to 52.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first article.