Skip to content

Meta

llama-3.3-70b-instruct benchmarks, ranks and price

Where Artificial Analysis, LiveBench, LMArena and Berkeley BFCL place llama-3.3-70b-instruct, what it costs through OpenRouter and Vercel AI Gateway, and the models either side of it on the consensus.

Updated Oct 1, 2026Full leaderboardSources and method

Read Oct 1, 2026

#66of 85

on the LLM leaderboard, with a consensus score of 32 across 2 leaderboards.

Source: Artificial Analysis, LiveBench, LMArena and Berkeley BFCL, combined by Rank.ai.

By leaderboard

Where each leaderboard places llama-3.3-70b-instruct

Artificial Analysis

Not measured by Artificial Analysis.

Intelligence Index: a composite of reasoning, knowledge, maths and coding evaluations.

LiveBench

Not measured by LiveBench.

Contamination-limited questions refreshed each release: reasoning, maths, coding, data analysis, language.

LMArena

#68of 79

Score 1318. Best run: llama-3.3-70b-instruct.

Blind head-to-head human votes, style controlled.

Berkeley BFCL

#25of 37

Score 31.90%. Best run: Llama-3.3-70B-Instruct (FC).

Function calling and tool use, single turn to multi turn and web search.

Price and context

–

No gateway lists it.

Nearby

The models either side of llama-3.3-70b-instruct

  1. #63Qwen3 235B A22B Thinking 250735
  2. #64gpt-4.1-nano-2025-04-1434
  3. #65Qwen3.6 27B33
  4. #66llama-3.3-70b-instruct32
  5. #67Mistral Small 432
  6. #68gpt-oss-120b31
  7. #69o3 Mini High30

Questions about llama-3.3-70b-instruct

How good is llama-3.3-70b-instruct?
llama-3.3-70b-instruct ranks #66 of 85 models on the Rank.ai LLM leaderboard, with a consensus score of 32 out of 100 across 2 leaderboards.
Is llama-3.3-70b-instruct better than Qwen3.6 27B?
Not on the consensus: Qwen3.6 27B scores 33 to llama-3.3-70b-instruct's 32. Compare them source by source below, since the leaderboards test different things.
Is llama-3.3-70b-instruct better than Mistral Small 4?
On the consensus, yes: 32 to 32.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first article.