Exact model configurations ranked by the official objective LiveBench release. This is a single-family capability orientation, not an absolute composite.
BFCL tool use configurations
109 models ranked by BFCL overall accuracy.
Claude-Opus-4-5-20251101 (FC) is out in front at 77.47%.
109 modelschecked Apr 13, 2026, 11:59 PM UTCdailyhow we measure this →CSV / JSON
Claude-Opus-4-5-20251101 (FC)
77.47%BFCL overall accuracy
- Ahead of second place
- 4.2%
- Held by the top three
- 5%
- Middle of the board
- 35.5%
- Below half the leader
- 64 of 109
GPT-5-nano-2025-08-07 (FC)
Ranks exact model-plus-inference configurations by official UC Berkeley BFCL overall accuracy.
Some explanation inputs are unavailable Missing: publicFrozenEvidence, comparablePriorRanking.
- Published rank
- #24
- Score
- 51.45 percent
- Sample size
- 11
100% confidence · 58% component coverage · as of Apr 13, 2026, 11:59 PM UTC
Component ledger
| Component | Value | Weight | Contribution | Evidence | Status |
|---|---|---|---|---|---|
| benchmarkVersion | v4 | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| bfcl evaluation total cost USD | — | 0% | 0 | 1 | available |
| bfcl irrelevance detection accuracy | — | 0% | 0 | 1 | available |
| bfcl latency mean seconds | — | 0% | 0 | 1 | available |
| bfcl latency p95 seconds | — | 0% | 0 | 1 | available |
| bfcl live accuracy | — | 0% | 0 | 1 | available |
| bfcl memory accuracy | — | 0% | 0 | 1 | available |
| bfcl multi turn accuracy | — | 0% | 0 | 1 | available |
| bfcl non live ast accuracy | — | 0% | 0 | 1 | available |
| bfcl overall accuracy | — | 100% | 51.45 | 1 | available |
| bfcl relevance detection accuracy | — | 0% | 0 | 1 | available |
| bfcl web search accuracy | — | 0% | 0 | 1 | available |
| interactionMode | FC | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| metrics | Structured | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| organization | OpenAI | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| rankingDirection | descending | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| reviewedAsOf | 2026-04-13 | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| sourceConfiguration | GPT-5-nano-2025-08-07 (FC) | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| upstreamRank | 24 | Unavailable | 24 | 0 | unavailableno_public_frozen_evidence |
new since prior snapshot
no_prior_published_snapshot
No component moved since the prior snapshot.
Source trail
BFCL evaluation total costbfcl evaluation total cost USD · observed Apr 13, 2026, 11:59 PM UTC$8.79100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL irrelevance-detection accuracybfcl irrelevance detection accuracy · observed Apr 13, 2026, 11:59 PM UTC89.1 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL mean latencybfcl latency mean seconds · observed Apr 13, 2026, 11:59 PM UTC10.36 seconds100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL p95 latencybfcl latency p95 seconds · observed Apr 13, 2026, 11:59 PM UTC23.56 seconds100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL live accuracybfcl live accuracy · observed Apr 13, 2026, 11:59 PM UTC59.44 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL memory accuracybfcl memory accuracy · observed Apr 13, 2026, 11:59 PM UTC24.73 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL multi-turn accuracybfcl multi turn accuracy · observed Apr 13, 2026, 11:59 PM UTC34.5 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL non-live AST accuracybfcl non live ast accuracy · observed Apr 13, 2026, 11:59 PM UTC68 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL overall accuracybfcl overall accuracy · observed Apr 13, 2026, 11:59 PM UTC51.45 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL relevance-detection accuracybfcl relevance detection accuracy · observed Apr 13, 2026, 11:59 PM UTC75 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL web-search accuracybfcl web search accuracy · observed Apr 13, 2026, 11:59 PM UTC72.5 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
More in Model rank
Capability-first rankings across frontier and open models.
Official style-controlled text human-preference ratings with confidence bounds, variance, vote counts, and a frozen dataset revision.
Hugging Face repository downloads and engagement for a fixed, reviewed model cohort.