Exact model configurations ranked by the official objective LiveBench release. This is a single-family capability orientation, not an absolute composite.
BFCL tool use configurations
109 models ranked by BFCL overall accuracy.
Claude-Opus-4-5-20251101 (FC) is out in front at 77.47%.
109 modelschecked Apr 13, 2026, 11:59 PM UTCdailyhow we measure this →CSV / JSON
Claude-Opus-4-5-20251101 (FC)
77.47%BFCL overall accuracy
- Ahead of second place
- 4.2%
- Held by the top three
- 5%
- Middle of the board
- 35.5%
- Below half the leader
- 64 of 109
o4-mini-2025-04-16 (FC)
Ranks exact model-plus-inference configurations by official UC Berkeley BFCL overall accuracy.
Some explanation inputs are unavailable Missing: publicFrozenEvidence, comparablePriorRanking.
- Published rank
- #21
- Score
- 53.24 percent
- Sample size
- 11
100% confidence · 58% component coverage · as of Apr 13, 2026, 11:59 PM UTC
Component ledger
| Component | Value | Weight | Contribution | Evidence | Status |
|---|---|---|---|---|---|
| benchmarkVersion | v4 | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| bfcl evaluation total cost USD | — | 0% | 0 | 1 | available |
| bfcl irrelevance detection accuracy | — | 0% | 0 | 1 | available |
| bfcl latency mean seconds | — | 0% | 0 | 1 | available |
| bfcl latency p95 seconds | — | 0% | 0 | 1 | available |
| bfcl live accuracy | — | 0% | 0 | 1 | available |
| bfcl memory accuracy | — | 0% | 0 | 1 | available |
| bfcl multi turn accuracy | — | 0% | 0 | 1 | available |
| bfcl non live ast accuracy | — | 0% | 0 | 1 | available |
| bfcl overall accuracy | — | 100% | 53.24 | 1 | available |
| bfcl relevance detection accuracy | — | 0% | 0 | 1 | available |
| bfcl web search accuracy | — | 0% | 0 | 1 | available |
| interactionMode | FC | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| metrics | Structured | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| organization | OpenAI | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| rankingDirection | descending | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| reviewedAsOf | 2026-04-13 | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| sourceConfiguration | o4-mini-2025-04-16 (FC) | Unavailable | — | 0 | unavailableno_public_frozen_evidence |
| upstreamRank | 21 | Unavailable | 21 | 0 | unavailableno_public_frozen_evidence |
new since prior snapshot
no_prior_published_snapshot
No component moved since the prior snapshot.
Source trail
BFCL evaluation total costbfcl evaluation total cost USD · observed Apr 13, 2026, 11:59 PM UTC$81.91100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL irrelevance-detection accuracybfcl irrelevance detection accuracy · observed Apr 13, 2026, 11:59 PM UTC83.91 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL mean latencybfcl latency mean seconds · observed Apr 13, 2026, 11:59 PM UTC3.71 seconds100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL p95 latencybfcl latency p95 seconds · observed Apr 13, 2026, 11:59 PM UTC9.33 seconds100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL live accuracybfcl live accuracy · observed Apr 13, 2026, 11:59 PM UTC66.1 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL memory accuracybfcl memory accuracy · observed Apr 13, 2026, 11:59 PM UTC34.19 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL multi-turn accuracybfcl multi turn accuracy · observed Apr 13, 2026, 11:59 PM UTC41.75 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL non-live AST accuracybfcl non live ast accuracy · observed Apr 13, 2026, 11:59 PM UTC37.73 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL overall accuracybfcl overall accuracy · observed Apr 13, 2026, 11:59 PM UTC53.24 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL relevance-detection accuracybfcl relevance detection accuracy · observed Apr 13, 2026, 11:59 PM UTC81.25 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
BFCL web-search accuracybfcl web search accuracy · observed Apr 13, 2026, 11:59 PM UTC75.5 percent100% confidence
- Source
- Berkeley Function Calling Leaderboard
- Grade
- A
- Freshness
- fresh
- Locator
{"csv":"https://gorilla.cs.berkeley.edu/data_overall.csv","repository":"https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard","leaderboard":"https://gorilla.cs.berkeley.edu/leaderboard.html","rowIdentity":"Model","rankingColumn":"Overall Acc"}
More in Model rank
Capability-first rankings across frontier and open models.
Official style-controlled text human-preference ratings with confidence bounds, variance, vote counts, and a frozen dataset revision.
Hugging Face repository downloads and engagement for a fixed, reviewed model cohort.