Rank.ai
Menu
Live
RANK.AI / PUBLISHED METHOD

BFCL v4 tool-use configuration ranking reviewed 2026-04-13

Ranks exact model-plus-inference configurations by official UC Berkeley BFCL overall accuracy.

Version
bfcl-v4-2026-04-13-official-v1-bfcl-tool-use-configurations-v1
Published
Jul 27, 2026
Live boards
1
Method ID
bfcl-tool-use-configurations

How the order is produced

official BFCL Overall Acc descending

When the method reruns

  • daily

Boards using this version

Only live, quality-gated boards appear. Each board opens on the exact published table and evidence path.

BoardPrimary measureRowsConfidenceFreshnessBoard versionCadenceEvidence
BFCL Tool-Use Configurationsbfcl-tool-use-configurationsbfcl-overall-accuracy109100% currentbfcl-v4-2026-04-13-official-v1-bfcl-tool-use-configurations-v1dailyOpen board →

Complete measurement ledger

These fields are stored with the methodology version. A changed contract must publish a new version before comparable movement can resume.

Source

bfcl-official

License

Apache-2.0

Evidence

Every ranking component links to the exact normalized observation and the immutable source-artifact record.

Identity

Rows retain BFCL's exact configuration label. Function-calling, prompt, thinking, and unspecified modes remain distinct systems.

Universe

All 109 exact model-plus-inference configurations in the SHA-pinned official BFCL leaderboard capture.

Benchmark

Berkeley Function Calling Leaderboard

Versioning

The official CSV URL, SHA-256 digest, exact row count, complete header, ordering, and reviewed date are code-pinned.

Limitations

  • This is an independent tool-use benchmark family, not an absolute model score.
  • Rows compare model-plus-inference configurations, not base weights alone.
  • Total cost and latency are BFCL workload context, not normalized market prices.
  • Scores may not generalize beyond BFCL V4 tool-use tasks.
  • A changed upstream CSV is suppressed until a new capture is reviewed and pinned.

Missing Data

Publication fails unless every configuration has every required metric. Optional format-sensitivity values are omitted when BFCL reports N/A; no value is imputed.

Rank Formula

official BFCL Overall Acc descending

Tie Breakers

  • official BFCL upstream rank ascending
  • canonical entity slug ascending

Reviewed As Of

2026-04-13

Primary Metric

bfcl-overall-accuracy

Optional Metrics

  • bfcl-format-sensitivity-max-delta
  • bfcl-format-sensitivity-standard-deviation

Benchmark Version

v4

Ranking Direction

descending

Required Metrics Per Entity

  • bfcl-overall-accuracy
  • bfcl-non-live-ast-accuracy
  • bfcl-live-accuracy
  • bfcl-multi-turn-accuracy
  • bfcl-web-search-accuracy
  • bfcl-memory-accuracy
  • bfcl-relevance-detection-accuracy
  • bfcl-irrelevance-detection-accuracy
  • bfcl-evaluation-total-cost-usd
  • bfcl-latency-mean-seconds
  • bfcl-latency-p95-seconds
BFCL v4 tool-use configuration ranking reviewed 2026-04-13 | Rank.ai