Benchmark

BFCL

The Berkeley Function Calling Leaderboard (BFCL) is the first comprehensive and executable function call evaluation dedicated to assessing Large Language…

Modality
text
Categories
reasoning, general, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Llama 3.1 405B InstructMeta0.8852024-07-23evidence
Llama 3.1 70B InstructMeta0.8482024-07-23evidence
Llama 3.1 8B InstructMeta0.7612024-07-23evidence
Nova 2 SonicAmazon0.7452025-12-02evidence
Nova LiteAmazon0.6662024-11-20evidence
Nova MicroAmazon0.5622024-11-20evidence
Nova ProAmazon0.6842024-11-20evidence
Qwen3 235B A22BQwen0.7082025-04-29evidence
Qwen3 30B A3BQwen0.6912025-04-29evidence
Qwen3 32BQwen0.7032025-04-29evidence
QwQ-32BQwen0.6642025-03-05evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.