Benchmark

BFCL-V4

Berkeley Function Calling Leaderboard V4 (BFCL-V4) evaluates LLMs on their ability to accurately call functions and APIs, including simple, multiple…

Modality
text
Categories
agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
15

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
LFM2.5-2.6BLiquid AI0.56882026-08-04evidence
LFM2.5-VL-3BLiquid AI0.3252026-08-12evidence
Nova 2 LiteAmazon0.6032025-12-02evidence
Nova 2 OmniAmazon0.5832025-12-02evidence
Nova 2 ProAmazon0.6162025-12-02evidence
Qwen3.5-0.8BQwen0.2532026-03-02evidence
Qwen3.5-122B-A10BQwen0.7222026-02-24evidence
Qwen3.5-27BQwen0.6852026-02-24evidence
Qwen3.5-2BQwen0.4362026-03-02evidence
Qwen3.5-35B-A3BQwen0.6732026-02-24evidence
Qwen3.5-397B-A17BQwen0.7292026-02-16evidence
Qwen3.5-4BQwen0.5032026-03-02evidence
Qwen3.5-9BQwen0.6612026-03-02evidence
Qwen3.7 MaxQwen0.752026-05-19evidence
Qwen3.7-PlusQwen0.7292026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.