Benchmark

BFCL v2

Berkeley Function Calling Leaderboard (BFCL) v2 is a comprehensive benchmark for evaluating large language models' function calling capabilities. It…

Modality
text
Categories
reasoning, general, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Llama 3.1 Nemotron Nano 8B V1NVIDIA0.6362025-03-18evidence
Llama 3.1 Nemotron Ultra 253B v1NVIDIA0.7412025-04-07evidence
Llama 3.2 3B InstructMeta0.672024-09-25evidence
Llama 3.3 70B InstructMeta0.7732024-12-06evidence
Llama-3.3 Nemotron Super 49B v1NVIDIA0.7372025-03-18evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.