Benchmark
BFCL v2
Berkeley Function Calling Leaderboard (BFCL) v2 is a comprehensive benchmark for evaluating large language models' function calling capabilities. It…
- Modality
- text
- Categories
- reasoning, general, tool_calling
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 5
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Llama 3.1 Nemotron Nano 8B V1 | NVIDIA | 0.636 | 2025-03-18 | evidence |
| Llama 3.1 Nemotron Ultra 253B v1 | NVIDIA | 0.741 | 2025-04-07 | evidence |
| Llama 3.2 3B Instruct | Meta | 0.67 | 2024-09-25 | evidence |
| Llama 3.3 70B Instruct | Meta | 0.773 | 2024-12-06 | evidence |
| Llama-3.3 Nemotron Super 49B v1 | NVIDIA | 0.737 | 2025-03-18 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.