Benchmark
BFCL_v3_MultiTurn
Berkeley Function Calling Leaderboard (BFCL) V3 MultiTurn benchmark that evaluates large language models' ability to handle multi-turn and multi-step…
- Modality
- text
- Categories
- reasoning, general, tool_calling
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 2
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| MiniMax M2.5 | MiniMax | 0.768 | 2026-02-12 | evidence |
| Nemotron Nano 9B v2 | NVIDIA | 0.669 | 2025-08-18 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.