Benchmark

Vending-Bench 2

Vending-Bench 2 tests longer horizon planning capabilities by evaluating how well AI models can manage a simulated vending machine business over extended…

Modality
text
Categories
reasoning, agents
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4.6Anthropic8017.592026-02-05evidence
Gemini 3 FlashGoogle3635.02025-12-17evidence
Gemini 3 ProGoogle5478.162025-11-18evidence
GLM-5.1Z.ai5634.412026-04-07evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.