Benchmark

AA-LCR

Agent Arena Long Context Reasoning benchmark

Modality
text
Categories
long_context, reasoning
Openness
unknown
Source
llm_stats
Reported scores
18

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Hy3Tencent0.7342026-07-06evidence
Kimi K2.5Moonshot AI0.72026-01-27evidence
MiniMax M2.1MiniMax0.622025-12-23evidence
Mistral Small 4Mistral0.7122026-03-16evidence
Muse Glimmer-30BMeta0.82026-08-10evidence
Nemotron 3 Super (120B A12B)NVIDIA0.58312026-03-11evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.6542026-06-04evidence
Nemotron 3.5 Lightning (30B A3B)NVIDIA0.522026-08-11evidence
Qwen3.5-0.8BQwen0.0472026-03-02evidence
Qwen3.5-122B-A10BQwen0.6692026-02-24evidence
Qwen3.5-27BQwen0.6612026-02-24evidence
Qwen3.5-2BQwen0.2562026-03-02evidence
Qwen3.5-35B-A3BQwen0.5852026-02-24evidence
Qwen3.5-397B-A17BQwen0.6872026-02-16evidence
Qwen3.5-4BQwen0.572026-03-02evidence
Qwen3.5-9BQwen0.632026-03-02evidence
Qwen3.6 PlusQwen0.6832026-04-02evidence
Solar Pro 4Upstage0.712026-08-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.