Benchmark

LongBench v2

LongBench v2 is a benchmark designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across…

Modality
text
Categories
long_context, reasoning, structured_output, general
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3DeepSeek0.4872024-12-25evidence
Kimi K2.5Moonshot AI0.612026-01-27evidence
MAI-Thinking-1Microsoft0.612026-06-02evidence
MiMo-V2-FlashXiaomi0.6062025-12-16evidence
MiniMax M1 40KMiniMax0.612025-06-16evidence
MiniMax M1 80KMiniMax0.6152025-06-16evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.6192026-06-04evidence
Qwen3.5-0.8BQwen0.2612026-03-02evidence
Qwen3.5-122B-A10BQwen0.6022026-02-24evidence
Qwen3.5-27BQwen0.6062026-02-24evidence
Qwen3.5-2BQwen0.3872026-03-02evidence
Qwen3.5-35B-A3BQwen0.592026-02-24evidence
Qwen3.5-397B-A17BQwen0.6322026-02-16evidence
Qwen3.5-4BQwen0.52026-03-02evidence
Qwen3.5-9BQwen0.5522026-03-02evidence
Qwen3.6 PlusQwen0.622026-04-02evidence
Qwen3.8 MaxQwen0.6632026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.