Benchmark

MRCR

MRCR (Multi-Round Coreference Resolution) is a synthetic long-context reasoning task where models must navigate long conversations to reproduce specific…

Modality
text
Categories
long_context, reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
7

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemini 1.5 FlashGoogle0.7192024-05-01evidence
Gemini 1.5 Flash 8BGoogle0.5472024-03-15evidence
Gemini 1.5 ProGoogle0.8262024-05-01evidence
Gemini 2.0 FlashGoogle0.6922024-12-01evidence
Gemini 2.5 FlashGoogle0.322025-05-20evidence
Gemini 2.5 ProGoogle0.932025-05-20evidence
MiMo-V2-FlashXiaomi0.4572025-12-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.