Benchmark

MRCR v2 (8-needle)

MRCR v2 (8-needle) is a variant of the Multi-Round Coreference Resolution benchmark that includes 8 needle items to retrieve from long contexts. This…

Modality
text
Categories
long_context, reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
23

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4.6Anthropic0.762026-02-05evidence
Gemini 2.5 Pro Preview 06-05Google0.1642025-06-05evidence
Gemini 3 FlashGoogle0.2212025-12-17evidence
Gemini 3 ProGoogle0.2632025-11-18evidence
Gemini 3.1 Flash-LiteGoogle0.6012026-03-03evidence
Gemini 3.1 ProGoogle0.2632026-02-19evidence
Gemini 3.5 FlashGoogle0.2662026-05-19evidence
Gemini 3.5 Flash-LiteGoogle0.2132026-07-21evidence
Gemini 3.6 FlashGoogle0.542026-07-21evidence
Gemini 3.7 FlashGoogle0.972026-08-13evidence
Gemma 3 27BGoogle0.1352025-03-12evidence
Gemma 4 12BGoogle0.4342026-05-23evidence
Gemma 4 26B-A4BGoogle0.4412026-04-02evidence
Gemma 4 31BGoogle0.6642026-04-02evidence
Gemma 4 E2BGoogle0.1912026-04-02evidence
Gemma 4 E4BGoogle0.2542026-04-02evidence
GPT-5.4 miniOpenAI0.3362026-03-17evidence
GPT-5.4 nanoOpenAI0.3312026-03-17evidence
GPT-5.5OpenAI0.742026-04-23evidence
GPT-5.6 LunaOpenAI0.4132026-07-09evidence
GPT-5.6 SolOpenAI0.9152026-07-09evidence
GPT-5.6 TerraOpenAI0.8962026-07-09evidence
Qwen3.8 MaxQwen0.9292026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.