Benchmark

CorpusQA 1M

CorpusQA 1M is a long-context question answering benchmark designed to evaluate models at approximately 1 million token contexts. Models are scored on…

Modality
text
Categories
long_context, reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V4-Flash-0423DeepSeek0.5932026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.6052026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.622026-04-23evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.