Benchmark

CritPT

CritPT is a challenging reasoning benchmark reported by Qwen for evaluating frontier mathematical and critical problem-solving capability.

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5.2Z.ai0.1672026-06-16evidence
Inkling-SmallThinking Machines Lab0.0832026-07-30evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.0312026-06-04evidence
Qwen3.7 MaxQwen0.1142026-05-19evidence
Qwen3.7-PlusQwen0.062026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.