Benchmark Radar
RSS

Most reported benchmarks in model cards

i

When an AI lab releases a model, it publishes a report listing the tests it ran. This counts how many of those reports mention each test. A test near the top is one almost everyone runs, which is not the same as a good one: labs keep running a popular test out of habit, even after the scores stop telling anyone much. A report counts once per test, even if it lists that test several times. Some reports publish their results as a picture rather than text, and we read those with software that can misread a digit, so the list at the bottom of this page links every count back to the report it came from.

When an AI lab releases a model, it publishes a report listing the tests it ran. This counts how many of those reports mention each test. A test near the top is one almost everyone runs, which is not the same as a good one: labs keep running a popular test out of habit, even after the scores stop telling anyone much.

  1. 01GPQA Diamond27 model cards
  2. 02Humanity's Last Exam20 model cards
  3. 03Terminal-Bench19 model cards
  4. 04SWE-bench Verified18 model cards
  5. 05AIME17 model cards

Scores over time

Benchmark reported scores over time

How to read this chart

Every value that could be read verbatim from a cited document, placed at the date that document was published rather than at any evaluation date. The line directly links actual observations that set a new reported record among the values shown; it neither holds a score between reports nor extends past the final record. Test versions and run conditions can differ, so this is a reported-record path, not a like-for-like trend. When no newer number could be read, the gap is marked rather than drawn through. Whether a benchmark has saturated stays a reading you make, not a score this panel prints.

Benchmarks by model card adoption

Each model card counts once per benchmark. A card reporting AIME in four configurations counts the same as a card reporting it once, so a long appendix cannot outweigh a different vendor. Organizations breaks the tie: the same count from six vendors is a shared standard, from one vendor a house style.

Audit the counts Model cards in the registry

The curated source list this ranking is computed from. Expand any card to see every benchmark it reports, grouped the way the source document groups them, so our data can be checked line by line against the original.