Benchmark Radar™
RSS Star —

Leaderboard

Benchmark Frontier

Which difficult benchmarks have been tested most

Each mark is one benchmark record. Height counts distinct models with numeric scores. Source model IDs preserve evaluated configurations; repeated score rows add no models. Source documents counts distinct cited reports or registry pages through the same evidence collection. Both counts appear on hover.

The side view projects reported scores and the selected counts onto the left wall, leaving out time. The gold staircase compares only benchmarks with verified score scales. Within the dated 2024+ cohort, a benchmark is Pareto when no other eligible benchmark has an equal or lower normalized score and an equal or greater selected count, with at least one strict advantage. Dates select the cohort; they do not determine dominance. Scores and counts are never multiplied. The score slice cannot promote a dominated benchmark onto the frontier.

Height uses log1p(count); ticks, tooltips and Pareto use raw counts. The axis covers the full scored 2024+ cohort and stays fixed while filtering. Only declared percentage metrics with a known selected count enter Pareto; lower-is-better percentages become 100 minus the original value. Hollow points flag an unverified score scale or an unknown count, which is never treated as zero. Overlapping caps are spaced apart; hover or focus traces their exact coordinates. Arrow keys move through benchmarks. A shared scale does not establish equivalent test protocols or prove a benchmark is solved.

The timeline starts on January 1, 2024, inclusive. Dates use the benchmark's release first, then its earliest numeric LLM score report. When a source dates score entries by model release, the earliest entry supplies a clearly labelled model-release proxy; this is not a verified score-publication date. Crawl timestamps and adoption-only mentions never supply dates. Known dates before 2024 are excluded. Records with no date remain individually visible in a labelled area.

  1. 01577 data points
  2. 02577 data points
  3. 03510 data points
  4. 04492 data points
  5. 05489 data points

Most documented benchmarks

Counts distinct source documents that record each benchmark. Model reports and registry pages use the same rule. Each document counts once per benchmark record. This measures documentation coverage, not benchmark quality.

  1. 01GPQA Diamond29 source documents
  2. 02Humanity's Last Exam22 source documents
  3. 03Terminal-Bench21 source documents
  4. 04SWE-bench Verified19 source documents
  5. 05AIME18 source documents
i

Counts distinct source documents that record each benchmark. Model reports and registry pages use the same rule. Each document counts once per benchmark record. This measures documentation coverage, not benchmark quality. Open the source-document list below to trace each count to its citations.

Benchmarks by source documents

Each source document counts once per benchmark record. Repeated scores or mentions in the same document add no citations. This table covers all sources.

Audit the counts Source documents in the catalog

Source documents from every registry. Expand a document to inspect the benchmarks it records and open the original evidence.