Today's radar
Benchmarks with this name
Daily briefing
Questions for today
Most reported benchmarks in model cards
i
When an AI lab releases a model, it publishes a report listing the tests it ran. This counts how many of those reports mention each test. A test near the top is one almost everyone runs, which is not the same as a good one: labs keep running a popular test out of habit, even after the scores stop telling anyone much. A report counts once per test, even if it lists that test several times. Some reports publish their results as a picture rather than text, and we read those with software that can misread a digit, so the list at the bottom of this page links every count back to the report it came from.
When an AI lab releases a model, it publishes a report listing the tests it ran. This counts how many of those reports mention each test. A test near the top is one almost everyone runs, which is not the same as a good one: labs keep running a popular test out of habit, even after the scores stop telling anyone much.
- 01GPQA Diamond27 model cards
- 02Humanity's Last Exam20 model cards
- 03Terminal-Bench19 model cards
- 04SWE-bench Verified18 model cards
- 05AIME17 model cards
Scores over time
Benchmark reported scores over time
How to read this chart
Every value that could be read verbatim from a cited document, placed at the date that document was published rather than at any evaluation date. The line directly links actual observations that set a new reported record among the values shown; it neither holds a score between reports nor extends past the final record. Test versions and run conditions can differ, so this is a reported-record path, not a like-for-like trend. When no newer number could be read, the gap is marked rather than drawn through. Whether a benchmark has saturated stays a reading you make, not a score this panel prints.
What the two layers say Stated findings
Benchmarks by model card adoption
Each model card counts once per benchmark. A card reporting AIME in four configurations counts the same as a card reporting it once, so a long appendix cannot outweigh a different vendor. Organizations breaks the tie: the same count from six vendors is a shared standard, from one vendor a house style.
Audit the counts Model cards in the registry
The curated source list this ranking is computed from. Expand any card to see every benchmark it reports, grouped the way the source document groups them, so our data can be checked line by line against the original.
Big picture
What we found
See what shows up most often and where it came from. Open the connections view when you want to look at a specific item.
Want more detail? See how everything connects
There are a lot of dots here. Pick one to see what it connects to, or open the matching results.
Recent activity
Signals over time
Counts describe discovery volume, not scientific quality.
New by domain
Daily evidence and attention volume
Category tags overlap. Each bar is an independent count, not a part of a stacked total.
Dev checker
Source mix counts ranked evidence after scoring. Fetch health counts raw records returned before scoring, so a source can be ok and still empty.
| Date | Coverage (UTC) | Evidence | Source mix | Categories | Events | Attention | Fetch health |
|---|
Dashboard unavailable
The validated data file could not be loaded.
Try refreshing, or inspect the latest daily Issue while the dashboard rebuilds.
Open daily Issues ↗