Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Every benchmark the catalog has found, with the paper, code, dataset and scores on record for it. Collected from public sources and rebuilt every day, so a number here is never older than the last run.
Why it exists
I kept running into new benchmarks while doing benchmark research. There was no one place that listed them, said what each one tests, and showed which models had been measured on it. So I built a collector that reads public sources every day and writes down what it finds, with the evidence attached.
What it collects
Two things, kept apart. Every day the collector reads arXiv, GitHub, Hugging Face, OpenReview, Semantic Scholar, Hacker News, and the labs' own release feeds, and writes down the benchmark work that appeared that day. Separately, the scored catalog is built from registered leaderboards and model reports: OpenCompass Hub, LLM Stats, Artificial Analysis, and the evaluation tables labs publish with a model. Each benchmark in that catalog keeps its own page with the paper, code, and dataset links it has, plus every score on record.
The rule the whole project follows
A number is only worth as much as the document it came from. Scores stay partitioned by the source that reported them and are never merged into a single ranking, because those sources measure under different conditions and say so. Where a field is missing, the page leaves it out rather than filling it in.
Who makes it
Benchmark Radar is built by Koutian Wu and open to contributions. The code is MIT-licensed, the data is downloadable in full, and every page is generated from the same public corpus, so anyone can check a claim against its source.
Where to start
- Browse every benchmark in the catalog, one page each
- Compare reported scores across the benchmark frontier
- Watch scores on a benchmark climb toward saturation
- Read the daily brief: what appeared today and what it means
- Query the whole catalog offline from the command line
- Cite the technical report, or read the publications behind it
Benchmark Radar elsewhere
- Source code on GitHub
- Technical report
- Download the full dataset
- Build log: every day of this project, written up
- Benchmark Radar: an evidence-first daily radar for AI benchmarks
- Why hard benchmarks should not become coding tricks
- The value of benchmarks and the ability to frame questions
- Google Scholar