Benchmark Radar
RSS

Today's radar

Daily briefing

Questions for today

Scoring rubric v5 · current

How priority is scored

Priority is the weighted mean of four components, each measured on a 0 to 100.00 scale. Every number below is read from the same definition the pipeline applies.

0.35 relevance + 0.20 evidence + 0.20 recency + 0.25 adoption

Relevance

weight 0.35

How squarely the title and the source's own description land inside the benchmark, evaluation, dataset, and data-quality taxonomy, after explicit low-value deductions.

  • 20 per taxonomy category matched
  • +5 per matched term, counting up to 2 terms per category
  • -50 for follower-count leaderboard
  • -30 for results dump or per-run index; suppressed
  • -30 for submission placeholder; suppressed
  • -25 for visualization-only companion; suppressed
  • -30 for sponsor-bait resource listing; suppressed
  • -30 for missing primary source URL: no primary source URL; suppressed
  • -15 for title-only provenance: no source description, attribution, artifact link, or public counter
  • Total deductions capped at 60
  • Clamped to 0-100

Evidence

weight 0.20

How directly the record is attested: a primary or structured record outranks a secondary mention, and named authors and linked artifacts add corroboration.

  • 10 baseline for any record that passed ingest
  • +40 from a primary or structured record (arXiv, First-party feed, OpenAlex, OpenReview, Semantic Scholar)
  • +30 from a structured artifact registry (GitHub, GitHub Release, GitHub Organization, Hugging Face, Kaggle Dataset, Zenodo)
  • +20 when the source names authors
  • +20 when another artifact URL corroborates it
  • Capped at 100

Recency

weight 0.20

How recently the artifact was published or materially updated within this scan's configured collection window.

  • 100 at release or first-discovery time
  • Age-based credit decays linearly across the configured 48-hour lookback
  • Age-based credit reaches 0 at 48 hours
  • released events retain 100% of their age-based recency
  • discovered events retain 100% of their age-based recency
  • prereleased events retain 75% of their age-based recency
  • updated events retain 50% of their age-based recency

Adoption

weight 0.25

The strongest available public uptake counter on a log scale. This measures attention, not scientific quality.

  • stars reaches 100 at 10000
  • citations reaches 100 at 1000
  • downloads reaches 100 at 100000
  • likes reaches 100 at 1000
  • downloads is capped at 60, the score of roughly 250 stars, because the counter accumulates without a human decision
  • likes is capped at 60, the score of roughly 250 stars, because the counter accumulates without a human decision
  • Uses the strongest available normalized counter
  • Clamped to 0-100

What this score does not claim

  • This is triage for a reader deciding what to open next. It is not peer review, a quality verdict, or an endorsement.
  • Relevance reads only the title and source-published description. Nothing this project writes about a record can earn it points.
  • Negative signals demote only named artifact patterns. Result indexes, submission placeholders, and visualization-only companions are suppressed; the selection funnel reports how many were removed.
  • Adoption measures attention, not correctness. Counters that accumulate without a human decision (downloads, likes) are capped below the top of the adoption scale; stars and citations are not.
  • This project's own repository (ktwu01/benchmark-radar) is excluded from its own ranking.
  • Attention observations are shown separately and are never quality-scored.
  • Watchlisted artifacts are retained whatever they score and sort first. Their rank reflects that request, not a higher score.
  • Update-driven recency is discounted from first announcements using the source's event kind. This cannot distinguish a material GitHub release from a packaging bump.
  • Structural checks use only missing or populated record fields as a provenance signal. They do not judge benchmark quality, novelty, or validity.

Every record matching at least one taxonomy category is retained. A score of 40.00 or above marks the item as recommended; it does not control inclusion. Watchlisted artifacts are also retained.