Benchmark Radar™
RSS Contact Star

Daily AI benchmark brief: 2026-09-22

453evidence observations
11sources represented
20public-attention signals

Daily briefing

  1. The new Curation-Bench evaluates coding agents as data curators: it fixes the model, training recipe, and test suite, then lets an agent inspect data, implement selection policies, run training, review noisy results, and revise its policy through a command line. Why it matters: If you are deciding whether an agent can replace part of a data-curation workflow, this design tests the iterative job rather than scoring a single proposed dataset. Fixing the model and training pipeline also makes differences more attributable to the agent’s curation policy. Evidence: E003. High confidence.
  2. The new XYEval framework converts existing agent benchmarks into “XY problem” tests, where a user proposes a plausible but misguided solution and the agent must recognize and explain the underlying problem instead of simply agreeing. Why it matters: If you evaluate assistants that advise users, ordinary task success can miss harmful compliance with a bad premise. XYEval adds a targeted comparison between baseline performance and performance under misleading user advice, changing model selection toward agents that diagnose requests rather than merely follow them. Evidence: E019. High confidence.
  3. The new VibeMemBench isolates whether persistent memory improves executable coding work, using 111 targets from 90 software repositories and 3,634 prior-work trajectories. It covers bug fixes, features, interface changes, and configuration tasks. Why it matters: If you are choosing a memory system for coding agents, recall scores alone do not establish that stored experience helps produce working code. This benchmark connects memory use to repository-level outcomes, allowing teams to compare memory implementations without treating retrieval accuracy as the final product metric. Evidence: E021. High confidence.
  4. The new BabelArena evaluates tool-using agents across languages with 16,146 instances derived from 702 canonical tasks. Its BabelFlow process translates runtime dependencies and task structure, then applies automated verification and human review to preserve the original evaluation meaning. Why it matters: If you deploy an agent outside English, directly translating prompts can inadvertently change tools, constraints, or grading. BabelArena offers task-aligned multilingual comparisons, making it possible to distinguish language-related failures from artifacts introduced by translation. Evidence: E022. High confidence.
  5. The new ChartJudgeBench tests multimodal models used as judges for chart-to-code systems, including 1,003 pairwise chart-comparison cases and 650 reasoning-judgment cases. Why it matters: If a visual model supplies rewards for training chart-generation systems, its errors can become training signals rather than merely evaluation errors. This benchmark lets teams assess the judge separately before using its scores for reinforcement learning, regression testing, or model selection. Evidence: E017. High confidence.
  6. Version 2.4 of the atomic-layer-deposition extraction pipeline is an update, not a new benchmark: it adds a second scorer that corrects defects involving temperature-unit conversion, chemical-name comparison, and other scoring behavior while retaining the earlier frozen scorer. Why it matters: If you compare extraction models or reproduce earlier results, the scorer version can change the measured outcome independently of the model. Publishing both implementations allows evaluators to preserve historical comparability while determining whether conclusions survive corrected scoring. Evidence: E078. Medium confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

medium confidenceNot enough evidence

The captured feed first observed many artifacts, including newly released benchmarks for video reasoning, automated data curation, gameplay, multimodal judging, agent communication, coding-agent memory, multilingual agents, and form-field detection.

Notable examples are AgentVidBench, Curation-Bench, GameHorizon Suite, ChartJudgeBench, XYEval, VibeMemBench, BabelArena, and mini-CommonForms. Their evaluations cover multi-step video questions, data-selection work, gameplay over different time spans, chart-code judging, resistance to misleading advice, useful agent memory, multilingual workflows, and document-form detection.

Takeaway: Today’s captured arrivals span both specialized datasets and methods for testing agents or automated judges. This is a catalog of this keyword-filtered feed, not evidence of a broader field trend.

Another reading: The supplied evidence is only a selected subset of artifacts first observed today, so this list is not exhaustive. “First observed” means new to the radar, not necessarily newly created, and some records describe benchmarks without establishing that all underlying assets are available.

  • S022 artifacts first observed by the radar today: 314

Which of today's arrivals document how they score an answer?

medium confidenceNot enough evidence

The clearest cases are DISCERN, DualLoop Evaluation, the pancreatic-cancer answer study, the updated ALD extraction pipeline, and Measuring the Checker. They expose grading artifacts, answer oracles, human-rating procedures, scorer code, or an explicit detection-based score.

DISCERN includes deterministic grading outputs, adjudication records, and regeneration scripts. DualLoop derives gold answers from a pixel oracle. The clinical study uses blinded evaluators and several quality dimensions. The ALD update releases frozen and corrected scoring implementations. Measuring the Checker scores test protocols by the share of injected faults they detect.

Takeaway: These arrivals provide more scoring transparency than records that merely report a leaderboard or claim evaluation. Their event types differ: some are releases, the ALD artifact is an update, and Measuring the Checker is a radar discovery.

Another reading: The packet contains summaries rather than complete methods, so it cannot establish whether every scorer is valid, reproducible, or appropriate. Other arrivals may document scoring in full text but are not identifiable from the selected evidence supplied here.

  • S022 artifacts first observed by the radar today: 314

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

low confidenceNot enough evidence

The registry contains cumulative download-movement entries for several previously tracked artifacts, with spans running from their first tracked observation through today.

A compliant artifact-by-artifact list cannot be provided because the packet supplies no artifact-level evidence identifiers for the named records.

Takeaway: The cited statistics contain the exact artifact names, cumulative download changes, and full tracked spans; these are not one-day changes and apply only to this captured feed.

Another reading: The registry itself identifies the artifacts and movements, so it may be operationally adequate, but it does not satisfy the required artifact-level evidence citation rule.

  • S025 downloads change for hf-benchmarks/transformers: 22,812.0 downloads
  • S026 downloads change for IntelligenceLab/LHTB-leaderboard: -15,587.0 downloads
  • S027 downloads change for vedangfake/chess-slm-benchmark: 7,929.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 7,668.0 downloads
  • S029 downloads change for vava22684/song-jury-leaderboard: -3,936.0 downloads
  • S030 downloads change for AlphaDojo/dojo_benchmark_kline: 1,984.0 downloads
  • S031 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 1,270.0 downloads
  • S032 downloads change for brettsp/stan-benchmark: 1,203.0 downloads

Which of that movement is corroborated by more than one data source?

high confidence

None of the registered download movements is corroborated by more than one data source in this captured feed.

Every cited download change came from Hugging Face alone, so another source did not independently confirm the same metric movement.

Takeaway: Treat these as single-source cumulative measurements across each artifact’s full tracked span, not independently confirmed movement or one-day change.

Another reading: Some tracked artifacts were sighted by more than one source, but independent sighting of an artifact does not mean that multiple sources measured the same changing metric.

  • S024 tracked artifacts today seen by more than one data source: 5
  • S025 downloads change for hf-benchmarks/transformers: 22,812.0 downloads
  • S026 downloads change for IntelligenceLab/LHTB-leaderboard: -15,587.0 downloads
  • S027 downloads change for vedangfake/chess-slm-benchmark: 7,929.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 7,668.0 downloads
  • S029 downloads change for vava22684/song-jury-leaderboard: -3,936.0 downloads
  • S030 downloads change for AlphaDojo/dojo_benchmark_kline: 1,984.0 downloads
  • S031 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 1,270.0 downloads
  • S032 downloads change for brettsp/stan-benchmark: 1,203.0 downloads

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Change the evaluation process, not the production stack: add task-specific stress tests and verify them locally before acting on headline results.

In this captured feed, newly released tests target agents accepting misleading advice, models used to score other models, and errors caused by noisy form labels. Teams facing those risks should add comparable checks to their own test sets, with human review and production-like data.

Takeaway: Use today’s releases as candidates for closing evaluation blind spots, especially around agent behavior and evaluator reliability. Do not treat release status, download activity, or public attention as proof that a test is valid or that a system should be replaced.

Another reading: These artifacts are supported mainly by their authors’ descriptions, not independent replication. Their tasks may not match a given product, so adding them indiscriminately could increase testing cost without improving decisions.

  • S017 records tagged benchmark: 312 count (multi-label)
  • S018 records tagged evaluation: 230 count (multi-label)
  • S020 records tagged agentic: 68 count (multi-label)
  • S022 artifacts first observed by the radar today: 314

What does today's evidence fail to show, and what would change the reading?

high confidenceNot enough evidence

The captured feed does not show a field-wide trend, independently verified benchmark quality, or production improvement.

There is no certified comparison window, some connectors were unavailable, and few captured artifacts had sightings from more than one source. Differences from other days could therefore reflect collection changes. The feed also lacks independent checks showing that highlighted tests predict real-world outcomes.

Takeaway: A stronger reading requires stable connector coverage and taxonomy, a certified comparison window, independent replication, accessible test data and methods, and head-to-head results on production-like tasks. Until then, use the feed for discovery rather than market or quality conclusions.

Another reading: The volume of captured records and contributions from several sources can still justify exploratory review. That breadth, however, does not overcome the missing comparability and validation needed for directional or quality claims.

  • S001 evidence records captured today: 453
  • S002 public attention observations captured today: 20
  • S006 records contributed by Hugging Face: 138
  • S007 records contributed by arXiv: 91
  • S008 records contributed by Crossref: 86
  • S009 records contributed by GitHub: 45
  • S010 records contributed by Zenodo: 30
  • S011 records contributed by Hugging Face Papers: 29
  • S012 records contributed by Kaggle Dataset: 18
  • S013 records contributed by OpenAlex: 9
  • S014 records contributed by First-party feed: 4
  • S015 records contributed by GitHub Organization: 2
  • S016 records contributed by GitHub Release: 1
  • S024 tracked artifacts today seen by more than one data source: 5

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 453 evidence records.

Briefing model: gpt-5.6-sol.

It read 133 of 453 records.

This is a keyword-filtered feed, not a representative sample of artificial intelligence work. Only 133 of 453 captured evidence records were supplied for analysis, although none of those 133 were dropped for size. Three connectors—Brave, OpenReview, and Semantic Scholar—were unavailable, and comparable history was insufficient to assess category-share movement.