Benchmark Radar™
RSS Contact Star —

Daily AI benchmark brief: 2026-09-29

758evidence observations
12sources represented
10public-attention signals

Daily briefing

  1. TraceDance is a new release that turns real agent deployment traces into targeted benchmarks for user-specified undesirable behaviors. It evaluates the model’s next action from a recorded decision point, rather than treating successful task completion as sufficient evidence of acceptable behavior. Why it matters: If you operate agents in production, this offers a way to convert observed failures into regression tests tailored to your deployment. That changes suite selection from relying only on fixed public tasks to testing behaviors that users actually encountered. Evidence: E009. Medium confidence.
  2. Certified Selective Automation of LLM Agent Evaluation is a new method for determining what fraction of trajectory reviews an automatic judge can handle while keeping its error rate within a stated budget. It accounts for correlated results from multiple agents attempting the same tasks, where ordinary independent-sample assumptions can produce misleading guarantees. Why it matters: If you are deciding how much human review to retain, this reframes judge adoption as a measurable risk-allocation decision rather than an all-or-nothing replacement. The task-level treatment also matters when comparing many agents on a shared benchmark. Evidence: E016. Medium confidence.
  3. WebPageBench is a new web-agent benchmark that verifies task completion from typed interface event logs instead of model judges or page scraping. It can also rerender the same task with one user-interface control changed while preserving the prompt and success conditions. Why it matters: If you evaluate web agents, this separates whether an agent achieved the required state change from how a judge interpreted the screen. Its controlled interface variants also let you test whether performance survives presentation changes rather than reflecting familiarity with one layout. Evidence: E015. Medium confidence.
  4. INSPIRE is a new benchmark for agents that search scientific literature from an open research problem. It hides the later target paper and its cited antecedents, limits evidence to a cutoff three months before that paper, and evaluates ranked results against graded papers from the realized research lineage. Why it matters: If you are choosing an evaluation for research assistants, this tests prospective discovery under a historical information boundary rather than retrieval from a disclosed candidate set. That better distinguishes useful search behavior from finding a known paper under easier conditions. Evidence: E002. Medium confidence.
  5. mu-bench is a new multilingual speech-transcription benchmark built from 4,270 caller utterances to an artificial-intelligence banking agent across five languages. Alongside word error rate, it releases an Utterance Error Rate judged on whether a transcript preserves meaning and calibrated against human ratings. Why it matters: If you are selecting speech recognition for voice agents, exact word matching can penalize harmless formatting differences while missing the operational importance of names, email addresses, and confirmation codes. This artifact provides a decision measure tied more directly to whether the agent received the caller’s intended information. Evidence: E003. Medium confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

medium confidenceNot enough evidence

In this captured feed, notable first-seen releases included PowerBench, Inspire, mu-bench, EmailBench, Claw-SWE-Bench, WebPageBench, AnyAppBench, AnesTRACE, MixBench-TS, READ-Bench, VoiceNet, and MULTISPEECH-BENCH.

These arrivals cover power-system reasoning, scientific search, multilingual speech, enterprise email, software engineering, web and mobile agents, anesthesia decisions, forecasting, historical-case retrieval, and voice understanding. New evaluation methods include meaning-preservation judging, executable assertions, interface-event matching, graded citation relevance, and expert-defined clinical criteria.

Takeaway: The selected evidence shows broad experimentation with task-specific benchmarks and more explicit checking methods. This is a partial view of artifacts first observed by the keyword-filtered radar, not evidence of field-wide direction or proof that every listed resource is available and complete.

Another reading: First observation by the radar does not establish that an artifact is new to the field. The RealCLI repository is labeled as a release, but its description says the dataset is still forthcoming, showing that feed appearance can precede practical availability.

  • S023 artifacts first observed by the radar today: 447

Which of today's arrivals document how they score an answer?

medium confidenceNot enough evidence

The clearest documented examples are Inspire, mu-bench, EmailBench, WebPageBench, and AnesTRACE. They respectively describe graded antecedent relevance, a human-calibrated meaning judge, executable assertions, event-log matching, and clinician-defined response criteria.

These artifacts explain what counts as success rather than merely saying that models are evaluated. Some compare returned literature with graded references, some check whether a transcript keeps the intended meaning, some run fixed checks, some verify recorded interface actions, and some grade clinical responses against expert-written requirements.

Takeaway: For engineers seeking inspectable scoring, WebPageBench and EmailBench emphasize deterministic checks, while Inspire, mu-bench, and AnesTRACE rely on graded references, model judging, or expert criteria. The packet does not expose every first-seen artifact, so this cannot be an exhaustive list.

Another reading: A summary can name a scoring mechanism without providing enough detail to reproduce it. AnyAppBench, for example, mentions a vision-language judge and a fixed failure taxonomy but does not explain a complete answer-scoring rule in the supplied text.

  • S023 artifacts first observed by the radar today: 447

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

medium confidenceNot enough evidence

The registry contains cumulative movement entries for previously tracked artifacts, with each span attached to the linked statistics, but the packet provides no citable artifact evidence.

These measurements cover each artifact’s entire tracked period, not merely the latest day. Because no evidence record IDs accompany them, I cannot safely provide an artifact-level list or independently validate the reported changes.

Takeaway: Treat the linked statistics as a provisional movement table for this keyword-filtered feed, pending artifact evidence with usable citations.

Another reading: The registry itself may be considered enough to identify the artifacts and spans, but doing so would violate the requirement for evidence citations on artifact-specific claims.

  • S026 downloads change for lmarena-ai/leaderboard-dataset: 28,711.0 downloads
  • S027 downloads change for RoboDojo-Benchmark/RoboDojo: 25,517.0 downloads
  • S028 downloads change for hf-benchmarks/transformers: 23,158.0 downloads
  • S029 downloads change for vedangfake/chess-slm-benchmark: 15,992.0 downloads
  • S030 downloads change for alexshpunt/explicit-edit-benchmark: 13,382.0 downloads
  • S031 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 5,760.0 downloads
  • S032 stars change for career-ops-hq/career-ops: 3,365.0 stars
  • S033 downloads change for NoeFlandre/benchmark-llms-landuse-relevance: 2,821.0 downloads

Which of that movement is corroborated by more than one data source?

high confidenceNot enough evidence

No movement can be classified as corroborated from the supplied citable material.

The registry reports that some tracked artifacts were seen by multiple sources, but explicitly warns that this does not mean those sources measured the same change. No registered movement statistic here is marked as corroborated.

Takeaway: Multiple sightings are not metric corroboration; confirmation requires the same movement to be reported by more than one source.

Another reading: A competing reading is that the metadata-level corroborated flag is sufficient, but the associated movement lacks a registered statistic and citable evidence, so it cannot support a grounded answer.

  • S025 tracked artifacts today seen by more than one data source: 11
  • S026 downloads change for lmarena-ai/leaderboard-dataset: 28,711.0 downloads
  • S027 downloads change for RoboDojo-Benchmark/RoboDojo: 25,517.0 downloads
  • S028 downloads change for hf-benchmarks/transformers: 23,158.0 downloads
  • S029 downloads change for vedangfake/chess-slm-benchmark: 15,992.0 downloads
  • S030 downloads change for alexshpunt/explicit-edit-benchmark: 13,382.0 downloads
  • S031 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 5,760.0 downloads
  • S032 stars change for career-ops-hq/career-ops: 3,365.0 stars
  • S033 downloads change for NoeFlandre/benchmark-llms-landuse-relevance: 2,821.0 downloads

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Within this captured feed, newly released work emphasizes testing agent behavior during execution, cross-application transfer, event-based task verification, and correlated evaluation errors, rather than relying only on final-task success or a single interface.

Add tests that inspect how an agent reaches an answer, rerun equivalent goals across different applications, verify actions against recorded events, and check automatic graders on groups of related tasks. TraceDance, AnyAppBench, WebPageBench, and Certified Selective Automation each propose one of these approaches, but remain newly released artifacts rather than independently validated standards.

Takeaway: Use these releases as design prompts for an internal evaluation gap review, not as proof that their methods are production-ready. Benchmark and evaluation tags cover much of today’s feed, but tags overlap and the feed is keyword-filtered rather than representative.

Another reading: The strongest competing reading is that no immediate process change is justified: these are largely self-described releases, and the packet provides neither independent replications nor evidence that adopting their methods improves production outcomes. Existing domain-specific tests may already cover the same failure modes.

  • S003 records with event kind released: 447
  • S018 records tagged benchmark: 499 count (multi-label)
  • S019 records tagged evaluation: 454 count (multi-label)
  • S021 records tagged agentic: 117 count (multi-label)

What does today's evidence fail to show, and what would change the reading?

high confidence

The feed does not establish a field-wide trend, validated benchmark quality, or corroborated momentum for tracked artifacts. Today lacks a certified comparison window, while reported download and star movements come from single sources and span each artifact’s full tracking period rather than one day.

Differences from earlier collection windows may reflect what the radar captured, not changes in AI research. The new benchmark papers describe their own methods, but this packet does not include independent reproductions, audits, or downstream deployment results. Public attention observations are also distinct from releases and updates.

Takeaway: The reading would change with matched collection windows under the same taxonomy and connector coverage, independent validation of the released benchmarks, repeated results across systems and domains, and corroboration of each usage metric by more than one source.

Another reading: A competing reading is that the breadth of benchmark and evaluation records still makes today useful for discovering evaluation ideas. That supports exploration, but not claims about growth, adoption, quality, or field-wide direction.

  • S001 evidence records captured today: 758
  • S002 public attention observations captured today: 10
  • S003 records with event kind released: 447
  • S004 records with event kind updated: 213
  • S025 tracked artifacts today seen by more than one data source: 11
  • S026 downloads change for lmarena-ai/leaderboard-dataset: 28,711.0 downloads
  • S027 downloads change for RoboDojo-Benchmark/RoboDojo: 25,517.0 downloads
  • S028 downloads change for hf-benchmarks/transformers: 23,158.0 downloads
  • S029 downloads change for vedangfake/chess-slm-benchmark: 15,992.0 downloads
  • S030 downloads change for alexshpunt/explicit-edit-benchmark: 13,382.0 downloads
  • S031 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 5,760.0 downloads
  • S032 stars change for career-ops-hq/career-ops: 3,365.0 stars
  • S033 downloads change for NoeFlandre/benchmark-llms-landuse-relevance: 2,821.0 downloads
  • S034 daily-average change in evaluation observations: -56.86 observations per day
  • S035 daily-average change in benchmark observations: -29.86 observations per day
  • S036 daily-average change in dataset observations: -26.43 observations per day
  • S037 daily-average change in agentic observations: -3.86 observations per day
  • S038 daily-average change in data_quality observations: -3.57 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 758 evidence records.

Briefing model: gpt-5.6-sol.

It read 149 of 758 records.

This is a keyword-filtered feed, not a representative sample of artificial intelligence work. The briefing examined 149 selected evidence records out of 758 captured records, although none of those 149 were dropped for size. Brave and OpenReview were unavailable, and there was insufficient comparable history to infer category-share movement. The findings describe source-authored releases; the supplied evidence does not independently validate their datasets, implementations, or reported results.