Benchmark Radar™
RSS Contact Star

Daily AI benchmark brief: 2026-09-15

560evidence observations
11sources represented
15public-attention signals

Daily briefing

  1. The new Sophea release evaluates a production Greek-English speech recognizer against nine simultaneous gates covering both languages, language identification, and hallucinations on non-speech audio. Across 23 training iterations, no training-data mix passed every gate; additional Greek noisy-speech exposure conflicted with preserving English identification. [E001] Why it matters: If you are setting release criteria for a multilingual speech product, this artifact supports evaluating all deployment gates jointly rather than selecting a checkpoint from average accuracy. It also makes training-data composition a release decision because improving one operating condition can violate another gate. [E001] Evidence: E001. High confidence.
  2. The newly released IWC-Bench tests generated web applications at runtime and instruments them with code coverage, which records which parts of an application actually execute. It addresses two scoring errors: static checks can credit unreachable features, while incomplete interaction can mistake an evaluator agent’s exploration failure for an application defect. [E008] Why it matters: If you are choosing a benchmark for web-application generation, this design lets you ask whether functionality works and whether the evaluator exercised it. That separation can change model rankings and debugging priorities compared with source-code inspection or unguided interaction alone. [E008] Evidence: E008. High confidence.
  3. The new receipt-based audit of document-question-answering agents moves evidence into harder-to-find conditions and scores statement-level provenance—an explicit source for each claim—rather than relying only on final-answer accuracy or confidence. The reported audit found that buried evidence reduced accuracy while increasing tool use and cost per correct answer. [E009] Why it matters: If you evaluate agents for financial review or other document-heavy work, this artifact offers a way to distinguish an answer containing supported facts from one that mixes supported and fabricated claims. It also makes evidence placement and cost part of the test rather than hidden properties of the task. [E009] Evidence: E009. High confidence.
  4. The newly released MODA General Attribute Suite evaluates fashion-attribute extraction in four separate tracks for garment crops, catalogue images, full-body photographs, and product text. Each track has its own fixed test set, input rules, metric, and leakage boundary, and the protocol does not average the tracks together. [E014] Why it matters: If you are comparing models for fashion search or cataloguing, this design prevents one aggregate score from concealing whether a model handles visible garments, inapplicable attributes, text inputs, or vocabulary mismatches. It supports selecting a model for the exact input regime the product will encounter. [E014] Evidence: E014. High confidence.
  5. A new thyroid-ultrasound benchmark groups exact and visually near-duplicate images before partitioning and evaluates frozen models separately against malignancy outcomes and ultrasound risk categories. This separates duplicate leakage from a second problem: treating agreement with a risk label as equivalent to predicting clinical malignancy. [E003] Why it matters: If you evaluate medical-imaging models across cohorts, this artifact shows that both duplicate control and endpoint definition affect what a score means. A model-selection decision can therefore depend on whether the intended use is outcome prediction or agreement with an imaging classification system. [E003] Evidence: E003. High confidence.
  6. The new V-ICAL Bench evaluates whether multimodal agents can learn executable behavior from demonstration videos, using 342 interactive tasks across 37 environments. The agent must transfer what it observes into actions in a new visual state and adjust those actions using environmental feedback. [E017] Why it matters: If you are choosing an evaluation suite for agents that learn from demonstrations, this benchmark tests more than video question answering: it measures whether observed behavior becomes a working policy in an interactive environment. That makes it relevant to product decisions involving tutorial-following or example-based task adaptation. [E017] Evidence: E017. High confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

high confidence

In this captured feed, notable first observations included IWC-Bench for generated web applications, V-ICAL for video-guided agents, VisInteract-Bench for imperfect visualization requests, MTAC-IFBench for coding agents, and domain-focused resources such as SALUTE, LegalRewardBench, MODA, MUSE-Bench, and BenchECG.

The arrivals test systems on practical tasks including building usable websites, learning actions from video, clarifying unclear requests, following coding instructions, answering legal questions with evidence, understanding fashion attributes, forecasting with mixed information, and interpreting heart recordings.

Takeaway: The first-observed pool was substantial. Benchmark, evaluation, and dataset tags were all common, but these labels overlap and must not be treated as separate portions of the feed. The packet also surfaced evaluation approaches based on duplicate control, small representative samples, and evidence receipts.

Another reading: First observed by this radar does not necessarily mean newly created or newly released. Some records were discoveries of earlier work, while other purported benchmark repositories explicitly describe themselves as preparation pipelines or incomplete releases, limiting confidence that every arrival is operationally ready.

  • S022 artifacts first observed by the radar today: 338
  • S017 records tagged benchmark: 363 count (multi-label)
  • S018 records tagged evaluation: 309 count (multi-label)
  • S019 records tagged dataset: 228 count (multi-label)

Which of today's arrivals document how they score an answer?

medium confidenceNot enough evidence

Clear examples in the supplied packet are the math-response dataset, which uses a structured rubric with written justification; the Lean evaluator benchmark, which uses the real compiler’s verdict; MODA, which assigns each track its own metric and forbids averaging across tracks; and MakerBench, which says grading is mathematical rather than delegated to a language-model judge.

These artifacts expose at least part of the decision rule instead of presenting only a leaderboard result. They respectively rely on human rubric dimensions, whether code compiles, separate task-specific measurements, or mechanically checkable mathematics.

Takeaway: These are the strongest scoring-documentation signals among the supplied arrivals. The receipt-based agent audit also proposes scoring individual claims against cited evidence and accounting for test conditions, but the packet does not provide enough detail to verify every formula, weight, or threshold.

Another reading: The packet contains summaries rather than complete protocols. MODA says each track has its own metric without giving the formulas here, and MakerBench only states that grading is mathematical. A complete arrival set and full artifact documentation are missing, so this is not an exhaustive or reproducibility-confirmed list.

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

low confidenceNot enough evidence

The registry reports download gains for huggingface-projects/drlc-leaderboard-data, hf-benchmarks/transformers, Weyaxi/huggingface-leaderboard, lmarena-ai/leaderboard-dataset, hf-audio/open-asr-leaderboard-results, and latency-sensitive-bench/benchmark-datasets, with declines for RoboDojo-Benchmark/RoboDojo and vava22684/song-jury-leaderboard.

These are cumulative download-counter changes across each artifact’s full tracked span, from the start date through the capture date shown in its cited statistic. They are not one-day changes, releases, or updates.

Takeaway: Within this captured feed, the listed artifacts had the largest measurable movements supported by registered statistics. The exact span differs by artifact and is provided in each cited statistic, ranging from late July, early August, mid-August, or early September through the capture date.

Another reading: The underlying artifact evidence records were not supplied with usable evidence identifiers. These single-source counters may also reflect counter corrections or collection behavior rather than changes in actual use, particularly where downloads declined.

  • S025 downloads change for RoboDojo-Benchmark/RoboDojo: -33,671.0 downloads
  • S026 downloads change for huggingface-projects/drlc-leaderboard-data: 32,780.0 downloads
  • S027 downloads change for hf-benchmarks/transformers: 15,236.0 downloads
  • S028 downloads change for Weyaxi/huggingface-leaderboard: 9,383.0 downloads
  • S029 downloads change for lmarena-ai/leaderboard-dataset: 4,244.0 downloads
  • S030 downloads change for vava22684/song-jury-leaderboard: -3,735.0 downloads
  • S031 downloads change for hf-audio/open-asr-leaderboard-results: 1,165.0 downloads
  • S032 downloads change for latency-sensitive-bench/benchmark-datasets: 1,098.0 downloads

Which of that movement is corroborated by more than one data source?

high confidence

None of the registered download movements is corroborated by more than one data source; every cited movement was measured only through Hugging Face.

An artifact appearing in several sources does not mean several sources measured the same download change. The registry explicitly marks each listed movement as uncorroborated.

Takeaway: Treat all listed movements as single-source observations within this keyword-filtered feed. Multi-source sightings elsewhere in today’s tracked set do not validate these particular metric changes.

Another reading: Corroborating measurements may exist outside the captured feed or under metrics not included in the registry, but the supplied material cannot establish that.

  • S024 tracked artifacts today seen by more than one data source: 11
  • S025 downloads change for RoboDojo-Benchmark/RoboDojo: -33,671.0 downloads
  • S026 downloads change for huggingface-projects/drlc-leaderboard-data: 32,780.0 downloads
  • S027 downloads change for hf-benchmarks/transformers: 15,236.0 downloads
  • S028 downloads change for Weyaxi/huggingface-leaderboard: 9,383.0 downloads
  • S029 downloads change for lmarena-ai/leaderboard-dataset: 4,244.0 downloads
  • S030 downloads change for vava22684/song-jury-leaderboard: -3,735.0 downloads
  • S031 downloads change for hf-audio/open-asr-leaderboard-results: 1,165.0 downloads
  • S032 downloads change for latency-sensitive-bench/benchmark-datasets: 1,098.0 downloads

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Make evaluation more deployment-like: use separate production gates, duplicate-controlled splits, runtime testing, difficult evidence placement, claim-level source receipts, and explicit abstention tests.

Today's releases show why one overall score can hide operational failures. The speech-recognition study used distinct production gates; IWC-Bench tests reachable behavior; the agentic audit checks whether claims trace to evidence; and LegalRewardBench tests grounding when retrieval is noisy or insufficient.

Takeaway: Add these checks to internal evaluations before changing models or shipping systems. Treat the feed's rising benchmark, evaluation, and dataset observations as prompts to inspect methods, not as proof that any released approach is better.

Another reading: These are newly released studies in a keyword-filtered feed, and their summaries do not establish independent replication or applicability to another deployment. Existing evaluations may already cover the relevant risks.

  • S033 daily-average change in benchmark observations: 60.0 observations per day
  • S034 daily-average change in evaluation observations: 55.43 observations per day
  • S035 daily-average change in dataset observations: 39.57 observations per day

What does today's evidence fail to show, and what would change the reading?

high confidenceNot enough evidence

The feed does not show a broadly superior model, benchmark, or evaluation method, nor does it establish representative adoption or independently verified performance movement.

Release and update volume shows publication activity, not quality or practical impact. Few tracked artifacts were seen across multiple sources, and a cross-source sighting does not mean the same performance measure was independently confirmed. Public attention is also too limited to establish adoption.

Takeaway: The reading would change with direct comparative results on relevant workloads, independent replications, corroboration of the same measures across sources, persistent composition shifts, and evidence from a more representative collection.

Another reading: A competing reading is that rising benchmark, evaluation, and dataset observations across comparable recent and prior windows reflect meaningful activity. Even so, observation volume alone cannot identify winners or justify deployment choices.

  • S001 evidence records captured today: 560
  • S002 public attention observations captured today: 15
  • S003 records with event kind released: 313
  • S004 records with event kind updated: 189
  • S005 records with event kind discovered: 58
  • S024 tracked artifacts today seen by more than one data source: 11
  • S033 daily-average change in benchmark observations: 60.0 observations per day
  • S034 daily-average change in evaluation observations: 55.43 observations per day
  • S035 daily-average change in dataset observations: 39.57 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 560 evidence records.

Briefing model: gpt-5.6-sol.

It read 154 of 560 records.

This is a keyword-filtered feed, not a representative sample of artificial intelligence work. The briefing received 154 selected evidence records from a 560-record corpus, so it did not review the full captured feed. Brave and OpenReview were unavailable, and the findings rely on source-authored release descriptions rather than independent reproduction.