Benchmark Radar™
RSS Contact Star —

Daily AI benchmark brief: 2026-09-28

505evidence observations
11sources represented
10public-attention signals

Daily briefing

  1. The new Finance Agents Benchmark packages financial due diligence as an end-to-end agent test: 50 tasks, 160 documents, 231 grading criteria, an execution harness, and a shared synthetic company data room. A companion release provides 600 complete runs with traces, answers, criterion-level verdicts, and judge explanations across four models and three trials per task. [E001, E005] Why it matters: If you are choosing an agent for document-heavy financial investigation, this release lets you compare not only final answers but also how agents searched, reasoned, and failed against explicit criteria. That supports debugging and model selection at the workflow level rather than through isolated question answering. [E001, E005] Evidence: E001, E005. High confidence.
  2. A new study of two-stage recommenders finds that fixed-candidate evaluation—giving every ranking model the same retrieved items—selected a different top pipeline from end-to-end evaluation in eight of ten benchmark instances. End-to-end evaluation instead tests each ranker on the candidates produced by its own retrieval stage. [E032] Why it matters: If you are selecting a complete recommendation pipeline whose retrieval components differ, a shared candidate set can reverse the selection decision by hiding whether useful items survived retrieval. The evaluation protocol therefore needs to match whether the product decision concerns a ranker alone or the full retrieval-and-ranking system. [E032] Evidence: E032. Medium confidence.
  3. A new audit reports cross-interaction leakage in three ephemeral group-recommendation benchmarks: held-out group-item labels were already present in every corresponding member’s individual interaction history, representing up to 23.4% of individual interactions. [E036] Why it matters: If you evaluate recommendations for newly formed groups, random splits on these benchmarks may test recall of information already exposed to the model rather than inference for a genuinely new group. This changes the decision about split construction and whether previously reported scores remain comparable after leakage is removed. [E036] Evidence: E036. Medium confidence.
  4. The new KNOWS benchmark tests browser agents on workflows that end in usable documents, presentations, or spreadsheets, rather than stopping after web search. Its tasks combine retrieval, synthesis, task decomposition, interface navigation, and production of a final artifact. [E014] Why it matters: If you are choosing an agent to perform office-style research, a search-answer benchmark does not test whether the system can organize findings and deliver them in the required application and format. KNOWS provides a closer test of the complete assistant workflow and exposes failures after information retrieval. [E014] Evidence: E014. Medium confidence.
  5. Efficient Safety Benchmarking via Item Response Theory proposes selecting safety questions by their estimated difficulty and informativeness instead of treating every item as equally useful. The study analyzes six safety benchmarks and reports that full-suite testing can require about 100,000 responses, many contributing little ranking information. [E031] Why it matters: If repeated safety testing is limited by inference cost, this design offers a basis for smaller adaptive test sets while retaining model-ranking signal. Item Response Theory is a statistical method that jointly estimates a model’s ability and each question’s difficulty, so the decision becomes which items distinguish the models under comparison rather than whether every model must answer every item. [E031] Evidence: E031. Medium confidence.
  6. The new MVVBench makes every question impossible to answer from any single designated camera view, while most questions also require evidence from more than one moment. It therefore isolates reasoning across cameras and time instead of allowing a model to succeed from one revealing frame. [E011] Why it matters: If you are selecting a vision-language model for multi-camera monitoring or navigation, ordinary video scores may not reveal whether it can connect an entity or event across viewpoints. MVVBench changes the test from recognizing visible content to reconstructing continuity across cameras and time. [E011] Evidence: E011. Medium confidence.
  7. The new ANT-A Benchmark preserves all five human annotations for each of 10,500 samples across seven text and multimodal tasks instead of collapsing them into a majority label. This retains cases where annotators disagree and supports evaluation against a distribution of judgments, sometimes called a soft label. [E018] Why it matters: If your application contains subjective or ambiguous classifications, a single majority answer can make reasonable alternatives appear wrong. ANT-A supports decisions about whether a model reflects the range of human judgments, is calibrated to disagreement, or merely matches the most common label. [E018] Evidence: E018. Medium confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

medium confidenceNot enough evidence

The captured feed first observed many artifacts today. Notable releases include finance-agent due-diligence tasks, agent-handoff decision data, multi-view video reasoning, ultrasound evidence grounding, browser-agent artifact production, clinical urgency assessment, safety evaluation using item-response modeling, and bias-aware language-model judging methods. [E001, E003, E011, E012, E014, E017, E031, E051]

The arrivals cover both new test material and new ways to evaluate systems. Examples range from financial research agents and medical urgency decisions to video reasoning, ultrasound grounding, web assistants, efficient safety testing, and methods that adjust for systematic judge bias. [E001, E011, E012, E014, E017, E031, E051]

Takeaway: Treat these as representative arrivals in this keyword-filtered feed, not a complete map of the field. The packet also includes datasets and reproducibility artifacts such as synthetic federated-learning evidence, annotation-disagreement data, security probes, and decision-change records. [E004, E018, E028, E040]

Another reading: The supplied evidence packet does not describe every artifact first observed today, so a complete inventory is unavailable. Some records are papers proposing methods, while others are dataset releases or repository discoveries; radar arrival does not establish that each artifact itself was newly created.

  • S022 artifacts first observed by the radar today: 335

Which of today's arrivals document how they score an answer?

medium confidence

The clearest examples are the Finance Agents Benchmark and its traces, which pair tasks with rubrics, grading criteria, verdicts, and judge explanations; Clankdar, which uses deterministic rules and records scorer versions; and the decision-model package, which supplies gold keys and supports re-scoring. [E001, E005, E055, E061]

These arrivals expose more than final scores. They provide expected answers, grading rules, pass-or-fail decisions, explanations, or the exact scorer version, making it possible to inspect how an answer became a result. Sys1 also supplies expected answers and decision outcomes under fixed cost policies. [E003, E009, E055, E061]

Takeaway: For answer-level auditability in this captured feed, start with Finance Agents Benchmark, Clankdar, Sys1, and the decision-model benchmark. The MCP security benchmark also publishes a fixed probe collection and per-target scorecards, although its summary does not fully spell out the scoring algorithm. [E001, E005, E009, E040, E055, E061]

Another reading: Repository summaries may overstate practical reproducibility: the packet confirms that scoring materials are described, but not that every rule, dependency, judge prompt, or edge case is fully documented. Several other arrivals mention rubrics or evaluation prompts without enough detail here to classify their scoring as auditable. [E014, E059]

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

medium confidenceNot enough evidence

The registry flags download gains for lmarena-ai/leaderboard-dataset, hf-benchmarks/transformers, vedangfake/chess-slm-benchmark, alexshpunt/explicit-edit-benchmark, Weyaxi/huggingface-leaderboard, and dreamdifferent/vam-cross-evaluation-artifacts; a download decline for vava22684/song-jury-leaderboard; and a star gain for career-ops-hq/career-ops.

These are cumulative changes across each artifact’s full tracked window, not changes from a single day. The cited statistics provide the exact spans, which range from mid-September through late September to late July through late September.

Takeaway: Within this captured feed, these are the registry-backed measurable movers. Artifact-level evidence records were not supplied, so the movements cannot be independently grounded with required evidence citations.

Another reading: The tracked-artifact packet contains additional metric changes, but those lack movement statistics in the registry and therefore cannot be reported quantitatively under the grounding rules.

  • S025 downloads change for lmarena-ai/leaderboard-dataset: 25,056.0 downloads
  • S026 downloads change for hf-benchmarks/transformers: 23,005.0 downloads
  • S027 downloads change for vedangfake/chess-slm-benchmark: 14,788.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 11,674.0 downloads
  • S029 downloads change for Weyaxi/huggingface-leaderboard: 7,346.0 downloads
  • S030 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 5,078.0 downloads
  • S031 downloads change for vava22684/song-jury-leaderboard: -4,081.0 downloads
  • S032 stars change for career-ops-hq/career-ops: 3,289.0 stars

Which of that movement is corroborated by more than one data source?

high confidenceNot enough evidence

None of the registry-backed movements listed above is corroborated by more than one data source; each movement statistic names only one source.

Seeing an artifact in several sources is not enough. More than one source must report the same metric movement, and the registry does not show that for the listed movers.

Takeaway: No corroborated metric movement can be established from the supplied movement statistics for this captured feed. Artifact-level evidence citations are also missing.

Another reading: The tracked-artifact packet marks modelscope/evalscope as corroborated across GitHub and GitHub Release, but its movement has no corresponding registry statistic or evidence citation, so it cannot be reported as a grounded corroborated movement.

  • S024 tracked artifacts today seen by more than one data source: 7
  • S025 downloads change for lmarena-ai/leaderboard-dataset: 25,056.0 downloads
  • S026 downloads change for hf-benchmarks/transformers: 23,005.0 downloads
  • S027 downloads change for vedangfake/chess-slm-benchmark: 14,788.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 11,674.0 downloads
  • S029 downloads change for Weyaxi/huggingface-leaderboard: 7,346.0 downloads
  • S030 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 5,078.0 downloads
  • S031 downloads change for vava22684/song-jury-leaderboard: -4,081.0 downloads
  • S032 stars change for career-ops-hq/career-ops: 3,289.0 stars

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Use today’s captured feed as a triage queue, not as a reason to replace established tests or infer a field-wide shift.

Review newly released tests that closely match your deployment. The captured feed includes finance-agent tasks with execution traces and grading material, plus browser workflows evaluated through completed work products. Pilot relevant cases alongside the tests you already run, and inspect the data, scoring rules, and execution setup before adopting them.

Takeaway: Add a task-matched benchmark review to today’s evaluation work, especially for agents, but require reproducible inputs, scoring, and repeated runs before results affect release decisions. This recommendation applies only to the keyword-filtered captured feed.

Another reading: The strongest competing reading is that no process change is warranted: these are newly released, author-described artifacts without independent validation in the packet. Teams without matching finance-agent or browser-workflow needs may gain nothing from reviewing them now.

  • S003 records with event kind released: 282
  • S004 records with event kind updated: 181
  • S005 records with event kind discovered: 42
  • S017 records tagged benchmark: 295 count (multi-label)
  • S018 records tagged evaluation: 257 count (multi-label)
  • S020 records tagged agentic: 46 count (multi-label)
  • S022 artifacts first observed by the radar today: 335

What does today's evidence fail to show, and what would change the reading?

high confidenceNot enough evidence

The captured feed does not establish a field-wide trend, benchmark quality, model improvement, or independently corroborated popularity movement.

There is no certified comparison window, and differences between recent periods may reflect collection changes. The download and star movements cover each artifact’s full tracked span rather than one day, and none is corroborated by more than one source. The feed also does not provide independent replication or enough performance evidence to rank the new releases.

Takeaway: The reading would change with stable comparable coverage, restored missing connectors, representative sampling, repeated independent evaluations, and per-metric confirmation from multiple sources. Until then, treat releases, updates, and attention observations as leads rather than proof of adoption or progress.

Another reading: The competing reading is that the feed’s breadth and volume are still useful for discovering evaluation candidates, even without trend certification. That supports prioritization, but it does not overcome the sampling, comparability, validation, or corroboration limits.

  • S001 evidence records captured today: 505
  • S002 public attention observations captured today: 10
  • S003 records with event kind released: 282
  • S004 records with event kind updated: 181
  • S005 records with event kind discovered: 42
  • S024 tracked artifacts today seen by more than one data source: 7
  • S025 downloads change for lmarena-ai/leaderboard-dataset: 25,056.0 downloads
  • S026 downloads change for hf-benchmarks/transformers: 23,005.0 downloads
  • S027 downloads change for vedangfake/chess-slm-benchmark: 14,788.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 11,674.0 downloads
  • S029 downloads change for Weyaxi/huggingface-leaderboard: 7,346.0 downloads
  • S030 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 5,078.0 downloads
  • S031 downloads change for vava22684/song-jury-leaderboard: -4,081.0 downloads
  • S032 stars change for career-ops-hq/career-ops: 3,289.0 stars
  • S033 daily-average change in evaluation observations: -50.0 observations per day
  • S034 daily-average change in dataset observations: -19.71 observations per day
  • S035 daily-average change in benchmark observations: -17.0 observations per day
  • S036 daily-average change in data_quality observations: -3.29 observations per day
  • S037 daily-average change in agentic observations: -2.43 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 505 evidence records.

Briefing model: gpt-5.6-sol.

It read 153 of 505 records.

This briefing covers a keyword-filtered, nonrepresentative feed. Only 153 of today’s 505 evidence records were injected into the briefing, so the findings do not represent a review of the full captured corpus. Twelve of fourteen evidence connectors were healthy; Brave and OpenReview were unavailable. Comparable history was insufficient to assess multi-day category-share changes.