Benchmark Radar™
RSS Contact Star

Daily AI benchmark brief: 2026-09-14

419evidence observations
10sources represented
15public-attention signals

Daily briefing

  1. The new KNOWS Benchmark evaluates web agents on 110 tasks that combine web research with producing documents, spreadsheets, and slide decks in Google Workspace. Each task includes the agent prompt and a structured rubric for grading both the finished artifact and its faithful use of retrieved sources. This moves evaluation beyond answering questions toward completing recognizable office work. Why it matters: If you are choosing an evaluation suite for workplace agents, KNOWS lets you test whether research survives into a usable, source-supported deliverable—not merely whether the agent navigates tools or returns a plausible final answer. Evidence: E001. Medium confidence.
  2. The new DiligenceProv benchmark evaluates financial due-diligence answers using synthetic deal rooms generated from a private truth ledger. It deliberately includes conflicting definitions, distractors, authority rules, and cases where evidence is silently absent, then scores answers at a fine-grained level for numbers, definitions, and supporting evidence. Why it matters: If you are evaluating acquisition-research systems, this design tests whether a model can distinguish an unsupported answer from a verifiable one under document-room imperfections. That is a different selection criterion from accuracy on reconciled public filings. Evidence: E005. Medium confidence.
  3. The new K-Bench tests whether knowledge removed from a large language model remains accessible after the model becomes an agent. Instead of treating a refusal as successful forgetting, it checks six exposed channels—including reasoning traces, tool calls, tool observations, and summaries—and separately places the secret in model weights, the prompt, or the retrieval store. Why it matters: If you are validating knowledge-removal or privacy controls for agents, final-answer testing can miss disclosures elsewhere in execution. K-Bench provides a way to choose controls based on end-to-end leakage rather than model behavior alone. Evidence: E015. Medium confidence.
  4. The new PhysCodeBench contains 1,200 expert-validated tasks for generating executable three-dimensional physics simulations from natural-language descriptions. Its evaluator inspects engine state using conservation-law residuals and expert assertions, rather than relying only on whether code runs or the rendered result looks plausible; it also supports comparisons across simulation engines. Why it matters: If you are selecting models to generate scientific simulation code, this separates programming success from physical correctness. A model that produces executable but physically wrong code can therefore fail explicitly instead of receiving credit for appearance or runtime alone. Evidence: E011. Medium confidence.
  5. The new ParaRecover benchmark evaluates how agents locate and recover from intermediate failures during multi-turn, parallel tool use. Its 10,626 instances cover 14 error types across planning dependencies, tool selection, and argument matching, exposing failures that can propagate between dependent execution branches. Why it matters: If you are comparing agents for workflows that call several tools concurrently, final task success does not explain whether recovery is dependable. ParaRecover supports model and orchestration choices based on where an agent detects an error and whether it repairs the affected process. Evidence: E020. Medium confidence.
  6. The new TraceJudgeBench audits large-language-model judges used to grade systems with citations or retrieved evidence. It tests content-equivalent pairs, citation removal, correctness conflicts, graded quality gaps, and increasingly forceful instructions to ignore presentation cues; the authors report that stronger anti-citation prompts can reduce bias while also reducing the judge’s ability to distinguish quality differences. Why it matters: If you use model-based judges for retrieval or agent evaluations, debiasing instructions may change both fairness and measurement sensitivity. TraceJudgeBench offers a way to test that trade-off before using judge scores to rank systems. Evidence: E019. Medium confidence.
  7. The new SynthSentry method screens a prospective training corpus for synthetic-data contamination without access to the generating model, generation history, or synthetic labels. Its corpus-level score combines changes in lexical diversity, rare phrase patterns, and variation in perplexity—how surprising text appears to reference models. Why it matters: If you are deciding whether an unknown-provenance corpus is suitable for training, this introduces a pre-training screening option rather than waiting to diagnose degradation after training. The captured evidence describes the signal, but not enough results to establish its reliability across unrelated corpora. Evidence: E010. Low confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

medium confidenceNot enough evidence

In this keyword-filtered feed, first-seen artifacts covered workplace web agents, financial diligence, physics-aware simulation, long-manual reasoning, medical fact verification, agent unlearning, benchmark compression, judge-bias diagnosis, tool-failure recovery, graph reasoning, and synthetic-data contamination detection.

Representative releases were KNOWS, DiligenceProv, PhysCodeBench, Tasks over Application Manuals, MedSNIP-Bench, K-Bench, ZipBench, TraceJudgeBench, ParaRecover, Graph Theory Bench, and SynthSentry. Other arrivals introduced paired-view cell-image evaluation and leakage-controlled antibody testing.

Takeaway: Today’s first sightings span both new task collections and evaluation methods that inspect intermediate behavior, physical correctness, evidence use, leakage, or measurement bias rather than relying solely on final-task success.

Another reading: This is not an exhaustive inventory because the supplied evidence packet contains only selected captured records. Keyword tagging also produces false positives: the AIREP mirror explicitly says it is neither a training dataset nor a benchmark despite receiving those tags.

  • S021 artifacts first observed by the radar today: 225

Which of today's arrivals document how they score an answer?

medium confidenceNot enough evidence

The clearest released arrivals with stated scoring logic are KNOWS, DiligenceProv, PhysCodeBench, K-Bench, and the land-use relevance benchmark. An updated CORE-LLM-Bench record also describes symbolically derived ground-truth answers and explanation structures.

KNOWS supplies a structured rubric for the produced artifact. DiligenceProv says answers are scored in atomic parts. PhysCodeBench checks execution, appearance, conservation rules, and expert assertions. K-Bench marks leakage when a secret appears in any exposed agent channel. The land-use benchmark recomputes classification measures from published predictions.

Takeaway: These records expose at least one concrete grading mechanism—a rubric, answer decomposition, executable checks, a leakage decision rule, published prediction metrics, or symbolic ground truth—rather than merely stating that an evaluation occurred.

Another reading: The evidence is insufficient for a complete list, and several summaries prove only that scoring exists. DiligenceProv’s description is truncated, while KNOWS and CORE-LLM-Bench do not expose full weighting, aggregation, or adjudication details in the supplied text.

  • S021 artifacts first observed by the radar today: 225

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

medium confidenceNot enough evidence

The registry records cumulative download movement for several datasets and cumulative star movement for a repository. Each statistic covers the artifact’s entire tracked span rather than a one-day change.

The movement statistics and their spans are available, but the supplied packet contains no E-labeled evidence records. Under the grounding rules, I therefore cannot safely name the individual artifacts or present their specific movements as verified artifact claims.

Takeaway: Use the cited movement statistics as provisional pointers, not a publishable artifact list. E-labeled source evidence is missing for every registered movement.

Another reading: The stat registry itself includes artifact labels, sources, and tracked spans, so it could be treated as sufficient metadata. However, the required artifact-level E citations are absent.

  • S024 downloads change for hf-benchmarks/transformers: 15,236.0 downloads
  • S025 downloads change for lmarena-ai/leaderboard-dataset: 4,244.0 downloads
  • S026 downloads change for latency-sensitive-bench/benchmark-datasets: 1,098.0 downloads
  • S027 stars change for confident-ai/deepeval: 1,029.0 stars
  • S028 downloads change for AlphaDojo/dojo_benchmark_kline: 1,002.0 downloads
  • S029 downloads change for GOD111111111/synthetic-timeseries-data: 787.0 downloads
  • S030 downloads change for brettsp/stan-benchmark: 609.0 downloads
  • S031 downloads change for OpenChainBench/benchmarks: 544.0 downloads

Which of that movement is corroborated by more than one data source?

medium confidenceNot enough evidence

None of the movement statistics cited above is marked as corroborated by more than one data source. Each is attributed to a single source.

For the movements that have registered statistics, no second source independently reported the same metric movement. The broader tracked-artifact data includes a corroboration flag outside these registered movement statistics, but its metric movement lacks both a registry statistic and E-labeled evidence.

Takeaway: No registered movement can be treated as corroborated in this captured, keyword-filtered feed. Additional source-matched metric records are needed for a complete answer.

Another reading: An artifact-level corroboration flag elsewhere in the packet suggests that some movement may have multi-source support, but the missing registered metric and evidence citations prevent verification under the stated rules.

  • S023 tracked artifacts today seen by more than one data source: 5
  • S024 downloads change for hf-benchmarks/transformers: 15,236.0 downloads
  • S025 downloads change for lmarena-ai/leaderboard-dataset: 4,244.0 downloads
  • S026 downloads change for latency-sensitive-bench/benchmark-datasets: 1,098.0 downloads
  • S027 stars change for confident-ai/deepeval: 1,029.0 stars
  • S028 downloads change for AlphaDojo/dojo_benchmark_kline: 1,002.0 downloads
  • S029 downloads change for GOD111111111/synthetic-timeseries-data: 787.0 downloads
  • S030 downloads change for brettsp/stan-benchmark: 609.0 downloads
  • S031 downloads change for OpenChainBench/benchmarks: 544.0 downloads

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Within this captured feed, evaluation observations rose across the recent comparable week. New releases emphasize weaknesses that final-answer scoring can miss, including intermediate tool failures, leakage across agent channels, long-horizon instruction following, and grounding in retrieved sources.

Run a limited internal trial that scores the steps an agent takes, not just whether its final answer looks correct. Check tool choices, recovery after errors, use of source material, and whether protected information appears anywhere in the workflow. Validate each candidate against your own tasks before adopting it.

Takeaway: Add process-level checks beside existing outcome tests, but do not replace production gates based on today’s releases alone. Select a task-matched candidate, reproduce its setup, and test whether it reveals failures your current evaluation misses.

Another reading: The competing reading is release churn rather than a practice-changing shift. Agentic observations declined across the recent comparable week, and cross-source sightings remained limited, so the feed does not show broad adoption or independent confirmation.

  • S032 daily-average change in evaluation observations: 24.14 observations per day
  • S035 daily-average change in agentic observations: -4.43 observations per day
  • S021 artifacts first observed by the radar today: 225
  • S022 artifacts seen today that the radar had already tracked: 194
  • S023 tracked artifacts today seen by more than one data source: 5

What does today's evidence fail to show, and what would change the reading?

high confidenceNot enough evidence

Today’s captured feed does not establish that any new benchmark is valid, representative, widely adopted, or linked to better deployed-system outcomes. Cross-source sightings were rare among artifacts seen today, and this keyword-filtered feed cannot support a field-wide conclusion.

Missing evidence includes independent reproduction, direct comparisons under identical settings, checks for contaminated or leaked test data, stable results across model versions, and validation on real production tasks. Restored coverage from unavailable sources would also reduce collection uncertainty.

Takeaway: The reading would strengthen if independent groups reproduced the results, multiple sources tracked the same artifacts and measurements, and teams showed that the evaluations predict failures or improvements in deployed systems. Until then, treat the records as leads for investigation rather than proof.

Another reading: The strongest competing reading is that a genuine feed-level shift may be underway: evaluation and dataset observations both rose across the recent comparable week with broad source coverage. Even so, that establishes increased captured activity, not benchmark quality or operational value.

  • S021 artifacts first observed by the radar today: 225
  • S022 artifacts seen today that the radar had already tracked: 194
  • S023 tracked artifacts today seen by more than one data source: 5
  • S032 daily-average change in evaluation observations: 24.14 observations per day
  • S033 daily-average change in dataset observations: 20.57 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 419 evidence records.

Briefing model: gpt-5.6-sol.

It read 120 of 419 records.

This radar is keyword-filtered and not representative of the AI field. The briefing received 120 of today’s 419 captured evidence records, so conclusions cover only that ranked subset. Brave Search and OpenReview were unavailable, and none of the supplied attention signals was observed today. Descriptions largely come from artifact authors and were not independently verified.