Benchmark Radar™
RSS Contact Star —

Daily AI benchmark brief: 2026-10-05

560evidence observations
12sources represented
18public-attention signals

Daily briefing

  1. The new “Passing the Test You Trained On” study evaluates 15 prompt-injection detectors on benign and injected outputs reconstructed from real agent benchmark tool calls. Detector rankings transferred poorly across benchmarks; the best detector on the public BIPIA benchmark caught only 2% of AgentDojo injections at the study’s stated operating point. [E001] Why it matters: If you are selecting a detector for a tool-using agent, public benchmark rank alone may not predict protection in your workflow. The study provides a concrete design for deployment-matched evaluation: replay actual tool-call traces, derive labels through differential replay, and compare detectors at a fixed false-alarm tolerance. Evidence: E001. High confidence.
  2. The newly released milipoint-runs rebuilds a radar dataset’s evaluation splits after finding that shuffled, overlapping windows let most test examples share frames with training examples. Holding out complete recording runs reduced the reported PointMLP identification accuracy from about 94% to 37–39%. [E046] Why it matters: If you evaluate models on windowed sensor streams, splitting individual windows can measure familiarity with neighboring frames rather than generalization to a new session. This artifact makes run-disjoint and participant-disjoint splits a concrete alternative when choosing protocols or reassessing published baselines. Evidence: E046. High confidence.
  3. OpenGameEval is a new benchmark that runs programming agents in reproducible, stateful Roblox Studio sessions. It scores executable checks on both the edited scene and simulated play, while separating observation tools from editing tools so that exploration behavior can be measured independently from final task completion. [E027] Why it matters: If you are choosing an agent benchmark for interactive software work, this design distinguishes an agent that inspected the environment effectively from one that happened to produce a passing final state. That separation can change debugging and model-selection decisions because failures can be attributed to exploration or editing. Evidence: E027. High confidence.
  4. Open-Endedness Bench is a new reference-free method for evaluating research agents from their execution records. Instead of judging only the final score, it examines how an agent forms hypotheses, runs experiments, and revises claims in response to evidence, including tasks where no known optimum or reference answer exists. [E039] Why it matters: If you evaluate agents on discovery or optimization tasks, outcome-only scoring cannot establish whether claims follow from executed experiments. This method offers a way to compare research process quality when conventional answer matching is unavailable. Evidence: E039. Medium confidence.
  5. OLMo-Detect is a new membership-inference benchmark for testing whether an evaluator can identify text used to train a language model. It spans pre-training, mid-training, and post-training, aligns member and non-member examples on stated distributional factors, and filters supposed non-members against the open OLMo 2 training corpus. [E024] Why it matters: If you are comparing training-data detection or privacy-audit methods, uncontrolled differences between members and non-members can inflate results. OLMo-Detect supplies stage-specific and corpus-checked comparisons, which can change which inference method appears suitable for an audit. Evidence: E024. High confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

high confidence

Notable first-seen releases covered agent security, scientific verification, forecasting, membership inference, production coding, user-interest grounding, dialogue compression, and evaluation of open-ended research processes.

Examples include EvoRiskBench for workspace-agent risks, MOF-VERIFY for materials hypotheses, LEAF for leakage-aware forecasting, OLMo-Detect for training-membership tests, OpenGameEval for game-development agents, GISTBench for user-interest verification, TPBench for dialogue compression, and Open-Endedness Bench for judging research behavior from execution records.

Takeaway: In this keyword-filtered feed, the first-seen set is unusually broad and includes both task-specific datasets and methods that inspect how agents work, not merely whether they reach a final result. These were first observed by the radar today, which does not establish that they were first released to the field today.

Another reading: Radar novelty is not field novelty. Some records may be mirrors, later indexing events, or versions published earlier, and unavailable sources could hide prior sightings. The evidence packet also highlights selected arrivals rather than describing every first-seen artifact.

  • S023 artifacts first observed by the radar today: 314

Which of today's arrivals document how they score an answer?

medium confidence

The clearest documented scoring approaches are compile verdicts, executable checks, field-level answer matching, deterministic rules, groundedness and specificity metrics, rubric scoring, and process evaluation from execution logs.

The MQL compile set records whether generated code compiles. OpenGameEval runs checks on edited scenes and simulated play. The invoice benchmark compares extracted fields with answer keys. PanduGizi supplies deterministic scoring rules. GISTBench measures supported and distinctive user interests. LICA uses a writing rubric. Open-Endedness Bench judges hypothesis formation, testing, and revision from agent records.

Takeaway: These arrivals expose more than a leaderboard result: they identify what evidence earns credit. The strongest machine-checkable examples are compilation, executable environment checks, and field-by-field comparison; the rubric and process-based methods offer richer criteria but require more judgment.

Another reading: The summaries do not consistently provide weighting, thresholds, aggregation rules, or complete implementation details, so documented criteria do not guarantee reproducible scoring. The MQL record also says generated completions are withheld, limiting independent inspection of scored outputs.

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

low confidenceNot enough evidence

The registry contains cumulative download-movement statistics for several previously tracked artifacts, with each statistic covering its full stated tracking span rather than a one-day change.

The movement records include both gains and declines, but the packet supplies no artifact-level evidence identifiers. Naming the artifacts as verified movers would therefore violate the evidence requirement.

Takeaway: The movement statistics are available through the cited stat entries, but compliant artifact-level identification requires corresponding evidence records.

Another reading: The stat labels themselves name the artifacts and spans, so they may be operationally sufficient; however, they do not replace the required artifact-level evidence citations.

  • S026 downloads change for open-llm-leaderboard/requests: -32,781.0 downloads
  • S027 downloads change for hf-benchmarks/transformers: 24,049.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 18,447.0 downloads
  • S029 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 8,439.0 downloads
  • S030 downloads change for vava22684/song-jury-leaderboard: -3,967.0 downloads
  • S031 downloads change for AlphaDojo/dojo_benchmark_kline: 3,691.0 downloads
  • S032 downloads change for NoeFlandre/geoparser-benchmark-results: 2,789.0 downloads
  • S033 downloads change for morzel85/synthetic-medical-document-recognition-benchmark: 2,018.0 downloads

Which of that movement is corroborated by more than one data source?

low confidenceNot enough evidence

The aggregate multi-source statistic counts independent artifact sightings, not confirmation that multiple sources measured the same movement metric.

The embedded tracking metadata marks some movements as multi-source, but those entries lack evidence identifiers and dedicated movement statistics in the registry. The corroborated subset therefore cannot be reported compliantly.

Takeaway: Artifact-level, per-metric evidence from each reporting source is missing and is needed to identify corroborated movement.

Another reading: The embedded metadata provides source lists and corroboration flags, which supports a competing reading that the subset is identifiable; the missing evidence citations still prevent a grounded answer.

  • S025 tracked artifacts today seen by more than one data source: 13

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Within this captured feed, separate new releases highlight weak transfer between benchmarks, future-information leakage, and failures that appear only in realistic operating contexts.

Do not choose a model or safety layer from one leaderboard alone. Replay production-like inputs, test on a separate dataset and operating context, check that test information could not enter training or retrieval, and report failures by task rather than only an overall score.

Takeaway: Make transfer and leakage checks release gates for evaluation pipelines. The feed’s recent benchmark and evaluation activity is higher than in the prior comparable window, but more releases do not make any individual result reliable without contextual testing.

Another reading: The cited records are new releases rather than independent replications, and several concern specialized domains. Their findings may not transfer to every system, so the strongest justified change is adding validation checks rather than rejecting existing benchmarks wholesale.

  • S003 records with event kind released: 329
  • S034 daily-average change in evaluation observations: 136.43 observations per day
  • S035 daily-average change in benchmark observations: 122.29 observations per day
  • S037 daily-average change in agentic observations: 29.43 observations per day

What does today's evidence fail to show, and what would change the reading?

high confidenceNot enough evidence

This captured feed does not establish fieldwide performance improvement, benchmark quality, real deployment adoption, or causes behind the higher volume of benchmark and evaluation observations.

More records do not mean better systems. The reported download movements come from single sources, cover different full tracking spans, and are not daily changes. A stronger reading would require independent reruns, shared test methods, production-like held-out tasks, and adoption signals confirmed by more than one source.

Takeaway: Treat today as a discovery queue, not a market or quality verdict. Broader source coverage, repeated results across unrelated teams, common measurement periods, and direct evidence that benchmark results predict deployed behavior would change the reading.

Another reading: The measurement windows are comparable, and several category averages are higher than in the prior window across broad source coverage. That supports a competing reading of genuinely greater publishing activity within this feed, though it still does not demonstrate better models or wider adoption.

  • S025 tracked artifacts today seen by more than one data source: 13
  • S026 downloads change for open-llm-leaderboard/requests: -32,781.0 downloads
  • S027 downloads change for hf-benchmarks/transformers: 24,049.0 downloads
  • S028 downloads change for alexshpunt/explicit-edit-benchmark: 18,447.0 downloads
  • S029 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 8,439.0 downloads
  • S030 downloads change for vava22684/song-jury-leaderboard: -3,967.0 downloads
  • S031 downloads change for AlphaDojo/dojo_benchmark_kline: 3,691.0 downloads
  • S032 downloads change for NoeFlandre/geoparser-benchmark-results: 2,789.0 downloads
  • S033 downloads change for morzel85/synthetic-medical-document-recognition-benchmark: 2,018.0 downloads
  • S034 daily-average change in evaluation observations: 136.43 observations per day
  • S035 daily-average change in benchmark observations: 122.29 observations per day
  • S036 daily-average change in dataset observations: 65.0 observations per day
  • S037 daily-average change in agentic observations: 29.43 observations per day
  • S038 daily-average change in data_quality observations: 2.43 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 560 evidence records.

Briefing model: gpt-5.6-sol.

It read 146 of 560 records.

This is a keyword-filtered feed, not a representative sample of artificial intelligence work. The briefing received 146 selected evidence records from a 560-record corpus, so it did not inspect the full captured corpus. Brave and OpenReview were unavailable, and the first-party feed reported a decoding error. Claims therefore apply only to the supplied evidence.