Benchmark Radar
RSS Contact

Daily brief

Daily AI benchmark brief: 2026-09-01

EleutherAI released Language Model Evaluation Harness v0.4.13 with fixes that can change prior scores: test questions could leak into their own few-shot…

daily briefAI benchmarksevaluation
515evidence observations
9sources represented
8public-attention signals

Daily briefing

  1. EleutherAI released Language Model Evaluation Harness v0.4.13 with fixes that can change prior scores: test questions could leak into their own few-shot demonstrations, a multiple-choice filter was faulty, and grouped standard errors were incorrect. The release also adds two execution backends and eight benchmark suites. Why it matters: If you use this harness for model comparisons, results from older versions may not be comparable with v0.4.13. Version-pinning, rerunning affected tasks, and documenting whether demonstrations were leakage-free now become part of deciding which scores can support a release or procurement decision. Evidence: E001. High confidence.
  2. The new real-time macroeconomic nowcasting evaluation tests whether large language model agents can estimate current indicators before official values are released. This design addresses a specific contamination problem: historical economic figures are widely published and may already exist in model training data. Why it matters: If you are evaluating models for time-sensitive forecasting, historical question answering can mistake memorization for forecasting ability. A pre-release protocol changes the model-selection decision by measuring predictions against information that was genuinely unavailable at evaluation time. Evidence: E003. Medium confidence.
  3. ASPIRE introduces a benchmark in which an agent receives only a vague capability goal while downstream tests remain hidden, requiring the agent to define what to learn and how to improve. S3Gym independently separates permissive self-testing from strict held-out evaluation, supporting a recurring concern in this feed: agents must not control the test used to claim self-improvement. Why it matters: If you are comparing self-improving agents, these designs distinguish optimizing a supplied score from identifying and closing capability gaps. Hidden or held-out tests make it harder for an agent to manufacture apparent progress by adapting only to its own chosen exercises or judgments. Evidence: E005, E006. High confidence.
  4. EvoSkill Injection is a new threat model for agents that generate, store, refine, and reuse skills. It targets the possibility that a malicious capability can enter this reusable skill pipeline and later execute as if it were a legitimate learned behavior. Why it matters: If your product lets agents preserve procedures across tasks, evaluating only the current prompt and response misses persistent compromise. This artifact adds the skill store and its generation history as evaluation targets, affecting decisions about provenance checks, red-team scenarios, and whether autonomous skill reuse is safe to enable. Evidence: E002. Medium confidence.
  5. A new pre-registered audit treats large language model essay judges as measurement instruments rather than relying only on agreement scores. Across 2,377 essays, 12 judges, four providers, and five version contrasts, it measures grader severity, halo effects, reliability, and score shifts between versions. Why it matters: If you use model judges for educational scoring or benchmark grading, agreement with humans alone may conceal a consistently harsh judge or a version-dependent scoring shift. Judge selection therefore depends on severity and stability checks, while production evaluation needs frozen versions or recalibration after upgrades. Evidence: E019. Medium confidence.
  6. ScienceArena is a new benchmark built from recent science olympiad competitions in physics, chemistry, and biology. It preserves open-ended, multi-step problems and process-credit rubrics, then calibrates automated model judging against expert grading rather than reducing each problem to final-answer accuracy. Why it matters: If you are choosing a scientific-reasoning evaluation, this design measures whether intermediate work earns the same partial credit that expert graders would assign. That changes comparisons between models that reach similar final answers through materially different reasoning processes. Evidence: E015. Medium confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

medium confidenceNot enough evidence

The captured feed first observed many artifacts, spanning agent behavior, scientific reasoning, medicine, vision, speech, safety, and evaluation design. Notable examples include ASPIRE, ScienceArena, ECGQuest, E-Commerce Bench, BIRD-History, SurgSkill-Bench, and SafeAtlas-VL.

Examples include hidden-task evaluation for self-improving agents, rubric-based grading of scientific answers, medical question data, long-running business simulations, database-query tasks using historical knowledge, surgical-skill assessment, and ordered multimodal safety judgments.

Takeaway: Today’s first observations show substantial breadth within this keyword-filtered feed. They include both newly released artifacts and items merely discovered by the radar; first observation does not establish that an artifact is globally new.

Another reading: This is not an exhaustive inventory. The packet contains selected evidence rather than documentation for every first-observed artifact, and some records published earlier were only discovered today. Full metadata for the remaining arrivals is missing.

  • S020 artifacts first observed by the radar today: 286

Which of today's arrivals document how they score an answer?

medium confidenceNot enough evidence

The clearest summary-level scoring descriptions appear in ScienceArena, R-GROUNDBENCH, PaperGym, AutoSciRub, SafeAtlas-VL, and the public AI-response evaluation sample.

ScienceArena gives partial credit through expert-audited rubrics and calibrates automated judging against medalists. R-GROUNDBENCH checks whether the selected molecule is correct. PaperGym converts paper-derived criteria into a reward. AutoSciRub verifies work against task-specific criteria. SafeAtlas-VL rates safety on an ordered scale. The response sample names dimensions such as relevance, tone, concision, severity, and preference.

Takeaway: These arrivals expose more than a final leaderboard result: they identify answer criteria, judgment targets, or verification procedures. Within the captured feed, ScienceArena provides the clearest account of grading open-ended answers.

Another reading: The supplied summaries generally omit complete formulas, weighting, aggregation rules, and evaluator prompts. They support identifying likely scoring approaches, but not confirming that each method is fully specified or reproducible. Unselected arrivals may contain additional scoring documentation.

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

low confidenceNot enough evidence

The registry contains several cumulative movement statistics over each artifact’s full tracked span, but the supplied packet has no E-tagged evidence records supporting those artifact-specific claims.

I cannot safely name the artifacts or restate their spans under the citation rules. The movement entries are present in the registry, but their evidence lists are empty.

Takeaway: Treat the registered movements as unresolved until artifact-level evidence records are supplied. They must not be interpreted as one-day changes.

Another reading: The registry labels may themselves appear sufficient to identify the artifacts and spans, but the required artifact-level evidence citations are missing.

  • S023 downloads change for RoboDojo-Benchmark/RoboDojo: 28,571.0 downloads
  • S024 downloads change for huggingface-projects/drlc-leaderboard-data: 18,639.0 downloads
  • S025 downloads change for hf-benchmarks/transformers: 6,022.0 downloads
  • S026 downloads change for AlphaDojo/dojo_benchmark_kline: 4,935.0 downloads
  • S027 downloads change for sselaine27/benchmark-research: 4,680.0 downloads
  • S028 downloads change for vava22684/song-jury-leaderboard: -3,371.0 downloads
  • S029 downloads change for lmarena-ai/leaderboard-dataset: -2,734.0 downloads
  • S030 downloads change for Weyaxi/huggingface-leaderboard: 2,030.0 downloads

Which of that movement is corroborated by more than one data source?

high confidenceNot enough evidence

None of the registered movement statistics is marked corroborated; each relies on one source. The packet also contains tracked metadata suggesting other corroborated records, but provides neither registered movement statistics nor E-tagged evidence for them.

Seeing an artifact in multiple feeds is not enough. More than one source must report the same metric movement, and the supplied registered movements do not meet that test.

Takeaway: No specific corroborated movement can be reported compliantly. Per-metric registered statistics and artifact-level evidence are missing for the records flagged elsewhere as corroborated.

Another reading: The multi-source artifact count could be read as corroboration, but its registry note explicitly says the sources may not have measured the same metric.

  • S022 tracked artifacts today seen by more than one data source: 5
  • S023 downloads change for RoboDojo-Benchmark/RoboDojo: 28,571.0 downloads
  • S024 downloads change for huggingface-projects/drlc-leaderboard-data: 18,639.0 downloads
  • S025 downloads change for hf-benchmarks/transformers: 6,022.0 downloads
  • S026 downloads change for AlphaDojo/dojo_benchmark_kline: 4,935.0 downloads
  • S027 downloads change for sselaine27/benchmark-research: 4,680.0 downloads
  • S028 downloads change for vava22684/song-jury-leaderboard: -3,371.0 downloads
  • S029 downloads change for lmarena-ai/leaderboard-dataset: -2,734.0 downloads
  • S030 downloads change for Weyaxi/huggingface-leaderboard: 2,030.0 downloads

What it means

What should someone building or evaluating AI systems do differently today?

high confidence

In this captured feed, the most decision-relevant item is a new lm-evaluation-harness release reporting correctness fixes that can change prior scores, while benchmark and evaluation observations rose across the comparable recent and prior windows.

If your evaluations use this harness or similar pipelines, freeze versions and configurations, check whether test examples entered few-shot prompts, validate multiple-choice filtering and reported error bars, then rerun affected baselines before comparing models. Adding new tests should not be confused with correcting old results.

Takeaway: Treat evaluation infrastructure as versioned measurement equipment: record the exact setup, add leakage checks, and rerun affected comparisons after adopting correctness fixes.

Another reading: The reported fixes may not affect your tasks, and this feed contains no independent reproduction showing which results materially change. Teams using unrelated evaluation paths may need no immediate change.

  • S031 daily-average change in benchmark observations: 118.71 observations per day
  • S032 daily-average change in evaluation observations: 71.0 observations per day

What does today's evidence fail to show, and what would change the reading?

high confidence

Today’s packet does not establish a field-wide shift, benchmark quality, model improvement, or adoption. This is a keyword-filtered feed, only a small subset of tracked artifacts appeared through more than one source, and the composition check found no material category pattern.

Release and update labels show that something appeared or changed, not that it works, is reproducible, or matters in production. The harness note identifies possible score-changing bugs, but the packet contains no independent reruns showing which published results change.

Takeaway: The reading would change with independent reproductions, before-and-after results on fixed test sets, verified dataset provenance and leakage checks, and persistent multi-source evidence tied to the same artifacts and measurements.

Another reading: A competing reading is that the comparable-window increases in benchmark and evaluation observations across broad source coverage reflect genuinely higher activity within this captured feed, even without artifact-level corroboration.

  • S001 evidence records captured today: 515
  • S022 tracked artifacts today seen by more than one data source: 5
  • S031 daily-average change in benchmark observations: 118.71 observations per day
  • S032 daily-average change in evaluation observations: 71.0 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 515 evidence records.

Briefing model: gpt-5.6-sol.

It read 117 of 515 records.

This radar is keyword-filtered and not representative of the AI field. The briefing received 117 selected evidence records from a captured corpus of 515, so it did not inspect every eligible record. Brave and Semantic Scholar were unavailable, and none of the carried-forward attention signals were observed today.