Benchmark Radar
RSS Contact

Daily brief

What changed in AI evaluation, and why it matters

One page per collection day: what new benchmarks and evaluations appeared, how strong the evidence was, and which claims the record does not support.

30 of 42 collection days

  1. Among today’s captured releases, EarlyEval introduces early outcome prediction: it estimates an agent’s final result from intermediate behavior and stops…

  2. The newly released InSight benchmark tests agents that must actively interact with visualizations to verify 21,349 claims, rather than answer once from a…

  3. EleutherAI released Language Model Evaluation Harness v0.4.13 with fixes that can change prior scores: test questions could leak into their own few-shot…

  4. PCFBench reflects a recurring push in this captured feed to inspect an agent’s process rather than only its final answer. It separately tests…

  5. The new INSIDER LLM Detection Benchmark evaluates models that may take harmful actions by comparing the model’s self-reported action log with an…

  6. The new NBPO benchmark-generations dataset publishes every decoded model response used in three judge-based comparisons, allowing another evaluator to…

  7. The new Same Model, Different Harness study holds the coding model and tasks fixed while changing how the agent harness manages conversation history and…

  8. OpenCompass v0.5.4 is a substantive harness update, adding native VLMEvalKit-based multimodal evaluation, multi-round inference with the Multi-IF…

  9. Across this captured feed, three new agent benchmarks make the evaluated unit an interactive model-plus-runtime system rather than a final answer…

  10. New release SUSVIBES evaluates 12 coding-agent settings on 186 real-world feature requests for which human developers previously committed vulnerable…

  11. No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a decision-useful finding in…

  12. No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a material finding in today’s…

  13. No material GPT insight: No material pattern was supported: the captured items did not show a sufficiently large, persistent, cross-source shift. Only 19…

  14. No material GPT insight: No material pattern cleared the feed’s persistence and cross-source thresholds today. Only 65 of 198 corpus evidence records were…

  15. No material GPT insight: No category changed far enough, persistently enough, and across enough sources to support a decision-useful pattern in this…

  16. No material GPT insight: No category changed far enough, persistently enough, or across enough sources to support a material finding in this captured…

  17. No material GPT insight: No category change was sufficiently large, persistent, and cross-source to support a material finding. Tracked metric movement…

  18. Several new releases in the captured feed redesign evaluation around conditions hidden by static averages: evolving evidence and temporal cutoffs…

  19. Agent benchmarking in the captured evidence is being designed around controlled execution, not just task sets: a new arena adds side-by-side and blind…

  20. In today’s captured feed, three independent new releases evaluate agents as operational systems: LCAB preserves complete repair sessions and hardware…

  21. Two new paper releases in this captured feed question whether aggregate agent scores measure deployable capability: one reports task interactions…

  22. Across today’s captured new releases, several benchmarks require inspectable process evidence rather than scoring only final outcomes: replayable…

  23. Several new releases in this captured feed bind benchmarks to deployment context: web-access APIs are compared across quality, latency, cost, and error…

  24. Multiple new releases in the captured feed move agent evaluation beyond final-task success toward inspecting collaboration graphs, requirement recovery…

  25. Several new releases in this captured feed evaluate whether a system reached an answer through valid evidence or execution, not merely whether the answer…

  26. Across this feed’s new releases, benchmark design is moving toward specialized, structurally difficult tasks: open-ended scientific extraction…

  27. Four newly released agent benchmarks in this captured feed expand evaluation beyond terminal task success: acquisition-stage privacy, learning across…

  28. Across this captured feed, four newly released benchmarks evaluate behavior beyond a static, well-specified task: proactive bug discovery without issue…

  29. Agentic artifacts rose to 26.3% of our captured feed over the last 5 days, against a 13.6% baseline across the prior 4 days (+12.7 percentage points).