<?xml version='1.0' encoding='utf-8'?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>Benchmark Radar daily brief</title>
    <link>https://benchmark-radar.org/blog/</link>
    <description>One page per collection day: what new benchmarks and evaluations appeared, how strong the evidence was, and which claims the record does not support.</description>
    <language>en</language>
    <atom:link href="https://benchmark-radar.org/blog/feed.xml" rel="self" type="application/rss+xml" />
    <lastBuildDate>Thu, 03 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Daily AI benchmark brief: 2026-09-03</title>
      <link>https://benchmark-radar.org/blog/2026-09-03/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-09-03/</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <description>Among today’s captured releases, EarlyEval introduces early outcome prediction: it estimates an agent’s final result from intermediate behavior and stops…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-09-02</title>
      <link>https://benchmark-radar.org/blog/2026-09-02/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-09-02/</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
      <description>The newly released InSight benchmark tests agents that must actively interact with visualizations to verify 21,349 claims, rather than answer once from a…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-09-01</title>
      <link>https://benchmark-radar.org/blog/2026-09-01/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-09-01/</guid>
      <pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate>
      <description>EleutherAI released Language Model Evaluation Harness v0.4.13 with fixes that can change prior scores: test questions could leak into their own few-shot…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-31</title>
      <link>https://benchmark-radar.org/blog/2026-08-31/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-31/</guid>
      <pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate>
      <description>PCFBench reflects a recurring push in this captured feed to inspect an agent’s process rather than only its final answer. It separately tests…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-30</title>
      <link>https://benchmark-radar.org/blog/2026-08-30/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-30/</guid>
      <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
      <description>The new INSIDER LLM Detection Benchmark evaluates models that may take harmful actions by comparing the model’s self-reported action log with an…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-29</title>
      <link>https://benchmark-radar.org/blog/2026-08-29/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-29/</guid>
      <pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate>
      <description>The new NBPO benchmark-generations dataset publishes every decoded model response used in three judge-based comparisons, allowing another evaluator to…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-28</title>
      <link>https://benchmark-radar.org/blog/2026-08-28/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-28/</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <description>The new Same Model, Different Harness study holds the coding model and tasks fixed while changing how the agent harness manages conversation history and…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-27</title>
      <link>https://benchmark-radar.org/blog/2026-08-27/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-27/</guid>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      <description>OpenCompass v0.5.4 is a substantive harness update, adding native VLMEvalKit-based multimodal evaluation, multi-round inference with the Multi-IF…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-26</title>
      <link>https://benchmark-radar.org/blog/2026-08-26/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-26/</guid>
      <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
      <description>Across this captured feed, three new agent benchmarks make the evaluated unit an interactive model-plus-runtime system rather than a final answer…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-25</title>
      <link>https://benchmark-radar.org/blog/2026-08-25/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-25/</guid>
      <pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate>
      <description>New release SUSVIBES evaluates 12 coding-agent settings on 186 real-world feature requests for which human developers previously committed vulnerable…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-24</title>
      <link>https://benchmark-radar.org/blog/2026-08-24/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-24/</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a decision-useful finding in…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-23</title>
      <link>https://benchmark-radar.org/blog/2026-08-23/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-23/</guid>
      <pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a material finding in today’s…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-22</title>
      <link>https://benchmark-radar.org/blog/2026-08-22/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-22/</guid>
      <pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No material pattern was supported: the captured items did not show a sufficiently large, persistent, cross-source shift. Only 19…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-21</title>
      <link>https://benchmark-radar.org/blog/2026-08-21/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-21/</guid>
      <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No material pattern cleared the feed’s persistence and cross-source thresholds today. Only 65 of 198 corpus evidence records were…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-20</title>
      <link>https://benchmark-radar.org/blog/2026-08-20/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-20/</guid>
      <pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No category changed far enough, persistently enough, and across enough sources to support a decision-useful pattern in this…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-19</title>
      <link>https://benchmark-radar.org/blog/2026-08-19/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-19/</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No category changed far enough, persistently enough, or across enough sources to support a material finding in this captured…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-18</title>
      <link>https://benchmark-radar.org/blog/2026-08-18/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-18/</guid>
      <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
      <description>No material GPT insight: No category change was sufficiently large, persistent, and cross-source to support a material finding. Tracked metric movement…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-17</title>
      <link>https://benchmark-radar.org/blog/2026-08-17/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-17/</guid>
      <pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate>
      <description>Several new releases in the captured feed redesign evaluation around conditions hidden by static averages: evolving evidence and temporal cutoffs…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-16</title>
      <link>https://benchmark-radar.org/blog/2026-08-16/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-16/</guid>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <description>Agent benchmarking in the captured evidence is being designed around controlled execution, not just task sets: a new arena adds side-by-side and blind…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-15</title>
      <link>https://benchmark-radar.org/blog/2026-08-15/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-15/</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>In today’s captured feed, three independent new releases evaluate agents as operational systems: LCAB preserves complete repair sessions and hardware…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-14</title>
      <link>https://benchmark-radar.org/blog/2026-08-14/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-14/</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>Two new paper releases in this captured feed question whether aggregate agent scores measure deployable capability: one reports task interactions…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-13</title>
      <link>https://benchmark-radar.org/blog/2026-08-13/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-13/</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>Across today’s captured new releases, several benchmarks require inspectable process evidence rather than scoring only final outcomes: replayable…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-12</title>
      <link>https://benchmark-radar.org/blog/2026-08-12/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-12/</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>Several new releases in this captured feed bind benchmarks to deployment context: web-access APIs are compared across quality, latency, cost, and error…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-11</title>
      <link>https://benchmark-radar.org/blog/2026-08-11/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-11/</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Multiple new releases in the captured feed move agent evaluation beyond final-task success toward inspecting collaboration graphs, requirement recovery…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-10</title>
      <link>https://benchmark-radar.org/blog/2026-08-10/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-10/</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Several new releases in this captured feed evaluate whether a system reached an answer through valid evidence or execution, not merely whether the answer…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-08</title>
      <link>https://benchmark-radar.org/blog/2026-08-08/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-08/</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Across this feed’s new releases, benchmark design is moving toward specialized, structurally difficult tasks: open-ended scientific extraction…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-07</title>
      <link>https://benchmark-radar.org/blog/2026-08-07/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-07/</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
      <description>Four newly released agent benchmarks in this captured feed expand evaluation beyond terminal task success: acquisition-stage privacy, learning across…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-06</title>
      <link>https://benchmark-radar.org/blog/2026-08-06/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-06/</guid>
      <pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate>
      <description>Across this captured feed, four newly released benchmarks evaluate behavior beyond a static, well-specified task: proactive bug discovery without issue…</description>
    </item>
    <item>
      <title>Daily AI benchmark brief: 2026-08-05</title>
      <link>https://benchmark-radar.org/blog/2026-08-05/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-05/</guid>
      <pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate>
      <description>Agentic artifacts rose to 26.3% of our captured feed over the last 5 days, against a 13.6% baseline across the prior 4 days (+12.7 percentage points).</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-08-04</title>
      <link>https://benchmark-radar.org/blog/2026-08-04/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-04/</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 212 evidence observations from 3 sources on 2026-08-04.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-08-03</title>
      <link>https://benchmark-radar.org/blog/2026-08-03/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-03/</guid>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 152 evidence observations from 3 sources on 2026-08-03.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-08-02</title>
      <link>https://benchmark-radar.org/blog/2026-08-02/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-02/</guid>
      <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 89 evidence observations from 2 sources on 2026-08-02.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-08-01</title>
      <link>https://benchmark-radar.org/blog/2026-08-01/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-08-01/</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 69 evidence observations from 2 sources on 2026-08-01.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-31</title>
      <link>https://benchmark-radar.org/blog/2026-07-31/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-31/</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 60 evidence observations from 2 sources on 2026-07-31.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-30</title>
      <link>https://benchmark-radar.org/blog/2026-07-30/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-30/</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 200 evidence observations from 3 sources on 2026-07-30.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-29</title>
      <link>https://benchmark-radar.org/blog/2026-07-29/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-29/</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 201 evidence observations from 3 sources on 2026-07-29.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-28</title>
      <link>https://benchmark-radar.org/blog/2026-07-28/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-28/</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 186 evidence observations from 3 sources on 2026-07-28.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-27</title>
      <link>https://benchmark-radar.org/blog/2026-07-27/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-27/</guid>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 30 evidence observations from 2 sources on 2026-07-27.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-26</title>
      <link>https://benchmark-radar.org/blog/2026-07-26/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-26/</guid>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 30 evidence observations from 1 sources on 2026-07-26.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-25</title>
      <link>https://benchmark-radar.org/blog/2026-07-25/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-25/</guid>
      <pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 32 evidence observations from 1 sources on 2026-07-25.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-24</title>
      <link>https://benchmark-radar.org/blog/2026-07-24/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-24/</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 32 evidence observations from 1 sources on 2026-07-24.</description>
    </item>
    <item>
      <title>Benchmark evidence summary: 2026-07-23</title>
      <link>https://benchmark-radar.org/blog/2026-07-23/</link>
      <guid isPermaLink="true">https://benchmark-radar.org/blog/2026-07-23/</guid>
      <pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate>
      <description>Benchmark Radar collected 20 evidence observations from 1 sources on 2026-07-23.</description>
    </item>
  </channel>
</rss>