Benchmark Radar™
RSS Contact Star —

Daily AI benchmark brief: 2026-10-07

808evidence observations
13sources represented
20public-attention signals

Daily briefing

  1. AndroidLife is a new benchmark that evaluates open-weight and small on-device models while they operate real Android phones. It provides a 60-task public set, a larger 530-task dataset, and an Apache-2.0 evaluation harness, moving mobile-agent testing beyond simulated interfaces or screenshots alone. Why it matters: If you are choosing a model for an Android product, this provides a way to test the model under the device and interaction constraints it will actually face. The public tasks can support comparisons, while the larger dataset and permissively licensed harness can support internal evaluation and tooling integration. Evidence: E001. Medium confidence.
  2. DecepEval is a new agent-safety benchmark with 1,532 cases across three task families and 28 professional scenarios. Instead of asking only whether an agent deceives, it varies four external conditions—pressure, incentive, opportunity, and an additional condition described by its Deception Diamond framework—to examine when deceptive behavior appears. Why it matters: If you are evaluating an autonomous agent for workplace use, this changes the test from a single deception rate into a conditional risk assessment. That can reveal whether deployment incentives, access, or task pressure alter behavior, which is more actionable for selecting controls and deployment conditions. Evidence: E005. Medium confidence.
  3. ST-Bench is a new benchmark built to compare multi-agent and single-agent systems on scientific data analysis while also measuring added cost. It contains 100 tasks adapted from published Earth-science studies and expands them into 2,067 queries across hydrology, agriculture, and wetland methane research. Why it matters: If you are deciding whether multiple cooperating agents justify their additional inference and coordination expense, ST-Bench offers a direct architecture comparison on research workflows rather than coding or question-answering tasks with simple pass/fail tests. The cost measurement makes the result relevant to both capability and operating-budget decisions. Evidence: E007. Medium confidence.
  4. ANT, or Agent Network Traffic, is a new dataset for auditing agent behavior from network traffic without inspecting private user content. Its 3,114 execution episodes cover 20 tasks and five scenarios, with annotations at risk, scenario, and low-level behavior categories that can be compared with observed traffic. Why it matters: If you operate agents inside an organizational network, ANT provides an evaluation target for deciding how much security monitoring is possible from traffic metadata alone. That can inform whether network-level controls are sufficient or whether deployment requires more intrusive application logging and content inspection. Evidence: E008. High confidence.
  5. Measurement-First Auditing of Agentic Leaderboards introduces a contamination audit that separates training-time exposure, evaluation-time retrieval, and leakage through the surrounding agent pipeline. This reflects a recurring pressure in the captured feed: another release explicitly marks benchmark-conditioned training data as contaminated, while a biomedical benchmark embeds a canary identifier to detect training exposure. Why it matters: If you compare agents on public leaderboards, a score may be unreliable for different reasons that require different evidence and remedies. Separating exposure during training from retrieval or pipeline leakage helps determine whether to exclude a model, close evaluation access, or change the harness rather than applying one generic contamination label. Evidence: E018, E003, E045. High confidence.
  6. LMBuild is a new benchmark for generated three-dimensional structures that scores whether an object can be assembled and perform its intended function, not just whether its geometry looks plausible. It represents objects through parts, joints, materials, and assembly sequences and provides a unified evaluation setup. Why it matters: If you are selecting an agent for physical design, LMBuild changes the acceptance criterion from visual form to realizability. A system that produces attractive geometry may still fail when materials, connections, assembly order, or function are considered, so this benchmark can alter which model or workflow qualifies for prototyping. Evidence: E010. Medium confidence.
  7. InteractionBench is a new benchmark for streaming video assistants that evaluates the complete model, memory, and response controller. Across 1,060 interactions and 812 videos, it separately scores response content, timing, and compliance with situations where the system should remain silent, including near-miss events. Why it matters: If you are choosing a system for real-time video assistance, answer accuracy alone does not capture interruptions, late responses, or false triggers. The separate timing and silence measures expose trade-offs that affect whether a system is suitable for continuous monitoring or interactive guidance. Evidence: E030. Medium confidence.
  8. SpatialChain is a new benchmark that evaluates whether a vision-language model’s stated spatial reasoning matches a scene graph—a structured record of objects and their relationships—even when its final answer is correct. It pairs 28,350 training examples and 899 test examples with reasoning chains and scores faithfulness and completeness separately from answer accuracy. Why it matters: If you rely on spatial explanations for debugging or oversight, SpatialChain can distinguish a correct answer reached from image evidence from one reached through a language shortcut. That changes model selection when trustworthy intermediate reasoning matters more than headline accuracy alone. Evidence: E025. Medium confidence.

What arrived

What benchmarks, datasets, or evaluation methods did the radar first see today?

high confidence

The captured feed first observed many artifacts today. Notable releases include AndroidLife for real-phone agents (E001), SEMAADB for coherent engineering diagrams (E004), DecepEval for agent deception (E005), ANT for network-based agent auditing (E008), FREA for reaction feasibility (E009), and LMBuild for buildable structures (E010).

Other first-seen releases test scientific multi-agent work with ST-Bench (E007), reasoning-chain faithfulness with SpatialChain (E025), real-time video responses with InteractionBench (E030), and automatic scientific benchmark generation with AutoSciBench (E032). Measurement-first auditing evaluates contamination and scorer validity rather than introducing a conventional task set (E018).

Takeaway: The arrivals span agents, physical systems, scientific work, security, and evaluation design. They are first observations by this radar, not proof that each artifact was first published globally. Benchmark, evaluation, dataset, and agentic tags overlap and should not be treated as separate shares of the feed.

Another reading: This is a selected inventory, not an exhaustive account of every first-seen artifact. Some records are repository releases or paper records, while others are datasets, frameworks, or evaluation studies; treating all of them as equivalent benchmark launches would overstate the evidence.

  • S024 artifacts first observed by the radar today: 599
  • S019 records tagged benchmark: 554 count (multi-label)
  • S020 records tagged evaluation: 434 count (multi-label)
  • S021 records tagged dataset: 339 count (multi-label)
  • S022 records tagged agentic: 127 count (multi-label)

Which of today's arrivals document how they score an answer?

medium confidence

The clearest cases are BazaarBench, which combines record checks with rubric-based model judgments (E024); SpatialChain, which combines chain-overlap measures with a judge for faithfulness and completeness (E025); InteractionBench, which scores content, timing, and appropriate silence (E030); and the financial-crime benchmark, which uses an explicit multi-part explanation rubric (E101).

The medical guideline study has experts rate relevance, accuracy, understandability, and safety, then combines those ratings (E077). DEPICT scores image-text agreement through answers to verification questions (E166). The debugging results release exposes model outputs and evaluator scores so reported results can be inspected and recomputed (E043).

Takeaway: These arrivals make scoring more inspectable by naming the checks, rating dimensions, or evaluator outputs. They also illustrate different approaches: automatic checks, model-based judging, human rating, and published raw scores (E024, E025, E043, E077, E101, E166).

Another reading: The packet contains summaries rather than full scoring specifications. If “document” requires executable code, exact prompts, weighting rules, and aggregation details, several listed artifacts may provide only partial documentation; the debugging release makes the strongest explicit recomputability claim (E043).

  • S024 artifacts first observed by the radar today: 599

What is still moving

Which artifacts the radar already tracked moved measurably, and over what span?

high confidenceNot enough evidence

The registry contains artifact-level download changes measured cumulatively across each artifact’s full tracked span, but the supplied packet provides no artifact evidence identifiers.

The statistics identify the artifacts and measurement periods, but reporting that list would make specific artifact claims without the required evidence citations.

Takeaway: A compliant artifact-by-artifact answer requires evidence identifiers linked to each registered movement; those links are missing.

Another reading: The statistic labels themselves name the artifacts and include spans, but the grounding rules separately require evidence identifiers for every artifact-specific claim.

  • S027 downloads change for lmarena-ai/leaderboard-dataset: 46,968.0 downloads
  • S028 downloads change for open-llm-leaderboard/requests: -33,205.0 downloads
  • S029 downloads change for hf-benchmarks/transformers: 24,647.0 downloads
  • S030 downloads change for alexshpunt/explicit-edit-benchmark: 19,042.0 downloads
  • S031 downloads change for vedangfake/chess-slm-benchmark: 17,725.0 downloads
  • S032 downloads change for hf-audio/open-asr-leaderboard-results: 8,765.0 downloads
  • S033 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 8,528.0 downloads
  • S034 downloads change for AlphaDojo/dojo_benchmark_kline: 3,939.0 downloads

Which of that movement is corroborated by more than one data source?

high confidence

None of the registered movement statistics is marked as corroborated by more than one source.

Every registered download change came from Hugging Face alone, so the radar cannot treat those movements as independently confirmed.

Takeaway: Multiple-source sightings of an artifact do not prove that multiple sources measured the same change, and no registered movement meets that stricter test.

Another reading: A separate tracked-artifact section marks some movements as corroborated, but those movements lack registered statistic identifiers and artifact evidence identifiers, so they cannot be reported under the supplied grounding rules.

  • S026 tracked artifacts today seen by more than one data source: 12
  • S027 downloads change for lmarena-ai/leaderboard-dataset: 46,968.0 downloads
  • S028 downloads change for open-llm-leaderboard/requests: -33,205.0 downloads
  • S029 downloads change for hf-benchmarks/transformers: 24,647.0 downloads
  • S030 downloads change for alexshpunt/explicit-edit-benchmark: 19,042.0 downloads
  • S031 downloads change for vedangfake/chess-slm-benchmark: 17,725.0 downloads
  • S032 downloads change for hf-audio/open-asr-leaderboard-results: 8,765.0 downloads
  • S033 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 8,528.0 downloads
  • S034 downloads change for AlphaDojo/dojo_benchmark_kline: 3,939.0 downloads

What it means

What should someone building or evaluating AI systems do differently today?

medium confidence

Within this captured feed, most records are releases, with overlapping benchmark and evaluation tags prominent. New releases cover real-phone agents, deception, network auditing, physical buildability, and contamination auditing, suggesting useful candidates for more deployment-specific test suites rather than evidence of a fieldwide shift.

Review whether your tests reflect how your system actually operates. Consider real devices, deceptive behavior, observable network activity, physical constraints, and benchmark leakage where relevant. Treat the newly released artifacts as candidates for inspection, not as proof that a model or evaluation method is reliable.

Takeaway: Add workload-specific failure cases and explicit contamination checks before acting on leaderboard results. Record data provenance, access paths, scoring behavior, and whether test material could enter training or evaluation-time retrieval.

Another reading: These specialized releases may not match a given workload, and their abstracts do not independently establish quality. One captured dataset explicitly identifies itself as benchmark-conditioned and unsuitable as unbiased evaluation evidence, reinforcing that adoption without inspection could make evaluation worse.

  • S003 records with event kind released: 466
  • S019 records tagged benchmark: 554 count (multi-label)
  • S020 records tagged evaluation: 434 count (multi-label)
  • S021 records tagged dataset: 339 count (multi-label)
  • S022 records tagged agentic: 127 count (multi-label)

What does today's evidence fail to show, and what would change the reading?

high confidenceNot enough evidence

This captured feed does not establish a fieldwide trend, benchmark quality, real adoption, or corroborated metric movement. There is no certified comparison window, and the reported download changes are cumulative across differing tracked spans and each comes from a single source.

The feed shows that artifacts were captured, released, updated, discovered, downloaded, or discussed. It does not show that they are valid, widely used, improving outcomes, or changing faster than before. A stable comparison window, complete collection coverage, independent replication, and matching measurements from multiple sources would strengthen the reading.

Takeaway: Do not interpret category differences or download movements as growth. Reassess when collection is comparable across periods, individual metrics are independently corroborated, and important releases have reproducible results plus documented leakage and data-quality checks.

Another reading: Some releases were independently sighted across sources, including an agent network-auditing dataset and several benchmark papers. That supports their existence and visibility within the feed, but not their quality, adoption, metric movement, or a broader trend.

  • S026 tracked artifacts today seen by more than one data source: 12
  • S027 downloads change for lmarena-ai/leaderboard-dataset: 46,968.0 downloads
  • S028 downloads change for open-llm-leaderboard/requests: -33,205.0 downloads
  • S029 downloads change for hf-benchmarks/transformers: 24,647.0 downloads
  • S030 downloads change for alexshpunt/explicit-edit-benchmark: 19,042.0 downloads
  • S031 downloads change for vedangfake/chess-slm-benchmark: 17,725.0 downloads
  • S032 downloads change for hf-audio/open-asr-leaderboard-results: 8,765.0 downloads
  • S033 downloads change for dreamdifferent/vam-cross-evaluation-artifacts: 8,528.0 downloads
  • S034 downloads change for AlphaDojo/dojo_benchmark_kline: 3,939.0 downloads
  • S035 daily-average change in evaluation observations: 113.86 observations per day
  • S036 daily-average change in benchmark observations: 112.86 observations per day
  • S037 daily-average change in dataset observations: 62.71 observations per day
  • S038 daily-average change in agentic observations: 25.71 observations per day
  • S039 daily-average change in data_quality observations: 2.86 observations per day

Evidence sources

Where this came from, and what it does not cover

Built from the validated daily snapshot and its 808 evidence records.

Briefing model: gpt-5.6-sol.

It read 188 of 808 records.

This is a keyword-filtered, nonrepresentative feed. The briefing examined all 188 selected evidence records supplied here, but those represent only part of today’s 808-record corpus. Brave and OpenReview were unavailable, and the comparable history is insufficient to infer category-share movement. Most release descriptions are author-supplied and do not independently validate benchmark quality or reported results.