What benchmarks, datasets, or evaluation methods did the radar first see today?
Notable first-seen releases included KaliBench for cybersecurity commands, Video-Index for auditing video shortcuts, VisionQ for criterion-based visual judging, TabJoinBench for table discovery, and SHAMS for Levantine Arabic pronunciation. New evaluation methods included layered contamination checks, a deployment-focused tracking protocol, and variance decomposition for agent leaderboards. Argo-Bench, AutoDataBench, and CUA-SWE arrived as discoveries rather than new releases.
The selected arrivals test whether systems can produce usable commands, genuinely understand video, judge images for the stated reason, find compatible tables, and handle speech varieties. Other work checks whether test data leaked into training, whether tracking systems recover from failures, and whether agent rankings reflect models, tasks, or surrounding software.
Takeaway: Within this captured feed, the clearest first-seen work emphasizes checks tied to observable behavior: executable outputs, explicit visual criteria, contamination audits, shared detector inputs, and separated sources of leaderboard uncertainty. These are notable examples from the supplied evidence, not a complete inventory of every first-observed artifact.
Another reading: First observed means new to this radar, not necessarily newly published or newly released. Discovery records may describe work published earlier, while the evidence packet covers only a selected subset of captured artifacts. Keyword matching can also overstate maturity; one dataset page explicitly says it is not a complete benchmark release.
- S023 artifacts first observed by the radar today: 364