What benchmarks, datasets, or evaluation methods did the radar first see today?
Within this captured feed, newly observed releases covered symbolic planning, multilingual factual knowledge, flood-response vision, agent runtime faults, temporal reasoning, harmful-output profiling, explanation assessment, agentic video, implicit timing, research-idea recovery, and aviation copilots.
Examples include PDDLCoder for verified planning specifications, FloodReasonBench for flood-scene segmentation, AGENTCHAOSBENCH for locating agent failures, TRACE for checking temporal reasoning, HarmProfile for characterizing harmful outputs, VideoGAIA for tool-assisted video tasks, Chronocooked for timing decisions, Reconstruction for recovering research ideas, and AeroCopilotBench for cockpit tasks. The feed also surfaced an Indic-language factual benchmark and CBX-Bench for assessing model explanations.
Takeaway: The arrivals span both new test material and new evaluation procedures rather than one dominant task. This is a discovery summary for the keyword-filtered radar, not a representative view of the field or a complete inventory from the selected evidence packet.
Another reading: “First observed” means first discovered by this radar, not necessarily newly created or newly available. The evidence packet is selected rather than exhaustive, and none of today’s tracked artifacts was independently sighted by multiple connectors, limiting confirmation.
- S016 artifacts first observed by the radar today: 136
- S018 tracked artifacts today seen by more than one connector: 0