What benchmarks, datasets, or evaluation methods did the radar first see today?
Today’s first sightings included released benchmarks for contamination-resistant language evaluation, agent-memory poisoning, runtime reasoning, surgical-video understanding, visual grounding, kinship generation, and mental-health responses. Released datasets covered forecasting, raw-video retrieval, and social attitudes. Separately, the radar discovered work on calibration and dynamic repository representations and first encountered an update to structured chest-scan reasoning evaluation.
The clearest common thread is making evaluation harder to game or easier to verify: use fresh text, derive answers from program execution, test consistency across related labels, vary prompts and environments, or check claims against cited evidence. Other arrivals broaden coverage into surgery, agriculture, historical documents, forecasting, raw video, and mental health. These are first sightings in this radar, not claims of first publication worldwide.
Takeaway: Treat today as a diverse intake rather than a field-wide shift. Particularly actionable sightings provide reproducible data, executable references, or explicit protocols, including Uncheatable Eval, SWE-Flux, AgroBench, SurgHiBench, and deterministic environmental-agent tasks. Because the feed lacks a certified comparison window, it cannot establish that these themes are becoming more common.
Another reading: The evidence packet does not contain every artifact first observed today, so this is a highlight set rather than a complete inventory. Some apparent arrivals were discoveries or updates rather than new releases, and collection changes or unavailable sources could materially alter the mix. Radar novelty should therefore not be read as field novelty.
- S022 artifacts first observed by the radar today: 196