What benchmarks, datasets, or evaluation methods did the radar first see today?
medium confidenceNot enough evidence
The captured feed first observed a large set of artifacts today, with prominent arrivals spanning benchmark leakage, adversarial robustness, remote sensing, natural conversation, infrastructure vision, grounded change descriptions, enterprise agents, legal-agent hallucinations, and clinical stopping decisions.
Examples include CGMIA for detecting code-benchmark contamination, a severity-calibrated federated-learning benchmark, WD-CD, Candor-LR, Infra-Bench CLS, Spot-the-Shift, Era by Eon, LexAgentHallu, and Cros. These test different problems rather than forming one comparable leaderboard.
Takeaway: Today’s first sightings emphasize specialized evaluation settings and methods that test more than final-answer accuracy, including leakage checks, controlled attacks, spatial grounding, exact answer keys, trajectory failures, and stopping reliability. This is a high-signal sample, not a complete inventory of every first-seen artifact.
Another reading: Several arrivals may be experimental, pilot, repackaged, or documentation-only resources rather than mature new benchmarks. AhiskaAI labels itself experimental, the CLSG release is a pilot, and EngIntervene says it republishes a frozen benchmark rather than introducing a new collection.
- S022 artifacts first observed by the radar today: 236
Which of today's arrivals document how they score an answer?
medium confidenceNot enough evidence
The clearest documented examples are Era by Eon’s computed answer keys, HalluDetector’s explicit contradiction-rule families, AhiskaAI’s quality, relevance, and correctness criteria, Smart-Eval’s rubric scoring, and the deindexing benchmark’s named score dimensions.
Other arrivals discuss scoring design without fully exposing it in the packet: the agent-benchmark audit centers ground-truth scoring, Spot-the-Shift proposes a dedicated evaluation protocol, and the black-box red-teaming framework uses human-validated model judges. EngIntervene explicitly says its scoring references are stored separately.
Takeaway: The strongest answer-scoring documentation in the supplied summaries favors reproducible keys, explicit rules, or named rubric dimensions. However, the summaries do not show enough implementation detail to verify thresholds, aggregation, judge prompts, tie handling, or whether every scoring component is publicly available.
Another reading: Mentioning criteria or judges is not the same as documenting a reproducible scoring procedure. AhiskaAI and Smart-Eval name scoring dimensions, while EngIntervene withholds references from the released input copy; full artifacts would be needed to determine whether their scoring can actually be reproduced.