What benchmarks, datasets, or evaluation methods did the radar first see today?
high confidenceNot enough evidence
Notable first-seen releases include PCFBench, LongPIBench, SkillSafetyBench, LoopArena, GMA, LongDS-Bench, WhatIfBench with PRISM, XHotpotQA, and GenIaC-SecBench. EASEL and ElephantBench were newly discovered by the radar rather than newly released.
The arrivals cover carbon-footprint workflows, prompt injection, agent safety, coding control, mobile assistants, long-running data analysis, counterfactual explanations, multilingual reasoning, infrastructure security, visual tool use, and disputed factual accounts. Several introduce datasets alongside evaluation procedures.
Takeaway: This captured feed shows broad evaluation coverage rather than one dominant subject. Treat the named artifacts as highlights from the first-observed pool, not as a complete inventory or a representative picture of AI research.
Another reading: First observed by the radar does not mean created today. The supplied evidence is a selected subset of the first-observed pool, some records have sparse descriptions, and unavailable search sources could hide duplicates or earlier appearances.
- S018 artifacts first observed by the radar today: 219
Which of today's arrivals document how they score an answer?
medium confidenceNot enough evidence
Clear examples include the network-topology framework, which compares predicted nodes and links with references and tests connectivity; WhatIfBench, whose PRISM method structures free-form explanations; the insider benchmark, which flags disagreement between independent logs; the datasheet benchmark, which combines source fidelity with tool traces; and GenIaC-SecBench, which applies common policy scanners to generated and human code.
These artifacts expose at least the basis of judgment: reference matching, structured analysis of explanations, log disagreement, checking extracted claims against sources, or automated security scans. That is more transparent than merely reporting a leaderboard result.
Takeaway: Use these as arrivals with documented scoring mechanisms, while checking their full papers or repositories before reproduction. The registry does not provide a count of all first-seen artifacts with explicit scoring documentation.
Another reading: The summaries often omit aggregation rules, thresholds, or implementation details. The financial reasoning arrival criticizes exact matching but its supplied text truncates the replacement method, so it cannot be confidently included among the clearly documented cases.
- S018 artifacts first observed by the radar today: 219