What benchmarks, datasets, or evaluation methods did the radar first see today?
medium confidenceNot enough evidence
In this keyword-filtered feed, first-seen artifacts covered workplace web agents, financial diligence, physics-aware simulation, long-manual reasoning, medical fact verification, agent unlearning, benchmark compression, judge-bias diagnosis, tool-failure recovery, graph reasoning, and synthetic-data contamination detection.
Representative releases were KNOWS, DiligenceProv, PhysCodeBench, Tasks over Application Manuals, MedSNIP-Bench, K-Bench, ZipBench, TraceJudgeBench, ParaRecover, Graph Theory Bench, and SynthSentry. Other arrivals introduced paired-view cell-image evaluation and leakage-controlled antibody testing.
Takeaway: Today’s first sightings span both new task collections and evaluation methods that inspect intermediate behavior, physical correctness, evidence use, leakage, or measurement bias rather than relying solely on final-task success.
Another reading: This is not an exhaustive inventory because the supplied evidence packet contains only selected captured records. Keyword tagging also produces false positives: the AIREP mirror explicitly says it is neither a training dataset nor a benchmark despite receiving those tags.
- S021 artifacts first observed by the radar today: 225
Which of today's arrivals document how they score an answer?
medium confidenceNot enough evidence
The clearest released arrivals with stated scoring logic are KNOWS, DiligenceProv, PhysCodeBench, K-Bench, and the land-use relevance benchmark. An updated CORE-LLM-Bench record also describes symbolically derived ground-truth answers and explanation structures.
KNOWS supplies a structured rubric for the produced artifact. DiligenceProv says answers are scored in atomic parts. PhysCodeBench checks execution, appearance, conservation rules, and expert assertions. K-Bench marks leakage when a secret appears in any exposed agent channel. The land-use benchmark recomputes classification measures from published predictions.
Takeaway: These records expose at least one concrete grading mechanism—a rubric, answer decomposition, executable checks, a leakage decision rule, published prediction metrics, or symbolic ground truth—rather than merely stating that an evaluation occurred.
Another reading: The evidence is insufficient for a complete list, and several summaries prove only that scoring exists. DiligenceProv’s description is truncated, while KNOWS and CORE-LLM-Bench do not expose full weighting, aggregation, or adjudication details in the supplied text.
- S021 artifacts first observed by the radar today: 225