What benchmarks, datasets, or evaluation methods did the radar first see today?
high confidence
Notable first-seen releases covered agent security, scientific verification, forecasting, membership inference, production coding, user-interest grounding, dialogue compression, and evaluation of open-ended research processes.
Examples include EvoRiskBench for workspace-agent risks, MOF-VERIFY for materials hypotheses, LEAF for leakage-aware forecasting, OLMo-Detect for training-membership tests, OpenGameEval for game-development agents, GISTBench for user-interest verification, TPBench for dialogue compression, and Open-Endedness Bench for judging research behavior from execution records.
Takeaway: In this keyword-filtered feed, the first-seen set is unusually broad and includes both task-specific datasets and methods that inspect how agents work, not merely whether they reach a final result. These were first observed by the radar today, which does not establish that they were first released to the field today.
Another reading: Radar novelty is not field novelty. Some records may be mirrors, later indexing events, or versions published earlier, and unavailable sources could hide prior sightings. The evidence packet also highlights selected arrivals rather than describing every first-seen artifact.
- S023 artifacts first observed by the radar today: 314
Which of today's arrivals document how they score an answer?
medium confidence
The clearest documented scoring approaches are compile verdicts, executable checks, field-level answer matching, deterministic rules, groundedness and specificity metrics, rubric scoring, and process evaluation from execution logs.
The MQL compile set records whether generated code compiles. OpenGameEval runs checks on edited scenes and simulated play. The invoice benchmark compares extracted fields with answer keys. PanduGizi supplies deterministic scoring rules. GISTBench measures supported and distinctive user interests. LICA uses a writing rubric. Open-Endedness Bench judges hypothesis formation, testing, and revision from agent records.
Takeaway: These arrivals expose more than a leaderboard result: they identify what evidence earns credit. The strongest machine-checkable examples are compilation, executable environment checks, and field-by-field comparison; the rubric and process-based methods offer richer criteria but require more judgment.
Another reading: The summaries do not consistently provide weighting, thresholds, aggregation rules, or complete implementation details, so documented criteria do not guarantee reproducible scoring. The MQL record also says generated completions are withheld, limiting independent inspection of scored outputs.