What benchmarks, datasets, or evaluation methods did the radar first see today?
medium confidenceNot enough evidence
The captured feed first observed the artifact count reported by the registry. Clear releases included OenoBench, OmniHandwritingOCR, FinSkillBench, SWE-bench Science, Thinkingbox, MemFuseBench, FM-Bench, and ContractScrub [E001, E004, E005, E010, E011, E013, E014, E027]. Dataset arrivals included the Missing-Premise Benchmark and a Korean retrieval-and-answering benchmark [E019, E022]. Evaluation-method arrivals included a lifecycle framework for model-based judges and work on measuring speech-benchmark optimization [E015, E064].
The arrivals cover specialized tests for wine knowledge, handwriting, finance agents, scientific coding, business workflows, memory, sports management, and legal review [E001, E004, E005, E010, E011, E013, E014, E027]. They also include reusable question collections and methods for maintaining automated judges or studying benchmark optimization [E019, E022, E015, E064].
Takeaway: Use this as a discovery list from the keyword-filtered captured feed, not as a complete launch list for the AI field. The supplied evidence contains examples but not the complete artifact-level list corresponding to the registry’s first-observed total.
Another reading: First observed does not necessarily mean released today. The Taiwan legal benchmark and safety-evaluation runs were published earlier than today but entered this captured evidence packet now [E020, E021]. Updated artifacts can likewise be newly tracked without being new releases. A complete artifact-level export is missing.
- S016 artifacts first observed by the radar today: 123
Which of today's arrivals document how they score an answer?
medium confidenceNot enough evidence
The clearest disclosures are FinSkillBench, which uses hidden expected results and task-specific verifiers [E005]; Thinkingbox, which evaluates the resulting backend state [E011]; FM-Bench, which uses a deterministic engine rather than a model judge or human rater [E014]; the Korean retrieval-and-answering benchmark, which separately measures retrieval and generated-answer quality [E022]; and an updated recommendation project that reports established ranking metrics [E052].
These artifacts explain what is checked after a system responds. FinSkillBench compares results through a checker tailored to each task [E005]. Thinkingbox inspects whether the workflow left the system in the correct final state [E011]. FM-Bench computes a final result inside its simulator [E014]. The Korean benchmark scores finding the right source and producing the answer separately [E022]. The recommendation project evaluates whether useful items appear near the top [E052].
Takeaway: Within this captured feed, documented scoring ranges from exact or task-specific verification to checking system state, deterministic simulation, retrieval-and-answer metrics, and ranked-result metrics [E005, E011, E014, E022, E052]. The packet does not support an exhaustive inventory of every arrival with a scoring specification.
Another reading: The evidence consists of shortened summaries, so linked documentation may specify scoring for additional arrivals. Conversely, some summaries describe ground truth, auditing, or evaluation goals without giving a complete scoring rule; OenoBench is an example [E001]. Full benchmark cards or papers are missing for an exhaustive determination.
- S016 artifacts first observed by the radar today: 123