Benchmark

CHIARO Benchmark

What is CHIARO?

Human-annotated benchmark of 1,000 sentences where each scene elicits opposing emotions in two agents, evaluated via macro-F1 for both LLMs and classifiers.

Released
2026-09-03
Evaluates
General AI, Language & Knowledge, cs.CL
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

CHIARO paper and code

No reported scores are on record for this benchmark yet.

Which sources cite CHIARO?

Related General AI, Language & Knowledge, cs.CL benchmarks