Benchmark

UTILMEM Benchmark

What is UTILMEM?

Evaluates long-term conversational memory systems on evidence utilization through 1,717 instances across five domains, testing reasoning, implicit retrieval…

Released
2026-08-31
Evaluates
General AI, Language & Knowledge, Reasoning
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

UTILMEM paper and code

No reported scores are on record for this benchmark yet.

Which sources cite UTILMEM?

Related General AI, Language & Knowledge, Reasoning benchmarks