Benchmark

SPIEval

What is SPIEval?

SPIEval is a human-curated benchmark for evaluating large language models as mobile assistants that retrieve and reason over personal information scattered…

Released
2026-08-11
Evaluates
General AI, Language & Knowledge, Reasoning
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

SPIEval paper and dataset

No reported scores are on record for this benchmark yet.

Which sources cite SPIEval?

Related General AI, Language & Knowledge, Reasoning benchmarks