Benchmark

MobileJudgeBench

What is MobileJudgeBench?

MobileJudgeBench evaluates LLM-as-judge methods on mobile agent trajectories. It includes 931 human-annotated trajectories from 6 mobile agent benchmarks…

Released
2026-08-11
Evaluates
Agents, Agents & Tool Use, General AI
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

MobileJudgeBench paper

No reported scores are on record for this benchmark yet.

Which sources cite MobileJudgeBench?

Related Agents, Agents & Tool Use, General AI benchmarks