Benchmark

LongJudgeBench

What is LongJudgeBench?

LongJudgeBench evaluates LLM-as-a-judge performance on long-form outputs across six datasets covering pointwise, pairwise, and listwise protocols, with…

Released
2026-06-01
Evaluates
General AI, Language & Knowledge, cs.CL
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

LongJudgeBench paper and code

No reported scores are on record for this benchmark yet.

Which sources cite LongJudgeBench?

Related General AI, Language & Knowledge, cs.CL benchmarks