Benchmark

BabelJudge Benchmark

What is BabelJudge?

BabelJudge audits LLM-as-a-judge reliability across languages and agent trajectories, measuring position bias, verbosity bias, order inconsistency, and…

Released
2026-06-21
Evaluates
Agents, General AI, Safety & Trustworthiness
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

BabelJudge paper and code

No reported scores are on record for this benchmark yet.

Which sources cite BabelJudge?

Related Agents, General AI, Safety & Trustworthiness benchmarks