Benchmark

Evaluation Agent

What is Evaluation Agent?

A promptable evaluation framework rather than a fixed test set: it plans its own queries per model, so two runs need not ask the same questions. It does report…

Released
2024-12-10
Evaluates
multimodal
Openness
unknown
Importer
Model reports
Reported scores
0

Source provenance

Evaluation Agent code

No reported scores are on record for this benchmark yet.

Related multimodal benchmarks