Benchmark

PluginEval

What is PluginEval?

PluginEval evaluates tool routing in LLMs via three decision types (missed, spurious, parameter errors) across difficulty levels, using deterministic…

Released
2026-08-09
Evaluates
General AI, Language & Knowledge, cs.AI
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

PluginEval paper

No reported scores are on record for this benchmark yet.

Which sources cite PluginEval?

Related General AI, Language & Knowledge, cs.AI benchmarks