Benchmark

Legal Agent Benchmark

The Legal Agent Benchmark (LAB) is Harvey's open-source benchmark for evaluating AI agents on complex, long-horizon legal work. Tasks are scored under an…

Modality
text
Categories
legal, reasoning, agents
Openness
unknown
Source
llm_stats
Reported scores
13

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.1332026-06-09evidence
Claude Opus 4.6Anthropic0.0422026-02-05evidence
Claude Opus 4.7Anthropic0.0712026-04-16evidence
Claude Opus 5Anthropic0.1172026-07-24evidence
Claude Sonnet 4.6Anthropic0.0542026-02-17evidence
Claude Sonnet 5Anthropic0.0582026-06-30evidence
Gemini 3 FlashGoogle0.02025-12-17evidence
Gemini 3.1 Flash-LiteGoogle0.02026-03-03evidence
Gemini 3.1 ProGoogle0.02026-02-19evidence
Gemini 3.5 FlashGoogle0.0082026-05-19evidence
GPT-5.4OpenAI0.0042026-03-05evidence
GPT-5.4 miniOpenAI0.02026-03-17evidence
GPT-5.5OpenAI0.0212026-04-23evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.