Benchmark

AlpacaEval 2.0

AlpacaEval 2.0 is a length-controlled automatic evaluator for instruction-following language models that uses GPT-4 Turbo to assess model responses…

Modality
text
Categories
reasoning, general, creativity, writing
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V2.5DeepSeek0.5052024-05-08evidence
Granite 3.3 8B BaseIBM0.62682025-04-16evidence
Granite 3.3 8B InstructIBM0.62682025-04-16evidence
IBM Granite 4.0 Tiny PreviewIBM0.35162025-05-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.