Benchmark
COLLIE
COLLIE is a grammar-based framework for systematic construction of constrained text generation tasks. It allows specification of rich, compositional…
- Modality
- text
- Categories
- reasoning, language, writing
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 10
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| GPT-4.1 | OpenAI | 0.658 | 2025-04-14 | evidence |
| GPT-4.1 mini | OpenAI | 0.546 | 2025-04-14 | evidence |
| GPT-4.1 nano | OpenAI | 0.425 | 2025-04-14 | evidence |
| GPT-4.5 | OpenAI | 0.723 | 2025-02-27 | evidence |
| GPT-4o | OpenAI | 0.61 | 2024-08-06 | evidence |
| GPT-5 | OpenAI | 0.99 | 2025-08-07 | evidence |
| Mistral Medium 3.5 | Mistral | 0.958 | 2026-04-29 | evidence |
| Mistral Small 4 | Mistral | 0.629 | 2026-03-16 | evidence |
| o3 | OpenAI | 0.984 | 2025-04-16 | evidence |
| o3-mini | OpenAI | 0.987 | 2025-01-30 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.