Benchmark
O.W.L. & N.E.W.T. Bench
What is O.W.L. & N.E.W.T. Bench?
Deliberately contaminated LoRA experiment with a 75-question Harry Potter evaluation set and training code that reproduces the leakage demonstration.
- Released
- 2026-09-08
- Evaluates
- Language & Knowledge, Software & AI Compute, cs.AI
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- ai-name-audit-deferred
- Reported scores
- 0
Source provenance
- Original evidence https://github.com/VSBDev/little-hermione
O.W.L. & N.E.W.T. Bench code and dataset
No reported scores are on record for this benchmark yet.