Benchmark

O.W.L. & N.E.W.T. Bench

What is O.W.L. & N.E.W.T. Bench?

Deliberately contaminated LoRA experiment with a 75-question Harry Potter evaluation set and training code that reproduces the leakage demonstration.

Released
2026-09-08
Evaluates
Language & Knowledge, Software & AI Compute, cs.AI
Openness
unknown
Importer
Claire Radar
Review status
ai-name-audit-deferred
Reported scores
0

Source provenance

O.W.L. & N.E.W.T. Bench code and dataset

No reported scores are on record for this benchmark yet.

Which sources cite O.W.L. & N.E.W.T. Bench?

Related Language & Knowledge, Software & AI Compute, cs.AI benchmarks