Benchmark

KernelBench Hard

KernelBench Hard evaluates agentic GPU kernel optimization on the hardest problem set. Each question is scored by the agent's submitted operator TFLOPs…

Modality
text
Categories
agents, code, systems
Openness
unknown
Source
llm_stats
Reported scores
1

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
MiniMax M3MiniMax0.2882026-06-01evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.