Benchmark

BenchCAD (with Python tool)

BenchCAD variant evaluated with access to a Python tool for programmatic CAD reasoning.

Modality
multimodal
Categories
multimodal, reasoning, code, vision
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-5.6 LunaOpenAI0.7392026-07-09evidence
GPT-5.6 SolOpenAI0.8342026-07-09evidence
GPT-5.6 TerraOpenAI0.7822026-07-09evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.