BenchCAD · leaderboard

Vision2Code

all four tasks

Same parts, matched tasks

Model performance across four matched programmatic-CAD tasks. Scoring is execution-grounded and objective — geometry by IoU, QA by ratio accuracy. Numbers are re-graded from submitted predictions, never self-reported. Pick a task to load its table below.

reproduce this task — full split
Loading…
frontier (proprietary) open weights control / baseline

* Self-reported — voxel IoU only, not re-graded. Anthropic: Mythos 5 / Mythos Preview / Opus 4.8 on the full set (Fable 5 / Mythos 5 system card), Sonnet 5 on a random 1,000-file subset, no tools (Sonnet 5 system card), Fable 5.1 on a random 1,000-file subset (Fable 5.1 / Mythos 5.1 system card, fig. 8.14.2.A). OpenAI: GPT-5.5 and GPT-5.6 Sol / Terra / Luna, no tools (GPT-5.6 launch table).

IoU · tools — the agentic setting: the model gets a Python sandbox to render, measure and iterate before submitting. Anthropic and OpenAI figures are vendor self-reported voxel IoU, not re-graded. Anthropic — a random 1,000-file subset, not the same split as the IoU-score beside it: Mythos 5 0.650 / Mythos Preview 0.610 / Opus 4.8 0.518 (Fable 5 / Mythos 5 system card, fig. 8.16.4.B), Sonnet 5 0.373 (Sonnet 5 system card), Opus 5 0.821 (Opus 5 system card, fig. 8.12.2.A), Fable 5.1 0.843 (Fable 5.1 / Mythos 5.1 system card, fig. 8.14.2.A — max effort, adaptive thinking, averaged over five runs). OpenAI — GPT-5.5 and GPT-5.6 Sol / Terra / Luna with a Python tool (GPT-5.6 launch table); split and attempt budget unstated. SpaceXAI — Grok 4.6 0.8055 (xhigh) and Grok 4.5 0.7771 (high), our own agentic runs on our scorer, not vendor-reported. GPT-6 Astra 0.959 — reported in OpenAI's GPT-6 Astra launch coverage (launch post); the same table quotes Fable 5.1 0.843, Opus 5 0.821 and Fable 5 0.675, which match Anthropic's published with-tools figures, so the split and metric line up. Secondhand: OpenAI's page was unreachable when this row was added and no no-tools figure has been published.

To add a model to the this leaderboard, run the command above and open a Model result submission issue on the code repo with your submission.jsonl. Submissions are re-graded before listing.