BenchCAD · leaderboard

Vision2Code

all four tasks

Same parts, matched tasks

Model performance across four matched programmatic-CAD tasks. Scoring is execution-grounded and objective — geometry by IoU, QA by ratio accuracy. Numbers are re-graded from submitted predictions, never self-reported. Pick a task to load its table below.

reproduce this task — full split
Loading…
frontier (proprietary) open weights control / baseline

* Self-reported — voxel IoU only, not re-graded. Anthropic: Mythos 5 / Mythos Preview / Opus 4.8 on the full set (Fable 5 / Mythos 5 system card), Sonnet 5 on a random 1,000-file subset, no tools (Sonnet 5 system card). OpenAI: GPT-5.5 and GPT-5.6 Sol / Terra / Luna, no tools (GPT-5.6 launch table).

To add a model to the this leaderboard, run the command above and open a Model result submission issue on the code repo with your submission.jsonl. Submissions are re-graded before listing.