In BenchCAD's image-to-code task1, a model gets four rendered views of a mechanical part and writes a CadQuery2 script that rebuilds it. It works in a harness loop, running code and reading the output over many rounds. Our standard harness is plain Python with CadQuery, and after each round it shows the model up to three renders of its current part.
We assumed those renders were essential for visual debugging. But when we ran the same model in mini-swe-agent3,4, which shows the target views only in the first turn and returns nothing but text after that, the score stayed flat and the API cost fell by about half.
The model's code explained why. It renders the part programmatically, computes the silhouette overlap with the target, and adjusts the geometry to push that number up, all in code. After the first turn, where it sees the target views, it never needs to look at an image again. A single scalar to optimise lets it make many fast, small iterations. On BenchCAD 2.0, which provides no rendering tool at all, GPT-6 Astra went as far as writing its own renderer and optimiser to run the same loop.
Seeing its own part does not help
Figure 1. One frontier model on the same 98 BenchCAD 1.0 parts in two harnesses, with the same tools and the same (high) reasoning effort.
Both harnesses let the model render its part the same way the target pictures were rendered; mini-swe-agent just never shows it the result. The scores differ by −0.002 (95% interval ±0.05). Switching the pictures off inside our own harness made no measurable difference either.
The model measures its part instead
We read the code from about 400 runs across the two harnesses. In about five of every six, the model wrote a script that renders its part and computes how much its silhouette overlaps a target picture, then edited the part and ran the comparison again. In almost every run that printed this overlap more than once, it went up, usually from about 0.8 to 0.98.
The model is hill climbing. It makes a change, keeps it if the number rises, and repeats. The number is an overlap of 2D outlines that stands in for the 3D answer, so we call it a 2D proxy.
Figure 2 shows a run of GPT-6 Astra on BenchCAD 2.0, where the task comes with no drawing tool.
Figure 2. GPT-6 Astra rebuilding a helical wind-turbine rotor from BenchCAD 2.0. It wrote a renderer for its own part and an optimiser that tunes the part's dimensions to fit the target silhouette. Its final part scores 0.907.
Other models do it too
We went through every run on 27 private BenchCAD 2.0 cases, for GPT-6 Astra and Claude Opus 5.5 at each reasoning-effort setting, and counted the runs that wrote their own drawing code.
Figure 3. Share of runs that write their own drawing code, by reasoning effort.
At low effort both models answer in a round or two and rarely draw. From high effort up, nine in ten runs or more write drawing code, and many use it to compare their part with the targets. Given room to think, both models settle on the same strategy.
Does climbing the proxy improve the 3D part?
If the 3D result comes from climbing the silhouette overlap, the model should do worse when its renders can no longer line up with the targets. On 100 BenchCAD 1.0 parts, we re-rendered three of the four target pictures 15° off, so that no part could match them pixel for pixel, and compared with a run that differed only in that.
Figure 4. The same 100 BenchCAD 1.0 parts, with the target pictures aligned to the model's renders and with three of them 15° off.
Without alignment the 3D score fell from 0.786 to 0.722, and the model needed about seven more rounds per part.
Where the climb stops
Most runs that rated their own match 0.95 or higher also scored above 0.9 in 3D, with a median of 0.97. Figure 5 shows four that did not.
Figure 5. Four BenchCAD 1.0 parts where the model rated its answer around 0.95 and the 3D score was lower, each next to the real part and cut open where they differ. We picked these to show where the silhouette score stops helping; most high-scoring runs don't look like this.
In all four, the details are wrong. The duct elbow is nearly solid where it should be hollow. The bellows has a thick wall where the real one is thin and folded. The venturi tube's bore runs straight instead of narrowing. The sprocket has fine gear teeth instead of sprocket teeth. Each of these barely changes the silhouette overlap but lowers the 3D score a lot.
Once the silhouette matched, the model's own score stopped telling it anything about the details, and its modelling of small features wasn't good enough to get them right without that signal. So it settled for a part that scored well on its proxy, a mild form of reward hacking5.
Experts also judge by appearance
We also checked whether people catch these differences. About 15 domain experts compared pairs of rebuilt parts against the original, about 10k votes in all. Each screen showed four rendered views of every part, with the 3D solids one click away.
Figure 6. How often each measure picks the same part as the expert (50% is chance). The orange bar compares images (2D); the blue bars compare the 3D solids. The dashed line is how often two experts agree.
Of the measures in Figure 6, expert choices agree most with DINOv26, an image model that scores how alike two pictures look (75%), well above the volume and surface measures CAD benchmarks usually report (63 to 66%). When the image score and the 3D score disagree about which part is better, experts side with the image score two times in three.
Image similarity can be fooled as well. A bevel gear with a second copy stuck to its side still scores 0.92 under DINOv2, about the same as the 0.93 it gives the same gear turned over.
What this means
BenchCAD 2.0's rebuild-from-pictures task is built against this kind of pixel matching. It gives no renderer and states only one camera position, (1, 1, 1); the other three views are taken from slightly different angles that are not disclosed. Lining up pixels alone isn't enough to score well.
Outline matching gets most parts right. The remaining gains are in fine detail such as wall thickness, bores and teeth, which takes a check more sensitive than outline overlap and better modelling of small features.
CAD benchmarks should score the solid, because image similarity can be fooled by a part that looks right, and expert review from pictures leans the same way.
Limits
- The scores, gaps and experiments come from one frontier model on BenchCAD 1.0's image-to-code task. For GPT-6 Astra and Claude Opus 5.5 we only counted what their code does.
- Figures 2 and 5 are chosen examples. Figure 2 is a single run, and Figure 5 shows low 3D scores on purpose.
Appendix A: Experiment details
The BenchCAD 1.0 results use about 300 image-to-code parts, each run once in Python tools at high reasoning effort with up to 100 rounds; the harness comparison and the camera test use subsets of about 100 parts. The 3D score is voxel IoU on a 64 × 64 × 64 grid. Intervals are 95% bootstrap intervals over parts. What each run did is read from its code and output.
Appendix B: The expert study
Experts saw the original part above two candidate reconstructions, each as four rendered views, with the 3D solids one click away, and picked the closer candidate or called it a tie. A small share of candidates were deliberately altered copies of the original, such as the doubled bevel gear. Figure 6 uses every vote where the expert picked a side, and counts agreement vote by vote: a measure agrees when it scores the expert's pick higher.
References
- Haozhe Zhang, Kaichen Liu, Miaomiao Chen, Lei Li, Shaojie Yang, Cheng Peng, Hanjie Chen. BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD. 2026. arXiv:2605.10865
- CadQuery. github.com/CadQuery/cadquery
- mini-swe-agent. SWE-agent team. github.com/SWE-agent/mini-swe-agent
- John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. arXiv:2405.15793
- Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David Krueger. Defining and Characterizing Reward Hacking. NeurIPS 2022. arXiv:2209.13085
- Maxime Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. TMLR 2024. arXiv:2304.07193
Acknowledgments
We thank Anthropic and OpenAI for the API keys and credits, and Ofir Press for his advice on setting up the evaluation environment.
Read more about the benchmark in Introducing BenchCAD 1.0.
Citation
If you use these results, please cite:
@misc{benchcad_2dproxy_2026,
title = {Frontier Models Reconstruct 3D by Hill-Climbing a 2D Proxy},
author = {{BenchCAD Team}},
year = {2026},
howpublished = {\url{https://benchcad.com/blog/2d-proxy.html}},
note = {Blog post}
}
@misc{benchcad2026,
title = {BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD},
author = {Zhang, Haozhe and Liu, Kaichen and Chen, Miaomiao and Li, Lei
and Yang, Shaojie and Peng, Cheng and Chen, Hanjie},
year = {2026},
eprint = {2605.10865},
archivePrefix= {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2605.10865}
}