blog · 2026-10-06

Frontier Models Reconstruct 3D by Hill-Climbing a 2D Proxy

Frontier models check the CAD parts they build by measuring how well the outlines match the target, and for most parts that is enough.

TL;DR
  • Frontier models don't need to look at what they build. On BenchCAD, a harness that never shows the model its own renders scores the same as one that does, at about half the cost.
  • They measure instead of look. The model renders its part in code, computes the silhouette IoU against the target views, and hill-climbs that number. GPT-6 Astra, given no renderer at all, wrote its own.
  • This 2D proxy gets most parts right. Where it breaks is fine detail: once the silhouettes line up, the model can stop at a part that matches in 2D but not in 3D, a mild form of reward hacking.
  • Human experts lean the same way. Across about 10k pairwise votes from about 15 domain experts, people judged a part more by how it looks in the renders than by its 3D geometry.

In BenchCAD's image-to-code task1, a model gets four rendered views of a mechanical part and writes a CadQuery2 script that rebuilds it. It works in a harness loop, running code and reading the output over many rounds. Our standard harness is plain Python with CadQuery, and after each round it shows the model up to three renders of its current part.

We assumed those renders were essential for visual debugging. But when we ran the same model in mini-swe-agent3,4, which shows the target views only in the first turn and returns nothing but text after that, the score stayed flat and the API cost fell by about half.

The model's code explained why. It renders the part programmatically, computes the silhouette overlap with the target, and adjusts the geometry to push that number up, all in code. After the first turn, where it sees the target views, it never needs to look at an image again. A single scalar to optimise lets it make many fast, small iterations. On BenchCAD 2.0, which provides no rendering tool at all, GPT-6 Astra went as far as writing its own renderer and optimiser to run the same loop.

Seeing its own part does not help

3D score same Python toolssees its own renders every round Python tools · 3D score 0.783 0.783 mini-swe-agent v2target views in turn 1 only mini-swe-agent v2 · 3D score 0.781 0.781 Cost per part −52% Python toolssees its own renders every round Python tools · cost 1× (reference) 1× mini-swe-agent v2target views in turn 1 only mini-swe-agent v2 · cost 0.48× of Python tools 0.48× Same 98 parts, model, tools and reasoning effort. Cost is relative to Python tools.

Figure 1. One frontier model on the same 98 BenchCAD 1.0 parts in two harnesses, with the same tools and the same (high) reasoning effort.

Both harnesses let the model render its part the same way the target pictures were rendered; mini-swe-agent just never shows it the result. The scores differ by −0.002 (95% interval ±0.05). Switching the pictures off inside our own harness made no measurable difference either.

The model measures its part instead

We read the code from about 400 runs across the two harnesses. In about five of every six, the model wrote a script that renders its part and computes how much its silhouette overlaps a target picture, then edited the part and ran the comparison again. In almost every run that printed this overlap more than once, it went up, usually from about 0.8 to 0.98.

The model is hill climbing. It makes a change, keeps it if the number rises, and repeats. The number is an overlap of 2D outlines that stands in for the 3D answer, so we call it a 2D proxy.

Figure 2 shows a run of GPT-6 Astra on BenchCAD 2.0, where the task comes with no drawing tool.

Pictures Astra drew for itself Round 9 puts its drawing next to the target Round 14 fits its outline to the target's overlap 0.964 Round 25 refits with more parameters overlap 0.969 Round 29 a last check before submitting The real part, and what Astra submitted the real part its answer its part score (3D) 0.907 In its overlays, teal is where both outlines overlap; red is only its drawing, blue only the target.

Figure 2. GPT-6 Astra rebuilding a helical wind-turbine rotor from BenchCAD 2.0. It wrote a renderer for its own part and an optimiser that tunes the part's dimensions to fit the target silhouette. Its final part scores 0.907.

Other models do it too

We went through every run on 27 private BenchCAD 2.0 cases, for GPT-6 Astra and Claude Opus 5.5 at each reasoning-effort setting, and counted the runs that wrote their own drawing code.

from high effort up,nine in ten runs or more draw 0% 25% 50% 75% 100% low medium high extra high max reasoning effort GPT-6 Astra Claude Opus 5.5 GPT-6 Astra · low effort · 3/25 runs (12%) GPT-6 Astra · medium effort · 20/27 runs (74%) GPT-6 Astra · high effort · 22/24 runs (92%) GPT-6 Astra · extra high effort · 25/27 runs (93%) GPT-6 Astra · max effort · 24/25 runs (96%) Claude Opus 5.5 · low effort · 0/27 runs (0%) Claude Opus 5.5 · medium effort · 17/27 runs (63%) Claude Opus 5.5 · high effort · 23/25 runs (92%) Claude Opus 5.5 · extra high effort · 25/27 runs (93%) Claude Opus 5.5 · max effort · 9/10 runs (90%) Astra 96% Opus 5.5 90% (max: 10 runs)

Figure 3. Share of runs that write their own drawing code, by reasoning effort.

At low effort both models answer in a round or two and rarely draw. From high effort up, nine in ten runs or more write drawing code, and many use it to compare their part with the targets. Given room to think, both models settle on the same strategy.

Does climbing the proxy improve the 3D part?

If the 3D result comes from climbing the silhouette overlap, the model should do worse when its renders can no longer line up with the targets. On 100 BenchCAD 1.0 parts, we re-rendered three of the four target pictures 15° off, so that no part could match them pixel for pixel, and compared with a run that differed only in that.

3D score −0.064 Pictures line uprenders can match the targets exactly aligned · 3D score 0.786 0.786 Three pictures 15° offno part can match them pixel for pixel offset · 3D score 0.722 0.722 Rounds per part +7 Pictures line uprenders can match the targets exactly aligned · 41 rounds per part 41 Three pictures 15° offno part can match them pixel for pixel offset · 48 rounds per part 48 Same 100 BenchCAD 1.0 parts and prompt; only the alignment of three target pictures differs.

Figure 4. The same 100 BenchCAD 1.0 parts, with the target pictures aligned to the model's renders and with three of them 15° off.

Without alignment the 3D score fell from 0.786 to 0.722, and the model needed about seven more rounds per part.

Where the climb stops

Most runs that rated their own match 0.95 or higher also scored above 0.9 in 3D, with a median of 0.97. Figure 5 shows four that did not.

Double-row sprocket its own score 0.948 · real 3D score 0.410 the real part its answer sprocket teeth, fine gear teeth instead Bellows its own score 0.948 · real 3D score 0.560 the real part its answer thin folded wall, thick wall Duct elbow its own score 0.981 · real 3D score 0.429 the real part its answer hollow duct, nearly solid answer Venturi tube its own score 0.978 · real 3D score 0.637 the real part its answer thin narrowing wall, thick straight bore

Figure 5. Four BenchCAD 1.0 parts where the model rated its answer around 0.95 and the 3D score was lower, each next to the real part and cut open where they differ. We picked these to show where the silhouette score stops helping; most high-scoring runs don't look like this.

In all four, the details are wrong. The duct elbow is nearly solid where it should be hollow. The bellows has a thick wall where the real one is thin and folded. The venturi tube's bore runs straight instead of narrowing. The sprocket has fine gear teeth instead of sprocket teeth. Each of these barely changes the silhouette overlap but lowers the 3D score a lot.

Once the silhouette matched, the model's own score stopped telling it anything about the details, and its modelling of small features wasn't good enough to get them right without that signal. So it settled for a part that scored well on its proxy, a mild form of reward hacking5.

Experts also judge by appearance

We also checked whether people catch these differences. About 15 domain experts compared pairs of rebuilt parts against the original, about 10k votes in all. Each screen showed four rendered views of every part, with the 3D solids one click away.

2D: compares images 3D: compares solids 50% 60% 70% 80% Image similarity, DINOv2 (2D) Image similarity, DINOv2 (2D) · 3182/4237 = 75.1% 75% Surface distance (3D) Surface distance (3D) · 2791/4246 = 65.7% 66% Volume overlap, rescaled (3D) Volume overlap, rescaled (3D) · 2244/3437 = 65.3% 65% Surface match (3D) Surface match (3D) · 2691/4156 = 64.7% 65% Volume overlap (3D) Volume overlap (3D) · 2633/4172 = 63.1% 63% two experts pick the same part 80%

Figure 6. How often each measure picks the same part as the expert (50% is chance). The orange bar compares images (2D); the blue bars compare the 3D solids. The dashed line is how often two experts agree.

Of the measures in Figure 6, expert choices agree most with DINOv26, an image model that scores how alike two pictures look (75%), well above the volume and surface measures CAD benchmarks usually report (63 to 66%). When the image score and the 3D score disagree about which part is better, experts side with the image score two times in three.

Image similarity can be fooled as well. A bevel gear with a second copy stuck to its side still scores 0.92 under DINOv2, about the same as the 0.93 it gives the same gear turned over.

What this means

BenchCAD 2.0's rebuild-from-pictures task is built against this kind of pixel matching. It gives no renderer and states only one camera position, (1, 1, 1); the other three views are taken from slightly different angles that are not disclosed. Lining up pixels alone isn't enough to score well.

Outline matching gets most parts right. The remaining gains are in fine detail such as wall thickness, bores and teeth, which takes a check more sensitive than outline overlap and better modelling of small features.

CAD benchmarks should score the solid, because image similarity can be fooled by a part that looks right, and expert review from pictures leans the same way.

Limits

  • The scores, gaps and experiments come from one frontier model on BenchCAD 1.0's image-to-code task. For GPT-6 Astra and Claude Opus 5.5 we only counted what their code does.
  • Figures 2 and 5 are chosen examples. Figure 2 is a single run, and Figure 5 shows low 3D scores on purpose.

Appendix A: Experiment details

The BenchCAD 1.0 results use about 300 image-to-code parts, each run once in Python tools at high reasoning effort with up to 100 rounds; the harness comparison and the camera test use subsets of about 100 parts. The 3D score is voxel IoU on a 64 × 64 × 64 grid. Intervals are 95% bootstrap intervals over parts. What each run did is read from its code and output.

Appendix B: The expert study

Experts saw the original part above two candidate reconstructions, each as four rendered views, with the 3D solids one click away, and picked the closer candidate or called it a tie. A small share of candidates were deliberately altered copies of the original, such as the doubled bevel gear. Figure 6 uses every vote where the expert picked a side, and counts agreement vote by vote: a measure agrees when it scores the expert's pick higher.

References

  1. Haozhe Zhang, Kaichen Liu, Miaomiao Chen, Lei Li, Shaojie Yang, Cheng Peng, Hanjie Chen. BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD. 2026. arXiv:2605.10865
  2. CadQuery. github.com/CadQuery/cadquery
  3. mini-swe-agent. SWE-agent team. github.com/SWE-agent/mini-swe-agent
  4. John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. arXiv:2405.15793
  5. Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David Krueger. Defining and Characterizing Reward Hacking. NeurIPS 2022. arXiv:2209.13085
  6. Maxime Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. TMLR 2024. arXiv:2304.07193

Acknowledgments

We thank Anthropic and OpenAI for the API keys and credits, and Ofir Press for his advice on setting up the evaluation environment.

Read more about the benchmark in Introducing BenchCAD 1.0.

Citation

If you use these results, please cite:

@misc{benchcad_2dproxy_2026,
  title        = {Frontier Models Reconstruct 3D by Hill-Climbing a 2D Proxy},
  author       = {{BenchCAD Team}},
  year         = {2026},
  howpublished = {\url{https://benchcad.com/blog/2d-proxy.html}},
  note         = {Blog post}
}