v0.2-dev · updated 2026-09-27 · 1 to 3 epochs per prompt

CADClamp Results

12 models × 7 CAD languages × 47 prompts, scored by a deterministic DfM engine against a reference solution per prompt

2026-09-26 harness 0.3 · engine 0.2.2 · spec 0.2.3 build123d · OpenSCAD · CadQuery · FreeCAD · Rhino 8 · Fusion · Blender OpenRouter · Claude Code CLI · Ollama
Correction, 2026-09-26. I found four bugs in the harness that produced the first v0.2-dev numbers and re-scored every run with them fixed. Each bug failed a part for a reason unrelated to whether it prints: a time limit that summed CPU seconds across every core, timeouts from an overloaded machine, OpenSCAD warnings treated as errors, and inside-out solids that every slicer repairs. The fixes only raised scores. 17 of the 44 cells in the published four-language table went up and none went down. I re-ran each model's saved code with no new model calls, and every changed sample records its old result.
Modelpublishedcorrected
claude-fable-5-10.9540.954
gpt-6-astra0.9350.941
gpt-6-luna-pro0.9120.928
gpt-6-sol0.9200.925
claude-opus-5-50.9200.920
grok-4.70.8800.896
gpt-6-luna0.8600.889
grok-4.60.8750.880
claude-opus-50.7900.809
gpt-5.10.2770.282
qwen2.5-coder:7b0.0130.024
Averages over the original four languages.
Spec 0.2.2, 2026-09-26. The v0.2 checks could not see inside a part: a ball-joint socket with no socket had the right size, volume and hole count, and scored 1.00. I added probes for every stated feature and proved each prompt's checks against mutants, copies of its reference with one feature broken, until every prompt rejected every mutant. Re-grading the saved parts failed 79 that had passed. I reviewed each by cross-section and every one misses a feature, most often a thread left out entirely. I no longer grade placement, so a correct part in the wrong place passes, but a part built upside down still fails. I did not rerun any model, and each lost 0.01 to 0.04. On 2026-09-27 the criterion checks joined the headline, scored against the reference, which moved no model by more than 0.013. Spec 0.2.3 then added a thread check (the prompts say right-hand single start), which failed three left-hand OpenSCAD threads from gpt-6-sol, kimi-k3 and grok-4.7. The tables on this page are the new ones.
How the headline works. A part earns its printability only if it passes every spec assertion: size, volume, hole count, and a probe for each stated feature. I then divide that printability by the reference solution's score on the same prompt, capped at 1. The reference solutions score 1.000 and a plain 20 mm cube scores 0.000, even though the cube's raw printability is a perfect 1.0.
headline = min(1, printability / reference printability)   if every spec assertion passes
         = 0                                                otherwise

Most cells are one attempt per prompt, with a 95% confidence interval of about ±0.08. opus-5.5 and fable-5.1 ran three epochs and astra two, so their cells are averages. The top three rows are a statistical tie. Underlined values are the best score in that column. "Valid" is the share of runs that produced a valid solid, averaged over the seven languages.

§ Key findings
Top of the table opus-5.5 (0.909), gpt-6-astra (0.900), and fable-5.1 (0.898) are a three-way tie on the seven-language average. The winner changes with the language: fable-5.1 in OpenSCAD and Fusion, astra in CadQuery, Rhino, and Blender, sol in build123d, and kimi-k3 in FreeCAD. Paired on the same prompts, opus-5.5 leads astra by 0.009 (95% interval −0.024 to +0.042).
Where models separate Most of the gap opens before a part exists. When a frontier model produces a valid solid, its printability lands near 0.9 whoever wrote it. The rows separate on whether the code runs and whether the part meets the spec.
Commercial CAD through MCP Rhino and Fusion spread the models furthest. gpt-6-luna scores 0.816 in build123d and 0.308 in Rhino. Its parts are not worse, but most of its Rhino code never runs. An API call that does not exist is the most common failure in both programs: 35% of Rhino errors and 39% of Fusion errors.
Chinese frontier kimi-k3 places ninth (0.726). It is level with the leaders on FreeCAD (0.946) but falls to 0.372 on Fusion. In the v0.1 grid it placed third, so the harder prompts and the commercial programs widen the gap.
One look fixes most of it With one repair turn, the three leaders reach 0.925 to 0.998 in Rhino and Fusion. An error message fixed all 23 crashed samples. A render of the part fixed parts that built wrong, and fable-5.1's Rhino score went from 0.697 to 0.953.
Freedom sinks weak models On the same 20 easy prompts, gpt-5.1's valid rate in build123d fell from 35% (v0.1) to 5%. Across all 47, 42 runs died with runtime errors, mostly invented API calls, and 16 answers reached for teardrop holes the print contract allows but does not require.
§ By language

build123d (Python, B-rep)

OpenSCAD (DSL, mesh)

CadQuery (Python, B-rep)

FreeCAD (Python, B-rep)

Rhino 8 (RhinoCommon, via Rhino MCP)

Fusion (Fusion API, via Fusion MCP)

Blender (bpy, headless)

§ What the models built

Every tile is a render of the exact STL that the engine scored. Green meets the spec, red is a valid part that misses it, and grey means the code produced no part.

Six prompts in Rhino: the three best models build all six parts, the three weakest mostly fail to run
The same six prompts in Rhino. The top three rows are the best models, the middle row is kimi-k3, and the bottom three are the weakest.
The same six prompts in OpenSCAD for the same six models
The same prompts in OpenSCAD, where the weaker models build more but still miss the harder parts.

The harness builds the Rhino and Fusion parts inside the real programs. These screenshots come from running the models' saved code again inside each program through its MCP server. Their volumes match the scored STLs within 0.1% in Rhino and 0.06% in Fusion.

gpt-6-astra t4-006 flanged bushing in Rhino
gpt-6-astra, t4-006, in Rhino
claude-fable-5-1 t5-001 print-in-place bearing in Fusion
claude-fable-5-1, t5-001 bearing, in Fusion
§ Repair rounds

Text repair sends the error output back once. Image repair shows the model four rendered views of its part, pass or fail, and lets it keep or fix the code. Both reuse the logged first answer, so only the second turn is new. Image repair ran on Rhino and Fusion only.

ModelLanguagesingle attempttext repairimage repair
gpt-6-astraRhino0.9030.9890.989
gpt-6-astraFusion0.7970.9040.925
claude-fable-5-1Rhino0.6970.7610.953
claude-fable-5-1Fusion0.9550.9970.997
claude-opus-5-5Rhino0.8610.9030.967
claude-opus-5-5Fusion0.8710.9130.998

Across the six languages it covered, text repair lifts the averages to 0.929 for astra, 0.921 for fable-5.1, and 0.911 for opus-5.5.

claude-fable-5-1 Rhino parts before and after image repair
fable-5.1 in Rhino: its first attempts (top) and the parts it rebuilt after seeing a render of them (bottom).
§ Checked by a real slicer

Every valid part sliced in OrcaSlicer 2.4.2 on five printers, each with its maker's own 0.20 mm profile and default supports: Bambu Lab P2S, Prusa CORE One, Creality K2 Plus, Elegoo Centauri Carbon 2, and Anycubic Kobra S1. 9,399 of 9,415 slices succeeded. 19 of the 47 reference solutions need some support, so I measure support against the reference for the same prompt.

Modelslices on all fiveneeds no supportsupport beyond referenceneeds support where reference needs none
grok-4.6100.0%80%0.0030%
grok-4.7100.0%79%0.0030%
claude-opus-5-5100.0%76%0.0010%
claude-fable-5-1100.0%76%0.0051%
gpt-6-luna-pro100.0%71%0.0091%
kimi-k399.6%71%0.0122%
claude-opus-598.4%71%0.0113%
gpt-6-astra100.0%68%0.0030%
gpt-6-sol100.0%67%0.0010%
gpt-6-luna100.0%65%0.0113%
qwen2.5-coder:7b100.0%61%0.0174%
gpt-5.199.1%54%0.02515%

When the engine's overhang check says fail, the slicer adds support on 98% of those parts. The engine misses 19 of 2,345 parts (under 1%), and the five printers agree on 97% of parts. These results are advisory and stay out of the headline.

§ Criterion checks (advisory)

Tier 5 prompts name the DfAM decision they test, and six checks grade it directly: bridge_span, fit_clearance, kinematic_sweep, load_orientation, bed_interface, living_hinge. I report them separately and keep them out of the headline. Each covers one or two prompts, so these are case studies, not rankings. Counts are over the four code-CAD languages.

fit_clearance: print-in-place bearing

Languages (of 4) scoring 0.98. Most models offset the part radially by 0.5 mm, which leaves about 0.35 mm across the 45° flanks (0.74). The grid ran on prompt text that allowed that reading, and v0.2.1 now says "measured perpendicular to the hub surface."

bed_interface: large flat parts

Languages (of 4) scoring 0.94 or better, meaning the bottom edge got a chamfer or radius. The reference leaves it plain (0.52) because the choice belongs to the model. The two checks rank the same models in nearly opposite orders.

§ Track C: real catalog parts, redesigned to print

Ten McMaster-Carr parts, from a spacer to a hinge and a rod end that print assembled. Each prompt describes the original part and the interfaces that must survive (bores, hole spacing, clearances), and the model redesigns the rest to print. Every prompt has a passing reference, and a solid block the size of the part fails every one. Each catalog photo sits next to the best part any model made.

McMaster-Carr catalog photos next to the best generated part for each of the ten Track C prompts
Catalog photo (left) and the best generated part (right), with the model and language that made it. Photos belong to McMaster-Carr.
Modelbuild123dOpenSCADCadQueryFreeCADBlenderRhinoFusionavg
claude-opus-5-50.8890.7750.8530.8410.4990.5630.8530.753
claude-fable-5-10.9030.7910.7160.6170.5950.3970.9520.710
gpt-6-astra0.8430.4850.8590.8440.7260.3760.7130.692
kimi-k30.3920.4390.6940.6920.3970.4500.2000.466

Each cell is ten prompts, so about ±0.2. opus-5.5 is the most even. fable-5.1 has the best cell (Fusion) and one of the worst (Rhino). Every model scores lower here than on v0.2 in Blender and Rhino.

§ Earlier: v0.1-dev

The v0.1 grid (15 models, build123d and OpenSCAD, 20 easy prompts, 3 attempts each, 2026-08-14) is in the project README. Its scores are raw printability with no reference solutions, so they are not on this page's scale.