Which models can actually do CAD?
5 models, 54 tasks, 675 attempts. Each attempt is graded on whether it built, on measured geometry, and on a checklist an AI judge answers from the renders. Click any number to open the attempt behind it.
Where the score and the reviewer disagree
Overall gives partial credit at every layer. Once every model can do most of a task the scores bunch: this run's five sit inside 2.4 points, and what little is left does not match what the reviewer saw. Strict reads the same evidence as a pass or a fail, and it does separate them.
All models
| # | Model | Quality | Cost and time | Overall | Strict | Human rank | Gate pass |
|---|---|---|---|---|---|---|---|
| 1 | Fable 5.1Anthropic | 95 | 68% | 49% | 98% | ||
| 2 | GPT-5.6 SolOpenAI | 93 | 68% | 43% | 95% | ||
| 3 | GPT-6 AstraOpenAI | 93 | 69% | 50% | 94% | ||
| 4 | Gemini 3.8 FlashGoogle | 93 | 71% | 61% | 96% | ||
| 5 | Opus 5Anthropic | 92 | 65% | 47% | 95% |
Overall gives partial credit at every layer, so the top of the board bunches. Strict counts only the attempts that got everything right. Gate pass is the share whose part was watertight, hole-correct and overlap-free. Hover any bar for its number. Ties break by cost, then time. In Quality, the bar marked with a person is one reviewer's ranking of the cases they ordered by hand. First place is 100%, last is 0%, ties share, and the tick marks 50%. That reviewer found every first attempt acceptable. Neither affects Overall or Strict.
Browse the tasks
Easy
Medium
Hard
608 bearing seat
Bottle neck adapter
Cable clip
Drawer divider
Extrusion 2020 end cap
Mesh add tabs
Mesh holder holes
Mesh rebuild blender handle
Mesh rebuild freecad bracket
Mesh rebuild freecad knob
Mesh rebuild scan
Mesh socket mount
Nema17 face mount
Photo cluttered real
Photo knob
Photo knob oblique
Photo lid
Photo motor bell
Photo resistor
Rpi4 tray
Bracket second mount
Enclosure build up
Knob then widen
Lid fit and finger hole
Photo lid then centred
How the score is made
Built 30%, measured geometry 50%, AI judge 20%. A failed gate zeroes geometry for that attempt. The AI judges are Opus 4.8 and GPT-5.6 Sol. Read the methodology.