partforge

Which models can actually do CAD?

partforge is an agentic coding harness for parametric CAD. An AI agent works with you to turn a prompt, and whatever guidance follows it, into a small program that generates a 3D model. Along the way it has the documentation, a reference library of published parts, and web search. After each step of a build it gets screenshots and measurements of the thing it has made. That makes this an unusually direct way to measure how the current models handle vision, tool use and reasoning about 3D geometry at once.

This first baseline ran 5 models over 54 tasks, 675 attempts in all. Each attempt is graded on whether it built, on measured geometry, and on a checklist an AI judge answers from the renders. Click any number to open the attempt behind it.

Quality
against
0%25%50%75%100%$0.00$0.30$0.60$0.90$1.200s50s1m 40s2m 30s3m 20s4m 10sFable 5.1GPT-5.6 SolGPT-6 AstraGemini 3.8 FlashOpus 5
Gemini 3.8 Flash sits alone in the top-left corner: first on human rank and the cheapest to run. 100% is first on every task, 0% is last, 50% is tied with the field.

What we found

Some of it surprised us. Cost and time spread much further than quality did: the field runs from $0.16 to $1.02 a task, and from 1m 17s to 3m 59s. Almost any current model can build a washer or a spacer. Ask for a motor spoke, a clamshell cover, or a lid that has to fit something, and the results vary a lot.

The thing we kept running into is that a person's idea of a good design does not always match the AI judges'. Better ergonomics and simpler forms are what set some models apart, and nothing in the score sees that directly.

This is the first of many runs, and we want to see how DeepSeek V4.1 and other open weight models do. On the strength of these results we have made Gemini 3.8 Flash the default model for partforge users. It costs about a sixth of what GPT-6 Astra does, it is quicker than every model that scored near it, and the human ranking came down in its favour more often than not.

Where the score and the reviewer disagree

Overall gives partial credit at every layer. Once every model can do most of a task the scores bunch: this run's five sit inside 2.4 points, and what little is left does not match what the reviewer saw. Strict reads the same evidence as a pass or a fail, and it does separate them.

Overall scoreHuman rankingFable 5.1: 1 by overall score, 3 by the reviewer's ranking, 2 places lowerGPT-5.6 Sol: 2 by overall score, 5 by the reviewer's ranking, 3 places lowerGPT-6 Astra: 3 by overall score, 2 by the reviewer's ranking, 1 place higherGemini 3.8 Flash: 4 by overall score, 1 by the reviewer's ranking, 3 places higherOpus 5: 5 by overall score, 4 by the reviewer's ranking, 1 place higher1Fable 5.12GPT-5.6 Sol3GPT-6 Astra4Gemini 3.8 Flash5Opus 5Gemini 3.8 Flash1GPT-6 Astra2Fable 5.13Opus 54GPT-5.6 Sol5

All models

Rank by
#ModelQualityCost and timeOverallStrictHuman rankGate pass
1Fable 5.1Anthropic
9568%49%98%
2GPT-5.6 SolOpenAI
9368%43%95%
3GPT-6 AstraOpenAI
9369%50%94%
4Gemini 3.8 FlashGoogle
9371%61%96%
5Opus 5Anthropic
9265%47%95%

Overall gives partial credit at every layer, so the top of the board bunches. Strict counts only the attempts that got everything right. Gate pass is the share whose part was watertight, hole-correct and overlap-free. Hover any bar for its number. Ties break by cost, then time. In Quality, the bar marked with a person is one reviewer's ranking of the cases they ordered by hand. First place is 100%, last is 0%, ties share, and the tick marks 50%. That reviewer found every first attempt acceptable. Neither affects Overall or Strict.

The tasks run from a plain washer to rebuilding a part from a photograph, and every one of them grades what the model built rather than what it said about it. The methodology page has the details.

Browse the tasks

How the score is made

Built 30%, measured geometry 50%, AI judge 20%. A failed gate zeroes geometry for that attempt. The AI judges are Opus 4.8 and GPT-5.6 Sol. Read the methodology.

Run baseline-2026-092026-09-10commit 41aec7daspend $339.70 + $48.02 AI judging, 7 unpricedOther runs: 2026-09-04T20-02-10-stage2-hard-cases, 2026-09-04T19-48-31-screen-14-fixed