partforge
Browse forgesLearnBlogSign in
Your forges
AdminMessagesSettings
← Blog

New model day: we tested both of them

September 23, 2026 · 5 min read · by Scott Sykora

LLM CAD Harness Evals

Anthropic released Opus 5.5 yesterday. A few hours later OpenAI released GPT-6 Sol and GPT-6 Luna. Yay 'new model day!' Two frontier labs, one afternoon.

We have a standing set of hard CAD tasks we run every model through for partforge, so we started a run yesterday evening. This is what came back.

The short version: Opus 5.5 is dramatically faster and cheaper than the Opus before it, GPT-6 Sol gives you more for the same money, Opus 5.5 is the closest anything has come to beating our default, and we are still shipping Gemini 3.8 Flash as the default anyway.

What we actually test

This eval was 26 hard modeling tasks. Things like a bearing seat with a counterbore, or a part described only by a photo. Each model tries each task three times, and every attempt is checked three ways: does the geometry measure correctly, did the model work sensibly to get there, and does it look right to a pair of independent reviewer models. Then a human (I) go through the results by hand, look at the renders, and rank the models against each other on every one of the 26 tasks.

It is a small set on purpose. Easy tasks stopped telling us anything useful months ago. All the frontier models pass them every time.

You can read the whole run, every attempt and every render, at partforge.ai/evals.

Opus 5.5 got dramatically better

This is the headline. On the same tasks, compared to Opus 5 two weeks ago:

  • About 40% cheaper per task.
  • Roughly a third of the time. Opus 5 averaged four and a half minutes of model time per task. Opus 5.5 averaged a minute and a half.

Anthropic also cut the list price about 20%, so some of that is the price tag and the rest is Opus 5.5 simply doing less thrashing to get to an answer.

It also posted the best overall score in this run, narrowly ahead of Gemini. It is not the best we have ever measured on these tasks (Fable 5.1 scored higher two weeks ago, and was not in this run), but it is the strongest showing from a model we would consider making the default.

One of our harder evals is one that asks the model to generate a motor spoke assembly based on complicated photo with only half the motor in the frame. Opus 5.5 got the closest to the actual geometry yet in all the models we've tested.

GPT-6 Sol quietly got better value

Sol is the interesting one, because on the surface almost nothing changed.

OpenAI cut the price about 20% on the way in and a third on the way out. But GPT-6 Sol also does more work per task than GPT-5.6 Sol did. It reads about half as much again, takes an extra step or two, and thinks a little harder on the tasks it finds difficult.

Those two cancel out almost exactly. A hard task cost us 25.7 cents on the old Sol and 25.2 cents on the new one. Call it the same.

What you get for that same quarter is a slightly better part. The score nudged up, and the share of attempts that built and ran cleanly held steady at 92%. So the honest summary is not "cheaper", it is "more model for your money", which is usually the better trade anyway.

Worth one piece of context: over the same two weeks, Gemini, which did not change at all, got 16% more expensive on these exact tasks as our test set and prompts moved around it. Holding flat in that window is better than it sounds.

So why are we still on Gemini?

Cost. The number we care about is not price per task, it is price per task that actually came out right. Gemini 3.8 Flash costs about $0.34 per solved task. GPT-6 Sol is about $0.48. Opus 5.5 is about $0.72, a bit over double Gemini. When you are iterating on a part, you do not run one prompt, you run fifteen, and that gap compounds fast.

It is still the most reliable. Opus 5.5 wins our overall score, but Gemini wins the two measures that matter most day to day: how often a part passes every single check with nothing wrong (64% versus 58%), and how often it builds and runs at all. In my own ranking pass the two were a hair apart, with Gemini barely ahead. Opus 5.5 is close enough that this could flip next month. It has not flipped yet.

And it feels faster, even when it is not. This one surprised us. Opus 5.5 finishes a task in less total time than Gemini does. It still feels slower to use.

Gemini works in lots of small steps, around eight per task, averaging 17 seconds each. Opus works in about four big ones, averaging 23 seconds each, and its slowest steps run past two minutes. Worse, Anthropic's models often send us no visible thinking at all during those stretches, so there is nothing to show you. You sit watching a still screen for two minutes and then the part changes all at once. We could improve that with an animating asterisk but we think showing real thinking and clues to what's going on is a real benefit to the user experience.

With Gemini, something on screen moves every fifteen or twenty seconds. You can see it working, and you can tell early when it has misunderstood you and stop it. A model that is quicker on paper but silent while it works does not feel quick and hurts the experience.

Try them yourself

None of this has to be our call. Every model here is selectable in partforge, including Opus 5.5 and both new GPT-6 models. Turn on advanced models in the settings, open the model picker next to the chat box and switch.

If you bring your own API key from Anthropic, OpenAI or OpenRouter, those runs go on your key and your bill, and we do not meter them at all.

We will run this again the next time somebody ships something. Which, at the current rate probably has already happened.

One honest caveat on the numbers above: the two runs we are comparing are two weeks apart, and our own test harness changed between them. Gemini, which did not change at all, came out 16% pricier and slower on the second run. So treat the before and after figures as a general trend rather than exact measurements. The head to head numbers within the new run are directly comparable.

Featured now

See all →
Layered Label

Layered Label

by Scott

Beer Tap Handle

Beer Tap Handle

by jake

Minimal Modular Storage Container

Minimal Modular Storage Container

by eugene

Steampunk Spider

Steampunk Spider

by Scott

E-Stop Enclosure

E-Stop Enclosure

by Scott

Hexagonal Chain Mail

Hexagonal Chain Mail

by Scott

Printable Nut & Bolt with configurable knob

Printable Nut & Bolt with configurable knob

by Scott

Propeller

Propeller

by Scott

Browse forges·Evals·Learn·Blog
·
Terms·Privacy·Contact·Your Privacy Choices

© 2026 Pixite Inc.