ABENE Agent Comparison
ABENE VHF-680 · Blender 5.2 · two agent runs compared

One milling machine, two agents

Claude Opus 5.5 and GPT-6 (astra) were each asked to model the ABENE VHF-680 CNC mill in Blender from a product photo and the manufacturer's brochure. Both got a recognisable machine. They took different routes, and the routes decide what the result can be used for.

Written by the Opus session after reading the GPT-6 process report and comparing it with its own build log. That makes one side an interested party. The judgments below try to stay with what the two reports and their renders show.

Product photo of the ABENE VHF-680
Reference. The product photo both agents worked from.
Opus render of the VHF-680
Opus 5.5. 129 meshes, 21 image textures, exported to USDZ.
GPT-6 render of the VHF-680
GPT-6. 800+ objects, 19 procedural Cycles materials, delivered as .blend.

At a glance

Opus 5.5GPT-6 (astra)
Driving BlenderLive Blender session through the MCP for Blender addon socket (execute_code)Blender CLI in background mode (blender -b --python script.py), one process per script
Desktop controlNoneBrief Computer Use attempt (@oai/sky). Did not reach the Python console and was not used for modelling.
Scale anchorBrochure: 900 × 500 mm table, 7 T-slots, X/Y/Z 800/510/475 mmThe same brochure values
GeometryParts merged into 129 meshes with UVs and baked bevels823+ separate objects with live bevel and weighted-normal modifiers, boolean window cut-outs
Materials27 PBR materials driven by 21 generated image maps (colour, roughness, normal) plus decals19 Principled BSDF materials with procedural noise bump, anisotropy and true transmission
Screen and keysLit HMI screen texture and legend texture on each key capPowered-off LCD, plain key caps
HierarchyGrouped by moving axis: Axis_Z → Y → X → A → C, spindle, lamp joints, pendantGrouped by assembly in numbered collections (enclosure, castings, table, head …)
Deliverable.blend, renders, ABENE_VHF-680.usdz (11.7 MB, Y-up, metres).blend (two versions) and three renders up to 2600 px
Validationpxr checks plus a USDZ round trip into a clean BlenderVisual review of renders; six defects found and fixed
ScriptsOne builder you can re-run (vhf680_model.py) plus texture and export scriptsTwo builders plus four repair scripts. The report warns that blind re-runs duplicate parts.
ReportBuilt around the five prompts, 18 images, 5.7 MB (mostly embedded fonts)Evidence-led, 4 images, all six scripts embedded with SHA-256, 0.9 MB

Where the approaches diverge

How each agent talked to Blender

Opus

Opus used the MCP for Blender addon running inside an open Blender. Every change landed in the same live scene, so fix, render and compare took seconds. The MCP tools were only registered partway through the session, so Opus sent the addon's execute_code command over its socket directly. That is the same call the MCP server makes.

GPT-6

GPT-6 started Blender headless for each script, saved the .blend, then rendered. It needs no addon and no running UI, and every run starts from a known file. Its early runs failed on relative paths, and it fell back to the CLI after the Computer Use attempt couldn't open a Python console.

Both are valid. The live socket is faster for iteration. The CLI is easier to automate in CI and doesn't depend on a third-party addon. In both cases Python did the modelling, not the GUI.

Materials: procedural vs. image-based

Opus

make_textures.py generates seamless image maps from filtered noise. UVs are scaled in metres per tile so texture density is consistent. The maps translate one-to-one into USD UsdPreviewSurface, so the USDZ looks much like the Blender scene.

GPT-6

Noise, bump, anisotropy and transmission are all built as Cycles shader nodes. This needs no texture files and reads very well in Cycles, especially on brushed and chromed parts. None of it survives a USD or glTF export without baking. The report says so itself under "Not established".

For a Blender still, GPT-6's choice is lighter and looks at least as good. For anything downstream, like USD, Omniverse, AR or a digital twin, the image-based route is the one that carries over.

Geometry and structure

Opus

Each part's pieces are merged into one mesh: 129 meshes, about 130 k triangles. Bevels are applied and normals hardened. The parent chain follows the machine's axes, so the table, head and pendant can be moved like the real machine.

GPT-6

Hundreds of small named objects (washers, screws, 72 graduation marks on the rotary table) keep live modifiers. Window openings are real boolean cuts in sheet-metal profiles. That gives crisp, convincing edges and a scene that is easy to edit, but heavy for real-time use. It has no motion hierarchy, and the report says the travel values were never turned into motion.

GPT-6 produced more small mechanical detail. Opus produced a structure that is ready for animation and simulation. Which matters more depends on the job.

Reading the photo

Opus

Opus worked from 4× crops of the photo. That produced the white head with black motor, the deep monitor housing on the pendant with the screen enlarged after user feedback, the handheld handwheel, and an iTNC-style screen. The camera angle is close to the photo's. Opus read the blue ring beside the head as a looped coolant hose, and the lower front as a stainless chip chute.

GPT-6

GPT-6 read the blue ring as a gear selector dial with a chrome bezel and numerals. Looking at the photo again, that is probably the better reading. Its open, dark area under the table is also closer to the photo. On the other side, it painted the head grey-green where the photo shows a light finish, and the pendant reads boxier.

Neither model matches the photo everywhere. Each got details the other missed. Close-ups of the two consoles and heads follow.

Opus pendant close-up
Opus pendant. Lit screen, legends on the key caps, handheld handwheel.
GPT-6 console close-up
GPT-6 console. Clean bevels and mounting detail, unlit screen.
Opus head close-up
Opus head. White casting, gearbox badge, blue hose loop.
GPT-6 spindle close-up
GPT-6 head. Blue selector dial, stepped spindle, beaded coolant hose.

How the work was reported

Opus

The build log is organised around the user's five prompts, quoted verbatim, with what was done after each and the renders it produced. It is easy to follow as a story. It does not include the scripts. Its 5.7 MB is mostly 36 embedded font files, which could be cut down a lot.

GPT-6

The report is written like an audit: which tool did what, what evidence survived, and what is verified versus not established. All scripts are embedded with hashes. It does not quote the user's prompts, so the conversation that drove the work is harder to see.

The GPT-6 report is more careful about evidence and more useful to someone who wants to reproduce the work. The Opus log better shows a reader how the work moved from prompt to result.

What each could take from the other

Opus would gain from

  • Real boolean cut-outs and finer hardware (screws, hinge pins, graduation marks) where the camera gets close.
  • Re-checking the blue head detail as a selector dial.
  • Embedding the build scripts with hashes, plus a "verified / not established" list in the report.
  • Embedding only the Latin font subset to bring the report under 1 MB.

GPT-6 would gain from

  • Baking procedural materials into image maps, or authoring them as images from the start, so the model survives export.
  • A USD/USDZ export with a round-trip check.
  • A parent hierarchy by moving axis so the brochure's travels can be animated.
  • One builder that can be re-run instead of builders plus stateful repair scripts.

Limits of this comparison