System 03 · AI systems

Semantic Commerce Architecture

Built and evaluated across ten models

In progress

A parent/child hub architecture that makes relationships between topics, materials, products and editorial content legible - built through competing model prototypes, scored against a fixed rubric across ten models.

StatusIn progress Current state4 live · 1 completing · 1 next Plan12-15 child hubs under one parent

Problem

Commerce catalogs hide their own structure

A commerce catalog typically exposes products and collections, but not the relationships between topics, materials, product families, applications and editorial content. Those relationships are what both conventional search and AI-driven retrieval systems use to understand what a site is actually about.

This system is part of the broader AI Operations architecture and builds on the catalog described in the International Commerce Transformation case study.

Architecture

One parent hub, planned children

Current state: 4 live, 1 completing, 1 next. Schema, breadcrumbs and explicit structured relationships are under development.

The architecture is designed to expose clearer relationships to search and AI-driven retrieval systems. Whether AI crawlers demonstrably understand the architecture yet is an open question - no such claim is made here.

Workflow

Competing prototypes, human selection

Each page is produced as four competing prototypes. The model pool is ten models: Google Flash 3.7, Claude Opus 5, Claude Fable 5, Claude Fable 5.1, GPT Terra 5.6, GPT Sol 5.6, GPT Luna 5.6, Grok 4.6, Kimi K3 and GLM 5.2, with Codex used heavily for execution and iteration. The models interpret the same project differently - which is the point: the human compares, selects and critiques rather than prompting one model and accepting its first answer.

  • Input
  • Autonomous action
  • Human judgment
  • Output
Agents work from the Shopify catalog, existing collections and articles, keyword and search data, the existing hub architecture, product relationships, and brand rules and tone.

Design and layout remain consistent across languages; copy is adapted independently for German, English, French and Spanish, respecting native search intent, terminology and keyword relevance - localization is independent adaptation per language, not translation.

Evaluation

Ten models, one rubric

Prototypes are not chosen by feel. Each page is built from four competing prototypes on identical inputs, scored against three fixed criteria. Same inputs, same reviewers - the last decision is mine.

  • Semantic quality - Does the structure explain the internal relationships? And does it only claim relationships the catalogue actually supports?
  • Visual identity - Does it hold the brand guidelines and the UI/UX standard?
  • Hub consistency - Does it read as part of the same system as the hubs already shipped?

The rubric exists because the judgement is mine and needs to be repeatable, not mood-dependent.

The model pool

Google Flash 3.7 · Claude Opus 5 · Claude Fable 5 · Claude Fable 5.1 · GPT Terra 5.6 · GPT Sol 5.6 · GPT Luna 5.6 · Grok 4.6 · Kimi K3 · GLM 5.2

Six runs, results
Leading models by rubric criterion across six runs
Criterion Leading
Semantic quality Kimi K3, Opus 5
Visual identity Luna 5.6 - almost always, usually on the first prototype
Hub consistency Luna 5.6, Fable 5
Aggregate, cost-normalised Luna 5.6 clearly ahead · Kimi K3 second · Claude models third

Not the result I expected. The frontier models lead where thinking about relationships matters - but not on the task as a whole, and not remotely on price.

Public benchmarks measure ability on tasks that are not this task. Six runs against a rubric of my own say more about my problem than any leaderboard.

Limits of the claim. Six runs, one task class, one brand voice - a small sample. What transfers is the method, not the result: the winner is this quarter’s standing, and model prices move.

The rubric is the counterpart to the metadata system’s calibration ladder: there, quality could be reduced to rules and supervision was withdrawn after three to four runs. Here it cannot, so the human stays and the rubric is what keeps them consistent.

Human role

What stays human

  • Ownership of the brief and the final call
  • Judgement: comparison, critique, direction

Agents execute and iterate; the human decides what good looks like. Estimated saving is ≥4 full working days (~30+ hours) per completed four-language hub - research, content, design, coding, responsive implementation, localization and iteration. This is an estimate, not a measured figure.

Related