System 03 · AI systems
Semantic Commerce Architecture
Built and evaluated across ten models
In progress
A parent/child hub architecture that makes relationships between topics, materials, products and editorial content legible - built through competing model prototypes, scored against a fixed rubric across ten models.
Problem
Commerce catalogs hide their own structure
A commerce catalog typically exposes products and collections, but not the relationships between topics, materials, product families, applications and editorial content. Those relationships are what both conventional search and AI-driven retrieval systems use to understand what a site is actually about.
This system is part of the broader AI Operations architecture and builds on the catalog described in the International Commerce Transformation case study.
Architecture
One parent hub, planned children
The architecture is designed to expose clearer relationships to search and AI-driven retrieval systems. Whether AI crawlers demonstrably understand the architecture yet is an open question - no such claim is made here.
Workflow
Competing prototypes, human selection
Each page is produced as four competing prototypes. The model pool is ten models: Google Flash 3.7, Claude Opus 5, Claude Fable 5, Claude Fable 5.1, GPT Terra 5.6, GPT Sol 5.6, GPT Luna 5.6, Grok 4.6, Kimi K3 and GLM 5.2, with Codex used heavily for execution and iteration. The models interpret the same project differently - which is the point: the human compares, selects and critiques rather than prompting one model and accepting its first answer.
- Input
- Autonomous action
- Human judgment
- Output
Design and layout remain consistent across languages; copy is adapted independently for German, English, French and Spanish, respecting native search intent, terminology and keyword relevance - localization is independent adaptation per language, not translation.
Evaluation
Ten models, one rubric
Prototypes are not chosen by feel. Each page is built from four competing prototypes on identical inputs, scored against three fixed criteria. Same inputs, same reviewers - the last decision is mine.
- Semantic quality - Does the structure explain the internal relationships? And does it only claim relationships the catalogue actually supports?
- Visual identity - Does it hold the brand guidelines and the UI/UX standard?
- Hub consistency - Does it read as part of the same system as the hubs already shipped?
The rubric exists because the judgement is mine and needs to be repeatable, not mood-dependent.
The model pool
Google Flash 3.7 · Claude Opus 5 · Claude Fable 5 · Claude Fable 5.1 · GPT Terra 5.6 · GPT Sol 5.6 · GPT Luna 5.6 · Grok 4.6 · Kimi K3 · GLM 5.2
| Criterion | Leading |
|---|---|
| Semantic quality | Kimi K3, Opus 5 |
| Visual identity | Luna 5.6 - almost always, usually on the first prototype |
| Hub consistency | Luna 5.6, Fable 5 |
| Aggregate, cost-normalised | Luna 5.6 clearly ahead · Kimi K3 second · Claude models third |
Not the result I expected. The frontier models lead where thinking about relationships matters - but not on the task as a whole, and not remotely on price.
Public benchmarks measure ability on tasks that are not this task. Six runs against a rubric of my own say more about my problem than any leaderboard.
Limits of the claim. Six runs, one task class, one brand voice - a small sample. What transfers is the method, not the result: the winner is this quarter’s standing, and model prices move.
The rubric is the counterpart to the metadata system’s calibration ladder: there, quality could be reduced to rules and supervision was withdrawn after three to four runs. Here it cannot, so the human stays and the rubric is what keeps them consistent.
Human role
What stays human
- Ownership of the brief and the final call
- Judgement: comparison, critique, direction
Agents execute and iterate; the human decides what good looks like. Estimated saving is ≥4 full working days (~30+ hours) per completed four-language hub - research, content, design, coding, responsive implementation, localization and iteration. This is an estimate, not a measured figure.
Related