Flagship 01 · AI systems

AI Operations

Repetitive knowledge work rebuilt around specialist AI agents - with orchestration, controlled autonomy and governance that keeps judgment human.

Systems3 · two autonomous, one in progress ModelSpecialist agents + orchestrator PrincipleAutomate work, not judgment

Leverage

220+ h Manual production avoided

Metadata across four languages. Human baseline ~7 min per localized entity.

~9.5 h/wk Analytical hours saved

Estimated average across the combined Google-platform workload.

≥4 days Saved per four-language hub

Estimate, ~30+ hours - research, content, design, code, localization.

Context

Repeatable knowledge work was the bottleneck

Three categories of work kept consuming disproportionate human capacity: repeated performance analysis across disconnected Google platforms, metadata production across a four-language catalog, and semantic content architecture that required research, design, code and localization for every page.

None of these are judgment problems. They are production problems with judgment checkpoints - which makes them candidates for a specific kind of AI architecture, not for generic chatbot use.

The centerpiece

Three systems, three different operating models

The mistake would have been one automation pattern applied three times. Each system earned a different level of autonomy, because each fails differently.

Read-only autonomous analysis

Google Performance Intelligence

Three specialist agents - Ads, GA4, Search Console - report to an orchestrator over MCP + OAuth. Cron-scheduled, fully autonomous, and physically unable to touch an ad account: permissions are read-only by design.

Human checkpoint: the external specialist executes. System →

Calibrated production automation

Multilingual Metadata Automation

Every field passed human Excel review until output survived 3-4 calibration runs per language unchanged. Only then was approval removed. Autonomy here was granted by evidence, not by optimism.

Human checkpoint: removed after calibration. System →

Human selection among alternatives

Semantic Commerce Architecture In progress

Roughly three competing model prototypes per page - the models genuinely interpret the brief differently. The human selects, critiques and directs; agents execute and iterate. Judgment never leaves the loop.

Human checkpoint: every page. System →

  • Input
  • Agent
  • Orchestrator
  • Governance gate
  • Human judgment
  • Autonomous action
The shared pipeline is the same; what differs is where the governance gate routes. That routing is the design decision.

Ownership

Who owns what

I owned

  • Operating model & architecture
  • Agent responsibilities & prompts
  • Orchestration design
  • Permission boundaries
  • Calibration & quality rules

AI / system executed

  • Data collection & analysis
  • Report generation
  • Metadata production
  • Prototype generation & iteration

External partner

  • Advertising execution authority - intentional
  • Budget & account changes stay human

Outcome & reflection

Autonomy is earned, not assumed

The measured leverage is real but deliberately bounded: 220+ hours of avoided metadata production, an estimated ~9.5 analytical hours saved weekly, an estimated four-plus working days saved per four-language hub. No financial attribution is invented - the output informs budget reallocations, performance investigations and optimization priorities, and that is the honest claim.

The less obvious result is organizational: attention moved from production to judgment. The interesting design question was never "what can AI do here?" but "where exactly should the human checkpoint sit?"

Systems & thinking