Company Confidential (Enterprise CPG)
Role Product & Design Lead
Scope Agentic pipeline & supervision UI
Key outcome Launch cycle: weeks → days
B2B SaaS / Agentic Pipeline / Content Operations

The supervision interface for an agentic content pipeline.

One brief in, 36 retailer-ready packages out. Six AI agents do the repetitive work. Humans own every decision that ships.

ShelfFlow turns a single product brief into retailer-ready content through an orchestrated pipeline: AI agents parse, draft, adapt and flag; four mandatory human gates decide what advances. The product problem wasn't generating content. It was making agentic work supervisable at enterprise scale.

I owned the product end to end: strategy and pipeline architecture, agent orchestration boundaries, state machine design, decision authority mapping and the full operator interface.

6Embedded AI agents, zero autonomous shipping
4Mandatory human gates, no bypass
36Content packages from 1 brief
18Granular human approvals per launch
ShelfFlow platform overview
ShelfFlow Commerce Workspace dashboard

The supervision surface. Operators track active briefs, pipeline status and AI agent activity across all launches from a single screen.

01 The Agentic Pipeline

Teams weren't failing at content creation. They were drowning in coordination. The fix wasn't a faster content tool: it was decision architecture with AI embedded inside it.

A single product launch at an enterprise CPG company means 6 SKUs × 3 retailers × 2 variants = 36 content packages, each with its own spec requirements, character limits and compliance rules. Producing this by hand took weeks: versioning lived in file names, approvals moved through email threads no one could reconstruct, and generative AI existed but wasn't trusted because nothing defined what it was allowed to decide.

Faster drafting wouldn't have solved it. The strategic insight: design the decision architecture first, then embed AI as an accelerator within it. ShelfFlow structures the pipeline as 6 stages with 4 mandatory human gates. Agents operate freely between gates. Nothing crosses a gate without a human.

Evidence Stakeholder interviews revealed that roughly 70% of launch delays originated not in content creation, but in approval ambiguity. "Who needs to sign off on this?" was the most frequently asked question. The workflow needed explicit gates with clear ownership, not faster drafting tools.

Six-stage pipeline

S1Brief IntakeProduct brief parsed, SKU matrix generated
S2AI Draft Generation◆ Gate 1: Brief approval
S3Human Review & Edit◆ Gate 2: Draft quality
S4Retailer AdaptationSpecs, character limits, compliance
S5Compliance & Legal◆ Gate 3: Compliance sign-off
S6Publish & Distribute◆ Gate 4: Final launch approval
02 Agent Orchestration

Six AI agents, each with a bounded scope. Every agent accelerates work. None can ship content.

ShelfFlow embeds 6 specialised agents across the pipeline, coordinated by a workflow orchestrator that routes tasks, schedules generation jobs and enforces gate rules. Each agent operates within a clearly defined boundary: it can draft, suggest, check or flag, but it cannot approve or publish. No agent can bypass a gate.

The agents absorb the repetitive, high-volume work that made manual workflows collapse: parsing briefs into structured SKU matrices, generating retailer-adapted copy within character limits, cross-checking claims against compliance databases, detecting contradictions across 36 parallel packages. The decisions that carry legal and brand accountability remain with humans.

Embedded AI agents

BRIEF PARSERExtracts SKU matrix, claims, and key messages from product brief
COPY DRAFTERGenerates retailer-adapted copy per SKU. Respects character limits and tone.
SPEC CHECKERValidates against retailer spec sheets. Flags violations before human review.
COMPLIANCE SCANNERChecks claims against regulatory databases. Flags risk, never clears it.
CONSISTENCY AUDITORCross-checks messaging across all 36 packages. Detects contradictions.
DIFF REPORTERSummarises what changed between versions. Enables faster re-approval.

Every agent drafts or flags. No agent approves. Zero autonomous shipping.

Generate content empty stateS4 · Generate / select SKUs, trigger agents
Building PDPs in progressS4 · Generating / A/B variants per retailer

Agents at work. The operator triggers generation and supervises progress in real time: per-package status, confidence signals and flags surfaced as they happen.

AI LAYER — SIX AGENTS, ONE ORCHESTRATOR the orchestrator routes tasks and enforces gate rules · agents draft, adapt, scan and flag — never approve, never publish BRIEF PARSER COPY DRAFTER DIFF REPORTER SPEC CHECKER COMPLIANCE SCANNER CONSISTENCY AUDITOR — cross-checks all 36 packages 1 brief S1 · Intake S2 · Draft S3 · Review S4 · Adapt S5 · Legal S6 · Publish G1 · brief G2 · draft G3 · compliance G4 · launch 36 packages18 approvals HUMAN AUTHORITY — NOTHING SHIPS WITHOUT A GATE OWNER CONTENT LEAD · G1 G2 COMPLIANCE OFFICER · G3 BRAND MANAGER · G4 OPERATOR · runs the workflow no gate authority generate & validate risk & consistency flags change reporting mandatory human gate — no bypass
AI AGENTS HUMAN GATES 1 brief S1 Intake S2 Draft S3 Review S4 Adapt S5 Legal S6 Publish BRIEF PARSER COPY DRAFTER DIFF REPORTER SPEC CHECKER CONSISTENCY AUD. COMPLIANCE SCAN CONTENT LEAD G1 · G2 COMPLIANCE OFF. G3 BRAND MANAGER G4 · launch 36 packages · 18 approvals generate & validate risk flags change reports mandatory human gate — no bypass OPERATOR runs the workflow · no gate authority

Agent orchestration and authority in one map. Six agents work the lanes between gates; four humans own the gates. The vertical axis is the trust architecture: AI drafts and flags above the line, human accountability signs below it. The operator runs the workflow but holds no gate authority.

03 Human Gates

Four gates, four roles, one question the system can always answer: who approved this?

In enterprise commerce, shipping wrong content is more expensive than a delayed launch: incorrect claims create regulatory exposure, mismatched specs trigger portal rejections. So the counterweight to agent speed is a hard gate model. Four mandatory gates, no bypass, each owned by a role with explicit authority enforced by the system, not by process docs. A Content Lead approves briefs and drafts but can't sign off compliance. A Compliance Officer owns Gate 3 but has no launch authority.

Authority mapping

CONTENT LEADOwns briefs, approves drafts (G1, G2)
COMPLIANCEOwns regulatory review (G3)
BRAND MANAGERFinal launch authority (G4)
OPERATORManages workflow, no gate authority

The decisions below defined the boundary between agent autonomy and human authority. Each one responded to a specific failure mode observed during research.

DecisionChosenRejectedWhy
AI authority Draft-only (never ships) Auto-publish with confidence threshold Enterprise compliance requires human sign-off on every claim that reaches the shelf. AI that ships autonomously creates regulatory liability. The value of AI here is speed-to-draft, not autonomy.
Gate model 4 mandatory gates, no bypass Flexible approval chains Flexibility introduces ambiguity. When a retailer or legal team asks "who approved this?", the system must return one unambiguous answer. Configurable chains make that impossible to guarantee.
Staged generation Brief → Draft → Adapt → Review Single-step generation Generating retailer-adapted content in one pass produces output that looks right but fails spec validation. Staged generation lets operators catch problems at each layer before they compound.
Batch operations Batch approve only after individual review Unrestricted batch approval v1 allowed batch-approving packages that hadn't been individually reviewed. Testing revealed operators rubber-stamped to save time. Added a "reviewed" flag requirement. Speed without review is liability.
AI AGENTS — draft, adapt, flag never approve · never publish Operator runs the workflow · no authority Content Lead owns G1 brief · G2 draft Compliance Officer owns G3 compliance Brand Manager owns G4 · launch hand-off at the gates NOTHING SHIPS WITHOUT A GATE OWNER
AI AGENTS — draft & flag never approve · never publish Operator runs workflow · no authority Content Lead owns G1 brief · G2 draft Compliance Officer owns G3 compliance Brand Manager owns G4 · launch

Decision authority map: who owns which decisions across the pipeline. Agents act autonomously between gates. Humans own every approval that ships.

04 Supervision & States

Every package is a state machine. Rejection isn't an error state. It's a structured feedback loop back to the agents.

Each of the 36 packages runs as an independent state machine: Draft, In Review, Approved, Published, Rejected or Blocked. Nothing else. Every transition requires an explicit human action or a system event, and every change is logged with who triggered it and why. This replaced the ambiguous in-between states of the old workflow: "probably approved", "waiting on someone", "I think legal saw it". Rejected packages route back to the previous stage carrying reviewer notes, which agents use to regenerate. Supervision, not babysitting.

Content package states

DraftAI-generatedAwaiting review
In ReviewHuman editingOwner assigned
ApprovedGate passedReady for next stage
PublishedRetailer-liveAudit trail complete
RejectedGate failedReturns to previous stage with notes
BlockedDependency unmetWaiting on external input
System rationale Every state transition is logged: who, when, why. Rejected packages carry reviewer notes back to the previous stage. No silent failures. No packages stuck in limbo. The state machine is the accountability layer.
System rationale How we measured the agents. Agent quality was tracked against human outcomes, not benchmarks: flag precision (how often a compliance flag survived human review), per-gate rejection rate as a drift signal, regeneration acceptance (how often a v2 shipped without further notes), and a review-time floor per cell — the rubber-stamp detector that triggered the batch guardrail. When rejection rates fell too low, we treated it as reviewer fatigue, not agent excellence.
Brief validation side-by-side diff

Gate 1: Brief validation. Side-by-side diff showing user brief vs. AI-enriched brief. Human confirms or edits before lock.

Review and approve matrix

Gate 4: Full matrix review. Per-cell approval across 18 human decisions, every SKU × retailer combination. "Approve for Launch" is the final gate.

◆ INTERACTIVE — SYNTHETIC DATA

Hold Gate 4 yourself.

A working slice of the decision architecture: approve variants, reject and watch the feedback loop, and try to batch-skip the review — the guardrails are live.

Open full screen ↗

05 Operational Scale

36 packages. 18 approvals. 200+ state transitions. The combinatorial reality that makes manual workflows collapse.

The core complexity of ShelfFlow isn't any single content package. It's the combinatorial explosion when you multiply SKUs by retailers by variants. This is where manual workflows break. Not at package #1, but at package #27, when reviewer fatigue sets in and the differences between Amazon and Walmart specs blur together. ShelfFlow's job is to make package #36 as reviewable as package #1: urgency signals, comparison tools and batch operations all exist to fight decision quality decay.

Scale math 6 SKUs × 3 retailers × 2 variants = 36 packages. Each package passes through 4 gates. That's 144 gate evaluations per launch. With roughly 50% requiring at least one revision cycle, the system manages 200+ state transitions per product launch. Without structured workflow, this volume becomes untrackable within the first week.
Edge cases and constraints The system had to handle partial approvals (2 of 3 retailers ready, 1 still in review), mid-cycle brief changes (product claims updated after generation), and cross-package dependencies (a compliance flag on one SKU potentially affecting all packages sharing the same claim). These weren't theoretical scenarios. They were weekly occurrences in the existing workflow.
1 brief ×6 SKUs ×3 retailers ×2 variants 36 packages 18 decisions launch every multiplication is content · every decision is human
1 brief ×6 SKUs ×3 retailers ×2 variants 36 packages 18 decisions launch

The content matrix explosion: 1 brief → 6 SKUs × 3 retailers × 2 variants = 36 content packages → 18 approval decisions → 1 launch.

Amazon PDP previewAmazon · Full-fidelity PDP preview
Sephora PDP previewSephora · Adapted to retailer format

Same SKU, different retailers, completely different specs. Previews render in each retailer's actual page layout so operators review what the customer will see, not an abstraction.

06 Reflection

The hardest problem wasn't the AI. It was making 18 approval moments feel fast instead of bureaucratic.

Five operational guardrails constrained every decision: AI assists, never ships. Every transition is auditable. The pipeline is the single source of truth: no side channels, no email approvals. Design for reviewer fatigue, not enthusiasm. Absorb the combinatorial complexity so humans can focus on judgement calls.

The pattern AI in enterprise workflows fails when it's positioned as a replacement for human judgement. It succeeds when it's positioned as an accelerator within an explicit decision architecture. The value of ShelfFlow isn't that AI writes copy. It's that the system makes the coordination overhead of 36 packages manageable, traceable and auditable. The workflow is the product. The AI is infrastructure.
What failed v1 had 6 decision gates, one per stage. Testing revealed that gates at Brief Intake and Retailer Adaptation added friction without adding value. Content leads were rubber-stamping them within seconds. I reduced to 4 gates positioned where genuine decision-making happens. The lesson: more governance doesn't mean better governance. Gates only work when they protect against real risk.
What success looks like A content team launches a new product across 3 retailers in days instead of weeks. Every approval has a timestamp and an owner. Every AI-generated draft has been reviewed by a human before it reaches a retailer portal. When compliance asks "who approved this claim for Walmart?", the answer takes seconds, not a forensic email search. The system doesn't make the work disappear. It makes the work structured enough to scale.
Next case study KODO Arbitrage →
×