Company RetailNorm
Role Founder · Product & Design
Key outcome 40x faster reporting
Commerce / Data Systems / End-to-End

Three platforms measure the same dollar three different ways. I built the common ruler.

RetailNorm normalizes cross-platform retail media data to one declared attribution standard, so mid-size agencies can compare what platforms designed to be incomparable. Founder-level ownership: research, validation, service design, interface, growth. Solo, AI-augmented build: validated concept to paying-user product in weeks.

40xReporting time: 2–4 h → 3 min
18.6%Median attribution inflation exposed
$200–600/moPre-sold before a line of code
4 weeksEvidence gathered before building
RetailNorm Normalization Engine
Try the live productUpload sample CSVs. Watch the engine work.
Open RetailNorm →
01 Context

Every platform grades its own homework. Agencies compare the grades.

Amazon Ads reports ROAS on a 14-day last-click window. Walmart Connect uses 30-day multi-touch. Criteo uses 7-day first-click. Each number is internally correct and externally meaningless: the measurement system changes with the party being measured, and the party being measured chose the system.

Mid-size agencies sit downstream of this with no tooling. Enterprise suites that solve it start at $3–10K/month; their real alternative is a planner, Excel, and a personal macro nobody else trusts. Budget decisions worth hundreds of thousands of dollars flow through that macro every Monday.

Scope Role: Founder & Product Designer
Scope: Full product (research to production)
Users: Mid-size media agencies
Work: Research, service design, IA, interface, growth, AI execution
One campaign, three rulers illustrative values · mechanics are real
Same $10K campaign, same 500 conversions, reported by three platforms. Every number below is "correct" inside its own measurement system.
Amazon Ads
14-day · last-click
ROAS 4.2
longest self-credit window
Walmart Connect
30-day · multi-touch
ROAS 5.1
credits partial touches
Criteo
7-day · first-click
ROAS 3.6
strictest of the three
Normalized · 7-day last-click, one ruler for all three
3.6
Amazon · −14%
3.4
Walmart · −33%
3.6
Criteo · unchanged
Normalized numbers are almost always lower, and that is the point: the gap between reported and normalized is the attribution inflation agencies were unknowingly paying for. Budgets move on comparability, not on whichever platform grades its own homework most generously.
"I need accurate cross-platform performance without spending half my Monday in Excel." The functional job. The real job was emotional: stop feeling like a spreadsheet operator in front of senior clients.
Persona + experience map
Ecosystem map
02 The Real Problem

A measurement problem, misdiagnosed as a workflow problem.

Everyone downstream treated this as an efficiency problem: planners just needed to be faster in Excel. That misdiagnosis is why it persisted. You cannot spreadsheet your way out of three incompatible attribution systems; the comparison itself is invalid before the first formula runs.

Framing it as a measurement-standards problem changed what needed to be built. Not a faster report generator: a normalization engine with one declared standard, and enough transparency that agencies would trust numbers that are consistently lower than the ones platforms hand them.

The constraint set was unusual and shaped everything: solo operator, no engineering team, mid-size customers priced out of enterprise tooling. Every architectural decision had to be maintainable by one person, which turned architecture into a design decision rather than an engineering detail.

Market positioning

Enterprise tools: $3K-10K/mo. Mid-size agencies: priced out. Their alternative: Excel. RetailNorm sits in the gap nobody was serving.

Competitive UX audit
03 Key Insight

Agencies don’t buy tools. They buy deliverables.

Concierge validation, manually delivering normalized reports before any product existed, surfaced it: agencies never asked about the dashboard. They asked when the next PDF was coming. The report is the unit of value; the interface is how the report gets made. Starting with the dashboard would have meant building the wrong product with full confidence.

The second insight is less comfortable: a normalizer that only ever outputs lower numbers is, commercially, a bad-news machine. It survives only if it can justify itself. That made explanation infrastructure, confidence scores, data flags and an auditable technical view, first-class product rather than documentation.

Evidence 4-week validation sprint. 8 interviews (6+ confirmed pain). Landing page got 20+ signups in 5 days. Concierge MVP: agencies requested normalized reports again the following week. 3+ agencies committed at $200–600/month before a product existed. Every feature in v1 traces to one of these signals. Nothing shipped on assumption.
04 What I Changed

What I built, and the order I refused to change.

1. A normalization engine with one declared standard

Upload CSVs from any platform, get comparable ROAS in seconds, normalized to 7-day last-click: deliberately the strictest of the three. Every normalized number lands lower than what agencies are used to seeing. That is not a side effect to soften; the gap between reported and normalized is the product’s clearest proof of value, and explaining it is the strongest trust moment in the entire experience.

2. One engine, two cognitive loads

Directors and planners need the same truth at different depths. The executive view gives three numbers and one plain-language insight; the technical view exposes decay parameters, z-scores and corrections, auditable line by line. Same engine, same confidence score in both: trust is never view-dependent.

3. Uncertainty as a visible property

Every analysis carries its confidence and its flags: “71%, 2 flags.” The counterintuitive bet of the product, showing imperfection to build credibility, and the one users cite most. Polished certainty is what the platforms sell. This is the alternative.

4. The PDF as the retention mechanism

Agencies judge the product by the artifact that reaches their client. If it looks like an Excel export, they rebuild it in PowerPoint and the tool becomes a data source. The report template got institutional-grade typography and structure because report quality, not feature depth, is what brings an agency back next Monday.

Branded report

This PDF is not a feature. It is the reason agencies come back every Monday.

Evidence before code 4-week sequence · each rung gates the next
01 8 problem interviewsis the pain real, weekly, and worth money? gate: 6+ confirmed pain → passed
02 Landing page, no productdoes the framing pull strangers? gate: 20+ signups in 5 days → passed
03 Concierge MVP: reports delivered by handwhat do they actually value? Answer: the artifact, not the tool gate: unprompted repeat requests → passed
04 Pre-saleswill they pay before it exists? gate: 3+ agencies at $200–600/mo → passed
05 Build: 4 must-haves onlyengine · dual view · confidence · PDF. Everything else waits for retention data live in weeks
The ladder only moves one way: no rung is skipped and no feature ships on assumption. The most consequential output of weeks 1–4 wasn't a spec, it was the discovery that the PDF is the product — made before opening a design tool.
The Monday morning, compressed
Before · manual normalization2–4 h per client, weekly
Export3 platforms, 3 formats Cleanheaders shift per report Mergefragile personal macros Adjusthand-tuned windows Sanity-checknobody trusts step 4 Rebuild deckevery single week
After · RetailNorm~3 min per client
UploadCSVs as exported, zero config Reviewnormalized ROAS + confidence + flags Exportclient-ready PDF
Everything outside Upload → Review → Export is progressive disclosure. The compression isn't automation of the old six steps: steps 2–5 stop existing because normalization happens inside the engine, on one declared standard, instead of inside a planner's macro.
Normalization pipeline
Architecture

Deliberately simple. The architecture is the feature.

Design system
Engine
Simulator
Alerts
Reports
One engine, two cognitive loads
Normalization engine
Every upload runs once, to one declared standard: 7-day last-click, the strictest of the three. Both views read the same result.
Executive view
for the director walking into the client call
  • 3 numbers, no jargon
  • one plain-language insight
  • the answer, not the math
Technical view
for the planner who has to defend the numbers
  • decay parameters per platform
  • z-scores and applied corrections
  • the math, auditable line by line
Both viewsshow the same confidence score and data flags — e.g. "71%, 2 flags". Uncertainty is never view-dependent: the director and the planner disagree on depth, never on trust.
A product that only ever outputs lower numbers than the platforms survives on one thing: being able to explain itself at whatever depth the reader needs. The dual view is that explanation, packaged twice.
Intelligence ships in stages, retention gates each one north star: reports exported / week / active user
live
Normalize
One ruler across platforms. Confidence scores and flags on every analysis.
gate to next: weekly export habit holds
next
Explain
Why the numbers moved: anomaly callouts and plain-language deltas inside the report.
gate to next: explanations survive client meetings
later
Recommend
Budget reallocation suggestions. Only credible once the first two layers earned trust.
requires trust the earlier layers built
The order is the strategy. Recommendations from a tool nobody trusts are noise; trust compounds from normalize → explain → recommend, and each layer must prove retention before the next consumes build time a solo operator doesn't have.
05 Key Decisions

What trade-offs did I make?

Every decision involved choosing between conventional wisdom and what the research actually showed:

DecisionChosenRejectedWhy
Normalization baseline 7-day last-click (strictest) 14-day (Amazon default) Strictest standard means every number drops, which forces the “why is this lower?” conversation: the moment the product earns trust or dies
Uncertainty display Visible confidence scores (71%) Hidden/clean 100% Transparent imperfection more credible than polished certainty
Architecture Single HTML file, no framework React + microservices One maintainer. A fix ships to production in 30 seconds. Every added layer is a liability I’d carry alone: architecture chosen as a design constraint, not a stack preference
Data input CSV upload API integrations OAuth, credentials, rate limits would triple scope. CSV takes 15 seconds. APIs are v2
06 Results

Comparability, priced for the agencies Excel was failing.

40xReporting time reduction (2–4 h → 3 min)
18.6%Median attribution inflation surfaced per analysis
3+Agencies paying $200–600/mo, pre-sold
71%Average confidence score, shown, not hidden
Validation Live in production at retailnorm.com. Concierge validation before any code. Pre-sales at $200–600/month before product existed. Growth system: cold email pipeline (38% open target), SEO glossary pages, 3 retention loops (weekly habit, multi-client expansion, report-driven word-of-mouth).

What RetailNorm is not: an attribution truth machine. No tool can recover what the platforms don’t expose. What it does is narrower and more defensible: one consistent ruler, applied identically to every platform, with its error bars showing. Agencies don’t need perfect attribution to allocate budget well. They need numbers that are wrong in the same way everywhere.

07 Takeaway

The most impactful work was deciding what I refused to build.

The pattern Design the output artifact before the interface. Make uncertainty visible instead of hiding it. Validate willingness to pay before opening a design tool. Treat architecture as a design decision when you’re the one maintaining it. And sequence trust deliberately: normalize before you explain, explain before you recommend.
What failed First version was an API (upload CSV, get JSON), but agencies don’t use endpoints. They use screens on Monday morning. Pivoted to dashboard-first in sprint 1. Also budgeted 1 week for CSV parsing, took 3 because Amazon headers change by report type. Also dark theme default was designer preference, not user context. Users needed light mode for client presentations.
Try the live productUpload sample CSVs. Watch the engine work.
Open RetailNorm →
Next case study GetYourGuide Checkout →
×