Five steps, and who owns each one.
-
01
Research & hypothesis
Analytics, session recordings, and competitive teardown synthesized into ranked hypotheses. AI assistance compresses the synthesis — reading hours of session data and surfacing patterns — while the judgment about what is worth testing stays human.
-
02
Variant build
Implemented against your real stack rather than mocked. This is the single largest source of delay in a traditional agency retainer, and the place where AI assistance produces the biggest compression.
-
03
Pre-launch QA
Cross-browser, cross-device, and tracking validation. Generating and running the check matrix is mechanical work that AI does faster and more exhaustively than a person under deadline.
-
04
Launch & monitor
Live in your platform with guardrail metrics watched from the first hours. Deciding to kill a test that is damaging revenue is a judgment call and stays human-owned.
-
05
Read & report
Called against pre-registered criteria. Statistical interpretation and the decision about what it means for the next hypothesis are the parts we will not hand to a model.
The rules we hold even when they cost us.
Pre-register the stopping rule
The criteria for calling a test are agreed before launch. Deciding when to stop after seeing the data is the most common way experimentation programs fool themselves, and it destroys credibility with finance permanently.
Report the losses
A program that only surfaces wins is not measuring, it is marketing. Losses are how a roadmap gets re-ranked, and hiding them makes the next hypothesis worse.
Validate tracking first
A meaningful share of engagements find a conversion event firing incorrectly. Every result computed before that discovery is worthless, so it gets checked before anything is launched.
Guardrails on every test
A variant that lifts signups while damaging refunds or activation is a loss. Tests carry guardrail metrics so a local win that hurts the business is caught rather than celebrated.
Cycle time — hypothesis to launch — is tracked on every engagement from day one and reported to you alongside results. Until a program has enough logged cycles to publish a defensible number, we describe the speed qualitatively rather than inventing a figure. If a CRO agency quotes you a precise average cycle time on a first call, ask them for the distribution.
See it run on your funnel.
The fastest way to judge a method is to watch it applied to something you know well.