What the retainer covers every month.
Hypothesis generation
A maintained backlog, re-ranked as results come in. Every hypothesis carries the metric it should move and the evidence behind it.
Variant build
Implemented against your real stack rather than mocked. This is the step that usually sits in a ticket queue for a week at a traditional agency.
Pre-launch QA
Cross-browser, cross-device, and tracking validation before a single visitor sees the variant. A test that fires wrong is worse than no test.
Launch and monitoring
Live in your platform with guardrail metrics watched from the first hours, so a variant that damages revenue is caught in hours rather than at readout.
Analysis and calling
Results called against pre-registered criteria, with the confidence and the caveats stated plainly.
Stakeholder reporting
Reporting your leadership can read without translation: what ran, what won, what it is worth, and what runs next.
Cycle time, reported.
The speed claim is the differentiator, so it carries an obligation: we measure it and show you.
Hypothesis to launch
Measured on every test from the first engagement onward, and reported alongside your results rather than quoted from a brochure.
Tests live per month
Against the number we modelled during qualification. If we are under it, that is a conversation, not a footnote.
Win rate and average lift
Tracked over the life of the program, because a high win rate with trivial lifts is a different problem from a low win rate with large ones.
Your traffic sets the ceiling on test velocity, not our capacity. We would rather commit to a realistic number modelled against your conversion volume during qualification than quote an impressive one and spend month three explaining why it never happened.
About the retainer.
How many tests per month should we expect?
It depends far more on your traffic than on our capacity — a funnel that needs three weeks to reach significance sets its own ceiling regardless of how fast we build. We model the realistic number against your actual conversion volume during qualification and commit to it in writing rather than quoting a generic figure.
What happens when a test loses?
It gets documented and it informs the next hypothesis. A losing test that was well-built is not a wasted month; it is a ruled-out branch. The programs that fail are the ones that stop testing after a loss, or quietly stop reporting them.
Who calls the result?
We do, against criteria registered before launch. Deciding the stopping rule after seeing the data is the most common way experimentation programs fool themselves, and it is the fastest way to lose a CFO's trust permanently.
Can we pause?
Yes. Experimentation programs have natural seasonality, and running tests through a period when your traffic mix is abnormal produces results you cannot trust anyway.
See what your funnel could sustain.
We will model the realistic test velocity against your actual conversion volume before proposing anything.