01 — Included

What the retainer covers every month.

Hypothesis generation

A maintained backlog, re-ranked as results come in. Every hypothesis carries the metric it should move and the evidence behind it.

Variant build

Implemented against your real stack rather than mocked. This is the step that usually sits in a ticket queue for a week at a traditional agency.

Pre-launch QA

Cross-browser, cross-device, and tracking validation before a single visitor sees the variant. A test that fires wrong is worse than no test.

Launch and monitoring

Live in your platform with guardrail metrics watched from the first hours, so a variant that damages revenue is caught in hours rather than at readout.

Analysis and calling

Results called against pre-registered criteria, with the confidence and the caveats stated plainly.

Stakeholder reporting

Reporting your leadership can read without translation: what ran, what won, what it is worth, and what runs next.

02 — What you get that is unusual

Cycle time, reported.

The speed claim is the differentiator, so it carries an obligation: we measure it and show you.

Metric 01

Hypothesis to launch

Measured on every test from the first engagement onward, and reported alongside your results rather than quoted from a brochure.

Metric 02

Tests live per month

Against the number we modelled during qualification. If we are under it, that is a conversation, not a footnote.

Metric 03

Win rate and average lift

Tracked over the life of the program, because a high win rate with trivial lifts is a different problem from a low win rate with large ones.

On the velocity promise

Your traffic sets the ceiling on test velocity, not our capacity. We would rather commit to a realistic number modelled against your conversion volume during qualification than quote an impressive one and spend month three explaining why it never happened.

03 — Questions

About the retainer.

How many tests per month should we expect?

It depends far more on your traffic than on our capacity — a funnel that needs three weeks to reach significance sets its own ceiling regardless of how fast we build. We model the realistic number against your actual conversion volume during qualification and commit to it in writing rather than quoting a generic figure.

What happens when a test loses?

It gets documented and it informs the next hypothesis. A losing test that was well-built is not a wasted month; it is a ruled-out branch. The programs that fail are the ones that stop testing after a loss, or quietly stop reporting them.

Who calls the result?

We do, against criteria registered before launch. Deciding the stopping rule after seeing the data is the most common way experimentation programs fool themselves, and it is the fastest way to lose a CFO's trust permanently.

Can we pause?

Yes. Experimentation programs have natural seasonality, and running tests through a period when your traffic mix is abnormal produces results you cannot trust anyway.

Start

See what your funnel could sustain.

We will model the realistic test velocity against your actual conversion volume before proposing anything.