Implementation worksheet · 6 min read
A Board-Ready Lifecycle Experiment Reporting Outline
One page, in this order: the decision or approval you are asking for; the portfolio (tests launched, read out, shipped, stopped for harm, inconclusive, still running); the program-level effect measured against a standing holdout, with its interval; the shipped changes, each with its own interval; and the caveats. Never add up individual test lifts to claim a total. Tests selected because they won tend to overstate their own effects, a selection bias Lee and Shen describe for the aggregated impact of launched experiments, and separate changes to the same journey don't simply add. If there is no standing holdout, say plainly that the program's total effect is unmeasured.
Scope: the head of growth or lifecycle preparing a quarterly board or executive update on lifecycle experimentation. Starting state: a list of experiments with results, and pressure to show one big number. Boundary: one quarter's lifecycle testing across channels; brand and long-term effects are named as unmeasured rather than estimated. Intended outcome: a page directors can read in two minutes that still stands when someone checks it.
Put it into practice
1. Lead with the ask
What should the board note or approve? Keeping the standing holdout for another quarter, funding a data project that makes more tests readable, stopping a channel. A results page with no ask invites a debate about the numbers instead of a decision.
2. Show the whole portfolio
Launched, read out, shipped, stopped for harm, inconclusive, still running. Inconclusive and stopped tests belong on the page: they show the process rejects things, which is what makes the wins believable.
3. Report the program effect from a standing holdout
A random share of customers excluded from all non-essential lifecycle messaging, kept for the quarter or longer. The difference between them and everyone else, with an interval, is the program's effect. The holdout has a cost, because those customers miss messages that may help them; choose the smallest share that gives an interval narrow enough for the decision, and review it every quarter.
4. Show shipped changes with their intervals
Each shipped test with its absolute effect and 95% interval, re-measured after launch where you can. The estimate for a test that was shipped because it won is more likely to be too high than too low.
5. Name what was stopped and why
A test stopped for harm (unsubscribes doubling, a complaint spike) is evidence that the guardrails work. One line each.
6. State the caveats in plain words
The measurement window, effects that can fade once the novelty wears off, seasonality, and what isn't measured at all, such as brand effects or anything beyond the window.
7. Worked example (illustrative, synthetic numbers)
Q3: 14 tests launched, 11 read out, 4 shipped, 2 stopped for harm, 5 inconclusive, 3 still running. The four shipped tests' individual lifts sum to +9.6 points of 30-day activation. The standing 5% holdout (2,000 new signups against 38,000) shows 31.0% activation against 34.1% for everyone else: +3.1 points, interval roughly +1.0 to +5.2. The page reports +3.1, explains why it is smaller than +9.6, and asks the board to keep the holdout for Q4.
Board reporting outline
Copy this structure into your review document and record your observed result for each row.
| Section | What it shows | Illustrative Q3 entry | Distortion to avoid |
|---|---|---|---|
| Ask | What the board should note or approve | Keep the 5% standing holdout for Q4 | A results parade with no decision |
| Portfolio | Launched, read out, shipped, stopped, inconclusive, running | 14 / 11 / 4 / 2 / 5 / 3 | Listing only the wins |
| Program effect | Standing-holdout difference with its interval | +3.1 points of 30-day activation (about +1.0 to +5.2) | Summing test lifts (+9.6 points) |
| Shipped changes | Each shipped test with its interval | Day-2 push: +1.8 points (+0.4 to +3.2) | Point estimates only |
| Stopped for harm | What was stopped, and the signal | SMS reminder cadence: unsubscribes doubled | Leaving it out |
| Inconclusive | Tests that couldn't separate effect from noise | 5 tests, each with the effect it could have detected | Reporting them as 'no effect' |
| Caveats | Window, novelty, seasonality, what isn't measured | 30-day window; brand effects not measured | Omitting them |
| Method appendix | Definitions, holdout design, analysis method | One paragraph and a link to the learning repository | Nothing a director could check |
A failure worth checking
The additive slide. Four shipped tests with lifts of +3.4, +2.6, +1.8 and +1.8 points are summed to '+9.6 points of activation this quarter'. The next quarter activation is up about 3 points, and the board asks where the rest went. Three things took it: tests picked because they won tend to have overestimated effects, two of the four changed the same onboarding journey so their effects overlapped, and one effect faded after launch. Verification: report the program effect from the standing holdout, and don't publish the sum of test lifts at all.
Common questions
Is a standing holdout worth the lost conversions?
It costs something: held-out customers don't get messages that may help them. Keep it as small as still gives a usable interval, rotate who is held out if the cost keeps falling on the same people, and weigh it against running the program with no measure of what it adds.
Should the page show p-values?
Show intervals. A director can read '+3.1 points, somewhere between +1.0 and +5.2'; a p-value on its own says nothing about how big the effect is.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.