Questera

Implementation worksheet · 6 min read

A Board-Ready Lifecycle Experiment Reporting Outline

One page, in this order: the decision or approval you are asking for; the portfolio (tests launched, read out, shipped, stopped for harm, inconclusive, still running); the program-level effect measured against a standing holdout, with its interval; the shipped changes, each with its own interval; and the caveats. Never add up individual test lifts to claim a total. Tests selected because they won tend to overstate their own effects, a selection bias Lee and Shen describe for the aggregated impact of launched experiments, and separate changes to the same journey don't simply add. If there is no standing holdout, say plainly that the program's total effect is unmeasured.

Scope: the head of growth or lifecycle preparing a quarterly board or executive update on lifecycle experimentation. Starting state: a list of experiments with results, and pressure to show one big number. Boundary: one quarter's lifecycle testing across channels; brand and long-term effects are named as unmeasured rather than estimated. Intended outcome: a page directors can read in two minutes that still stands when someone checks it.

Put it into practice

1. Lead with the ask

What should the board note or approve? Keeping the standing holdout for another quarter, funding a data project that makes more tests readable, stopping a channel. A results page with no ask invites a debate about the numbers instead of a decision.

2. Show the whole portfolio

Launched, read out, shipped, stopped for harm, inconclusive, still running. Inconclusive and stopped tests belong on the page: they show the process rejects things, which is what makes the wins believable.

3. Report the program effect from a standing holdout

A random share of customers excluded from all non-essential lifecycle messaging, kept for the quarter or longer. The difference between them and everyone else, with an interval, is the program's effect. The holdout has a cost, because those customers miss messages that may help them; choose the smallest share that gives an interval narrow enough for the decision, and review it every quarter.

4. Show shipped changes with their intervals

Each shipped test with its absolute effect and 95% interval, re-measured after launch where you can. The estimate for a test that was shipped because it won is more likely to be too high than too low.

5. Name what was stopped and why

A test stopped for harm (unsubscribes doubling, a complaint spike) is evidence that the guardrails work. One line each.

6. State the caveats in plain words

The measurement window, effects that can fade once the novelty wears off, seasonality, and what isn't measured at all, such as brand effects or anything beyond the window.

7. Worked example (illustrative, synthetic numbers)

Q3: 14 tests launched, 11 read out, 4 shipped, 2 stopped for harm, 5 inconclusive, 3 still running. The four shipped tests' individual lifts sum to +9.6 points of 30-day activation. The standing 5% holdout (2,000 new signups against 38,000) shows 31.0% activation against 34.1% for everyone else: +3.1 points, interval roughly +1.0 to +5.2. The page reports +3.1, explains why it is smaller than +9.6, and asks the board to keep the holdout for Q4.

Board reporting outline

Copy this structure into your review document and record your observed result for each row.

Board reporting outline
SectionWhat it showsIllustrative Q3 entryDistortion to avoid
AskWhat the board should note or approveKeep the 5% standing holdout for Q4A results parade with no decision
PortfolioLaunched, read out, shipped, stopped, inconclusive, running14 / 11 / 4 / 2 / 5 / 3Listing only the wins
Program effectStanding-holdout difference with its interval+3.1 points of 30-day activation (about +1.0 to +5.2)Summing test lifts (+9.6 points)
Shipped changesEach shipped test with its intervalDay-2 push: +1.8 points (+0.4 to +3.2)Point estimates only
Stopped for harmWhat was stopped, and the signalSMS reminder cadence: unsubscribes doubledLeaving it out
InconclusiveTests that couldn't separate effect from noise5 tests, each with the effect it could have detectedReporting them as 'no effect'
CaveatsWindow, novelty, seasonality, what isn't measured30-day window; brand effects not measuredOmitting them
Method appendixDefinitions, holdout design, analysis methodOne paragraph and a link to the learning repositoryNothing a director could check

A failure worth checking

The additive slide. Four shipped tests with lifts of +3.4, +2.6, +1.8 and +1.8 points are summed to '+9.6 points of activation this quarter'. The next quarter activation is up about 3 points, and the board asks where the rest went. Three things took it: tests picked because they won tend to have overestimated effects, two of the four changed the same onboarding journey so their effects overlapped, and one effect faded after launch. Verification: report the program effect from the standing holdout, and don't publish the sum of test lifts at all.

Common questions

Is a standing holdout worth the lost conversions?

It costs something: held-out customers don't get messages that may help them. Keep it as small as still gives a usable interval, rotate who is held out if the cost keeps falling on the same people, and weigh it against running the program with no measure of what it adds.

Should the page show p-values?

Show intervals. A director can read '+3.1 points, somewhere between +1.0 and +5.2'; a p-value on its own says nothing about how big the effect is.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with Questera

Discuss your workflow →