Questera

Implementation worksheet · 7 min read

A Lifecycle Experiment Prioritization Worksheet

Before scoring any idea on impact or ease, work out whether it can be read at all. Weekly eligible volume, the baseline rate and the smallest effect worth acting on give a rough sample size per arm (a rule of thumb for comparing two proportions, as presented by van Belle: about 16 × p(1 − p) ÷ difference², for 80% power and a two-sided 5% test), and from that the weeks to a readout. Ideas that can't read out within the quarter are marked 'not testable as designed': pool audiences, pick a more frequent outcome, or ship without a test and say so. Rank the testable ones on what changes if they win, build cost, customer risk and collisions with other tests.

Scope: the lifecycle lead or analyst ranking experiment candidates across journeys and channels. Starting state: a backlog of ideas scored on gut feel. Boundary: randomized tests of lifecycle messaging with a yes-or-no primary outcome; continuous outcomes such as revenue per user need a different formula. Intended outcome: a ranked queue in which every item has a decision attached and a realistic readout date.

Put it into practice

1. Write each idea as a decision

'If the day-2 push wins, we send it to all trials.' An idea with no decision attached produces a result nobody acts on. Drop it or rewrite it.

2. Use eligible volume per week, not segment size

Take the number of people who reach the journey step each week from its entry data. A segment of 200,000 feeding a step that 1,500 people reach a week is a 1,500-a-week test.

3. Set the smallest effect worth acting on

From the decision's cost and value, not from hope: the difference in percentage points below which you wouldn't ship even if it were real. It drives sample size more than anything else, because the required sample grows with the inverse square of the difference. Halve it and you need about four times as many people.

4. Estimate sample and weeks with the rule of thumb

Per arm, about 16 × p̄(1 − p̄) ÷ d², where p̄ is the average of the baseline and target rates and d is the difference between them. Double it for two equal arms and divide by weekly volume. Unequal splits, such as a 90/10 holdout, need more people in total. van Belle presents this as an approximation; in the examples below it lands within about 2% of the standard normal-approximation formula, which is close enough for ranking. Confirm the ideas you launch with a proper power calculation.

5. Mark what can't be read, then choose what to do

Past the quarter: pool similar audiences, choose a more frequent outcome, accept a larger minimum effect, or ship without a test, with a monitoring plan and a note that it was never tested. Don't run it underpowered to get a verdict. An underpowered 'no difference' is mostly a statement about sample size.

6. Rank the testable ideas

Decision value, build cost, risk to customers if the variant is bad, reversibility, and collisions: two tests changing the same journey step for the same people can't be read separately.

7. Worked example (illustrative, synthetic numbers)

Day-2 push for trials: 4,000 eligible a week, baseline 25%, worth acting on at +2 points: about 7,700 per arm, roughly four weeks. Win-back subject line: 1,500 a week, baseline 29%, +3 points: about 3,770 per arm, five weeks. Renewal-reminder timing: 120 a week, baseline 80%, +5 points: about 920 per arm and fifteen weeks, so not readable this quarter. SMS cart reminder: 900 a week, baseline 8%; +1 point needs about 28 weeks, but +2 points (about 3,280 per arm) needs about seven.

Experiment prioritization worksheet

Copy this structure into your review document and record your observed result for each row.

Experiment prioritization worksheet
CandidateDecision if it winsEligible per weekBaseline and smallest effectPer arm and weeks (rule of thumb)Status
Day-2 push for trialsSend it to all trials4,00025%, +2 pointsAbout 7,700; about 4 weeksTestable: rank 1
Win-back subject lineAdopt the winning line1,50029%, +3 pointsAbout 3,770; about 5 weeksTestable: rank 2
SMS cart reminderAdd an SMS step for opted-in users9008%, +2 pointsAbout 3,280; about 7 weeksTestable after the consent audit
SMS cart reminder at +1 pointAdd the SMS step for a smaller gain9008%, +1 pointAbout 12,400; about 28 weeksNot testable as designed
Renewal-reminder timingMove the reminder from 30 to 45 days out12080%, +5 pointsAbout 920; about 15 weeksNot this quarter: pool with monthly plans or ship with monitoring
Welcome email copy refreshReplace the current copy3,000Outcome is opensNot computedRedefine the outcome as activation first

A failure worth checking

The null that wasn't. A renewal-timing test tops the backlog on an impact score, runs for six weeks at 120 eligible a week and reports no significant difference, and the team concludes timing doesn't matter. With about 360 per arm and a baseline near 80%, the same rule of thumb says the test could reliably detect only differences of roughly 8 points or more. It said nothing about the 5-point effect the team cared about. Verification: before interpreting any null result, compute the effect the achieved sample could detect, and record the test as inconclusive, not negative, when that effect is larger than the one you cared about.

Common questions

Is the rule of 16 accurate enough?

For ranking, yes. It assumes equal arms, 80% power and a two-sided 5% test, and van Belle presents it as an approximation. Before launch, confirm with an exact power calculation for your actual split and any interim looks you plan.

Why not rank with an impact, confidence and ease score?

Use one if it helps the conversation, but after the readability check. A high-impact idea that can't produce a readable result this quarter isn't a test; it's a decision you'll make without evidence, and it's better to say so.

Basis and scope

This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.

Continue with Questera

Discuss your workflow →