Implementation worksheet · 7 min read
A Retention Uplift Claim Verification Worksheet
Check the claim in order and stop at the first check it fails: what 'retained' means (event, unit, window), what it was compared with, whether the group was selected for unusually low recent activity, whether the metric and window were fixed before the results, the absolute difference with an interval, which base any percentage refers to, whether the effect lasts to the retention window that matters, and whether the comparison group was reached some other way. At-risk and win-back programs are the most exposed, because their entry rule selects people at a low point, and people selected at a low point tend to drift back toward their usual level whether or not you message them.
Scope: an analyst, lifecycle lead or buyer reviewing a retention claim from an internal team, an agency or a vendor. Starting state: a slide with one number. Boundary: claims that lifecycle messaging caused re-engagement or retention; not forecasts. Intended outcome: one of four verdicts (supported, supported with rewording, unverified, contradicted) with the reason written down. Unverified is not false; it means the evidence offered can't carry the claim.
Put it into practice
1. Restate the claim with a unit, window and population
'Improved retention by 35%' becomes: which customers, what counted as retained, over how many days, and 35% of what. This step often changes the claim's shape: the restated version can turn out to be a return rate among recipients rather than a difference between two groups.
2. Identify the comparison
From strongest to weakest: a concurrent random holdout drawn from the same eligible group; a staggered or matched comparison; a before-and-after on the same group; recipients compared with non-recipients. The last is not a comparison of the program at all, because who received or opened a message was not random.
3. Check for selection at a low point
If people entered because their activity was unusually low (no session in 21 days, no order in 60), expect some of them to come back on their own. Barnett and colleagues describe regression to the mean as natural variation that can look like real change, most noticeable when follow-up is measured on a subgroup selected by its baseline value. A holdout drawn from the same selected group drifts the same way, which is why the holdout difference, not the return rate, is the effect.
4. Check that the metric was fixed before the results
Ask when the definition, window and segment were chosen, and how many alternatives were examined. If three windows and five segments were looked at and the best one reported, the headline is the largest of fifteen noisy numbers, and a number picked for being largest is biased upward. Ask to see all fifteen.
5. Convert to an absolute difference with an interval
Points, not percent, with a 95% interval and both arm sizes. A relative figure is only readable beside its base: '+6 points from 29%' and '21% relative' describe the same result, and only the first lets a reader picture it.
6. Check persistence at the window that matters
A win-back message can bring someone back for a session without changing whether they are still a customer at day 90. If the claim says retention, measure retention. Effects of something new can fade over time, so a 30-day return effect doesn't establish a 90-day one.
7. Check the holdout stayed clean
Did support, sales or another journey contact the held-out customers? Contact that reached them through another route pulls their outcome toward the treated group's, which tends to shrink the measured difference. Record it rather than discarding the test.
8. Worked example (illustrative, synthetic numbers)
Claim: 'Win-back journey brought back 35% of at-risk customers, improving retention by 35%.' Eligible: customers with no session in 21 days; 6,000 received the journey and 1,500 were held out at random. Returned within 30 days: 2,100 (35.0%) versus 435 (29.0%), +6.0 points, interval roughly +3.4 to +8.6. Still active at day 90: 19.5% versus 18.0%, +1.5 points, interval roughly -0.7 to +3.7. Verdict: supported with rewording. 'The journey raised 30-day return by about 6 points; a 90-day retention effect was not demonstrated.'
Retention claim verification worksheet
Copy this structure into your review document and record your observed result for each row.
| Check | Question | Illustrative finding | Verdict for this check |
|---|---|---|---|
| Definition | What counts as retained, for which unit, over what window? | 'Returned' means one session within 30 days; the claim says 'retention' | Reword: re-engagement, not retention |
| Comparison | Compared with what? | Random 20% holdout from the same at-risk group | Pass |
| Selection | Were people selected for unusually low recent activity? | Yes: 21 days without a session | Return rate is not the effect; use the holdout difference |
| Pre-registration | Fixed before results? How many alternatives were examined? | Three windows and five segments examined; best one reported | Fail until all fifteen are shown |
| Absolute effect | Difference in points, with interval and arm sizes | +6.0 points (+3.4 to +8.6); 6,000 vs 1,500 | Pass |
| Percentage base | What is the '35%' a percentage of? | The treated group's return rate, not a lift | Fail as worded |
| Persistence | Does the effect hold at the retention window? | +1.5 points at day 90 (-0.7 to +3.7) | Not demonstrated |
| Contamination | Was the holdout reached another way? | Support's manual outreach reached both groups at similar rates | Recorded; no adjustment |
| Overall | Which verdict, and what sentence is supportable? | 'Raised 30-day return by about 6 points' | Supported with rewording |
A failure worth checking
The flat reference group, used in either direction. One reviewer is shown that non-targeted customers 'stayed flat' while the targeted group 'rose', and accepts the claim; another notices the non-targeted group rose too, and rejects it. Both skipped a step. A reference group selected differently can only be compared on magnitude: if the targeted group moved 6 points and the reference group 0.4 points over the same window, the targeted group changed more. That is a description, and it establishes cause in neither direction. Only a random holdout from the same eligible group answers the causal question. Verification: recompute the holdout difference from raw data, starting from everyone assigned, and confirm that assignment happened before anyone was messaged.
Common questions
What if there is no holdout at all?
Then the most the claim can say is 'coincided with', and at-risk programs need that caveat more than most, because regression to the mean produces an apparent recovery on its own. The honest verdict is unverified. The fix is to hold out a random slice of the next cohort.
Is a 30-day return effect worth anything if the 90-day effect wasn't demonstrated?
It can be, if returning sessions have value of their own, such as orders. But it's a different claim, and a program sold on retention should be judged on retention. 'Not demonstrated' isn't 'zero' either: an interval from -0.7 to +3.7 points leaves room for a small effect a larger test could find.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.