Implementation worksheet · 6 min read
A Customer Journey Before-and-After Evidence Pack
Assemble five things: both journey versions as exported definitions, one eligibility rule applied to both periods, a weekly outcome series covering several weeks on each side of the switch, a comparison series the change didn't touch, and a log of everything else that moved in the window. With those you can say whether the outcome changed at the switch by more than it was already moving, and by more than the comparison series moved. You can't say the redesign caused it. If a large decision rides on the answer, run the old and new versions side by side with random assignment instead.
Scope: a lifecycle team that replaced a live journey (onboarding, win-back, renewal) without holding anyone out, and now has to show what happened. Starting state: two dates and two averages. Boundary: one journey, one primary outcome, observational evidence only. Intended outcome: a pack a skeptical reader can check, graded by the strongest statement it supports.
Put it into practice
1. Export both versions instead of describing them
Entry rules, exit rules, every message with its channel and delay, and any branching. 'v2 is shorter and adds push' hides the detail that explains results later, such as an exit rule that now removes activated users a day earlier.
2. Hold eligibility constant across the switch
If v2 also changed who enters (a new signup source, a tighter filter), you are comparing two populations, not two journeys. Compare people who would have qualified under both rules, and report how many fell outside.
3. Build weekly series, not two averages
Weekly eligible counts and outcome rates for several weeks before and after, so a reader can see the trend and the normal week-to-week swing. Eight weeks on each side is a rule of thumb for weekly data, not a requirement; at minimum, cover one full cycle of anything that recurs, such as month-end billing.
4. Add a comparison series the change didn't touch
A segment or market that stayed on v1, or eligible users excluded from both versions. Compare the size of the change in each series over the same weeks. A bigger move in the journey's series is consistent with the redesign mattering; it doesn't prove it, because the groups differ in ways you didn't choose.
5. Keep a change log and check tracking continuity
Product releases, pricing, acquisition mix, holidays and, above all, changes to how the outcome is tracked. A tracking fix that ships the same week as the journey change can produce the entire apparent effect.
6. Grade the conclusion
Pick the strongest statement the pack supports: 'coincided with a change larger than the pre-period swing and the comparison series', 'coincided with a change within normal variation', or 'not interpretable', for example when tracking changed at the switch. Causal verbs are not on the list.
7. Worked example (illustrative, synthetic numbers)
Onboarding v1 (five emails over 14 days) was replaced by v2 (three emails, an in-app checklist and a day-2 push). Before: 9,600 signups over eight weeks, 25.2% activated within 14 days, weekly rates 24.1% to 26.3%. After: 9,800 signups, 27.9%, weekly 26.8% to 29.0%, so every week after the switch sat above every week before it. Partner-channel signups, excluded from both versions, went from 22.0% to 22.6%. Sampling noise alone puts the +2.7-point difference at roughly +1.5 to +3.9. Grade: coincided with a change larger than the pre-period swing and the comparison series' movement.
Before-and-after evidence pack
Copy this structure into your review document and record your observed result for each row.
| Section | What goes in | Illustrative content | Weak if |
|---|---|---|---|
| Version exports | Both journey definitions, exported | v1: five emails over 14 days; v2: three emails, in-app checklist, day-2 push | Only a summary sentence exists |
| Eligibility | One entry rule for both periods | Self-serve signups; partner channel excluded | The rule changed at the switch |
| Outcome definition | Event, window, unit and tracking source | Workspace activated within 14 days, from the product database | Tracking changed at the switch |
| Before series | Weekly counts and rates | Eight weeks, 9,600 signups, 25.2% (weekly 24.1% to 26.3%) | Only one average is shown |
| After series | Weekly counts and rates | Eight weeks, 9,800 signups, 27.9% (weekly 26.8% to 29.0%) | It stops after the first good weeks |
| Comparison series | An unexposed group over the same weeks | Partner-channel signups: 22.0% → 22.6% | There isn't one |
| Change log | Everything else that moved | Paid-search share of signups up from week 3 after; pricing page change in week 5 after | It is empty |
| Sampling noise | Interval on the period difference | +2.7 points, roughly +1.5 to +3.9 | It is presented as the whole uncertainty |
| Conclusion grade | Strongest statement the pack supports | Coincided with a change larger than pre-period variation and the comparison series | A causal verb is used |
A failure worth checking
The tracking fix on switch day. v2 goes live in the same week that an analytics fix stops a web bug which had been dropping some activation events. The weekly series shows a clean step up at the switch. The comparison series, measured from the same events, rises too, but by less, because partner signups are mostly on mobile. The pack looks strong and says nothing about the journey. Verification: compare event volumes by source and platform around the switch date, and recompute the outcome from a record the tracking change couldn't affect, such as bookings or payments in the product database.
Common questions
Does the sampling interval tell us how confident to be?
About one source of error only. It shows how much the difference would vary with different people in the same weeks. It says nothing about trends, mix changes or tracking, which is what the comparison series and change log are for. Read it as a lower bound on the uncertainty.
Can this pack replace a holdout next time?
No. It's the method for when a holdout wasn't possible. If the next redesign will drive a large decision, run both versions at the same time with random assignment, and keep the pack for changes that can't be split.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.