Implementation worksheet · 6 min read
A Marketing Agent Change-Management Checklist
Treat every change to a marketing agent as a release: a new model version, edited instructions, a new tool or permission, a new data source, a changed guardrail, updated brand or product facts. Each gets a change record, a test chosen by change type, a staged rollout and a rollback trigger written in advance. Instruction and model changes are replayed against a frozen set of past cases and compared with the current version's output; permission and tool changes go through the permission matrix before any test; data-source changes need freshness and consent checks. No change goes straight from edit to full autonomy: shadow first, then review-required, then the level the agent had before.
Scope: the team running an AI marketing agent in production that drafts copy, proposes segments or schedules sends within a permission matrix. Starting state: the agent works, and changes to it are made in a settings screen by whoever notices a problem. Boundary: changes to the agent itself, not the campaigns it produces or the incidents it causes, which have their own processes. Intended outcome: every change recorded, tested in proportion to its risk, and reversible.
Put it into practice
1. Define what counts as a change
Model version, system instructions and prompts, tools and their permissions, data sources, guardrail thresholds such as frequency caps or quiet hours, and the brand and product facts it draws on. Include changes you didn't make: if the provider can update the model behind the name you call, pin a version where the provider allows it, and treat any unpinned update as a change.
2. Write the change record before the change
What changes, why, who approved it, the version identifiers before and after, the planned test and the rollback step. A change that can't be written down this way isn't ready.
3. Keep a frozen replay set
A few hundred past cases (briefs, segment requests, send decisions) with the outputs your team approved, plus every edge case behind a past incident: missing time zones, suppressed contacts, sensitive segments, expired offers. Freeze it so results from different changes are comparable, and add to it deliberately.
4. Test by change type
Rule checks run automatically on every replay case: no suppressed contact included, frequency caps respected, no forbidden claims. Quality gets a blind comparison on a sample, current output against new, with the version hidden from reviewers. Permission and tool changes need a permission-matrix decision before any test runs.
5. Roll out in stages
Shadow: the new version proposes while the current one acts, and the differences are reviewed. Then review-required for a set period or number of actions. Then the level the agent had before. Record the version on every action, so the audit trail shows which version did what.
6. Write the rollback trigger first
A specific condition, such as any rule-check failure, any send outside the approved segment, or complaints above the pre-change level, and the person who can pull it without asking anyone. A trigger debated after the problem appears is a trigger pulled late.
7. Worked example (illustrative, synthetic numbers)
A model upgrade is proposed. On a 200-case replay set the new version excludes suppressed contacts in 200 of 200 cases, but schedules two sends inside quiet hours where the contact's time zone was missing and it defaulted to UTC. The rollout is blocked; the fix holds any contact without a time zone for review, and the rerun passes 200 of 200. Reviewers prefer the new drafts in 88 of 150 blind pairs, a modest preference. Seven days in shadow produce 412 proposals, 23 of them materially different from the current version's, and all 23 are reviewed. Fourteen days at review-required follow before the previous autonomy level is restored.
Change-type test matrix
Copy this structure into your review document and record your observed result for each row.
| Change type | Example | Required tests | Rollout path | Rollback trigger |
|---|---|---|---|---|
| Model version | Provider model upgrade | Replay rule checks; blind review sample | Shadow, review-required, previous level | Any rule-check failure |
| Instructions or prompt | New tone guidance | Replay rule checks; blind review sample | Shadow, then review-required | Brand rejections above the pre-change level |
| New tool | Ability to schedule sends | Permission-matrix entry first; replay with tool calls | Review-required only | Any unapproved tool call |
| Permission level | Draft-only to scheduling | Blast-radius and reversibility review; incident history | Review-required for a fixed period | Any send outside the approved segment |
| Data source | New product-usage table | Freshness, identity join rate, consent fields present | Shadow | Stale data or missing consent fields |
| Guardrail threshold | Frequency cap from 3 to 4 a week | Collision test on replay; complaint trend | Staged by segment | Complaints or unsubscribes above baseline |
| Brand or product facts | New pricing page | Fact check against the source; replay claims | Shadow | Any outdated claim in output |
| Unplanned provider update | Model behind an alias changes | Detected by version logging; rerun the replay set | Treat as a new change: shadow first | Output drift on replay |
A failure worth checking
The silent upgrade. The team calls its model provider through an alias that points to the latest version. The provider updates the model behind it, nothing changes in the team's own settings, and change management never triggers. Push titles grow longer and are cut off on lock screens, and a claim the old version avoided starts appearing in drafts. Verification: log the resolved model version on every agent action, alert when it changes, and rerun the replay set against the new version as if the team had made the change itself.
Common questions
How big should the replay set be?
Big enough to cover every rule you enforce and every edge case behind a past incident; coverage matters more than count. For the blind review, report the counts, and treat a small preference as a small preference.
Do we need all this for a one-line prompt edit?
The path can be short: replay rule checks and a day in shadow. It shouldn't be skipped, because one-line edits are the changes most likely to bypass review.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.