The checklist that took us from pilot to full rollout
Going from one channel and five agents watching closely to company-wide coverage doesn't happen by flipping a switch. Here's the checklist we actually followed, and the specific gating logic behind it that we think generalizes well beyond the particular team it was built for.
Pilot mode
Pilot mode looked like this: one channel, five agents reviewing every single automated reply before it sent, for three weeks. That review step wasn't overhead — it was how the team built enough trust to eventually turn it off. Reviewing every single reply, rather than a sample, was a deliberate choice specifically for the pilot phase, because the goal wasn't statistical confidence about an average quality level, it was building the specific, concrete trust that comes from a human actually seeing what the Agent would have said, repeatedly, before deciding to let any of it go out unreviewed.
Three weeks turned out to be roughly the right amount of time for patterns to emerge clearly — long enough to see the Agent handle a reasonably representative range of real questions, short enough that the pilot didn't drag on past the point where the reviewing agents had already learned what they needed to learn from it.
The four gates
The checklist that got us to full rollout had four gates, each requiring a week of clean data before moving on: resolution rate above 90%, zero escalations for the same unresolved reason twice, sub-5% override rate from reviewers, and a Content Gap Analyzer report with nothing older than three days. Each gate targets a different kind of readiness, deliberately: resolution rate measures raw capability, the repeated-escalation-reason gate measures whether known problems are actually getting fixed rather than recurring, override rate measures whether human reviewers still feel the need to intervene, and the gap-report gate measures whether the documentation itself is being kept current rather than allowed to drift.
The repeated-escalation-reason gate is worth calling out specifically, because it's the one most teams don't think to include on their own first attempt at a rollout checklist. A team can hit acceptable resolution and override rates while still having the same specific gap causing repeated escalations week after week — that gate exists specifically to catch that pattern, forcing a fix before allowing expansion rather than letting a known, recurring problem simply get diluted by an otherwise healthy average.
Each gate unlocked one more channel and removed one more layer of review, in that order — never both at once. Teams that tried to expand channels and reduce oversight in the same week were the ones that had to roll back. That specific sequencing rule — one change at a time, never two simultaneously — exists because it's the only way to know, when something does go wrong after an expansion, which specific change actually caused it. A team that expands to a new channel and reduces review depth in the same week, and then sees a quality dip, has no clean way to know which of those two changes was responsible.
What eight weeks looked like
Eight weeks after the pilot started, the same five agents were covering four channels with spot-checks instead of full review, and the team backfilled the freed-up time into the two channels that still needed a human touch. That backfilling detail matters as much as the headline efficiency gain — the point of the rollout wasn't to reduce headcount, it was to redirect the same five people's attention toward the parts of the job that genuinely still needed a human, rather than spreading them thin across everything equally regardless of how much automation had actually reduced the need for oversight in a given area.
Looking back at the eight-week arc as a whole, the team's own retrospective take was that the checklist's real value wasn't any single gate — it was having gates at all, defined in advance, rather than making an ad hoc judgment call about readiness at each step. An ad hoc judgment call is vulnerable to optimism bias exactly when a team is excited about early results; a predefined, objective gate isn't.
Common questions
Are these specific gate thresholds — 90% resolution, sub-5% override — universal, or should they be tuned per team? We'd treat the specific numbers as a reasonable starting point rather than a universal standard — what matters more than the exact threshold is having the same four categories of readiness represented by some gate, tuned to what your own team considers an acceptable bar in each category.
What happens if a team passes three gates but stalls on the fourth for an extended period? That's useful information rather than a failure — a persistent stall on one specific gate, in our experience, almost always points to a specific, identifiable underlying issue (usually a documentation gap or a particular ticket type the Agent handles poorly) worth investigating directly rather than just waiting longer for the number to improve on its own.
What almost went wrong at gate three
The rollout described above reads cleanly in retrospect, but gate three — sub-5% override rate from reviewers — came closer to failing than any of the others, and the specific reason why is worth describing in more detail than a clean success story usually includes.
Override rate had been trending down steadily through the first few weeks, comfortably on track to clear the 5% threshold, until it spiked unexpectedly in week four, driven almost entirely by one specific reviewer who was overriding roughly three times as often as the rest of the team on the exact same category of ticket.
The initial assumption, reasonably enough, was that this reviewer had simply set a stricter personal bar than everyone else — a normal, expected kind of variance across individual reviewers rather than a signal about the Agent itself. Digging into the specific overrides before accepting that explanation, though, revealed something different: this particular reviewer happened to be the one most familiar with a recent, narrow policy exception that hadn't yet made it into the written documentation, and was correctly overriding the Agent on cases where that undocumented exception applied — cases the Agent had no way to know about, because nobody had told the Knowledge Base yet.
Once identified, this took the same fix we've described elsewhere in this series for a documentation gap: write the exception down, in this case a specific paragraph about a temporary policy adjustment tied to a supply issue affecting one product line. Once added, override rate on that specific ticket category dropped immediately, and the team's aggregate override rate cleared the gate the following week without further intervention.
The lesson we took from almost failing this gate wasn't about the specific exception — it was about not accepting the first plausible explanation for a metric moving in the wrong direction without checking it against the actual underlying cases. "This reviewer is just stricter" was a reasonable-sounding story that happened to be wrong, and treating it as settled without digging further would have either delayed the rollout unnecessarily while waiting for override rate to "naturally" improve, or worse, led someone to coach a reviewer to be less strict on cases where being strict was actually correct.
We've since added a specific step to how we investigate any gate that's trending the wrong direction: before accepting a plausible surface-level explanation, pull the specific cases behind the number and check whether they share an underlying cause that's fixable, the way this one turned out to be, rather than assuming the metric itself is telling the whole story on its own.
Common questions
How rigid should the four-gate order be — is it ever appropriate to work on two gates simultaneously if they seem unrelated? We generally recommend strict sequencing even between seemingly unrelated gates, specifically because of the diagnostic clarity benefit described in this piece — when something goes wrong after a change, being able to point to exactly one thing that changed is valuable enough to be worth the extra calendar time sequential gating costs.
What happens if a team passes all four gates in the pilot channel but then struggles significantly when expanding to a second channel with different characteristics? This is expected and normal enough that we recommend treating each new channel expansion as its own miniature version of the same gated process, rather than assuming success in one channel guarantees success in a structurally different one.
How do you handle a gate that's genuinely ambiguous to measure, like "zero escalations for the same unresolved reason twice," when different reviewers might categorize an escalation's underlying reason differently? We standardize this with the same tagging taxonomy described elsewhere in this series for failure categories, specifically so "same underlying reason" is judged against a consistent, shared definition rather than each reviewer's individual interpretation.
Is there a risk that reducing review depth too aggressively after clearing the override-rate gate creates a false sense of security if reviewer attention was itself part of what kept quality high? This is a genuine risk worth watching for, which is why override rate continues to be monitored even after review depth is reduced, rather than treating the gate as a one-time checkpoint that, once passed, no longer needs attention.
How do you know if your own team's version of this checklist has the right gates for your specific situation, versus needing different or additional ones? Start with these four as a baseline, since they cover the four distinct readiness dimensions described in this piece, and add a team-specific gate only if there's a concrete, identified risk the existing four don't capture — resist adding gates speculatively, since more gates also means a slower rollout.
Does the review depth ever need to increase again after being reduced, or is it strictly one-directional? It's not strictly one-directional — a team that notices quality dipping after reducing review depth should be willing to temporarily increase it again while investigating, and treating that as a normal, expected part of the process rather than a failure of the original rollout plan.
How do you decide the order in which channels get added during expansion? We recommend adding channels roughly in order of how similar their ticket patterns are to the channel the pilot already succeeded on — a channel with very different, unfamiliar ticket patterns is a bigger unknown to add next than one that closely resembles what's already been validated.
The near-miss at gate three is, in retrospect, the most instructive part of this whole rollout, more than the parts that went smoothly — a metric moving in an unexpected direction is genuinely useful information, but only if someone digs into the specific cases behind it rather than accepting the first plausible-sounding explanation. We'd rather a team's rollout hit one real snag that gets properly diagnosed than sail through cleanly on a set of gates nobody had to actually investigate — the diagnosis itself is where a team learns something durable about their own documentation and their own Agent's behavior.
One thing we'd emphasize for any team adapting this checklist to their own rollout: the specific four gates described here reflect what mattered for this particular team's risk profile and ticket mix, and we'd expect a team in a different industry, with a different risk tolerance, to reasonably weight the four dimensions differently, or even define a fifth gate specific to a risk this team simply didn't have. The value of the framework is in having explicit, predefined gates measuring genuinely distinct readiness dimensions, not in these four specific thresholds being universally correct for every team that adopts this approach.