All articlesEngineering

Designing a fallback plan for when your AI Agent doesn't know

Marcus Ibe10 min read
Engineering?

Confidence scores tell you how sure a model is — not whether it's right. We learned that the hard way after an Agent confidently answered a billing question with a policy from two versions ago, and the fallback design that came out of that incident is one of the more load-bearing pieces of how we build agents today.

Why confidence scores aren't enough

68% confidence → ask

The uncomfortable part of that incident wasn't that the Agent got something wrong — every system gets things wrong sometimes. It's that the Agent's own confidence score for that specific answer was high, because the outdated policy document it retrieved was clearly and unambiguously worded, just no longer accurate. High confidence measures how clearly a source supports an answer, not whether that source is actually the current, correct one — those are genuinely different questions, and conflating them is exactly the mistake we'd made.

That distinction is what led us away from treating confidence score as a single binary gate — answer if above threshold, escalate if below — and toward a graduated response that treats different confidence bands differently, rather than assuming a single cutoff can capture the full range of appropriate caution.

The three tiers

The fix wasn't a smarter model. It was a three-tier fallback: below 90% confidence the Agent still answers but flags the reply for review; below 70% it asks a clarifying question instead of guessing; below 50% it hands off to a human immediately with the full conversation attached. Each tier represents a different kind of appropriate response to a different kind of uncertainty, rather than treating all uncertainty below a single line the same way.

The middle tier — asking a clarifying question rather than either answering outright or escalating outright — turned out to be the one doing the most unappreciated work. A meaningful share of moderate-confidence situations resolve cleanly once the Agent asks one specific question, because the underlying uncertainty was often about which of two similar policies applied, a question a human customer can usually answer immediately and definitively.

What each tier logs

Each tier logs why it fired — which source it almost used, what confidence score triggered the handoff — so the weekly review isn't guesswork either. The "which source it almost used" detail matters specifically because it's what let us catch the outdated-policy problem systematically rather than one incident at a time: reviewing a batch of flagged-for-review answers from the top tier surfaced a pattern of the same outdated document being retrieved repeatedly, which pointed directly at a documentation problem rather than a model problem.

That logging discipline is the same underlying practice we've described elsewhere in this series as "Why AI Said That" — the fallback tiers aren't just a response mechanism, they're also a structured data source for finding and fixing the underlying causes of uncertainty, rather than just managing the symptom conversation by conversation.

Since shipping this, wrong-but-confident answers dropped to nearly zero, and the handoff rate settled around 4%, which turned out to be the number our support lead actually wanted to see — low enough that the Agent is doing real work, high enough that genuinely uncertain cases are reliably reaching a human rather than being confidently guessed at.

Common questions

How did you pick the specific thresholds — 90%, 70%, 50% — rather than some other set of numbers? We started with rough estimates and tuned them against actual outcomes, watching how often a human reviewer agreed with an answer at each confidence band and adjusting the boundaries until the bands actually corresponded to meaningfully different real-world accuracy rates, rather than being arbitrary round numbers.

Does the middle tier's clarifying question ever annoy customers who feel like the Agent should just know the answer? Occasionally, but less than we expected — most customers respond well to a specific, well-targeted clarifying question, and the alternative, a confidently wrong answer, produces far worse customer sentiment on the occasions it happens, based on how we've compared outcomes across both.

Watching the middle tier work on a real batch of tickets

The middle tier — asking a clarifying question rather than answering outright or escalating outright — is the hardest of the three to design well, because a badly-worded clarifying question can feel more frustrating to a customer than either a direct answer or a direct handoff to a human would have been. Getting the wording right took real iteration, and reviewing a specific batch of real tickets from that iteration is more informative than describing the tier in the abstract.

An early version of the middle-tier clarifying question was generic: "Can you provide more detail about your issue?" — technically appropriate for a moderate-confidence situation, but unhelpful in practice, because it put the burden on the customer to guess what specific detail was actually missing, rather than the Agent using its own uncertainty to ask something targeted.

The revised version asks a specific question derived from exactly what made the Agent uncertain in the first place: if the ambiguity was between two similar policies, the clarifying question names both candidates and asks which one applies — "just to confirm, are you asking about the standard 30-day return window, or the extended holiday return window that applies to orders placed after a certain date?" — rather than asking the customer to volunteer information the Agent could have surfaced itself.

That specific change — from a generic request for more detail to a targeted question naming the actual candidates the Agent was uncertain between — measurably improved how often the middle tier resolved cleanly on the first follow-up message, rather than requiring a second or third round of back-and-forth. Customers, it turns out, answer a specific either-or question much faster and more accurately than an open-ended request to elaborate.

We now treat "does the clarifying question name the actual candidates causing the ambiguity, or does it just ask generically for more information" as a specific quality bar for the middle tier, checked during any review of how it's performing on a given topic. A generic clarifying question, even one grammatically fine and polite, is treated as a design defect worth fixing, not an acceptable fallback.

The broader principle this batch of tickets taught us, beyond the specific fix, is that a fallback tier isn't really "safe by default" just because it avoids a wrong answer — a badly designed fallback can still produce a poor customer experience even while technically doing its job of avoiding the worse outcome of a confident mistake. Each tier needs its own design quality bar, not just a correct triggering threshold.

Common questions

How do you decide whether a specific customer complaint about the middle tier's clarifying question reflects a genuine design flaw versus an unusually impatient customer? We look for the pattern across multiple customers on the same specific question wording, not a single complaint — an isolated complaint is more likely individual impatience, while the same complaint pattern recurring across different customers on the same clarifying question points to an actual wording problem worth fixing.

Does the three-tier system ever need a fourth tier for an especially high-risk category, rather than the same three bands applying everywhere? For our highest-risk action types, like anything touching account security, we do push the effective thresholds more conservative rather than adding a literal fourth tier — more cases route to human review at confidence levels that would clear the general threshold comfortably, which achieves a similar effect to a stricter tier without adding structural complexity.

How often do you find that a case escalated at the lowest confidence tier turns out to have actually been something the Agent could have handled correctly? Occasionally, and we track this specifically as a signal that the bottom threshold might be set slightly too conservative for that particular topic — a consistently over-cautious threshold on a well-documented topic is itself a tuning opportunity, distinct from the risk of an under-cautious one.

Does adding this fallback system reduce the number of write actions the Agent takes overall, compared to not having it? It shifts the mix rather than reducing overall volume meaningfully — cases that would have been confidently (and sometimes wrongly) resolved before now get appropriately routed to a clarifying question or a human, while cases that were already being handled correctly continue to be handled the same way.

How do you communicate to customers, in the middle tier specifically, that the Agent is asking a clarifying question because of genuine uncertainty rather than just following a script? We don't explicitly tell customers "I'm uncertain" in those words, since that framing tends to undermine trust unnecessarily — instead the clarifying question itself, framed around the specific candidates causing the ambiguity, communicates genuine engagement with the actual question rather than reading as generic script-following.

Do these thresholds need to be different for different ticket types, or is one set of tiers enough across the board? We've found ticket-type-specific tuning helps at the margins, particularly for high-stakes categories like account security, where we deliberately push more cases into the top human-review tier even at confidence levels that would clear the general threshold comfortably.

How often should a team revisit whether their fallback tiers are still calibrated correctly? We recommend the same weekly-review cadence as the Content Gap Analyzer, since the two tools tend to surface related problems — a rising rate of top-tier flags on a specific topic is often the same underlying documentation gap the Content Gap Analyzer would also catch, just approached from a different angle.

Stepping back from the specific tiers and thresholds, the underlying principle we'd want any team to take from this piece is that uncertainty deserves a graduated response, not a binary one. A system that only knows how to either answer confidently or hand off entirely is missing the entire middle range of situations where the right move is neither — where a single well-targeted question, asked by the system itself rather than punted to a human, resolves the uncertainty faster and more pleasantly for the customer than either extreme would have.

We'd also add a note on how this system interacts with the traceability work described elsewhere in this series: every fallback-tier decision gets the same full trace treatment as any other action, meaning a reviewer looking back at why a specific case landed in the middle tier rather than the top or bottom tier can see the exact confidence score and the specific competing sources that drove that placement, not just the outcome. That visibility is what let us catch and fix the generic-versus-specific clarifying-question problem described earlier — without it, we'd have known the middle tier was firing, but not why its wording wasn't working as well as it should have.

If we had to summarize the actual shift in thinking this system represents, it's moving away from asking "was the model confident enough" as a single yes-or-no gate, and toward asking "what does this specific level of confidence actually call for" as a graduated design question — a small reframing that, once made explicit, changes how every subsequent confidence-related decision in the system gets designed, not just the three thresholds described in this piece.

We'd also note that none of the three tiers described here are meant to be static forever — as a deployment's documentation improves and its Knowledge Base matures, the thresholds themselves are worth periodically revisiting, since a threshold set conservatively during an early, thinly-documented phase may reasonably be relaxed once the underlying source material has genuinely caught up, the same way we've described documentation maturity evolving elsewhere in this series.