All articlesProduct

Why we built 'Why AI Said That' before adding more features

Elena Marsh10 min read
Productwhy?

The first time a customer disputed a refund the Agent had issued, the person reviewing it had no way to see why the Agent thought the refund was valid — just the outcome. That review took forty minutes of digging through logs to reconstruct, and it's the reason we built "Why AI Said That" before shipping a single new channel integration.

The story is worth telling in full, because the order we built things in — traceability before expansion — is a decision we'd make again, and one we think most teams get backwards.

Refund approvedsource: refund-policy.md · 96% confidence

The dispute that started it

The dispute itself was ordinary: a customer claimed the refund they'd received was for the wrong amount, and the person reviewing the case had nothing to go on except the final outcome recorded in the system. They had to manually reconstruct, from raw application logs never meant to be read by a human trying to answer a specific question quickly, what the Agent had actually checked before issuing that refund.

Forty minutes is a long time to spend confirming a decision that, it turned out, had been correct all along — the Agent had applied the right policy to the right order. The cost wasn't that the decision was wrong. The cost was that nobody could tell it was right without an investigation disproportionate to the size of the original question.

What the trace actually shows

"Why AI Said That" attaches the full reasoning trace to every automated reply and action: which Knowledge Base sources it read, which ones it discarded and why, and the exact confidence score at the moment it decided to answer instead of asking a human. None of that is generated after the fact for the sake of the feature — it's a record of the actual decision process, captured as it happens, which is the only way the trace can be trusted as a genuine explanation rather than a plausible-sounding reconstruction.

The "discarded and why" part turned out to matter as much as the sources actually used. A reviewer questioning a decision often wants to know not just what the Agent relied on, but whether it considered an obvious alternative and ruled it out for a specific reason — and a trace that only shows the winning source, without the runner-up it passed over, answers half the question a skeptical reviewer actually has.

Why we built it before new channels

We built this before shipping any new channel integrations, because every feature after it — Shopify, WhatsApp, Discord — would have multiplied the number of places a team couldn't explain what happened without it. Adding traceability to one channel and retrofitting it across three more later is a meaningfully harder engineering problem than building it once, early, as a foundation every new channel inherits automatically.

It's also the feature that convinced skeptical teams to turn on write actions at all. Nobody hands an Agent the ability to issue refunds until they can see exactly why it would issue one, and several of our largest deployments today started as read-only, trace-visible pilots specifically because the team wanted to build trust in the explanation before trusting the action.

That pattern — trust the explanation first, then trust the action — showed up consistently enough that we now recommend it as the default rollout path rather than treating it as a cautious exception. Teams that skip straight to write actions without first watching the trace on read-only answers tend to have a rougher first month, simply because they're learning to trust the system and learning what "wrong" looks like for the first time, simultaneously.

Common questions

Does showing the full reasoning trace ever confuse a non-technical reviewer more than it helps? We designed the default view to summarize in plain language first — what source, what confidence, what decision — with the full technical trace available underneath for anyone who wants it, rather than defaulting to the raw version for everyone.

How long is the trace retained, and does that create its own storage burden? Retention follows the same policy as conversation data generally, which we've written about separately — full detail for a defined window, a redacted summary retained longer for audit purposes without holding onto anything unnecessarily.

What the trace looks like on a real disputed case

A more recent dispute, well after "Why AI Said That" existed, is a useful contrast to the original forty-minute investigation that motivated building it. A customer disputed a partial refund, claiming they were owed the full amount instead. The reviewer opened the trace and had a complete answer within about ninety seconds.

The trace showed exactly what the Agent had checked: the order record, confirming the purchase; the specific policy clause it had matched against, which covered partial refunds for a particular category of delayed-but-eventually-delivered shipments; and the confidence score at the moment it decided the partial-refund policy, rather than the full-refund policy, was the correct match. It also showed the full-refund policy clause as a source the Agent had considered and specifically discarded, along with a brief note on why — the order didn't meet the threshold for that specific clause, which required a shipment delay beyond a certain number of days that this particular order didn't reach.

That last detail — seeing the alternative the Agent considered and ruled out, not just the one it acted on — was what let the reviewer answer the customer definitively and quickly, rather than needing to independently re-derive whether the full-refund policy might have applied. The trace had already done that comparison, visibly, as part of the original decision.

The reviewer's response to the customer, informed directly by the trace, explained specifically why the partial-refund policy applied rather than the full one, citing the actual shipment delay figure from the order record. The customer accepted the explanation without further escalation — not because they'd necessarily wanted to hear it, but because the explanation was specific and grounded in their own actual order details rather than a generic policy citation.

Compare that ninety-second resolution to the original forty-minute investigation that started this whole effort, and the difference isn't really about speed for its own sake — it's about what the reviewer was able to tell the customer. A reviewer who has to guess at the Agent's reasoning after the fact can only report the outcome and ask the customer to trust that it was correct. A reviewer with the actual trace can explain the specific reasoning, which is a meaningfully different, more convincing thing to hand back to a skeptical customer.

We've since started sampling disputed-case resolutions specifically to measure this: how often does having the trace let a reviewer give a specific, cited explanation rather than a general reassurance? That rate has stayed consistently high since the feature shipped, and it's become one of the metrics we point to internally when deciding whether a new capability needs its own dedicated traceability before it's allowed to touch anything customer-facing.

Common questions

Does building full traceability slow down the Agent's actual response time to the customer, since it has to log all of this in real time? The logging happens asynchronously alongside the response generation, not as a blocking step before a reply is sent, so it doesn't add meaningful latency to what the customer experiences — the trace is available for review essentially immediately after the response, without the customer having waited any longer for it.

How detailed does a trace need to be before it's actually useful, versus just noise a reviewer has to wade through? We deliberately layer this — a plain-language summary by default, with the full technical detail available underneath for anyone who wants to go deeper — specifically so a reviewer isn't forced to wade through low-level detail for a straightforward case, while still having it available for a genuinely complex dispute.

Does every automated reply get a trace, even simple ones like answering a basic FAQ question? Yes, without exception — the discipline only works if it's universal, since a policy of "trace the complex ones, skip the simple ones" requires someone to correctly judge in advance which ones will turn out to matter, and that judgment is exactly the kind of thing that's easy to get wrong before a dispute happens.

How do you prevent the trace itself from becoming another place sensitive customer information ends up stored insecurely? The same redaction-at-ingestion discipline we've described in more depth elsewhere in this series applies directly to trace data — a trace records that a check happened and its outcome, not the raw sensitive value the check was performed against.

Has "Why AI Said That" ever revealed a systemic problem, beyond individual disputed cases, the way the Content Gap Analyzer does for documentation gaps? Yes — reviewing traces in aggregate, rather than one dispute at a time, is exactly how we've caught patterns like the outdated-policy issue described elsewhere in this series, where the same specific stale source kept getting selected across multiple, seemingly unrelated conversations before anyone had connected the individual instances into a pattern.

Does building this slow down how fast new features ship, since every new capability now needs to be traceable too? It adds a real, deliberate constraint to how new capabilities get built — anything without a traceable decision path doesn't ship — and we consider that constraint a feature of the process, not friction to route around.

Can customers themselves see the reasoning trace, or is it internal-only? Today it's an internal tool for the team operating the Agent, though we've had more than one customer ask for a simplified, customer-facing version of it specifically for disputed cases, which is on our radar as a natural extension of the same underlying capability.

It's worth acknowledging that traceability, on its own, doesn't guarantee every decision it explains was actually correct — it guarantees that a wrong decision, if one happens, is explainable rather than mysterious. That's a meaningfully different and more modest claim than "this feature prevents mistakes," and we think being precise about that distinction matters, because overselling what traceability does risks a team trusting an Agent's decisions more than the trace itself actually warrants.

What it does reliably deliver, in every case we've reviewed, is turning the question "was this right" from an investigation into a lookup — and given how much of the cost of the original forty-minute incident was investigation time rather than anything else, that's the specific, bounded value we'd point to rather than a broader claim about correctness the feature was never designed to guarantee on its own.

We'd also flag a specific trap we watched a different team fall into after building something similar to this: treating the existence of a trace as sufficient, without also building the habit of actually reviewing traces proactively rather than only pulling them up reactively when a dispute happens. A trace nobody looks at until there's already a problem is still valuable in that reactive moment, but it misses the earlier, larger value this piece has tried to emphasize — catching a systemic issue, like a stale document repeatedly getting selected, before it generates enough individual disputes to become obviously visible on its own.

The teams getting the most value from this kind of traceability sample their own traces proactively, on a standing schedule, the same way they'd review any other quality metric — not waiting for a customer to force the question by disputing an outcome.

The team member who originally spent those forty minutes reconstructing the first disputed decision now reviews a sample of traces every week as part of a standing habit, not because anything requires it, but because she's the one who felt the cost of not having it most directly, and she's also the one who's caught the most systemic issues since, simply from having built the instinct to look.

We'd genuinely recommend any team building an Agent capable of taking real actions build something like this before adding a single additional capability, rather than treating it as a nice-to-have to circle back to once things are already running — the cost of retrofitting traceability across several already-shipped action types is real, and we paid a version of that cost ourselves on the handful of capabilities that predated this feature.