All articlesInsights

What we log (and what we don't) from every AI conversation

Diego Ferreira10 min read
Insights

Every automated action needs a trace — which source it used, what confidence score it hit, what API call it made. That's non-negotiable for the reasons we've written about before. But "log everything" and "respect customer privacy" pull in opposite directions the moment payment details or health information enters a conversation, and where exactly to draw that line took real, deliberate work to get right.

The tension

The tension isn't hypothetical or rare — it comes up constantly, because customer support conversations routinely surface exactly the kind of sensitive information a full, unredacted trace would otherwise capture by default: a card number mentioned while explaining a billing dispute, a date of birth offered to verify identity, sometimes health information relevant to a return or accommodation request. A logging system built purely for traceability, with no privacy discipline applied to it, would capture all of that by default, simply because it's present in the conversation.

Resolving that tension required treating privacy not as a filter applied after the fact, but as a decision made at the point of ingestion — before anything sensitive ever touches storage, encrypted or otherwise, rather than trusting a later cleanup pass to catch everything reliably.

What we redact and when

We redact before we store, not after: card numbers, SSNs, and a configurable list of patterns per customer get stripped from the trace at ingestion, so they never touch a database in the first place — not even encrypted. The distinction between "redacted at ingestion" and "redacted after storage, on some later cleanup pass" matters more than it might initially seem: anything ever written to a database, even briefly, creates a window where a backup, a replication lag, or an access-control misconfiguration could expose it, however small that window is. Redacting before the write closes that window entirely rather than narrowing it.

The configurable-pattern-list detail matters too, because sensitivity isn't identical across every customer or industry — a healthcare-adjacent deployment needs a different, broader redaction pattern list than a general e-commerce one, and building the system to allow that configuration rather than assuming one universal pattern list fits every use case was a deliberate design choice made specifically because we'd seen how much sensitive-data categories vary across our actual customer base.

What survives

What survives is the reasoning, not the raw sensitive content: "verified last 4 digits match" instead of the card number itself, "confirmed date of birth" instead of the date. Enough to audit the decision, not enough to reconstruct the customer's data. This is the specific design goal that makes the whole approach work: a reviewer investigating a disputed decision needs to know that verification happened and on what basis, but doesn't need — and shouldn't have access to — the actual sensitive value that was checked against.

That distinction between "proof that a check happened" and "the raw data the check was performed on" is the same underlying principle as the "Why AI Said That" traceability feature we've described elsewhere in this series, applied specifically to the privacy dimension rather than the general reasoning dimension — both are about making a decision auditable without requiring full, unrestricted access to everything that went into it.

Retention is short by default too — 90 days for full conversation transcripts, indefinite for the redacted trace. Teams that need longer transcript retention for compliance reasons can extend it, but it's opt-in, not the default. The default being short, rather than long, reflects a deliberate bias: most support conversations don't need to be retrievable in full detail indefinitely, and defaulting to short retention with an explicit opt-in for longer retention puts the burden of justifying extended retention on the specific compliance need, rather than defaulting to keeping everything indefinitely just in case.

Common questions

How do you handle sensitive information that doesn't match a known pattern — something novel that the redaction rules weren't built to catch? This is an ongoing area of active work rather than a fully solved problem; we treat pattern coverage as something that needs continuous review and expansion, informed by periodic audits of what actually appears in real conversations, rather than assuming an initial pattern list is complete and permanent.

Does redacting at ingestion ever interfere with the Agent's own ability to do its job — for instance, needing to reference a card number to process a refund? The redaction applies specifically to what gets logged for later human review, not to what the Agent uses in the moment to complete an action — the live transaction still has access to what it legitimately needs, through a properly scoped, audited API call, while the stored trace of that transaction reflects only that the check happened, not the raw value used.

A specific audit that tested the redaction rules

The clearest test of whether the redaction-at-ingestion approach actually works came from a deliberate internal audit, run specifically to try to find sensitive information that had slipped through into stored traces despite the rules meant to catch it.

The audit sampled a large batch of stored, redacted traces and searched specifically for patterns that looked like they might be unredacted sensitive data — sequences of digits in the length range of a card number or a national ID, common formats for dates of birth, and a few other heuristics — flagging anything that matched for manual review rather than assuming the redaction rules had caught everything by design.

The vast majority of flags turned out to be false positives — order numbers, tracking numbers, and other legitimate numeric identifiers that happened to be the right length to trigger the heuristic but weren't actually sensitive. That's an expected and acceptable cost of running the audit this way; a heuristic loose enough to reliably catch real misses will also flag plenty of harmless matches, and reviewing those false positives is the price of the confidence the audit provides.

The audit did find a small number of genuine misses, and every one of them traced back to the same underlying cause: a customer volunteering sensitive information in a format the existing redaction patterns hadn't been built to recognize — a card number written with unusual spacing, for instance, that didn't match the specific digit-grouping pattern the redaction rule was looking for.

Each of those genuine misses got two fixes: an immediate correction to the specific stored trace, removing the sensitive content after the fact for that instance, and a broadened redaction pattern going forward so the same formatting variation wouldn't slip through again. Neither fix alone would have been sufficient — fixing the one instance without broadening the pattern leaves the same gap open for the next customer who happens to format things the same unusual way; broadening the pattern without fixing the existing instance leaves a known problem sitting in storage.

We now run this same audit on a standing quarterly cadence rather than treating the original pass as a one-time validation, specifically because customer-provided formatting is genuinely varied and new variations show up over time as new customers, in new contexts, volunteer sensitive information in ways the existing patterns weren't built to anticipate.

This is the same underlying discipline we've described elsewhere in this series applied to a different problem: a system that redacts sensitive information by design is not the same claim as a system that's been checked, repeatedly and skeptically, to confirm the redaction is actually working as intended — and we treat the second claim as the one actually worth making, which requires the ongoing audit, not just the original design.

Common questions

How do you handle a jurisdiction with data-retention or right-to-erasure requirements that conflict with the default 90-day full-transcript window? The default window is a starting point, not a fixed constraint — retention policy is configurable per deployment specifically to accommodate jurisdiction-specific requirements, including shorter mandatory deletion windows or explicit erasure-request handling where required.

Does the redaction process ever need to be different for voice conversations versus text, given that sensitive information might be spoken rather than typed? Yes — voice introduces its own detection challenge, since redaction has to work against a transcribed version of speech rather than typed text directly, and transcription itself can introduce errors that affect how reliably a redaction pattern matches, which is an area we treat as needing its own dedicated testing separate from the text-based patterns described in this piece.

How do you audit whether the redaction rules are being applied consistently across every integration and channel, rather than just the ones tested most thoroughly? The quarterly audit described in this piece runs across all channels and integrations by design, specifically to avoid a gap where redaction is well-tested in one channel but was never verified in a newer, less-scrutinized one.

What's the actual process if a customer explicitly asks what specific data is stored about their own conversations? We support data-access requests as a standard operational capability, returning what's actually stored — which, by design, is the redacted trace and any transcript still within the retention window, not the sensitive raw values that were never stored in the first place.

Does having strict redaction-at-ingestion rules ever prevent the Agent from doing its job correctly, if it needs some piece of sensitive information to complete a legitimate action? The redaction applies to what gets logged for later review, not to what the Agent uses live to complete an action — a legitimate action that needs a sensitive value uses it through a properly scoped, audited call in the moment, while the stored trace of that action records only that the check happened, which is the same distinction described earlier in this piece between proof-of-check and the raw checked value.

How do you balance this privacy-by-default approach against a customer or regulator sometimes needing the full original transcript for a specific investigation? The 90-day full-transcript retention window is specifically sized to cover the vast majority of legitimate investigation needs, which tend to surface relatively quickly after an incident; for the rarer cases needing longer retention, that's exactly what the opt-in extended retention exists for, applied deliberately rather than as an ambient default.

Is this redaction-at-ingestion approach something customers can audit or verify themselves, rather than just trusting that it's implemented correctly? We treat this as a fair and important question, and our answer is that the redaction logic itself, and what it does and doesn't capture, is documented and available for review as part of a customer's own security and compliance evaluation — we don't consider "trust us" sufficient for a claim this consequential.

If we had to state the single principle underneath every specific rule described in this piece, it would be this: log enough to prove a decision was reasoned and correct, and no more than that. Everything else — the specific redaction patterns, the retention windows, the quarterly audits — is an implementation detail in service of that one principle, and it's the principle, not any specific rule, that should guide a team through the cases this piece hasn't explicitly covered.

It's worth adding a note on how this interacts with the trust-metric work described elsewhere in this series: the tone-based assessment described there runs against the same redacted trace this piece describes, not against raw, unredacted transcripts, which means the trust metric itself was designed from the start to work within the same privacy constraints rather than requiring an exception to them. Building both systems with the same underlying privacy discipline in mind, rather than retrofitting one to respect constraints the other ignored, avoided a class of awkward compatibility problems we've seen other teams run into when privacy and measurement get built as separate, uncoordinated efforts.