Behind the scenes: how we instrument our AI features

Behind the scenes: how we instrument our AI features

Tracing every span

We trace every LLM call with OpenTelemetry, including prompt content, model id, and downstream tool calls. One assistant turn is one trace; every tool invocation is a child span, and a retry is a sibling, not a replacement.

That shape matters more than it sounds. When someone reports "the assistant did something strange", the question is almost always which of four tool calls returned the surprising thing — and a flat log cannot answer it.

The data is privacy-scrubbed before it lands in the warehouse, but full traces stay available for support sessions, behind an access grant that expires.

What we put on a span

The attribute set is deliberately small. Every attribute is a thing someone has to maintain, and a thing that might leak.

  • Model id and provider, plus the resolved model when a gateway rewrote it.

  • Token counts in and out, and the cached-prefix count separately.

  • Latency to first token and to completion — they fail differently and averaging them hides both.

  • Tool name, argument shape (keys only, never values), and whether the call was auto-approved or held for a human.

  • The eval verdict, attached after the fact by the nightly job.

Privacy comes before the warehouse

Scrubbing runs at the collector, not at query time. The rule is that a raw prompt never crosses the boundary into analytics storage, because "we will filter it on read" is a promise that survives exactly one incident.

Tool arguments are scrubbed too — that took a second pass to get right, since half the interesting arguments are record ids and the other half are free text a user typed.

Workspaces can turn the whole layer off. A handful of our regulated customers do, and they lose the support-session replay along with it. That is the honest trade and we say so in the settings copy.

Evals are the other half

Traces tell you what happened. They do not tell you whether it was any good. For that we run a nightly eval suite — about 180 cases, half of them regressions harvested from real support tickets.

Each case is a fixed input, a tool environment, and an assertion that is usually weaker than you would like: not "the answer equals this string" but "the answer cites the right record and does not invent a second one".

A model upgrade that improves aggregate latency and loses two eval cases does not ship. That rule has held twice and cost us a week both times.

Observability tells you the system ran. Evals tell you it was right. Teams that buy the first and skip the second end up with beautiful dashboards of a product nobody trusts.

What a support session looks like

A customer reports that the assistant assigned a task to the wrong person. Support requests trace access for that conversation; the grant lasts four hours and is written to the audit log with the ticket number.

The trace shows four spans: a search that returned six people, a disambiguation the model skipped, the write, and the approval card the user clicked through in two seconds.

The fix, in that case, was not in the model. The search tool was returning people without saying which were inactive, and the approval card showed a name without a role. Both were our bugs, and neither would have been findable from a log line saying the task was created.

What the instrumentation costs

Traces are about 4KB each after scrubbing, and we keep every one for seven days before sampling down to 5%. At current volume that is a few hundred euro a month of storage and roughly the same again in collector compute.

The eval suite costs more, and less predictably: a full run is about 180 model calls, nightly, plus every re-run triggered by a prompt change. We cap it per day and skip the run when nothing in the prompt or tool layer changed.

Both numbers are small against the support time they save. We track them anyway, because the version of this that gets switched off in a cost review is the version nobody put a number on.

What we got wrong first

We sampled at 10% for the first two months, to control cost. Then we spent a fortnight chasing a bug that was, of course, in the unsampled 90%. We now keep every trace for seven days and sample only what survives past that.

We also logged prompt content before the scrubber existed, on the assumption we would add it "next sprint". Deleting those rows took longer than writing the scrubber would have.