Atomiq logoAtomiq
B2B SaaS

Support agent with a real eval harness

A support team of thirty-one was absorbing a ticket volume growing faster than headcount could follow. An earlier chatbot pilot had been quietly switched off after it started confidently inventing billing policy. We rebuilt the thing around measurement first: a graded set of twelve hundred real tickets, then a retrieval-backed agent scored against that set every night.

Sector
B2B SaaS
Engagement
AI Automation
Duration
7 weeks
Team
2 engineers, 1 lead
Stack
Client AWS, existing Zendesk instance
of tickets resolved with no human touch
68%

of tickets resolved with no human touch

median first response
12 hrs → 40 min

median first response

answers escalated as incorrect, down from 9%
0.4%

answers escalated as incorrect, down from 9%

from scoping to production cutover
7 weeks

from scoping to production cutover

The problem

What was actually wrong.

Ticket volume was growing at roughly 40% year on year against a support team that could not grow at the same rate. First-response time had drifted past twelve hours and was still moving in the wrong direction.

The previous pilot had failed in the way these usually fail: it demoed well on a dozen hand-picked questions and then produced fluent, wrong answers about refunds and contract terms once it met real traffic. Nobody could say how often it was wrong, because nothing was being scored.

The knowledge that support actually relied on was scattered across a public help centre, an internal Notion space and, in practice, four long-tenured agents.

What we did

The approach, in the order it happened.

01

Build the scoring set before the system

We pulled twelve hundred resolved tickets spanning the full intent distribution, not just the common ones, and had senior agents grade the correct resolution for each. That set became the contract: nothing shipped without moving the score.

02

Retrieval over the sources that were actually authoritative

Help centre and internal documentation were indexed with provenance attached, so every answer could cite the paragraph it came from. Where the two sources disagreed — which they did, often — the conflict surfaced as a documentation defect rather than being averaged into a plausible sentence.

03

Refuse rather than guess

Billing, contractual and account-security intents were routed to a human by classification, before generation. The agent's job on those is to recognise them and hand over, and it is scored on that separately.

04

Nightly evaluation in their CI

The harness runs against the graded set every night and on every prompt or model change. A regression fails the build, which is what makes it safe for their engineers to keep changing the system now that we have gone.

Outcome

What changed. Measured against the baseline agreed at scoping, over at least a full quarter of production running.

Results

  • Sixty-eight per cent of inbound tickets now resolve without a human, measured over the first full quarter after cutover.
  • Median first response fell from twelve hours to forty minutes, with the improvement holding across the two seasonal peaks since.
  • The support team stopped growing with volume and reallocated two headcount to a customer-education function.
  • Their engineers have shipped eleven changes to the agent since handover without our involvement, each gated on the evaluation score.
The eval harness was the part we would have skipped. It is the only reason we still trust the thing six months later.
VP Customer Operations
Get started

Let’s find out what AI is actually worth to you.

A 30-minute call, no deck. Bring the workflow that frustrates you most and we will tell you whether it is worth automating.

Typically replies within one business day.