
Customer support
Reference buildFirst response in under a minute, without a confidently wrong answer
A support agent that resolves routine tickets end to end through narrowly scoped tools, states its confidence before acting, and escalates rather than guessing.
At a glance
- Duration
- 17 weeks
- Team size
- 1 person
- Engagement
- New build
- Project type
- AI & automation
- Industry
- B2B SaaS
The situation
The challenge
Routine tickets — order status, refunds, password resets — consume the same triage attention as complex ones. But an agent that is occasionally confidently wrong damages customer trust more than a slow human queue does, so autonomy had to be earned rather than assumed.
What we did
Narrow the tools, enforce the limits in code, and make escalation the default whenever confidence is low.
The calls that mattered
Tool-scoped access, not system access
Order lookup, account lookup, refund issuance with a hard cap, ticket updates, escalation — each validated in code independent of the model's reasoning. The limits are not instructions the model can talk itself out of.
Confidence stated before consequences
The agent declares confidence before any consequential action, and anything below a calibrated threshold escalates automatically instead of proceeding.
What changed
- of tickets auto-resolved
- 68%of tickets auto-resolved
- first-response time
- <60sfirst-response time
- escalation accuracy
- 96%escalation accuracy
- customer satisfaction
- +22%customer satisfaction
- 68% — no human involvement
- <60s — down from over four hours
- +22% — post-rollout
68% of tickets resolved without a human, first response under a minute instead of four hours, 96% escalation accuracy and a 22-point satisfaction improvement.
Services used
- Tool-scoped support agent with code-enforced limits
- Confidence-based escalation and human handoff
- Full reasoning and tool-call audit trail
What we would do differently
Every project has one of these. Publishing it is the point — a case study with no regrets in it is marketing, not evidence.
The valuable engineering was not the reasoning quality, it was the tool design and the hard limits enforced outside the model. Starting with conservative thresholds and relaxing them against real evaluation data is what made it deployable.
More work
- Internet Money · Crypto financial services$10M+monthly transaction volume
One wallet interface across five blockchains
A multi-chain crypto platform where every chain hides behind one interface, with a tamper-evident audit trail underneath it. $10M+ moves through it monthly.
Read case study - MagicTask · Project management SaaS100K+concurrent users
From latency spikes at 5,000 users to 100,000 concurrent
A gamified project management platform re-architected from a single Node server into an event-driven system that holds 100,000+ concurrent connections without the reward engine touching the critical path.
Read case study