AI Customer Service: What It Does, What It Doesn’t
Field notes · Contrarian
AI chat for customer service: what it does, what it doesn’t
The vendor demo looks impressive. The pilot works. The bot handles the first tier of enquiries and your ticket volume to human agents drops by thirty percent in the first month.
Six months later, you notice something odd. The tickets that reach a human are angrier than they used to be. Not more of them – angrier. And the CSAT on human-handled tickets, which used to be one of your stronger numbers, is drifting down.
The bot is not the problem. The seam is the problem – the moment the AI customer service layer hands off to a person, and how the customer feels in that transition.
Most conversations about AI customer service focus on what the bot can handle. The more useful conversation is about what happens when it cannot. Because deflection is where the vendor makes their case, and the handover is where the customer decides how they feel about the brand.
Here is what AI genuinely moves in a support operation, what it quietly breaks, and what the seam has to do to work.
What AI genuinely moves
AI-supported customer service works, in specific, real ways.
Repeatable, information-retrieval questions. Where is my order. What is your returns policy. What is the size chart. These questions have a right answer that lives in your system somewhere, and a well-configured AI can produce that answer faster than any human. On this category of enquiry, deflection rates of sixty to eighty percent are realistic in a mature deployment, and customers are usually happy with the outcome because they got their answer quickly.
First-language handling in secondary markets. Real-time translation, scoped to your markets, lets a support operation respond in the customer’s language faster and more consistently than an unaided team could. This is genuinely useful – not because the translation is perfect, but because the response time in the customer’s language shortens dramatically.
Draft responses for human agents. The version of AI most operators underuse is the one that drafts a suggested response for a human to review and edit. This shortens agent handle time significantly without the risks of unsupervised bot output. It is less glamorous than a full-deflection bot, and it is often the better investment.
Ticket classification and routing. Consistent tagging at scale is a boring, valuable thing. AI does it well. This is where the first fifteen percent of AI customer service value in a support operation usually comes from, and it rarely makes the vendor’s front page.
None of these is trivial. All of them are real, measurable, and – when configured well – they hold up over time.
What AI quietly breaks
The failure modes are less discussed and more common than the vendor decks suggest.
Overconfident answers on the edge cases. An AI answering “where is my order” when the answer is in the system is a win. An AI answering “where is my order” when the answer requires reading a note about a customs delay, or a supplier change, or a specific customer’s history, is a risk. Confident wrong answers on the edge cases are more damaging than slow correct ones, because they close the ticket with a wrong resolution and the customer walks away misinformed.
Register mismatch on emotional tickets. A customer who is angry, upset, or grieving does not want a chirpy bot. They want a human. Many AI deployments do not detect emotional escalation reliably, and when they miss it, the customer’s first three exchanges before they reach a human are exchanges with a system that is not reading the room. Those three exchanges damage the relationship in ways the eventual human reply cannot fully recover.
The lesson here has been learned publicly. In 2024, the fintech Klarna declared it would cut headcount by nearly half and put AI in charge of 75% of customer service interactions. Eighteen months later, the company reversed course and restarted hiring customer service roles, with the CEO acknowledging publicly that customers wanted the option of a human. The MIT Sloan researchers who study this describe it in a framework called EPOCH – Empathy, Presence, Opinion/Judgment, Creativity, Hope – the five capabilities where humans are still ahead of AI. Their research is consistent: AI complements human customer service workers more than it replaces them.
The seam problem. When the AI hands off to a human, does the human get context? Does the customer have to explain again? Does the tone shift feel abrupt? Most deployments handle these badly, at least initially, and the human-handled tickets get worse because they are inheriting a bad handoff. This is usually the source of the “angrier tickets” symptom described at the top of this article.
Silent drift in FCR and CSAT. A bot deflects thirty percent of volume. Great. What happens to the thirty percent it deflected? Some are genuinely resolved. Some are the customers who gave up and did not come back. The dashboard cannot tell the two apart from deflection rate alone, and a lot of operations do not measure the difference. This is exactly the silent churn pattern at a smaller scale – customers who walked away without a complaint.
The seam problem
The seam is the moment the customer transitions from AI to human. It is the most important design decision in an AI-supported support operation, and the least discussed.
A well-designed seam does three things.
One: the human agent sees the full conversation the customer has had with the AI, and reads it before replying. Not a summary – the actual transcript. Summaries lose the tone and the specific words the customer used.
Two: the customer does not have to explain again. The first thing the human says references what the customer said to the AI, so the customer feels heard rather than restarted.
Three: the seam is triggered by the right signals. Not just “customer asked to speak to a human,” but also emotional register shifts, keyword patterns that suggest escalation, and length of conversation without resolution. A bot that only escalates on explicit request will keep customers in a conversation they wanted to leave three exchanges ago.
Getting the seam right is not a bot-configuration problem. It is an operational design problem. It requires human agents trained on how to open a conversation the AI started, and it requires the AI to be configured with escalation triggers that are more sensitive than the vendor default.
Technology supports the work. It does not run it.
What “AI-supported” looks like when it works
The working shape, in most SME support operations we have seen, is this.
The AI handles the top few intents that have clear right answers, and does so with a hand-off threshold that is deliberately conservative – it escalates before the customer has to ask.
Human agents own everything with judgement, register, or emotional weight. The AI drafts responses for these agents to speed their handle time without replacing their judgement.
Real-time translation supports the human agents across the markets served, so the operation feels first-language to the customer without the staffing overhead of native coverage in every language for every shift.
The dashboard tracks not just deflection but post-deflection outcomes – return contact rate for deflected tickets, CSAT specifically for the seam, and the tone-drift check across the handoff. Deflection without those metrics is a vanity number. This is the same discipline that runs through the metrics that matter for small teams: measure what the customer actually experienced, not what the dashboard reported.
The operation is designed so that a customer who reaches a human never has to explain what they already said to the bot. That single design commitment prevents most of the failure modes above.
You can see the shape of how we approach AI-supported delivery on the Chat Tool page – where the platform is described in the same terms we design it in: technology supporting the team, not replacing it.
When AI should not answer at all
There are categories of contact where an AI should not attempt a resolution. The list is short and it is worth being clear about it.
Complaints and public-review-risk tickets. These need human judgement, human tone, and a considered path to resolution. An AI attempting to handle a complaint is one screenshot away from becoming the complaint.
Cancellation and churn-risk contacts. These are commercial conversations, not information-retrieval conversations. The AI cannot make the retention decision, and a poor first response accelerates the churn.
Bereavement, illness, or personal hardship. Customers occasionally mention personal circumstances that require a human response. The AI cannot recognise these reliably, and the tone risk is severe.
Anything with legal or regulatory weight. Refunds tied to consumer rights, data access requests, disputes. These need a human on the record.
None of this means the AI customer service layer is a bad idea for the rest of the operation. It means the AI needs to know which conversations it should not be in – and hand them over cleanly, before the damage.
The uncomfortable observation
The measure of a good AI customer service implementation is not what it handles on its own. It is how the customer feels when it does not.
Vendors sell on deflection rates because deflection rates are big numbers and big numbers close deals. Deflection rates are also easy to hit at the expense of the metrics that actually matter – resolution quality on the deflected tickets, CSAT on human-handled ones, silent churn on customers who never came back.
An AI-supported operation that reads deflection rate alone will look like a success and be a slow failure. An operation that reads the seam metrics – handover CSAT, post-deflection return contact rate, tone continuity across the seam – will know quickly whether the AI customer service layer is genuinely helping or quietly costing the business its best customers.
The line we hold on this, in every deployment: technology supports the work. It does not run it.
Where to start
If you are thinking about AI customer service in your support operation, the useful sequence is not “which vendor.” It is: which intents does it make sense to route to AI, what does the seam look like when it hands off, and how will you know whether it is working – beyond deflection rate.
If a founder-led, honest conversation about that sequence is what you need – with a team that has designed and run AI-supported operations for clients around the world – that is where Support Solutions starts.
Talk to us about Support Solutions
AI-supported customer service designed the way you would design it yourself – with the seam done properly, the metrics that matter tracked from day one, and human judgement kept where it belongs. Technology supports the work. It does not run it.
Explore Support Solutions →