If you are considering AI in customer service, the choice is rarely between AI and no AI, but between letting the AI answer the customer directly and letting it write drafts the team reviews. Vendors sell the first. Most of those who have succeeded started with the second. This article explains why, when full automation is still right, and how you can tell a flow is ready.
Why do reviewed AI drafts beat full automation at the start?
Because at the start you do not know what the AI does not know, and the only way to find out without customers finding out for you is to have a person read the answers first. The approach is called human-in-the-loop: the AI proposes, a person approves, edits or rewrites, and only then does the reply go to the customer.
The best-documented study of this setup is Generative AI at Work by Brynjolfsson, Li and Raymond, which followed 5,179 customer service agents as an AI assistant suggesting replies was rolled out in stages. Agents could use, edit or ignore the suggestions. Productivity, measured as issues resolved per hour, rose by 14 percent on average and by 34 percent for new and less experienced agents, while the customers' tone in the conversations became more positive and staff turnover fell.
What the study shows is that the gain does not require removing the person. It comes from the answer to the common questions already being written when the agent opens the ticket. What remains is to read, adjust and send, and that takes seconds when the draft is right.
How do reviewed drafts build trust and train the knowledge base?
Every review is a measurement point. When the agent sends the draft unchanged, you know the knowledge behind it is right. When she edits it, you know the article needs correcting or completing. When she rewrites from scratch, you have found a knowledge gap, a question that recurs without an approved answer. None of those signals come from a bot that has already answered the customer.
That is how the knowledge base gets trained: not by the model learning anything, but by the content it answers from getting better week by week. The drafts point out which articles are missing, which are wrong and which are written so that the AI misreads them. How to find those gaps systematically is described in the article on knowledge gaps.
Trust is built in both directions. The team sees what the AI would have said before the customer does, and learns where it is reliable and where it is not. Customers, 71 percent of whom according to Salesforce's State of the AI Connected Customer think a human should check what AI produces, get a reply someone has actually stood behind. The same report shows that the share of customers who trust companies to use AI ethically has fallen to 42 percent, from 58 percent in 2023. That trust is not won back with a bot that guesses.
When is full automation justified?
Full automation is justified for narrow flows where the answer already exists in an approved article, the data can be fetched from a system and the cost of an error is low. Order status, opening hours, delivery terms, how to start a return within policy, password resets. These are flows where a correct answer is the same every time and where the customer would rather have it at 10 pm than the next morning.
The criteria, in the order you should check them:
- The answer exists in a reviewed article, with the exceptions described.
- The data the answer needs, such as order number and delivery status, can be fetched directly from the business system or e-commerce platform. What that takes is described on the integrations page.
- A wrong answer is cheap to correct and does not harm the customer. A wrong "the parcel arrives tomorrow" is annoying; a wrong "you are entitled to a refund" is a promise.
- The customer can reach a person at any time, and the AI hands over by itself when the question goes beyond the article.
- The volume is high enough that the automation saves time the team can actually use for something else.
Gartner expects agentic AI to resolve 80 percent of common customer service issues without human involvement by 2029. The important word is common. The forecast is about the narrow flows, not the ticket where the customer got the wrong item, paid with a gift card and wants to exchange for something out of stock.
What are the risks of releasing everything at once?
The biggest risk is that the AI confidently answers wrongly in your name, and that you discover it when the customer has already acted on the answer. In 2024 a Canadian tribunal held Air Canada to what the company's chatbot had promised a customer about a bereavement discount, even though the policy said otherwise. The tribunal rejected the argument that the bot was a separate legal entity, as reported by CBC. The amount was small. The principle was not: what the company says through its bot, the company has said.
The second risk is counting on a saving that does not arrive. Klarna, which in 2024 said its AI assistant handled two thirds of customer service chats, announced in May 2025 via Bloomberg that it was again recruiting people to customer service, because the focus on cost had, according to the CEO, produced lower quality. Gartner predicts that half of the organisations that planned major customer service headcount reductions because of AI will abandon those plans by 2027, and that over 40 percent of agentic AI projects will be cancelled in the same period, because of cost, unclear value or inadequate risk control.
The third risk is regulation. From 2 August 2026, the AI Act's transparency requirements apply, meaning the customer must be informed that they are interacting with an AI system at the latest at first contact. The personal data in the tickets is also covered by GDPR whoever replies; what that means in practice is in the article on GDPR and AI in customer service. None of the risks is an argument against AI. They are arguments for knowing what the AI answers before it answers everyone.
How do you measure when a flow is ready to automate?
A flow is ready when the drafts for that specific question are sent unchanged almost every time for several weeks, without customer satisfaction or the share of customers coming back differing from the replies the team wrote themselves. It is a decision per question, not for the whole inbox.
| Measure per question | What it says | Ready when |
|---|---|---|
| Share of drafts sent unchanged | Whether the knowledge behind them is right | Close to one hundred percent for at least four weeks, with sufficient volume |
| Share of drafts fully rewritten | Whether there are knowledge gaps | Close to zero |
| CSAT on tickets with AI drafts | Whether the customer notices a difference | Same level as the team's own replies |
| Share of customers coming back on the same question | Whether the reply actually resolved the question | Same level as the team's own replies |
The levels above are our rules of thumb, not an industry standard. What matters is that you set the threshold in advance, per question, and that whoever owns the article makes the decision. Once the flow is released you keep spot-checking a share of the replies and keep the option to switch automation off for that question if the measures slip. That is how Supportifier works: drafts in the inbox first, automation per question when the numbers hold, and always a handover to a person when the answer is missing from the knowledge base. How a chatbot is built on the same principle is described in the article on chatbots that do not guess.
What to do
- Pick the ten most common questions in the inbox and check that each has a reviewed article with exceptions. Write the ones that are missing.
- Let the AI write drafts for those questions, and only those, for two weeks. Every draft goes through an agent.
- Tag every draft: sent unchanged, edited or rewritten. Three buttons or three tags will do.
- Go through the edited and rewritten drafts every Friday. Correct the articles, not the AI.
- Compare CSAT and the share of customers coming back for the AI drafts with the team's own replies.
- When a question holds your thresholds for four weeks, and meets the criteria for narrow flows above, release that question. One at a time.
- Keep spot checks and a clear route to a person. Tell the customer when it is an AI replying.
Common questions
Do agents get slower from reviewing drafts?
No, not when the drafts are built on a reviewed knowledge base. Reading and approving a correct reply takes seconds; writing it from scratch takes minutes. The Generative AI at Work study found 14 percent more issues resolved per hour with AI suggestions the agents themselves chose to use. If the team gets slower, it is because the drafts are poor, and then it is the knowledge base that needs fixing.
Do we have to tell the customer the reply was written by AI?
When the AI answers the customer directly, yes. From 2 August 2026 the EU AI Act requires that anyone interacting with an AI system is informed of it, at the latest at first contact. A draft an agent has read, edited and sent in their own name is a reply from the agent. Regardless of the law, you gain from openness; according to Salesforce, 72 percent of customers think it is important to know whether they are talking to an AI.
How many tickets can be automated?
It depends entirely on how large a share of the tickets are narrow, recurring questions with an answer that exists in an approved article. For an online store, order status, delivery and returns within policy are often a large part of the volume, but the share varies with range and season. Measure it in your own ticket history before promising anything. Gartner's 80 percent forecast applies to common issues in 2029, not to your inbox today.
What happens if the AI answers wrongly in a reviewed draft?
Then the agent should catch it and correct it, and that is the whole point of the review. Log the error, find the cause in the article the AI answered from and correct it there. If the article is correct but the reply still wrong, it is a question where automation should not be considered until the cause is clear. An error caught in review costs a minute. The same error at the customer costs trust.
Sources
- Generative AI at Work — Brynjolfsson, Li and Raymond, NBER, 2023
- State of the AI Connected Customer — Salesforce, 2025
- Gartner Predicts Agentic AI Will Autonomously Resolve 80% of Common Customer Service Issues Without Human Intervention by 2029 — Gartner, 2025
- How can I mislead you? Air Canada found liable for chatbot's bad advice on bereavement rates — CBC News, 2024
- Klarna Turns From AI to Real Person Customer Service — Bloomberg, 2025
- Gartner Predicts 50% of Organizations Will Abandon Plans to Reduce Customer Service Workforce Due to AI — Gartner, 2025
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 — Gartner, 2025
- Transparency obligations under Article 50 of the AI Act — European Commission, 2025