AI in customer service10 min read

Why reviewed AI drafts beat full automation at the start

Short answer

At the start you do not know what the AI does not know. Drafts that an agent reviews before sending give the time saving on the common questions, but every edit shows where the knowledge base falls short, and errors never reach the customer. Full automation is justified for narrow flows where the answer exists in an approved article, the cost of an error is low and the drafts have been accepted unchanged for several weeks.

If you are considering AI in customer service, the choice is rarely between AI and no AI, but between letting the AI answer the customer directly and letting it write drafts the team reviews. Vendors sell the first. Most of those who have succeeded started with the second. This article explains why, when full automation is still right, and how you can tell a flow is ready.

Why do reviewed AI drafts beat full automation at the start?

Because at the start you do not know what the AI does not know, and the only way to find out without customers finding out for you is to have a person read the answers first. The approach is called human-in-the-loop: the AI proposes, a person approves, edits or rewrites, and only then does the reply go to the customer.

The best-documented study of this setup is Generative AI at Work by Brynjolfsson, Li and Raymond, which followed 5,179 customer service agents as an AI assistant suggesting replies was rolled out in stages. Agents could use, edit or ignore the suggestions. Productivity, measured as issues resolved per hour, rose by 14 percent on average and by 34 percent for new and less experienced agents, while the customers' tone in the conversations became more positive and staff turnover fell.

What the study shows is that the gain does not require removing the person. It comes from the answer to the common questions already being written when the agent opens the ticket. What remains is to read, adjust and send, and that takes seconds when the draft is right.

How do reviewed drafts build trust and train the knowledge base?

Every review is a measurement point. When the agent sends the draft unchanged, you know the knowledge behind it is right. When she edits it, you know the article needs correcting or completing. When she rewrites from scratch, you have found a knowledge gap, a question that recurs without an approved answer. None of those signals come from a bot that has already answered the customer.

That is how the knowledge base gets trained: not by the model learning anything, but by the content it answers from getting better week by week. The drafts point out which articles are missing, which are wrong and which are written so that the AI misreads them. How to find those gaps systematically is described in the article on knowledge gaps.

Trust is built in both directions. The team sees what the AI would have said before the customer does, and learns where it is reliable and where it is not. Customers, 71 percent of whom according to Salesforce's State of the AI Connected Customer think a human should check what AI produces, get a reply someone has actually stood behind. The same report shows that the share of customers who trust companies to use AI ethically has fallen to 42 percent, from 58 percent in 2023. That trust is not won back with a bot that guesses.

When is full automation justified?

Full automation is justified for narrow flows where the answer already exists in an approved article, the data can be fetched from a system and the cost of an error is low. Order status, opening hours, delivery terms, how to start a return within policy, password resets. These are flows where a correct answer is the same every time and where the customer would rather have it at 10 pm than the next morning.

The criteria, in the order you should check them:

  1. The answer exists in a reviewed article, with the exceptions described.
  2. The data the answer needs, such as order number and delivery status, can be fetched directly from the business system or e-commerce platform. What that takes is described on the integrations page.
  3. A wrong answer is cheap to correct and does not harm the customer. A wrong "the parcel arrives tomorrow" is annoying; a wrong "you are entitled to a refund" is a promise.
  4. The customer can reach a person at any time, and the AI hands over by itself when the question goes beyond the article.
  5. The volume is high enough that the automation saves time the team can actually use for something else.

Gartner expects agentic AI to resolve 80 percent of common customer service issues without human involvement by 2029. The important word is common. The forecast is about the narrow flows, not the ticket where the customer got the wrong item, paid with a gift card and wants to exchange for something out of stock.

What are the risks of releasing everything at once?

The biggest risk is that the AI confidently answers wrongly in your name, and that you discover it when the customer has already acted on the answer. In 2024 a Canadian tribunal held Air Canada to what the company's chatbot had promised a customer about a bereavement discount, even though the policy said otherwise. The tribunal rejected the argument that the bot was a separate legal entity, as reported by CBC. The amount was small. The principle was not: what the company says through its bot, the company has said.

The second risk is counting on a saving that does not arrive. Klarna, which in 2024 said its AI assistant handled two thirds of customer service chats, announced in May 2025 via Bloomberg that it was again recruiting people to customer service, because the focus on cost had, according to the CEO, produced lower quality. Gartner predicts that half of the organisations that planned major customer service headcount reductions because of AI will abandon those plans by 2027, and that over 40 percent of agentic AI projects will be cancelled in the same period, because of cost, unclear value or inadequate risk control.

The third risk is regulation. From 2 August 2026, the AI Act's transparency requirements apply, meaning the customer must be informed that they are interacting with an AI system at the latest at first contact. The personal data in the tickets is also covered by GDPR whoever replies; what that means in practice is in the article on GDPR and AI in customer service. None of the risks is an argument against AI. They are arguments for knowing what the AI answers before it answers everyone.

How do you measure when a flow is ready to automate?

A flow is ready when the drafts for that specific question are sent unchanged almost every time for several weeks, without customer satisfaction or the share of customers coming back differing from the replies the team wrote themselves. It is a decision per question, not for the whole inbox.

Measure per questionWhat it saysReady when
Share of drafts sent unchangedWhether the knowledge behind them is rightClose to one hundred percent for at least four weeks, with sufficient volume
Share of drafts fully rewrittenWhether there are knowledge gapsClose to zero
CSAT on tickets with AI draftsWhether the customer notices a differenceSame level as the team's own replies
Share of customers coming back on the same questionWhether the reply actually resolved the questionSame level as the team's own replies

The levels above are our rules of thumb, not an industry standard. What matters is that you set the threshold in advance, per question, and that whoever owns the article makes the decision. Once the flow is released you keep spot-checking a share of the replies and keep the option to switch automation off for that question if the measures slip. That is how Supportifier works: drafts in the inbox first, automation per question when the numbers hold, and always a handover to a person when the answer is missing from the knowledge base. How a chatbot is built on the same principle is described in the article on chatbots that do not guess.

What to do

  1. Pick the ten most common questions in the inbox and check that each has a reviewed article with exceptions. Write the ones that are missing.
  2. Let the AI write drafts for those questions, and only those, for two weeks. Every draft goes through an agent.
  3. Tag every draft: sent unchanged, edited or rewritten. Three buttons or three tags will do.
  4. Go through the edited and rewritten drafts every Friday. Correct the articles, not the AI.
  5. Compare CSAT and the share of customers coming back for the AI drafts with the team's own replies.
  6. When a question holds your thresholds for four weeks, and meets the criteria for narrow flows above, release that question. One at a time.
  7. Keep spot checks and a clear route to a person. Tell the customer when it is an AI replying.

Common questions

Do agents get slower from reviewing drafts?

No, not when the drafts are built on a reviewed knowledge base. Reading and approving a correct reply takes seconds; writing it from scratch takes minutes. The Generative AI at Work study found 14 percent more issues resolved per hour with AI suggestions the agents themselves chose to use. If the team gets slower, it is because the drafts are poor, and then it is the knowledge base that needs fixing.

Do we have to tell the customer the reply was written by AI?

When the AI answers the customer directly, yes. From 2 August 2026 the EU AI Act requires that anyone interacting with an AI system is informed of it, at the latest at first contact. A draft an agent has read, edited and sent in their own name is a reply from the agent. Regardless of the law, you gain from openness; according to Salesforce, 72 percent of customers think it is important to know whether they are talking to an AI.

How many tickets can be automated?

It depends entirely on how large a share of the tickets are narrow, recurring questions with an answer that exists in an approved article. For an online store, order status, delivery and returns within policy are often a large part of the volume, but the share varies with range and season. Measure it in your own ticket history before promising anything. Gartner's 80 percent forecast applies to common issues in 2029, not to your inbox today.

What happens if the AI answers wrongly in a reviewed draft?

Then the agent should catch it and correct it, and that is the whole point of the review. Log the error, find the cause in the article the AI answered from and correct it there. If the article is correct but the reply still wrong, it is a question where automation should not be considered until the cause is clear. An error caught in review costs a minute. The same error at the customer costs trust.

Sources

Rickard Collander

By

Rickard Collander

Rickard har arbetat med kundservice och Customer Success i snart tjugo år, på både köparsidan och leverantörssidan. Han började på Gjensidige med ansvar för kundtjänst och telemarketing, var i knappt fem år kundservicechef på Bonnier Tidskrifter med ansvar för avtal, servicemål och kvalitet i en outsourcad kundservice i alla kanaler, och har därefter arbetat som managementkonsult på Omnisale och som chef över Telias outsourcade kundservice. Han har varit CCO på kontaktcenterbolaget Releasy och senast Head of Customer Success på Scania, där han ledde kundframgång och support för digitala tjänster globalt. Under de senaste åren har han arbetat med AI-implementationer i bolag som Dold Adress, Omnio och Axfina. Han grundade Successifier och arbetar med hur kundserviceorganisationer bygger skalbara arbetssätt, mäter rätt saker och fångar risker innan kunder lämnar.

LinkedIn (opens in a new tab)
  • ai
  • automation
  • human-in-the-loop
  • knowledge base

This article is also available in Swedish: Varför AI-utkast som granskas slår full automation i början

Next step

Start from your everyday work.

Tell us which question takes time today. We go through what a first step could look like.

A clear first step beats a big promise.

See how the knowledge work, the review and the first channel fit together.

About the knowledge analysis
From the same question
to a better answer.
Book a walkthrough