The short version
Draft, don’t send. Automate the retrieval and the first pass, keep the human on the outbound. Measure handling time rather than deflection rate. And expand to auto-sending only for narrow, well-understood categories where being wrong is cheap and reversible.
Why not just auto-reply?
Because the cost of a wrong answer is asymmetric. A support reply is a statement by your company to a customer: it can commit you to a refund, contradict your terms, or land badly on someone already angry. A drafted reply that a human glances at and sends costs a few seconds and eliminates that entire class of risk.
There is also a practical argument. You do not know your accuracy until you have watched it on real traffic for a few weeks, and drafting gives you that measurement for free — every edit a human makes is a labelled example of where it was wrong. Teams that auto-send from day one get the same education, just from customers instead.
What should be automated first?
The three things that happen before anyone writes a word:
- Categorisation. What kind of request is this, how urgent, which team.
- Retrieval. Pull the order, the account, the previous conversation and the relevant policy into the ticket, so nobody is searching three systems.
- Drafting. A suggested reply, grounded in what was retrieved, in your tone.
In most inboxes the actual typing is a minority of the handling time. The bulk goes on working out what the customer means, finding the context, and deciding who owns it. Automating that is where the saving is, and none of it requires trusting the system to speak to anyone.
How do we stop it inventing policies?
By grounding it in your actual documentation and requiring it to cite what it used. A drafted reply that shows which policy paragraph and which order record it drew on is checkable in a glance; one that reads fluently with no provenance has to be verified from scratch, which removes the saving.
The related discipline is teaching it to decline. A draft that says “I could not find a policy covering this — escalating” is a good outcome. A system tuned to always produce a confident reply will confidently invent a returns window, and the customer will hold you to it.
What about tone and brand voice?
Easier than people expect and worth doing properly, because a reply that sounds nothing like your company is the thing customers notice first. Supply real examples rather than adjectives: twenty genuinely good past replies teach tone far better than a paragraph describing it.
Pay particular attention to the awkward categories — complaints, refusals, apologies. Generic systems default to a customer-service register that reads as insincere in exactly the moments where sincerity matters. If you only curate examples for one category, curate those.
What should we measure?
Handling time per ticket, first-contact resolution, and the edit rate on drafts. Those three tell you whether it is working.
Not deflection rate. Optimising for how many customers never reach a human is how you end up with a system that traps people in loops and a satisfaction score that quietly collapses. The deflection number goes up and the business gets worse, which is the worst kind of metric — one that improves while the thing it proxies for degrades.
The edit rate is the most useful early signal. If agents are rewriting eighty percent of drafts, the retrieval or the tone is wrong and you should fix that before expanding scope. If they are sending most drafts unchanged, you have a candidate for narrow auto-sending.
When is auto-sending actually safe?
For narrow categories where the answer is factual, the source is authoritative, and being wrong is cheap and reversible. Order status, delivery tracking, opening hours, password resets, confirming receipt of something.
Keep a human on anything touching money, anything from a customer already complaining, anything legal or contractual, and anything the system was not confident about. Those categories are a small share of volume and nearly all of the risk — which is the whole reason the split works.
What does this do to the support team?
Changes the shift rather than ending it, if it is done honestly. The volume of routine handling drops substantially, and what remains is the harder work: the angry customer, the case that does not fit the policy, the problem nobody has seen before.
Two things follow. Support becomes a more demanding job rather than a less demanding one, which has implications for who you hire and what you pay. And the people who were handling the routine volume are your best source of the exception rules the system needs, so involving them early is both decent and practical — a team that expects to be replaced does not explain the edge cases.
What do we need before starting?
Less than most vendors suggest. Your help centre or policy documents in a form something can read, access to whatever holds order and account data, and a few hundred past tickets with the replies that were actually sent. That last set is the most valuable asset you have and most teams do not realise it — it is simultaneously your tone guide, your training data and your test set.
What you do not need is tidy documentation. Part of the value of the first pass is discovering which of your policies are contradictory or missing, because those are the same gaps that currently cause agents to escalate to a supervisor.
How long before it is useful?
Triage and retrieval usually earn their keep quickly, because they are low-risk and the saving is immediate. Drafting takes longer to become genuinely useful, because tone and accuracy both need iteration against real tickets.
Plan for a period where agents are correcting drafts more often than sending them and treat that as the work rather than as failure. Every correction is a labelled example, and the edit rate falling week on week is the clearest evidence you have that it is converging.
What happens when a customer realises it is automated?
In the drafting model, nothing, because a person did send it and did read it. That is a meaningful honesty advantage over auto-reply and worth preserving.
If you do move to auto-sending for narrow categories, say so plainly in the message rather than hoping nobody notices. Customers accept an automated order-status reply without complaint; what they object to is discovering that a message signed with a human name was not from one. The reputational cost of being caught disguising it is far higher than the cost of disclosing it.