Back
Blog / 
AI Agents

What agents change in AI-written replies

written by:
David Eberle
A printed reply with three kinds of pencil edit and a magnifier resting on its margin.

When an AI drafts a reply and a person edits it before sending, the edit is the most honest evaluation you will get. It is produced by someone who knows the customer, the product and the policy, on a real case, with their own name going on the message. It is easy to throw away: a team that tracks only whether the draft was used, and perhaps how long the agent took, never looks at what changed.

This article is about looking. It gives you a taxonomy for coding edits, a way to run the analysis without a research team, and a way of reading edit rates as one input into the decision about an intent's next evaluation stage. It is deliberately not a tone-of-voice guide; the existing tone of voice guide for support teams covers style. Tone is one of five things agents change, and usually not the most important one.

Where this comes from. The taxonomy comes out of three conversations Typewise held in 2026: a demo a colleague ran for a retailer whose advisers judge AI drafts from two systems, a reference call in which an established customer explained its specialist set-up to a prospective one, and a measurement deep-dive with a logistics group piloting our writing assistant. I was in the last two. The observations are reported practice, anonymised; the taxonomy and the decision rule are our recommendation, and the worked month is invented.

Why "used the draft" is the wrong metric

A draft that was sent unchanged may be excellent, or may have been sent by someone in a hurry who did not check it. A draft that was heavily edited may be bad, or may have been a good starting point for a case that needed personal judgement. Usage tells you about adoption. It does not tell you about quality, and it does not separate adoption from blind acceptance, which is the thing your risk owner actually worries about.

The retailer in the demo had done the useful version of this. For two months one system produced suggested replies that advisers could change; the team tracked the modification rate per request type and used it to decide what could be automated directly and what was not ready. With a second system, advisers graded each draft as acceptable, not acceptable or needs improvement. Both are reported practices from one team, and both are exactly the signal this article is about.

Edits carry the signal usage lacks, provided you classify them. Five categories cover most of what people change.

The edit taxonomy

CategoryWhat changed (the symptom)What it tells youFix, once the cause is traced
FactsA date, amount, policy term, product detail or status was wrong or missingThe most serious symptom. It does not say where the error came from: the cause may be a wrong or stale source, retrieval picking the wrong passage, the model misreading or inventing, missing conversation context, or a reviewer who was mistakenDepends on the traced cause: correct the source, fix retrieval or connect the data, adjust instructions or the model setup, carry the missing context, or recalibrate reviewers; until traced, keep the intent in reviewed drafts
RelevanceThe draft answered a different question, missed one of several questions, or ignored the customer's situationA symptom with several possible causes: routing, missing context, retrieval of the wrong material, or the model's reading of the requestDepends on the traced cause: improve routing, carry prior context, fix retrieval, test multi-question cases
VerbosityText removed or compressed; the draft said too much or repeated itselfUsually a style symptom: length and structure defaults do not match the channel; confirm it is not hiding a relevance problemChange instructions per channel; shorter defaults for chat
ToneWording softened, warmed, formalised or made more directStyle mismatch with the brand or the customer's stateTone instructions and examples; see the tone guide
Next actionThe agent added or changed what happens next: a question, a handoff, an action taken in a system, a commitmentThe draft stopped short of resolving; the real work was outside the text, or the design did not allow the actionConnect the action and grant it where appropriate, or decide the intent stays human

The order matters. A team that edits mostly for tone has a style problem, which is cheap. A team that edits mostly for facts has a serious problem whose cause still has to be found: an edit is a symptom, and a wrong fact can come from a stale or wrong source, from retrieval picking the wrong passage, from the model misreading or inventing, from missing conversation context, or from a reviewer who was themselves mistaken. Investigate before assigning the fix. A team that edits mostly for next action has discovered that the reply was never the hard part of the ticket.

How to run the analysis

You do not need tooling beyond a spreadsheet and an hour a week.

  1. Sample. Take a representative sample of drafts per intent per week, say twenty drawn at random from all drafts the intent produced, edited and unedited alike. The denominator is all drafts, not the edited ones; an edit rate computed from edited drafts only says nothing about the population. Have a second person check the unedited drafts in the sample for whether they should have been edited; that is your blind-acceptance check.
  2. Code. For each edited draft, record which categories were edited (several can apply), a one-line note, and, for fact and relevance edits, where the error came from once traced (source, retrieval, model, context, reviewer). Two coders on the first batch, compared, so the categories mean the same thing to everyone.
  3. Count per intent, with uncertainty. Edit rates aggregated across all intents are meaningless; a 40% edit rate may be 5% on order status and 90% on complaints. Report the rate and the dominant category per intent, with the sample size next to it: twenty drafts cannot distinguish a 5% rate from a 15% one, so small samples are a reason to keep observing, not a verdict.
  4. Look at the unedited sends. If reviewers find facts that should have been corrected in drafts that went out untouched, your usage metric was flattering you. This is the number to show whoever signs off on expanding scope.
  5. Decide. Use the rule in the next section.

Some teams add a simple adviser verdict at review time, such as acceptable, not acceptable, needs improvement. That is useful alongside the taxonomy, not instead of it: the verdict says how bad, the category says why.

From edit rates to the next evaluation

Edit data is one input to the decision to move an intent from reviewed drafts towards autonomous handling. It is never the whole decision. A workable reading, over a defined observation period such as two months:

  • Candidate for the next, supervised evaluation stage: low edit rate in a representative sample, with the remaining edits tone or verbosity only, nobody corrected a fact and nobody added an action. That earns a supervised live evaluation with outcome tests, not a switch to autonomous handling. Promotion still depends on representative outcome tests over the repeat-contact window, a sample large enough that the low rate is not noise, the permissions and integration reliability the action needs, and the risk owner's sign-off.
  • Not ready: any meaningful rate of fact or relevance edits. Trace the cause (source, retrieval, model, context or reviewer) and fix it; automating now ships the error.
  • Not an automation candidate yet: next-action edits dominate. The text is fine and the work is elsewhere; connect the action first, or keep the intent with people.

Combine this with outcome data before promoting anything. A draft that agents rarely edit can still produce a reply the customer comes back about. Reopen rates for AI-handled versus human-handled tickets are the check, and they need the same denominator and observation window to be comparable; see what counts as a resolution. If the comparison shows AI-handled tickets reopening more often than human-handled ones, that is a reason to look at the fact and relevance categories again and trace their causes, not a verdict on the approach. The same retailer had run that comparison in its reporting and found reopen rates for AI-handled tickets noticeably higher than for tickets handled only by people; the service lead had not yet completed a ticket-by-ticket analysis of the reopened tickets, and the two sets were not shown to be comparable, which is exactly why we treat such a result as a reason to trace causes.

What the edits also reveal about the workflow

Beyond individual intents, patterns across edits point at workflow decisions.

  • Batching questions. If agents routinely merge several one-question-per-message drafts into one message that asks everything at once, the troubleshooting flow is correct and irritating. The AI owner at the consumer-products customer told the prospective customer on the reference call that their specialists had initially asked one basic troubleshooting question per message, customers took it badly, and the fix was to instruct the specialist to ask the relevant checks together in one email. Instruct the AI to ask the relevant checks together.
  • Snippets versus full templates. Where agents replace a draft assembled from short snippets with an existing long template, the template is faster for that case. The logistics group's project lead told us her team had decided against turning long templates into snippets because finding the template was quicker; snippets stayed for short sections. She also wanted the calculation behind any time-saving figure in writing, including what benchmark it assumed, because employee representatives would ask. The assistant is useful for the short, varied replies and not for the long standard ones. Measure the counterfactual per ticket type before claiming a saving.
  • Review burden. Time spent reading and editing drafts is real work. If the edit analysis shows most drafts need substantive changes, the assistant is adding effort for that intent; this is a plausible outcome in multilingual or technical settings where drafts need checking against specialised knowledge. Remove it from those intents rather than defending the average.
  • Who reviews. Reviewing and tuning AI output is a distinct skill from handling customers, and not every agent wants it. The consumer-products AI owner said so from experience: thinking in systems is not a skill every agent has, and one agent in her team had grown into it. The people who produce the most useful edit notes are candidates for a defined reviewer role; see the operating model.

A worked, hypothetical month

Invented numbers to show the shape of the output.

IntentDrafts sampledEditedDominant categoryUnedited sends with a missed factDecision
Order status80 (random, all drafts)9%Verbosity0Candidate for supervised evaluation with outcome tests; not yet autonomous
Return eligibility80 (random, all drafts)31%Facts (return window by country)3Not ready; trace the cause (traced here to stale country data) and fix it
Delivery complaint60 (random, all drafts)72%Next action (carrier claim)1Keep human; connect carrier claim first
Product compatibility40 (random, all drafts)55%Relevance2Not ready; traced to missing product attributes and to retrieval picking the wrong variant

Only one intent earns the next evaluation stage, and the reasons the others do not are three different problems with three different owners. The sample sizes are small, so each rate carries wide uncertainty; the table is a diagnostic snapshot, not a measurement of the population.

Limits

Edits are evidence about drafts a person saw; they say nothing about conversations the AI handled alone, which need outcome data instead. Edits are symptoms, not diagnoses: the category tells you what was wrong, not yet why. Coding is subjective at the margins, which is why two coders calibrate first. Small samples from a few intents are not an edit-frequency study of the industry, and a sample of edited drafts only cannot give a population edit rate; the sample has to be drawn from all drafts. And a low edit rate can mean trust or inattention; the unedited-send check is what distinguishes them.

FAQ

Is a low edit rate enough to automate an intent?

No. A low edit rate in a representative sample, with only tone or verbosity edits, earns the next supervised evaluation stage. Autonomous handling still depends on representative outcome tests over the repeat-contact window, a sample large enough that the rate is not noise, the permissions and integration reliability the action needs, and the risk owner's sign-off. A low edit rate alone can also mean drafts were sent without being read.

How many drafts do we need to code?

Enough per intent to see the dominant category and to narrow the uncertainty on the rate: a fixed weekly random sample drawn from all drafts of the intent, edited and unedited, over an observation period such as two months. Report the sample size with every rate. Aggregate rates across intents hide the decision you are trying to make.

Agents mostly edit for tone. Is that a problem?

It is the cheapest category to fix: instructions and examples. The categories to worry about are facts, which need the cause traced (source, retrieval, model, context or reviewer), and next action, which means the reply was never the hard part of the ticket.