Back
Blog / 
Strategy & ROI

Building a defensible business case for customer-service AI

written by:
David Eberle
Two stacks of paper slips balanced on a level beam, with one slip set aside.

There is one way a customer-service AI business case fails that is worth guarding against first. A pilot reports a saving, say a large percentage of writing time, and the case multiplies it by headcount and salary. Finance then asks three questions: writing time out of what? for which tickets? and where did the saved time actually go? If the pilot cannot answer, the case is not wrong, it is undefined.

This guide is a method for making it defined. It walks through the reconciliation steps that have to happen before a currency sign appears, with a worked hypothetical model you can copy and replace with your own numbers. It assumes you have already decided roughly what to automate; if not, start with how much of your customer service AI can actually automate, which this article builds on.

Where this comes from. The method comes out of four conversations Typewise held in 2026 with teams running a Typewise pilot: three pilot reviews with travel and logistics customers where the business case was argued over live, and a review with an online retailer. I was in all four. The observations are anonymised and marked as reported; the steps are our recommendation; the worked model is invented.

Step 1: agree the unit of work

Messages, tickets and cases are different things, and AI tools, help desks and finance each tend to count a different one.

  • A message is one inbound or outbound communication.
  • A ticket is a help-desk record. One customer case often spans several tickets: the original request, a follow-up, a reopen after delivery.
  • A case is the customer's problem, from first contact to the point where nothing further is needed.

An assistant that measures time saved per message says nothing about the case. A help desk that reports ticket volume says nothing about repeat contacts. Write down the conversion you will use, for example an observed average of 1.6 tickets and 3.4 messages per case in your own data, and use the case as the unit for the money question. Value is created when a case is finished, not when a message is sent.

This is not a theoretical distinction. In one pilot review, a travel company's service lead walked me through the business case she had built at the start: a per-ticket handling time that covered opening, reading, thinking and writing, against which our measurement counted writing time only, a fraction of her figure. In the same conversation it came out that one customer case can be a single ticket or several, depending on whether the customer phones or keeps writing, and that nobody could say the average. Neither side was wrong; we were counting different things.

Step 2: baseline per cohort, per channel, per season

A single "average handling time" hides most of what the business case needs.

  • Cohorts. Agents who split their day between phone and email have a different per-ticket baseline from those who only write. Specialists handling complex tasks differ again. The travel team above had colleagues who mostly phone, colleagues who mostly write, and one who handled several times the email volume of the others; a single average would have described nobody. Baseline each cohort the AI will actually touch.
  • Channels. Chat, email and contact-form requests carry different volumes and different handling patterns. A successful chat pilot does not validate email; the scenarios have to be calculated separately.
  • Seasons. Volume in a quiet month can be half of what a peak month brings. A baseline measured in one period and applied to the whole year misstates both the saving and the capacity you need.

The baseline is also where the counterfactual lives. A logistics group's project lead told us her team had decided against turning their long email templates into snippets, because finding the existing template was quicker than assembling the message; snippets stayed useful for short sections only. That is a reported decision about one team's workflow, and it changes the counterfactual for that cohort: an assistant that rebuilds those emails from snippets may be slower there and faster only for the short, varied replies. Ask what the current fastest path is for each ticket type before you assume the AI beats it.

Step 3: separate three kinds of saving

Keep these apart in every table you produce, because they do not add up and they are not interchangeable.

  1. Writing time. The minutes an agent spends composing a reply. This is what assistant-level tools tend to measure, commonly modelled from feature usage and session timing rather than from a controlled before-and-after. It is real, but it is a share of a share.
  2. Handling time. Everything the agent does for the ticket: reading, researching in other systems, deciding, writing, updating records. If research in the reservation or order system stays manual, a large writing-time saving becomes a modest handling-time saving. In a second travel pilot review, the sponsor asked where the promised overall saving was. The team lead's answer was that for most requests staff still had to look things up in the reservation system before answering, so the assistant shortened the writing step only, and that some mailboxes could not use it at all because contact-form requests arrived without any body text. The saving was real and smaller than the headline, and the research step was the reason.
  3. Whole-case resolution. The case finishes without a person. This is the only saving that removes work rather than shortening it, and it is the one that needs a clean definition of "resolved", including reopens and partial handoffs. See what counts as a resolution for that definition.

A worked hypothetical

The numbers below are invented to show the mechanics. They are not a benchmark, not a Typewise result and not a prediction for your team.

Illustrative assumption for this example only: in this cohort each case is recorded as exactly one ticket, so tickets and cases coincide and no repeat tickets inflate the count. Where your help desk records several tickets per case, convert to cases first (step 1) and run the same arithmetic on cases. Assume an email cohort with an average handling time of 9 minutes per ticket: 3 minutes reading and researching, 4 minutes writing, 2 minutes updating records. The cohort handles 10,000 tickets a month. To keep the two savings lines from double-counting, split the volume into disjoint sets first: 1,500 tickets (15%) of one type are candidates for autonomous handling, and the remaining 8,500 are the assisted pool, of which the assistant is used on 60%, that is 5,100 tickets. No ticket appears in both sets.

LineAssumptionResultNote
Writing-time saving reported by the pilot40% of writing time1.6 min per ticketOf a 4-minute writing step
As a share of handling time1.6 of 9 min18% per touched ticketResearch and record updates unchanged
Assisted tickets (disjoint from the autonomous set)60% of the 8,500 non-autonomous tickets5,100 ticketsSome ticket types cannot use the assistant at all
Gross writing-time minutes5,100 × 1.6 min8,160 minUnrounded
Review cost for AI drafts5,100 × 0.5 min2,550 minReading a draft is not free
Net assisted saving8,160 − 2,5505,610 min = 93.5 hoursAbout 6.2% of the cohort's 90,000 handling minutes

A headline of "40% time saved" has become about 94 hours a month of capacity. That is still worth having. It is also a fraction of what the headline implied, and the reconciliation is what makes the number survive a finance review.

Now add the second line for whole-case resolution on the disjoint autonomous set. Treat the 1,500 tickets of that type as autonomous attempts: the AI reached a verified candidate resolution on each (an outcome confirmed by the performing system, no pending work), and the repeat-contact window then has to pass. Suppose 150 of them (10%) are reopened inside the window and 1,350 become verified resolutions once the window has matured. Two separate things follow from that, and they should not be merged. The resolution state is 1,350 resolved and 150 known unresolved out of 1,500 attempts (1,350 / 1,500 = 0.90), using the definitions in what counts as a resolution. The net human-effort saving is computed across all 1,500 attempts against the comparable counterfactual: the 9-minute baseline is what the same attempted tickets would have cost people over the same horizon, with all residual work included on both sides. So 1,500 × 9 = 13,500 baseline minutes, minus the human work the reopened share actually consumed, 150 × 12 = 1,800 minutes, gives 13,500 − 1,800 = 11,700 net minutes, or 195 hours. That net figure is an effort calculation over all attempts; it is not a claim that all 1,500 resolved autonomously. Tickets with no verified outcome, where the customer simply went silent, are neither attempts that resolved nor savings; they sit in an unknown bucket until audited.

The two lines together give 5,610 + 11,700 = 17,310 minutes, or 288.5 hours a month, computed before any rounding and only then rounded for the slide. Each line has a different confidence, a different measurement and a different risk, and because the sets are disjoint, and tickets equal cases in this example, the lines can be added without counting a ticket or a case twice.

Step 4: subtract the residual human work

Every automated or assisted ticket can leave work behind:

  • a draft that is reviewed and edited rather than sent as is;
  • a case the AI handled "partially", where a person completes the action;
  • a reopen, where the customer returns because the answer did not hold;
  • exceptions the AI routed to a queue, which someone now has to triage.

Pilot dashboards may show the share of conversations the AI participated in. That is not the share it finished. Before you value anything, inspect the partially resolved cases and the human time they consumed; early figures without a stable denominator and an audited outcome should not enter the business case.

Step 5: value the capacity honestly

Freed hours are capacity, not cash, until something changes.

  • Cash saving happens if you reduce agency or overtime spend, avoid planned hires, or redeploy people into revenue work with a measurable result.
  • Capacity gain happens if the same team handles growth, clears backlog, or takes over tasks from more expensive colleagues. That is valuable but it is not a cost line that disappears.

Two practical rules follow. First, use your real labour mix. When we put a case built on a domestic salary benchmark in front of an online retailer, they corrected it on the spot: part of their service team sits in a lower-cost country, so the per-agent figure was far below what we had assumed, and their plan was for senior staff to handle complex work while junior roles were absorbed over time. A domestic benchmark overstates savings for a team like that and understates the value of freeing senior people from routine work. Second, give the freed capacity a named destination before go-live: advisory work, outbound contact, a backlog, a service-level improvement. The travel service lead above already had hers, advisory work and outbound calls; that was her plan, not a measured result, which is exactly what a destination is before go-live. Capacity without a destination tends to evaporate into the same work spread thinner, and the case then cannot be audited afterwards.

Step 6: put the costs on the same page

A defensible case shows the full cost side over the same horizon as the savings: licence or usage fees, integration and data-readiness work, the time your own team spends reviewing drafts and maintaining instructions and knowledge, evaluation effort, and the people who own the system after the pilot. Choose a horizon, commonly twelve months, and state it. Then show the net per month, with the writing-time, handling-time and whole-case lines still visible, so the reader can see which line the result depends on.

Pricing and costs differ by vendor and contract; use the actual figures you have been quoted rather than any example here.

The worksheet

Before the case goes to a decision-maker, every item should have a number and a source.

  • [ ] Unit of work agreed (message, ticket, case) and the conversion between them from your own data.
  • [ ] Baseline handling time per cohort, per channel, with the seasonal range.
  • [ ] The counterfactual for each ticket type: the fastest current path, including templates.
  • [ ] Writing-time, handling-time and whole-case savings on separate lines, each with its measurement method, computed on disjoint ticket sets so they can be added, with unrounded arithmetic shown before rounded totals.
  • [ ] Coverage: the share of volume each line applies to, and the types it cannot apply to.
  • [ ] Residual human work: review time, partial handoffs, reopens, exception triage.
  • [ ] Resolution definition: verified outcome, no pending work, repeat-contact window passed; unknown silent cases excluded from the resolved count.
  • [ ] Labour mix used, and the named destination for freed capacity.
  • [ ] Full cost side over a stated horizon, including internal ownership time.
  • [ ] A before-and-after measurement plan, not only a dashboard percentage.

Common ways the case goes wrong

  • Different denominators. The baseline measures handling time, the pilot measures writing time, and the slide multiplies one by the other.
  • One channel standing in for all. Chat results presented as an email forecast.
  • Dashboard participation presented as resolution. High AI involvement with high partial handoff.
  • Pilot enthusiasm as evidence. Early excitement in the team is not a result; the honest conclusion comes after the outcome is measured, and some teams find the saving smaller than expected once research time is counted. In the second travel review, the sponsor and a manager said as much themselves: the team's early enthusiasm had dipped, and they wanted an honest conclusion before acting as a reference for anyone.
  • Savings with no destination. Hours that nobody redeployed and nobody can find later.

Limits of this method

It produces a conservative case. If your cohort's work is mostly writing, or the AI finishes whole cases at scale, the reconciled number will be closer to the headline. The point is not to shrink the result but to make each part of it traceable. The model also ignores second-order effects, such as faster responses improving customer outcomes; include those only with their own measurement.

FAQ

Why is writing-time saving not the same as handling-time saving?

Writing is one step of handling a ticket. If reading, researching in other systems and updating records stay manual, a large saving on the writing step is a much smaller saving on the whole ticket. Model them on separate lines.

Can we use the pilot's dashboard percentage directly in the business case?

Only once you know what it measures and over which denominator. Dashboard figures can be modelled from feature usage or show AI participation rather than finished cases. Reconcile them against a before-and-after measurement of actual effort first.

How should we value freed capacity?

As cash only where spend actually falls or a hire is avoided; otherwise as capacity with a named destination, such as backlog, growth or redeployed work. State which it is, and use your real labour mix rather than a generic salary.