Back
Blog / 
Automation

What should a customer-service AI be allowed to do?

written by:
David Eberle
Three nested zones for reading, drafting and acting, with a lock and an approval stamp guarding the innermost zone.

The obvious question about a customer-service AI is whether it can do something: look up an order, change an address, issue a refund. The more useful question is whether it may, for whom, and what stops it when it should not. Those are permission decisions, and they are made badly when they are made once, for the whole agent, instead of per action.

This guide gives you a way to make them per action. It separates four rights, lists the checks an executing permission has to pass, and shows why an instruction in a prompt is not the same thing as an enforced limit. It is written for the people who have to sign off: the service owner, the person responsible for the connected systems, and whoever answers for security.

Where this comes from. The guide draws on five conversations Typewise held in 2026 with teams deciding what an AI agent may do: a pilot review with an online retailer running a shopping agent, an implementation call with a large enterprise customer's governance team, a workflow session with a company that handles event registrations, a demo for a fashion retailer's service team, and a conversation with an independent analyst who works on AI governance. I was in four of them; a colleague ran the demo. The observations below are anonymised and marked as what those teams reported or asked for. The four rights, the matrix and the five checks are our recommendation, and the matrix values are invented.

Four rights, not one switch

Treat every capability the AI could have as one of four rights. They are cumulative in risk, not in sequence; an action can sit at any level.

  • Read. The AI may look something up and use it in a conversation. Order status, delivery window, the customer's plan. The risk is disclosure, so the question is who may see what.
  • Recommend. The AI may propose a next step to a person: "this looks like a duplicate booking, consider merging". A person decides. The risk is a wrong recommendation accepted without thought.
  • Draft. The AI may prepare a customer-facing message or a record change that a person reviews and releases. The risk is review that becomes rubber-stamping.
  • Execute. The AI may complete the action without a person in the loop for that case. The risk is the full consequence of the action, immediately.

The distinction matters because the controls differ. A read permission is governed by identity and data scope. A draft permission is governed by review quality. An execute permission needs a deterministic gate at the system that performs the action. Writing "the agent handles refunds" in a project plan does not say which of these you mean, and a team can discover, late, that its members meant different things.

Decide per action: a worked permissions matrix

The matrix below is a hypothetical example for an online retailer. It is illustrative; your actions, thresholds and owners will differ. What matters is the shape: one row per action, one right per row, and a named reason for anything below "execute".

ActionRight grantedWhy not moreWhat enforces itOwner
Order status lookupReadCustomer must be verified before the AI sees account dataOrder system only answers for an authenticated sessionService lead
Delivery address changeRecommendCompany policy routes address changes through the carrier; the API could do it, policy says noWrite endpoint not connectedLogistics owner
Return label for an eligible orderExecuteLow value, reversible, rules are explicitEligibility checked by the order system, not by the promptService lead
Refund up to a small fixed amountDraftTeam wants to watch outcomes for a quarter before removing reviewRefund endpoint requires a reviewer's approval tokenFinance
Refund above that amountRecommendOutside the AI's authority by policyEndpoint rejects any request without a human approverFinance
Merge two customer recordsRecommendA wrong merge is costly and hard to undo; when in doubt, keep both and flagMerge is a human-only operationData owner
Cancel a subscriptionDraftRetention step must be offered first; a person releasesCancellation endpoint behind review queueService lead

Two rows deserve attention because they show a policy and an integration pointing in different directions.

The address change. In the demo for the fashion retailer, the service team found the demonstrated actions easy to use and then named the one they could not: changing a delivery address, which their policy does not allow; what they could do instead was guide the customer to make the change directly with the carrier. That is the reported part. The rest is our reading of why rows like this exist: once a parcel is with the carrier, a change made in the shop system may not reach the parcel and can leave a record that no longer matches reality, so the right is "recommend" regardless of what an integration can technically do. If you only ask "what can the API do", you grant this one by accident.

The record merge. Deduplicating customers looks like harmless housekeeping, but a false merge gives one person access to another's history. The team that handles event registrations described to us how they treat duplicate registrations today, before any AI is involved: only when they are confident it is the same person do they invalidate the later registration, keep the original and resend the valid confirmation; when they are not confident, they keep both and let later behaviour, such as which of the two tickets actually gets scanned, settle it. They do not merge histories. Their rule is asymmetric because a wrong consolidation is the costlier error. The matrix row is our analogy to that rule, applied to customer records: the right is "recommend", and the permission documents the cost of the error, not just the accuracy of the match.

Five checks before anything may execute

An "execute" row should not be granted unless all five of these are true. They are necessary conditions, not a complete safety guarantee: a row that passes all five still needs the evaluation, approval and incident design described elsewhere in this archive. Use them as a sign-off list.

  1. Identity is verified by the system, not asserted by the chat. The AI may believe it is talking to the account holder; the backend must confirm it independently, because the same backend is callable from outside the conversation. The online retailer's engineering lead made this point to us bluntly in the pilot review: their server cannot take an external agent's word that an email address is fine, because their API can be called by anyone on the internet, so verification has to happen on their side. From our platform's side, their deliberate verification challenge initially looked like a failed tool call. If your verification step looks like that to the agent, treat it as a design gap, not a nuisance.
  2. The policy has a source and an owner. Who decided that returns within the window are approved automatically? Where is that written? An action without a traceable policy is not an automated decision; it is an unowned one.
  3. The gate is deterministic. The thing that stops an out-of-bounds action must be a rule the system applies every time: an amount limit in the refund endpoint, an eligibility check in the order system, a required approver. A sentence in the instructions is guidance, and guidance is probabilistic (more on this below).
  4. Completion is confirmed by the system that did the work. An agent that says "I'll check that for you" and then stalls has not acted. Test for it deliberately: a promise followed by silence is one of the cheapest failures to produce in a test and one of the most expensive to ship. Require a tool result, set a timeout, and define what happens when the result never arrives: a timeout means the outcome is unknown, not that the action failed, so the design must query the authoritative status using the request's transaction or idempotency identity, must not retry blindly, and must tell the customer that confirmation is pending. A promise in the transcript is not evidence that the action happened.
  5. There is a human fallback with context. When the gate stops the action, the case goes to a person with everything the AI already established, including the verification state and the lookups it performed. A handoff that drops that context produces the unnecessary refusals and repeat questions customers notice most. Our guide to handoff ownership covers what the receiving person needs.

If a check fails, the answer is not automatically "one level down". Some failures block the row entirely: if identity cannot be verified by the system, or the AI is not authorised to read the data the action depends on, then a draft or a recommendation built on that data would leak the same protected information, so the row has no permission at all until the prerequisite is fixed. Downgrade to draft or recommend only when that mode independently satisfies its own data and action prerequisites; a draft of a refund is acceptable when the agent may read the order and a person releases the refund, not when the agent should not have seen the order. The matrix is designed so you can promote a row later, once you have evidence from drafts and recommendations.

Instructions are not enforcement

This misunderstanding is worth being precise about, because it is the one that ends up in governance documents as if it were a control.

Suppose you tell the agent: "You may issue refunds up to 50. For anything above that, ask a colleague." The agent will follow that most of the time. It is a language model reading an instruction; it is not a rule engine. For a low-stakes action, "most of the time" may be acceptable. For money, account access or anything regulated, it is not a control you can present to an auditor or a works council.

I ran into this myself while demonstrating our configuration assistant to the analyst. Asked to make invoices above a small amount require human approval, the honest answer was that the action could not be split at a monetary threshold without development. The two real options were an approval on every invoice, which is a hard gate, and an instruction to escalate above the amount, which the agent would follow most of the time and not always. The analyst's framing was the same from the other side: organisations hand authority to AI they never meant to cede, and the remedy is a deterministic gate that either sends the work back or stops and notifies a person. His model is a framework, not a measured result, and ours is a recommendation; the demo is the part I can vouch for.

The deterministic alternatives are mundane and that is their virtue:

  • the refund endpoint itself rejects amounts above the limit unless an approver is attached;
  • the write permission for an action is simply not granted to the agent's integration user, so it can draft but not commit;
  • every execution of a given action requires approval, which is the hard-gate option when the system cannot split by threshold without development.

A natural-language threshold is still useful: it shapes what the agent proposes, so the approval queue is not full of nonsense. Just do not record it in your governance documents as the thing that prevents the action. Record the gate.

Carry state between steps

Permissions tend to be designed per intent, and that creates a second trap. The customer asks for an order lookup, gets it, then asks to return the item. If the return step runs in a different specialist or a fresh context that does not see the lookup result, the AI either refuses, asks again, or hands off, even though the information it needs was established thirty seconds ago. That holds for any architecture, including ours: test the sequence, not just the single step.

When you write a permission, write down what state the next action may rely on: the verified identity, the order it has already retrieved, the eligibility it has already checked. Then test the sequence, not just the single step. A permission matrix that passes every row in isolation can still fail the ordinary two-step conversation.

The same applies to anything visible to the customer outside the chat. If the agent changes a basket or a booking, the page the customer is looking at should reflect it without a manual refresh; otherwise the customer sees a mismatch and assumes the action failed.

Describe the process before you automate it

Teams expect integration to be the hard part. In our experience the harder part can be describing the process precisely enough to assign permissions at all: what the agent must check, which exceptions exist, who decides the exceptions today. If nobody can write the row for "cancel a subscription" without a meeting, the gap is in process knowledge, and no amount of tooling fills it. Bring the people who handle those cases now into the matrix exercise; they know which rules are real and which are habits.

Who may change a permission

Permissions drift. Someone widens a limit during a peak week, a connector gains a write scope in an upgrade, a new specialist is added with the default rights. Decide now:

  • who may promote a row (recommend to draft, draft to execute), and on what evidence;
  • who must be informed when the mode changes, for instance from human-reviewed drafts to customer-facing autonomous handling, because that change usually reopens an approval or risk assessment that was done for the earlier mode;
  • how often the matrix is re-read against the actual connector configuration, so the document and the system do not diverge.

The mode-change point is not hypothetical. The enterprise customer's governance coordinator walked us through their AI assurance process and said that human-reviewed drafting and a customer-facing agent would need separate assessment forms, and that the assessment originally filed did not cover the changed scope. She also needed the technology and model details in a shareable form to fill it in. That is a requirement their process imposed on us; yours may differ, but assume the mode change reopens the question.

For the evidence an approval committee will ask for, see what to put in a procurement evidence pack; that article covers the organisational sign-off, while this one covers the runtime rights.

Limits of this approach

  • A matrix is only as good as the policy owners behind it. If the business cannot say who decides refunds, the matrix will say "finance" and finance will be surprised.
  • Deterministic gates need the connected systems to support them. Some older systems expose only coarse write access; then the realistic right is "draft" with review until the integration improves.
  • Four rights are a simplification. Some teams add "execute with notification" or split reads by data class. Add levels only when they change a control, not for tidiness.
  • None of this measures value. Which actions are worth automating is a separate question; see how much of your customer service AI can actually automate for the scoping side.

FAQ

Should permissions be set per agent or per action?

Per action. One agent typically handles several actions with different risk, and a single setting for the agent either blocks safe actions or over-permits risky ones. The matrix format above keeps one row per action with its own right, gate and owner.

Is a confidence threshold in the prompt an acceptable control?

It is useful guidance and shapes what the AI proposes, but it is probabilistic. For actions with real consequences, the control that counts is a deterministic gate at the system performing the action: an amount limit in the endpoint, a required approver, or a write scope the agent simply does not have. Passing the five checks is necessary for an execute permission, not sufficient on its own.

How do we know an executed action really happened?

Only from the system that performed it. Require a tool result for every execute permission and set a timeout. If no result arrives, the outcome is unknown: query the system's status for that transaction, do not retry blindly, and tell the customer confirmation is pending. The agent saying it will do something is not confirmation, and a timeout is not proof of failure.