Back
Blog / 
Implementation

The evidence pack that gets customer-service AI approved

written by:
David Eberle
One bound folder with five coloured tabs on a round table with five empty chairs.

A customer-service AI deployment is approved, in practice, by four or five people who each need a different thing: the security assessor needs to know where data goes, the IT owner needs to know what touches the stack, procurement needs the scope to match the contract, the service owner needs to know who is accountable when it is wrong, and someone in legal or risk needs to see that the decision was made deliberately. When those needs arrive one at a time, each answer waits for the next question. One evidence pack lets the same people assess the same scope together and cuts the avoidable back-and-forth.

This guide describes that pack. It is about organisational approval before deployment, which is a different question from what the AI may do at runtime; the runtime side is covered in what a customer-service AI should be allowed to do. It is also not a legal compliance checklist; the archive has separate pieces on the EU AI Act and GDPR, and nothing here is legal advice.

Where this comes from. The pack is assembled from five conversations Typewise held in 2026 with teams taking customer-service AI through security, governance or procurement: two security-assessment sessions with a large enterprise customer, an implementation call with a retailer whose data policy restricts cloud processing, a discovery conversation with a public-sector organisation subject to EU procurement rules, and a monthly review my colleague held with a retail customer whose pilot had been stopped for formal procurement. I was in two. What the approvers asked for is reported; the structure of the pack is our recommendation.

Two gates, not one

Technical readiness and organisational permission are separate gates, and they are easily confused. A solution can be built, tested and working in simulation and still not be allowed to go live, because the internal implementation of regulatory or policy requirements has not been completed. In regulated sectors the governance gate can take longer than the technical one. The public-sector organisation's AI lead told my colleague their first use cases were built and essentially ready on the platform they already had; the limit was regulation and its internal implementation, not the technology. He also said that, under EU tendering, a better feature elsewhere was nowhere near enough to justify adding a second supplier. Plan for both gates from the start.

The pack below is for the second gate.

What the pack contains

1. The actual data flow

A diagram and a table, not a paragraph. For each step of the workflow: which data leaves which system, where it is processed, whether it is anonymised and at which point, what is stored, for how long, and where. Security teams will ask specifically whether personal data reaches the vendor before anonymisation; if it does, say so and say what controls apply. The enterprise customer's security reviewer asked exactly where in the flow anonymisation happens, in a walkthrough I attended, and the answer shapes how the assessment proceeds. If the design redacts identifiers upstream, explain how identifiers needed in the reply are restored, because a redaction service that strips the order number also strips the ability to answer about the order. The retailer's integration team wanted to run anonymisation on their side, because anything touching the cloud needed a special request from above; our engineer's caveat was that identifiers stripped upstream have to be restored in the reply, and the customer's own idea was that many of their replies might not need personal details at all. They asked for small diagrams of both options for their security team, which is why this section asks for a diagram. Both remained options for that team to judge; neither was an approved architecture at that point. Sometimes the simplest design is to confirm that certain replies do not need personal details at all.

Be precise about what the vendor's cloud provider's certification covers. An infrastructure report covers infrastructure controls; application-level controls such as identity, access and logging are a separate set of questions, and the assessor will ask for a demonstration (a real single sign-on and multi-factor login, a joiner-and-leaver record) rather than a certificate. The enterprise reviewer said as much: the cloud provider's report covered infrastructure controls only, and application and identity questions remained open.

2. Read and write scope, per workflow

List every system the AI reads from and every one it can write to, per workflow, with the scope of each. "Reads order status for an authenticated customer" and "creates a return label for eligible orders" are scopes; "integrates with the shop" is not. Note which actions are human-reviewed drafts and which are autonomous, because the approval questions differ and an assessment done for the reviewed mode does not automatically cover the autonomous one. If the scope later moves from assisting employees to acting towards customers, expect to reopen the assessment: retention, anonymisation and risk differ between the two modes, and one assessment may not cover both. The enterprise customer's coordinator asked in the same walkthrough whether one assessment could cover both the reviewed-draft assistant and the proposed autonomous agent, given that everything in the first was anonymised; the two were treated as separate scopes.

3. Documentation the assessor can use

The coordinator filling in an internal assessment form needs technology and model details in a shareable form: what models are used, where they run, what is retained, what the sub-processors are. Provide the questionnaire answers as documents, not as a call. In one of the assessment sessions the governance tool was unavailable, so the assessors logged findings in an offline sheet and asked for business justification and action plans by email. Teams in that position still need a written confirmation trail; make the written trail the default.

4. Exceptions, decided and owned

A security review produces findings. Each needs an explicit decision. The assessors' form offered two responses per finding, accept as a risk with a justification or accept with remediation and a date, and some findings fell outside our responsibility or overlapped with others. We list more options than the form did, because in our view the options are wider than "accept": remediate before go-live with a dated plan and an owner; accept the residual risk, which only the authorised risk owner may do, with a written justification, and only where no mandatory requirement forbids it; remove the affected scope from the deployment; defer the deployment until the finding is resolved; or reject the deployment. Each decision needs a named service owner and a named security owner. Some findings will be outside the vendor's responsibility and some controls will overlap; the pack should show the decision for each, not a promise to fix everything. A control score without an accountable decision is not an approval, and a recorded decision or a remediation plan is not an authorisation to go live either: go-live is a separate sign-off by the people named in the next section, taken once the mandatory findings are closed.

5. Named sign-offs

One line per role: who approves the business value, who approves the fit with the technology strategy, who approves the security exceptions, who owns the contract scope. In many organisations the business sponsor holds the budget but the technology owner must agree before anything enters the stack; neither can decide alone. If employee representatives must be consulted, the pack also needs the metric definitions they will ask for, in writing: what is measured about individual users, how sessions and timeouts are defined, whether any users are excluded from reporting and why. A dashboard introduction is not a definition.

6. A phased evaluation with stop criteria

The approval is easier to give when it is bounded. Describe the phases (offline testing, supervised live handling, limited autonomous handling), the scope of each, the evidence that permits moving between them, and what stops the deployment. The go-live guide covers the design; the pack needs the summary with dates and owners.

Procurement scope: agree it before the pilot

Two patterns cause avoidable friction.

The first is the pilot that outgrows its approval. The retail customer's service manager told my colleague the planned quick pilot had been stopped because the spend it implied for the following year required formal procurement, that the tender scope had grown to a full ticketing tool, and that an attempt to use the pilot as evidence had been stopped as well. Agree the eventual scope and the procurement path before investing in a quick pilot, even if the pilot stays small.

The second is the incumbent problem in regulated procurement. Where a supplier is already approved and a public or regulated tender would be needed to add another, a better feature is rarely enough to justify the switching cost of the process itself. If you are the buyer, be honest with yourself about this early; if you are evaluating alternatives, ask what would make a second supplier worth the process.

A pack checklist

  • [ ] Data-flow diagram and table: what leaves, where it is processed, anonymisation point, retention, location.
  • [ ] Cloud-provider versus application-level controls separated; demonstrations arranged.
  • [ ] Read and write scope per workflow; reviewed versus autonomous mode marked.
  • [ ] Trigger for re-assessment when the mode or scope changes.
  • [ ] Technology and model documentation in shareable form; sub-processors listed.
  • [ ] Every finding with an explicit decision (remediate, accept residual risk by the authorised owner within mandatory requirements, remove scope, defer, or reject), with owner and date.
  • [ ] Go-live sign-off recorded separately from the findings decisions, after mandatory findings are closed.
  • [ ] Named sign-offs: business value, technology fit, security exceptions, contract scope.
  • [ ] Metric and timeout definitions in writing for employee representatives where required.
  • [ ] Phased evaluation with scope, evidence for progression and stop criteria.
  • [ ] Procurement path agreed for the eventual scope, not only the pilot.

Limits

The pack describes what approvers commonly ask for in our experience; your organisation's process will add items and may order them differently. It does not establish compliance with any regulation, and the opinions of customers or vendors about what regulators require are not legal advice. Confidential findings from any specific assessment belong in that organisation's records, not in a public article, which is why this guide stays at the level of structure.

FAQ

Our vendor's cloud provider is certified. Is that enough for security approval?

Treat it as insufficient on its own. Infrastructure certification covers infrastructure controls. Assessors will ask separately about application-level controls such as identity, access, logging and anonymisation, and will often want a demonstration rather than a document.

We were approved for an employee-assist use case. Does that cover a customer-facing agent?

Treat it as a new assessment unless your security team confirms otherwise. Retention, anonymisation and risk differ between a reviewed-draft mode and autonomous customer-facing handling.

What do we do with security findings we cannot fix?

Decide each one explicitly: remediate with a dated plan, accept the residual risk (only the authorised risk owner may, and only within mandatory requirements), remove the affected scope, defer the deployment, or reject it, each with a named service owner and security owner. An undecided finding blocks approval. A decided finding is a prerequisite for approval, not the approval itself; go-live still needs the named sign-offs once mandatory findings are closed.