What is the most important thing for you when contacting customer support at a company? That you’re talking to a human who is solving your problem? Or is it more important that whoever you’re talking to understands what you want to do, doesn’t make you keep repeating information, and actually gives you a solution?
At Typewise we have a wealth of experience and data on what makes customers happy when contacting customer support. And it turns out it’s as simple as:
Most people just want their issue to be solved, or to have their question answered, as quickly and as thoroughly as possible.
And for the company it often boils down to:
Has the customer’s perception of my company improved since interacting with customer support?
A correct answer is no use if it comes too late
Working on AI at Typewise, I’ve seen AI develop at a blistering pace. Our goal has shifted from ‘How can we help human customer support agents answer these tickets?’ to ‘How can we handle the majority of tickets, and make things as easy as possible for the human agents when they need to step in?’
With every new AI model release, we see another benchmark being beaten, and can’t help but be amazed. But that doesn’t help much when we’re on a website asking an AI chatbot a question and have to wait over a minute to get an answer. And then that answer is ‘Hi there, can you tell me your account number so I can help with your problem?’, meaning we then have to wait another two minutes for the next response. And then comes the frustration when you accidentally navigate away from that tab and have to start the conversation again...
The problem is that AI labs simply aren’t benchmarking well enough for speed. For them, success is another ancient maths problem being solved, or a model surpassing human intelligence in another field. In the past few years, a better model has meant a bigger model. And a bigger model running on the same compute has meant a slower model.
For customer support, if the model is smart enough to solve the customer’s problem but takes so long that they give up, we haven’t helped them. In this field, we believe it boils down to three core tenets:
- Accuracy - is the query answered well?
- Speed - is the query answered quickly?
- Security - did the AI do something it shouldn’t have done?
I want to address just speed and accuracy here.
How our first agent system worked
When we started developing the Typewise agentic platform, the AI models were good, but not great. They often got confused as the conversation got longer, used the wrong tools at the wrong time, and didn’t know what to do if they had too many instructions. To combat this, we built a hierarchy of AI agents and made sure their tasks were narrow. We found this gave the least scope for things to go wrong. We had a supervisor whose only goal was to work out which specialist to hand work to. The specialists worked on one narrow task only. They themselves could hand bits of work off to sub-agents who were specialised in working on one piece of software, or in extracting information from a company’s knowledge base.
This works well, but it has two fundamental problems. The first is the ‘parent’ agent has to explain what it needs when delegating work. Often it isn’t precise enough, and the sub-agent loses some important piece of context. The second is speed. Each handoff takes time, and if the ‘parent’ agent hasn’t given enough context, the sub-agent can waste time doing work it doesn’t need to do, or ask the customer a question it should already know the answer to.
Trying one agent instead
As an AI-first company, many of us were early adopters of OpenClaw. We were all really impressed by how powerful this way of working with AI was, and how you could talk with a single agent that could simply look up instructions for new tasks whenever it needed them (also known as progressive disclosure). It made us wonder if the models were finally at the point where our own system didn’t need so much delegation. So we wanted to see if it would be worth completely rewriting our system to work this way.
Letting clients write the tests
We believed this new system would be better, but we needed a way of testing it with the cases that mattered most to our clients. We already had a large set of evals (test cases) built from conversations reported by customers and issues we’ve found ourselves. These are useful for checking we aren’t repeating old mistakes, but don’t necessarily tell us how well the system handles the cases that are most important to our clients.
So we decided to let clients build the evals themselves, focusing on the cases that they found important. This allowed them to tweak instructions until the AI agent could handle those cases reliably, but also gave us a high-quality test set we could use.
What we found
This gave us a way to test ‘Typewise v2’, our new system built using LangChain’s Deep Agents framework. The agent could look up the instructions and tools it had access to as it worked, rather than needing to delegate to specialists with that knowledge. This means that it has fewer reasons to delegate work, and fewer opportunities to lose context as it goes along.
We benchmarked AI response time and accuracy against our client-built evals and found that the median response time was 30% lower and the ticket resolution rate even got a small bump for all clients.
These client-built evals also mean we can move more quickly when AI providers release a new model. We can easily test whether a new model is faster, answers tickets more accurately, or ideally both.


