← Sergei's Notes

AI Service Quality Is the Next Bottleneck

By Sergei Ponomarev · July 27, 2026

Everyone I talk to is building something. A support bot, a sales assistant, an agent that books appointments or quotes prices. The tools are cheap, the demos are dazzling, and a working prototype now takes a weekend. Building has stopped being the hard part.

Here is what nobody is doing: checking whether any of it actually works for the customer.

I don't mean "works" in the engineering sense — the model responds, the integration holds, the dashboard shows green. I mean works the way a customer defines it: I asked a question and got the right answer, I was told the truth about what this thing can and cannot do, and when it failed, a human picked up the pieces without making me beg. By that definition, a very large share of deployed customer-facing AI simply doesn't work. MIT's much-quoted finding — 95% of generative AI pilots producing zero measurable results — is not a story about weak models. The models were fine. It's a story about services that nobody tested from the outside before pointing them at paying customers.

I spent about twenty years on exactly this problem, before AI made it fashionable. My background is quality assessment of public services: nationwide monitoring programs, independent evaluation, mystery shopping. The method was almost insultingly simple. You don't ask the office whether it serves people well — every office believes it does. You walk in as an ordinary person, ask the ordinary questions, and write down what actually happens. The gap between what organizations report about themselves and what a mystery shopper experiences was reliably enormous. It paid my salary for two decades because it never closed.

Now replace the office clerk with a chatbot, and notice what changed: nothing. The company still believes its service is excellent. The dashboard still says response time is three seconds. And the customer still walks away having spent forty minutes in a loop, promised a refund the bot had no authority to issue. The gap survived the technology transition perfectly intact. What didn't survive is the accountability: a human clerk who lies to a customer gets a manager called on them; a bot that lies generates a ticket that averages out in the metrics.

This is why I think the next real AI business opportunity sits on the verification side, not the building side. It's the classic gold rush logic — when everyone digs, sell shovels — except here the shovel is trust. Consider what is quietly assembling around this: the EU AI Act's transparency rules took effect in August, with fines that get a CFO's attention. US states are stacking up their own disclosure laws. Insurers are starting to price AI incidents. Procurement departments are beginning to ask vendors, in writing, "how do we know your agent won't promise our customers things we can't deliver?" Every one of those pressures converts, eventually, into a purchase order for somebody who can independently answer the question: does this AI service actually do what the company claims?

Answering it properly takes three things, and none of them is a benchmark. First, the claims themselves have to exist in checkable form — what the bot is allowed to do, what it must refuse, when a human takes over. If that isn't written down anywhere, the test has nothing to test against; you can't audit a service against a vibe. Second, somebody has to walk in the customer's door: real conversations, awkward questions, the stubborn follow-up a polite demo never includes — mystery shopping, updated for machines. Third, the results have to leave a record the customer's side can keep. A receipt, essentially: what was asked, what was promised, what actually happened.

Notice that none of this requires a data science degree. It requires the discipline of an auditor and the mindset of a demanding customer — a profession that quality-of-service people have been practicing for decades on humans. The skills transfer almost one to one. I'd go further: people who spent careers in service standards, consumer protection, and independent assessment are sitting on exactly the expertise this market is about to need, and most of them don't know it yet.

For business owners, the practical takeaway is less romantic. Before your AI faces a customer, someone who doesn't report to you should try to break it the way a customer accidentally would. Not a red-team of prompt engineers — a patient stranger with a real problem and no goodwill. What that exercise finds in an afternoon is routinely the difference between "our AI saves us money" and "our AI is generating complaints we haven't read yet."

The bots will keep multiplying either way; that race is over and everyone won. The scarce thing now is knowing which of them can be trusted with your customers — and being able to prove it. That's the bottleneck. And in business, the bottleneck is always where the money pools next.