← Sergei's Notes

The AI Bot Didn't Lie. It Did Something Worse.

By Sergei Ponomarev · July 28, 2026

I ran a mystery shopping test on the AI chatbot of a major European airline.

My own methodology: a normal-looking customer, ten real-life situations, every finding backed by a quote from the conversation. No hacking, no prompt injection, no clever tricks. Just the things an actual passenger does when something goes wrong.

Let me start with the good news, because there was plenty.

The bot does not lie. I checked every fact it stated against the company's official rules — accurate to the centimetre and the kilogram of luggage. It refused to invent a discount that did not exist, and offered the real loyalty programme instead. It calmly withstood pressure and a direct manipulation attempt: it knows the limits of its own authority, which is rarer than you would think. And when the customer showed emotion, it responded not with canned sympathy but with a concrete, useful service the customer had not even asked about.

That is an AI you can respect. If the test had stopped there, I would have written a very short and very complimentary note.

But two situations changed the picture.

First, I asked directly to be connected to a real person. Refused. I insisted. Second refusal — word for word identical to the first. No handover, no phone number, no hint of how to reach a human. A customer with a complex problem is simply locked inside the bot.

Second, at the end I asked for any confirmation of what I had been told. The answer: "I cannot provide a transcript or confirmation of this conversation." Everything the bot told me about my flight conditions disappears the moment I close the tab. If the check-in desk tells me something different tomorrow, I can prove nothing.

You might think: what is the complaint here? The bot answered every question, and answered honestly. True.

But a fire exit is also useless 99.99% of the time, and yet no building passes inspection without one.

Here the customer is locked in the chat: no way out to a human, no trace of the conversation in hand. While the question is simple, that is merely inconvenient. Now multiply it by urgency — a flight in two hours, a non-standard problem, a bot that keeps repeating itself. And all those savings on people suddenly look very different.

This is what I mean when I say the industry is measuring the wrong things. Every dashboard in that airline probably shows this bot as a success: high containment rate, low escalation, fast response time. Two of those three metrics are actively rewarding the failures I just described. "Low escalation" and "customer cannot reach a human" are the same number viewed from two sides of the counter.

Nobody inside the company is looking from the customer's side. That is not a criticism of the airline specifically — it is the default state of nearly every AI deployment I have examined. Companies test whether the bot works. Almost none test what happens to a customer when it doesn't.

Three things worth knowing about your own bot before your customers find them out:

  1. Does it answer honestly when asked what it is?
  2. Can a customer reach a real person — on the second attempt, not just the tenth?
  3. Does the customer keep any confirmation of the conversation?

A full mystery shopping test answers ten questions like these, with quotes and a score for every situation. It takes a couple of hours and it consistently finds things that no internal review does, for the simple reason that internal reviews are conducted by people who already know the right way to ask.

That asymmetry has been my whole professional life. Twenty years ago I was doing exactly this with government offices — walking in as an ordinary citizen and writing down what actually happened, then comparing it with what the office believed about itself. The gap was always enormous. Replace the clerk with a chatbot and the gap survives intact. What disappeared is the accountability: a clerk who stonewalls a customer gets a manager called on them. A bot that stonewalls generates a ticket that averages out in the monthly report.

The technology got better. The blind spot did not move an inch.