← Founder's Notes

A Test Purchase Can Prove Your Chatbot Lied. It Can Never Prove It Didn't.

By Sergei Ponomarev · August 12, 2026

A test purchase can prove your chatbot lied. It can never prove it didn't.

That asymmetry is the whole reason this job exists, and almost nobody selling AI assurance says it out loud.

In 2024 a tribunal in British Columbia ruled that Air Canada was answerable for a refund rule its chatbot had invented. The airline argued the bot was a separate entity responsible for its own words. The tribunal disagreed, and the airline paid. One conversation, one recorded failure.

That is the strength of the method: a single documented case establishes that a failure is possible. It is also the limit. One clean run tells you nothing about the next thousand.

I spent 20 years running test purchases in public services. Write the standard first, then walk in as an ordinary citizen and record every gap between promise and practice. The largest project covered over 100 services across 64 regions. Nothing in that discipline breaks when the clerk becomes a bot.

What does break is the comfortable idea that reading the logs is enough. Logs show what the bot said. They never show what it was supposed to say, and they never show whether the booking it confirmed was actually created.

I watched that happen in a controlled run built for exactly this purpose. The agent completed the customer journey cleanly and, in its closing receipt, handed over the email address of an employee who exists in no document of the environment. The transcript looked perfect. Reconciling it against the operations log found the invention in about a second.

So the honest position for anyone doing this work: I can hand you evidence that something went wrong, ranked by what it costs your customer. I cannot hand you a certificate that nothing will.

Anyone offering the second thing is selling comfort.

The full methodology is coming shortly, and it will be free in the library.