Your AI agent has been refusing refunds for weeks. You found out from a bank chargeback.
The developer fixes it in an hour. Then the real questions start.
When did it begin?
How many customers were turned away?
Which of them went to their banks?
And the returns statistics that were wrong the whole time: the warehouse rebuilt its forecast on them, marketing decided the new packaging was a hit.
The bot is fixed. Nothing else is.
I have been arguing for a while for AI receipts. More precisely, for a full system of four parts that reinforce one another: a public AI policy, an AI service passport, an AI receipt for every interaction, and the rules that hold them together.
So I ran a thought experiment.
Two identical online retailers, the same agent, the same failure, the same day. One has the system implemented. The other has the usual set: terms on the website, a system prompt, a dialogue log.
To be plain: the system does not prevent the failure. Both companies get hit. Agents make mistakes and updates change behaviour, and no document stops that.
What it changes is everything after.
One company investigates against a map prepared in advance: which rule was in force on which day, which interactions applied the wrong one, who received the wrong output, who decides when to resume. The other spends weeks reconstructing its own control environment before it can even start.
The full case is published in the library: the investigation step by step at both companies, twelve questions any such investigation has to answer, and an honest list of what the exercise does not prove.
Read the full case in the library
Questions are welcome. Write to me and I will walk you through the details.
