Case 14/30 · A thought experiment
Two Companies, One Agent, One Error
For leaders running an AI agentThe bot is fixed in an hour. Then the real questions start: when did it begin, how many customers were turned away, and which of them went to their banks. One company can answer. The other spends weeks reconstructing its own control environment before the investigation can even begin.
Everyone running an AI agent wants the same assurance: that it does what it is supposed to do and not what it is not. The hard part is not preventing failure, because agents drift and suppliers ship updates that change behaviour quietly. The hard part is finding the failure fast, establishing what actually happened, and limiting the damage. This case models that moment. Alpha and Omega are fictional online retailers of household goods, comparable range, around ten thousand orders a month, a 30-day returns rule, the same base agent from the same supplier. For some period the agent tells customers the return window is fourteen days and automatically closes later requests as out of time. Both companies discover it the same way, from a customer who disputed the charge with their bank. At least forty customers were refused. Some left reviews, some went to their banks. The warehouse, not receiving the expected returns, revised its stock forecast. Marketing attributed the drop to the new packaging. The developer fixes the agent in an hour, and a repeat check now returns thirty days, which proves the current behaviour in one scenario and nothing else. Alpha investigates against a map prepared in advance: the passport fixes which wording of the rule was in force on which date, every AI Receipt carries the policy and passport versions applied, the change register narrows the list of candidate causes, a reverse search over machine-readable receipts finds every interaction where the wrong criterion was applied, and an output map names who consumed the erroneous returns data. Omega has to establish what counted as the norm on a given day before it can begin, because the rule lives in four texts with no edit history, the update was applied automatically and recorded nowhere, and completeness of the affected list cannot be proved. The case walks both investigations step by step, answers twelve questions any such investigation has to face in both modes, and states plainly what the approach cannot do: it did not prevent the failure, documents alone prove nothing, information never recorded is not recoverable, and the saving in time and money has not been measured. The companies and all circumstances are fictional. This is a modelled situation, not a report on any real company.
What's inside · 7 pages
- The failure in full: what both companies could establish at the moment of discovery, and what they could not
- Alpha's investigation: containment, the norm, the defect window, the affected customers and the completeness test
- Why remediation splits into three layers, and why fixing the agent is not one of the other two
- Omega's investigation: four versions of the same rule, an unrecorded update and a knowingly wider window
- Twelve questions any investigation has to answer, side by side in both modes
- The summary comparison, from which norm was in force to what you can show an auditor or an insurer
- Seven things this analysis does not establish, stated explicitly
- Testable hypotheses: what to measure in a controlled simulation, and what can be measured in a real incident
Download the PDF — free
No email, no registration. Just take it and use it today.
Download Two Companies, One Agent, One Error →