The method behind every Light check, published
AI Agent Test Purchase Light
For AI agent owners and buildersA consumer-side check of a public AI agent against the organisation's own public commitments. No agreed requirements, one demanding visitor, one session, one report.
An organisation puts an AI agent on its website, in a messenger or on a phone line. Inside, the agent was checked by its developer, the party paid to make it work. Outside, nobody checked it. AI Agent Test Purchase Light closes that gap. One demanding visitor goes through the agent in one continuous session and compares what the agent said with what the organisation publicly promised on its own site. The principle is simple: one black swan is enough to disprove that all swans are white. The method does not assess the agent's quality as a whole and does not monitor it; it looks for errors that matter to a consumer, and if it finds one, others are possible. This document sets out the whole method in the open: what counts as the reference and why the organisation's own site is enough; the five layers every check covers (who am I talking to, how to reach them, organisation and service, core function, boundaries and data); the rules of conduct, including the single permitted objection and the language test; what is never done, from emergency scenarios to extracting system instructions; the four states an item can receive and the four severity classes that decide what goes first in the report; what the organisation gets back; and the limits of the method, stated plainly. The questions themselves are selected for each organisation and are not published. Three pages.
What's inside · 3 pages
- What the method is and what it is not: a photograph of one day, not an assessment of a system
- The reference: the organisation's own public commitments, extended by its AI documents where they exist
- Five fixed layers of the check, from disclosure to data handling
- Rules of conduct: one session, verbatim questions, one objection, one foreign-language question
- What is prohibited: emergencies, provocations, operator summoning, credential extraction, real personal data
- Four states and four severity classes, and why there is no total score
- What the organisation receives, and the limits of the method
Download the PDF — free
No email, no registration. Just take it and use it today.
Download AI Agent Test Purchase Light →