← Library

The method, published in full

AI Agent Test Purchase

For anyone checking an AI agent

A method for evaluating AI services through the customer's eyes — carried over from twenty years of test purchases in public services.

This is the author's method of Sergei Ponomarev, PhD, set out in full. A test purchase is an old discipline: write down what a service is supposed to do, then go through it as an ordinary customer and record every gap between the promise and the practice. This guide carries that method over to AI agents, where the human shopper becomes an AI shopper, the clerk becomes the company's agent, and the analysis becomes a reconciliation between what the agent said and what the system actually recorded. It sets out the six steps of a check — standard, script, run, reconciliation, report, fix and re-test — and the three documents that serve as the norm: a public AI Policy, an AI Service Passport per service, and an AI Receipt per interaction. It covers the five axes along which a check is designed (external or internal, manual or automated, partial or full, comprehensive or targeted, standard or specialised), how evidence is captured so the results can be believed, and how findings are scored. It is equally clear about the boundaries: one run documents a failure but never establishes how often it happens, an external check answers what the customer got rather than why, and the analyst is an AI too and can be wrong. The method is carried over from twenty years of the author's own work on service quality: nationwide monitoring of public services, hundreds of independent assessments and test purchases. The reference pilot, run jointly with the Swiss AI metrology company NeoMundi, is published in full as open source, code, prompts, transcripts and defects included.

What's inside · 35 pages

  • Why reading chat logs proves almost nothing, and what does prove a bot broke a promise
  • The three-document standard: a public AI Policy, an AI Service Passport, and an AI Receipt
  • External and internal test purchases: how to check any bot, with or without access
  • The five axes of designing a check, and how to choose the right combination
  • How one bot tests another: the AI shopper, the AI analyst, and where the human methodologist stays irreplaceable
  • Recording and validation: what counts as evidence, and how to capture it before the website changes
  • Real cases: an airline bot that closed every route to a human, and a tribunal that made a company answer for its chatbot's invented rule
  • Honest limits: why a single test purchase proves guilt but never innocence, and what that means for monitoring
  • Ethics and law, the budget, partners, and a readiness checklist to work through before you start

Download the PDF — free

No email, no registration. Just take it and use it today.

Download AI Agent Test Purchase