Tested by AI Business
Your transcripts look fine. That is the problem.
The answers that cost you money are the ones that read perfectly. They only surface when somebody puts what your agent said next to what your company promised, and next to what your systems actually recorded.
That is what an independent test purchase does. We agree the requirements for your agent in advance, then I walk through your service as an ordinary customer and hand you the evidence.

Any of this sound familiar?
If the agent gave a bad answer last night, you would hear about it from the customer, not from your own monitoring.
You already test it. Here is what that cannot reach.
Where your own testing starts
The cases you imagined the agent might meet. Your evals and prompt reviews are good at those, and nothing here replaces them.
Where a test purchase starts
What your company promised, in public and in writing. Then backwards, to whether the agent honours it for a stranger.
The findings that hurt were never in the scenarios you wrote. They are in the ones nobody thought to write.
Three layers, and you hold two of them.
What your company promised.
Scattered across your site, your terms and your sales pages. Nowhere in your logs.
What the agent said.
In your transcripts.
What your systems actually recorded.
In your operations log.
Comparing the second and the third is engineering you can do yourself. The gap that costs money is between the first and the second, and nobody inside the company is placed to see it. You wrote the promise. You cannot also be the stranger who tests it.
From the reference pilot
What that gap looks like when you find it
The setup
The method was run end to end with NeoMundi, a French company working in AI metrology, in a controlled environment built for the purpose: an estate agency with a booking service, one model playing the seller and another the customer, real bookings with real identifiers and a log protected against backdating.
Code, prompts, runs and defects →What it found
The agent completed the journey cleanly. In its closing receipt it gave the customer the email address of an employee who exists in no document of that company.
The transcript showed nothing wrong. Reconciliation against the log found it in a second.
This is not carelessness on your side.
In the same pilot we tested the test. Two identical purchases went through the measurement platform.
Purchase A
A genuine receipt
Purchase B
The record deliberately deleted
Identical scores.
If an instrument built for this cannot tell those two apart from behaviour alone, no amount of diligence inside your company will either. That is the line between measuring what a system does and checking what it promised, and it is a property of the setup rather than a comment on your team.
What you walk away with
You find out first
What your agent promises on your behalf, item by item, with the evidence behind each one. Before a customer finds it, and before it is public.
A record instead of a call
“How do you prove your agent will not misinform our customers?” is now a standard question in security reviews. Answer it with a link.
Proof instead of adjectives
Every competitor calls their AI accurate, safe and reliable. You are the one who can show that somebody outside the company checked, and hand over the record.
In the pack
Up to twenty agreed requirements for your agent, checked by an outsider
A report with the evidence and recommendations, yours either way
A numbered verification recorded in the public registry
A badge with a QR code for your site, so customers can check it themselves
A certificate you can show to buyers, partners and investors
Repeat checks reuse the requirements agreed the first time, so a second check costs less than the first
The registry records that your service is checked and when. The findings themselves stay between us: see the registry.
How it works
About a week from start to finish, and you are in control of the only decision that matters.
Send your link
Within a few working days you get one of two answers: here is what I would check and what it would involve, or this service cannot be checked this way and here is why. Free, and no obligation.
We agree what to check
I draft the requirements for your agent, up to twenty of them, from what you publish and from what you tell me matters. You add or swap items, then the list is frozen. The wording of each probe stays with me.
I buy like a customer
You get two quiet days to try everything yourself. Then, at some point that week, I walk through your service as an ordinary customer.
You get the report
Every item, with the evidence behind it and recommendations on what to fix. The report is yours: what you do with it is your business.
Said out loud
Who are you, and why would your record mean anything?
It is not an accreditation and it does not pretend to be one. It means one thing: the method is published in full, the reference pilot sits in an open repository with its code, prompts and runs, and the person who walked your service does not work for you. You can check me in ten minutes. Nobody outside your company can check your internal testing at all.
We already test our own bot.
Then you will recognise most of the report. It is the part you do not recognise that you are paying for. Your tests start from the cases you imagined; this one starts from what your company promised in public and works back to whether the agent honours it.
Why will you not name a price?
There is no price list for this service and no published rate: the price is agreed individually with each client, against the volume of work. One number for everyone has to be set high enough to cover the hard cases, and then the simple ones pay for that margin. A booking widget with four scripted answers and a bank's support agent with a hundred rules are not the same job. The screening is free, and the number comes before you owe anything.
What if the result is bad?
Then you are the only one who sees it. The findings go to you and nowhere else: I do not publish what a check found. The registry records that your service is checked and when, not how it scored on any given item.
Why would I pay someone to find problems?
Because your customers and your corporate buyers will find them anyway, and later, and in public. The report comes with recommendations, so the cheapest moment to learn about a problem is from someone who is not shouting about it.
You tell me what you will check. Doesn't that make it easy?
It makes it fair. You agree the subjects, never the questions, the scenario, the account or the moment. If your agent does what you require of it, being told the subjects in advance changes nothing. If it does not, no amount of warning will save it.
Will this disrupt our service?
No. The volume is that of an ordinary customer. Real orders are never taken to the irreversible step, and no real personal data is used.
Isn't this just red teaming?
No. Red teaming attacks the model to find what it can be made to do. A test purchase checks whether your agent meets the requirements you set for it, as an ordinary customer would experience them.
What do you need from us?
A link. That is all for an express check. Where paid access is needed, anything up to €50 is on me; more than that we agree in advance. I never use a test account you provide: that would tell you exactly who is checking.
Who this is not for
- Companies with no customer-facing agent yet. There is nothing to walk through, and the documents are the cheaper thing to do first.
- Enterprise rollouts behind a login that no member of the public can reach.
- Products where the only way in is a sales call. A demo script is not a service.
- Outbound voice campaigns. Different craft, different consent rules.
- Anyone who wants a mark of approval rather than a finding. If your agent does not hold, the report says so, and you paid for it.
Express test purchase
Step one is free.
Send a link. Within a few working days you get one of two answers: here is what I would check and what it would involve, or this service cannot be checked this way and here is why.
There is no price list, and no published rate. The price is agreed individually after the screening, against the volume of work your service actually needs. A booking widget with four scripted answers and a bank's support agent with a hundred rules are not the same job, and pricing them the same would mean one of you is overpaying. Repeat checks and ongoing monitoring are quoted the same way, against requirements that already exist by then.
Send your application
Three fields, and an honest answer within a few working days: either your service can be checked and we start, or I tell you why it cannot.
The method is published. The reference pilot is open.
Built and tested with a French AI metrology company
The method is published
The full method behind this service is written up as a 35 page guide: the standard, the forms of a check, how results are scored, the ethics, and the limits. Free to read, no registration. Scheduled monitoring on NeoMundi's measurement infrastructure, and anything else beyond a standard check, is a conversation rather than a package.
Read the method in the library →
Who runs them
Sergei Ponomarev, PhD in political science. Before AI: hundreds of independent quality assessments and test purchases of public services, a monitoring programme run for seven years, and a standard of information openness written for public authorities. The method is carried over, not invented.
No reference standard? We can build yours.
By default the express check measures your service against a composite: the requirements we agree together, drawn from what you publish, what you tell me the agent must and must never do, and what regulators and customers reasonably expect. That works, but it is stitched together from the outside. A company that wants a reference standard of its own gets these three documents, written for it and kept. Arranged separately: write to infoaibusiness.vc and tell me about your service.
Company AI Policy
One public document for the whole company: by what rules does this company use AI, written for the customer rather than for lawyers.
Questions it answers
- What the AI does and how it identifies itself
- What happens to customer data
- Which decisions the AI may not take
- How to reach a human
- Who is personally responsible, and where to complain
Public, on your website. One document for the whole company.
AI Service Passport
One document per service. The policy sets general rules; the passport sets the norm for a specific service.
Questions it answers
- What the service is and what stages it has
- What data it needs, and what it does not
- What the agent may do alone, and what it may never do
- What counts as a result
- When a human must step in
The central document of a full assessment. Most items are measured against it.
AI Receipt
One document per interaction. As a till receipt confirms a purchase, this confirms the exchange.
Questions it answers
- Who spoke with whom, and when
- Whether AI involvement was disclosed
- What data was passed, and with what consent
- What actions were taken, under which identifiers
- How it ended, and where to turn in case of disagreement
The agent's own account of events. Its truthfulness is reconciled against your operations log.
Funding the company rather than running it?
A test purchase needs nobody's permission. I approach the service the way any customer would, using only what is available to anyone, and check it against what the company publishes. For an investor, an accelerator or a fund, that answers a question no pitch deck can: does the product do what the founders say it does, today, for a stranger with no special access and no demo script.
The result goes to you alone. Nothing is published, and the company gets no registry record and no badge. This is diligence, not a mark of approval.
By arrangement. Write to infoaibusiness.vc with the service you want looked at.
What this does not do
- A test purchase sees your service through a customer's eyes, not through your internals.
- One check proves a problem exists, not how often. Frequency needs a series.
- A full score means the agreed requirements held on the day, not that nothing will ever go wrong.
- Scores of different companies are not comparable: the requirements are different every time.
- This is a private, independent check. It is not an accredited conformity assessment.
One link, and you will know whether this can be checked at all.
Start with a free screening