Tested by AI Business

Your transcripts look fine. That is the problem.

The answers that cost you money are the ones that read perfectly. They only surface when somebody puts what your agent said next to what your company promised, and next to what your systems actually recorded.

That is what an independent test purchase does. We agree the requirements for your agent in advance, then I walk through your service as an ordinary customer and hand you the evidence.

AI Tested badge with a QR code leading to the public registry

Any of this sound familiar?

You have never read a full week of your agent's conversations. Only the ones a customer complained about.
Nobody in the company can write down, on a single page, what the agent is not allowed to do.
Your agent explains your refund rules from memory, and nobody has put its answer next to the policy you actually publish.
A corporate buyer asked how you prove the agent will not misinform their customers, and the honest answer was a paragraph of adjectives.

If the agent gave a bad answer last night, you would hear about it from the customer, not from your own monitoring.

You already test it. Here is what that cannot reach.

Where your own testing starts

The cases you imagined the agent might meet. Your evals and prompt reviews are good at those, and nothing here replaces them.

Where a test purchase starts

What your company promised, in public and in writing. Then backwards, to whether the agent honours it for a stranger.

The findings that hurt were never in the scenarios you wrote. They are in the ones nobody thought to write.

Three layers, and you hold two of them.

01

What your company promised.

Scattered across your site, your terms and your sales pages. Nowhere in your logs.

02

What the agent said.

In your transcripts.

03

What your systems actually recorded.

In your operations log.

Comparing the second and the third is engineering you can do yourself. The gap that costs money is between the first and the second, and nobody inside the company is placed to see it. You wrote the promise. You cannot also be the stranger who tests it.

From the reference pilot

What that gap looks like when you find it

The setup

The method was run end to end with NeoMundi, a French company working in AI metrology, in a controlled environment built for the purpose: an estate agency with a booking service, one model playing the seller and another the customer, real bookings with real identifiers and a log protected against backdating.

Code, prompts, runs and defects →

What it found

The agent completed the journey cleanly. In its closing receipt it gave the customer the email address of an employee who exists in no document of that company.

The transcript showed nothing wrong. Reconciliation against the log found it in a second.

This is not carelessness on your side.

In the same pilot we tested the test. Two identical purchases went through the measurement platform.

Purchase A

A genuine receipt

Purchase B

The record deliberately deleted

Identical scores.

If an instrument built for this cannot tell those two apart from behaviour alone, no amount of diligence inside your company will either. That is the line between measuring what a system does and checking what it promised, and it is a property of the setup rather than a comment on your team.

What you walk away with

You find out first

What your agent promises on your behalf, item by item, with the evidence behind each one. Before a customer finds it, and before it is public.

A record instead of a call

“How do you prove your agent will not misinform our customers?” is now a standard question in security reviews. Answer it with a link.

Proof instead of adjectives

Every competitor calls their AI accurate, safe and reliable. You are the one who can show that somebody outside the company checked, and hand over the record.

In the pack

Up to twenty agreed requirements for your agent, checked by an outsider

A report with the evidence and recommendations, yours either way

A numbered verification recorded in the public registry

A badge with a QR code for your site, so customers can check it themselves

A certificate you can show to buyers, partners and investors

Repeat checks reuse the requirements agreed the first time, so a second check costs less than the first

The registry records that your service is checked and when. The findings themselves stay between us: see the registry.

How it works

About a week from start to finish, and you are in control of the only decision that matters.

01

Send your link

Within a few working days you get one of two answers: here is what I would check and what it would involve, or this service cannot be checked this way and here is why. Free, and no obligation.

02

We agree what to check

I draft the requirements for your agent, up to twenty of them, from what you publish and from what you tell me matters. You add or swap items, then the list is frozen. The wording of each probe stays with me.

03

I buy like a customer

You get two quiet days to try everything yourself. Then, at some point that week, I walk through your service as an ordinary customer.

04

You get the report

Every item, with the evidence behind it and recommendations on what to fix. The report is yours: what you do with it is your business.

Said out loud

Who are you, and why would your record mean anything?

It is not an accreditation and it does not pretend to be one. It means one thing: the method is published in full, the reference pilot sits in an open repository with its code, prompts and runs, and the person who walked your service does not work for you. You can check me in ten minutes. Nobody outside your company can check your internal testing at all.

We already test our own bot.

Then you will recognise most of the report. It is the part you do not recognise that you are paying for. Your tests start from the cases you imagined; this one starts from what your company promised in public and works back to whether the agent honours it.

Why will you not name a price?

There is no price list for this service and no published rate: the price is agreed individually with each client, against the volume of work. One number for everyone has to be set high enough to cover the hard cases, and then the simple ones pay for that margin. A booking widget with four scripted answers and a bank's support agent with a hundred rules are not the same job. The screening is free, and the number comes before you owe anything.

What if the result is bad?

Then you are the only one who sees it. The findings go to you and nowhere else: I do not publish what a check found. The registry records that your service is checked and when, not how it scored on any given item.

Why would I pay someone to find problems?

Because your customers and your corporate buyers will find them anyway, and later, and in public. The report comes with recommendations, so the cheapest moment to learn about a problem is from someone who is not shouting about it.

You tell me what you will check. Doesn't that make it easy?

It makes it fair. You agree the subjects, never the questions, the scenario, the account or the moment. If your agent does what you require of it, being told the subjects in advance changes nothing. If it does not, no amount of warning will save it.

Will this disrupt our service?

No. The volume is that of an ordinary customer. Real orders are never taken to the irreversible step, and no real personal data is used.

Isn't this just red teaming?

No. Red teaming attacks the model to find what it can be made to do. A test purchase checks whether your agent meets the requirements you set for it, as an ordinary customer would experience them.

What do you need from us?

A link. That is all for an express check. Where paid access is needed, anything up to €50 is on me; more than that we agree in advance. I never use a test account you provide: that would tell you exactly who is checking.

Who this is not for

  • Companies with no customer-facing agent yet. There is nothing to walk through, and the documents are the cheaper thing to do first.
  • Enterprise rollouts behind a login that no member of the public can reach.
  • Products where the only way in is a sales call. A demo script is not a service.
  • Outbound voice campaigns. Different craft, different consent rules.
  • Anyone who wants a mark of approval rather than a finding. If your agent does not hold, the report says so, and you paid for it.

Express test purchase

Step one is free.

Send a link. Within a few working days you get one of two answers: here is what I would check and what it would involve, or this service cannot be checked this way and here is why.

There is no price list, and no published rate. The price is agreed individually after the screening, against the volume of work your service actually needs. A booking widget with four scripted answers and a bank's support agent with a hundred rules are not the same job, and pricing them the same would mean one of you is overpaying. Repeat checks and ongoing monitoring are quoted the same way, against requirements that already exist by then.

Start with a free screening

Send your application

Three fields, and an honest answer within a few working days: either your service can be checked and we start, or I tell you why it cannot.

Only live services are checked. If yours is still being built, come back when customers can use it.

Screening is free. You pay only after I confirm your service can be checked.

The method is published. The reference pilot is open.

Built and tested with a French AI metrology company

The method is published

The full method behind this service is written up as a 35 page guide: the standard, the forms of a check, how results are scored, the ethics, and the limits. Free to read, no registration. Scheduled monitoring on NeoMundi's measurement infrastructure, and anything else beyond a standard check, is a conversation rather than a package.

Read the method in the library →
Sergei Ponomarev

Who runs them

Sergei Ponomarev, PhD in political science. Before AI: hundreds of independent quality assessments and test purchases of public services, a monitoring programme run for seven years, and a standard of information openness written for public authorities. The method is carried over, not invented.

No reference standard? We can build yours.

By default the express check measures your service against a composite: the requirements we agree together, drawn from what you publish, what you tell me the agent must and must never do, and what regulators and customers reasonably expect. That works, but it is stitched together from the outside. A company that wants a reference standard of its own gets these three documents, written for it and kept. Arranged separately: write to infoaibusiness.vc and tell me about your service.

Company AI Policy

One public document for the whole company: by what rules does this company use AI, written for the customer rather than for lawyers.

Questions it answers

  • What the AI does and how it identifies itself
  • What happens to customer data
  • Which decisions the AI may not take
  • How to reach a human
  • Who is personally responsible, and where to complain

Public, on your website. One document for the whole company.

Funding the company rather than running it?

A test purchase needs nobody's permission. I approach the service the way any customer would, using only what is available to anyone, and check it against what the company publishes. For an investor, an accelerator or a fund, that answers a question no pitch deck can: does the product do what the founders say it does, today, for a stranger with no special access and no demo script.

The result goes to you alone. Nothing is published, and the company gets no registry record and no badge. This is diligence, not a mark of approval.

By arrangement. Write to infoaibusiness.vc with the service you want looked at.

What this does not do

  • A test purchase sees your service through a customer's eyes, not through your internals.
  • One check proves a problem exists, not how often. Frequency needs a series.
  • A full score means the agreed requirements held on the day, not that nothing will ever go wrong.
  • Scores of different companies are not comparable: the requirements are different every time.
  • This is a private, independent check. It is not an accredited conformity assessment.

One link, and you will know whether this can be checked at all.

Start with a free screening