← Founder's Notes

12 US State DMV AI Bots, the Same 10 Questions: Scores from 100 to 6

By Sergei Ponomarev · September 28, 2026

12 US state DMV AI bots. The same 10 questions. Scores from 100 to 6.

One topic for all of them: REAL ID.

The questions an ordinary resident would ask: what to bring, is there a form, how much, how long, what if my passport expired, how do I reach a person.

One rule for grading: the correct answer is whatever the state's own website says.

No legal interpretation, no expert opinion.

If a bot contradicts its own agency, that is the finding.

What we found:

❗ Personal data. One bot asked for the driver's licence number, although the same chat warns users not to enter personal information.

❗ Wrong facts. Bots gave a wrong phone number, promised the card "within a few business days" when the agency says 30 to 45 days, and said a temporary paper document works at the airport, while the state's own page says the TSA does not accept it.

❗ Wrong topic. Asked for the REAL ID application, one bot offered the vehicle title form and the boat registration form.

❗ Breakdowns. A visitor was shown an unfinished developer placeholder. Another was told to log in where no login exists. One bot answered "Sorry, I didn't get that" to 8 of 10 questions.

❗ False premise. When we slipped one into a question, not one bot agreed with it. Only 5 of 12 corrected it plainly.

And the best AI bot got every question it was scored on right. The gap is not between AI and no AI. It is between bots on the same task.

Why this experiment is unusual.

Public services have used mystery shopping for decades. We applied the same logic to AI agents.

  • One scenario, many agents. A single test can be written off as a bad day. Twelve side by side cannot.
  • The answer key is each agency's own website, saved before the conversations and fingerprinted with SHA-256.
  • Every answer, screenshot and source is kept. A script computes the scores: run it again and you get the same ranking.
  • No accounts, no personal data, no red teaming. Only the questions real residents ask.

This is a pilot: one conversation per bot, one day, 12 of 50 states. It is not a verdict on any agency. But it shows that public AI assistants can be measured the way we measure public services.

The full report is in our Library, free, 21 pages with every answer, every source and the state-by-state ranking: US DMV Chatbots and REAL ID.