AI Test Purchase · Pilot on 12 state government chatbots
US DMV Chatbots and REAL ID
For public-sector AI owners and researchersTwelve US state DMV chatbots, the same ten questions, every answer checked against the agency's own website. Scores ranged from 100 to 6.
A single test of one chatbot is easy to dismiss as a bad day. A comparison is not. This report applies the logic of mystery shopping, long used to check shops and public counters, to AI assistants on government websites. Twelve US state DMV chatbots that could be tested comparably from Europe, without an account, personal data or a live operator, each received the same ten questions about REAL ID in one conversation on 24 September 2026. Answers were graded only against the same agency's public website, with no legal interpretation: if a bot contradicts its own agency, that is the finding. The core reference pages were saved before the conversations and fingerprinted with SHA-256; every answer, screenshot and source is kept, and a script computes the scores so that anyone rerunning it gets the same ranking. Scores ranged from 100 to 6. Eight of twelve systems had one substantive factual error or risk, from a wrong phone number and a wrong timeline to vehicle and boat forms offered as the REAL ID application. Three more failures do not fit a factual-error count: a developer placeholder shown to a visitor, a login demand where no login exists, and a dialogue that answered 'Sorry, I didn't get that' eight times. No bot agreed with a planted false premise, but only five corrected it plainly. The report is a descriptive pilot, one conversation per system on one day, and not a verdict on any agency. Twenty-one pages with the full answers, sources and checksums in the appendices.
What's inside · 21 pages
- A descriptive combined ranking of 12 state DMV chatbots, from 100 to 6
- Eight substantive errors, each set against the agency's own published wording
- Service failures a factual-error count misses: a leaked placeholder, a phantom login, a dialogue that never started
- How the bots handled a false premise and a request for personal data
- Which questions separated strong bots from weak ones
- The method: one scenario, the agency's site as the only answer key, SHA-256 evidence, reproducible scoring
- Every answer and its source for all twelve states, and the limits of the pilot stated plainly
Download the PDF — free
No email, no registration. Just take it and use it today.
Download US DMV Chatbots and REAL ID →