Same AI-bot, same question, three sessions: who leads this organisation?
First session: it didn't know. Second: the right name. Third: it said it was there to talk about the condition, not the organisation.
I ran a test purchase of the patient-education AI assistant of a US health nonprofit: 27 questions, as an ordinary visitor, three sessions over ten days, every answer checked against what the organisation publishes on its own website.
The medical limits held every time. It never diagnosed, never approved changing a dose, refused medical documents, and made up nothing about a drug name I invented.
That is the hard part, and it was built well.
Several basic facts about the organisation were unpredictable.
Who leads it, who sits on the board, which email to write to: unanswered, answered, unanswered again. A wrong founding year appeared in the second session and was still there in the third.
Privacy never landed at all.
The policy says plainly what happens to a conversation: stored for up to 40 days, only general themes extracted, used to improve the chatbot.
In three sessions the bot didn't say it once.
The lesson is not about this bot.
Ask an AI assistant the same question in different sessions and you may get different answers, and the organisation behind it is often the last to know.
One check gives you a snapshot. Repeated checks show whether the behaviour holds.
If you run an AI assistant, send me a message.
Together we'll agree a clear list of what it must get right for your business, and I'll test it against that list on a regular basis, so you're the first to know when your assistant starts to drift.
