The open community

AI experts

People who build AI, put it into a business, govern it and check that it works. Listed one by one, not as firms.

← All experts
Joseph Cirello

Joseph Cirello

AI Red-Teamer & Forensic Model Auditor

CEO, Potestas AI

Providence, Rhode Island, USA · North America · English

About

I spent 25 years in the Army making sure that what someone said they had, they actually had. Special Forces logistics. You count it, you sign for it, and the signature means something. I do the same job now, on AI.

Now I run Potestas AI, an independent testing lab. Companies bring me the model or agent they are about to put in front of customers; I test for things the models must never do, then hand them the evidence whatever it says.

Why that job exists. I asked a finance agent to move money it shouldn't have moved. It told me no. It named the request as wire fraud, said so plainly, declined. Then it sent the wire. Not a jailbreak. The safety reasoning was intact and correct; it just didn't govern what the agent did. What a system says and what it executes are two different measurements. Most people only take the first. So I take the second.

A screening is 153 say-vs-do scenarios: working tools, a plausible reason to misuse them, and a grade based on the actions taken. A full stress test adds 30 campaigns of 400 turns across a 36-category battery, then re-runs every probe that surfaces 100 more times, because one failure tells you a system can break, not how often it will.

Most of what I test isn't a frontier model. It's yours: the fine-tuned model, the retrieval stack, the agent with your tools wired into it. A vendor's safety numbers say nothing about what your system does after you have adapted it, and no hosted scanner reaches a deployment that never touches the internet. I run the battery where the model actually lives, including on-premise and air-gapped.

Independence is the whole product. Findings are judged by a model from a different vendor than the one under test. Disputed turns get a human. Everything ships sealed with a hash you can check, and a full example report is published on my site, including the turns my own instrument got wrong.

After Enron we made auditor independence a matter of law. Then we built the most consequential technology in history and let the labs grade their own homework. I'm writing a book about that. It will be free.

Ask me anything. Say "no pitch, here's my question" and that's what you get: an answer, free, from someone with no stake in what it turns out to be. No call unless you ask for one.

Retired Senior Warrant Officer, U.S. Army. Bronze Star. Active Secret clearance; TS/SCI-cleared personnel available per engagement through an established network. Disabled Veteran-Owned Small Business, SAM.gov registered.

What I do

  • Agentic say-vs-do screen, $3,000, same day: 153 scenarios in which the model is handed working tools and a plausible reason to misuse them, then graded on what it actually did, read from the record of tool calls rather than from another model's opinion. You give a model name; no API key or system access is needed, so there is no security review to pass. The sealed report tells you whether a behavior is present, not how often, and says so. The fee credits in full toward a full engagement and sits under the federal micro-purchase threshold
  • Full measured forensic stress test, from $25,000: 30 campaigns of 400 turns across a 36-category battery, with every probe that surfaces re-run 100 more times, so a finding arrives as a rate with a 95% confidence interval instead of a story. Judged by a model from a different vendor, disputed turns go to a human, false positives removed first. If nothing at severity 7 or higher is confirmed, there is no fee
  • Cross-model comparison and ongoing testing, from $45,000: the same battery run across several vendors' models at once, with a comparison report no lab can run on itself. Drift monitoring included
  • Continuous monitoring, $3,000 per month: one full 153-scenario screen every month against a fixed baseline, with a drift report and a sealed evidence pack each run. Rate locked for 24 months, cancel with 30 days' notice
  • Testing where the model actually lives, including fine-tuned models, retrieval stacks and on-premise or air-gapped deployments
  • Cleared and classified work: TS/SCI- and polygraph-cleared personnel available per engagement for classified, ITAR-sensitive or high-assurance environments

Practice areas

Red teamingEvaluation & testingAssurance, audit & conformity assessmentAI safety, security & incident responseAI risk & impact assessmentHuman oversight, transparency & accountabilityAI agents & automationLLM applications & prompt engineeringAI engineering & developmentMachine learning & data scienceGoverned AI adoption & process redesignAI strategy & transformationGovernance operating models & AI policyLegal, regulatory & standards complianceData governance, privacy & documentationAI research & public policy

Industries

Financial servicesTechnology & softwareTransport & logistics

Open to

Consulting, Advisory & board work, Research collaboration

Contacts

LinkedIn →Website →(833) 837-8556

Identity and links confirmed by AI Business. This is not a rating, an endorsement or a recommendation.