Benchmarks
Our own data — not the market's press releases
Most AI coverage repeats what companies say about themselves. These are the numbers we produce and maintain ourselves: what models cost, who actually earns money, and whether customer-facing AI does what it promises. Free to read, free to cite — attribution is all we ask.
Models
LLM Leaderboard
Every frontier and open-weight model that matters, compared on price per million tokens, context window, and public ELO where it exists. Updated as models ship, not once a quarter.
Money
AI Revenue Leaderboard
Who actually earns money in AI, ranked by run-rate revenue, valuation and revenue per employee — with the accounting caveats that make the numbers honest, and corrections when the press gets it wrong.
Tools
Tool Directory & Comparisons
Reviewed tools with pricing, target user and the one feature that justifies the bill — plus head-to-head comparisons within each category.
Service quality
AI Service Check
Mystery shopping for customer-facing AI: ten real situations per bot, checked against the company's own published rules, with a quote as evidence for every finding.
In progress
In progressAI Receptionist Test
Independent test calls to the AI receptionists small businesses actually buy. Does it quote the right price, book the appointment, admit it's AI, and hand you to a human when it can't help?
Visibility
AI Visibility Benchmark
How AI search engines read a website: llms.txt quality, crawler access, schema, citation readiness. Run it on your own domain and get the scoreboard by email.
Why we publish benchmarks
Every company will tell you its AI is fast, accurate and helpful. Almost nobody checks from the outside. The method behind these benchmarks comes from twenty years of independent service assessment — nationwide monitoring, quality standards and mystery shopping — now pointed at AI. More about the author →
If you want the methods themselves rather than the results, they are free in the Library — no email, no registration.