Field note · What we found
We checked five AI statistics we were about to publish. Three had no source.
This is not a think piece about AI hype. It is the audit log of five specific numbers we were about to put in front of buyers, and what happened when somebody was made to find the primary source for each.
The short answer
Of five statistics we intended to publish, one was misattributed to Gartner and traced to a software vendor’s press release, one was stale and had quietly stopped supporting its own argument, and three had no traceable source at all — with the largest public dataset contradicting one of them outright. Two survived. They are the two now on our site.
Every AI agency page carries numbers. Ours did too. In September 2026 we made somebody find the primary source for each one before it shipped.
Two survived intact. Here is the whole audit, including the one that was on our own homepage.
1. "A 1,445% one-year surge in demand for multi-agent systems, says Gartner"
Verdict: misattributed. Removed.
This was on our homepage. Chasing it back, the figure appears in a press release issued by Belitsoft, a software development company. We could find no Gartner publication carrying it.
Borrowing an analyst's name for a vendor's number is worse than making one up, because the borrowed authority is what makes it persuasive. It is gone.
2. "40% of enterprise apps will run task-specific AI agents by 2026, says Gartner"
Verdict: true, but misworded. Corrected and dated.
This one is genuinely Gartner's — press release of 26 August 2025, predicting 40% by 2026, up from less than 5% in 2025. But Gartner published "feature", and we had written "run".
One verb. It is the difference between quoting an analyst and paraphrasing one, and only one of those is checkable. It now matches, and carries the date.
3. "Desktop computer-use scores in the low 80s on OSWorld"
Verdict: stale, and it had stopped making its own point. Replaced.
The current leaderboard figure is 86.1%. More importantly, the human baseline on OSWorld is 72.36% — so agents now score above humans on that benchmark.
We had been using "low 80s" to imply a ceiling. At 86.1% against a 72.4% human baseline it implies the opposite. The number had drifted past the argument it was hired to make.
4. "Across 500,000+ cold emails, AI personalisation performed 25% worse than sending nothing"
Verdict: no source found. Contradicted. Struck.
We could not locate any study matching this claim. What we did find was Hunter's State of Email Outreach 2026, drawn from 31 million emails, reporting that two custom attributes lift reply rate by 56% — 5.6% against 3.6%.
That is not a small discrepancy. It points the other way, from a dataset sixty times larger than the one our claim cited.
5. "One human running two AI seats books 1.9× more meetings per dollar"
Verdict: no source found. Struck.
Nothing. No study, no dataset, no vendor even claiming it. It had presumably been absorbed from a conference slide.
What replaced them
| Claim | Source | Weight |
|---|---|---|
| 40% of enterprise apps will feature task-specific agents by end of 2026 | Gartner, 26 Aug 2025 | Analyst, primary |
| 86.1% best agent on OSWorld, vs 72.36% human baseline | OSWorld leaderboard | Public leaderboard, primary |
| 30% of long-horizon cross-application tasks completed unaided | CMU, arXiv:2412.14161 | Preprint, primary |
| Manually edited emails beat fully automated by 18% | Hunter, 31M emails | Vendor dataset, large, method stated |
The method, so you can do it to us
- Find the primary source, not the article citing it. Most "according to Gartner" claims resolve to a blog citing a blog citing a press release.
- Check the date. Benchmarks move. A true 2024 number can be a false 2026 one.
- Check the wording against the original. "Run" for "feature" is how a checkable claim becomes an unfalsifiable one.
- Ask who profits. A vendor dataset is not worthless — Hunter's 31 million emails are worth more than most academic samples — but it must be labelled.
- Write the date you checked it, so the next person knows how stale it is without repeating the work.
Every figure on our site now carries a visible source and the date it was last checked, and our build fails if one does not. That is not a promise. It is a test that runs.
Questions people actually ask
How do you check whether an AI statistic is real?
Trace it to the primary source rather than the article citing it, check the publication date against how fast the field moves, compare the wording to the original, note who profits from the claim, and record the date you checked.
Is the 1,445% multi-agent growth figure from Gartner?
We could find no Gartner publication carrying it. The figure appears in a press release issued by Belitsoft, a software development company. We removed it from our own site for that reason.
Does AI personalisation improve cold email reply rates?
The largest public dataset we found — Hunter’s State of Email Outreach 2026, covering 31 million emails — reports that two custom attributes lift reply rate from 3.6% to 5.6%, a 56% improvement. It also found manually edited emails outperform fully automated ones by 18%.
Sources
- Gartner, Gartner Predicts 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026, Up from Less Than 5% in 2025, 26 August 2025. — analyst, primary re-checked yearly
- OSWorld-Verified leaderboard — 86.1% best agent (Qwen3.8-Max); 72.36% human baseline from the original OSWorld paper. Re-checked 8 September 2026, still the top score. source ↗ — public leaderboard, primary re-checked every 30 days
- TheAgentCompany, Carnegie Mellon, arXiv:2412.14161 — 30% of tasks completed autonomously. source ↗ — preprint, primary
- Hunter, State of Email Outreach 2026 — 31 million emails; two custom attributes 5.6% vs 3.6% reply; manually edited 5.2% vs fully automated 4.4%. — vendor dataset, large, method stated re-checked every 6 months
- Belitsoft press release, Multi-Agent Systems Surge 1,445%… — the actual origin of the figure we had credited to Gartner. — vendor press release — the point of the story
Checked 8 September 2026. Next scheduled check 8 October 2026. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.
If you find an unsourced number on our site, tell us.
We will either produce the source or remove the claim, and say which in the note. That is the whole positioning, and it only works if it is testable.
Talk to us about this