https://www.noxia.co.uk/field-notes/the-default-changed-on-15-september · printed from noxia.co.uk · sources checked 24 September 2026
Field note · Being found
On 15 September the default changed. Your site may have stopped talking to AI.
A setting you did not choose, on a service you may not know you use, changed on 15 September. If your site sits behind Cloudflare, was added since then and carries ads, the assistants that fetch pages to answer questions may now be turned away at the door — and the only symptom is an absence.
- Being found
- AI search
- Crawlability
- Being cited
- Cloudflare
The short answer
On 15 September 2026 Cloudflare began offering every new domain one of two presets. A site without ads allows everything. A site that carries ads keeps search crawling but sets “Disallow AI Training” — which accountable crawlers from Apple, Google and Microsoft respect while still indexing for search — and blocks AI agents on pages with ads. Existing customers’ settings were migrated from their earlier choices. The fix is a setting, not a rebuild.
On this page · 6 sections
Most firms do not know whether they are behind Cloudflare. They know they have a website, that somebody set it up, and that it works. Cloudflare sits in front of roughly a fifth of the web as a proxy — it is what your domain points at before your actual site answers — and it is switched on by default by a great many hosts and agencies without ever being mentioned in the invoice.
That matters this month because on 1 July 2026 Cloudflare announced a change to what it does by default, and on 15 September 2026 it launched a version of it — narrower than the announcement, and more precise.
Why would anyone block a crawler?
Because the exchange stopped being an exchange. The old bargain with a search engine was legible: it took your page and sent you readers. The new bargain is that a model takes your page, answers the question itself, and sends you nobody. Cloudflare's own framing when it began this in 2025 was blunt about the ratio: getting traffic from OpenAI, it said, was 750 times more difficult than from “the Google of old”, and from Anthropic 30,000 times — figures it gave without saying how it measured them.
There is also plain waste in it. Cloudflare says more than half of AI crawler traffic is re-fetching pages that have not changed. A small firm's site being re-read forty times a week by bots is not a compliment; it is bandwidth.
So the policy is not "no AI". It is: say which one you are. The version that launched makes the distinction explicit. “Disallow AI Training” publishes the preference in robots.txt, and the accountable mixed-use crawlers — Cloudflare names Apple’s, Google’s and Microsoft’s — keep indexing for search while refusing to train. Agents, which fetch a page to act on it, are turned away from pages with ads.
Who is actually affected
New domains added from 15 September, and the preset they get depends on whether the site carries ads. Existing customers’ settings were migrated from what they had chosen before: a site that had blocked AI bots, or blocked training on pages with ads, now carries “Disallow AI Training”. The July announcement said the new default would also reach every free-plan site; the launch post does not say so, and we are not repeating it.
So, plainly: this was not a retroactive switch-off. It is a change to what a new site gets before anybody configures it. But "new site behind Cloudflare" describes a great many small professional firms in the UK, including almost every site built since mid-September by an agency that uses Cloudflare as a matter of course.
How to find out in four minutes
- Find out whether you are behind Cloudflare at all. A DNS lookup on your domain answering with Cloudflare nameservers, or a response header naming it, settles it. If you do not run your own DNS, your web developer knows in one sentence.
- Read your own robots.txt. Type your domain followed by
/robots.txt. It is a plain text file. If it names crawlers you have never heard of and disallows them, somebody made a decision on your behalf. - Check the AI crawler setting in the Cloudflare dashboard, under the bot management controls. It is a toggle, not a project.
- Then check whether you were ever readable in the first place, which is a different and more common problem. The visibility checker reads a page the way an answer engine does and tells you what it could and could not find, and the stack check asks the same question of the software behind it.
What this does not tell you
It does not tell you that being crawled is the same as being cited. It is not. A model that can reach your page still has to find a sentence worth lifting — attributable, dated, specific — and most professional-services websites do not contain one. Being blocked is a problem you can fix with a toggle.
Being uncitable is a problem you fix by writing differently, which is the argument in the note on invisibility and the reason every figure on this site carries a source, a date and a weight.
It also does not tell you that allowing every crawler is correct. A firm whose content is its product may reasonably decide that being summarised without attribution is a bad trade. The point is that it should be a decision. On 15 September, for a lot of people, it stopped being one.
Questions people actually ask
Did Cloudflare block AI crawlers on all websites?
No. From 15 September 2026 each new domain is offered a preset: a site without ads allows everything, and a site that carries ads keeps search crawling, sets “Disallow AI Training” and blocks AI agents on pages with ads. Existing customers’ settings were migrated from what they had chosen before. Search crawling is allowed either way.
How do I know if my website is behind Cloudflare?
Check whether your domain uses Cloudflare nameservers, or look at the HTTP response headers your site returns — Cloudflare adds its own. If a web developer or agency set the site up, they will know in one sentence. Many UK small-business sites are behind it without the owner having chosen it.
Does the new Cloudflare default stop me appearing in Google?
No. Search crawling stays allowed, and Google is one of the accountable crawlers that honours “Disallow AI Training” while still indexing for search. What the default can affect is whether AI agents — assistants that fetch a page at answer time to act on it — can reach your pages that carry ads.
Should a small professional firm allow AI crawlers?
It depends on whether being quoted is worth more to you than being copied. For a consultancy, broker or agency whose website is a shop window rather than the product, being readable and citable is usually worth more than the content is worth as training data. For a publisher whose archive is the business, the calculation is different. The change is reversible either way.
Sources
- Cloudflare, “Have it both ways: stay discoverable in search while disallowing AI training”, 15 September 2026: “Beginning September 15, customers onboarding a new domain will be offered one of two preset configurations, depending on whether the site earns money from advertising”; for an ad-supported site, search allowed, training set to the new “Disallow AI Training” and agents blocked on pages with ads; the setting “lets you easily stay indexed for search while refusing to let that same crawler train on your content”, honoured by accountable mixed-use crawlers from Apple, Google and Microsoft; existing customers’ settings migrated from their earlier choices. blog.cloudflare.com ↗ — primary; the company’s own launch post re-checked yearly
- TechCrunch, “Cloudflare’s new policy pushes AI companies to pay for publishers’ content”, 1 July 2026 — the plan as announced, which the launch narrowed: from 15 September 2026, Cloudflare’s default settings block “mixed-use” crawlers from ad-supported pages, applying to new customers, new sites from existing customers, and all free-tier users; site owners can change the setting. It reports Cloudflare’s position, in the reporter’s words rather than a quotation, that most website owners want their content discoverable via search and often through AI services, but want protection against their intellectual property being given away for free. Also cited: more than 50% of AI crawler traffic is re-fetching unchanged pages. techcrunch.com ↗ — secondary; trade press reporting the July announcement, superseded by the launch above re-checked yearly
- Cloudflare, “Content Independence Day: no AI crawl without compensation!”, 1 July 2025 — the announcement that established the policy direction a year earlier: “changing the default to block AI crawlers unless they pay creators for their content”. On the ratio: “With OpenAI, it’s 750 times more difficult to get traffic than it was with the Google of old”, and “With Anthropic, it’s 30,000 times more difficult”. blog.cloudflare.com ↗ — primary; the company’s own blog. The post gives no method for either figure, and we have not seen them independently reproduced re-checked yearly
- The four-minute check, the reading of “default” as the operative word, and the claim that crawlability and citability are different problems are ours. — our own argument, labelled as such
Checked 24 September 2026. Next scheduled check 24 September 2027. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.
Cite this note
Noxia, “On 15 September the default changed. Your site may have stopped talking to AI”, Field notes, 18 September 2026; sources checked 24 September 2026. https://www.noxia.co.uk/field-notes/the-default-changed-on-15-september
Find out what a crawler can actually see.
Our visibility checker reads a page the way an answer engine does and reports what came back — what it could find, what it could quote, and what was hidden behind script. It is free and it runs in your browser. If the answer is worse than you expected, the retainer is what fixes it and keeps it fixed.
Talk to us about thisRead next
Being found
Your firm may be invisible to ChatGPT. Here is how to check in ten minutes.
Crawlability scored 9.5 as a citation factor, the top score of twenty-three and ahead of every content signal. Plenty of adviser sites fail at that first hurdle without anyone knowing.
Being found
Your old website is still in Google, and it is sending people nowhere.
A rebuilt site leaves its old URLs in Google for months. Search Console reports them as excluded, which alarms people, and is usually correct.
Being found
Search referrals to finance content fell 23% in one quarter.
Across participating UK publishers, organic Google referrals fell 7.1% in a quarter. Finance and business fell 23%, how-to guides over 20%, while opinion and commentary rose 4%.