TL;DR
- Cloudflare launched a Disallow AI Training setting on 15 September 2026 that writes a no-training preference into robots.txt while Googlebot, Applebot and Bingbot keep crawling for search.
- Block now stops those three crawlers completely, including for search, reversing the July plan in which blocking AI training would have blocked them automatically.
- At Google the setting works through a Disallow rule for Google-Extended, which Google documents as having no effect on Search inclusion or ranking.
- Bing is not covered yet. Microsoft support for a robots.txt no-training preference is targeted for early 2027, and the current opt-out is the NOARCHIVE meta tag.
- Block AI Bots and Managed Robots.txt are being deprecated, and existing Training blocks migrate to Disallow AI Training automatically.
Cloudflare has introduced a Disallow AI Training setting that writes a no-training preference into a site's robots.txt while leaving Googlebot, Applebot and Bingbot free to crawl for search. Search Engine Journal reported the change on 15 September 2026. The practical effect is that site owners who want their content kept out of AI model training no longer have to accept losing search crawling as the price, and the older all-or-nothing Block AI Bots toggle is being retired.
The reason this matters now is that Cloudflare had previously signalled the opposite. As Search Engine Journal notes, the plan described in July was that from 15 September, sites blocking AI training would also block Googlebot, Applebot and Bingbot, because all three crawl for both search and model training. That plan has changed. Block still exists, and it is now the total option: choosing Block stops those three crawlers completely, search included.
What Cloudflare changed on 15 September 2026
Disallow AI Training is a new option inside Cloudflare's Training control. The Training control is one of three, alongside Search and Agent, and Cloudflare first described the Disallow AI Training option in August. When Disallow AI Training is enabled, crawlers that serve both search and training stay allowed for search, provided Cloudflare has labelled the operator behind them as Accountable. Training-only crawlers are blocked.
Two settings that used to exempt mixed-use crawlers no longer do. Block and Block on pages with ads now apply to them as well. Cloudflare's stated reason for the earlier exemption was that blocking those crawlers could damage a site's search visibility, which is exactly the trade-off the new Disallow AI Training option is meant to remove.
Most sites are not expected to do anything. Existing Training selections of Block or Block on pages with ads migrate to Disallow AI Training. Sites that had used either blocking option under the older Block AI Bots toggle end up with Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent. Block AI Bots itself, along with Cloudflare's Managed Robots.txt feature, is set to be deprecated. New domains that monetise with ads are offered Disallow AI Training as their Training preset.
Here is how the settings line up after the change, based on what Cloudflare described.
| Setting | What it does now |
|---|---|
| Disallow AI Training | Publishes a no-training preference in robots.txt. Mixed-use crawlers from Accountable operators keep search access. Training-only crawlers are blocked. |
| Block | Stops mixed-use crawlers entirely, including Googlebot, Applebot and Bingbot crawling for search. |
| Block on pages with ads | Now also applies to mixed-use crawlers, which it previously did not. |
| Block AI Bots (old toggle) | Being deprecated. Sites that used it are migrated to Allow for Search plus Disallow AI Training plus Block on pages with ads. |
What Accountable means, and who currently qualifies
Accountable is a label Cloudflare created after talks with crawler operators that began in July. A mixed-use crawler keeps search access under Disallow AI Training only if its operator carries that label, and Cloudflare set four requirements an operator has to meet or commit to meeting.
- A way to opt out of AI training through robots.txt or a comparable standard.
- A way to opt out of AI summaries, set with the operator now and through Cloudflare next year.
- URL-level visibility into which pages were made available for training, plus metrics on how the content appeared in search.
- An assurance that opting out of training will not affect traditional search results.
Cloudflare says Apple, Google and Microsoft meet those requirements, each with some features available now and commitments carrying deadlines for the rest. It also lists the relevant crawlers from Amazon, Anthropic, Meta and OpenAI as Accountable, on the grounds that those companies run separate search and training crawlers. Their training crawlers remain blocked under Disallow AI Training, which is the intended outcome rather than a gap.
How the setting maps to Google, Apple and Bing
The setting is not one universal switch. It resolves to a different underlying mechanism at each of the three search companies, and the coverage is uneven.
At Google, Disallow AI Training works through a Disallow rule for Google-Extended, the robots.txt token Google provides for opting content out of Gemini model training. Google's own crawler documentation states that Google-Extended does not affect a site's inclusion in Google Search or its ranking. That is the load-bearing detail for anyone nervous about touching robots.txt: the token was built to be neutral to search.
Google-Extended is also not the control for AI Overviews or AI Mode. A separate Search Console setting governs whether a site appears in AI Overviews, AI Mode and Discover's generative AI features, and Google's help page says that setting does not affect AI training. The two are independent, and confusing them is the easiest way to get the wrong outcome: a site can be opted out of training and still be summarised in AI Overviews, or the reverse.
At Apple, the setting works through a Disallow rule for Applebot-Extended. Apple's documentation says Applebot-Extended does not crawl pages and is not considered in search ranking. Keeping content out of AI-generated answers to broad knowledge questions in Siri and Search is a separate job, and Apple's documented route for that is the nosnippet meta tag.
Bing is the gap. Choosing Disallow AI Training will not send Bing a no-training preference through robots.txt yet, because Microsoft still has to add support for it, and Cloudflare says that work is in progress. Cloudflare's earlier Training block did not affect Bingbot either, so nothing has regressed, but nothing has improved for Bing either. Bing's current training opt-out is the NOARCHIVE meta tag. Bing's documentation says content marked NOARCHIVE is not used to train Microsoft's generative AI models, and it is also not linked in Chat and Copilot, which means the Bing opt-out carries a visibility cost that the Google and Apple tokens do not.
| Operator | Mechanism used | Status |
|---|---|---|
| Disallow rule for Google-Extended | Live. Google documents that it does not affect Search inclusion or ranking. | |
| Apple | Disallow rule for Applebot-Extended | Live. Apple documents that it is not considered in search ranking. |
| Bing | robots.txt no-training preference | Not yet supported. Microsoft support targeted for early 2027. Current opt-out is the NOARCHIVE meta tag. |
What is coming next, with the dates Cloudflare gave
Cloudflare set out three follow-on items. Google is planning to roll out URL-level transparency tools for Google-Extended in the coming weeks. Apple is working on its own URL-level tool for next year. Microsoft's support for a robots.txt no-training preference is targeted for early 2027, which makes Bing the slowest of the three by a wide margin.
Cloudflare's own next focus is AI summaries rather than training. The company says its goal is a single Cloudflare setting that controls how much content is included in summaries, instead of the current situation where the preference has to be set operator by operator, and it puts that target at early next year.
What to check in your own setup
The migration is automatic, which is the risk: a setting changed on your behalf is a setting nobody on the team has read. A short check is worth doing.
- Open the Training control in Cloudflare and read what your domain is actually set to now, rather than what it was set to in July. A domain that was on Block AI Bots has been moved to a three-part configuration.
- Fetch your own robots.txt and read it. If Disallow AI Training is active you should see a no-training preference in the file, and Googlebot should not be disallowed.
- Confirm nothing is sitting on Block by accident. Block is now the setting that stops Googlebot, Applebot and Bingbot outright, and the same word meant something narrower before 15 September.
- Check your Search Console generative AI setting separately if AI Overviews and AI Mode appearance is what you actually care about. Google-Extended does not control it.
- If Bing and Copilot matter to your traffic, decide on NOARCHIVE deliberately. It stops training but also removes Chat and Copilot links, so it is a different trade from the Google and Apple tokens.
- Log the date you checked. Cloudflare has now changed the behaviour of these controls twice in three months, and a dated note is what tells you which version you configured against.
If crawler access and AI citation behaviour are already part of how you report on organic performance, this belongs in that same review. Sites that want the visibility side of the same question examined can start with a technical SEO audit, and the citation side sits with generative engine optimisation work.
What the announcement did not say
Several things a site owner would want to know are not in the source material, and it is worth naming them rather than filling the space with inference.
Cloudflare did not publish a figure for how many domains the automatic migration touches. It did not say whether the Accountable label can be revoked from an operator that misses one of its committed deadlines, or what happens to sites relying on that operator if it is. It did not publish the exact robots.txt syntax the Disallow AI Training preference writes, beyond describing it as a no-training preference. And it gave no measurement at all of what opting out of training does to a site's presence in AI answers over time, which is the question most publishers actually care about. The training opt-out and the summary opt-out are different controls, and only the first one is being changed here.
What this means for Thai marketers
Most Thai sites of any size sit behind Cloudflare, so this change reaches further into the local market than the average platform announcement. The important part is the one that removes a bad trade. Until now, a Thai publisher or e-commerce site that wanted its product copy and articles kept out of model training risked doing damage to its Google crawling in the process, and for a market where Google organic is the dominant discovery channel, that was rarely worth it. Disallow AI Training separates the two decisions.
A second point is specific to bilingual sites. Thai-language content is comparatively scarce in training data, which cuts both ways: it is more valuable to a model, and it is also the content most likely to be the reason a Thai brand is mentioned at all in an AI answer. Opting out of training does not opt a site out of being cited in AI Overviews or AI Mode, because Google controls those through Search Console instead. A Thai site that wants the citation without the training can, on Google's documentation, have both.
The Bing gap matters less here than it would in the United States, given Bing's share of Thai search, but Copilot is a different question from Bing search, and the NOARCHIVE trade should be made on purpose rather than by default.
Frequently asked questions
Does Disallow AI Training hurt my Google rankings?
Google's crawler documentation says Google-Extended does not affect a site's inclusion in Google Search or its ranking, and Disallow AI Training works through a Disallow rule for Google-Extended. Cloudflare also made the assurance that opting out of training will not affect traditional search results one of its four Accountable requirements. The setting to be careful with is Block, which does stop Googlebot entirely.
Does this stop my pages appearing in AI Overviews and AI Mode?
No. Appearance in AI Overviews, AI Mode and Discover's generative AI features is controlled by a separate Search Console setting, and Google's help page says that setting does not affect AI training. They are two independent controls and this change only touches the training one.
Do I need to change anything if I already blocked AI bots?
Cloudflare says most customers do not. Existing Training selections of Block or Block on pages with ads migrate to Disallow AI Training, and sites that used the older Block AI Bots toggle are moved to Allow for Search, Disallow AI Training and Block on pages with ads. It is still worth opening the dashboard and reading the result, because Block AI Bots and Managed Robots.txt are being deprecated.
Is this live for Bing?
Not yet. Cloudflare says choosing Disallow AI Training does not send Bing a no-training preference through robots.txt because Microsoft still has to add support, which Cloudflare puts at early 2027. Bing's current training opt-out is the NOARCHIVE meta tag, which also removes links from Chat and Copilot.
Which AI companies keep search access under this setting?
Cloudflare says Apple, Google and Microsoft meet the Accountable requirements, and it also lists the relevant crawlers from Amazon, Anthropic, Meta and OpenAI as Accountable because those companies run separate search and training crawlers. The training crawlers of those operators are still blocked under Disallow AI Training.
Where this leaves site owners
The useful summary is that the choice has been split in two. Training and search were bundled together in one crawler for years, which made every opt-out decision a gamble on organic traffic. Cloudflare has now written the split into a dashboard setting, Google and Apple have documented tokens that back it, and Microsoft has not caught up. Nothing here changes how a page ranks or whether an answer engine cites it, so the practical work is unchanged: read your current setting, read your robots.txt, and keep the training decision separate from the visibility decision. Teams that want help mapping both sides of that against their own traffic can talk to our AI SEO team.







