Report: Cloudflare's AI crawler blocking is returning 403s to Googlebot, weeks before the September deadline

Report: Cloudflare's AI crawler blocking is returning 403s to Googlebot, weeks before the September deadline

SEOAugust 6, 2026
By Antonio Fernandez

Search Engine Journal reported on 4 August 2026 that a user in the r/TechSEO subreddit found Googlebot and Bingbot getting HTTP 403 responses on the site's sitemap after Cloudflare's AI bot blocking was switched on. The report is unverified. Search Engine Journal states directly that it has not been established whether the cause is a genuine technical issue or a user misconfiguration. Google's John Mueller replied in the thread and asked to be contacted directly so the case could be investigated.

What the report describes

The control named in the report is Cloudflare's AI Crawlers & Scrapers setting, and within it the Block AI Training option. According to the account published by Search Engine Journal on 4 August 2026, enabling that option produced 403 responses for Googlebot and Bingbot when they requested the site's sitemap. A 403 is a refusal by the server. When the refused file is the sitemap, the crawler loses the map it uses to reach everything else on the site.

That is the extent of the reported symptom. Cloudflare has not confirmed a fault, and the report does not present a reproduction on other sites. One thread is one data point.

Why a search crawler sits inside an AI setting at all

Cloudflare classifies Googlebot and Bingbot as Search + Training bots, meaning it treats them as crawlers that serve both search indexing and AI training. That classification is the link between an AI policy decision and organic search. A publisher blocking AI training is making a content licensing decision, and because the same crawler does both jobs, that decision has a route to indexing.

The 15 September 2026 date in Cloudflare's documentation

Cloudflare's documentation says that starting 15 September 2026, mixed-purpose crawlers that combine search and training functions will also be blocked by all configurations that block AI training. That published date is what makes the report unexpected: by Cloudflare's own schedule, this treatment was not due to begin until September, and the report is dated 4 August 2026.

The facts on record are short enough to list in full.

The 15 September 2026 date in Cloudflare's documentation
DetailWhat is on record
Setting named in the reportCloudflare's AI Crawlers & Scrapers control, Block AI Training option
How Cloudflare classifies Googlebot and BingbotSearch + Training bots, serving both search and AI training
Reported symptomHTTP 403 returned to Googlebot and Bingbot on a sitemap request
Documented start for blocking mixed-purpose crawlers15 September 2026, per Cloudflare documentation
Google's response in the threadJohn Mueller asked to be contacted directly so the case could be investigated

What has not been established

Search Engine Journal is explicit that it is not known whether the 403s came from a genuine technical issue or from a user misconfiguration. Nobody in the reporting has said Cloudflare shipped the September change early. Mueller's request to be contacted directly is a request for evidence, not a confirmation. What is on record is that one site owner saw something odd and Google wants to look at it.

How to check your own site

If the AI Crawlers & Scrapers setting is on for a domain you care about, the verification is cheap and needs nobody's approval.

  • Pull server or CDN logs, filter to Googlebot and Bingbot user agents, and look for 403 status codes. Start with the sitemap URL, since that is the file the report names.
  • Open Search Console, read the Sitemaps report for fetch errors, then check the Pages report for a rise in URLs that could not be crawled.
  • Find out who controls the Cloudflare account and when the AI blocking option was last changed. That change date is what you compare against any crawl anomaly.

If the timing lines up, escalate it with real log lines attached. Crawl status codes sit inside the standard scope of a technical SEO audit, so this class of problem tends to surface there without a special investigation. The broader question of which crawlers you want reading your content, and what blocking AI crawlers costs you in AI search visibility, deserves a decision of its own rather than a firewall toggle made in isolation.

What this means for Thai marketers

Cloudflare sits in front of a large share of Thai WordPress sites as the default CDN and WAF, and it is usually configured by the hosting provider or the developer rather than by the marketing team. A well-meaning decision to block AI bots can be taken by someone who never thinks to mention it to the site owner. That is the practical exposure here: not a confirmed Cloudflare fault, but a setting with search consequences that can be changed by a person outside the search conversation. If organic impressions drop without a content or ranking explanation, ask who touched the CDN and on what date. Keeping that line open between the developer and whoever runs SEO costs nothing and closes a real blind spot.

Frequently asked questions

Should I turn off Cloudflare's AI blocking right now?

No, not on the strength of one unverified report. Nothing has been confirmed as a Cloudflare fault, and blocking AI training crawlers is a deliberate content decision that many publishers have made for their own reasons. Read your logs first and act on what they actually show.

Is this a confirmed Cloudflare bug?

No. Search Engine Journal states that it is not established whether the 403s came from a genuine technical issue or a user misconfiguration, and the report contains no Cloudflare confirmation of a fault.

Why would blocking AI bots affect Googlebot?

Because Cloudflare classifies Googlebot as a Search + Training bot, treating it as a crawler that serves both search and AI training. A rule aimed at training therefore has a path to a crawler that also handles indexing.

What is scheduled for 15 September 2026?

Cloudflare's documentation says that from that date, mixed-purpose crawlers combining search and training functions will also be blocked by all configurations that block AI training. Anyone running the setting has until then to decide what they want it to do.

Has Google confirmed the blocking is real?

Google has not. John Mueller responded to the thread asking to be contacted directly so the case could be investigated. The source does not state any further Google position.

The useful action from this report is a small one: find out whether the AI blocking setting is on for your domain, then read the logs for Googlebot 403s on the sitemap. If you would rather have that checked by someone who reads crawl logs regularly, Relevant Audience can fold it into a technical review.

Antonio Fernandez

Antonio Fernandez

Founder and CEO of Relevant Audience. With over 15 years of experience in digital marketing strategy, he leads teams across southeast Asia in delivering exceptional results for clients through performance-focused digital solutions.

Share to:
Copy link: