On Tuesday a setting on your website changes, and nobody asked you about it.

That's 15 September, the day Cloudflare's new defaults for AI crawlers take effect. Cloudflare announced them on 1 July, so there's been ten weeks of notice, in a blog post most site owners had no particular reason to read.

Here's what flips. One switch for AI bots becomes three: Search, Agent and Training. From Tuesday, Training and Agent start out blocked on pages that display ads, and Search starts out allowed. Cloudflare's own wording is "For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads."

Who it happens to matters more than most of the coverage lets on, and the answer is filed in a strange place. The blog post everybody links to mentions new domains and stops there. The part about existing customers sits in Cloudflare's press release from the same day: "On September 15, 2026 these changes will also be made for all existing free customers that have not changed their settings by September 15, 2026 in their dashboard."

So it's new customers, new domains added by existing customers, and free zones where nobody has touched the settings. If you're on a paid plan and you configured your bot settings at some point, you keep what you chose and nothing flips under you. More than 20% of the web sits behind Cloudflare's network by its own count, and only a slice of that slice is in scope.

Angel Castro put that gap better than I would have.

There's a second thing most owners don't know, and it's why I'm writing this four days out rather than next month: Cloudflare classifies Googlebot as a crawler that does two jobs, Search and Training. Block Training, and on ad-carrying pages the block catches Googlebot too.

That's in the announcement, in Cloudflare's words, and it isn't a fault.

Three categories, and why the split is basically sensible

Start with what the three actually cover, because the names carry more than they look like they do.

Cloudflare describes Search as anything that "collects or indexes your content, so it can answer questions about it later". That's Googlebot reading your services page so it turns up when somebody searches for what you do.

Agent is "acting, usually in real time, on a person's behalf, to get something done right now". Someone asks ChatGPT who installs solar in Newcastle, and a bot goes off to read your page while they wait for an answer.

Training is "a crawler taking your content to train or fine-tune a model". Your words end up as weights inside something that will later answer that same question without sending anyone to you at all.

Three different transactions, and until now most owners had one lever, far too blunt for the choice they actually wanted to make.

Here's what convinced me the split is reasonable. Owners have been drawing this exact distinction by hand for months, with the crude tools they had. TechnologyChecker.io pulled Cloudflare Radar's robots.txt data in early September, a 31 August snapshot of 4,223 files from Cloudflare's top-domains sample, and for each crawler compared how many files name it under a disallow against how many under an allow. Bytespider sits at 5.80 to 1. CCBot at 4.08. ClaudeBot at 2.39, GPTBot at 2.33. Then look at the answering bots from those same companies: ChatGPT-User at 1.12, Googlebot at 1.11, and OAI-SearchBot at 0.94, named in more files as allowed than as blocked. It's Cloudflare's dataset, but the files in it were written by site owners, and the direction is unmistakable.

Same firms, opposite treatment, depending entirely on what the bot is for. Their summary of it: "Publishers are not blocking 'AI'. They are blocking training and allowing answering".

So Cloudflare isn't inventing a distinction. It's codifying one that site owners had already drawn for themselves.

Somebody is measuring the change as it happens, which I'll be reading with interest. Aleyda Solis flagged it, and hers was the only well-known SEO account I found measuring the thing rather than explaining it.

SeenSure is a commercial monitoring vendor selling a product for exactly this problem, so weigh their numbers accordingly, and the sample isn't a random slice of the web: 894 hosts pulled from Certificate Transparency logs plus 152 from a European agency directory. The baseline still says something worth knowing. On 2 September, before anything changed, "sites behind Cloudflare already refused AI crawlers roughly four to twelve times more often than sites that are not", depending on the crawler. For ClaudeBot it was 25.2% of the Cloudflare-fronted sites against 5.5% of the others.

One more piece of the default is worth having straight. The whole thing hinges on "pages that display ads", and Cloudflare hasn't said how it decides a page does. The blog post gives the reasoning, which is defensible enough: "An ad is a signal that a website owner meant for a person to land there and see it." What it doesn't give is the test.

One bot, two jobs

Googlebot crawls to index your pages for Search. It also crawls to feed Google's AI products. One user agent, two purposes, and no way for you to accept one and refuse the other at the door. Cloudflare puts Googlebot, Bingbot and Applebot in the mixed-use pile for exactly that reason.

Cloudflare is blunt about why it minds: "Google has access to about 2x more information than leading AI companies because Google leverages a mixed-use bot." The implication does the rest. You can't take part in Google's search without also taking part in Google's AI, because the same crawler does both jobs on the same visit.

That's a strong hand to hold when Google accounts for roughly 88% of referral traffic, by Cloudflare's own count. Refusing Googlebot isn't a setting. It's a business decision.

By Cloudflare's count, mixed-use crawlers make up over 36% of crawler activity, and it calls pure search crawling "a small and declining share". The mixed pile is the big pile.

The headline number Cloudflare led with needs correcting. At announcement, Cloudflare said "52% of crawler requests are now for AI training as of June 2026, up from 22% in Spring 2025." TechnologyChecker.io, which tracks the Radar series, then found that Cloudflare had reclassified it, moving a block of requests out of Training and into Mixed Purpose. They went back and re-queried it: "The 2026 high is 47.25% in June", and running the identical window again on 1 August returned 45.94% for training rather than the 52.3% originally published. By August the corrected series had training down to 40.15% and still falling. None of the traffic changed. The labels did.

I prefer the revised figure, because of where those six points went. They didn't disappear. They moved into mixed purpose, which is the precise category this whole change is built around.

Google's defence is legitimate and you should hear it. Google-Extended isn't a crawler at all; it's a robots.txt token, and Googlebot does the fetching either way. What the token does is let you keep your content out of training for future Gemini models, and out of grounding in Gemini Apps and the Vertex AI API. Google's documentation calls it "a standalone product token" and is unambiguous about the trade-off, or rather about the absence of one: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." That's a real control, it's free, and it's been there for a while.

Cloudflare's counter is about the bot, not the token: "Most leading AI companies separate discovery crawlers from training crawlers, making it relatively simple for publishers to enable content access for one purpose or the other. Google does not." Opting out of Google-Extended doesn't touch what Googlebot itself feeds, and you don't have to take Cloudflare's word for that. Google's own documentation on AI features says "AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search." Both statements are true at once. That overlap is exactly what the new defaults expose.

What nobody has done since 1 July is split a crawler. OpenAI had already separated its bots before any of this, GPTBot for training against OAI-SearchBot and ChatGPT-User for answering, and so had Anthropic, with ClaudeBot for training against Claude-SearchBot and Claude-User. Google still runs Googlebot as one bot doing two jobs, and I couldn't find any official Google statement responding to Cloudflare's mixed-use classification.

The argument itself got heated on Hacker News in late July, where the announcement drew 194 points and 158 comments. One commenter, dannyw, called Google's approach "manifestly predator, unfair, and IMO illegal". Another, sneak, made a point that cuts against Cloudflare just as hard: "There is no technical mechanism whereby it is actually possible to allow people to read your webpage and not use it for other things."

Two very large companies are arguing about crawler design. Free-tier zones are the ground they're arguing on.

Then people started hitting it

The mechanism didn't creep up on anybody. Cloudflare put it in the announcement, in one sentence: "Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training". Search Engine Journal ran the story on 2 July, the day after the announcement, under the headline "Cloudflare's AI Crawler Rules Can Block Googlebot": "Starting September 15, under Cloudflare's strictest-rule change, sites that block Training will also block combined crawlers like Googlebot, Applebot, and Bingbot." On Hacker News in late July, Simon Willison called it "the big news here" and quoted Cloudflare's paragraph in full.

Not a malfunction, then. A specification, on the record from day one.

The warning went out anyway and still didn't reach the people who needed it.

Jonathan Bird, a web developer in Brisbane, got called into a site in July whose organic traffic had, in his words, "fallen off a cliff". The cause: "The IT provider had enabled the option to block bots from crawling the site." Search Console showed clicks and impressions collapsing from around 21 July and staying down for a fortnight. Three weeks after the fix, the site was still only about halfway back.

He's anonymised both the client and the provider, and one detail matters for accuracy: the switch involved was the older one-click AI bot block rather than the new Training category. Same underlying rule, earlier interface. Somebody with admin access ticked a box that sounded sensible and took the site's Google visibility with it.

Then August. Someone posted to r/TechSEO under the title "Has anyone else run into Cloudflare AI Crawlers blocking Googlebot/Bingbot?" The poster, u/TheChiuaua, runs a multilingual pet-care blog, and their description was specific: "When I set AI Training = Block, both Googlebot and Bingbot start receiving HTTP 403 responses when trying to fetch my sitemap. As soon as I disable the AI Training block, the sitemap is accessible again."

They named a window, 2 August at 16:00 UTC through to the following afternoon, and listed Search Console's 4XX crawl errors across it, the posts sitemap among them, failing on a 403.

Now the obvious objection, because it's a good one. Plenty of people test crawler behaviour by spoofing a Googlebot user agent, and Cloudflare turns away fake Googlebots deliberately. u/reggeabwoy said exactly that: "Cloudflare isn't blocking Googlebot - it's blocking your test bot trying to spoof Googlebot".

The poster had pre-empted it in the post itself: "To clarify, this is not based only on a test that spoofs the Googlebot user-agent." Search Console had recorded the errors independently, and the Cloudflare dashboard itself, in their words, "appeared to show Googlebot and Bingbot as blocked".

And it wasn't just them. u/YaroslavSyubayev, in the same thread: "I disabled the block training option, because that blacklists Googlebot."

So that's three operators I can point to across two months: Bird's client in July, then two more in August. Cloudflare's own community forum adds a handful more. Since early July it has carried at least four threads whose titles say the same thing, verified Googlebot or Bingbot receiving 403s under the "Block AI training crawlers" rule, though the forum turns away automated browsers, so I've read the titles and not the threads. It's scattered rather than systematic, and I'm not going to dress it up as a wave.

John Mueller commented in the thread, and his words are worth reading precisely: "I'm happy to take a look at what's happening on Google's side if you want to DM me some details". That's a request for details, not a confirmation of anything.

Search Engine Journal came back to the subject on 4 August and hedged: "It's unclear whether this is a fluke or user error". That line gets passed around as though SEJ doubted the mechanism. It didn't. It had published the mechanism a month earlier, under a headline naming Googlebot. The open question was narrower, whether this particular operator's 403 responses were what they looked like, and at the time it was the right question to ask.

Which brings me to the reason all three cases could happen before the deadline. The most restrictive rule already applies today. Any site with a Training block switched on, including through the older one-click toggle, already turns away mixed-purpose crawlers. Tuesday doesn't invent the behaviour. It changes who gets it without choosing it.

What to check before Tuesday

Three questions. Ten minutes.

Which plan is the zone on, and how old is it? Free tier with no preference set gets the new defaults automatically, and so does any domain added from Tuesday onwards, whatever the plan. Paid, with settings already configured, does not, whatever the panicky version of this story says.

Do the pages carry ads? A brochure site, a services site, a SaaS marketing site with no ad units: the ad-page default barely reaches you. I'll say the honest version out loud. For a decent share of readers, Tuesday is a non-event.

What do the three settings say right now? Open the zone, then Security, then Settings, and find the AI crawler controls. For each of Search, Agent and Training you can "block on all pages, block only on pages that display ads, or choose not to block". Read them instead of assuming, particularly if somebody enabled a block for AI bots a year ago and never went back. One warning: Cloudflare's AI Crawl Control documentation, last updated in late July, still describes the old single-category model and doesn't mention the three categories. The dashboard is ahead of the docs.

I ran this over our own properties this week, expecting to report what our dashboard said. There wasn't one to report on.

webcoda.com.au returns Server: cloudflare with a Sydney CF-RAY. The rest of the headers say it's a HubSpot CMS site, and the Cloudflare is HubSpot's, not ours. Whatever the three settings say for that site, they aren't in a dashboard Webcoda can open. Then there's ai-checker.webcoda.com.au, which returns Server: Microsoft-IIS/10.0 and X-Powered-By: ASP.NET. It's on Azure with no Cloudflare anywhere near it, so Tuesday does nothing to it.

Two properties, one company. One sits behind Cloudflare with the switches in somebody else's hands; the other never had the switches. Neither is what I'd assumed before I looked, which is the argument for going and looking rather than reasoning about it from the outside.

If your broader question is whether you're accidentally shutting AI crawlers out altogether, we've covered that separately.

Modern digital illustration showing a website interface being approached by both vibrant human silhouettes and translucent digital AI agent silhouettes, representing dual-audience SEO.
Related Article10 min read

Your Website Has AI Visitors and You're Ignoring Them

AI bots visit your website every day. Most Australian businesses are accidentally blocking them. Here's how to check and why it matters now.

Read full article

Defaults are policy

The thing I keep coming back to isn't the Googlebot mechanism. That's classification working as specified, and Cloudflare said so in public ten weeks ago.

It's that a decision I'd have opinions about arrives as a default I didn't set.

I've been building and running websites for about twenty years. For most of that time, the settings that shaped what a site did were mine to get wrong. A slice of that now sits with whoever runs the layer in front of it, and their defaults are policy whether or not I read the announcement.

I don't think Cloudflare is wrong here. The category split beats the blunt switch it replaces, and the robots.txt numbers say owners wanted that distinction long before anyone offered it to them. But a better default and a decision of mine aren't the same thing. The gap between them is worth ten minutes before Tuesday.

---

Sources
  1. Cloudflare. "Your site, your rules: new AI traffic options for all customers." 01/07/2026. https://blog.cloudflare.com/content-independenc...
  2. Cloudflare. "Content Independence Day, one year on: building the business model for the agentic Internet." 01/07/2026. https://blog.cloudflare.com/agentic-internet-bo...
  3. Cloudflare. "Cloudflare allows the agentic Internet to flourish with a simple philosophy: your content, your rules." Press release. 01/07/2026. https://www.cloudflare.com/press/press-releases...
  4. Cloudflare. Changelog: new AI traffic options. 01/07/2026. https://developers.cloudflare.com/changelog/pos...
  5. Southern, Matt G. "Cloudflare's AI Crawler Rules Can Block Googlebot." Search Engine Journal. 02/07/2026. https://www.searchenginejournal.com/cloudflares...
  6. Hacker News. Discussion thread on Cloudflare's new AI traffic options, 194 points and 158 comments. 25/07/2026. https://news.ycombinator.com/item?id=49052564
  7. Bird, Jonathan. "Cloudflare Blocking Googlebot: How to Find and Fix the 403 (2026)." 20/08/2026, updated 07/09/2026. https://jonathanbird.com.au/blog/cloudflare-blo...
  8. u/TheChiuaua and others. "Has anyone else run into Cloudflare AI Crawlers blocking Googlebot/Bingbot?" r/TechSEO. August 2026. https://www.reddit.com/r/TechSEO/comments/1ve92...
  9. Montti, Roger. "Report That Cloudflare AI Bot Blocking Prevents Googlebot From Indexing Sites." Search Engine Journal. 04/08/2026. https://www.searchenginejournal.com/report-that...
  10. Cloudflare Community. "Googlebot blocked by 'Block AI training crawlers' before September 15", one of at least four threads since July 2026 reporting verified Googlebot or Bingbot receiving 403s under that rule (also threads 938043, 941022 and 944091). Titles read via search; the forum blocks automated browsers. https://community.cloudflare.com/t/googlebot-bl...
  11. TechnologyChecker.io. Robots.txt AI crawler blocking report, analysing Cloudflare Radar's top-domains robots.txt data. Snapshot of 4,223 files dated 31/08/2026. https://technologychecker.io/blog/robots-txt-ai...
  12. TechnologyChecker.io. AI crawler statistics, including analysis of the reclassified Cloudflare Radar crawl-purpose series. Updated 03/09/2026. https://technologychecker.io/blog/ai-crawler-st...
  13. SeenSure. "September 15" baseline study of 1,046 websites. 04/09/2026. https://seensure.com/research/september-15
  14. Google. Google-Extended, in Google's crawler and fetcher documentation. Accessed 10/09/2026. https://developers.google.com/crawling/docs/cra...
  15. Google. "AI features and your website." Google Search Central documentation. Accessed 10/09/2026. https://developers.google.com/search/docs/appea...
  16. OpenAI. "Overview of OpenAI crawlers." Accessed 10/09/2026. https://developers.openai.com/api/docs/bots
  17. Anthropic. "Does Anthropic crawl data from the web, and how can site owners block the crawler?" Accessed 10/09/2026. https://support.claude.com/en/articles/8896518-...
  18. Aleyda Solis (@aleyda). Post on the SeenSure baseline study of 1,046 websites. 09/09/2026. https://x.com/aleyda/status/2097652258431242407
  19. Angel Castro (@tweetangelc). Post on which Cloudflare customers the new defaults reach. 08/09/2026. https://x.com/tweetangelc/status/20971280710651...