Cloudflare’s AI crawler rules change today, September 15, 2026, but the important story is not simply that Cloudflare is “blocking AI bots.”
The real change is more specific: Cloudflare is separating automated traffic into Search, Agent and Training categories, giving website owners more control over what different kinds of crawlers can do. For publishers, the distinction could matter just as much as the block itself because search crawling and AI training do not have the same value to a website owner.
That creates a practical question for anyone running a blog, business site or publication: Can you stop AI training without also cutting off the search crawlers that help people find your content?
Cloudflare’s current system is designed to make that choice more precise, although it is not equally simple for every crawler yet.
What changed on September 15?
Cloudflare now lets customers manage AI-related crawler traffic by behavior rather than relying on one broad “AI bot” decision.
The three categories are:
Search: crawlers that collect or index content so it can be discovered or used to answer questions later.
Agent: automated activity acting in real time on behalf of a person, including chat-fetch systems and browser-use agents.
Training: crawlers that take website content to train or fine-tune AI models. Cloudflare says this category can also include mixed-purpose crawlers that perform both Training and Search.
Each category can be configured with three basic choices: Block on all pages, Block on pages with ads, or Allow.
So when you hear that Cloudflare’s AI crawler rules are changing today, think granular control, not one giant switch.
Why mixed-use crawlers are the difficult part
The hardest case is a crawler that does more than one job.
Imagine you run a technology blog. You want Google to crawl your articles so they can appear in Google Search. At the same time, you may not want that same crawling infrastructure to make your content available for AI-training purposes.
If one crawler represents both activities, a blanket block can create an unwanted tradeoff: you may stop the AI-related use, but you may also interfere with search crawling.
Cloudflare specifically identifies Googlebot, Applebot and Bingbot as mixed-purpose crawlers in its current explanation of the problem.
That is why Cloudflare's new model matters.
The question is no longer only:
“Do I trust this bot?”
It becomes:
“Which job is this bot performing, and which of those jobs am I willing to allow?”
Blocking a crawler is not the same as disallowing AI training
This is probably the most important distinction for website owners to understand.
Cloudflare's documentation says that a full Block can apply to mixed-use crawlers, including those used for both Search and Training. That means a strict block can affect access that is useful for discovery.
Cloudflare is also developing a Disallow Training approach through Bot Preference Sync. In that model, Cloudflare can publish the relevant preference in robots.txt while cooperating mixed-use crawlers can continue to access the site for search indexing.
Here's a simple example.
Suppose your site publishes smartphone reviews. You want:
Search: allowed
AI training: not allowed
AI agents: allowed
A broad crawler block could be too aggressive because it may interfere with search access. A training-specific preference is designed to communicate a narrower choice.
The catch is important: robots.txt is a preference mechanism, not a technical access-control system. Cloudflare's own documentation and earlier work on AI crawler control make that distinction clear. A crawler operator has to honor the signal for it to work as intended.
That means you should not confuse “I published a Disallow rule” with “the crawler is technically incapable of accessing my content.”
Googlebot makes the distinction especially important
Google provides its own mechanism called Google-Extended.
Google says Google-Extended is a robots.txt control that publishers can use to manage whether content Google crawls may be used for training future generations of Gemini models and for certain grounding uses. Google also says using Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal in Search.
That creates an important distinction between Google's systems.
If you block Googlebot, you are dealing with Google's Search crawling.
If you restrict Google-Extended, you are controlling a separate use of content associated with Google's AI systems.
So a website owner should not treat “Googlebot” and “Google-Extended” as interchangeable names for the same crawler.
They serve different control purposes.
What does Cloudflare mean by “Accountable” crawlers?
Cloudflare has also introduced an Accountable concept for crawler operators that provide greater transparency and controls around how their crawlers use publisher content.
According to Cloudflare, operators handling mixed Search-and-Training use can receive this treatment when they provide mechanisms such as a way to honor a no-training preference, controls around AI summaries, visibility into which URLs were made available for training, and evidence about search usage. Cloudflare says operators that do not meet its transparency requirements do not receive the same benefit when a site owner disallows training.
The underlying idea is straightforward.
If a crawler wants broad access to published content, Cloudflare wants the operator to explain what it is doing and give publishers meaningful choices.
That is less about “good bots versus bad bots” and more about identifiable purpose and verifiable behavior.
What happens to AI-only training crawlers?
This is generally easier to manage than mixed-use traffic.
Cloudflare says its system can distinguish AI-training crawlers from other categories, and its broader framework is designed to let publishers make separate decisions for Search, Agent and Training traffic.
For a publisher, the practical difference is significant.
An AI-training crawler may consume your article to contribute to model development without serving the same direct referral function as a search engine.
A search crawler, by contrast, exists specifically to make content discoverable.
An agent can sit somewhere in between from a publisher's perspective: it may fetch an article to answer a user's question without necessarily producing a traditional search click.
Those are three different relationships with the same webpage.
Cloudflare's policy framework is an attempt to let website owners treat them differently.
What about the September 15 defaults?
Cloudflare's current documentation says the updated defaults taking effect on September 15, 2026 apply to new domains: Training and Agent traffic are blocked on pages that display ads, while Search remains allowed.
Cloudflare's earlier July announcement described the same direction and said multi-purpose crawlers combining Search and Training would be subject to the more restrictive applicable behavior when a customer chooses to block training.
This detail matters because it is easy to misread the change as “every Cloudflare website is being switched to the same AI-blocking configuration today.”
That is not what the current documentation says.
If you already use Cloudflare, the right move is to check your actual AI bot policy settings rather than assume the September 15 date automatically changed your entire site.
Existing sites should check their settings, not guess
For an existing website, there is a simple practical rule: inspect the configuration that Cloudflare is actually applying to your domain.
Cloudflare currently exposes these settings under its AI bot policy controls, where Search, Agent and Training can each be configured separately.
That matters because two sites can use Cloudflare and still have very different crawler behavior.
Common mistake: treating the old “Block AI Bots” switch as a complete policy
Why it happens: The old setting was easy to understand: block AI bots or don't.
Better approach: Think in terms of the three current behaviors—Search, Agent and Training—and decide separately what you want each category to do.
Cloudflare says the legacy Block AI bots option is being deprecated on September 15, 2026.
The modern configuration is more granular precisely because “AI bot” is no longer one clearly defined type of traffic.
What does this mean for a site like a technology blog?
Let's say you run a technology publication and publish an original article explaining a new AI model.
You probably want Google to discover that article.
You may also want an AI assistant to find it when a reader asks a question and the system can send a useful referral.
But you may not want the article to become training data for another model without your permission or without a relationship that provides value to the publisher.
That is exactly where purpose-based crawler controls become useful.
The bigger question is no longer simply whether the web should remain open.
It is what kind of access creates value for the person publishing the content?
For ad-supported sites, that question becomes even more important because a human visitor and a crawler can consume the same webpage while producing very different economic outcomes.
Why Cloudflare is making this change
Cloudflare's stated position is that the economics of web publishing are changing as AI systems become a larger part of how people discover and consume information.
In its July announcement, Cloudflare said automated agents and bots accounted for more than half of all web requests in its own measurements and argued that publishers need more control over AI access to their content.
Cloudflare's later reporting said that, in its measurements for June 2026, 52% of crawler requests were for AI training, up from 22% in spring 2025, while mixed-use crawlers represented more than 36% of activity. Those are Cloudflare's network observations, not a universal measurement of the entire internet.
The exact percentages are less important for a website owner than the underlying shift.
Crawling is no longer mainly about building traditional search indexes.
A growing share of automated access is connected to AI training, AI answers and agentic interactions.
That changes what a publisher needs to know about the traffic reaching its pages.
Cloudflare is also moving toward AI traffic measurement
Cloudflare is not positioning the September 15 changes as only a blocking feature.
The company has introduced tools such as BotBase and Attribution Business Insights to help customers understand which bots are interacting with their content and, in some cases, how that activity relates to referrals.
Cloudflare has also discussed making AI crawling more efficient by reducing unnecessary re-fetching and improving visibility into how content is consumed and cited.
That points to a larger trend: publishers are likely to care increasingly about identity, purpose, frequency, attribution and value, not just raw request counts.
What website owners should actually do today
You do not need to panic because September 15 has arrived.
Instead, work through the problem in a practical order.
First, check whether your domain is using Cloudflare's current Search / Agent / Training policy controls.
Second, decide what matters most to your site.
If search traffic is important, be careful with broad blocking rules that can stop a mixed-use crawler entirely.
Third, look at your robots.txt strategy.
If you use a preference designed to limit AI training, understand which crawler operators support that preference and remember that robots.txt communicates intent; it is not an enforcement mechanism by itself.
Fourth, do not assume that a Google-related AI control automatically controls Google Search.
Google explicitly says Google-Extended does not affect Search inclusion or Search ranking. Controls for what appears in Google's AI features in Search are a separate issue and can involve controls such as nosnippet, data-nosnippet, max-snippet or noindex, depending on the outcome you want.
Troubleshooting: “I blocked AI training, but I’m worried about losing Search traffic”
Problem: You want to restrict AI-training use without disappearing from search.
Why it happens: Search and Training can sometimes be associated with the same mixed-use crawler.
What to check: Whether you are using a broad crawler block or a more specific training preference, and whether the relevant crawler operator supports that preference.
How to fix it: Avoid assuming that one setting controls every use case. Review Cloudflare's AI bot policy, your robots.txt, and the crawler-specific controls provided by search platforms such as Google.
The bigger issue is control over the web's content economy
Cloudflare's September 15 policy change is really about something bigger than bot management.
The web was built around a relatively simple exchange: websites publish information, crawlers organize it, and users arrive through links.
AI introduces several new forms of access.
A model can train on content.
A search system can index it.
An AI answer engine can summarize it.
An agent can retrieve it for a user.
All four activities can touch the same webpage, but they do not necessarily provide the publisher with the same value.
That is why Cloudflare's decision to separate Search, Agent and Training is more significant than a simple anti-bot feature. It reflects an emerging attempt to create a more precise permission model for AI-era web traffic.
What remains uncertain
The system is not finished.
Cloudflare's own documentation shows that crawler support and behavior are not perfectly uniform. Microsoft, for example, is described in Cloudflare's material as having existing controls while additional robots.txt support is still being developed.
There is also a fundamental limitation to all preference-based systems: the usefulness of a robots.txt instruction depends on whether the crawler operator respects it. Cloudflare itself has described content signals as expressions of publisher preference rather than technical countermeasures against scraping.
So the industry is still working out the rules.
The technology is changing faster than the conventions that govern access to published content.
The real significance of Cloudflare's AI crawler rules
The most important thing to understand about today's change is simple:
Cloudflare is moving away from treating AI crawling as one thing.
Search, training and agents can now be treated as different categories with different policies. That gives website owners a more useful choice than a single all-or-nothing AI block.
For publishers, that distinction could become increasingly important as search engines and AI systems compete for access to the same information.
Your website may need Google to crawl your content.
You may want some AI systems to discover it.
You may want to restrict training.
And you may have a completely different view of an autonomous agent acting on behalf of a user.
Those are not contradictory positions.
They are different permissions.
Cloudflare's September 15 changes do not settle the larger debate over who should be allowed to use public web content, or under what business terms. But they do make one principle more explicit: website owners increasingly need to decide not only who can reach their content, but what that access is actually for.
0 Comments
Have a question, feedback, or something to add? Share your thoughts below.