Cloudflare Disallow AI Training Separates Search Visibility From Model Training
Cloudflare has launched Disallow AI Training, a new control designed to preserve search indexing while refusing model training by the same crawler. Until now, publishers often faced a blunt choice when dealing with mixed-use bots: accept both purposes or block the crawler completely.
The change covers Applebot, Googlebot, and Bingbot, which Cloudflare classifies as mixed-use crawlers because each can support search and AI-related uses. Cloudflare now calls these bots “Accountable” when their operators provide controls, reporting, and assurances that a training opt-out will not reduce traditional search visibility.
That designation establishes a shared operating model among Cloudflare, Apple, Google, and Microsoft. Yet it does not create a binding technical standard. The new system combines dashboard controls, robots.txt instructions, company commitments, and Cloudflare’s enforcement against other training crawlers.
This is the central tradeoff. Publishers gain a simpler way to express consent without disappearing from search, but much of the protection still depends on crawler operators honoring that choice.
Cloudflare Disallow AI Training Changes the Default Choice
Cloudflare is turning one overloaded blocking decision into separate choices for search, training, and user-directed agents.
Cloudflare’s AI crawler controls classify automated activity by behavior rather than treating every AI-related request alike. Search crawlers build indexes, training crawlers collect material for model development, and agents retrieve pages while acting for a user.
Those categories matter because the economic effects differ. A search engine normally shows links that can send a person to the original website. A training process can absorb information without producing an immediate visit. An agent might retrieve a page and deliver its contents without the user ever opening it.
A crawler can fall into more than one category. Googlebot, Applebot, and Bingbot are important examples because their search functions make them difficult for publishers to block. Removing their access can eventually affect indexing, freshness, and discoverability.
Cloudflare’s earlier controls accommodated that problem by excluding mixed-use crawlers from some training blocks. That protected search visibility, but it also left publishers without a direct way to refuse the training component through the same setting.
The new crawler control model introduces Disallow AI Training as the middle option. It publishes a no-training preference through robots.txt, keeps Accountable mixed-use crawlers available for search, and blocks other crawlers associated with training.
Cloudflare says training-only crawlers operated by Amazon, Anthropic, Meta, and OpenAI can be blocked without affecting their separate search crawlers. Those companies use different bots for different purposes, making enforcement more straightforward at the network level.
Cloudflare is also changing what its stronger settings mean. “Block” and “Block on pages with ads” now apply to mixed-use crawlers, including Googlebot, Applebot, and Bingbot. Selecting either option can therefore affect search as well as training.
That distinction makes configuration more consequential. Disallow AI Training communicates a limited preference while preserving search access. Block denies the crawler itself, regardless of whether a particular request supports search or training.
Existing configurations are being migrated toward the new controls. A domain that previously used a general AI blocking option will typically retain search access while moving its training preference to Disallow AI Training. Domains with existing granular policies will also have their practical selections carried forward.
For new ad-supported domains, Cloudflare recommends allowing search, disallowing training, and blocking agents on pages where ads appear. New domains without advertising receive a less restrictive recommendation that allows all three categories.
These presets are recommendations, not permanent rules. Site owners can alter them during onboarding or later through Security Settings. The controls are available across Cloudflare plans and operate at the domain level.
The result is a clearer decision tree. A publisher can permit ordinary indexing, reject training, and choose a separate policy for agents. That structure better reflects how automated systems now interact with websites.
It also makes incorrect configuration easier to diagnose. If a publisher selects Block and later loses crawler access, the consequence follows directly from the chosen setting. Disallow AI Training is intended for the narrower objective of keeping search while refusing model development use.
This is not merely a renamed bot switch. It changes the unit of control from crawler identity alone to the combination of identity, disclosed purpose, and operator behavior.
Why Mixed-Use Crawlers Put Publishers Under Pressure
The conflict exists because a crawler that delivers valuable search traffic can also collect material for an entirely different commercial purpose.
The open web’s traditional exchange was relatively easy to understand. Publishers allowed search engines to crawl their pages, and search engines returned links, excerpts, and potential visitors. Advertising, subscriptions, sales, and reader relationships depended on some of those users arriving.
Generative AI complicates that exchange. A model can use collected material during training, while an AI answer can satisfy a query before the reader visits any cited source. Search, model development, and answer generation therefore create different forms of value.
Cloudflare’s own measurements illustrate why publishers are concerned. The company reported that training represented 80 percent of classified AI crawling over a 12-month period. During the following six-month view, training’s share rose to 82 percent, while search represented 15 percent and user actions represented 3 percent.
Those figures describe traffic observed and classified by Cloudflare, not the entire web. They still show that training activity can dominate the automated demand reaching content providers.
A publisher that blocks a training-only crawler faces a manageable calculation. The block can stop the unwanted collection without removing a major search indexer. Distinct bots such as GPTBot and OAI-SearchBot make those purposes easier to separate.
Mixed-use crawlers create a harder problem. If the same crawler supports search and training, an infrastructure-level block cannot determine what will happen after the content reaches its operator. Blocking access protects the material but also removes the search function.
Allowing access preserves discovery, yet it requires another mechanism to restrict downstream use. That is the gap the Cloudflare Disallow AI Training setting attempts to close.
Cloudflare began addressing this gap with Content Signals, a proposed vocabulary placed in robots.txt. The content signal policy distinguishes three declared uses: search, ai-input, and ai-train.
The search signal covers indexing and traditional search results. It does not include AI-generated summaries. The ai-input signal concerns real-time use by AI systems, including retrieval and grounding. The ai-train signal addresses model training and fine-tuning.
A site can therefore publish search=yes and ai-train=no. It can leave ai-input unspecified if its owner has not decided how generative answers should use the content.
This separation is important because a missing preference should not be interpreted as permission or refusal. Cloudflare’s policy treats an omitted signal as neutral rather than guessing what the publisher intended.
However, Content Signals are expressions of preference. They are not barriers that physically prevent a scraper from downloading a page. Cloudflare has previously advised publishers to combine such signals with bot controls or firewall rules when technical enforcement is required.
The new Accountable designation tries to bridge those layers. It identifies crawler operators that Cloudflare says provide, or have committed to provide, specific controls and transparency. The requirements include a training opt-out, an AI-summary opt-out, URL-level visibility, and protection for traditional search ranking.
Apple, Google, and Microsoft meet that threshold through different combinations of current features and time-bound commitments. The designation does not mean their implementations are identical. It means Cloudflare believes each operator has accepted the same basic responsibilities.
That creates pressure on other crawler operators. A company that wants broad access can now be compared against a published baseline for consent, inspection, and search neutrality. Separating bot identities remains one way to meet that baseline, but it is no longer the only route.
Publishers also face a new operational responsibility. Search visibility, AI training, answer generation, and agent access now require separate policies. A single “block AI” decision no longer captures the business tradeoffs.
A news site funded by page views might reject both training and agent access on advertising pages. A retailer could value qualified AI referrals even if their volume is lower. A documentation site might welcome real-time AI retrieval while refusing long-term model training.
The relevant question is no longer whether AI bots are good or bad. It is which use justifies access, what value returns to the publisher, and whether that use can be verified.
Apple, Google, and Microsoft Share a Model, Not One Implementation
The three companies support the same principle, but their controls remain technically uneven and arrive on different timelines.
Apple already lets publishers address training through Applebot-Extended. A site owner can disallow that user agent in robots.txt while continuing to permit Applebot’s ordinary search functions.
Apple’s documentation says the Applebot-Extended preference does not affect how a site appears in search results. It also supports page-level mechanisms related to generative output, including nosnippet and labels for paywalled material.
Those tools do not yet provide every element in Cloudflare’s Accountable framework. Cloudflare says Apple lacks URL-level inspection for this purpose. Apple has reportedly shared details of an in-progress solution expected next year.
The current Applebot controls therefore provide a functioning separation between search and training, but incomplete visibility into what happened after access. Cloudflare is accepting a commitment to close that gap.
Google uses a similar extension model. Publishers can block Google-Extended without blocking Googlebot. Google-Extended is a control token that governs certain generative AI uses rather than a separate crawler that always makes its own requests.
That detail matters. Googlebot might still retrieve the content for search, while Google uses the Google-Extended preference to determine whether the material can support covered AI systems. The publisher controls downstream use through policy rather than a separate network identity.
Google says opting out through Google-Extended does not affect inclusion or ranking in Google Search. Its training opt-out guidance documents the relationship between Googlebot and Google-Extended.
Google also provides search performance reporting and controls related to generated search experiences. Cloudflare says Google is working on additional URL-level transparency associated with Google-Extended, with a launch expected within weeks.
Microsoft is the least complete of the three under this specific mechanism. Bing supports granular webmaster controls, and publishers can use the NOARCHIVE meta tag to limit certain uses of cached or displayed content.
Microsoft says NOARCHIVE does not remove a page from search ranking. Site owners can also use Bing Webmaster Tools for content removal and URL management.
However, Bingbot does not yet automatically honor Cloudflare’s domain-level no-training preference through robots.txt. Cloudflare says Microsoft is developing that capability for early 2027.
Until then, selecting Disallow AI Training does not automatically communicate the intended restriction to Bing through the new workflow. Publishers seeking an immediate Bing restriction still need to use Microsoft’s existing tools and page-level metadata.
Microsoft has described its AI control options as a way to preserve search discovery while limiting how content appears in generative experiences. Yet Cloudflare’s unified switch is not fully connected to those controls today.
This implementation gap is the most important caveat in the launch. Cloudflare presents Applebot, Googlebot, and Bingbot under one Accountable label, but only Apple and Google currently expose the specific extended-user-agent route for training preferences.
Microsoft’s inclusion rests partly on a dated commitment. That may be reasonable for establishing a cooperative standard, but publishers should understand the difference between available enforcement and promised compatibility.
The shared model also depends on each operator’s interpretation of training. Pretraining a new model, fine-tuning an existing system, grounding a live answer, and generating a search summary are separate activities. A “no training” choice does not necessarily reject all AI-mediated use.
Cloudflare explicitly treats AI input and AI summaries as different questions. That prevents one preference from silently covering unrelated uses, but it also means the new control is narrower than its simple dashboard label might suggest.
A publisher can reject model training while remaining eligible for AI-generated search summaries. Another publisher can use operator-specific summary controls while allowing training. These choices can produce different traffic and attribution outcomes.
The Accountable framework is therefore best understood as a minimum contract. It asks operators to separate purposes, honor preferences, provide inspection, and avoid punishing a training refusal in traditional search.
It does not make Apple, Google, and Microsoft technically interchangeable. Nor does it guarantee that every AI feature offered by those companies falls within the same opt-out.
The immediate value comes from consolidation. Cloudflare customers receive one place to express a common preference, while the operators map that preference to their existing or forthcoming controls.
The long-term value depends on whether those mappings become transparent enough for a publisher to audit at the URL level.
The Setting Is a Consent Signal, Not Proof of Compliance
Cloudflare has simplified the instruction, but it cannot prove every downstream use stopped merely because a site published that instruction.
This limitation begins with robots.txt. The file was designed as a voluntary crawler protocol, not an access-control system. Compliant crawlers read it and adjust their behavior. An operator that ignores it can still request publicly available pages unless another control blocks the traffic.
Cloudflare can enforce decisions at its network edge when it recognizes a crawler. That makes blocking stronger than a preference signal. Yet enforcement depends on reliable identification, and user-agent strings can be copied by unrelated bots.
Verified bot programs reduce that risk by checking request sources against information supplied by operators. Cryptographic authentication could offer stronger proof, but adoption remains incomplete across the crawler market.
Even a verified request reveals who retrieved a page, not necessarily every later use of its contents. A crawler operator must maintain internal separation between search indexing, model training, answer generation, and other processing.
Cloudflare’s Accountable requirements address that trust gap through reporting and commitments. URL-level visibility should help publishers see which pages became available for training and how content appeared in search.
The key word is “should.” Apple’s inspection system remains in development, Google’s additional tools are forthcoming, and Microsoft’s domain-level robots.txt support is targeted for early 2027.
The designation therefore combines current capabilities with future promises. Cloudflare is not claiming that all four requirements have identical production implementations today.
Publishers should also avoid reading “Disallow AI Training” as a universal legal resolution. Copyright exceptions, contractual terms, jurisdictional differences, and past collection remain separate issues. A new preference cannot retroactively remove material from an existing model.
The setting governs prospective crawler behavior as implemented by participating operators and Cloudflare. It does not confirm that previously collected copies have been deleted. It also does not establish how a trained model might retain or reproduce information.
Another uncertainty concerns classification. Cloudflare assigns Search, Training, and Agent behaviors based partly on operator disclosures and other observed information. A crawler’s stated purpose can change, and one service can support several products.
When classifications lag behind product changes, a policy might permit more activity than a publisher expects. Transparent change logs and independent monitoring will matter as much as the initial dashboard design.
AI summaries expose an even larger gap. Training determines whether content contributes to model development. Summaries determine whether current content is transformed into an answer that can reduce the need for a visit.
Cloudflare cites research showing that AI summaries are already common in search behavior. A search behavior study found that users were less likely to click result links when an AI summary appeared.
That is not the same problem as training. A publisher could successfully reject training yet still lose visits when a search product summarizes freshly indexed material.
Cloudflare says Accountable operators must offer an AI-summary opt-out directly and eventually through Cloudflare. Its next goal is more granular control over how much content a summary can include.
That plan recognizes a weakness in binary consent. Allowing a short quotation with a clear link is different from allowing a detailed answer that replaces the source. Both might technically count as summary use.
Business models also change the acceptable balance. An advertising publisher needs visit volume because impressions support revenue. A retailer might accept fewer visits if AI referrals produce more purchases. A subscription publisher might value attribution and reader recognition more than raw clicks.
Cloudflare has cited third-party estimates suggesting that AI referrals can convert at higher rates than traditional search referrals. Those figures vary by dataset and methodology, so they should not be treated as a universal offset for lost traffic.
The real measurement problem is causal. A publisher must know which crawler accessed a URL, which product used it, whether a summary appeared, how much content it displayed, and whether the interaction produced a visit.
Most organizations lack that complete chain today. Server logs reveal requests, while search consoles reveal impressions and clicks. Neither alone establishes how content flowed through an AI product.
The new controls improve agency before they deliver full accountability. They let publishers declare a narrower policy and apply stronger blocks to crawlers that do not qualify for the Accountable exception.
They do not eliminate the need for monitoring. Publishers should inspect search coverage, crawler logs, referral patterns, and the public outputs of major AI products after changing their settings.
For teams that maintain research, documentation, or institutional memory, crawler policy is only one layer of information governance. A searchable knowledge base can preserve source context internally even when external platforms summarize the public version.
The skeptical conclusion is straightforward. Cloudflare has created a credible control surface, but compliance remains a system of technical enforcement, voluntary standards, operator policy, and future transparency.
Calling a crawler Accountable raises the expected standard. It does not make the underlying trust problem disappear.
Three Signals Will Show Whether the New Model Works
The next test is whether Cloudflare’s shared policy produces measurable behavior, not whether more companies endorse its language.
The first signal is Microsoft’s promised support for a domain-level no-training preference in robots.txt. Cloudflare says that capability is targeted for early 2027.
A working implementation would close the largest current gap among the three mixed-use crawler operators. It would let the same Cloudflare setting communicate a training refusal to Bing without requiring a separate NOARCHIVE deployment or removal workflow.
A delay would weaken the Accountable designation because one of its most important participants would still rely on manual or page-level alternatives. The implementation should also clarify which Microsoft AI uses fall within “training” and which remain governed by separate controls.
The second signal is URL-level reporting from Apple and Google. Publishers need more than a confirmation that a domain preference exists. They need to know which pages were accessed, which uses were permitted, and whether the preference changed later processing.
Google’s promised additions related to Google-Extended provide an early test. Apple’s planned inspection capability provides a longer one. Useful reporting should be specific enough to compare crawler access with search performance and AI visibility.
A generic dashboard count would offer limited accountability. Page-level records, understandable purpose labels, and stable historical data would strengthen Cloudflare’s claim that publishers can make informed decisions.
The third signal is Cloudflare’s work on AI-summary controls. Training opt-outs solve only one part of the publisher conflict. Search-generated answers can affect traffic even when no model-training permission exists.
Cloudflare’s planned control over how much content appears in a summary is more ambitious than a simple opt-out. It would require operators to interpret a shared preference consistently and expose enough data for publishers to evaluate the result.
Success would strengthen the broader principle behind Cloudflare Disallow AI Training: access should be purpose-specific, measurable, and changeable by the content owner. Failure would leave publishers navigating separate controls across every search and AI provider.
The practical response today is to treat the launch as a policy upgrade, not a set-and-forget guarantee. Site owners should review the migrated settings for each domain, especially if they previously enabled a broad AI blocking option.
They should confirm whether Search remains allowed and whether Training now shows Disallow AI Training rather than Block. A full Block selection can stop Applebot, Googlebot, and Bingbot, which creates a different outcome from publishing a no-training preference.
Teams should also document why each category is allowed or refused. Search, training, and agents serve different purposes, so the policy should reflect the site’s revenue model and audience relationship.
After any change, monitor crawler responses and indexing. A fall in search coverage might indicate that the wrong setting was selected or that another firewall rule is overriding the preference.
The same review should cover important subdomains. Documentation, support centers, blogs, and application pages can sit behind different configurations even when they share a parent brand.
Cloudflare’s launch matters because it replaces an artificial binary with a more realistic choice. A site should not have to donate material for model development merely to remain visible in a conventional search index.
Still, the system’s credibility will come from verifiable outcomes. Microsoft must complete its integration, Apple and Google must deliver useful inspection, and Cloudflare must turn summary control into something publishers can measure.
For now, Cloudflare Disallow AI Training gives website owners a clearer instruction and a safer middle path. The next question is whether the largest crawler operators make that instruction observable enough to trust.
Review your domain’s three crawler policies, record the intended outcome, and watch search coverage after any change. If traffic remains stable while training access declines, the shared model has passed its first practical test.



