top of page

Cloudflare Rolls Out Granular AI Crawler Controls

Cloudflare announced new controls that let website owners choose how different AI systems interact with their pages. The update replaces simple block-or-allow switches with three distinct categories for automated visitors.

Site operators can now label traffic as search, agent, or training. Each label carries its own default behavior and can be overridden per page or per path. The change also includes a toggle that stops any crawler from reaching ad-supported sections of a site.

The Old Binary Rule No Longer Fits

Until this week Cloudflare offered a single AI crawler block that treated every automated visitor the same. Many publishers found the approach too blunt. Search engines still needed access for indexing, while training runs consumed bandwidth without returning traffic.

The new system splits automated requests into three buckets. Search crawlers retain normal indexing access. Agent crawlers, which act on behalf of users, receive limited permission that ends once the user session closes. Training crawlers face the strictest default: they are blocked from any page unless the owner explicitly allows them.

How the Three Categories Work in Practice

When a request arrives, Cloudflare checks the user agent string and any declared purpose header. The edge rules then apply the matching policy. Owners can change the policy for any path without touching the rest of the site.

Ad revenue pages receive an extra layer. A single dashboard switch marks every advertising template so that training crawlers cannot fetch those URLs even if a broader rule would otherwise permit them. The setting does not affect search or agent traffic.

Why Publishers Asked for More Control

Large language model operators continue to gather data at scale, a trend driven by surging demand for training datasets across the AI industry and heightened debates over web content ownership. Cloudflare Blog notes that training runs often arrive from the same IP ranges that also serve user-facing agents. A single on-off switch forced publishers to choose between losing search visibility or accepting unlimited training access. Similar patterns appear in coverage from outlets tracking AI-data conflicts.

The new labels let each business set different priorities. News sites can keep search open while blocking training. Documentation sites can allow agent access for logged-in customers without opening the same pages to bulk downloads.

Early Publisher Reactions

Several large media companies tested the controls in closed beta. One operator reported that training requests dropped by more than half after the ad-page toggle was enabled. The publisher noted, “The granular options finally let us protect ad inventory without cutting off search referrals.” A second beta participant, a mid-sized news outlet, observed a 40% reduction in unwanted agent sessions after applying path-level rules, preserving referral traffic while curbing bulk scraping. Smaller sites have fewer resources to monitor logs. For them the default blocks on training crawlers provide immediate relief without extra configuration. One independent publisher described seeing “training requests vanish from the weekly summary within days of enabling defaults,” confirming the dashboard’s category breakdowns matched their server logs.

Limits That Remain

The system still depends on accurate identification. A crawler that mislabels itself can bypass the intended rules. Cloudflare says it will update its detection lists monthly, yet perfect coverage is impossible. Models such as OpenAI’s GPTBot, Anthropic’s Claude crawler, and Google-Extended are among those expected to be most directly affected by the new training category.

Some training operators have begun routing requests through residential proxy networks. These requests look like ordinary user traffic and fall outside the new controls. Site owners who want to stop them must rely on rate limiting or challenge pages that also affect real visitors.

What Site Owners Should Watch Next

Three signals will show whether the controls become standard practice. First, how many top sites flip the ad-page protection switch within the next month. Second, whether major training operators publish updated user-agent strings that match the new categories. Third, whether Cloudflare releases usage data that compares request volumes before and after the change.

Publishers who rely on ad revenue have the clearest incentive to test the new options immediately. Those who depend on search referrals may prefer to start with narrower rules and adjust after reviewing the weekly reports.

The settings sit inside the existing Cloudflare dashboard under Security, AI Crawl Control. No additional contract or pricing tier is required.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

For better AI experience,

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page