On September 15, 2026, Cloudflare implemented a new technical rule for all newly onboarded domains: on pages containing ads, AI training crawlers and Agent crawlers are blocked by default, while search crawlers remain allowed. Hybrid crawlers such as Googlebot, Applebot, and BingBot that combine search and training functions are handled under a “strictest restriction rule”—if the training category is blocked, the entire crawler is intercepted. The rule applies to new customers, newly created sites, and all existing free users; existing configurations for paying users are temporarily retained. Cloudflare co-founder and CEO Matthew Prince said: “Internet traffic is already dominated by non-humans; we must move faster.”
Three Types of Crawlers, Three Levels of Permission: An Executable Classification Framework
The new rule breaks “AI crawlers” into three behavioral categories. Search crawlers build content indexes for queries; Agent crawlers perform real-time actions on behalf of users, including browser agents for ChatGPT and Claude; Training crawlers permanently absorb content into model weights. The three behaviors have different commercial consequences, and previous technical blocking tools could not distinguish among them.
The three-tier content-use permissions introduced at the same time include: immediate allows interaction but not storage; reference (the default) allows indexing, excerpting, and backlinking; full allows summarizing and reproducing the full text. Site owners can precisely specify how a given type of crawler may use their content. The identity determination for “crawlers that have passed identity checks” also changes: passing the check is only a precondition for access; whether a crawler is allowed through depends on whether its category is permitted.
The Essential Upgrade in Payment Logic: From Pay Per Crawl to Pay Per Value
Cloudflare launched Pay Per Crawl last year, billing by actual crawl count, but data showed that more than 50% of AI crawler traffic came from repeated requests to pages that had not changed. Pay Per Use moves the billing anchor from “did it crawl?” to “did it use?”: settlement is triggered only when content actually appears in search results or an Agent actually accesses paid content. Cloudflare has launched pilots with Ceramic.ai and You.com. The former compensates publishers when their content appears in search results; the latter allows Agents to pay instantly for high-value content on demand.
Reallocation of Interests Among Parties
For content creators and publishers, the new rule rebuilds a compensation mechanism after the implicit “crawlers in exchange for traffic” pact collapsed. AI large models scrape corpora at scale but intercept traffic; Cloudflare’s category-based blocking technically seals off the channel, forcing AI companies to choose between “training corpus” and “traffic referral.”
For AI labs, the cost of acquiring training data will continue to climb. As websites migrate or switch to the new default settings, freely crawlable “clean training corpus” will shrink. Ways out include negotiating direct licenses with content platforms, accepting Pay Per Use billing, or turning to synthetic data and already licensed corpora.
For Agent developers, reliance on real-time web access faces disruption. Agent-type crawlers are blocked by default on ad-bearing pages, increasing the friction cost of calling third-party content unless the platform has established an authorization relationship in advance. The You.com pilot shows that the marginal operating cost of Agents is no longer zero, which will affect product pricing and commercialization paths.
Small website owners face a dilemma: blocking training crawlers will not bring user traffic back, but it may harm their visibility on AI search platforms. This contradiction depends on each site’s scale and business model.
The End of a Thirty-Year Agreement: The Essential Difference Between Declaration and Enforcement
robots.txt, which emerged in 1994, is a “gentlemen’s agreement” lacking technical enforcement. Cloudflare embeds “crawler category permissions” into a mandatory node through which 20% of global internet traffic passes, turning “permission” from a documentary declaration into a firewall rule. Cloudflare data shows that machine traffic has, for the first time, exceeded total human internet traffic, providing the backdrop for default blocking.
Forward-Looking Judgment
Training-data licensing agreements will quickly commercialize. Large content platforms will sign bulk licensing agreements with AI companies to bypass per-use billing; small and medium-sized websites will become beneficiaries of the “default blocking” camp. Rising training costs will accelerate industry divergence: platforms with large amounts of user-generated content will be limitedly affected, while challengers reliant on public-web crawling will face higher barriers to entry, creating a moat effect for existing leading models.
As a hybrid crawler, Googlebot faces the risk of being blocked en masse under the “strictest restriction” clause. If Cloudflare’s rules are widely adopted, Google will be forced to physically separate its search crawler from its training crawler, which will affect the data-source structure of its AI Overviews. The actual scale of compensation in the Ceramic.ai and You.com pilots will be a key test of whether Pay Per Use can achieve commercial deployment.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接