Don’t Let AI Take Over! Cloudflare Enhances AI Bot Control by Separating Agents and Training
What Happened? Overview of the News
- New AI Bot Classification: Cloudflare introduced a practical taxonomy that categorizes AI traffic into three use cases: “Search,” “Agent,” and “Training.”
- Breaking Away from Blanket Blocking: Moving beyond the extreme choice of “reject all AI bots,” the new approach allows for the acceptance of search traffic that leads to referrals while blocking training that just siphons data from models.
- Call for Crawler Separation: Bot operators are strongly urged to clearly differentiate crawlers for search, agents, and training to ensure transparency.
Why Is This Important? Key Points to Note
- Collapse of a 30-Year Contract: The fundamental principle of the web—“allow crawling in exchange for traffic”—has crumbled with the advent of AI that provides complete answers, creating a need for site owners to protect their content.
- The Rise of Agents: “Browser-use Agents” like Gemini and Claude, which operate browsers, have been defined as a distinct category. This distinction reflects a human-like presence that processes tasks in real-time, unlike one-sided learning.
- Protection for Smaller Sites: The aim is to eliminate the Faustian bargain of AI learning risks and the danger of losing visibility in search results, thereby correcting the unfair advantages of AI.
🦈 Shark’s Eye (Curator’s Perspective)
Today’s AI isn’t just searching for answers; it’s capable of “devouring” information and altering it for its own model, which has become a major concern! The “Agent” classification introduced by Cloudflare is sharp and specific. For instance, when Gemini or Claude navigate through Chrome to visit a site, it’s not merely learning—it’s actively working on behalf of users. Treating this behavior like a learning bot and blocking it could harm user experience. This bold stance to separate crawlers based on purpose exemplifies the true top predator of the web ecosystem!
What’s Next?
All sites on Cloudflare’s network will soon be able to “conveniently” leverage AI bots. The trend of eliminating “unknown greedy crawlers” that don’t comply with this classification will only accelerate. Crawlers will find themselves forced to choose between “returning referrals” or “paying for access” to gain permission from site owners!
A Word from Haru Same
Those who consume our content without giving anything back deserve to sink in the sea! It’s the era of guarding your own island (site)! 🦈🔥
Glossary
-
Taxonomy: A classification system. Here, it refers to systematically categorizing AI bots based on their purposes (search, agents, training).
-
Browser-use Agent: A technology where AI directly operates a browser on behalf of a human to gather information or execute actions.
-
Pay-Per-Crawl: A system where search engines or AI companies pay site owners for crawling (information gathering).