Cloudflare announced on September 15 a new security setting called Disallow AI Training, designed to solve a problem that has quietly reshaped the economics of the open web: the mixed-use crawler trap. For years, the same bot that indexed a site for Google Search also fed that content into Google's AI training pipeline, and site owners had no way to block one without blocking the other. Cloudflare's fix lets a publisher stay indexed in search while refusing to let that same crawler train an AI model on their articles. Apple, Google, and Microsoft have committed to honoring the new preference.
We at TechInKenya have written before about how exposed publishers, including Kenyan ones, are to Google's AI Overviews and AI Mode rollout, and how the platform's zero-click search behavior has hammered referral traffic across the industry. This announcement is the first concrete infrastructure-level response to that specific complaint, not from Google itself, but from the network layer that sits in front of roughly one in five websites globally. It deserves a clear-eyed look at what it actually changes, and what it does not.
What the Setting Actually Does
Cloudflare's dashboard previously offered publishers a blunt choice: Allow, Block, or Block on pages with ads. None of these granular settings could separate a crawler's search function from its training function when the two were bundled into a single bot, which is exactly how Googlebot, Applebot, and Bingbot operate. Selecting Block meant losing search visibility entirely. Leaving everything on Allow meant every article was fair game for model training.
The new Disallow AI Training option changes that. It publishes a no-training preference directly in a site's robots.txt file (specifically targeting tokens like Google-Extended and Applebot-Extended) while leaving the underlying search crawler untouched. According to Cloudflare, training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI are blocked entirely under this setting, since those companies run separate bots for search and training and therefore never presented the mixed-use dilemma in the first place.
Cloudflare paired the new control with what it calls an "Accountable" designation, a set of four requirements a crawler operator must meet or formally commit to: an opt-out mechanism for training, a separate opt-out for AI summaries, URL-level visibility into which pages were used for what, and assurance that opting out of training will not affect search rankings. Apple, Google, and Microsoft currently qualify. Bing's implementation is partial: it does not yet honor a robots.txt-based no-training signal, relying instead on an older NOARCHIVE meta tag while Microsoft builds the more complete mechanism, targeted for early 2027.

Cloudflare's own data on how site owners are currently behaving is instructive on its own. Company figures show that almost no one wants to block search (under one percent of sites on the network choose to), but roughly one in six sites already enables some mechanism to block AI training. That gap is the entire reason a blunt "Block AI" toggle was never going to satisfy publishers: they were never against being found, only against being harvested for free.
Why This Is Progress, Not a Fix
The framing in Cloudflare's own announcement is worth sitting with: this addresses training, not AI summaries. Those are two separate problems for publishers, and the second one is arguably the more damaging of the two in the short term.
Training determines whether a publisher's archive gets baked into a model's weights, a slow-moving, largely invisible harm that mostly matters for long-term content value and licensing leverage. AI summaries, by contrast, determine whether a reader who searches for an answer ever clicks through to the publisher's site at all, and that is the mechanism actively cutting into referral traffic and ad revenue right now, this month, on every search that returns an AI Overview.
Cloudflare is explicit that a site-wide opt-out for summaries exists with some operators today (Google's toggle to exclude content from generative search results, and Apple's nosnippet directive, for instance), and that Cloudflare's own centralized control over summary inclusion is not expected until "early next year." Until then, publishers who disallow AI training but leave summaries untouched will keep losing the referral traffic that pays the bills, even as their content stops flowing into new model checkpoints. Blocking training protects the archive. It does nothing for this week's page views.
There is also a structural reality buried in Cloudflare's own reasoning that publishers should not gloss over: even a fully enforced, universally honored AI training opt-out does not reverse the behavioral shift already underway. Cloudflare cites Pew Research figures showing that more than half of search users now read AI summaries, and that those users are over 40 percent more likely to end their search session without clicking a link. Once a reader's habit shifts toward asking a chatbot instead of opening a browser tab, opting a single site out of training does not pull that reader back. The user does not need your specific archive to have been in the training set. They need Google, OpenAI, or Anthropic's model, trained on the aggregate of the entire internet, to already know enough to answer without you.
The Coordination Problem
This is where the logic in Cloudflare's post runs into a limit it does not fully address, and where the argument for a genuinely collective response becomes harder to avoid. An individual Kenyan publisher opting out of AI training does not remove that publisher's category of knowledge from the models consumers use daily. A general tech explainer, a "how to register a business in Kenya" guide, or a product comparison article exists in dozens of near-identical forms across the web. If one Nairobi tech site opts out and forty competitors do not, the models still learn the same underlying facts from the sites that stayed in, and users still get their answer without visiting anyone specifically.
A training opt-out only regains leverage at scale, when enough publishers in a given content category withdraw that the remaining training data becomes noticeably thinner or more dated. That kind of coordinated action has historically required either an industry body, a licensing consortium, or regulatory pressure, none of which currently exist in a meaningful form for East African digital publishers the way, for instance, South Africa's Competition Commission has begun to apply pressure on Google over journalism sustainability.
None of this makes Cloudflare's announcement unimportant. Removing the mixed-use bundling problem is a genuine, structural improvement: publishers can now make a training decision without gambling their search visibility on it, and that alone removes a coercive dynamic that has existed since generative AI crawlers first appeared. But the tool only restores a choice. It does not restore the traffic, and it does not, by itself, change the incentive that is currently making that traffic less valuable to begin with. Whether Kenyan publishers benefit from that restored choice will depend less on what box they tick in a Cloudflare dashboard and more on whether they, and their regional peers, ever act on it together.
Comments