Should Ecommerce Sites Block AI Crawlers? Cloudflare's New Defaults, Read for Merchants
On September 15th, Cloudflare changes its AI crawler defaults for a large number of sites. Almost every guide you will read about this is written for publishers, whose interests are the opposite of yours. If you sell products, blocking AI crawlers is closer to a self-inflicted wound than a defensive measure. Here is how to think about it.
There is a real fight going on about AI and content, and it has produced some genuinely good tooling. Cloudflare has spent the last year building controls that let site owners separate search crawling from AI training from live AI answers, and charge for the last two if they want to.
That tooling was built for the people making the loudest noise: news publishers, media companies, and anyone whose business model is advertising against content that AI systems are now summarising for free. Their incentive is clear. If an AI answers the question, nobody visits the page, nobody sees the ad, and the business dies.
Your incentive as a merchant is close to the exact opposite. You do not sell page views. You sell products. An AI system reading your catalogue and recommending your product to a buyer is not extracting value from you, it is doing your marketing. The danger for merchants is not being read too much. It is being read too little, or not at all.
The problem is that the defaults, the guides, and the general mood on this topic are all set by publishers. If you follow them without thinking, you will quietly opt your store out of AI discovery at exactly the moment it starts to matter.
What Actually Changes on September 15th
Cloudflare now sorts AI crawlers into three categories rather than treating them as one undifferentiated blob:
- Search crawls to build a traditional search index and sends you traffic in return.
- Agent fetches your page in real time to answer a user's live question or complete a task on their behalf.
- Training takes your content to train or fine tune a model.
From September 15, 2026, Training and Agent bots will be blocked by default on pages that carry advertising. Search stays allowed. The change applies to new Cloudflare customers, new sites set up by existing customers, and all existing free plan customers.
Read that scope carefully, because it is narrower than the headlines suggest. The trigger is pages carrying advertising. A typical ecommerce store does not run display ads on its product pages, so a typical store is not caught by this default.
Two groups of merchants should not relax. First, anyone running a content or blog section monetised with display ads or a hybrid content-commerce model. Second, anyone on a free Cloudflare plan who has never looked at these settings, because "all existing free customers" is a wide net and defaults have a way of applying more broadly than you assumed.
Either way, the useful response to a default changing is not to panic about the default. It is to stop having a default at all, and make a deliberate choice.
The Three Settings, and What a Merchant Should Set Them To
Alongside the crawler categories, Cloudflare has a Content Signals Policy that adds machine readable preferences to your robots.txt. Three signals, each answering a different question about what a bot may do with your content after it has fetched it.
Cloudflare applies a managed default across millions of domains: search=yes, ai-train=no, and ai-input left neutral until you decide.
Here is how I would think about each one if you sell physical products.
| Signal | What it permits | Merchant setting | Why |
|---|---|---|---|
search | Traditional search indexing | Yes | This is not a debate. Blocking search is blocking your own store. |
ai-input | Feeding your page into a model for a live answer, including AI Overviews, AI Mode, and shopping agents | Yes | This is the one that matters, and the one people get wrong. Setting it to no removes you from AI powered product discovery. |
ai-train | Using your content to train or fine tune a model | Your call, defaults to no | Low stakes either way for most stores. Your product catalogue is not valuable training data. Leave it at no and move on. |
If you take one thing from this post, take this: ai-input is not the same as ai-train, and merchants have very different interests in the two.
The publisher's case against ai-input is strong. Their article gets summarised, the reader gets the answer, nobody clicks, no ad is served. Value out, nothing back.
The merchant's case is the reverse. A shopping agent reading your product page and recommending your item to a buyer produces a sale. That is not value extraction, that is the entire point. Every piece of work in product feed optimisation and agentic checkout readiness exists to make your products more legible to exactly these systems. Turning ai-input off undoes it.
The Googlebot Trap
This is the part that could actually hurt someone, and it is buried in most coverage.
Googlebot is a combined crawler. The same bot that indexes you for search also feeds Google's AI systems. Cloudflare applies the most restrictive matching rule to any multi-purpose crawler, so if you enable Training blocking on ad-monetised pages, you can end up blocking Googlebot on those pages too.
That is a search visibility problem wearing an AI policy costume. A merchant who reads a publisher's guide, flips on Training blocking to feel safe, and runs display ads on a content section could take an organic traffic hit on that section and spend weeks looking for the cause in the wrong place.
The operator take: the risk here is not AI companies taking something from you. It is you switching something off in a moment of caution and not connecting it to a traffic decline six weeks later.
Whatever you decide, write down what you changed and when. I have lost more time to undocumented infrastructure changes than to almost anything else, and this is a category of change that produces delayed, hard to attribute symptoms.
When Charging AI Crawlers Does Make Sense for a Merchant
Cloudflare's pay per crawl work, and the Monetization Gateway and Wallets infrastructure built on top of it, let you put a price on a resource instead of blocking it. Reporting suggests this is evolving from charging per crawl toward paying per use, where you would be compensated when your content appears in an AI answer or when an agent buys premium information for a task. Broad availability is expected later in 2026 without a firm date.
Ignore the mechanism for a moment. The question underneath is simple: does this piece of content earn its keep by being found, or by being paid for?
For a merchant, almost everything is in the first category:
- Product pages, category pages, feeds: free, always. This is your shop window.
- Buying guides and comparison content: free. These exist to be found and to route people to your products.
- Support docs and sizing information: free. An agent that can answer "will this fit" is an agent that can close a sale for you.
The narrow exception is content with standalone value that does not lead to a product sale. Original research, proprietary datasets, category benchmarks, pricing analysis. If you have genuinely invested in producing something a competitor or a model would want and it does not drive your own revenue, charging for machine access is defensible.
I sell a productised market intelligence report, so I have had to think about this for my own business rather than in the abstract. My conclusion so far is that the technology is ahead of the demand. There is no evidence yet that agents are buying content at volume in most categories, and building a machine native sales channel before buyers exist is a good way to spend a quarter on nothing. I am watching it rather than shipping against it.
What To Actually Do
| Action | Effort | When |
|---|---|---|
| Log into Cloudflare and find your current AI crawler and bot settings. Screenshot them. | 15 minutes | Before September 15 |
Check your robots.txt for Content Signals lines and confirm ai-input is not set to no | 10 minutes | Before September 15 |
| Identify whether any of your pages carry advertising, which is what triggers the new default | Low | Before September 15 |
| If you run ads anywhere, check whether Training blocking would catch Googlebot on those pages | Medium | Before September 15 |
| Review your security plugin and WAF rules for blanket bot blocking that predates any of this | Medium | This quarter |
| Write down what you changed and the date, somewhere you will find it again | Minutes | Every time |
| Decide your position on paid machine access to original research, if you have any | Thinking, not doing | No rush |
The Short Version
The AI crawler debate has been framed by people whose content is their product. Your content is your shop window. Those are different businesses and they call for opposite settings.
Let search crawl you. Let AI systems read you for live answers, because that is where product discovery is going. Decline model training if you like, because it costs you nothing either way. And be very careful about any blanket setting that treats all machine traffic as a threat, because the machines are increasingly how your customers find you.
The merchants who lose here will not lose because an AI company took something from them. They will lose because they made themselves invisible and did not notice for two quarters.
Not sure what your stack is currently blocking?
A bot layer audit is part of my AI Commerce Readiness Assessment: what your CDN, WAF, and security plugins actually block today, whether AI and shopping agents can reach your product data, and what to change before the defaults change for you.
Book a Readiness AssessmentWritten by Stuart Elrick, who has spent over a decade running multi-channel ecommerce operations and now advises merchants on AI commerce readiness. No vendor sponsorship, no affiliate relationship with any company named in this article.





