Local Business

Should your business block or welcome AI crawlers?

Not every AI bot does the same job. Some train models, some power AI search answers, and some fetch a page because a person asked. Here is what robots.txt can and cannot do, a simple way to decide by business type, and the policy we run on our own site.

September 25, 2026 7 min read AIcrawlersrobots.txtSEOsmall businessCloudflare

If you have looked at your website’s traffic logs lately, you have probably seen names like GPTBot, ClaudeBot, and PerplexityBot. These are AI crawlers, and every business with a website now has a decision to make about them, whether it makes it deliberately or not.

The headlines tend to frame this as a single question: block the AI companies or let them take your content. In practice it is three separate questions, because AI bots do three different jobs. Once you separate them, the right answer for most small businesses gets much clearer.

Three kinds of AI bot

Training crawlers collect pages to train or fine-tune AI models. Your content becomes part of what the model learned, usually without a link back to you. OpenAI’s GPTBot, Common Crawl’s CCBot, and Google-Extended (a control token rather than a separate crawler, which tells Google whether your content may be used for its AI models) fall into this group.

Search and answer-engine crawlers build the index that AI search tools answer from. When someone asks ChatGPT search, Perplexity, or a similar tool for “an accountant in Dartmouth who handles small business payroll”, these crawlers are how the tool knows you exist. OAI-SearchBot and PerplexityBot are examples. This is the AI equivalent of being in Google’s index.

User-triggered fetchers visit a page because a person asked an assistant to look at it right now: “summarize this company’s services page” or “check their opening hours”. ChatGPT-User and Perplexity-User are examples. There is a human on the other end of every one of these visits.

Vendors rename and add bots regularly, and some use one name for more than one purpose, so treat any list, including this one, as a snapshot. Each major vendor publishes its current bot names and what they do, and those pages are the reference to check.

How robots.txt works

robots.txt is a plain text file at the root of your site (yoursite.com/robots.txt) that tells crawlers what they may and may not fetch. A rule looks like this:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

That example blocks OpenAI’s training crawler from the whole site while allowing its search crawler. You can target any bot by name, block specific folders instead of the whole site, and set a default for everything else with User-agent: *.

A newer addition, Content Signals, lets you state preferences more directly, for example that your content may be used for search and AI answers but not for training. Support for it varies by vendor, so treat it as a clear statement of intent rather than a switch that every bot obeys.

What blocking actually achieves

robots.txt is voluntary. It is a request, not a lock. The large, reputable AI companies say they honour it, and in general they do. Scrapers that do not care about your preferences will ignore it, and a Disallow line does nothing to stop them.

Blocking is not retroactive. If a model was trained on your pages last year, adding a rule today stops future collection from crawlers that respect it. It does not remove anything already collected.

Blocking the wrong bot costs you visibility. If you block an answer engine’s search crawler, that tool has a harder time finding, reading, and citing you. Some may still mention your business from other sources, such as directories or reviews, but they will be working from someone else’s description of you instead of your own pages.

Network-level controls are stronger. If you really need to keep bots out, the enforcement happens at your hosting or network layer. Services like Cloudflare offer bot management that can identify and block known AI crawlers before they reach your site, regardless of whether they read robots.txt. Those controls are only as good as the vendor’s detection, and they can be blunt: a one-click “block AI bots” setting may block the answer-engine crawlers you actually want along with the training ones. Check exactly what a setting covers before you turn it on.

A simple decision framework

Most small businesses sit in one of these groups.

Local service businesses (trades, clinics, law and accounting firms, restaurants, retail). Your website exists to be found and to convince people to contact you. Your service descriptions and opening hours are not valuable as training data to protect; they are valuable because people read them. Welcome search and answer-engine crawlers and user-triggered fetchers. Whether you allow training is a matter of preference, and for most of these businesses there is little to lose either way.

Businesses whose content is the product (publishers, course creators, paid research, subscription newsletters). Here the trade-off is real. You may want AI search tools to find and cite you while declining training, so block training crawlers by name, allow search crawlers, and keep paid content behind a login, which is the only protection that reliably works. If this is your whole business, network-level bot controls are worth the time to configure carefully.

Businesses with sensitive or regulated material on public pages. The answer is not a robots.txt rule. Anything that should not be read by a machine should not be on a public page at all. Move it behind authentication. robots.txt can even draw attention to folders you list in it.

If you are unsure, start from the local-service default: welcome the bots that help people find you, decide on training separately, and revisit the decision once a year.

What we do on our own site

We welcome all three kinds. Our own robots.txt allows every major AI crawler by name, including the training crawlers, and publishes a Content Signals line that permits search, AI answers, and training. We made that choice because our site exists to explain what we do and how we think, and we would rather an AI tool describe us from our own words than from a guess. We also keep Cloudflare’s managed AI-bot blocking turned off, so the robots.txt rules are the policy that actually applies.

One practical detail: our robots.txt and our ai.txt are generated from a single list in our code. Before that, the two files had drifted apart and listed different bots, one with a name that did not exist. If you maintain more than one policy file, keep them in sync from one source.

That is our answer for a services business that wants to be found. Yours may differ, and that is fine. The point is to decide on purpose.

A short checklist

  • Look at your current robots.txt (yoursite.com/robots.txt) and confirm what it says, if anything
  • Check whether your host, CDN, or website builder has an AI-bot blocking setting turned on, and what it covers
  • Decide separately on training crawlers, answer-engine crawlers, and user-triggered fetchers
  • Name the bots you care about explicitly instead of relying only on User-agent: *
  • Keep anything sensitive behind a login, not behind a robots.txt rule
  • Keep robots.txt and any other policy files consistent
  • Put a yearly reminder in the calendar to review the list, because bot names change

Where to go from here

If you want to see what your site currently tells AI crawlers, our free AI-readiness audit checks your crawler policy along with structured data and the other signals answer engines read, and emails you the results. If you would rather have the gaps closed for you, the AI readiness sprint is the fixed-scope version of that work.

Or send us a short brief about your site and what you are unsure about, and we will reply in writing within one business day.

— Newsletter

Get the writing by email.

An occasional note from the team — case studies, new free tools, engineering essays. Never daily.

Three fields, no tracking. Privacy policy.

Esc