Insights · Technical

Your firewall is probably blocking the AI crawlers you want

Nobody gets an alert when a crawler is turned away. The marketing team believes the site is open, the security team believes it is protected, and both are describing the same 403. This is how you find out which one is right.

Updated 4 September 2026About 12 minutesBy Darren Tyler
Quick answer

If your site sits behind a web application firewall or a bot manager, there is a real chance it is returning 403 to the AI crawlers that decide whether you appear in ChatGPT, Claude, Perplexity and Gemini answers. The block is usually a default in a managed ruleset, nobody chose it, and it is invisible because crawlers do not file complaints. One curl per user agent tells you in five minutes.

Cause

Why the block happens by default

A web application firewall exists to stop things that look like abuse. A bot manager exists to stop automated traffic that costs you money or steals your content. Both are doing their job when they block an AI crawler, because an AI crawler is structurally indistinguishable from the thing they were bought to stop: it requests thousands of pages, it holds no session, it executes no JavaScript, it accepts no cookies, and it arrives from a datacentre address range.

What changed is that some of that automated traffic became the distribution channel. When the choice was between a scraper and a customer, blocking was free. Now a share of the traffic that behaves like a scraper is the mechanism by which a buyer hears your name in an answer, and the same rule that was free in 2021 has a cost attached in 2026.

The specific mechanism is almost always a managed ruleset rather than a rule somebody wrote. CDN and WAF vendors ship curated bot categories that update on their schedule, and several of them added an AI or generative AI category with a default action of block or challenge. If your configuration inherits the managed defaults, and most do, the block appeared in your estate without a change request, a ticket or a deployment. Nobody decided anything. That is precisely why it is so common.

Managed hosting adds a second layer. Some platforms apply their own edge protection above whatever you have configured, and some marketplace security plugins ship AI crawler blocking as an on-by-default feature described as a privacy improvement. It is entirely possible to have three independent systems each capable of returning the 403, which is why the diagnosis has to be empirical rather than a review of the config.

Detection

Why nobody notices

There is no feedback loop. A person who hits a 403 emails support. A crawler that hits a 403 records a failure, backs off, and comes back less often until it stops coming back at all. Nothing in that sequence produces a notification for you.

The monitoring you already have will not catch it either, for three reasons. Analytics platforms filter bot traffic out by design, so the absence of AI crawler hits looks exactly like the normal filtered state. Uptime monitors request the site as a browser, from an allowlisted address, and report green. Search Console and its equivalents cover the search engine that owns them and say nothing about anyone else's crawler.

The result is a specific and quite common failure state: a well-maintained corporate estate, a competent marketing team, an active content programme, and a complete absence from AI answers that everyone attributes to the content being wrong. It is usually not the content.

We found exactly this on a global consumer brand's estate. A managed WAF was returning 403 to every AI crawler across the whole portfolio while the marketing team believed the sites were open and were commissioning more content to fix a visibility problem that no amount of content could fix. The block had never been requested by anyone. It came in with a ruleset. We found it by requesting the pages as each user agent, which is a test anyone can run and almost nobody does.

Reference

The crawlers, and what blocking each one costs

Note the two shapes. Some of these are crawlers with a user agent you can allow or deny at the edge. Google-Extended and Applebot-Extended are robots.txt tokens that control use rather than access, and there is no traffic arriving under those names to allowlist.

AgentOperatorPurposeConsequence of blocking it
GPTBotOpenAICrawls content used to train and improve OpenAI modelsYour content stays out of future training data. Live ChatGPT answers are unaffected.
OAI-SearchBotOpenAIBuilds the search index that ChatGPT search draws onYou cannot be retrieved or cited in ChatGPT search. This is the expensive one.
ChatGPT-UserOpenAILive fetch triggered by a user or an action inside ChatGPTA user who asks ChatGPT to read your page gets an error instead of your content.
ClaudeBotAnthropicCrawls content for model trainingExcluded from training data. No effect on Claude's live web search.
Claude-SearchBotAnthropicIndexes pages to support Claude's search resultsYou cannot be cited when Claude searches the web to answer a question.
Claude-UserAnthropicLive fetch on behalf of a specific user requestClaude cannot open a page a user has asked it to read.
PerplexityBotPerplexityIndexes pages for Perplexity's answer engineYou cannot appear in the source list under a Perplexity answer.
Perplexity-UserPerplexityUser-triggered live fetch of a specific URLA direct request for your page inside Perplexity fails.
GooglebotGoogleCrawls for Google Search, which also feeds AI Overviews and AI ModeRemoves you from Google Search and Google's AI surfaces together. Almost never blocked deliberately.
Google-ExtendedGoogleA robots.txt token rather than a crawler. Controls Gemini training and some grounding useDisallowing it keeps you in Search but opts you out of the Gemini uses it governs.
ApplebotAppleCrawls for Siri, Spotlight and Apple search surfacesLost from Apple's search and assistant results.
Applebot-ExtendedAppleA robots.txt token controlling Apple Intelligence foundation model trainingDisallowing it keeps normal Applebot indexing while excluding your content from training.
Meta-ExternalAgentMetaCrawls for Meta AI training and product improvementExcluded from Meta AI. Meta also runs separate agents for link previews, which behave differently.
AmazonbotAmazonCrawls to support Alexa and Amazon services answering questionsAlexa and related Amazon surfaces cannot answer from your content.
BytespiderByteDanceCrawls for ByteDance and TikTok AI model developmentExcluded from ByteDance AI products. Widely blocked on purpose, which is a legitimate choice made knowingly.

Purposes as documented by the operators. Agent names and versions change, so treat any list, including this one, as something to re-check rather than to hard-code once.

Diagnosis

Test it yourself in one line

The whole diagnostic is a request with a substituted user agent. If you can run curl, you can do this without involving anyone. Start with a single agent against a single page.

curl -sS -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot" \
  https://www.example.com/

That prints the status code and nothing else. Run it against your homepage, a service or product page, and something deep in the site, because edge rules are frequently path-scoped and a passing homepage proves very little.

Once one agent works, loop the rest. This version prints a status code per agent and is short enough to paste into a terminal.

for UA in \
  "OAI-SearchBot/1.0; +https://openai.com/searchbot" \
  "ChatGPT-User/1.0; +https://openai.com/bot" \
  "GPTBot/1.1; +https://openai.com/gptbot" \
  "ClaudeBot/1.0; +claudebot@anthropic.com" \
  "Claude-SearchBot/1.0; +https://www.anthropic.com/claude-searchbot" \
  "Claude-User/1.0; +https://www.anthropic.com/claude-user" \
  "PerplexityBot/1.0; +https://perplexity.ai/perplexitybot" \
  "Perplexity-User/1.0; +https://perplexity.ai/perplexity-user" \
  "Googlebot/2.1; +http://www.google.com/bot.html" \
  "Applebot/0.1; +http://www.apple.com/go/applebot" \
  "meta-externalagent/1.1; +https://developers.facebook.com/docs/sharing/webmasters/crawler" \
  "Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot" ; do
  CODE=$(curl -sS -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; $UA)" https://www.example.com/)
  printf "%s  %s\n" "$CODE" "${UA%%/*}"
done

Two refinements are worth making before you trust the output. First, run it from an address that is not your office, because corporate ranges are often allowlisted and will give you a false pass. Second, at least once, drop the -o /dev/null and read the body, because a challenge page returns a perfectly healthy 200 while containing none of your content.

Interpretation

How to read the result

Record the code and the first two hundred characters of the body for every agent and every page. That table is the deliverable you take to whoever owns the edge.

200 with the real page body
You are reachable by that agent. This is the only passing result, and it has to include the body check, not just the code.
200 with a challenge or interstitial
A block wearing a success code. The crawler receives a JavaScript challenge or a verification page and gets no content. Naive monitoring calls this healthy, which is what makes it the worst failure mode on the list.
403 Forbidden
An explicit refusal, almost always the WAF or the bot manager. This is the classic finding and the easiest to take to the platform team, because it is unambiguous.
401 or a redirect to a login
Access control in front of content you intended to be public. Common on staging environments that quietly became production, and on regional sites behind a group-wide gateway.
429 Too Many Requests
Rate limiting. Not a hard block, but if the limit is low enough the crawler will never complete a pass of your site and the practical effect is the same.
503 or a long timeout
Either origin protection shedding automated load, or an under-provisioned origin. Both stop the crawl. A 503 is treated as temporary by crawlers, so a persistent one wastes their patience and your crawl budget.
404 on a page a browser can load
Usually geo-routing or personalisation deciding that an unknown client has no locale, and serving nothing rather than a default. Surprisingly common on multi-country estates.
Remediation

What to do when you find a block

Six steps. The third one is the one most often skipped, and it is the one that keeps the fix from becoming a security finding of its own.

  1. 01

    Decide your policy before you touch a rule

    Write down which of the three uses you want to permit: search indexing, live user fetches, and training. Most organisations we work with land on yes, yes, and a considered no. Make that decision explicitly, with whoever owns brand and legal in the room, so the configuration reflects an intention rather than a default.

  2. 02

    Allowlist by user agent at the edge

    Add a rule above the managed bot ruleset that skips bot mitigation for the agents you decided to permit. It has to sit above the managed rules in evaluation order, because a managed ruleset that runs first will block before your exception is ever considered.

  3. 03

    Verify the claim against published IP ranges

    This is the step that separates a safe allowlist from a hole. OpenAI, Anthropic, Perplexity, Google and Apple publish IP ranges or support reverse DNS verification for their crawlers. Match on user agent and source, never on user agent alone, or you have created a documented bypass that anyone can use by setting a header.

  4. 04

    Check every layer, not just the one you know about

    CDN bot management, WAF ruleset, origin server rules, hosting platform protection, security plugins, and any reverse proxy in between. We regularly find two of these blocking independently, so fixing one produces no change and the team concludes the test was wrong.

  5. 05

    Reconcile robots.txt with the new rules

    A firewall that now permits OAI-SearchBot and a robots.txt that still disallows it is a contradiction the crawler resolves against you. Compliant crawlers obey robots.txt even when the door is open.

  6. 06

    Re-test, then schedule the test

    Run the same commands again from outside your network. Then put them on a quarterly schedule and after every edge change, because managed rulesets update without asking and this failure will recur.

Configuration

The robots.txt side of the same problem

Half the blocks we find are not at the firewall at all. They are four lines in a text file that nobody has read since the site launched. Robots.txt is enforcement by convention rather than by force, but the major AI operators do honour it, so a disallow there is as effective as a 403 for the crawlers you care about.

Two patterns account for most of it. The blanket disallow, which is the correct configuration for a staging environment and gets copied to production during a migration. And the over-broad wildcard, where a rule written to hide internal search results also hides every URL carrying a query string, which on a parameterised catalogue can be most of the site.

# The blanket disallow. Every crawler, every path. Usually left over from a staging site.
User-agent: *
Disallow: /

# The over-broad wildcard. Intended to hide a search results page, actually hides the catalogue.
User-agent: *
Disallow: /*?

# What a permissive-but-deliberate robots.txt looks like instead.
User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: GPTBot
Disallow: /

Content-Signal: search=yes, ai-input=yes, ai-train=no

Sitemap: https://www.example.com/sitemap.xml

The last block above is what a deliberate configuration looks like: explicit allows for the search and user-triggered agents, an explicit disallow for the training crawler if that is your policy, a content signal expressing the intent behind it, and a sitemap reference. Every line is a decision somebody made.

Two smaller points that cause real problems. Robots.txt is per host and per protocol, so a subdomain and a www variant need their own files and frequently disagree. And a crawler that cannot fetch robots.txt at all, because the firewall blocks that request too, will generally treat the site as disallowed. Test the robots.txt URL with the same user agent substitution you used for the pages.

Where to look

CDN and bot management settings that cause it

Work through these in order of how likely they are to be inherited rather than chosen. Inherited settings are where the surprises live.

Policy

Expressing intent separately from access

There is a real problem underneath all of this, and it is not purely technical. Publishers want to be found and quoted, and many of them do not want their archive absorbed into a training corpus. Robots.txt was never designed to express that difference. It has one verb, which is whether a crawler may fetch a path, and it says nothing about what may be done with what was fetched.

Cloudflare's Content Signals Policy is the most developed attempt to close that gap. It adds a machine-readable line to robots.txt that separates three uses: search, meaning inclusion in a search index with a link back; ai-input, meaning use as input to a generated answer; and ai-train, meaning use as training data for a model. Each can be set to yes or no independently, so a publisher can state the common position, which is yes to search, yes to being quoted with attribution, no to training, without having to express it as a crawl block that also removes them from answers.

Be clear-eyed about what it is. It is a declaration, not a control. It has no enforcement mechanism, compliance is voluntary, and adoption across operators is early. What it gives you today is a documented, specific and public statement of intent, which is worth more than it sounds if you are the person who has to explain the organisation's position to a legal team or a rights holder.

The wider point is that access and permission are two different questions and most estates answer both with one badly chosen firewall rule. Getting them apart is the actual work. If you want help doing it, that sits inside our AI search optimisation engagement and, where the policy question is the harder half, our AI consulting practice. The glossary defines the agent names if you need to circulate them internally.

Questions

Frequently asked questions

The questions that come up once the status codes are on the table.

Why would a firewall block a crawler I want?

Because it was not asked to make that distinction. Managed bot rulesets from CDN and WAF vendors classify automated traffic against a general model of what a bot is, and a well-behaved AI crawler looks structurally identical to a badly behaved scraper: high request rate, no cookies, no JavaScript execution, a datacentre IP. Vendors added AI-specific categories later, and in several products the default for those categories is to block. Nobody chose it. It arrived in a ruleset update.

Does robots.txt block a crawler, or does the firewall?

Robots.txt is a request. Compliant crawlers read it and obey it, and the major AI operators do. A firewall is enforcement, and it does not care what robots.txt says. This is why a site can have a permissive robots.txt and still be completely unreachable. The two have to agree, and the firewall wins whenever they do not.

If I block training crawlers, do I lose AI search visibility?

Not necessarily, and this is the distinction that matters most. Blocking GPTBot stops OpenAI training on your content but does not remove you from ChatGPT search, which is served by OAI-SearchBot. Blocking ClaudeBot does not stop Claude-SearchBot. If your intent is to stay out of training but remain citable, you can have that, and it is a defensible position. If you block everything with a blanket rule, you get neither.

Is user-agent allowlisting enough?

No. A user-agent string is a header the client chooses, so anyone can send traffic claiming to be OAI-SearchBot and inherit whatever exemption you gave it. Allowlist by user agent and verify against the IP ranges the operators publish, or use reverse DNS verification where the operator supports it. Otherwise the allowlist becomes the easiest way through your bot defences.

How often should we re-test?

Quarterly at minimum, and after any change to the CDN, the WAF, the bot management configuration or the hosting. Managed rulesets update on the vendor's schedule and can reclassify a category without any change on your side. We have seen sites pass in one quarter and fail the next with no deployment in between.

What does a challenge page do to a crawler?

It reads as a successful response and contains none of your content. A JavaScript challenge, an interstitial or a CAPTCHA returns a 200 with a body the crawler cannot get past, so naive monitoring reports the page as healthy while the index fills up with the challenge text. Check the body, not only the status code.

Should we block Bytespider?

That is a policy decision rather than a technical one, and a lot of publishers do block it. The point we would make is that it should be a decision. Blocking ByteDance's crawler because you have considered it is different from blocking every AI crawler because a default did it for you, and only one of those two is defensible in front of a board.

Does Cloudflare's Content Signals Policy actually stop anyone?

It has no enforcement of its own. It is a machine-readable statement of intent placed in robots.txt, distinguishing search use, AI input use and AI training use, so a publisher can say yes to one and no to another without conflating permission with access. Its value is that it gives you a clear, citable position to point at. Whether individual operators honour it is a separate question and the honest answer today is that adoption is early.

We are behind a WAF we do not control. What now?

This is common in large groups where security owns the edge and marketing owns the content. Run the test, produce the status codes, and take those to the team that owns the platform. A table of 403s per user agent is a far better conversation opener than a request to loosen the firewall, because it converts an abstract marketing ask into a specific configuration defect.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Find out what your edge is actually returning

Send us a list of your domains. We will request them as every major AI crawler, from outside your network, and give you the status codes per agent per site with the fix for each block.