Skip to main content
§01 · Blog / AEO

Your robots.txt is perfect. Your CDN is blocking ChatGPT.

Ben LittleFounder, WhyIQPublished 7 August 202612 min read

Every AEO checklist you have read is a list of things that earn a citation. Almost none of them list the things that prevent one, which is backwards, because the blockers win. A page an AI engine cannot fetch cannot be cited however well it is written. There are five structural blockers: a blocked search-tier crawler, a CDN or firewall quietly overriding your robots.txt, JavaScript-only content, an answer buried under a long intro, and no third-party corroboration anywhere on the web. Clear those and the clever tactics start to matter. Skip them and nothing else does.

I know the second one from the inside. In May 2026 our robots.txt was immaculate, every AI crawler explicitly allowed, and our citation rate was sitting near zero because Cloudflare's managed AI-bot toggle was returning a 403 at the edge before any crawler reached the file. Nothing in the file was wrong. The file was not the control. That is the problem with reading your robots.txt and calling it done, and on 15 September it gets a wrinkle worth understanding before somebody sells you a panic about it.

Comic panel: a proud webmaster holds up an immaculate scroll labeled ROBOTS.TXT ALLOW ALL while a huge bouncer labeled CDN turns a queue of crawler robots away at a velvet rope before they can reach the door.
The file said come in. Nobody told the bouncer.

What Actually Prevents an AI Citation?

Five things, ranked by how thoroughly they finish you.

One: the engine's search crawler is disallowed in robots.txt. Two: your CDN or firewall blocks it at the edge regardless of what robots.txt says. Three: the content only exists after JavaScript runs, and most AI crawlers do not run JavaScript. Four: the answer sits four sections down, below a warm-up narrative nobody quotes. Five: nothing on the rest of the web corroborates you, so there is no third-party page for an engine to cite instead.

This ordering is not a hunch. A meta-analysis of 54 studies scored crawler and URL accessibility as the single strongest citation factor at 9.5 out of 10 (Zyppy, 2026). It is near-binary: fail it and nothing further down the list can compensate. Structured data, for scale, scores 5.6 on the same measure, and Google states plainly that it is not required for generative AI search. Most AEO budgets are spent in inverse order to that evidence, which is how you end up with a schema audit on a site the crawlers cannot reach.

9.5 / 10

evidence score for crawler and URL accessibility, the top-ranked citation factor across a 54-study meta-analysis. Structured data scores 5.6. Zyppy 54-study meta-analysis, 2026

Key takeaway

Prevention beats optimisation. Work the blocker list before any tactic list, because a page that cannot be fetched or quoted cannot be cited, and no amount of first-paragraph craft changes that.

Why Isn't Your robots.txt Enough?

Because robots.txt is a request and your CDN is a wall. Different layers, and only one of them answers the door.

Cloudflare's managed AI-bot toggle, Bot Fight Mode, and any WAF rule that catches non-browser user agents all act before your origin sees the request. The crawler gets a 403 and your file gets no say. Citation just stops, and since volatility here is high anyway, it reads as weather rather than a wall.

Ours ran that way for an unknown period in May 2026, caught only because we read the served file rather than the one in our repo. So that is the check now. Fetch your live robots.txt over the public internet: if the first line is a comment block about conditions of access instead of your own directives, a managed rule is rewriting it. Then request a real page with a crawler user agent from outside your network and confirm a 200, not a challenge. Cloudflare now sorts AI bots into Search, Agent and Training and lets you handle each separately, so a refusal can also arrive as an HTTP 402.

A blocked crawler produces no error, no alert, and no gap in your analytics. It produces silence, and silence in this channel is indistinguishable from a slow month.

What Changes on 15 September 2026?

Less than the headlines suggest, and one thing more dangerous than either.

From 15 September, domains newly onboarding to Cloudflare get Training and Agent crawlers blocked by default, but only on pages that display ads, and the Search tier remains allowed. It does not change existing domains and it is not a blanket AI block. Search is the tier that gates citation, so the default itself does not stop you being cited. If you see a post claiming Cloudflare is switching off AI citation on that date, close it.

The real hazard is quieter and already live. Cloudflare applies the strictest matching rule to a crawler serving several purposes, and Googlebot crawls for both search and training. A site that switches on Training-tier blocking can therefore block Googlebot, and lose AI Overviews and classic organic search together. That is the most damaging easily-made mistake in this area, and it is made by people being careful, not careless. Blocking training crawlers is defensible if your content is the product. Blocking search-tier crawlers removes the only path most brands have to citation at all.

Comic panel: a large net labeled TRAINING BLOCKED sweeps up a group of AI crawler robots, and tangled in the middle of the net is a startled Googlebot, while behind it two doors labeled AI OVERVIEWS and ORGANIC SEARCH slam shut.
One switch, two doors. Googlebot crawls for both search and training, and the strictest rule wins.

Which Crawler Actually Gates Citation?

The search-tier one. Almost every robots.txt argument on the internet is about the wrong bot.

Each major engine runs a split fleet. OAI-SearchBot indexes for ChatGPT search while GPTBot collects training data, and they are independent: block GPTBot and you can still be cited, block OAI-SearchBot and you cannot. Anthropic runs the same split with Claude-SearchBot and ClaudeBot. Perplexity uses PerplexityBot. Google's AI Overviews and AI Mode draw from the ordinary Googlebot index.

Which makes Google-Extended the most expensive misunderstanding in the category. It is a training opt-out token only. It does not affect AI Overview eligibility in either direction, so disallowing it costs you nothing there and allowing it wins you nothing. Two more asymmetries are worth knowing: OpenAI documents that ChatGPT-User, the user-triggered fetcher, does not respect robots.txt, so a rule there is not a control. Anthropic's equivalent, Claude-User, does respect it, and a hit from it is the clearest signal that Claude is actively recommending your page to a person.

Key takeaway

Before you argue about whether to block AI crawlers, check which one you are blocking. The training-tier decision is a content-licensing question. The search-tier decision is a visibility question, and they deserve different answers.

Comic panel: two side-by-side gateways under a sign reading ONE OF THESE DECIDES IF YOU GET CITED. A robot badged OAI-SEARCHBOT is padlocked behind bars holding a golden CITATION ticket, while a near-identical robot badged GPTBOT strolls whistling through the open gate beside it, and a pleased-looking webmaster stands between them holding the key.
Near-identical names, opposite jobs. Most robots.txt arguments are about the wrong bot.

What Gets ChatGPT to Cite You?

Being quotable in your first 300 words, and being fetchable by OAI-SearchBot. Not ranking in Bing.

The Bing thing needs retiring properly, because it is still in circulation. In early 2025 it was fair to say ChatGPT leaned on Bing and that most of its citations matched Bing's top results. That is obsolete as architecture. Mid-2026 instrumentation shows retrieval fanning out across several commercial fetch networks, with the mix varying by user cohort and changing week to week, plus OpenAI's own index and cache. Bing is one gated experiment among several, not the foundation. Optimising for Bing rank as a ChatGPT proxy is now a strategy pointed at a pipe that may not be in the path.

What holds is duller and more durable. Allow OAI-SearchBot. Answer the actual question in the first 150 to 300 words, in language a model can lift without editing. Roughly 44.2 percent of citation extractions come from the first 30 percent of body text (AirOps, 2026), so answer placement is a bigger lever than answer length. Comparative and alternatives coverage performs strongly here, both yours and other people's.

44.2%

of AI citation extractions come from the first 30 percent of a page's body text. Where the answer sits beats how long the page is. AirOps, 2026, 548,000-page study

What Gets Perplexity to Cite You?

Community presence and recency. Schema effort here is close to wasted.

Perplexity runs its own crawler and index with a heavy recency bias, and extracts at passage level, so every paragraph has to stand on its own. It can cite content published within the hour and skews hard to current-year pages. Its declared crawl footprint is small: Googlebot reaches around 167 times more unique URLs than PerplexityBot (Cloudflare, January 2026), which is a useful corrective if you assumed it was crawling the whole web.

Reddit remains its single largest source domain. Note the correction there, because the widely quoted 46.7 percent figure was a pre-litigation peak measured on a top-citations denominator, and the share now sits at roughly a fifth to a quarter. Still the biggest single domain on any engine, still volatile enough that you should quote it softly. We covered the mechanics of that in why Reddit outranks you. What prevents a citation here: blocking PerplexityBot, stale pages with no dateModified, corporate pages with zero community corroboration, and astroturfed presence, where Reddit's own detection is the binding constraint rather than Perplexity's.

What Gets Claude to Cite You?

Being findable in Brave. This is the most neglected AI-visibility action available, because almost nobody does it.

Claude's web search is powered by Brave, which Anthropic lists as a subprocessor, and independent measurement in June 2026 put Claude-versus-Brave citation overlap at around 79.2 percent across roughly 400 queries. Brave runs one of the few fully independent major indexes, around 30 billion pages, with its own robots-respecting crawler. So there is a second index in this market with a fraction of the competition, and the entire SEO industry is queuing at the other door.

There is no Brave webmaster console, so the sequence is short. Confirm Bravebot is not caught by a blanket disallow in robots.txt or at your CDN, the trap almost everyone falls into because nobody blocks Bravebot on purpose. Submit at search.brave.com/submit-url. Part of Brave's index comes from its Web Discovery Project and is outside your influence; the rest is ordinary ranking. Beyond that, Claude rewards a neutral authoritative register with clear author attribution, and punishes superlatives harder than any other engine. Write the way you would want to be quoted.

Comic panel: an enormous queue of SEO workers stretches toward a giant ornate door labeled GOOGLE, while a few metres away an unattended side door labeled BRAVE stands wide open with a single Claude robot strolling through it.
One index, almost no queue, and a robot that keeps walking through it.

What Gets Google AI Overviews to Cite You?

SEO fundamentals, done on a page the engine can actually quote. But the top-10 assumption has quietly expired.

Google's own position is that there are no additional requirements to appear in AI Overviews or AI Mode: no AI files, no special markup. What it asks for is non-commodity content with a unique point of view, indexed, snippet-eligible and server-rendered. Retrieval is grounded in the core Search ranking systems plus query fan-out, where the model fires several related queries at once. So answer the questions adjacent to the main one, not just the literal one.

The shift worth pricing in is the top-10 assumption. The coupling between AI Overview citation and organic rank is weakening fast: top-10 share of AIO citations has fallen from roughly 76 percent to about 38 percent, and around 60 percent now come from URLs outside the top 20. Rank top-10 and you will be cited is no longer a safe restatement of Google's it-is-still-SEO thesis. The gate is moving from do you rank toward can the engine fetch and quote you. Gemini shares the index and the guidance; treat them as one surface.

76% to 38%

fall in the share of Google AI Overview citations that come from organic top-10 results. Around 60 percent now come from outside the top 20. AI Overview citation analysis, 2026

What Should You Fix First?

Five moves, in this order. The first two are unglamorous and they are the ones that decide it.

One: confirm no AI search crawler is blocked, in robots.txt and at your CDN, checked against the live served file rather than the one in your repo. Two: server-render whatever you want quoted, because most AI crawlers do not execute JavaScript and will read a blank page instead, as we measured in AI crawlers can't read your website. Three: answer the buyer's real question in the first paragraph, in plain language. Four: build third-party corroboration, because around 85 percent of AI citations point at domains other than the brand's own. Five: keep it fresh, quarterly at minimum. Pages off that cadence are around three times more likely to lose citations.

Why Isn't One Check a Measurement?

Because the engines are probabilistic, and you are reading one sample off a noisy distribution.

Volatility runs 40 to 60 percent month to month, only about 30 percent of brands stay visible from one answer to the next, and only 20 percent remain present across five consecutive runs (AirOps, 2026). Run the same prompt twice and you can get two different answers citing two different sources, minutes apart. So a single check that shows you cited proves little, and a single check that shows you absent proves less.

One query on one engine on one day is a sample, not a measurement, which is why our own AI Radar runs each prompt several times and reports a band rather than a verdict. The full sequence, with effect sizes on every move, is the AI Citability Playbook.

~85%

of AI citations point at domains other than the brand's own. Brands earn roughly 10 percent of their citations from their own site. Cross-engine citation analysis, 2026

Comic panel: a small crawler robot stands defeated at the first of five barred gates in a dark corridor, labeled CRAWLER, CDN, RENDER, ANSWER and PROOF. The first two are slammed shut, the last three stand open, and a brilliant glowing webpage floats unreachable at the far end.
Five gates. The last three do not matter until the first two are open.

Frequently asked questions

Why is my site not being cited by AI?

In most cases it is a blocker, not a quality problem. Check five things in order: whether the engine's search crawler is allowed in robots.txt, whether your CDN or firewall is blocking it independently of robots.txt, whether the content is server-rendered rather than JavaScript-only, whether the answer appears in the opening of the page, and whether anyone else on the web mentions you.

Does blocking GPTBot stop ChatGPT from citing me?

No. GPTBot is OpenAI's training crawler and is independent of search. The crawler that gates citation is OAI-SearchBot. You can block GPTBot and still be cited, or allow GPTBot and never be cited because OAI-SearchBot is blocked. The same split applies at Anthropic, where ClaudeBot is training and Claude-SearchBot is search.

Can Cloudflare block AI crawlers even if my robots.txt allows them?

Yes, and this is the failure mode robots.txt cannot show you. Cloudflare's managed AI-bot controls, Bot Fight Mode, and custom WAF rules all act at the edge, returning a 403 before the request reaches your file. Our own citation rate was silently sunk this way in May 2026 while our robots.txt read perfectly.

What is Cloudflare changing on 15 September 2026?

From that date, domains newly onboarding to Cloudflare get Training and Agent crawlers blocked by default, but only on pages that display ads, and the Search tier stays allowed. It does not change existing domains and it is not a blanket AI block. Search is the tier that gates citation, so the default does not stop you being cited.

Does Google-Extended affect AI Overviews?

No. Google-Extended is a training opt-out token only. Disallowing it does not remove you from AI Overviews or AI Mode, and allowing it does not help you appear there. The crawler that matters for both is Googlebot, and the controls that limit your appearance are noindex, nosnippet, data-nosnippet, and max-snippet.

How do I get cited by Claude?

Claude's web search is powered by Brave, which Anthropic lists as a subprocessor, and independent measurement puts Claude-versus-Brave citation overlap at around 79 percent. So the neglected lever is Brave Search presence. Confirm Bravebot is not caught by a blanket disallow, submit at search.brave.com/submit-url, and write in a neutral, authoritative register.

Why does my AI citation status keep changing?

Because citation is probabilistic, not deterministic. Volatility runs 40 to 60 percent month to month, only around 30 percent of brands stay visible from one answer to the next, and only 20 percent remain present across five consecutive runs (AirOps, 2026). A single check is one sample, not a measurement.

Go deeper

For the retrieval mechanic behind every citation, see how AI search decides what to cite. For the tactic that tops every checklist and earns nothing, see does llms.txt actually work. For where all of this sits against ordinary SEO, see answer engine optimization.

Find out which blocker is on your page.

WhyIQ checks crawler access, server-rendered content, answer placement, freshness and the rest of the citation signals on any URL. Free first scan, results in about 2 minutes.

Get my free WhyIQ Score