Skip to main content
§01 · Blog / AEO

What each AI engine actually wants before it cites you.

Ben LittleFounder, WhyIQPublished 11 August 2026Last updated 11 August 20265 min read

Everyone talks about AI search like it is one machine with one rulebook. It is five machines, and they disagree about what deserves a citation. I have watched people spend a quarter optimizing for a preference ChatGPT retired a year ago.

Here is the whole answer, engine by engine. ChatGPT wants its own search crawler let in and something quotable in your first 300 words. Perplexity wants community proof and a recent timestamp. Claude wants you findable in Brave, which almost nobody optimizes for. Google AI Overviews and Gemini want unglamorous, excellent SEO plus a point of view worth quoting.

All five agree on one thing, and it outranks every preference below: if the crawler cannot fetch your page, nothing else matters. Crawler accessibility rated 9.5 out of 10 as the strongest citation factor across a 54-study meta-analysis (Zyppy, 2026). Every figure here is checked against multiple named sources, because this field is full of confident statistics that died 18 months ago.

Comic panel: five robot doormen guard five club doors labelled ChatGPT, Perplexity, Claude, Google, and Gemini, each holding a different rulebook, while a founder tries to hand all five the same flyer.
Five machines, five rulebooks, one flyer that only works at some of the doors.

~85%

Share of AI citations that point to sites OTHER than the brand's own domain. Most of this game is played off your own website. Cross-engine citation analyses, 2026

What Does ChatGPT Want?

A crawler it can send, and a sentence it can quote.

Wants: OAI-SearchBot allowed in robots.txt and at your CDN. That is the search-indexing crawler and the only official lever you have. It wants the query answered directly in your first 150 to 300 words, because it cites what it can quote rather than what merely ranks. Comparative and alternatives content performs well, yours and other people's.

Does not want: a blocked OAI-SearchBot, which can leave you surfacing as a plain link but never as a cited source. Content that only exists after JavaScript hydration. An answer buried under a narrative intro.

Worth retiring: optimizing Bing rank as a ChatGPT proxy. Retrieval now fans out across several commercial fetch networks and varies by user cohort. And blocking GPTBot, the training crawler, does not remove you from search. They are independent controls.

What Does Perplexity Want?

Proof that you exist somewhere other than your own website.

Wants: genuine community presence. Reddit is still its single largest source, roughly a fifth to a quarter of citations in mid-2026 measurements, down from the 2025 peak but bigger than any other domain on any engine. It wants freshness badly enough to cite a page published within the hour, and it extracts at passage level, so every paragraph has to stand on its own.

Does not want: a blocked PerplexityBot, stale pages, or corporate pages with no third-party corroboration behind them. Zero-markup Reddit threads routinely out-cite schema-rich vendor pages here, which tells you exactly where the effort belongs. And do not fake it: astroturfed community activity gets detected and burned by the community itself long before the engine cares.

Zero-markup Reddit threads routinely out-cite schema-rich vendor pages on Perplexity. That tells you where the effort belongs.

What Does Claude Want?

To find you in Brave, which is the cheapest neglected move in AI visibility.

Wants: presence in Brave Search. Anthropic lists Brave as a subprocessor, and an independent June 2026 measurement across roughly 400 queries put Claude-to-Brave citation overlap at 79.2 percent. Brave runs the only fully independent major index, around 30 billion pages, and almost nobody optimizes for it. Claude also wants a neutral, authoritative register and server-rendered HTML.

Does not want: a blanket robots.txt disallow, because that blocks Bravebot along with everything else and quietly closes your path into Claude's answers. Allow Claude-SearchBot and Bravebot both. It does not want superlative-heavy marketing copy either. A model tuned for caution reads that as untrustworthy, so write the way you would want to be quoted.

Comic panel: a huge queue of identical marketers snakes outside a club labelled Google, while next door the Brave club stands empty with a sign reading Claude drinks here and one founder walking in unopposed.
Everyone is queued for the same door. Brave has almost nobody in line.

What Do Google AI Overviews and AI Mode Want?

What Google has always wanted, done properly.

Wants: nothing new. Google's June 2026 guidance is blunt that there are no additional requirements, no new files, no AI-specific markup. It grounds in the core Search index and then fans out to related queries, so pages that answer the questions around your question get pulled in more often. It wants a genuine point of view, because commodity restatement gives it no reason to pick you over ten identical pages.

Does not want: noindex, nosnippet, or max-snippet set to zero, which suppress the quote outright, and Googlebot blocked at the CDN.

Worth correcting: Google-Extended is a training opt-out and does not affect AI Overview eligibility. And the comfort that ranking top 10 is the ticket is fraying fast. Around 60 percent of AI Overview citations now come from URLs outside the organic top 20 (AirOps, 2026).

What Does Gemini Want?

The same things, because it grounds through the same Google index.

Wants: whatever you already did for AI Overviews. The honest summary is that Gemini is not a separate optimization project, and anyone selling you one should explain what they think is different.

Does not want: everything on the Google list above. One distinction is worth knowing: Google-Extended controls training inclusion but not real-time grounding citations on the standard API. The one documented exception is the Gemini Enterprise Agent Platform, whose grounding docs exclude pages that disallow Google-Extended. That caveat is enterprise-only, so do not generalize it.

One quirk if you use a tracking tool: Gemini returns redirect URLs in its grounding citations, so domain attribution is guesswork. Honest tools exclude Gemini from share-of-voice rather than guess.

Key takeaway

Do the shared work first: unblock every AI search crawler in robots.txt AND at your CDN, server-render what you want quoted, and answer the question in your first paragraph. Then diverge. Quotability for ChatGPT, community and freshness for Perplexity, Brave for Claude, fundamentals and fan-out for Google and Gemini.

FAQ

Does blocking GPTBot stop ChatGPT from citing me?

No. GPTBot is OpenAI's training crawler. Citation runs through OAI-SearchBot, the search-indexing crawler, and the two are independent controls. You can opt out of training and stay citable, or accidentally do the reverse. Check both lines separately.

Do schema markup and llms.txt get me cited?

Barely, and no. Structured data scored a small positive 5.6 out of 10 across 54 studies (Zyppy, 2026). For llms.txt, 97 percent of files received zero requests across 137,210 domains (Ahrefs, 2026). Keep both if cheap. Budget for neither.

Do I need to rank in Google's top 10 to appear in AI Overviews?

Less every month. Around 60 percent of AI Overview citations now come from URLs outside the organic top 20 (AirOps, 2026). Ranking still helps, but the gate is shifting toward whether the engine can fetch and quote you cleanly.

Why does my AI citation appear one week and vanish the next?

Because engines sample. Citation volatility runs 40 to 60 percent month over month, and only 30 percent of brands visible in one answer are still visible in the next (AirOps, 2026). A single run is a sample, not a measurement.

Comic panel: a founder proudly frames a screenshot of an AI answer citing their brand, while the live screen behind them re-rolls the same question and cites a competitor instead.
The screenshot on the wall and the live answer behind it are already telling different stories.

One last thing that applies to all five. Citation is probabilistic: volatility runs 40 to 60 percent month over month, and only 30 percent of brands visible in one answer are still visible in the next (AirOps, 2026). A screenshot of an engine praising you is a sample, not a measurement, and so is a screenshot of it ignoring you.

If you want to see which engines already cite you and which competitors hold your answer slots, WhyIQ AI Radar measures it weekly against real engine answers, flat from $29/mo. For the mechanism underneath all of this, read how AI search decides what to cite, and for the full checklist of moves, the AI Citability Playbook.

Stop guessing which engines cite you.

WhyIQ AI Radar runs your buyers' real questions through ChatGPT, Perplexity, Claude, Gemini, and Google AI every week and reads back who each one cited. Five engines flat from $29/mo.

See who cites you