There is a question people keep posting in SEO forums, reworded slightly every time, and it always has the same shape. How do I get ChatGPT to recommend my business? One thread title in r/SEO put it less politely: ChatGPT acts like my business does not exist. Underneath every version of that question is the same moment. Somebody opened a chat window, typed their own company name, and sat there with their jaw slightly tight while a machine decided what to say about them in front of an audience of one.
Then they closed the tab and carried that answer around all week like it meant something. It does not mean much. AI search engines are probabilistic by design, so the same question asked twice returns different answers and cites different sources. Only 30% of brands stay visible from one AI answer to the next, and only 20% are still there across five consecutive runs of the same question (AirOps, 2026). Ask once and you have not measured your AI visibility. You have flipped a coin and written down which way it landed.
30% / 20%
of brands stay visible from one AI answer to the next, and across five consecutive runs of the same question. AirOps, 2026 State of AI Search

This post is about why that happens, what it costs you, and how to run the check properly this afternoon without paying anyone. Including us. The free protocol is in section four and it works whether or not you ever touch our product.
Why Does ChatGPT Give a Different Answer Every Time?
Because it was built to. This is not a bug anyone is racing to fix.
Four things stack up. Sampling: language models pick each next word from a probability distribution rather than reading from a lookup table, so two runs diverge the moment one token lands differently. Query fan-out: modern AI search does not run your question, it runs a handful of rewritten questions derived from yours, and that rewrite is generated fresh each time. Live retrieval: the engine fetches pages at answer time, and the web moved between your two attempts. Personalisation and routing: your account, region, chat history and whichever model version you were routed to all shift the result.
None of these are exotic. Together they mean the same question asked twice is genuinely two different questions, answered from two different snapshots of the web, by a machine designed to be a little unpredictable. If you want the retrieval mechanics underneath this, we wrote them up in how AI search decides what to cite.
Key takeaway
The variance is architectural, not accidental. Any tool or habit that treats one AI answer as a verdict is reading noise and calling it a signal.
So How Bad Is the Drift, Actually?
Bad enough that the monthly picture moves as much as the hourly one.
Citation volatility runs 40% to 60% month over month (AirOps, 2026). Inside a single sitting, the answer-to-answer numbers above apply. And it is not only who gets named that moves, it is which pages get pulled: only 38% of Google AI Overview citations now come from the organic top ten, down from 76% a year earlier (Zyppy, 2026), with roughly 60% coming from outside the top twenty entirely (AirOps, 2026).
40-60%
how much citation sets shift month over month, on top of the run-to-run variance inside a single sitting. AirOps, 2026
The version of this that people actually notice is not run-to-run drift, it is the gap between engines. One r/SEO post this year is titled, near enough, ChatGPT loves us and Gemini apparently has no idea we exist. Same brand, same week, two machines, opposite verdicts. Nothing had gone wrong. The person had simply sampled two different systems and been surprised that they disagreed, which they do constantly.
So the ground under you moves in two directions at once. The engine is unstable run to run, and the rules about which pages it prefers are unstable quarter to quarter. A single screenshot captures neither.

What the One-Off Check Actually Costs You
Two bad decisions, and they point in opposite directions.
The first is the false alarm. You check on a Tuesday, you are not mentioned, and you spend the next fortnight and a chunk of budget rewriting pages that were fine. The engine had simply rolled a different set of sources that morning. You have now optimised against a random number.
The second is worse, because it feels good. You check, you are named first, you screenshot it for the team Slack, and you stop worrying. Three of the next four runs would not have named you at all. You have declared victory on a 25% result and moved your attention somewhere else.
Both mistakes come from the same root: treating a sample of one as a state of the world. The honest version of the sentence is not "ChatGPT recommends us." It is "ChatGPT recommended us once, on one machine, on one afternoon."
One check is not a measurement. It is an anecdote with a screenshot attached.
The Free Protocol: How to Check Properly in One Afternoon
You do not need a tool for this. You need a spreadsheet, an hour, and the discipline not to stop early.
1. Write ten buyer questions, not your brand name. Nobody types "is Acme any good" into ChatGPT. They type "best project management tool for a small agency" or "how do I stop invoices going late." Write the questions a buyer asks before they know you exist. Those are the ones worth winning, and the ones your brand-name vanity check never touches.
2. Add two brand-defence questions. "What is Acme" and "Acme reviews." You want to know what the engine says when somebody already has your name, because that is the answer a warm lead gets.
3. Strip your own footprint. Log out, or use a temporary chat with memory and personalisation off. Otherwise the engine is answering a question it has learned from you, and you are measuring your own history rather than the market.
4. Run every question five times. Five, not one. Fresh chat each time. This is the entire point of the exercise and the step everybody skips.
5. Do it on more than one engine. ChatGPT, Perplexity, Claude, Google AI Overviews and Gemini retrieve differently and cite differently. Being invisible on one and dominant on another is the normal case, not an edge case, and we broke down what each engine wants before it cites you separately.
6. Record three things per run, not one: were you mentioned at all, were you cited with a link, and which domain got cited instead of you. That third column is the most useful data you will collect all month.
Key takeaway
Ten questions, five runs, five engines is 250 rows. That is the smallest honest measurement of your AI visibility. Anything less is a mood.

Reading Your Results Without Lying to Yourself
Stop counting yes or no. Start counting out of five.
If you were named in two runs out of five, your answer is not "we are cited." It is "we show up about 40% of the time on this question, on this engine." That is a rate, and a rate is something you can track week over week. A yes-or-no flips constantly and tells you nothing about direction.
Then read the third column, the one with the competitors in it. If the same domain keeps appearing on the questions you lose, you have not found a content problem. You have found the page that currently owns your category's answer slot, which is a far more useful thing to know. Around 85% of AI citations point to sites other than the brand being discussed, and brands earn only about 10% of their citations from their own domain (AirOps and Foundation, 2026). The answer slot usually belongs to somebody else. Finding out whose is the work.
~85%
of AI citations point to third-party domains. Brands earn only around 10% of their citations from their own site. AirOps and Foundation, 2026
One more trap worth naming here, because it is the other question people keep asking. Tracking referral traffic from ChatGPT is not the same measurement as tracking citations, and the two are easy to conflate because both feel like "am I showing up." Most people who read an AI answer never click anything, so your referral number will always understate how often you were actually named. Track both. Just never let one stand in for the other, because a month where citations doubled and referrals stayed flat is a good month that looks like a flat one.
Where the Protocol Falls Apart
I want to be straight about this, because the protocol above is genuinely good and it will still fail you by about week three.
It falls apart on arithmetic. Twelve questions, five runs, five engines is 300 manual queries. That is most of a working day, and it captures one week. Do it weekly, as you should, and you have invented a part-time job that produces a spreadsheet nobody else on the team can read.
It falls apart on consistency. You will quietly drop from five runs to three, then to one. You will skip the engine you find annoying. You will reword a question because the old one felt clumsy, and destroy your own trend line doing it, because a reworded question is a new question with no history.
And it falls apart on memory. Six weeks in, somebody asks whether things are improving, and the honest answer is that you cannot tell, because weeks two and four are missing and week three used different wording.
None of that makes the protocol wrong. Run it. It is the fastest way to find out whether you have a problem at all. It just does not survive contact with a calendar.
What Automating It Looks Like
This is the part where I tell you we built the machine that does it, so read this section with appropriate suspicion.
WhyIQ AI Radar runs that protocol on a schedule. You pick the buyer questions rather than inheriting a keyword list. It sends each one to all five engines, ChatGPT, Perplexity, Claude, Gemini and Google AI Mode, reads back what each one actually cited, and stacks the result week over week so you get a trend instead of a snapshot. On the Agency tier it runs every question three times and averages them, which turns "cited: yes" into "cited on 2 of 3 passes." That is 40 questions times 5 engines times 3 passes, 600 real AI queries per client every week.
The three-pass band exists for exactly the reason this post exists, and we made the full argument for multi-pass measurement elsewhere. As far as we can find, no funded competitor in this category advertises a multi-run confidence measure, which means most tools in the space sell you the same coin flip you could have run yourself, formatted more attractively.
Two honest limitations before you get excited. Off-page work takes four to eight weeks to show up in citations, so nothing you read in week one is a verdict on anything. And our reading of Google AI Mode currently runs in English only. The free check is eight sample questions across all five engines, no account, about three minutes. It is a snapshot, which is precisely what this post spent 1,500 words telling you not to trust, so treat it as a starting position rather than a score.
Most AI visibility tools sell you the same coin flip you could have run yourself, formatted more attractively.
What to Do When the Number Is Low
Resist the urge to rewrite your homepage. That is almost never where the problem is.
Start with whether the engines can reach you at all. Crawler and URL accessibility scores 9.5 out of 10 as a citation factor across a 54-study meta-analysis, the highest of any signal measured, while structured data scores 5.6 (Zyppy, 2026). And check your CDN, not just your robots.txt. Ours was immaculate in May while Cloudflare quietly returned a 403 to every AI crawler at the edge, which is a humbling way to learn that the file is not the control. The rest of that list is in why AI does not cite you.
Then go where the citations actually live. Since roughly 85% of them land on somebody else's domain, the lever is third-party presence: review sites, comparison pages, and the communities your buyers already read. Brand mentions across the web correlate with AI citation at r=0.664, about three times as strongly as backlinks at r=0.218 (Ahrefs, 2025), and unlinked mentions count.
Then keep it fresh. 83% of citations on commercial questions go to pages updated within the last twelve months, and pages not refreshed quarterly are three times more likely to lose the citations they had (AirOps, 2026).
Key takeaway
Access first, third-party presence second, freshness third. Homepage copy sits somewhere below all three, which is the opposite of how most AEO budgets get spent.

The One Sentence Version
Ask once and you learn nothing. Ask five times and you learn your rate. Ask five times a week for two months and you learn your direction, which is the only thing that was ever worth knowing.
The rest is screenshots. If you want the wider picture of how this fits against ordinary search work, start with answer engine optimization, and if you want to know how we measure any of it, that is on our science page.

Frequently asked questions
How do I get ChatGPT to recommend my business?
Mostly by being present on pages other than your own. Around 85% of AI citations point to third-party sites, so review listings, comparison pages and the communities your buyers read move the number more than your homepage does. Before any of that, confirm the engine's search crawler can actually reach you.
Does ChatGPT give the same answer every time?
No, and it is not meant to. Models sample each word from a probability distribution, AI search rewrites your question into several sub-questions before retrieving, and the pages it fetches change between runs. Identical prompts routinely return different answers citing different sources.
Why does ChatGPT act like my business does not exist?
Usually one of three things: a crawler is blocked at your CDN or in robots.txt, nothing on the wider web mentions you so there is no third-party page to cite, or you asked once and caught a run that happened to omit you. Check the third possibility first, because it is free.
How do I check if ChatGPT mentions my brand?
Write ten questions a buyer would ask before they know you exist, run each one five times in a fresh logged-out chat, and record whether you were mentioned, whether you were cited with a link, and which domain got cited instead. Five runs matters more than ten questions.
How many times should I run the same prompt to get a reliable reading?
At least five per question per engine. One run is a coin flip. Three starts to show a rate. Five is the smallest sample that survives ordinary volatility, which is why measurement tools that run multiple passes report a band such as two of three rather than a yes or no.
How do I track traffic from ChatGPT?
Filter referrals from chatgpt.com and the other assistant domains in your analytics. Be aware this measures a different thing from citation: most people who read an AI answer never click, so referral traffic understates how often you were named. Track both, and do not treat one as a proxy for the other.
Should I block ChatGPT from crawling my site?
Understand which crawler you are blocking first. Training crawlers such as GPTBot and ClaudeBot are separate from search crawlers such as OAI-SearchBot and Claude-SearchBot. Blocking the training crawler is a legitimate editorial choice. Blocking the search crawler removes you from citation entirely.
How does ChatGPT decide which websites to recommend?
It retrieves pages at answer time and quotes what it finds, rather than reading a fixed ranking. Ignore the recurring claim that it simply runs Bing: retrieval now routes through several commercial fetch pipes and varies by user cohort. The one lever you control is allowing OAI-SearchBot, its search crawler, which is separate from GPTBot.