Skip to main content
§01 · Blog / AEO

An AI visibility score answers one of two questions. Most tools won't tell you which.

Ben LittleFounder, WhyIQPublished 9 August 202611 min read

An AI visibility score is a number describing how visible your brand is in AI-generated answers, the recommendations ChatGPT, Perplexity, Gemini, and Google AI hand your buyers. If you are shopping for one, here is the distinction the sales pages skip: every score answers one of two different questions. Either it predicts whether your pages are ready to be cited, computed from on-page signals, or it measures whether you actually are being cited, computed by running real prompts through real engines and reading back the answers. Both are legitimate. They are not interchangeable, and a vendor who will not say which one you are buying is selling you the confusion.

The category is consolidating fast. Search demand for "AI visibility" has grown every month for a year, the question-form searches have more than tripled, and the big SEO suites now ship toolkits under exactly that name. The stakes behind the number are real: across a 57.2-million-citation analysis, brands earned only 10.15 percent of AI citations from their own domains (Foundation and AirOps, 2026), so most of what an AI visibility score measures is happening somewhere you do not control. Which makes this the right moment to get precise about what the number can and cannot tell you, because the two kinds of score fail in opposite ways, and confusing them costs real money in both directions.

Comic panel: a puzzled marketer stands between two large gauges, one labeled READY and one labeled CITED, showing different readings for AI visibility.
Two gauges, two questions. The label on the dial is the part most dashboards leave off.

What Does an AI Visibility Score Measure?

"Am I ready to be cited?" and "Am I actually being cited?" sound like the same question. They are not even close.

The first question is about your pages. Can the engines fetch them? Does the opening paragraph carry a self-contained answer? Is there evidence density, author attribution, a fresh date? Those signals are readable from the page itself, in minutes, without asking any engine anything. Score them and you get a readiness number: a structural prediction of how citable the page is. WhyIQ's page scanner produces exactly this kind of number, our AI Citability Index, and we are strict about its label: it predicts. It is not a count of citations, and we say so on our methodology page in those words.

The second question is about outcomes. When a buyer asks Perplexity "best tools for X," who gets cited, you or the competitor whose thread holds the answer slot? The only way to answer it is to do what the buyer does: run the real prompt through the real engine and read the real answer. That is a measurement, and it is a different instrument entirely. It knows nothing about your schema markup. It knows what happened. The mechanics of how engines pick those sources, retrieval, grounding, query fan-out, are covered in how AI decides what to cite.

Comic split panel: an inspector approves a glowing blueprint labeled PREDICTION on one side, while on the other side a finished building labeled MEASUREMENT has its lights on with visitors walking in.
A blueprint inspection and a walk through the finished building. Useful in that order, never interchangeable.

What Can a Readiness Score Actually See?

Readiness scores are fast, cheap, and diagnostic. Their weakness is structural: they can only see the page.

A good readiness score earns its keep as a fix list. The evidence-backed signals are real: crawler and URL accessibility is the top-ranked citation factor across a 54-study meta-analysis (Zyppy, 2026), answer placement matters because 44.2 percent of citation extractions come from the first 30 percent of body text (AirOps, 2026), and statistics density lifted visibility by roughly 40 percent in the Princeton GEO study (KDD 2024). Scoring those signals tells you what to repair before the engines visit, which is exactly what a pre-flight inspection is for.

Now the ceiling. The strongest single correlate of actually being cited is not on the page at all: brand mentions across third-party sites correlate with AI visibility at r=0.664, roughly three times the correlation of backlinks (Ahrefs, 2026). A readiness score cannot see your Reddit presence, your review-site footprint, or the listicle that names your competitor and not you. So a page can score 90 for readiness and still be invisible in real answers, because the engines mostly cite what other people say about you. Perfect blueprint, empty building. That is not a flaw in readiness scoring; it is the boundary of what page signals can know, and any vendor whose readiness number implies otherwise is overclaiming.

r=0.664

correlation between third-party brand mentions and AI visibility, about 3x backlinks. The strongest signal is one no page scan can see. Ahrefs, 75,000 brands, 2026

What Can Only Measurement Tell You?

Which engine. Which prompt. Which competitor took the slot. None of that exists in a page scan.

A measured AI visibility score is built from real engine runs, so it can answer the questions that decide budgets. Per engine: you might be cited regularly by Perplexity and absent from ChatGPT, and those two problems have different fixes, because the engines favor different sources. Per prompt: visibility on "your brand review" and visibility on "best tool for your category" are different businesses; the first is brand defense, the second is discovery, and only measurement can keep their denominators apart. Per slot: when you are absent, measurement shows who is present instead, which turns a score into a target list. The competitor holding your answer slot on a Reddit thread is a piece of information no readiness scan will ever produce.

Measurement is also what closes the loop on everything else you do. Fix your pages, earn the mentions, then watch whether the citations actually arrive, engine by engine, week by week. Without the read-back step, every AEO tactic is an act of faith. With it, tactics get retired when they do not move the number, which is exactly how we ended up deleting a check from our own product, a story we told in the llms.txt post.

Comic panel: a detective robot with a magnifying-glass eye inspects a podium labeled BEST TOOL FOR X where a rival mascot stands holding a flag reading REDDIT THREAD, while engine robots watch from a bench.
When you are absent, measurement shows who is standing there instead. That is a target list, not a score.

A readiness score tells you what to fix. A measured score tells you whether it worked. Confuse them and you either fix blind or celebrate a forecast.

Why Does My AI Visibility Score Keep Changing?

Ask an engine the same question twice and you can get two different answers. Your score inherits that.

This is the measurement category's own honesty test. AI engines are probabilistic: the same prompt, on the same engine, minutes apart, can cite different sources. A tool that runs each prompt once and prints "cited" or "not cited" is presenting one dice roll as a fact, and next week's opposite roll as a trend. The flapping you see in single-shot dashboards is often not your visibility changing. It is the sampling error changing.

The fixes are boring statistics. Run each prompt multiple times and report a band: cited in two of three passes says something a single pass cannot. Track weekly trends rather than daily snapshots, because citations move on the timescale of crawls and off-page mentions, not hours. And when an engine cannot be measured, say so, rather than printing a zero that looks like a finding. We wrote up the multi-pass approach in the 3-pass confidence band, including the worked example where a single-shot check would have reported the same page as both winning and losing inside one afternoon.

Comic panel: a robot oracle answers the same question from three identical visitors, telling one CITED, one ABSENT, and one CITED again, while dice tumble in its display.
Same question, same engine, three answers. A single-shot score is one roll of these dice.

What Is a Good AI Visibility Score?

You can audit a score in five questions, and none of them require seeing the vendor's code.

First: is this a prediction or a measurement? The vendor should answer in one sentence. Second: what is the denominator? A cited rate means nothing until you know it is "share of answers to non-branded buyer questions," and a score that mixes brand-name prompts into the denominator is flattering itself, because an engine citing you when someone types your own name is defense, not discovery. Third: what is the sample? How many prompts, how many engines, how many passes, how often. Fourth: what happens when measurement fails? The honest answer is a gap in the data, never a fabricated zero. Fifth: can you see the receipts, the actual prompts run and sources cited, or only the aggregate?

The big suites entering this category ship capable tools, and this is not an argument against any of them. It is an argument for reading the methodology page before trusting the dial. The same five questions apply to our own products, which is why the answers are published, and why the two numbers we produce are labeled as what they are: the Citability Index predicts, and WhyIQ AI Radar measures. We keep the line sharp because the day a forecast gets presented as a measurement, the number stops meaning anything.

Comic panel: a buyer holds a glowing lantern up to the fine print beneath a giant score numeral 87, illuminating the questions SAMPLE SIZE, DENOMINATOR, and HOW MEASURED.
Any dial can print a number. The five questions are what tell you whether to believe it.

Key takeaway

Five questions audit any AI visibility score: prediction or measurement? What denominator? What sample? What happens on failure? Where are the receipts? A vendor who answers all five plainly is selling you an instrument. Anything less is selling you a dial.

How Do I Check My AI Visibility for Free?

The right order: measure first to see where you stand, then use readiness to decide what to fix.

Start with a measured check, because it tells you whether there is a problem worth working on and who currently owns your answer slots. WhyIQ's AI Radar free check runs 8 buyer-intent prompts in your category through ChatGPT, Perplexity, Claude, Gemini, and Google's AI Mode, and reads back what each engine actually cited, you, or the competitor who took the slot. Then run a readiness scan on the pages you want cited, and work the fix list: crawler access, first-paragraph answers, evidence density, freshness. Then, and this is the part most teams skip, measure again. Visibility work has a 4 to 8 week lag while engines re-crawl and third-party mentions accumulate, so the loop is measure, fix, wait, measure, not scan once and hope.

However you run it, keep the two questions separate in your reporting the way you now know to keep them separate in your tooling. "Our pages are ready" and "the engines cite us" are the start and the end of the same journey, and the distance between them, the off-page work, the mentions, the weeks of waiting, is where AI visibility is actually won.

Frequently asked questions

What is AI visibility?

AI visibility is how often and how prominently a brand appears in AI-generated answers when buyers ask engines like ChatGPT, Perplexity, Gemini, or Google AI questions in its category. It covers being cited as a source, being named in the answer text, and holding the recommendation slot a competitor would otherwise hold. It is the AI-search analogue of ranking in classic search results.

What is an AI visibility score?

An AI visibility score is a single number summarizing that visibility. The critical distinction is what the number is made of. Readiness scores predict how citable a page is from on-page signals like crawler access, answer clarity, and schema. Measured scores run real buyer prompts through real engines and count the citations that actually happened. Both are legitimate; they answer different questions, and a vendor should tell you plainly which one you are looking at.

Is an AI visibility score a prediction or a measurement?

It depends on the tool, and that is the first thing to check. If the score is computed by analyzing your page (fast, no engine queries), it is a prediction of readiness. If it is computed by querying ChatGPT, Perplexity, and the other engines with real prompts and reading back the cited sources, it is a measurement. A useful test: ask the vendor which prompts were run, on which engines, how many times. A readiness score cannot answer that question, because it never ran any.

What is a good AI visibility score?

There is no universal benchmark, and any score without its denominator is unreadable. The honest questions are: what share of answers to non-branded buyer questions cite you (single digits is a normal starting point; category leaders reach 25 percent or more), and is the number computed over enough samples to be stable? Beware scores inflated by brand-name prompts: an engine citing you when someone types your own name is defense, not discovery, and mixing the two flatters the number.

Why does my AI visibility change between checks?

Because AI engines are probabilistic. The same prompt on the same engine can produce different answers, and different cited sources, minutes apart. A single-shot check is one sample from a noisy distribution, so week-to-week flapping is often measurement noise, not real movement. Tools handle this by running each prompt multiple times and reporting a band (cited in 2 of 3 runs) or by tracking a trend over weeks rather than reacting to one snapshot.

How do I check my AI visibility for free?

Run both kinds of check. A readiness scan analyzes your page's citation signals in about two minutes and predicts what is blocking you. A measured check runs real category prompts through real engines and shows who actually got cited, you or a competitor, on each one. WhyIQ offers both free: the page scan scores readiness, and the AI Radar free check runs 8 buyer prompts through 5 engines and reads back the results.

Go deeper

For the citation mechanic itself, see how AI decides what to cite. For why single-shot tracking misleads, see weekly vs daily AI citation tracking.

Stop guessing your AI visibility. Measure it.

The WhyIQ AI Radar free check runs 8 real buyer prompts through ChatGPT, Perplexity, Claude, Gemini, and Google's AI Mode, and reads back exactly who each engine cited. No account, results emailed in about 3 minutes.

Run my free AI visibility check