How AI Engines Decide Which Websites to Cite in Answers
By Joe Della Mora — Founder, GroundScore

How AI engines decide which websites to cite is not a mystery, and it is not a black box you can only appease with vibes. Citation is the visible end of a pipeline: the engine interprets a question, retrieves candidate pages, selects the passages that answer best, and writes a grounded answer that attributes claims to sources. Your site survives every stage or it appears nowhere.
That pipeline view is what this post adds: not another list of tips, but a map of where sites get filtered out — and which fix belongs to which stage, so you stop applying content fixes to access problems and vice versa. This is the capstone to our fundamentals series, and it ties the earlier guides together. We will walk the pipeline stage by stage, then map the fixes.
What happens between a question and an answer?
Direct answer: When you ask an AI engine a question, it interprets what you want, retrieves candidate pages from an index or a live web search, selects the specific passages that best answer the question, and then writes an answer grounded in those passages — citing the sites they came from.
The details differ by engine — ChatGPT's search mode, Perplexity's always-on citations, Claude's web search, Gemini's grounding in Google — but the shape is shared, because they are all solving the same problem: a language model generating from memory alone would be stale and unreliable, so the engine grounds its answer in retrieved documents and shows its receipts.
Think of it as four gates in sequence. Query understanding decides which searches get run, and therefore which topics you could even be considered for. Retrieval builds the candidate list — pages the engine can actually fetch and index. Passage selection narrows candidates to the specific paragraphs worth building an answer from. Grounding and attribution decides which of those passages get relied on and credited in the final text.
The gates multiply rather than add. Perfect content behind a blocked crawler is invisible; a fully crawlable site with no quotable passages is retrieved and then ignored. This is why AI visibility work that focuses on one favorite tactic so often disappoints — the pipeline fails at its weakest stage, not its strongest. If you want the one-line summary of this entire post: find your weakest gate and fix that one first.

Query understanding: which questions put you in play
Direct answer: Before any retrieval happens, the engine turns the user's question into search intents — what actually needs looking up. If your content never matches the questions buyers ask, because it is organized around your product names instead of their problems, you are filtered out before the race even starts.
This stage is easy to overlook because nothing about your site is being evaluated yet. The engine takes "what should I look for when hiring a deck builder?" and decomposes it into things to search for: qualification signals, cost factors, red flags, local considerations. Those derived intents — not your keywords — determine which searches run and which corners of the web get considered.
You influence this stage indirectly, through the topics your site can be found for. A site organized entirely around its own vocabulary — brand names, service package titles, "solutions" — matches poorly against the plain-language intents engines derive from real questions. A site with pages built around buyer questions matches naturally, because the page and the intent share a shape.
The practical work here is question research: writing down what customers actually ask before they choose a business like yours, in their words, and making sure each important question has a page whose job is answering it. This is also where honest expectations start. If nobody asks the engines questions in your market, no downstream fix manufactures demand; and if the questions being asked are not the ones you have answered, the pipeline ends for you here, silently, on every one of them.
Retrieval: making the candidate list
Direct answer: Retrieval builds the shortlist: the engine pulls candidate pages from its index or from a live web search. Sites that block AI crawlers, hide their content behind JavaScript, or were never indexed do not make the list. No retrieval means no citation, no matter how good the content is.
This is the most mechanical stage and the most brutal, because failure here is total and invisible. The engine can only select passages from pages it holds in its hands — pages its crawlers fetched, or its search partner indexed, or its user-triggered fetcher can grab live. Everything else does not exist for that answer.
Three failure modes dominate. Blocked crawlers: robots.txt rules or firewall settings that turn away GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, or Google-Extended — sometimes deliberately, more often as a forgotten leftover. JavaScript-only rendering: pages whose text materializes only when a browser executes scripts, which many fetchers do not. And plain absence from search indexes, since several engines retrieve through conventional search infrastructure.
Building GroundScore, the pattern I keep seeing when we scan sites is failure exactly here: an owner deep into content strategy while a security plugin quietly serves errors to every AI crawler. It is the first thing worth checking because it is checkable in minutes and fixable in an afternoon — our guide to checking whether AI crawlers can access your website walks through the full audit, agent by agent.
Pass this gate and you are a candidate. Nothing more. Candidacy is where the interesting competition begins.
Passage selection: winning the quote
Direct answer: From each candidate page, the engine extracts the passages most likely to answer the question. Clear, self-contained paragraphs that state an answer directly beat long, meandering sections. This is where structure pays: a page can be relevant overall and still offer nothing clean enough to quote.
Here the unit of competition shrinks from the page to the paragraph. The engine is not ranking sites; it is hunting, inside each retrieved page, for spans of text that resolve the user's question — and comparing your best paragraph against the best paragraph from every other candidate.
That changes what "good content" means in a specific, writable way. A passage wins when it is self-contained (it makes sense lifted out of the page), direct (it answers in its first sentence rather than its fifth), and concrete (it commits to specifics a generated answer can use). A passage loses when the answer is smeared across twelve paragraphs, buried under a preamble, or hedged into mush.
Structure is how you help the engine find your best passages. Headings that state the question each section answers, opening paragraphs that answer it immediately, FAQ blocks with standalone answers, and schema markup that labels what each piece of content is — Organization, Article, FAQPage — all reduce the work between your page and a quotable span. Our guides to writing content AI engines actually cite and the schema types that matter most for AI search cover both halves of this stage in depth. The habit that captures it: write every important section so its first paragraph could be quoted alone, because that is precisely what the engine is trying to do.
Grounding and attribution: the final cut
Direct answer: In the final stage the model writes its answer and attributes claims to sources. Passages that are specific, verifiable, and consistent with what other sources say are safer to cite. Content that is vague, contradictory, or unsupported tends to get paraphrased without credit — or dropped from the answer entirely.
The last gate is about trust under scrutiny. The engine is about to put its name on an answer and point at your site as evidence, so the selected passages get weighed as evidence: Do they commit to checkable specifics? Do they agree with what other retrieved sources say? Does the site itself present a consistent identity — the same business name, location, and claims here as everywhere else the engine has seen?
Corroboration does quiet, heavy work at this stage. A claim that appears only on your site is a claim; a claim echoed by directories, review platforms, and independent pages is a fact the engine can lean on. This is why entity consistency — boring, unglamorous agreement between your site, your listings, and your profiles — shows up so often in the accounts of what engines reward. Inconsistency does not get you penalized so much as quietly routed around: your material informs the answer as unattributed background while a steadier source collects the citation.
The uncomfortable, useful truth: paraphrase-without-credit is the fate of generic content. When your passage says what every other candidate says, the engine needs none of you in particular. Surviving the final cut means being the source that said something specific enough to need attributing. The broader signal picture — authority, freshness, corroboration — is covered in our guide to the signals AI engines weigh before citing your website.
Which fixes map to which stage?
Direct answer: Map each fix to the stage it unblocks: question-shaped pages feed query understanding, robots.txt and firewall fixes unblock retrieval, schema and direct-answer writing win passage selection, and entity consistency plus corroboration survive grounding. Fix in pipeline order — an unblocked crawler comes before a better paragraph.
Diagnosis before treatment. Most wasted AEO effort is a stage mismatch: rewriting paragraphs when crawlers are blocked, or buying links when the real failure is that no page answers the question being asked.
| Stage | What filters you out | Where to start |
|---|---|---|
| Query understanding | Content organized around you, not buyer questions | Question research; one page per real question |
| Retrieval | Blocked crawlers, JavaScript-only pages, no index presence | robots.txt and firewall audit; render check |
| Passage selection | Nothing standalone or direct enough to quote | Direct-answer openings; FAQ blocks; schema |
| Grounding | Vague claims, inconsistent entity facts | Specifics with sources; consistent NAP everywhere |
Work the table top to bottom. Retrieval fixes are fast and verifiable, so clear them first even though they feel less strategic than content work. Passage-level rewriting is the slow, compounding middle. Grounding-level consistency is mostly housekeeping — but housekeeping that decides close calls.
And measure at the end of the pipe, because citation is the only stage you can observe directly from outside. Ask the engines your buyer questions, log who gets cited, fix the weakest stage, and ask again. Everything upstream is inference; presence in answers is the result.

Frequently asked questions
Do AI engines just cite whatever ranks first on Google?
No, though the overlap is real — several engines retrieve through conventional search infrastructure, so ranking well helps you get retrieved. The divergence happens at passage selection: engines cite the paragraph that answers best, which is often not on the top-ranked page. Strong rankers with no quotable passages get retrieved and then passed over.
Why does a competitor get cited when my site ranks higher?
Almost always a passage-level loss. Ranking measures your page's overall standing; citation goes to the specific paragraph that resolves the question most directly. If your competitor opens a section with a clean forty-word answer and your equivalent material is scattered across a long page, the engine quotes them and merely retrieves you.
Do ChatGPT, Claude, Perplexity, and Gemini all decide citations the same way?
The pipeline shape is shared — interpret, retrieve, select, ground — but implementations differ: which crawlers they run, which indexes they draw on, how visibly they cite. That is why fixes at the pipeline level transfer across engines, while any tactic promising one specific engine's citation is overfitting to details that shift.
Can I pay an engine or a vendor to be cited?
No AI engine currently sells citations in organic answers, and any vendor guaranteeing specific citations is promising an outcome nobody controls — generated answers vary run to run even with no changes on the web. What you can legitimately buy or build is pipeline work: access fixes, structure, question-shaped content, and honest measurement.
How long until fixes show up in AI answers?
It depends on the stage. Retrieval fixes take effect as engines recrawl — typically days to a few weeks. Passage-level rewrites need recrawling plus recurring sampling to confirm, so think in weeks. Grounding-level consistency compounds slowest. Judge movement across a fixed question set over a month or more, never from one answer.
How do I find out which stage is filtering my site out?
Work backward from the visible end. Ask the engines ten real buyer questions and log whether you appear; if not, check retrieval (can crawlers fetch your pages?), then passage quality (does any paragraph answer those questions directly?), then consistency. A free AI visibility check runs the measured version of this in about a minute.
The bottom line
AI engines decide citations the way a careful researcher assembles sources: figure out what is being asked, gather what can be gathered, pull the clearest statements, and credit the ones sturdy enough to stand on. Your site is competing at all four stages at once, and it is only as visible as its weakest one. The map beats the tips: diagnose the stage, apply its fix, and re-measure — in that order.
Find your weakest stage today: run a free AI visibility check and get a scored read on your access, structure, and presence in about a minute.
How visible is your site in AI search?
Check your AI visibility score in seconds — free, no account needed.
Check your score