How it works
How an AI answer engine actually works
When someone asks ChatGPT, Perplexity or Google's AI a question, the engine does two separate jobs before it answers, and each one rewards something different. Get that split clear and you understand most of what it takes to be named in the answer.
Here's what those two stages are, what each one actually rewards, and — honestly — where the evidence is solid and where we're still testing our own bets.
Last updated
Stage one: retrieval, getting into the running
First, the engine gathers a set of candidate pages that look relevant to the question, pulling them from its own index or a live web search. It casts a wide net. Nothing gets quoted yet; this stage is only about being considered at all.
Retrieval rewards coverage. The more pages you have matching more of the questions people ask, the more often you land in that candidate pool. If you've published a lot, this is where the sheer volume helps — and it genuinely does help here.
Stage two: selection, being the one it uses
Then the model reads those candidates and picks the few it actually builds the answer from. It's looking for the source it judges most credible and most specific to that exact question. Raw volume counts for nothing at this point.
What counts here is substance: a clear point of view, real detail, a named human behind the writing, and the same claims turning up elsewhere to back them. Retrieval gets you considered. Selection decides who makes the answer. They're two different tests, won in two different ways.
What “being in the answer” really means
“Showing up” isn't one thing either. When people say they want to be “quoted” by AI, they usually mean something looser than the model reproducing their exact words. What matters is your business or your name appearing inside the answer a buyer sees. That happens in a few forms, worth separating because they're worth very different amounts:
- Recommended by name.The answer puts you forward as an option: “for this, tools like yours, Y and Z.” Someone asked who they should use, and the AI named you. It's the closest thing to a warm lead.
- Named or mentioned.Your name appears in the answer text, whether or not it's linked: “their approach to this is…”. You're in the conversation.
- Cited as a source.The engine lists you underneath the answer: the numbered footnotes in Perplexity, the source links in ChatGPT, the cards in Google's AI Overviews. This is also where you can earn a click, and you can watch it arrive tagged
utm_source=chatgpt. - Influenced but invisible.Your content shaped the answer, but you're not named or linked. Common, and the weakest outcome, because a buyer sees no credit at all.
The first three are what selection is really deciding: which sources make it into a form a reader can see. Being retrieved makes you a candidate. Being named, cited or recommended means you survived selection and showed up.
Why we concentrate, rather than spread thin
There are two ways to use content for AI visibility, and both produce plenty of it over time. What separates them is where all that depth goes.
One approach spreads thin: a lightweight page for each keyword, across hundreds of them, so you're present on as many questions as possible. Ours concentrates — we pile a lot of substantial, person-voiced content onto the smaller set of questions your buyers actually ask AI, so you become the recognised authority on them. Especially across a team, that's a high volume of content. It's just aimed at a focused target rather than sprayed across everything.
That maps straight onto the two stages. Spreading wide helps you get retrieved on more queries. Concentrating deep helps you get selected on the ones that matter, because a business can't be the authority on 400 topics at once. Spread that thin and you're a weak candidate on every one of them. Concentrate, and you can be the specific, quotable source on the handful of questions that turn into customers.
There's a cadence point underneath it too. Dumping hundreds of pages at once reads like a content farm, risks Google's scaled-content-abuse flag, and goes stale fast. Building steadily, in your own people's voices, on topics you deepen over time, builds a track record and stays current — which is exactly what selection rewards.
What's solid, and what we're still testing
The two-stage picture — retrieval then selection — is genuinely how these systems work. That part isn't a marketing claim.
That selection favours substance, citations and named humans is what the evidence points to. Princeton's GEO research found that adding citations, quotations and statistics to a page measurably increased how often AI answer engines pulled it into their responses. Ahrefs' analysis of AI citations shows that ranking well on Google and getting cited by AI are two different jobs, won in different ways.
The bigger claim — that concentrated depth beats broad coverage for citation specifically — is our reasoned bet: tested in the open, not settled fact. We think the mechanism points that way, and we'll say so plainly if the results ever tell us otherwise.
One honest limit sits underneath all of it: plenty of AI answers don't credit their sources visibly at all, so good content sometimes still lands in that invisible fourth bucket. That's why our promise is that you're more likely to be the named source, never that a citation is guaranteed.
Frequently asked questions
What’s the difference between being retrieved and being cited?
Retrieval is the first stage, where an AI answer engine gathers candidate pages that look relevant to a question. Citation happens in the second stage, selection, where the model picks the few sources it actually uses in the answer. Being retrieved makes you a candidate; being cited means you were chosen.
Does publishing more pages get me cited more?
More pages help you get retrieved more often, because coverage is what the retrieval stage rewards. But volume does little for selection, which decides who actually appears in the answer. Selection rewards substance, a clear point of view and corroboration, so more thin pages rarely means more citations.
Does an AI have to quote my exact words for it to count?
No. What matters is your business or your name appearing in the answer a buyer sees: recommended, named or cited as a source. Being reproduced word for word happens occasionally, but it isn’t the thing that drives value.
Can Ghostart guarantee I’ll be cited by AI?
No, and nobody honestly can. Many AI answers don’t credit their sources visibly, so even strong content is sometimes used without a mention a reader can see. Ghostart works to make you the more likely named source on the questions that matter, and reports honestly on whether it’s moving.
Is “depth beats breadth” a proven fact?
The two-stage mechanism is well established. That concentrated depth beats broad coverage for citation specifically is a reasoned hypothesis Ghostart is testing in the open, pointed to by findings from Princeton’s GEO research and Ahrefs, not a settled result.
See where you stand in AI answers
Add your website and Ghostart checks whether ChatGPT, Claude, Gemini, Perplexity and Google's AI name you for the questions your customers are actually asking, then helps you publish the content that earns the mention.
Sources
- Aggarwal et al., “GEO: Generative Engine Optimization” (Princeton, IIT Delhi, Georgia Tech & Allen Institute for AI; KDD 2024) — arxiv.org/abs/2311.09735
- Ahrefs' analysis of AI citations versus search rankings — ahrefs.com/blog/search-rankings-ai-citations
- The retrieval-then-selection description reflects how retrieval-augmented answer engines are documented to work.
ChatGPT, Claude, Gemini, Google and Perplexity are trademarks of their respective owners (OpenAI, Anthropic, Google and Perplexity AI). Ghostart is not affiliated with, endorsed by, or sponsored by any of them.