How to get cited in AI answers: what ChatGPT, Perplexity and Google AI Overviews actually select

In short
What gets cited is not the best page, it is the passage that lifts out cleanest. This overview sorts the levers by evidence: what provably works, what is only claimed, and what is demonstrably useless. What's left is the uncomfortable finding that visibility in AI answers is less a technical question than a question of whether the content actually exists. scaile produces exactly that content.
Key takeaways
The most reliable levers are unspectacular: reachability for AI crawlers, current content, sourced claims, and presence on the sources these systems already trust.
Evidence density is measurable. Content with citations, sources and statistics gained up to around 40 percent more visibility in generative answers in controlled tests, with the biggest effect for pages that weren't ranking well before.
llms.txt is currently useless. A study of 137,000 websites found that 97 percent of these files were never fetched, and Google states outright that it doesn't use the file.
Freshness beats age, but only barely: 75 percent of cited pages were updated within the last year, while the average cited document is nearly three years old. The winners are old pages that are kept up.
In the end, what decides is whether the answer to your customer's question is actually written. scaile writes exactly those answers, checks every claim, and puts them in front of your team for approval, within 14 days of onboarding.
To get cited in AI answers, you need three things: your pages have to be reachable by the AI providers’ crawlers, they have to answer a specific question completely and with evidence, and people have to be talking about you outside your own domain. The third point surprises most people, but it’s the biggest one: for brand-related questions, only around 23 percent of citations come from the brand’s own website, and nearly half come from third-party sources.
What does little for you, by contrast: backlinks counted the classic way, big PR names, and the much-discussed llms.txt file.
All three of the things that do work are production, not configuration, which is why most programmes stall after the technical fixes. scaile runs them as one managed pipeline: the commercial question set mined first, research from your own knowledge base, every fact checked, your team approving before anything publishes, and published pieces refreshed at the same URL as they age. Live in 14 days.
How AI systems choose their sources
Generative systems don’t answer every question with a search. On ChatGPT, only about 18 percent of all conversations trigger a web search at all, measured across more than 730,000 conversations. The rest is answered from what’s already inside the model. For your visibility, that means there are two ways in, the index and the model itself, and they work differently.
When a search does happen, the system breaks the question into several sub-questions and retrieves passages for each one. Google officially calls this process query fan-out and uses it in both AI Overviews and AI Mode. The consequence is that your page doesn’t need to rank for the question that was asked, it needs to rank for one of the sub-questions the system derives from it. On Google, only 38 percent of cited pages now come from the top 10 of the actual query, while 31 percent sit beyond position 100.
Important for practice: ChatGPT no longer relies mainly on Bing. Overlap with Bing results fell from 26 to 8 percent, while overlap with Google rose to 33 percent. Anyone who ties their AI visibility to a single search engine is measuring the wrong thing.
The levers, ranked by evidence
Proven and documented
Reachability for the AI crawlers. OpenAI documents its own bots, including OAI-SearchBot for search. Blocking them in robots.txt or through a firewall doesn’t make you invisible, it makes you not present at all. The most common mistake in practice: the bot is allowed, but the CDN blocks it by IP range.
Indexing and snippet eligibility on Google. For AI Overviews, Google states exactly one hard requirement: the page has to be indexed and eligible for a snippet. Google also states explicitly that no new machine-readable files and no special markup are needed.
Evidence density in the text. In controlled tests, citations, source references and statistics increased visibility in generative answers by up to around 40 percent. The effect is unevenly distributed: pages that were previously sitting around position five gained by far the most, while the already-leading page lost share. For smaller providers, that’s the best news in this whole topic.
Well supported by third-party data
Freshness. Three-quarters of the pages cited by AI systems were updated within the last year, with a median of 5.6 months. At the same time, the average cited document is nearly three years old. Together, those two facts point to one clear instruction: maintaining existing, established pages beats constantly publishing new ones. This is exactly why scaile updates published pieces at the same URL instead of creating new ones.
Presence outside your own domain. For brand-related questions, only around 23 percent of citations come from a brand’s own content, 48 percent from earned mentions, and the rest from commercial sources. Across all queries, the share of a brand’s own pages is higher, around 57 percent. The lesson isn’t to neglect your own website, it’s to know that part of the work happens outside it, in trade publications, comparison lists, forums and directories.
Answer-ready structure. A passage gets cited when it can be lifted out without context. In practice that means the answer directly under the question, headings that are real questions, and evidence in the same paragraph as the claim.
Disproven or ineffective
llms.txt. An analysis of 137,000 websites found that 97 percent of the deployed files were never fetched. Google states plainly in its own documentation that Search doesn’t use it. The file doesn’t hurt, but it replaces none of the other measures.
Backlinks as the main factor. The correlation between classic backlink metrics and mentions in AI answers is close to zero in the available studies. Links remain relevant for classic search, they’re just not the lever here.
Structured data as a citation guarantee. Numbers circulating online, like a threefold higher citation chance from FAQ markup, don’t hold up in any credible study, and Google explicitly disputes the underlying idea. Schema helps indexing and never hurts. It doesn’t buy citations.
Registering with ChatGPT. There’s no directory you enter your company into. The only real submission path is product data for ChatGPT Shopping. Since February 2026 there has been advertising in ChatGPT, but it’s kept separate from the answers and doesn’t change them.
What this means for your content
Add up the proven levers and a simple order emerges. First, technical reachability, which is a day’s work. Then the question of whether a good answer to the topics you want to be named for actually exists on your website. And finally, the work outside your own domain.
The second point is where most programmes fail. Not because it’s hard to understand, but because it’s work: dozens of concrete customer questions, each with a complete, sourced, current answer, in your own company’s language. A measurement tool shows you this gap cleanly, it doesn’t close it.
That’s exactly what scaile is built for. The engine surfaces the questions your customers ask right before a purchase decision, researches from your knowledge base and current sources, checks every fact, and puts every article in front of your team for approval. Along the way it learns from every edit, and we explicitly recommend adding your own knowledge: the number, the customer case, the judgement only you have. That’s the part no competitor can copy, and for AI systems it’s the reason your page gets cited instead of someone else’s.
The free AI Visibility Check shows where you stand today. If you are already visibly missing from answers, this diagnosis narrows the cause down in half an hour. And because the content this produces is editorial by nature, it is worth knowing where the EU AI Act’s labeling duty lands: human-reviewed articles are exempt.
FAQ
How do I get into ChatGPT’s answers?
Allow the OpenAI crawlers, make sure you have sourced, current answers to your customers’ questions, and get mentioned on third-party pages ChatGPT already cites. There’s no submission form.
Can I register my company with ChatGPT?
No. A real submission path exists only for product data under ChatGPT Shopping. Everything else comes from content and mentions.
Does llms.txt do anything?
Not as things stand. 97 percent of these files were never fetched in one study, and Google says it doesn’t use them.
How long does it take before a page gets cited?
For a tightly scoped set of questions, first movement shows up within weeks, and a solid share of visibility takes months. Updated established pages move faster than brand-new ones.
Does the same thing count for AI answers as for Google?
Partly. Indexing and technical cleanliness are a shared prerequisite, but source selection differs significantly: on Google, only 38 percent of cited pages now come from the top 10 of the query asked.
Do I need to be active on Reddit?
Useful, but not required, and it varies sharply by system. Reddit is the single most-cited source across all systems combined, while on Microsoft Copilot it barely factors in.
We know the three levers. Who actually pulls them every month?
That is the part that decides the outcome, and it is a cadence problem: the answers have to keep being written, kept current and kept sourced. scaile runs it as managed infrastructure, with a named AI Search Strategist, research from your knowledge base, fact-checking on every claim and your team approving before publication, live in 14 days. Building Radar doubled its qualified inbound leads in 90 days on it.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Lewis et al. Why answers are built from retrieved passages.
- GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024. The effect of evidence density and the uneven effect favoring lower-ranked pages.
- Google Search Central: AI features and your website. Indexing and snippet eligibility as the only requirement, no special files needed, query fan-out.
- OpenAI: crawler overview. OAI-SearchBot, GPTBot and ChatGPT-User.
- Ahrefs on llms.txt adoption. 97 percent of files were never fetched.
- Ahrefs on rankings and AI citations. Only 38 percent of cited pages sit in the top 10 of the query.
- Ahrefs on the age of cited content. Nearly three years on average.
- Omniscient Digital on where AI citations come from. Share of owned, earned and commercial sources for brand-related questions.
- Profound, analysis of more than 730,000 ChatGPT conversations and 11.8 billion citations. Share of queries triggering search, the shift from Bing to Google, source distribution. Published via tryprofound.com.



