How to get cited by ChatGPT: the retrieval mechanics
The short answer
Four things in order: make sure AI crawlers can read your site and your content is server-rendered, restructure content into self-contained extractable claims, publish original data that makes you the origin of a fact, and build presence on third-party properties like review platforms and practitioner communities, which carry the largest share of citations.
The question is usually asked as a content question. It is a retrieval question, and the difference determines whether anything you do works.
You are not persuading a model to like your brand. You are trying to be the source it reaches for when it needs to support a claim, and then the passage it can actually lift. Two separate hurdles, and most advice addresses neither.
What happens when someone asks a buying question
Simplified, but accurate enough to act on.
1. The assistant decides whether it needs external sources. For most vendor questions it does, because the answer changes over time.
2. It issues search queries, usually several, and usually not the user's exact wording.
3. It retrieves a set of candidate pages, heavily weighted toward what already ranks.
4. It reads them and extracts passages that support an answer.
5. It composes an answer and attributes claims to the sources it actually used.
Two implications fall straight out of this.
Step 3 means ranking still matters enormously. The candidate set is largely drawn from search results. If you do not rank for the queries the assistant generates, you are not in the pool, and nothing about your page content can rescue that. This is the single most common misunderstanding in the field, and it is why defunding SEO to fund AEO backfires, as covered in GEO vs SEO.
Step 4 means extractability decides the rest. Being in the pool gets you read. Being liftable gets you cited.
Hurdle one: be retrievable
Unblock the crawlers. Check `robots.txt` today. A meaningful share of B2B SaaS sites block GPTBot, ClaudeBot, PerplexityBot, Google-Extended or CCBot, almost always from a blanket rule added during the 2023 scraping panic that nobody revisited. Commissioning AEO work while blocking the crawlers is paying to stay invisible.
Server-render. If your content only appears after JavaScript hydration, many crawlers see an empty shell. Verify with view-source or `curl`, never with devtools, which shows you the hydrated DOM and will give you a false pass.
Rank for the query behind the question. Assistants reformulate. Someone asking "what should we use if our outbound stopped working" generates queries about outbound tooling and reply rates. Rank for the underlying informational queries, not the conversational phrasing.
Hurdle two: be liftable
A passage gets cited when it is self-contained, factually complete, and attributable.
Most B2B content fails this for a specific reason: it is written to flow. The key claim is distributed across three paragraphs and depends on the two above it. Excellent to read top to bottom. Useless when a model needs one block that stands alone.
The rewrite patterns that work:
Answer first, elaborate second. The direct answer in the first sentence under the heading, then the explanation. Models lift the first block under a matching heading far more often than anything below it.
One complete claim per paragraph. No claim should require the previous paragraph to make sense.
Headings phrased as real questions. How buyers ask them, not how your category writes them.
Numbers with their conditions attached. "10x demand lift" is unusable, because a model cannot responsibly cite an unqualified figure. "10x lift in combined inbound and outbound demand across 1,400+ target accounts" survives extraction because the qualification travels with the number.
Comparison tables. Disproportionately extracted, because the structure maps directly onto comparative questions, which is what most vendor research is.
The test: pull any paragraph out of context. Does it still assert something complete and checkable? If not, it will not be cited no matter how well the page ranks.
Hurdle three: be corroborated
The hurdle nobody plans for.
A model weighting sources for a "best X" question discounts self-interested ones. Your own site is a party to the question. Review platforms, comparison articles, practitioner threads and customer write-ups read as independent.
This is why a brand's own domain typically accounts for a minority of its citations, with review sites and forums carrying a much larger share.
The uncomfortable consequence: the highest-leverage AEO work happens on properties you do not own. Review platform presence with recent detailed reviews. Getting into the comparison content that already ranks. Genuine participation in practitioner communities, which are unusually good at detecting and punishing marketing.
Slow, largely uncontrollable, and the biggest single share of the outcome.
The strongest lever: be the origin of a fact
Models preferentially cite the source of a specific claim rather than a page repeating it.
If you publish the only measured number on a topic, you become the default citation for that topic, and it compounds: other pages quote you, models see corroboration across sources, and your position strengthens without further work.
This is the one lever most teams cannot pull, because it requires having something to say that nobody else has. It does not require a research department. A benchmark from your own product data, an honest before-and-after from your own work, a survey of your customers, or a teardown with real measurements all qualify.
What does not qualify is a roundup of other people's statistics, which is what most "data-driven" B2B content actually is. Aggregating other people's numbers makes you a repeater, and repeaters do not get cited. The originals do.
What does not work
Worth stating plainly, because these consume budget:
- Keyword density. Retrieval is semantic. You are tuning a mechanism that is not operating.
- Publishing volume. Twenty pages saying the same thing do not increase citation probability. One page with a citable claim beats all twenty.
- `llms.txt`. A proposed standard for pointing models at your content. Adoption by major engines remains unconfirmed and Google has indicated it does not use it. An hour to add, possibly useful later, not a strategy.
- Schema markup on its own. Genuinely valuable for Google AI Overviews, which lean on structured data. Much less influential for chat assistants. Worth doing, wrong thing to lead with.
- Mentioning your brand more often on your own pages. Models are not counting brand mentions on your domain. Density on a self-interested source is not evidence.
The order to work in
1. Unblock crawlers and confirm server-rendering. Days. Everything else depends on it.
2. Restructure your best existing pages into extractable claims. One to three months. Cheaper and higher-return than new content.
3. Publish original data. The strongest lever you control directly.
4. Build third-party presence. Slowest, largest share of outcome, compounds over quarters.
Baseline your citation rate before step 1. A team that does two quarters of this work and then starts measuring cannot demonstrate any of it. Half a day of setup buys years of provable outcomes, and the method is in how to track LLM citations.
Frequently asked questions
How do you get cited by ChatGPT?
Four things in order: make sure AI crawlers can read your site and your content is server-rendered, restructure content into self-contained extractable claims, publish original data that makes you the origin of a fact, and build presence on third-party properties like review platforms and practitioner communities, which carry the largest share of citations.
Does ranking on Google affect AI citations?
Yes, substantially. Assistants build their candidate source set largely from search results, so pages that do not rank are rarely in the pool to be cited at all. The pages most cited by LLMs are largely the pages that already rank well.
What makes a passage citable by an LLM?
Being self-contained, factually complete and attributable. It must assert something checkable without depending on surrounding paragraphs. Answer-first structure, one claim per paragraph, question-shaped headings, numbers with their qualifying conditions attached, and comparison tables all extract well.
Why doesn't ChatGPT cite my own website more?
Because a model weighting sources for a vendor question discounts self-interested ones. Your own domain typically accounts for a minority of your citations, with review platforms and community forums carrying a much larger share, since they read as independent corroboration.
Does llms.txt help you get cited?
There is no confirmed evidence that major engines use it, and Google has indicated it does not. It costs about an hour to add and may become useful later, but it is not a strategy and should not displace crawler access, content structure or third-party presence.
---
*NomiOS is RZLT's GTM and ABM engine. Point it at a target and get back finished, branded work built on a real read of that company.*
[See how NomiOS works →](https://nomios.rzlt.io)
Questions
Frequently asked
- How do you get cited by ChatGPT?
- Four things in order: make sure AI crawlers can read your site and your content is server-rendered, restructure content into self-contained extractable claims, publish original data that makes you the origin of a fact, and build presence on third-party properties like review platforms and practitioner communities, which carry the largest share of citations.
- Does ranking on Google affect AI citations?
- Yes, substantially. Assistants build their candidate source set largely from search results, so pages that do not rank are rarely in the pool to be cited at all. The pages most cited by LLMs are largely the pages that already rank well.
- What makes a passage citable by an LLM?
- Being self-contained, factually complete and attributable. It must assert something checkable without depending on surrounding paragraphs. Answer-first structure, one claim per paragraph, question-shaped headings, numbers with their qualifying conditions attached, and comparison tables all extract well.
- Why doesn't ChatGPT cite my own website more?
- Because a model weighting sources for a vendor question discounts self-interested ones. Your own domain typically accounts for a minority of your citations, with review platforms and community forums carrying a much larger share, since they read as independent corroboration.
- Does llms.txt help you get cited?
- There is no confirmed evidence that major engines use it, and Google has indicated it does not. It costs about an hour to add and may become useful later, but it is not a strategy and should not displace crawler access, content structure or third-party presence.