How to track LLM citations: the method behind an 87 to 850 result
The short answer
Build a fixed set of 30 to 50 real buyer questions, run them monthly across the assistants your buyers use in clean logged-out sessions, and log four things: whether you were named, your share of all vendor mentions, how you were framed, and which domains the assistant cited. Repeat each prompt several times per run, since responses are non-deterministic.
When we say we took a client from 87 LLM citations to over 850, the obvious question is how we counted.
It is a fair question, and most AI visibility claims do not survive it. This is the method, including what it cannot tell you.
Why your analytics tool cannot do this
The instinct is to look for LLM traffic in Google Analytics. That measures the wrong thing.
Referral traffic only counts the buyers who clicked through. Most AI-assisted research never produces a click. The buyer asks which vendors to consider, reads the answer, notes three names, and goes directly to the two they recognise. You were either named or you were not, and no click ever happened either way.
Being named is the outcome. The click is a side effect. If you measure only clicks, you are measuring a fraction of the value and you will conclude the channel is small.
So citation tracking is not an analytics problem. It is a sampling problem. You cannot observe every conversation, so you construct a representative set of questions and measure your presence across it, the same way polling works.
Step 1: build the prompt set
Thirty to fifty prompts. This is the whole foundation, and a lazy prompt set produces numbers that mean nothing.
Write questions, not keywords. Nobody types "abm platform" into an assistant. They type a paragraph about their situation.
Cover four shapes, roughly evenly:
| Shape | Example |
|---|---|
| Category | "What tools help B2B teams run account-based marketing?" |
| Problem-first | "Our outbound reply rates have collapsed and we sell six-figure deals. What should we change?" |
| Comparison | "What are the main differences between the leading ABM platforms?" |
| Alternatives | "What are good alternatives to [known incumbent] for a 40-person SaaS company?" |
Two rules that decide whether the set is any good:
Never include your brand name in a prompt. Asking "is NomiOS good for ABM" measures nothing. The whole question is whether you appear *unprompted*. A branded prompt guarantees a mention and quietly inflates every number you report.
Write them as your buyer speaks, not as your marketing speaks. If your category page says "revenue orchestration" and your buyers say "getting more meetings", the buyer's phrasing goes in the set. You are sampling real demand, not your own positioning.
Freeze the set once written. Changing prompts between runs destroys comparability, and the temptation to swap in prompts you now rank for is strong. Version it. If you must add prompts, start a second cohort and track it separately.
Step 2: run it on a schedule
Monthly across the assistants your buyers actually use. For most B2B SaaS that is ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews.
Practical notes that change the numbers more than people expect:
- Log out, or use a clean session. Personalisation and memory will show you a flattering result that no stranger sees.
- Run each prompt more than once. Responses are non-deterministic. A single run is a sample of one. Three runs and record how many named you.
- Keep the raw response text, not just a yes/no. The framing matters as much as the mention, and you will want to re-read these later.
- Same day of the month, roughly the same time. These systems change under you constantly.
Monthly is enough. Weekly produces noise you will misread as trend.
Step 3: the four numbers
| Metric | Definition | What it tells you |
|---|---|---|
| Citation rate | % of prompts where you are named at all | Raw visibility. The headline number. |
| Citation share | Your mentions as a share of all vendor mentions | Competitive position. Rising while citation rate is flat means rivals are losing ground. |
| Framing | Named as a leader, a viable option, or a caveat | Quality of the mention. Being named as "a smaller alternative" is not the same win as being named first. |
| Source mix | Which domains the engines cite when answering | The only one that tells you what to do next. |
Source mix is the most actionable metric and almost nobody tracks it. Every other number tells you where you stand. Source mix tells you where to go.
Log every domain the assistants cite across your prompt set, then count. If a review platform appears in 60% of answers and your own domain in 10%, you have just been handed your workplan. Investing in your blog when the models are reading G2 and Reddit is optimising a surface that is not being read.
This is the same logic as reading B2B buying signals properly: the dashboard everyone stares at is rarely the one that changes a decision.
Step 4: instrument the referral traffic anyway
Not as your primary metric, but because it tells you something citations cannot: whether the mention converted.
LLM referrers are identifiable in analytics. Volume will look disappointing. Judge it on conversion rate rather than sessions, because these visitors arrive pre-qualified by a recommendation rather than cold from a search result. A few hundred visitors that convert like referrals from a trusted partner is a different asset to a few thousand that bounce.
What this method cannot tell you
Stating the limits, because a measurement method that oversells itself is worse than none.
It is a sample, not a census. You are measuring 40 questions out of an unbounded space. A rising citation rate is real evidence, but it is not "we appear in 62% of all conversations about our category". Report it as what it is.
It cannot attribute revenue. You will not connect a citation to a closed deal. What you can do is ask new pipeline how they first heard of you and watch the share that says an AI assistant. That is self-reported and imperfect, and it is still the best signal available.
It is volatile. Answers change week to week with no change on your side. Model updates, index refreshes and retrieval changes all move it. Read the quarterly trend and ignore the monthly wobble. Reporting a bad month as a regression will get you chasing noise.
Position within an answer is soft. Being named third in a list is worth less than being named first, but the ordering is not stable enough to optimise directly. Track framing instead.
What moved the number for us
On the engagement that produced 87 to 850+, the tracking itself did not create the lift. It told us where to spend.
The source mix showed the engines were overwhelmingly citing third-party properties rather than the client's own domain. That single observation redirected the budget away from publishing more posts and toward third-party surface and original data, which is where the increase actually came from. We cover the levers in the parent piece on answer engine optimization for B2B SaaS, and the specific mechanics of citation in how to get cited by ChatGPT.
Take the baseline before you start the work. This is the mistake that cannot be undone later. A team that spends two quarters on AI visibility and then starts measuring has no way to prove any of it worked. Half a day of prompt-set construction now buys you the ability to demonstrate the outcome for years.
Frequently asked questions
How do you track LLM citations?
Build a fixed set of 30 to 50 real buyer questions, run them monthly across the assistants your buyers use in clean logged-out sessions, and log four things: whether you were named, your share of all vendor mentions, how you were framed, and which domains the assistant cited. Repeat each prompt several times per run, since responses are non-deterministic.
Can you see LLM traffic in Google Analytics?
You can see referral traffic from assistants, but it captures only the buyers who clicked through. Most AI-assisted research produces a named shortlist and no click, so referral data understates the channel substantially. Use it to judge conversion quality, not overall visibility.
How many prompts do you need to track AI visibility?
Thirty to fifty is enough for a stable signal. Fewer than twenty makes month-to-month movement indistinguishable from noise. Spread them across category, problem-first, comparison and alternatives-shaped questions, and never include your own brand name in a prompt.
How often should you measure LLM citations?
Monthly. Weekly produces volatility you will misread as trend. Read results as a quarterly trend line, because answers shift week to week from model and index updates you do not control.
What is the most useful AI visibility metric?
Source mix, meaning which domains the assistants cite when answering your prompt set. Citation rate tells you where you stand. Source mix tells you which properties to invest in next, which is the only metric that changes what you do on Monday.
---
*NomiOS is RZLT's GTM and ABM engine. Point it at a target and get back finished, branded work built on a real read of that company.*
[See how NomiOS works →](https://nomios.rzlt.io)
Questions
Frequently asked
- How do you track LLM citations?
- Build a fixed set of 30 to 50 real buyer questions, run them monthly across the assistants your buyers use in clean logged-out sessions, and log four things: whether you were named, your share of all vendor mentions, how you were framed, and which domains the assistant cited. Repeat each prompt several times per run, since responses are non-deterministic.
- Can you see LLM traffic in Google Analytics?
- You can see referral traffic from assistants, but it captures only the buyers who clicked through. Most AI-assisted research produces a named shortlist and no click, so referral data understates the channel substantially. Use it to judge conversion quality, not overall visibility.
- How many prompts do you need to track AI visibility?
- Thirty to fifty is enough for a stable signal. Fewer than twenty makes month-to-month movement indistinguishable from noise. Spread them across category, problem-first, comparison and alternatives-shaped questions, and never include your own brand name in a prompt.
- How often should you measure LLM citations?
- Monthly. Weekly produces volatility you will misread as trend. Read results as a quarterly trend line, because answers shift week to week from model and index updates you do not control.
- What is the most useful AI visibility metric?
- Source mix, meaning which domains the assistants cite when answering your prompt set. Citation rate tells you where you stand. Source mix tells you which properties to invest in next, which is the only metric that changes what you do on Monday.