You assume that getting cited by AI is a content problem: write clearer pages, add schema, make sure ChatGPT can read what you already publish. Those steps fix being read. They do not fix being found, and the evidence points to two different starting points for each.
Most citation advice treats an AI engine like a search engine that reads your site more carefully. It is not. It is a system that goes looking for information on its own terms, decides who is worth reading, and only then writes an answer. Fixing your own pages addresses the last of those three steps. For where this sits inside the wider practice, see What Is Generative Engine Optimisation? The evidence below is mostly about the first two steps, not the last.
The engine searches before it reads
Before an AI engine answers a buying question, it usually runs its own searches first, not the words the person typed. It writes several queries, runs them, and works from what comes back. In Study 02, SearchIntel captured 12,264 of those machine-written searches: ChatGPT, Claude and Gemini, 177 category questions in the UK and Ireland, each run five times, captured on 23 and 24 July 2026. On 52% of the 89 questions in one market and 66% of the 88 in the other, the engine's own searches named companies before it had read a single page.
That first round of searching is not neutral. It is already a shortlist. Being named in it predicts what happens next: across the 438 cases in the study where a competitor's name turned up in the machine's own searches, that competitor was mentioned in the answer 87% of the time. The other 13%, 56 cases, were named in the search and then left out of the answer. Named is not the same as recommended. It is most of the way there. A brand that never appears in that first round is not competing on the merits of its page. It is not competing at all, because the search that would have found the page was never run.
One company in the study made the point starkly. Its own searches by the engine turned it up on 45 of the 88 home-market questions, 222 times across the five runs, and once across the border on the other market's 89 questions. Same company, same website, same service. Asked about it directly in the weaker market, the engine described it accurately. Known, described correctly, not considered. No amount of on-page content closes that gap, because the gap opens before the engine reads a page at all. Whatever changed between the two markets happened somewhere the brand's own site cannot see: the sources that make an engine associate a name with a category in the first place.
Where the citations actually come from
Once the engine has a set of results, it reads them, and what it reads is mostly not the brand's own site. Muck Rack's May 2026 edition of What Is AI Reading?, built on more than 25 million links cited by ChatGPT, Claude and Gemini, put earned media at 84% of AI citations, a figure that has held between 82% and 89% across all three editions of that study since July 2025.
| Source type | Share of AI citations | Source |
|---|---|---|
| Earned media (trade press, reviews, forums, independent coverage) | 84% (82-89% across three editions since July 2025) | Muck Rack, May 2026 |
| Of which, journalism specifically | 27% | Muck Rack, May 2026 |
| Paid or advertorial content | 0.3% | Muck Rack, May 2026 |
The breakdown is not published at the market or category level we would need to plan a UK campaign from it, so read it as the shape of the whole, not a formula for one industry. What it rules out is clearer than what it proves. Paid placement is not the route: 0.3% of citations trace to it, which is close to none. A brand's own domain is not where most of the reading happens either. Study 02 reaches the same conclusion from a different angle: machine-written searches on one side, citation sources on the other. Two different studies, two different measurements, one direction.
It is worth being precise about what "journalism" means inside that 84%. Muck Rack reports it as 27% on its own, which reads as a subset of the wider earned-media figure rather than an additional slice, articles specifically, not the press mentions folded into review sites or forum threads. The distinction matters if you are deciding where to put effort: a mention in a trade publication and a mention in a Reddit thread both count as earned media, but they are not interchangeable, and the study does not tell us the split between them.
The four things that follow
Getting cited is not a checklist you finish once. It is four separate jobs, and the order matters more than any of them individually. Do them out of order and you spend the budget on the step that helps least right now.
- Be a candidate before you are a citation. Study 02 shows the engine's own searches, not your content, decide who gets a chance to be read. Consistent, accurate mentions of your brand on the sources the engine already trusts are what put you in that first round. This is upstream of everything else on this list, and it is the step most citation advice skips, because it is not something you can fix on your own site.
- Be present on the sources it reads. Muck Rack's figures say where that presence needs to be: earned media, not your own domain. Trade press, review platforms, forums and independent coverage, the places someone writes about you without being asked to. A mention that reads as a genuine account from an independent source is worth more here than a mention you paid for or wrote yourself. The 0.3% figure for paid content confirms it.
- Be extractable once the engine arrives. A page that answers the question plainly, states what you do, for whom, and since when, and carries accurate structured data, is easier for an engine to quote from. This step only matters after the first two. A well-built page that no search ever turns up is not read, and a citation-rich brand with a confusing page still loses some of what the first two steps earned it.
- Measure per platform, not once. Study 01 found that of the brands recommended across six platforms on 100 UK buying questions, only 5.3% were named by all six, and 46.1% were named by exactly one. The average recommended brand appeared on 2.26 of the six. Two platforms agreed on the brands they recommended just 31.9% of the time on average, and the two flagship chatbots, ChatGPT and Claude, agreed least, at 22.2%. Even inside Google, AI Mode and AI Overviews cited the same URL only 13.7% of the time, across 540,000 query pairs (Ahrefs, December 2025). A check on one platform tells you about that platform, and nothing about the other five.
Where that leaves a marketing team with a limited budget is a sequencing question rather than a spending question. The candidate step and the presence step both run through the same channel, third-party sources, so they can be worked together: the same trade press pitch or review-platform push that builds earned-media presence is also the input the engine's own searches draw on. The extraction step is comparatively cheap and can happen in parallel, on the pages most likely to be found once the first two steps start working. The measurement step is not a later phase. It has to run from day one, because without it there is no way to tell whether the first three are doing anything at all.
Why the volume makes the effort worth it
AI Mode alone passed one billion monthly users in its first year, Google said at its May 2026 product event (Google, May 2026). That is one surface out of six this applies to. The question for most brands is no longer whether AI-generated answers are being read at meaningful volume. It is whether their brand is in the answer when someone asks.
What we cannot yet say
Our two studies are UK and Ireland: one on 100 questions, one on 177, captured on the model versions current in June and July 2026. Muck Rack's figures are global and are not broken out by market or category. Neither tells you the exact citation ratio for your industry or your country, only the direction recorded so far, and that direction is consistent across both of our studies and Muck Rack's three editions. We do not have a UK-specific citation-share breakdown by category, and we are not going to publish one we have not run. If your category behaves differently, that is a real possibility we have not tested for, not a claim we are making either way. The same caution applies to the Ahrefs figure: it measures URL overlap between two Google surfaces, not brand overlap, and we have not re-run it ourselves.
None of this is a content strategy you can finish. It is closer to an ongoing check than a project with an end date. See whether the engine's own searches find you at all. See what the sources it reads say about you, and whether that changes month to month. Fix the page once a search actually lands on it. Then run the check again next quarter, because the answers move and the engines do too, and a citation programme that stops checking is just a guess with better formatting. The methodology behind both studies is public, and a free check gives you a first reading today.