
What does ChatGPT cite in 2026: help docs, preprints and vendor pages in one measured category
On 10 August 2026 we listed the domains holding AI citations across the 58 topics we track in one category — AI visibility and search-optimisation tooling. The eight our tracker names as holding the most ground held between 5 and 9 topics each, and four of those eight are not content marketing publishers at all. One is a preprint server. One is OpenAI's own help centre. One is a competitor's help documentation. One is a consumer technology magazine.
That is a smaller and narrower thing than the citation indices you have probably read this year, and we want to be exact about how much smaller: 58 topics, one category, one reading taken on one day, measured mostly against ChatGPT. It cannot tell you what the whole web looks like. What it can tell you is what the competition for a citation slot actually consists of inside a single commercial category — and in this one, four of the eight domains our tracker recorded as holding the most ground are not publishing marketing content at all.
That distinction is the reason we are publishing a small reading at all. The large reports measure the open web, which is the right way to answer "what does an assistant tend to cite". The narrower question a person actually acts on is "who is holding the specific questions my buyers ask, and what kind of page are they" — and that question is answered inside products rather than in public. Profound's own feature pages offer to find the prompts where competitors outrank you and to show every cited URL for each prompt it monitors. What is missing is a reading of one category published in the open: every report we found while checking this SERP measures the internet-wide pattern instead, which is a different question and a bigger one. A category is where content budgets get spent, and a category is small enough to check yourself in an afternoon — which is what the last section of this article is for.
This piece is the report. If you want the practical companion — the things a page needs before an assistant will quote it — that is our write-up on getting citations from ChatGPT, and we are not going to re-explain it here.
Which domains hold the citations in this category
Profound and arXiv each held 9 of the 58 topics, Semrush held 7, and Otterly held 6, in the reading our citation tracker returned on 10 August 2026. The set is the 58 topics we track in the AI-visibility and search-optimisation category, and the measurement is which domains were cited in answers to those questions, mostly on ChatGPT.
| domain | topics it held, of 58 | what the domain is | |---|---|---| | tryprofound.com | 9 | an AI-search visibility vendor | | arxiv.org | 9 | an open-access preprint server | | semrush.com | 7 | an SEO and marketing platform | | otterly.ai | 6 | an AI-search monitoring vendor | | help.openai.com | 5 | OpenAI's own help centre | | scrunch.com | 5 | an AI-visibility platform | | techradar.com | 5 | a consumer technology publication | | help.otterly.ai | 5 | that vendor's own help documentation |
Read the counts carefully, because the shape of the number matters more than its size. Each figure is the number of topics on which that domain was cited, not a share of anything. More than one domain can be cited in a single answer, so these counts overlap and they do not add up to 58. Our tracker names these eight and does not publish the tail beneath them, so this is not a claim that no other domain reached five topics — only that these eight were the ones it recorded as holding the most ground on the day we looked.
On 14 of the 58 topics, no domain was cited at all, and the tracker's own label for those rows is that there was no citation to win. That is what the reading records; our read of why is that no assistant went searching, on the mechanism OpenAI states about its own product and which the next section quotes. That leaves 44 topics on which anybody held anything, which is the real size of the ground the eight domains above are competing over.
Two of the eight rows belong to one company: otterly.ai held 6 topics and help.otterly.ai held 5 in the same reading. That is one organisation holding ground through two different kinds of page — a marketing site and a knowledge base — and it is easy to miss when a list is sorted by count. We report it as the count it is and not as a verdict on anybody; the point is that a row in a domain list is not reliably a separate competitor.
Why half of that list is not marketing content
A citation slot is not reserved for the companies selling into a category, and four of these eight domains were built to record or report something rather than to argue an article's way into an answer. Those four are a preprint archive, two support-documentation sites and a technology magazine. arXiv is an open-access archive of nearly 2.4 million scholarly articles, and it does not peer-review them, which it states on its own front page. help.openai.com is OpenAI's support documentation for its own product. help.otterly.ai is the same object for a monitoring tool: onboarding guides, prompt-research instructions, the ordinary furniture of a knowledge base. TechRadar is a consumer technology publication — not a company blog and not a vendor.
We had assumed, before we listed them, that a category like this one would be held almost entirely by the marketing arms of the tools in it. Four of the eight are exactly that: Profound and Otterly sell AI-search visibility monitoring, Scrunch sells an AI-visibility platform, and Semrush is a long-established SEO and marketing suite that now sells AI-search features too. All four build in this space and all four publish about it, so their presence is unremarkable. The other four were the surprise, and they are the useful half.
The practical consequence is a change in who you think you are competing with. If you are writing an article to win a citation slot on a question like these, your competition on the day we looked was not only the four vendors' blogs. It was also a preprint archive, a model-maker's support pages, and a general-interest technology magazine — three kinds of page you cannot displace by writing a better blog post, because they are not the same kind of object.
It is worth being precise about when any of this is even in play. A citation exists only when the assistant actually goes and searches; OpenAI's own announcement of ChatGPT search says that "ChatGPT will choose to search the web based on what you ask, or you can manually choose to search by clicking the web search icon" (openai.com, published 31 October 2024, updated 5 February 2025). Where no search runs, there is no slot to hold, which is what the 14 topics above are.
What a reference page has that an article usually does not
Each of the four non-marketing domains is the place a fact lives rather than a place a fact is discussed, and we think that shared property is why they hold ground — a read, we should say immediately, and not a measurement. Our reading records the domain that was cited on each of the 58 topics, not the page. So what follows is an argument about kinds of domain, and it is weakest exactly where one domain publishes many kinds of page. A preprint is the paper itself, and arxiv.org publishes nothing else. A help centre is the vendor's own statement of how its product behaves, which is all help.openai.com and help.otterly.ai are for. Not one of those three is a summary of somebody else's summary. TechRadar is the row the domain does not tell you about: it runs reviews, news, buying guides and explainers, and because we did not record which of its pages were cited, whether the 5 topics techradar.com held share the property at all is an open question rather than a finding.

That reading is consistent with what we have measured about which sentences get lifted out of pages, and the method and the samples for that work are set out on our research page rather than repeated here. But consistent-with is not the same as demonstrated, and we have not run a test that isolates page type as a variable. If someone has, we would like to read it.
It is also worth naming what this does not license, because the obvious misreading is expensive. It does not mean documentation outranks articles in general — we did not measure that, and four of the eight domains holding the most ground here are ordinary vendor marketing sites. It does not mean a page becomes citable by being formatted to look like documentation. And it says nothing at all about categories we have not read.
What you do with it is your call and it depends on things we cannot see. In some categories there is a documentation site you own and have never treated as publishable material. In others there is no reference-shaped page available to you at all, and the honest answer is that this particular door is closed. We are not going to tell you which of those you are in.
How to read an AI citation index, including this one
Any AI citation index becomes readable once you have three answers about it: measured where, over what set, and on what date. That includes the reading above, and it includes every report we quote below. Every apparent disagreement we have looked at between two of these documents dissolved once all three answers were on the table, because the two reports were not measuring the same thing in the first place.

Take two of the other reports that rank for this question, both of which are considerably larger than ours. Otterly's AI Citations Report 2026 states its scope on the page: "Analysis of 1+ million AI citations across ChatGPT, Perplexity, and Google AI Overviews from January-February 2026", last updated 1 February 2026 — their sample, their engines, their two months. 5WPR's State of AI Citations 2026, dated May 2026, is explicit that it is a different kind of document again: it "synthesizes findings from the largest publicly available citation datasets", including a set of more than 680 million tracked citations, and names each dataset it draws on.
Those three answers are different from ours in every position. They measured across several engines; we measured mostly one. They measured the open web; we measured 58 questions inside one commercial category. They measured over months; we took one reading on one day. None of that makes any of them wrong — a million citations across three engines answers a question ours cannot touch, which is what the general pattern of AI citation looks like. It also cannot answer ours, which is who is actually holding the slots for the specific questions a buyer in one category asks. We are the small, narrow instrument here, and where an instrument is narrow it should say so before it says anything else.
A fourth question catches more disagreements than the first three combined: what, exactly, is being counted? A citation with a link, a brand mention without a link, a domain, a page and a quoted passage are five different units of measurement, and citation reports move between them freely and often silently. The 5WPR synthesis makes the gap explicit for one engine, reporting that ChatGPT "mentions brands roughly 3.2 times more often than it cites them with links" — their figure, from their synthesis, dated May 2026. Our own counts above are topics on which a domain was cited, which is a sixth unit again. Two honest reports using different units will produce numbers that look contradictory and are not.
There is one place where two of these very different measurements point the same way, and we found it useful. 5WPR's report notes that in the B2B SaaS CRM category, "TechRadar alone accounts for 8.86% of category citations" — a different dataset, a different category and a different method from ours. Our own separate reading put techradar.com on 5 of our 58 topics. Neither of those numbers corroborates the other arithmetically, and we are not going to pretend they do. What they share is a shape: a general-interest technology publisher holding category ground in a commercial software category where you would expect vendors to hold it.
Read our three questions as our view of what matters, not as a neutral standard. We build a measurement tool, and a company that builds a measurement tool has an obvious interest in persuading you to interrogate measurements. We think the questions survive that bias — they are the same three we have to answer about ourselves, and the answers above are not flattering to our sample size — but you should know where they come from.
We have also gone the other way with a competitor's data when it deserved it. When we cross-checked Surfer's exported ChatGPT source lists against our own independent measurement, with different prompts and a different method, their data held up; the method and the row count for that check are in our write-up on data accuracy and we are not going to re-run the numbers here. It is the same discipline in both directions: state the scope, then report what you found.
What this measurement cannot tell you
This reading covers 58 topics in one commercial category, taken on one day, mostly against one assistant — that is the largest limit on it, and it governs every sentence above. The rest sit beside their claims in the sections above; they are collected here because they are load-bearing rather than decorative.
One category, one day. 58 topics inside AI visibility and search-optimisation tooling, read on 10 August 2026. Nothing here establishes that another category behaves this way, and we have not tested one.
Mostly one engine. Our measurement work is predominantly ChatGPT; Gemini, Perplexity and Google's AI surfaces are sampled far more thinly in our data and behave differently, as our research page sets out with its samples. Perplexity in particular cites far more freely than the engine we measured most.
Topics, not volume. Every count above is a number of topics on which a domain was cited. It is not a share of an engine's citations, not a traffic figure, and not a ranking.
A reading is a moment. A domain that held five topics on 10 August 2026 may hold none today; we have not re-measured, and we are not going to imply currency we do not have. The same caution applies to every index in this article, ours included, and it is why the date belongs beside every number.
One reading is one run. All eight counts in the table above come from a single pass taken on 10 August 2026, and we have published that asking the same assistant the same question twice can return different citations. That makes the smaller counts on this table the softer ones — a domain recorded on five topics in one pass rests on less than one recorded on nine, and we have not put a number on how much a second pass would move either. The shape, that four of the eight are not marketing publishers, is the part we would expect to survive a re-run. The exact figures are not, and we have not re-run them.
Correlation, not cause. We can say which domains held slots. We cannot say that being a preprint server or a help centre caused it, and the page carrying our citation research says so plainly: we observe what cited pages share; we cannot yet prove that adding those properties causes citation.
You cannot reproduce this from what we have published. We have not published the 58 questions behind this reading, the exact wording each one went in as, or how many times each was asked. You can see the counts and the date; you cannot audit them. That is a limit on this reading and not a disclaimer about it, because the alternative is asking you to take a number on trust — which is the one thing the section above tells you not to do for anybody else's index either.
And our own position is deliberately absent. The same reading contains our own citation count for these 58 topics, and we are not quoting it, because it is dated 10 August 2026 and predates every article we have published since. Quoting a stale number about ourselves would be exactly the thing this article is asking you not to accept from anyone else. When we re-take the reading, we will publish what it says.
What people actually ask us
Six questions come back whenever we show the 10 August 2026 reading to someone who works in this category, and the answers below are the ones we give them. Each is written to be taken whole.
Should I publish on arXiv, then? Almost certainly not, and that is not the lesson. In the one case we examined closely — who held the citations for a single generative-engine-optimisation keyword, checked on 3 August 2026 — arXiv was cited because the paper that named the field is hosted there and was being used as the primary source it is. We have not traced the other topics arXiv held, so we cannot say that explains all nine. Submitting marketing material to a preprint archive is not a strategy; noticing that primary sources beat commentary about them is.
Can I get into a help centre I do not own? No, and that is the point of naming it. Two of the eight domains holding the most ground in our reading are vendors' own documentation, which makes those slots structurally unavailable to every company except the one that owns the product. If a meaningful share of the citation slots in your category are held by pages nobody outside those companies can publish, the honest read is that the winnable share is smaller than the topic count suggests — and knowing that before you commission twelve articles is worth more than any tactic.
Is a bigger index a better one? It depends what you are asking it. A million citations across three engines describes general behaviour far better than 58 topics ever could. Fifty-eight topics inside your own category describe your competition far better than a million citations across the open web ever could. The mistake is not reading a big index; it is reading a big index as though it answered a question about your specific market.
Does this transfer to my category? We do not know, and we would not guess. What we can tell you is how to find out, because the procedure is not difficult and it is at the end of this article.
What about Gemini, Perplexity and Google's AI answers? Our measurement is mostly ChatGPT, so treat everything above as a statement about that surface first. The engines demonstrably differ in how freely they cite and in which kinds of source they favour, which is one reason the reports quoted above disagree with each other as much as they do.
Why only 58 topics? Because that is what we have actually tracked in this category, and we would rather publish the set we ran than round it up into something that sounds like a study. The set grows as the weekly sweeps run, and when the reading changes we will publish what changed — including if it contradicts the table above. A number that only ever moves in the flattering direction is not a measurement.
What to do with this tomorrow
You can run this same reading on your own category in an afternoon, using 20 to 30 questions and no tooling beyond a notebook. Take the questions a buyer in your market would genuinely ask an assistant, put each one to the engine you care about, and record every domain it cites. Then sort those domains into four piles: primary sources and reference material, vendor documentation, publishers, and marketing pages. Those four piles are our categories, chosen because they are the split that turned out to matter in our own reading; if a different cut fits your market better, use that one.

The ratio between those piles is the useful output, and you can read it in an afternoon. If marketing pages dominate, an article is genuinely in the competition and the question becomes what makes yours worth lifting. If documentation and reference material dominate, as they did in half of our category, then some of those slots are not available to you at any quality level, and the sensible response is to spend less effort on the topics they hold and more on the ones they do not.
Record the domain and the exact page for every citation, never just the brand. It is on this list because our own reading did not do it: we recorded domains, which is why the page-type argument above is a read rather than a finding, and it is the gap we would close first in our own instrument. The split between a company's marketing site and its documentation was half of what we found on 10 August 2026, and a brand-level tally would have hidden even that. Keep the questions that returned no citation at all, rather than dropping them. Those are not failures of your content; they are questions where no slot exists, and they belong in a different pile from the ones you lost.
Write down the date you did it, and the exact set of questions you used. Then do it again in a month. The single most useful property of a citation reading is not its size — it is that you can repeat it and see what moved.
Sources
All URLs checked 4 September 2026.
- LiamVi citation standing, reading dated 10 August 2026 — our own tracker. Its question set and prompt method are not published; the counts and the date are what it publishes.
- LiamVi research — our cited-versus-ignored keyword work, its samples and its limits — https://liamvi.com/research/
- LiamVi, How to get citations from ChatGPT — https://liamvi.com/blog/how-to-get-citations-from-chatgpt
- LiamVi, Best accurate data platform for AI search optimization (the vendor cross-check) — https://liamvi.com/blog/best-accurate-data-platform-for-ai-search-optimization
- LiamVi, Generative engine optimization companies (who held one keyword) — https://liamvi.com/blog/generative-engine-optimization-companies
- OpenAI, Introducing ChatGPT search, 31 October 2024, updated 5 February 2025 — https://openai.com/index/introducing-chatgpt-search/
- arXiv — https://arxiv.org/
- OpenAI Help Center — https://help.openai.com/en/
- OtterlyAI — https://otterly.ai/ · OtterlyAI knowledge base — https://help.otterly.ai/
- OtterlyAI, The AI Citations Report 2026, last updated 1 February 2026 — https://otterly.ai/blog/the-ai-citations-report-2026/
- 5WPR, The state of AI citations 2026, May 2026 — https://5wpr.com/research/state-of-ai-citations-2026/
- Profound — https://tryprofound.com/ · Answer Engine Insights feature pages (competitors, citations) — https://www.tryprofound.com/features/answer-engine-insights/competitors · https://www.tryprofound.com/features/answer-engine-insights/citations
- Scrunch — https://scrunch.com/
- Semrush — https://www.semrush.com/
- TechRadar — https://www.techradar.com/