AI & Brand Visibility

Why AI does not cite your website, and what the machine is reading instead

By Vantage Branding·Reviewed by ·24 September 2026·10 min read

Most brands are not cited by AI engines because the engines are not reading their website when they answer the question. They are reading what other people have published about the brand, and they weight those sources far above anything the brand publishes about itself.

Owned content still matters, because it is the reference that accurate third-party sources get built from and checked against. But the practical work of AI visibility is not on-page work. It is the editorial work of correcting and seeding the sources the machine actually reads. This article sets out the evidence for that, and what to do about it.

What an AI citation is, and why being mentioned is not the same as being cited

A citation is a source an AI engine retrieves, uses and attributes when it composes an answer. A mention is the engine naming a brand from what it already holds, with no link and no source attached. The two are routinely confused, and the difference decides what a brand can do about its visibility.

The distinction is measurable. 5W Public Relations reports in its 2026 research synthesis that ChatGPT mentions brands roughly 3.2 times more often than it cites them with links. A brand can therefore be discussed constantly inside an AI answer and almost never be the page the answer was built from. It has awareness without attribution, and no way to correct the account through its own publishing.

What a citation is not: a ranking. Citation is not a reordering of the Google result set with a summary on top. It is a separate retrieval and selection process, running on a different source pool, with different preferences per engine. Treating it as SEO with extra steps is the most common and most expensive mistake in the category.

A brand can be discussed constantly by an AI engine and almost never linked by it

Consider what this does to the standard remedy. A company notices it is absent from AI answers, commissions better pages, adds schema, writes an llms.txt, restructures the copy into clean question and answer blocks. All of that is correct. None of it changes who the engine asks.

The closest everyday parallel is a credit file. It is a document about you, compiled largely from what third parties have reported, consulted by people making decisions about you, and you did not write a word of it. You can supply accurate information and ask for errors to be amended. You cannot author it. A company's machine-readable reputation now works the same way, and most marketing budgets are still built on the assumption that the brand is the author.

An AI engine will talk about your brand all day. Getting it to read your website is a different problem, and a harder one.

The 2026 datasets agree from different directions that third-party sources outweigh owned ones

Three independent bodies of evidence published in 2026 point the same way, from different methods and different commercial positions.

5W Public Relations published The State of AI Citations 2026 in May 2026, a synthesis of the largest publicly available citation datasets, including more than 680 million tracked citations across ChatGPT, Claude, Perplexity, Gemini, Google AI Overviews and Google AI Mode. Two findings from it matter most here. First, an Ahrefs analysis of 15,000 queries found only 12 per cent of URLs cited by AI tools overlap with Google's top ten organic results, meaning the large majority of AI citations come from pages that do not rank on page one. Second, entity presence across Wikidata, Wikipedia where notability supports it, and four or more authoritative third-party platforms is associated with a 2.8 times increase in citation likelihood. 5W is a public relations firm, and the conclusion happens to favour the service it sells. Naming that is more useful than ignoring it, because the second dataset is not a PR firm's.

Otterly.ai published its citation study on 1 February 2026, built on more than one million citations across ChatGPT, Perplexity and Google AI Overviews gathered in January and February 2026. Its per-engine figures show how differently the surfaces behave. Google AI Overviews gave brand domains 59.8 per cent of citations, ChatGPT 44.7 per cent, and Perplexity 28.9 per cent. The same brand, the same website, and roughly double the brand-source share on one engine compared with another. The study also found that 73 per cent of sites carry technical barriers blocking AI crawler access, which is a reminder that the owned layer can still fail outright.

The third source is academic rather than commercial. Don't Measure Once: Measuring Visibility in AI Search (GEO), a preprint by Schulte, Bleeker and Kaufmann of the University of St. Gallen posted to arXiv on 8 April 2026, tested four verticals across ChatGPT, Perplexity, Gemini and Google AI Mode between January and March 2026. It found that the standard error on a brand's estimated detection rate only falls below 0.10 at around seven runs, and that source overlap between consecutive days can fall to between 34 and 42 per cent. A single check of AI visibility is not a measurement. It is one draw from a distribution.

Citation share is concentrating into a small set of domains, which makes the target smaller and the work more specific

The same research shows the source pool is narrow and getting narrower. 5W's synthesis reports Wikipedia at roughly 47.9 per cent of ChatGPT's top ten source share, and Reddit at roughly 46.7 per cent of Perplexity's. Only an estimated 11 per cent of domains are cited by both engines. Within categories the concentration is sharper still, with single trade publications taking high single-digit shares of an entire vertical's citations.

That concentration cuts both ways. It means a brand cannot be everywhere, which is a relief. It also means that in any given category there are perhaps five to ten sources that decide how a company is described, and being wrong on one of them is expensive. Concentration also makes the picture unstable: 5W documents ChatGPT's Reddit citation share collapsing from roughly 60 per cent of responses to roughly 10 per cent within weeks in September 2025. A strategy resting on one source is fragile by construction.

Exhibit 1: The owned-to-earned ladder

Control falls as influence rises. That inversion is the whole argument.

RungSource classWhat it controlsWho controls it
1Your own site, schema and llms.txtWhether a correct, unambiguous reference description exists at allYou, entirely
2Structured and reference records: directory listings, registries, professional bodies, business profilesWhether your basic facts resolve consistently across sourcesYou, by application and maintenance
3Third-party editorial: roundups, trade press, interviews, analyst notesHow you are characterised, in whose wordsThem, influenced by you
4Community and discussion sourcesWhether the characterisation is corroborated by people with nothing to gainNobody, in any direct sense

Most budgets are concentrated on rung one. Most citations are drawn from rungs three and four.

We have been misdescribed by AI for months while publishing the correct description on our own site

Vantage has run a weekly, dated AI citation and entity accuracy audit since July 2026. The findings are worth stating plainly, because they are a first-hand account of the mechanism the studies describe, and anyone can reproduce them by running the same queries.

A widely read third-party agency roundup describes Vantage as a specialist in FMCG, retail and packaging. That is wrong. Vantage is a Singapore brand consultancy specialising in brand research, strategy, and identity design for ambitious organisations across Southeast Asia, with particular depth in healthcare, finance, government, and cultural-institution branding, and it does not position itself as a packaging or FMCG specialist. The error has persisted across months, and it propagates: at least two further publications reproduce it. Throughout that period the correct description has been published on vantagebranding.com.sg in plain copy, in structured data and in llms.txt. It has not displaced the third-party error in a single engine.

The control case sits in the same audit. A separate, actively maintained roundup describes Vantage accurately, and where an engine has that source available, it gets the description right. One engine, queried directly, returned an accurate description and resolved a professional certification detail correctly and unprompted. The variable was never the quality of the owned page. It was which third-party source the engine had to hand.

A vague position produces a vague machine description, because averaging is what a model does with thin evidence

This is the part that makes AI visibility a brand problem rather than a technical one.

A model composing a description of a company is reconciling several sources of varying quality. Where those sources say specific and consistent things, it repeats them, often close to verbatim. Where they say loose, overlapping, category-level things, it averages. The output of averaging is the category, not the company. A firm described by six sources as a full-service creative agency offering branding, marketing and digital will be returned as exactly that, because there is nothing in the evidence to distinguish it from four hundred others.

So the uncomfortable question is not whether a brand is optimised for AI. It is whether the brand is specific enough to survive being summarised by something that has no interest in flattering it. A position that only holds when a human reads the whole website is not a position. Third parties can only describe a company in terms the company has given them, and a model can only repeat what the third parties wrote down.

Correcting an entity is editorial work, not technical work

When a description is wrong, the instinct is to publish a better page. The evidence says to go upstream instead.

Exhibit 2: The entity correction sequence

  1. Find the upstream source, not the copies. Wrong descriptions propagate. Run the query across several engines, collect the pages that carry the error, and identify which one the others are reproducing. Correcting a copy achieves nothing.
  2. Check whether the source is maintained. Read the page's own structured data for a modification date. A page edited last month is a live editorial relationship. A page abandoned in 2023 is a different and much harder problem, and may be better addressed by displacing it than by amending it.
  3. Write the correction as a factual amendment, not a complaint. State the specific error, supply the canonical description verbatim, and keep it to a paragraph. Give the editor something they can paste. Where a correct listing already exists elsewhere, point to it as the shape of an accurate entry.
  4. Request the change, and expect most requests to go unanswered. This is editorial outreach. Reply rates are modest, the work is cumulative, and the ones that land are worth the ones that do not.
  5. Re-query on a fixed schedule and record the result. Per the St. Gallen preprint, a single query tells you almost nothing. Run the same prompts repeatedly, log the answers with dates, and judge the change over weeks rather than days.

None of this is quick, and nobody should promise a timeline for it. What it is, is the only lever that has moved anything in the audit.

Why this matters more in Southeast Asia than in larger markets

The source pool an AI engine can draw on for a Singapore company is thin. Where a US software firm might be described from dozens of trade publications, review platforms and analyst notes, a Singapore consultancy is often described from a handful of agency roundups and directory listings, several of which are maintained by parties with commercial interests of their own. That is an observation from the audit rather than a measured claim, but the consequence is straightforward: each individual source carries more weight here, and a single error propagates further and survives longer.

There is an opportunity in the same fact. A thin source pool is easier to influence than a crowded one. In a market where five or six pages effectively decide how a category is described, becoming the source that other sources check is achievable work for a mid-sized firm, in a way it would not be in a larger market.

What to do, in the order that actually moves the description

Fix the owned layer first, because it is cheap and because everything else refers to it. Confirm AI crawlers can reach the site at all, given the 73 per cent barrier rate Otterly recorded. Publish one canonical description of the organisation and use it without variation everywhere.

Then audit the third-party layer. List every page an engine cites when describing the company, classify each as accurate, wrong or absent, and note which ones are maintained. That list, not a keyword list, is the working document.

Then correct and seed, in that order. Wrong descriptions on maintained pages are the highest-value target, because the fix is a paragraph. Absent listings on reference records come next, because they are within the company's direct control. Third-party editorial is slowest and most valuable, and it is earned by having something specific to say rather than by asking.

Then measure repeatedly. Fixed prompts, fixed schedule, several runs, logged with dates. And keep the underlying question in view: if the description coming back is vague, the problem is usually not the machine. It is that the position is vague enough to average.

For the on-page half of this problem, which remains necessary even though it is not sufficient, see how to get your brand mentioned by ChatGPT. For the wider shift this sits inside, read how AI is changing branding. For an assessment of what the GEO evidence does and does not support, see does generative engine optimisation work. The argument that a vague position produces a vague machine description is developed further in the brand positioning framework, and the distinction between what a company intends and what the market reports back is the subject of brand identity versus brand image.

Frequently asked
questions

Why does ChatGPT talk about my brand but never link to my website?
Mentions and citations are different behaviours. A mention comes from what the model already holds about your brand; a citation is a source it retrieved and attributed while composing the answer. 5W's 2026 synthesis reports ChatGPT mentions brands roughly 3.2 times more often than it cites them with links. So a brand can be named frequently without any traffic or any attribution. Changing that means becoming present in the sources the engine retrieves from, which are mostly not your own website.
Do AI engines use the same sources as Google?
Largely not. An Ahrefs analysis of 15,000 queries, cited in 5W's 2026 report, found only 12 per cent of URLs cited by AI tools overlap with Google's top ten organic results. The engines also differ sharply from one another, with only an estimated 11 per cent of domains cited by both ChatGPT and Perplexity. Ranking well in Google is useful, but it does not buy AI citation, and a single strategy will not cover the surface.
Does adding schema and an llms.txt file make an AI engine cite me?
They make you legible and they remove barriers, which matters: Otterly.ai found 73 per cent of sites carry technical obstacles blocking AI crawler access. But legibility is a precondition, not a cause. No structural change to your own site compels an engine to select it as a source when it is weighting third-party and community pages more heavily. Treat the owned layer as the floor, and expect the movement to come from elsewhere.
How do I correct a wrong description of my company in AI answers?
Find the upstream source rather than the pages copying it, check whether it is actively maintained, and send a short factual amendment with your canonical description supplied verbatim. Then re-query the engines on a fixed schedule and log the results with dates. Expect a modest reply rate and treat it as cumulative editorial work. Publishing a better page on your own site, on its own, has not been observed to displace a persistent third-party error.
How long does an entity correction take to show up in AI answers?
There is no reliable published figure, and any consultancy quoting one is guessing. The honest answer is that it depends on three things: whether the source page is updated at all, how often the engines recrawl it, and how many other pages still carry the error. Corrections to maintained pages move first; propagated copies lag behind and may never be cleared. Measure it over weeks with repeated queries rather than looking once and drawing a conclusion.
Is this search engine optimisation, or is it something else?
It overlaps with SEO at the technical layer and diverges everywhere else. The retrieval pool is different, the selection behaviour differs by engine, and the decisive work is editorial and positional rather than on-page. It sits closer to reputation and entity management than to keyword work. It is also, ultimately, a brand strategy question, because what third parties can accurately say about a company is determined by how clearly that company has defined itself.

Wondering what AI says about your brand?

Tell us a little about your brand, and we will be in touch soon.

What brings you here?

Vantage does brand strategy and identity. We do not run media, performance marketing or campaign execution.

We will get back to you soon.

Job and survey scams. Vantage does not recruit for or operate any remote “online task” or “brand survey” work, and will never ask you to pay to earn or release commissions. Anyone offering paid task or survey work on our behalf is not us. Report suspected scams to ScamShield — call 1799 or visit scamshield.gov.sg.