The question comes up in every US marketing meeting now, usually in a slightly irritated tone: where is ChatGPT even getting this? A founder in Austin showed us an answer that recommended three competitors and cited a four-year-old forum thread, a review site profile his company had never claimed, and a trade publication he had never pitched. Nothing on his own site appeared anywhere in it.

That is not a bug. It is the shape of the system. A ChatGPT recommendation is not a ranking of pages — it is a paragraph written from a handful of sources the model either learned during training or opened at the moment of the question. So the useful question is not "how do I rank?" but "which sources does it actually read, and what do they say about me?"

Over the past year several independent groups have tried to answer that empirically, at real scale — hundreds of millions of citations across ChatGPT, Perplexity, Gemini and Google's AI surfaces. The studies disagree on decimals and agree on the shape. Below is that shape, and what a US brand should do about each layer.

The one caveat that keeps people out of trouble

Every number in this space is an estimate of a moving target. Citation studies sample prompts, and the prompt set decides the answer: a study weighted toward commercial software queries finds review platforms everywhere, one weighted toward informational questions finds encyclopedic reference everywhere. Different tools also measure different surfaces — ChatGPT with search on cites visibly, while ChatGPT answering from memory names brands with no citations at all.

So treat the figures below as direction, not dosage. The ordering has held steady across studies for over a year, which is what makes it worth planning around.

1. Encyclopedic reference: the entity layer

Wikipedia is the single most consistently cited domain in ChatGPT analyses, with one widely-circulated synthesis putting it at roughly 48% of ChatGPT's top-10 most-cited sources and around 16% of all citations, and it appears at the top of essentially every credible study of the same question.

Its influence is bigger than the citation count suggests, because reference pages do not just get quoted — they define. Wikipedia and Wikidata are where a model settles what your company is: category, founding, ownership, headquarters, what it is known for. When that entity record is thin, stale or contradictory, the model hedges or substitutes a competitor it can describe confidently.

What to do, carefully. Do not write your own Wikipedia article; it is against the rules, it gets reverted, and it can leave a permanent conflict-of-interest trail. Do make yourself citable by independent sources so an editor eventually has something to work from, and do get your Wikidata item and other structured identifiers correct, because those are editable in good faith and machine-read constantly.

2. Community discussion: where "best X" gets decided

Reddit is the other giant, and it is the one that offends marketers most. Analyses put Reddit near the top of AI citation sources overall — roughly 40% of top citations in one cross-engine index and around 47% of Perplexity's top-10 — and its influence on ChatGPT is heaviest exactly where money is: recommendation and shortlist questions.

The reason is structural. A Reddit thread titled "best CRM for a 10-person agency" contains what no vendor page contains: multiple named options, argued trade-offs, prices people actually paid, and complaints. That is precisely the material a model needs to write a balanced recommendation, so retrieval keeps landing there. Quora, Stack Exchange and category-specific forums play the same role at smaller volume.

What to do. Not astroturfing — it is detectable, against platform rules, and a deleted thread cites nothing. What works is being genuinely present: a founder or engineer answering with a disclosed affiliation, in the subreddits where your buyers actually are, over months. It also pays to know which threads currently get cited for your category and whether the information in them about you is simply out of date, since a factually wrong 2022 comment can outrank your entire site in an answer.

3. Video: the fastest-moving layer

The notable change in 2026 is video. Trade coverage this year reported YouTube overtaking Reddit as the leading social citation source across AI search surfaces, and YouTube shows up strongly in Google's AI features in particular.

This favors a specific kind of content: demos, walkthroughs, hands-on comparisons and conference talks — video with a real transcript full of specific claims. A polished brand film with three sentences of narration gives a model nothing. A twelve-minute unedited product walkthrough with named features, prices and limitations is extremely quotable.

For US B2B brands this is currently the most under-competed layer. Most categories have almost no honest, well-transcribed video comparison content, and the cost of producing it is a screen recording and an hour of preparation.

4. Review and comparison platforms: the facts layer

For software and services, G2, Capterra, TrustRadius, Software Advice and their vertical equivalents are cited heavily — and their profiles double as structured fact sheets: category, pricing tier, integrations, company size, ratings.

This is the cheapest fix on the list and the one most often left half done. An unclaimed profile with 2023 pricing, a wrong category tag and eleven reviews is actively working against you. A claimed profile with current pricing, correct integration list and a steady flow of compliant, real reviews feeds the model the same facts your site says, which is exactly the corroboration it looks for.

The same applies to directories and marketplaces in non-software categories: the app store listing, the manufacturer directory, the professional association roster, the state licensing record. Boring pages, disproportionate influence.

5. Trade press, not national press

One of the more counterintuitive findings in recent audits: several major national newspapers barely feature in ChatGPT's top-cited domains, while specialist trade and industry publications punch well above their traffic. Part of this is licensing and access — hard paywalls limit what retrieval can read — and part is fit: a model answering "best warehouse management system for cold storage" needs a trade outlet, not a general news front page.

The practical implication is that a placement strategy built around three national logos may deliver less AI visibility than fifteen credible placements in the outlets your buyers' industry actually reads. Prioritize publications that are open to crawlers, publish detailed comparative or technical coverage, and are updated often.

6. Your own site: smaller role, non-negotiable

Owned pages are usually a minority of citations in a recommendation answer, but they are what the model checks to confirm specifics — and if it cannot confirm them, it softens or drops you.

Four pages carry nearly all of that weight: a pricing page with actual numbers or a stated band, an integrations and compatibility page (an enormous share of buyer follow-ups are "does it work with Salesforce / NetSuite / Epic / QuickBooks"), a comparison page that names real alternatives and admits where you lose, and documentation, which models trust because it is precise and unpromotional.

Also make sure retrieval can reach them. OpenAI runs GPTBot for training, OAI-SearchBot for search citations and ChatGPT-User for live browsing, controlled separately in robots.txt — and CDN bot rules at Cloudflare or Akamai frequently block them regardless of what robots.txt permits. Facts hidden behind client-side JavaScript, images, PDFs or a form are functionally invisible.

How to find your own source list in an afternoon

Do not plan against the industry averages. Every category has its own citation set, and yours is discoverable in about three hours.

Write thirty prompts the way your buyers would say them — situations, constraints and budgets, not your brand name. Run each in ChatGPT with search enabled and record four fields: were you named, in what position, which sources were cited, and what those sources said about your category. Repeat on a schedule.

You will almost always find eight to fifteen domains recurring. That list is your work order, ranked by how often it appears. Everything else in AI visibility is downstream of it — and re-running the same prompts monthly is how you find out whether anything you did mattered, since aggregate GA4 traffic will not tell you.

Who does this work, and how their front pages sell it

If you want outside help, the positioning on a firm's homepage tells you a lot about what you will actually receive. Screenshots were captured on 11 September 2026; homepages change often.

llmrecommend.com — proof on one keyword before you commit

Disclosure: llmrecommend.com is our own brand, which is why it is listed first and labelled as ours.

It was built around the objection we hear most from US marketers: answer-engine retainers ask for six months of faith before anything is provable. The scope is deliberately narrow — one high-intent keyword, one engine. The team pulls the current answer, lists every source it cites, identifies the missing first-hand evidence, and publishes on the sources that engine already trusts. A shared dashboard shows daily whether your brand is in the answer, in what position, and from which cited source. First milestone at day 30, full milestone only if presence holds across 60 days, no result no invoice.

Stated plainly: one keyword on one engine is a proof, not a program. It is the right first purchase if you want evidence before budget, and the wrong one if you need category-wide coverage across four models this quarter.

Front page of llmrecommend.com
llmrecommend.com — our own brand; outcome-based, single-keyword proof.

Omniscient Digital. B2B SaaS content specialists with unusually rigorous published thinking on organic growth. The right call when the underlying problem is that nothing you own is worth citing yet.

Front page of Omniscient Digital
Omniscient Digital — B2B SaaS content and organic growth.

Siege Media. Long-established content and digital PR shop; strongest where the gap is third-party coverage and original data that trade publications will actually pick up. Ask how they measure citations rather than placements.

Front page of Siege Media
Siege Media — content and digital PR for earned coverage.

Graphite. Data-led growth firm working with product-led and marketplace companies, treating measurement as an engineering problem. Suited to teams that want visibility instrumented properly rather than reported in a slide.

Front page of Graphite
Graphite — data-led growth for product-led companies.

The short version

Reference pages decide what you are. Community threads and video decide whether you get shortlisted. Review platforms and trade press corroborate. Your own site confirms the specifics — price, fit, compatibility — or fails to.

None of that is optimizable with a trick, and none of it can be bought as placement: there is no ad inventory inside a ChatGPT recommendation, so anyone guaranteeing a position is selling something they do not control. What works is unglamorous and durable — be described the same way everywhere, be quotable, be corroborated by sources you do not own, and check the actual prompts instead of the traffic chart.