All guides

AI Overviews vs Perplexity vs ChatGPT: what each one actually cites

Last updated August 2026 · 90 min first run, 30 min per repeat · Free. Perplexity, ChatGPT and Google AI Overviews all answer on free tiers. · Comfortable

AI Overviews vs Perplexity vs ChatGPT: what each one actually cites

Three AI engines answer the same buying question and cite three different kinds of source. Perplexity shows a numbered list you can read straight off the page. ChatGPT browses and names a handful inline. Google AI Overviews mostly pulls from pages already ranking underneath it. Those are three different games, and a single audit run tells you almost nothing about any of them, because the answers shuffle between runs. This is the test, the comparison, and the logging sheet that turns it into a trend.

Ask three AI engines the same buying question and you get three answers that agree on the substance and disagree completely on who gets credit for it. That disagreement is the entire subject of this guide, because it is where the work is.

If you have not measured your own visibility yet, start with the 5-question AI visibility audit, which scores whether you get named at all. This guide is the layer under that one: not whether you appear, but what kind of source each engine reaches for, and therefore which change to your content is worth making.

What you'll have when you're done

  • 15 logged answers: your 5 buying questions across Perplexity, ChatGPT and Google AI Overviews
  • A domain tally showing which kind of source each engine favours in your niche
  • A per-engine decision: the one content change that plausibly moves that engine, and the ones that do not
  • A run log with dates, so run 2 and run 3 mean something
  • One sentence you can say to a client or a founder about where you actually stand
  • [SCREENSHOT: the three-tab citation log with two dated runs side by side]

Before you start

  • Your 5 buying questions. The questions a customer types the week before they choose, none of them containing your brand name. If you do not have these written down, write them first. Everything downstream is only as good as these five lines.
  • Perplexity, free tier. Citations are visible without an account.
  • ChatGPT, free tier. Use a temporary chat so memory does not shape the answer.
  • Google, in an incognito window. Note your country and city, because AI Overviews vary by location.
  • A spreadsheet, and about 90 minutes.
  • Optional: Claude or ChatGPT for the tallying step, which is dull and mechanical and worth handing over.

The three citation behaviours

Here is the shape of the thing before you run it, so you know what you are looking at.

PerplexityChatGPTGoogle AI Overviews
Where citations appearNumbered list beside the answer, always visibleInline links, when it browsedLink cards next to and under the overview
How many you can seeUsually the most of the threeUsually a handfulA small set, expandable
Does it always browseEffectively yes, retrieval is the productOnly when the question needs itYes, it is a search surface
Typical source mixRoundups, review sites, community threads, docsA shorter mix, often the strongest-known namesPages already ranking for that query
Overlap with classic rankingsLowerLowerHighest of the three
What moves itBeing quotable, and being mentioned on cited third-party sitesSame, plus being a well-known named entityRanking, plus a clean quotable passage

Read the last row as the summary. AI Overviews is still substantially an SEO problem. Perplexity and ChatGPT are substantially a mentions-and-quotability problem. Teams that treat all three as one channel end up doing SEO work and wondering why Perplexity never changes.

A caveat that belongs here and not in a footnote: these behaviours are product decisions, and product decisions change. Everything below is a method for measuring the engines yourself, precisely so you are not relying on a description someone wrote six months ago, this one included.

Step 1: Fix the conditions before you run anything

Every result is worthless if run 2 is not comparable to run 1. Write the conditions at the top of the sheet and do not change them between runs.

RUN CONDITIONS
Date:            2026-08-29
Questions:       Q1-Q5, verbatim, unchanged since run 1
Perplexity:      logged out, new thread, default model
ChatGPT:         temporary chat, no memory, default model
AI Overviews:    incognito, google.com, location: Tokyo, JP
Order:           Perplexity, then ChatGPT, then Google

Three rules that decide whether the exercise is science or vibes:

  1. Never edit a question between runs. Not for typos, not for phrasing. A changed question is a new question and the history resets.
  2. Clean context every time. Memory, history and personalisation all shift answers, and they shift them in the direction of what you have already looked at, which is the exact bias you are trying to measure around.
  3. Log the model version if the engine shows it. When a result jumps, this is usually why.

Check it worked: your sheet has a run-conditions block filled in and 5 questions typed out in full, and you have not started asking yet.

Step 2: Run all 15 and log the sources

For each question in each engine, capture three things: whether your brand is named anywhere in the answer, the ordered list of every domain cited or linked, and one line on what kind of answer it gave.

One row per source, not one row per answer. Long sheet, easy pivots:

run_datequestionenginepositiondomainsource_typebrand_named
2026-08-29Q1perplexity1g2.comreviewno
2026-08-29Q1perplexity2reddit.comcommunityno
2026-08-29Q1chatgpt1g2.comreviewno
2026-08-29Q1aio1competitor.combrand ownno

source_type is the column that pays for the whole exercise. Use six values and nothing else: review, roundup, community, news, brand own, docs.

On AI Overviews specifically: run the query, screenshot the overview, then scroll and note the top 5 organic results too. You want the overlap number, because that is the single most decision-relevant fact about this engine in your niche. If 4 of the 5 cited pages are also in the top 5 organic, your AI Overviews strategy is your SEO strategy and you can stop treating it as a separate project.

Fifteen answers gives you somewhere between 40 and 80 source rows. Hand the tallying to a model:

Citation tally by engine
Below are 15 AI answers to buying questions in one niche. Each is
under a heading like "Q1 / perplexity". Extract every domain cited
or linked, in order.

Output:
1. A table by ENGINE: engine, total citations, unique domains,
   and the count of each source_type (review, roundup, community,
   news, brand own, docs).
2. A table by DOMAIN: domain, times cited, which engines, which
   questions. Sort by times cited, descending.
3. The domains cited by all three engines. Label this "consensus set".
4. Domains cited by exactly one engine. Label this "engine-specific"
   and say which engine.
5. One sentence per engine: what kind of source it appears to prefer
   in this niche, based only on the counts above.

Do not include any domain that is not in the text below. Do not
infer, rank or editorialise beyond what the counts support.

ANSWERS:
[paste all 15]

Check it worked: the tally gives you a consensus set and an engine-specific set, and you can name the top three domains for each engine without looking again.

Step 3: Read the result per engine

Now the tally becomes decisions. What you are looking for is not your score, it is the shape of each engine's preference.

If AI Overviews shows high overlap with organic results. Your lever is ranking for the question, plus making sure the passage at the top of that page directly answers it. Both parts matter: rank without a quotable passage often gets you the click-under rather than the citation. Write the answer in the first two sentences under a heading that is the question itself.

If Perplexity leans on reviews and community threads. Your own site is barely in this game. The lever is being named on the pages it already cites. Get listed on the review platform that appears most, ask five real customers to review you there, and answer the specific community threads that came up in your run, properly, as somebody who knows the subject. A thread cited today keeps being cited.

If ChatGPT names a short list of strong sources and never you. This usually reads as an entity problem rather than a content problem. It is naming things it is confident it can describe. The fix is upstream: one clear sentence about what you are, live in identical wording across your site, your schema and your profiles, and then third-party mentions that repeat it. The three-lever GEO playbook is the long version of that fix.

If the same domain appears across all three engines. That is the highest-value target on your list. One placement there registers three times.

What the tally showsThe change worth makingThe change not worth making
AIO overlaps organic heavilyRank for the query, answer in sentence oneChasing Perplexity-specific tactics
Perplexity favours reviews and forumsListings, reviews, real thread answersPublishing more blog posts
ChatGPT names few, strong sourcesEntity clarity plus third-party mentionsKeyword work on your own site
A domain appears in all threeGet named on that one pageSpreading effort across 20 domains

Check it worked: you can finish this sentence for each engine, out loud: "for [engine], the thing worth doing next is X, because in my run it cited mostly Y."

Step 4: Why one run tells you almost nothing

This is the part most audits skip and it is the part that decides whether your numbers mean anything.

Run the same question twice in the same engine, an hour apart, in clean sessions. You will usually get an answer that agrees with itself on substance and cites a partly different set of sources. Nothing about you changed in that hour. The causes are all upstream: the retrieval pass surfaces a slightly different set, the index has moved, the model got updated, or there is plain sampling variation in what gets quoted.

The consequences are unforgiving:

  • A single run cannot tell you a source is important. It tells you a source was reachable once.
  • A single run cannot tell you a change worked. If your score goes from 3 out of 15 to 5 out of 15 after a month of effort, that gap is inside the noise you would see doing nothing.
  • A single run absolutely cannot support a percentage. "You appear in 40% of AI answers" is a coin flip with a decimal point on it.

What does survive:

  1. Repeat sources. Domains cited in run 1 and run 2 and run 3. That is your real competitive set, and it stabilises fast, usually by the third run.
  2. The source-type mix per engine. The exact domains shuffle. The proportion of reviews to community threads to brand sites is much steadier, and it is what you are actually planning against.
  3. Direction over three runs. Three consecutive runs pointing the same way is a signal. Two is a line drawn through two points.
The discipline is boring and it is the whole thing: same 5 questions, same 3 engines, same conditions, once a month, logged with a date. Everything interesting in this dataset lives in the difference between run 1 and run 3, and none of it is visible on the day you first run it.

Step 5: Log runs so you get a trend

Three tabs, one sheet.

Tab 1: runs. One row per source, the long table from step 2, appended forever. Never overwrite a run.

Tab 2: summary. One row per run date per engine:

run_dateengineanswersbrand_namedunique_domainspct_reviewpct_communitypct_brand_ownpct_roundup
2026-08-29perplexity501839281716
2026-08-29chatgpt50933113422
2026-08-29aio51111895518

Tab 3: targets. Every domain seen twice or more, with the runs it appeared in, its source type, a realistic path to being mentioned there, and the date you acted. This tab is the actual work list, and it is the only one anybody else needs to see.

Three numbers to read each month, and no others:

  1. Named count out of 15. Blunt, comparable, honest.
  2. Engine spread. 5 out of 15 concentrated in one engine is a completely different position from 5 spread across three. The first is an engine-specific quirk, the second is real presence.
  3. Repeat-source count. How many domains carried over from last run. Rising means your target list is firming up, and the work gets more efficient every month.

What not to track: how many pages you published, how many pitches you sent, whether rankings moved. Those are inputs, and inputs are what people report when the output has not moved.

Check it worked: two dated runs exist in tab 1, tab 2 has six summary rows, and tab 3 has at least five domains with a next action and a date.

Where this is going

The method above works, and it is 30 minutes a month forever, done by hand, by a person who could be doing something else. That is exactly the shape of a job that should be a scheduled job: fixed questions, fixed engines, fixed sheet, run on a timer, with the history kept and the trend drawn for you.

That is what I am building CiteScore for. Same 5 questions, same 3 engines, run on a schedule, sources tallied and typed automatically, and the only thing surfaced is what changed since last time. The reason it exists is the whole of step 4: the useful signal in this data lives across runs, and nobody reliably keeps a manual log going past month two.

Run it by hand first anyway. The manual runs are where you learn what a good source looks like in your niche, and that judgment is the part no dashboard hands you.

FAQ

Why does Perplexity cite so much more than ChatGPT?

Because citation is Perplexity's product surface, not a feature bolted onto it. Every answer is assembled from a retrieval pass and the numbered sources are rendered next to the sentences they support, so you can read the full list off the page without asking. ChatGPT names sources when it browses for an answer, and inline, which usually means a shorter visible list. That difference is about interface and product design, not about one engine reading more of the web than the other.

Do AI Overviews just cite whatever ranks first?

Not exactly first, but the overlap with the classic ranking results on the same page is high, and it is the highest of the three engines. In practice this is the useful takeaway: for AI Overviews, normal search visibility is still the main input, so the SEO work you already do is the lever. For Perplexity and ChatGPT the lever shifts toward being quotable and being mentioned on third-party sites the engines already trust.

How many runs before I believe a result?

Three, spread over at least three weeks, same questions and same sheet. Answers shuffle between runs for reasons that have nothing to do with you, including model updates, index freshness, personalisation and plain sampling variation. A source that appears in one run is noise. A source that appears in all three is your actual competitive set. Anyone who audits once and reports a percentage is reporting a coin flip.

Should I use logged-out or logged-in sessions?

Logged out and in a clean window, every time, so memory and history do not shape the answer. On ChatGPT use a temporary chat. On Perplexity, a new thread with no personalisation. For AI Overviews use an incognito window and note your location, because the overview varies by country and sometimes by city. Write the conditions at the top of every run so a later run is comparable to this one.

Can I just buy a tool that does this?

You can, and several exist, but run it by hand once first. The manual run is what teaches you what a good source looks like in your niche, which is the judgment a dashboard cannot hand you. Once you know what you are looking at, automating the runs is worth it because the value is entirely in the repetition. That automation is what I am building CiteScore for: same questions, same engines, scheduled, with the trend line kept for you.

My brand appears in zero answers. Where do I start?

Look at what did get named rather than at your own absence. Tally the domains, classify them, and find the ones that are realistically reachable: directories, review platforms, community threads and roundups. Getting named on the pages the engines already cite moves you faster than anything you can publish on your own site, because your site is one vote and those pages are the ones being read.

Related guides

Get the next build in your inbox

One email a week: the newest guides, plus one thing I only share with the list.

No spam. Unsubscribe anytime.

Build alongside others

Join the free community and share what you're shipping.

Jordan Hong Tai

Jordan Hong Tai

I've scaled products to over 500K users, and now I build AI systems in public from a balcony in Tokyo.