GEO Research

Who AI recommends, and why.

We ask the engines real buyer questions and log exactly who they name and what they cite. Then we publish the raw data so you can check us. Everything below is measured — nothing here is estimated, illustrative, or rounded to look better.

Every number is measured

No estimates, no "representative" figures. If we couldn't measure it, it isn't in the report.

The dataset is public

The raw CSVs ship with every report. Recompute anything you doubt.

We publish our failures

Hypotheses that didn't survive significance testing get named, not quietly dropped.

The 2026 AI Citations series

Six reports, one method

Real buyer questions, run through ChatGPT, Perplexity and Gemini — first 39 questions on July 1, then 250 (including the identical 39) on July 16. Each report is a different true thing in the same measured data.

No. 1

The 2026 State of AI Citations

56%

Ask three engines the same buying question and all three share even one recommendation just 56% of the time. Their full lists overlap ~7%. "Rank on AI" is really three separate games.

All 3 engines + sourcesRead →
No. 4

One Engine's Favourite, Everyone Else's Unknown

79%

The most-recommended brand in the study is named only by ChatGPT — zero from Perplexity, zero from Gemini. And 79% of the 646 brands exist in exactly one engine's world.

All 3 enginesRead →
No. 2

In Every Room, Never on the Podium

87% / 0

Reddit is cited in 87% of Perplexity's answers — and was the top source for none of them. Meanwhile 86% of cited domains appeared exactly once. The tail is the market.

Perplexity only · n=39Read →
No. 3

Can Your Own Site Get Cited by AI?

p=0.004

Yes — 31% of the time. We tested four theories about why. Three were statistical noise, and we publish them anyway. The one that survived was structure: FAQ blocks and tables.

Perplexity only · n=39Read →
No. 5 · the drift report

The Reddit Collapse

87% → 18%

Fifteen days, identical questions, identical pipeline: Reddit fell from 87% of Perplexity's answers to 18%, and the top-cited source changed on 31 of 39 questions — while every structural finding held. The rules are stable; the winners are not.

Perplexity only · identical 39Read →
No. 6

The Webflow Agency AI Visibility Index

2 of 376

Fifty Webflow buyer questions produced 376 named providers — 97% known to exactly one engine, and just two (Finsweet, Flow Ninja) known to all three. Of 1,464 real Webflow agencies, 4.2% were ever named.

All 3 enginesRead →
The caveat that governs this whole series. ChatGPT and Gemini do not expose their citation URLs — that's their API, not a gap in our collection. So any finding about sources (Reddit, the long tail, page anatomy) is Perplexity-only, n=39, and every report says so at each claim. Findings about which brands get named can use all three, because names come out of the answer text. Anyone telling you what "AI" cites without naming the engine is guessing.
Check our work

The dataset, and the script that proves it

The 2026 AI Citations Dataset

The raw data behind every report above. Published so you can check our numbers rather than take our word for them — if you recompute something and get a different answer, we want to hear about it.

queries.csv · 39 rows · the questions we asked
multi39.csv · 117 rows · July 1 baseline
multi250.csv · 750 rows · July 16 unified scan (incl. identical 39)
serp250.csv · 2,078 rows · Google top-10 per question
questions250.csv · 250 rows · the full question set
reddit-threads.csv · fetch log (all blocked — disclosed)
multi39_enriched.csv · 39 rows · top-cited pages, coded
verify-numbers.py · recomputes every published figure

python3 verify-numbers.py recomputes every published figure in all six reports — including the Fisher's exact p-values — straight from those CSVs, and exits non-zero if any of them has drifted. We run it before we publish anything.

It exists because we needed it. On 15 July 2026 we found four figures in our own first edition had gone stale — superseded by a recompute that never made it back into the report, while the old numbers quietly propagated further because they'd had longer to spread. A worked example was wrong too. Prose drifts; the CSV doesn't. Every number is now an assertion against this data. The full correction log is in report No. 1's changelog.

Download the dataset →

Does AI recommend you?

Same method, pointed at your domain. We'll ask all three engines your buyers' questions and show you exactly who they name instead of you.

Run the free audit →