Yes — about a third of the time. We set out to document a "self-published penalty" and the data refused to cooperate. So instead we tested four theories about what separates the brand-owned pages that get cited from the third-party pages that do. Three of the four were noise. This report includes them anyway.
39 top-cited pages · 12 self-published vs 27 third-party · Fisher's exact, two-tailed · July 2026 · Perplexity only — see scope note
Of 39 top-cited pages, 12 were self-published — a brand's own page, ranking for a question that brand commercially cares about. 27 were third-party: listicles, directories, review sites, comparison posts.
So the pessimistic story — "AI won't cite you, only the listicles that outrank you" — is wrong. It cites you roughly a third of the time. The interesting part is what's different about those pages.
We tested four hypotheses. Here is every result, including the failures:
| Trait | Self-published | Third-party | p (Fisher) | |
|---|---|---|---|---|
| FAQ blocks or data tables | 75% | 22% | 0.004 | significant |
| Contains original data | 42% | 26% | 0.455 | noise |
| Listicle / "best-of" format | 75% | 78% | 1.000 | noise |
| Has a named author | 42% | 44% | 1.000 | noise |
One trait clears the corrected threshold: structured content. Brand-owned pages that get cited are more than 3× as likely to carry an FAQ block or a data table as the third-party pages that get cited.
We think the mechanism is boring and mechanical. An FAQ block is a question with its answer directly underneath. A table is a set of facts with labels attached. Both are trivially extractable. A retrieval engine looking for something quotable finds it faster in a table than in a paragraph. Third-party listicles don't need that help — they win on being a listicle. Your page apparently does.
Structure isn't decoration. It's how the engine finds the quote.
Original data didn't separate them. 42% vs 26% looks like a gap. At n=12 vs n=27 it's p = 0.455 — comfortably inside the range you'd get from coin flips. Elsewhere in our research original data is a live advantage. On this question at this sample size we can't claim it, so we aren't claiming it.
Format didn't separate them. 75% vs 78% listicle. p = 1.000. Cited pages are overwhelmingly listicles regardless of who publishes them. Being a "best-of" post is table stakes, not an edge.
Bylines didn't separate them. 42% vs 44%. p = 1.000. There's a great deal of GEO advice built on author/E-E-A-T signals. In this data a named author does nothing to distinguish a cited brand page from a cited third-party page. That's not proof it doesn't matter. It's proof we didn't measure it mattering.
And length ran the wrong way. Cited brand-owned pages had a median of 3,158 words; third-party pages ran 3,901. If length were the lever, self-published pages would be longer, not shorter.
An agency's own site was cited nearly half the time. An ecommerce brand's own site was cited once in thirteen questions. If that holds at scale it's worth more than every on-page tactic in this report combined.
But we have to be straight with you: agency vs ecommerce is p = 0.073. It does not clear significance at 13 questions per category. It's the single most interesting thing in this dataset and the single best argument for running the study bigger. We're running it bigger.
FAQ blocks and data tables are the one thing that separated cited brand pages from cited third-party pages at significance. They're also cheap.
They may well matter. They didn't measure as mattering here, and we're not going to pretend otherwise.
If you sell DTC, your own blog cleared 8% in this sample — third-party placement is likely the higher-leverage play. If you're an agency, your own site is a real lane.
Cited brand pages were shorter than cited third-party pages. Length is not the lever.
| Metric | Scope | Value |
|---|---|---|
| Top-cited pages analysed | Perplexity, 39 questions | 39 |
| Self-published | Perplexity, 39 questions | 12 / 39 = 31% |
| Third-party | Perplexity, 39 questions | 27 / 39 = 69% |
| FAQ/tables — self-published | n=12 | 9 = 75% |
| FAQ/tables — third-party | n=27 | 6 = 22% |
| FAQ/tables — significance | Fisher, 2-tailed | p = 0.004 |
| Original data — self-pub vs third-party | n=12 / n=27 | 42% vs 26% · p = 0.455 |
| Listicle — self-pub vs third-party | n=12 / n=27 | 75% vs 78% · p = 1.000 |
| Named author — self-pub vs third-party | n=12 / n=27 | 42% vs 44% · p = 1.000 |
| Median word count — self-published | n=12 | 3,158 |
| Median word count — third-party | n=27 | 3,901 |
| Self-published rate — agencies | n=13 | 6 = 46% |
| Self-published rate — B2B SaaS | n=13 | 5 = 38% |
| Self-published rate — ecommerce | n=13 | 1 = 8% |
| Agencies vs ecommerce | Fisher, 2-tailed | p = 0.073 |
verify-numbers.py recomputes every figure above straight from the CSVs.Yes — 12 of the 39 top-cited pages (31%) were the brand’s own page, not a third-party listicle (Perplexity, July 2026).
Structure. Cited brand pages carried FAQ blocks or data tables 75% of the time versus 22% for cited third-party pages — p=0.004, the only significant trait of the four we tested.
Not measurably here. Named authors: 42% vs 44% (p=1.000). And cited brand pages were shorter than third-party ones — median 3,158 vs 3,901 words — so length is not the lever.
Same method, pointed at your domain. We'll ask the engines your buyers' questions and show you whether your pages — or someone else's — come back.
Run the free audit →