
On July 7, 2026, we ran a controlled citation audit: 30 real buyer and how-to queries submitted to ChatGPT (gpt-5.5), Perplexity (sonar), Gemini (gemini-3.5-flash), and Claude (claude-sonnet-4-6), each with live web search enabled, plus Google AI Overviews pulled from live US desktop SERPs. That produced 150 query-surface observations: 140 generated-answer observations and 10 observations with no Google AI Overview. The analysis counted 524 distinct cited domains across the generated answers. These figures are self-reported by SCALZ.AI; response-level raw files are not yet publicly downloadable or independently reproduced.
We did this for a selfish reason. SCALZ.AI sells answer engine optimization, and we wanted a truthful baseline of where we stand before claiming anything to clients. The honest answer: we appeared in exactly 1 of 150 query-surface observations. That number is in this study, alongside everyone else's, because a benchmark you cannot fail is not a benchmark.
In this sample, the most visible pattern was structural. Engines rarely cited business service pages directly and more often surfaced Reddit threads, YouTube videos, agency directories, and ranked lists. A single-day, agency-focused sample cannot establish a universal preference, but it gives service businesses a testable reason to measure third-party and comparison surfaces alongside their own pages.
How Did We Run This Citation Study?
We submitted 30 fixed queries, split into buyer intent, how-to, and local vertical categories, to four LLM engines with web search enabled and to Google AI Overviews, on a single day using the DataForSEO AI Optimization and SERP APIs. All 150 query-surface checks returned usable collection records. Of those, 140 contained generated answers and 10 recorded no Google AI Overview. We normalized cited URLs to their root domains for the published analysis.
The query set mirrors real agency-buyer behavior in three buckets of ten. Buyer intent: queries like best answer engine optimization agency and best AI SEO agency. How-to: queries like what is answer engine optimization and how to get cited by ChatGPT. Local and vertical money queries: Tampa SEO agency, HVAC SEO company, dental SEO company, and similar. Each query ran once per engine, giving one clean snapshot per engine per query.
Engine versions matter for reproducibility, so here they are: ChatGPT ran gpt-5.5, Perplexity ran sonar, Gemini ran gemini-3.5-flash, and Claude ran claude-sonnet-4-6, all with web search or grounding active. Google AI Overviews came from live desktop SERPs, US location, English. AI Overviews appeared on 20 of the 30 queries; the other 10 returned standard results with no AI block.
One technical note for anyone replicating this: Gemini masks its citation URLs behind vertexaisearch redirect links, so raw URL logging undercounts real sources. We recovered the true domains from the citation titles Gemini returns alongside each redirect. Counts in this study are unique domains per response, so a domain cited five times in one answer counts once for that answer.
Which Domains Get Cited Most by AI Engines?
Reddit leads with 46 citation appearances across 22 of 30 queries, followed by YouTube with 44 across 24 queries. The top brand-owned domain is firstpagesage.com with 25 appearances, earned through published agency rankings. LinkedIn, Semrush, and the Clutch directory round out the leaders. In total, 524 unique domains were cited.
The full top seven: reddit.com with 46 appearances, youtube.com with 44, firstpagesage.com with 25, thriveagency.com with 21, semrush.com and linkedin.com with 15 each, and clutch.co with 14. Appearances means the number of query-surface observations citing that domain at least once, out of 150. Nearly a third of all AI answers we collected cited Reddit somewhere.
Look at what those domains have in common. Two are user-generated content platforms. One is a directory. One is a professional network. The two agencies on the list, First Page Sage and Thrive, both publish ranked lists of agencies, including lists they appear on themselves. Not one classic service page or brochure-style agency site cracked the leaderboard.
The long tail is brutal. Of 524 unique domains cited, the median domain appeared exactly once. Citation share in AI answers is concentrated at the top in a way that resembles early-2010s SERP click curves, except the winners are aggregation and discussion surfaces rather than the businesses themselves.
Why Do Reddit and YouTube Dominate AI Citations?
In this sample, community discussion and video were retrieved often. Reddit pages contain multiple public viewpoints, while YouTube transcripts can provide crawlable explanations. The study did not test why a model selected any source, so these are observed patterns rather than proof of an engine-wide preference.
Reddit was cited in 46 of the 150 query-surface observations, roughly 31 percent. It won 22 of 30 queries in at least one engine, and it led both the buyer intent bucket, with 20 appearances, and the how-to bucket, with 22. In this sample, buyer-intent answers frequently cited comparison threads and rarely cited an agency statement about itself.
YouTube's 44 appearances came disproportionately from Gemini, which cited it 20 times, more than any other engine cited any domain. That is a Google ecosystem effect, but it is not only that: Perplexity cited YouTube 14 times and AI Overviews 10. Video transcripts are dense, structured, spoken-language answers, which is precisely the format retrieval systems excerpt well.
The uncomfortable implication for service businesses: two of the biggest citation surfaces cannot be bought or directly controlled. You earn presence there by participating credibly, publishing genuinely useful video, or being discussed favorably by people you do not pay. Astroturfing Reddit is a brand-destroying shortcut, and engines and moderators both keep getting better at detecting it.
Who Wins Buyer-Intent Queries Like Best AEO Agency?
Publishers of ranked lists were prominent in this snapshot. After Reddit and YouTube, firstpagesage.com had 13 buyer-intent appearances, minuttia.com and yesoptimist.com had 11 each, and thriveagency.com had 8. Each publishes best-of rankings, but this observational study does not establish that format as the sole cause of citation.
This is the most actionable finding in the study. For the best-agency queries in this sample, ranked comparison documents appeared more often than agency service pages. The collection shows what was cited on July 7, not how each model evaluated every candidate source. First Page Sage has industrialized this: 25 total appearances across 13 different queries, the most cited agency domain in the entire dataset.
The pattern holds for the emerging AEO-specific players too. Domains like ziptie.dev, onely.com, contently.com, and aeoengine.ai each picked up multiple citations on generative engine optimization queries by publishing opinionated, structured content about the discipline itself: definitions, methodologies, and yes, rankings of other providers.
There is an integrity question here worth answering plainly. Ranking yourself first in your own list is a credibility tax that readers and engines will eventually price in. The durable version of this play is an honest ranked list with transparent methodology, real competitor strengths, and disclosed authorship. That is the version we are building for our own site, and this study is the methodology section.
Who Wins Local and Vertical Agency Queries?
Directories were prominent in this snapshot. On queries such as Tampa SEO agency and HVAC SEO company, clutch.co had 14 appearances, followed by thriveagency.com with 13, firstpagesage.com with 9, designrush.com with 7, and goodfirms.co and Semrush Agency Partners with 6 each. The study did not test whether engines treat any directory as pre-vetted.
The vertical bucket behaves differently from the other two. Reddit drops to 4 appearances because there are fewer high-quality threads comparing, say, dental SEO companies. Into that vacuum step the directories: Clutch, DesignRush, GoodFirms, and Semrush's agency marketplace. The retrieved results often had the shape of provider lists, and directories supplied that format in this sample.
Thrive Internet Marketing Agency deserves study here. Its 21 total appearances come largely from city-targeted pages and best-companies listicles for individual metros. It is running the ranked-list play and the local landing page play simultaneously, and both are earning AI citations on money queries in market after market.
For an agency or local service business, complete and accurately categorized directory profiles can support discoverability beyond the company website. In this sample, directories were cited on buyer queries where individual provider sites were less visible. That observation supports testing directory coverage, not treating any profile as a guaranteed citation path.
How Do the Five Engines Differ in Citation Behavior?
Substantially. ChatGPT cited the fewest sources, averaging 2.5 unique domains per answer with a scattered long tail. Gemini averaged 10.3 with heavy YouTube bias. Perplexity averaged 9.2, favoring Reddit and LinkedIn. Claude averaged 7.0 and leaned on agency listicles. AI Overviews averaged 5.5 across the 20 queries where they appeared.
ChatGPT is the outlier. With gpt-5.5 and web search, its answers cite few sources and no domain dominated: its most cited domains appeared in just 2 responses each. Practically, that means ChatGPT citation presence is harder to engineer and noisier to measure than the other engines, and single-run ChatGPT audits will swing more between runs.
Gemini and Perplexity are the citation-rich engines, but they reach for different shelves. Gemini's top sources were YouTube with 20 and Reddit with 12, consistent with Google ecosystem grounding. Perplexity's were Reddit with 21, YouTube with 14, and LinkedIn with 11, making it the engine most influenced by professional and community discussion. Claude's top sources were First Page Sage and Thrive with 6 each, then SEOProfy with 5: it retrieves and trusts edited, ranked, long-form comparisons.
Google AI Overviews sit in the middle and remain the most volatile surface: present on 20 of 30 queries in our snapshot, with Reddit and YouTube again on top at 13 and 10 followed by the same listicle publishers. Google's own documentation on AI features in Search is worth reading for how it frames source selection; the operational takeaway from our data is that no single engine strategy covers you. Engine mix is a real strategic variable now.
What Did We Learn About Our Own Citation Presence?
That we are at the starting line. SCALZ.AI appeared in 1 of 150 query-surface observations: the Google AI Overview for best SEO company in Orlando cited scalz.ai 4 times via our Orlando page, even though that page was not in the organic top 10. Zero citations across ChatGPT, Perplexity, Gemini, and Claude.
We are publishing that number on purpose. Every agency selling AEO should be able to answer one question: where do you show up in AI answers, measured how, as of when? Our answer is 0.7 percent of query-surface observations on a 30-query buyer set as of July 7, 2026. That is the baseline this program will be judged against, publicly, in follow-up runs.
The single hit is instructive, though. The Orlando AI Overview cited our page four separate times while ranking it outside the organic top 10. Citation selection and organic ranking are correlated but not identical systems: a tightly structured city page with direct answers can be quoted by the AI layer even when it loses the blue-link fight. That asymmetry is the entire opportunity in AEO right now.
The run demonstrates a measurement approach. This one-day, 150-observation audit produced a competitor leaderboard, surface-level citation patterns, and an internal baseline. Collection time and cost will vary by tool, model, query count, and review process. We will rerun the same fixed query set on a recurring cadence, because a single snapshot proves presence, not stability, and stability is what clients are actually buying.
What Are the Limitations of This Study?
Four big ones. It is a single-day snapshot of non-deterministic systems. It tests one model per engine, not every tier users encounter. It is US-English only. And unique-domain counting understates how heavily an engine leans on one source within a single answer. Treat the numbers as directional, not permanent.
Non-determinism is the honest headline caveat. Ask the same engine the same question tomorrow and you will get a different answer with partially different citations. The 150-observation sample reduces some of that noise, and similar patterns across Reddit, YouTube, listicles, and directories make those surfaces reasonable candidates for another controlled run. But any specific per-query citation, including our Orlando one, can vanish in the next model update.
Model selection shapes results too. We tested the web-search-enabled tiers of each engine because those are the citation-producing modes, but consumers on free tiers may hit different models with different retrieval behavior. Our probe testing also found that a model without web search interpreted the acronym AEO as Authorized Economic Operator, a customs term. Retrieval does not just add citations, it disambiguates entire topics.
Finally, the query set is agency-centric by design because that is our market. The domain leaderboard would look different for e-commerce, healthcare, or B2B software queries, though we would bet the structural pattern of community content, rankers, and directories over brand sites survives across most service verticals. That is a testable claim, and we intend to test it.
How Can You Replicate This Audit for Your Own Brand?
Fix a query set of 20 to 30 questions your buyers actually ask, run each through the web-search tiers of the major engines plus Google AI Overviews, log every cited domain, and score your presence as a share of responses. Collection time and cost depend on the tools, models, query count, and review process.
Start with the query set, because it is the part you must never change between runs. Split it three ways: buyer-intent queries with your category and qualifiers, how-to queries your prospects ask on the way to hiring someone, and local or vertical money queries naming your cities and niches. Write them the way a customer types, not the way your industry talks. Twenty to thirty queries can be a manageable starting point, but the set must match the decisions the study is intended to support.
For collection, you can run queries by hand in each interface and paste citations into a sheet, which works but does not scale past one run. We used the DataForSEO AI Optimization API for the four LLM engines and its SERP endpoint for AI Overviews, logging raw responses to files so the analysis is rerunnable. Whatever tooling you pick, record the model version, the date, and whether web search was on, because results are not comparable without those three facts.
Then score three things: your citation rate as the share of responses citing your domain at least once, your text mention rate for unlinked brand mentions, and the competitor leaderboard of domains cited most on your query set. The leaderboard is usually the most valuable output because it tells you which surfaces, lists, directories, or communities are collecting the citations you want, which is exactly where your next quarter of AEO work should point.
What Are the 7 Takeaways for Getting Cited by AI?
Build for the surfaces engines actually cite. That means ranked lists and original data on your own site, real presence on Reddit, YouTube, and LinkedIn, complete directory profiles, engine-specific tactics, and a fixed-query measurement habit. Service pages did not lead citation counts for the buyer queries in this snapshot.
- Publish honest ranked lists with transparent methodology. List publishers won the buyer-intent bucket outright, and the most cited agency in this snapshot appeared through ranking content rather than a service page.
- Publish original data. A study with inspectable evidence can be easier to evaluate and cite than an unsupported opinion. Publish the protocol, observations, counting rules, and limitations.
- Complete and maintain directory profiles on Clutch, DesignRush, GoodFirms, and Semrush Agency Partners. Directories collected 33 combined appearances on local and vertical money queries.
- Build a credible Reddit participation habit in your niche. Reddit was the most cited domain in this snapshot. Participate only when you can add genuine, policy-compliant value, and measure whether that surface matters in your own market.
- Treat YouTube as crawlable answers, not branding. YouTube was the second most cited domain in this snapshot and appeared often in Gemini results. Test question-focused titles, chapters, and accurate transcripts without assuming any one format guarantees retrieval.
- Match tactics to engines. This snapshot associated Perplexity with community and LinkedIn sources, Claude with long-form comparisons, Gemini with Google ecosystem assets, and ChatGPT with a diffuse source set. Re-test before applying those patterns to another market.
- Measure with a fixed query set on a recurring cadence. One snapshot is a baseline; only repeated runs distinguish real citation presence from one-time luck.
Sources and further reading
These are the primary sources referenced in this article. Each is an authoritative documentation page or publication we verified before citing.
- Google's own documentation on AI features in Search: engine differences section, framing how Google describes source selection for AI features
- DataForSEO AI Optimization API: methodology section, the API used to collect LLM responses

