How sources behind AI recommendations are manufactured
An independent analysis of 380 software categories examines how AI-integrated search systems build recommendations. It sent 760 queries—one per category to Perplexity/sonar and one to Perplexity/sonar-pro through OpenRouter—and collected 7,534 citations. 59.8% of cited domains were outside the 100,000 most visited sites, while 23.4% did not appear in Tranco’s top million.
The central finding is not that a few famous sites control answers: the ten most-cited sources represented only 17.3% of citations. The striking pattern lies at the periphery, where 751 of 2,055 cited domains were unranked and had more recent Wayback captures. Some appeared designed to be consumed by retrieval systems, with automated pages and descriptions explicitly oriented toward model grounding.
The report identifies three apparently related domains that published 215,128 automated “best software” pages, and highlights guideflow.com, whose blog was cited 194 times across 96 categories despite not being a directory or operating in those markets. The authors compare sitemaps, templates, DNS, registration dates, and archived pages, while noting that shared infrastructure is circumstantial evidence, not proof of common ownership.
The broader lesson is that AI answers can rely on recent, high-volume content designed for machines without making its authority obvious. The study measured only Perplexity—Google was excluded—so its findings should not automatically be generalized to every search engine or model.