What data sources do AI models use to answer questions about brands?
Modern AI engines draw on four source categories: the training corpus (public web at the model's knowledge cutoff), retrieval-augmented generation (live pages fetched at query time by engines like Perplexity, ChatGPT Search, and Google AI Overviews), structured knowledge bases (Wikidata acting as an entity hub), and user-generated content: Reddit, YouTube, and forums, which has become the most-cited source category in AI-generated answers. A reputation program focused only on Google search results reaches the first category partially and largely misses the rest.
Modern AI engines do not answer from a single place. When a model is asked about a brand, it draws on several distinct source categories, each of which gives a reputation program a different point of leverage.

The four source categories
- 1. The training corpus
- The body of text the model was built on, fixed at its training cutoff: web pages, news archives, books, Wikipedia, and structured datasets. This is the model’s baseline understanding of an entity. Because this knowledge is baked into the model’s weights, it persists even when no live retrieval is performed, but it ages, and what gets included or upweighted during training is not publicly disclosed by model providers.
- 2. Retrieval-augmented generation (RAG)
- Live web pages fetched at the moment the user asks a question, rather than relying on the training corpus alone. Engines including Perplexity, ChatGPT Search, and Google AI Overviews work this way. What ranks on the live web at query time therefore shapes the answer in real time, meaning a brand’s current web presence matters independently of what the model learned at training.
- 3. Structured knowledge bases
- Databases of entity facts, Wikidata being the most prominent, that act as a central hub linking an entity’s representations across sources and languages. A single structured record can surface across many engines and locales, making entity accuracy in these databases a distinct reputation lever.
- 4. User-generated content (UGC)
- Reddit threads, YouTube, and platform-specific forums. This category has become a mainstream AI citation source: a 2026 analysis of roughly 30 million sources found Reddit the most-cited domain in AI-generated answers, followed by YouTube and LinkedIn. A separate 2026 report found the share of AI citations attributed to social media climbed consistently from late 2025 through early 2026, topping 9% of all citations. Classic SEO programs that ignored these channels leave a significant gap in AI-visible brand narrative.
Why a Google-only program misses most of this
The training corpus is set at the model’s cutoff and is not something a single search campaign moves in real time. The retrieval layer pulls from the live web at query time. Structured knowledge sits in entity databases. And the user-generated layer lives in communities: Reddit, YouTube, forums, that classic SEO has tended to ignore. Influencing what AI engines say about a brand means working across all four categories.
Last reviewed: 19/05/2026