How does ChatGPT decide which brands to mention?
Two pathways. Training-data associations: brands consistently described across many trusted sources get durably linked to their category in the model's weights. Live retrieval (ChatGPT Search): OAI-SearchBot fetches current pages, favoring crawlable, answer-structured, authoritative sources — with a documented preference for third-party listicles when users ask for 'best X' recommendations.
The two compound: retrieval surfaces you today, and the coverage that earns retrieval becomes tomorrow's training data. Citation engineering means feeding both pipelines deliberately.
The five-step citation program
One: technical doors open — GPTBot and OAI-SearchBot allowed, everything server-rendered, no content locked behind JS interaction. Two: answer-first money pages with 40–60 word answers and honest numbers (pricing pages get cited most). Three: entity clarity — Organization schema, identical one-line brand description everywhere, sameAs links binding your web presence together. Four: third-party saturation — industry listicles, Crunchbase/Clutch-class directories, Reddit and Quora presence, trade press. Five: original data — publish the benchmark or stat report your niche lacks, because unique numbers are the #1 citation magnet.
- robots.txt: allow GPTBot, OAI-SearchBot
- Answer blocks + real pricing on money pages
- One canonical brand description, used everywhere
- Listicles, directories, Reddit, trade press — repeatedly
- Publish original stats competitors must cite
Measuring mention rate honestly
Define 15–25 prompts your buyers actually use ('best web3 marketing agency', 'how much does crypto PR cost', 'who builds custom AI agents'). Run them monthly in fresh sessions, logging mention (named at all), citation (linked as source) and position (first recommendation or afterthought). Expect noise between runs; judge the 90-day trend.
This scoreboard is exactly what we run for Chalk Labs itself — selling GEO while being invisible to the engines would be the emperor's-new-clothes version of this business.
Questions we hear about this
OpenAI has never confirmed it, and Google explicitly says such files are unnecessary for its AI features. Ship one as a five-minute hedge (we do), but the leverage is in crawlability, answer structure, entity clarity and third-party presence.
Almost always third-party footprint: they appear in more listicles, directories and discussions with consistent descriptions. The model isn't evaluating your product; it's reflecting the corpus. The fix is presence-building — digital PR, directories, community — not homepage rewrites alone.
Not directly — there's no ad product for organic answer mentions. Indirectly, sponsored listicle inclusions in crawlable publications do feed the corpus. Treat clearly-labeled sponsored presence as a legitimate, minor input; treat anyone selling 'guaranteed ChatGPT rankings' as a scam.
Perplexity is retrieval-first with visible citations on every answer, so page-level answer quality and freshness matter even more, while training-data associations matter somewhat less. The same program covers both; the measurement set should include each engine separately.