Category playbooks
What ChatGPT searches before it names a brand

A brand that everyone in its home market knows can be entirely absent from the English-language answer. Ask ChatGPT for the best compact air purifier for a bedroom, the best sunscreen for oily skin, the best rice cooker for a small kitchen, and you get four or five names. The brand that outsells several of them at home is often not among them.
Most advice about fixing that starts with what to publish. It is the wrong starting point, because it assumes you know what ChatGPT is looking for. You are not competing for the question your buyer typed, and you are not competing on one index.
What does ChatGPT search when someone asks for a recommendation?
Not the sentence the person wrote. Google documents the technique as query fan-out, defined in its own guidance as “a set of concurrent, related queries generated by the model to request more information.” Its worked example: a question about a weedy lawn produces searches like best herbicides for lawns and how to prevent weeds in lawn, neither of which the person typed.
One prompt becomes several searches, and most of what they return is never quoted
- 15,000Prompts a person typed
- 43,233Searches the model actually issued
- 548,534Pages retrieved from those searches
- about 15%Pages that were cited
32.9% of the pages that did get cited appeared only in the results for a fan‑out query — not for the prompt the person typed.
Your page is competing for the generated queries. A third of the pages that earned a citation in the pipeline above appeared only in results for one of them. Content that answers the headline question and nothing around it is absent from most of the searches actually run.
Keyword volume is close to useless here. 95% of those generated queries have no monthly search volume in conventional keyword tools. A page planned around a tool showing 300 searches a month was planned against 5% of the strings that decide the outcome.
The second wave carries names. Fan-out is iterative: the queries generated after the first round of results contain the entities that round surfaced. A brand absent from the first wave is not in the vocabulary available to the second. This is the mechanism behind the pattern every brand notices — ChatGPT describes you accurately when your name is in the question, and cannot find you when it is not.
The wrong response is a page per variation. Google’s guidance names that directly: building pages for every search variation, done mainly to manipulate rankings or AI responses, is scaled content abuse, and its advice is to “focus on what your users want, and avoid overdoing it.” What survives that rule is one document that covers the subtopics a question decomposes into.
The firms publishing fan-out research also sell fan-out tooling, and their headline ratios for how often each engine splits a prompt disagree with each other by more than an order of magnitude. The mechanism is Google’s own documentation. The frequencies are not settled, and Searchd does not repeat them.
Which search index does each AI engine use?
Different ones, and the consoles you already watch report on none of them.
Four assistants, four different indexes, and separate switches for training and retrieval
| Assistant | Index it searches | Training agent | Retrieval agent |
|---|---|---|---|
| ChatGPTOpenAI recommends allowing OAI-SearchBot. Whether that is what fills the index is unresolved. | Its own, plus other pipelinesobserved, not disclosed | GPTBot | OAI-SearchBot |
| ClaudeBlocking Claude-User stops retrieval of your page in response to a question. | Bravelisted on Anthropic's subprocessor page | ClaudeBot | Claude-SearchBot |
| PerplexityPerplexity states PerplexityBot is not used to crawl for foundation models. | Its own, and contestedno settled account | no separate agent | PerplexityBot |
| GoogleGoogle-Extended also governs grounding, and does not affect Search inclusion. | Google'sthe one uncontested row | Google-Extended | Googlebot |
Claude’s web search runs on Brave’s index. ChatGPT serves results from several pipelines, one of which is an index of OpenAI’s own — identified by instrumenting 1,200 answers, where only 1.5% of the URLs appeared in Bing’s top 20 for the same queries and 24% of the titles were longer than Bing’s 75-character cap (Search Engine Land, 17 August 2026). Google’s surfaces read Google’s index. Perplexity runs its own crawler, and how much else it draws on is disputed.
So “we are indexed” is four separate facts, and a brand can be healthy in Search Console and absent from the index ChatGPT or Claude is reading. The cheapest first check Searchd runs on a new brand is not a tool: search their product pages on Brave and see whether they are there at all.
The second check is that the retrieval agents can reach you. The block, when there is one, is usually accidental — a robots.txt line or CDN bot rule written in 2023 or 2024 against AI training, before the retrieval agents existed under their current names. It did what it was asked, and it also removed the brand from the index that decides whether the engine can quote it. Check from outside your own network, because your office IP is often allowlisted:
for ua in "GPTBot/1.4" "OAI-SearchBot/1.4" "ClaudeBot" "Claude-User" "PerplexityBot/1.0"; do
printf '%-22s %s\n' "$ua" \
"$(curl -sI -A "$ua" -o /dev/null -w '%{http_code}' https://yourbrand.com/products/hero)"
done
Anything other than 200 on a page you expect to be quoted is the finding. A 403 is a bot rule. A 200 that returns an empty shell is a rendering problem: these engines do not reliably run your JavaScript, so a specification that appears only after hydration is a specification they do not have.
Two honest limits on this step. The live fetchers sit outside robots.txt by design — OpenAI documents that its rules may not apply to ChatGPT-User because a person initiated the request, and Perplexity says Perplexity-User generally ignores it. And allowing the retrieval agent is necessary rather than sufficient: in the study above, canary pages with server logs showed OAI-SearchBot nearly silent while the index kept serving results, so what fills that index is not established. Watch your own logs rather than assuming the documented agent is the one that came.
Where do you start, and what do you measure first?
Fix a set of questions before you change anything, and record what comes back today. Twenty to forty questions a buyer would ask without knowing your name — a prompt panel — run on each engine you care about, with the date and model version noted.
Do this first because it is the only step that cannot be done later. Once you have shipped a change the before is gone, and every claim you make afterwards about what worked is a story rather than a measurement. Searchd publishes its own at /results — zero, at launch, which is what a new domain should expect.
The method, including why a single check tells you almost nothing, is in how to check if ChatGPT mentions your brand.
Does your brand name resolve to one thing?
An engine has to decide that the home-market company name, the trading name on the invoice, the Instagram handle and the romanised spelling on a retailer page are one brand. That decision is entity resolution, and it is the quietest of the six failures. When it cannot, the evidence for you is divided among several thin identities instead of pooled behind one.

This is ground a cross-border brand loses that a domestic competitor never had to defend. A US brand has one spelling. You have a Hangul or Kana original, at least one romanisation, a legal entity that differs from the trading name, and a distributor who chose a third form on a marketplace listing.
Pick one string. Use it byte-identically on your own site, in your schema, on every marketplace listing and in every directory record. Then link the profiles that prove it is you — sameAs in your organisation schema exists for exactly this, and it is an ownership assertion rather than a guess, so only put profiles in it that are actually yours.
A one-hour job that everything else compounds on, which is why it sits above the expensive work rather than below it.
Is there a sentence on your page worth quoting?
A model does not quote a page. It quotes a passage. The properties that decide whether a paragraph survives being pulled out of its page — its citability — are mechanical.
sample
Lifted out of the page, this survives — the four that decide extraction
The Haruwell Daily Fluid is ratedSPF 50+, tested at the full2 mg/cm² application rate to the2019 edition of ISO 24444.
This does not
As mentioned above, it performs well against the standard discussed, and this is what makes it such a strong choice for the shopper described.
- Names the brandno unresolved “we” or “the company”
- Carries a numbera value that can be checked
- Carries a dateso staleness is visible
- Stands aloneno “as mentioned above”
“Lightweight daily protection with a comfortable finish” is grammatical, accurate and useless. “SPF 50+ PA++++, tested at the full 2 mg/cm² application rate to ISO 24444, chemical filters, no white cast” is a sentence ChatGPT can lift and attribute. The second is not better writing. It is writeable only because someone went and got the numbers.
The move is the same in every category. An air purifier states its CADR figure and the room size that figure was measured against. An appliance states its capacity in both the unit its home market uses and the unit the buyer uses, and the voltage it ships with. A device states the standard, its edition, and the test conditions.
For a brand whose English pages were translated rather than written, this is usually the step with the most to gain. Translation optimises for fidelity to the original sentence. It does not introduce the checkable claim that was never in the original either.
Who owns your product facts — you, or a marketplace listing?
When ChatGPT needs your specification it uses whichever source it retrieved, not whichever source you would prefer. For a cross-border brand that is frequently a marketplace listing written by a distributor two revisions ago.
Your product page
- Weight limit
- 3.5–15 kg
- Newborn insert
- Included
- Last updated
- This season
Distributor’s marketplace listing cited
- Weight limit
- 3.5–15 lb
- Newborn insert
- Sold separately
- Last updated
- Two versions ago
What the assistant then states“It holds 3.5 to 15 pounds, and the newborn insert is a separate purchase.”
The failure has a recognisable shape: units converted wrongly, an age or size range from a superseded revision, an accessory listed as included when it is now sold separately, a reformulation missing from an ingredient list. In a regulated category, an engine stating your product’s rating wrongly is a compliance-shaped problem with your name on it, and the correction is not a website edit — it is a request to a retailer with its own queue.
Find them first: run your panel questions, read the sources the engine cites, and open every one that is not yours.
Why do competitors keep appearing instead of you?
Because the answer is assembled from a reading list — the corpus — and most of it is not yours to write.
Seven surfaces one answer is assembled fromsample
- Your siteyours to write
- Amazon listingyours to correct
- Retailer pagesyours to correct
- Review round-upssomeone else's
- Reddit threadssomeone else's
- YouTube reviewssomeone else's
- Korean reviewswritten, wrong language
How many you write
For a brand whose home market is not English-speaking, this is where the home advantage inverts. The round-ups, comparison threads and long-form reviews that anchor English answers often exist — in Korean, or Japanese, or Chinese. The engine answering an American buyer in English is not reading any of it. The English record is thinner than the brand’s actual reputation, and closing that gap depends on other people publishing, which is why it takes months while the steps above take days.
Does schema markup or an llms.txt file help?
The most expensive part of this work is the part sold most confidently and supported least.
Raw before and after, treated pages only
+43%
Same pages, after matched control pages
+2.4%
| Google AI Overviews | −4.6% | significant, about 1 in 2,500 by chance |
|---|---|---|
| Google AI Mode | +2.4% | not distinguishable from zero |
| ChatGPT | +2.2% | not distinguishable from zero |
Schema markup did not lift citations in the only controlled test we could find. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against matched controls and reported roughly +2.4% on AI Mode and +2.2% on ChatGPT — indistinguishable from noise — and −4.6% on AI Overviews (Ahrefs, 11 May 2026). One caveat that vendors on both sides skip: every page in that study already carried 100+ AI Overview citations, so it shows schema will not lift a page that is already visible. It does not show what schema does for a page an engine has never retrieved. Keep structured data for what it demonstrably does, and do not buy it as a citation lever.
llms.txt is ignored by Google Search. Its guidance states you do not need “new machine readable files, AI text files, markup, or Markdown to appear in Google Search,” and that keeping such files “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them” (Google, updated 10 July 2026). The file is cheap, so keep one if another system reads it. Do not pay for one as a visibility deliverable.
Guaranteed placements cannot be delivered by anyone who does not run the engine. Google’s guidance says so in its own words: “no third-party tool has access to our internal ranking or AI systems.”
More of what we checked and rejected, with the sources, is on our evidence page.
How will you know whether any of it worked?
Re-run the same panel, and compare against pages you deliberately left alone. Engine answers vary between runs, so a single before-and-after on a single page is a coin flip presented as a result. A control group is what separates “we changed something and the number moved” from “we changed something.”
Expect the first weeks to read as zero. A new or newly-corrected corpus takes time to be retrieved, and treating early noise as failure is how teams abandon the cheap steps before they have had a chance to matter.
Who does this work — your team, or an agency?
You can do all of it yourself. The first four steps need an afternoon each and a person willing to read robots.txt. The fifth is a queue at someone else’s company, so it needs someone who will keep asking. The sixth needs outreach sustained for months, and it is the one that quietly slips.
Searchd does this as an engagement — the panel, the index and access audit, the entity work, the passage rewriting, the marketplace corrections and the third-party outreach — and publishes its own number while doing it. What that engagement guarantees, what it only commits to, and what it refuses to promise are set out on the audit page and the retainer page. The short version is that we guarantee the layer we control and measure the rest, because the alternative is selling a mechanism nobody outside the engine can see.
Related questions
Do we need a page for every question an AI might ask?
No, and Google names that pattern as a spam policy violation. Its guidance says building pages for every search variation, mainly to manipulate rankings or AI responses, falls under scaled content abuse and is ineffective as a strategy. The version that works is one page that genuinely covers the subtopics a question decomposes into, which is why this page is organised as sections rather than split into nine thin ones.
If we already rank on Google, are we not covered?
Only for Google's own surfaces. Google's index feeds AI Overviews and AI Mode. Claude's web search runs on a different index, ChatGPT serves results from pipelines including one of its own, and Perplexity operates its own crawler. Search Console and Bing Webmaster Tools report on neither of the last three, so a brand can be indexed everywhere it checks and absent from the index ChatGPT or Claude is actually reading.
Should we block AI crawlers to protect our content?
That is a legitimate business decision, but make it per agent rather than as one rule. Training crawlers and retrieval agents are separate tokens at every engine, so a blanket block written against training also removes you from the index that decides whether the engine can quote you. If you want out of training but still quotable, block the training agent and allow the retrieval one.
How long before anything changes?
The first four steps are days of work each, and their effect shows in the next panel run. Correcting a retailer's listing runs at that retailer's pace, usually weeks. The last step is months, because it depends on other people publishing. Anyone quoting you a date for a citation is quoting a date for someone else's behaviour.
Is this just SEO with a new name?
For Google's surfaces, Google says it is: its own guidance states that optimizing for generative AI search is optimizing for the search experience, and thus still SEO. That settles one surface. ChatGPT, Claude and Perplexity retrieval are not covered by that guidance, and the steps here that touch those engines — which index they read, their separate retrieval agents, the entity they resolve you to — are where the work stops looking like an SEO deliverable.