Category playbooks

What ChatGPT searches before it names a brand

Short answer. ChatGPT does not search the sentence your buyer typed — it generates several searches of its own and answers from what those return. The engines also read different indexes: Claude reads Brave, Google reads Google, and ChatGPT serves from pipelines including one of its own. Searchd works the order that follows from that: baseline panel, retrieval access, one resolved brand name, a quotable specification, corrected marketplace facts, then the third-party corpus.

A brand that everyone in its home market knows can be entirely absent from the English-language answer. Ask ChatGPT for the best compact air purifier for a bedroom, the best sunscreen for oily skin, the best rice cooker for a small kitchen, and you get four or five names. The brand that outsells several of them at home is often not among them.

Most advice about fixing that starts with what to publish. It is the wrong starting point, because it assumes you know what ChatGPT is looking for. You are not competing for the question your buyer typed, and you are not competing on one index.

What does ChatGPT search when someone asks for a recommendation?

Not the sentence the person wrote. Google documents the technique as query fan-out, defined in its own guidance as “a set of concurrent, related queries generated by the model to request more information.” Its worked example: a question about a weedy lawn produces searches like best herbicides for lawns and how to prevent weeds in lawn, neither of which the person typed.

One prompt becomes several searches, and most of what they return is never quoted

  1. 15,000Prompts a person typed
  2. 43,233Searches the model actually issued
  3. 548,534Pages retrieved from those searches
  4. about 15%Pages that were cited

32.9% of the pages that did get cited appeared only in the results for a fan‑out query — not for the prompt the person typed.

Source: AirOps, measuring ChatGPT. The bars are the counts, drawn to scale; the last is a share of the row above it.

Your page is competing for the generated queries. A third of the pages that earned a citation in the pipeline above appeared only in results for one of them. Content that answers the headline question and nothing around it is absent from most of the searches actually run.

Keyword volume is close to useless here. 95% of those generated queries have no monthly search volume in conventional keyword tools. A page planned around a tool showing 300 searches a month was planned against 5% of the strings that decide the outcome.

The second wave carries names. Fan-out is iterative: the queries generated after the first round of results contain the entities that round surfaced. A brand absent from the first wave is not in the vocabulary available to the second. This is the mechanism behind the pattern every brand notices — ChatGPT describes you accurately when your name is in the question, and cannot find you when it is not.

The wrong response is a page per variation. Google’s guidance names that directly: building pages for every search variation, done mainly to manipulate rankings or AI responses, is scaled content abuse, and its advice is to “focus on what your users want, and avoid overdoing it.” What survives that rule is one document that covers the subtopics a question decomposes into.

The firms publishing fan-out research also sell fan-out tooling, and their headline ratios for how often each engine splits a prompt disagree with each other by more than an order of magnitude. The mechanism is Google’s own documentation. The frequencies are not settled, and Searchd does not repeat them.

Which search index does each AI engine use?

Different ones, and the consoles you already watch report on none of them.

Four assistants, four different indexes, and separate switches for training and retrieval

AssistantIndex it searchesTraining agentRetrieval agent
ChatGPTOpenAI recommends allowing OAI-SearchBot. Whether that is what fills the index is unresolved.Its own, plus other pipelinesobserved, not disclosedGPTBotOAI-SearchBot
ClaudeBlocking Claude-User stops retrieval of your page in response to a question.Bravelisted on Anthropic's subprocessor pageClaudeBotClaude-SearchBot
PerplexityPerplexity states PerplexityBot is not used to crawl for foundation models.Its own, and contestedno settled accountno separate agentPerplexityBot
GoogleGoogle-Extended also governs grounding, and does not affect Search inclusion.Google'sthe one uncontested rowGoogle-ExtendedGooglebot
Agent tokens are each vendor's own, read from OpenAI, Anthropic, Perplexity and Google on 8 September 2026. Tokens change; re-read them before you write a rule.Index attributions are weaker evidence and differ per row. OpenAI's own index was identified by instrumenting 1,200 ChatGPT answers and is not disclosed by OpenAI (Search Engine Land, 17 August 2026). Brave appears on Anthropic's subprocessor list under Web Search; Anthropic has not confirmed the relationship on the record, and we read that listing through secondary reporting rather than the page itself. Perplexity's mix is disputed between sources that sell competing accounts of it, so we do not resolve it here.

Claude’s web search runs on Brave’s index. ChatGPT serves results from several pipelines, one of which is an index of OpenAI’s own — identified by instrumenting 1,200 answers, where only 1.5% of the URLs appeared in Bing’s top 20 for the same queries and 24% of the titles were longer than Bing’s 75-character cap (Search Engine Land, 17 August 2026). Google’s surfaces read Google’s index. Perplexity runs its own crawler, and how much else it draws on is disputed.

So “we are indexed” is four separate facts, and a brand can be healthy in Search Console and absent from the index ChatGPT or Claude is reading. The cheapest first check Searchd runs on a new brand is not a tool: search their product pages on Brave and see whether they are there at all.

The second check is that the retrieval agents can reach you. The block, when there is one, is usually accidental — a robots.txt line or CDN bot rule written in 2023 or 2024 against AI training, before the retrieval agents existed under their current names. It did what it was asked, and it also removed the brand from the index that decides whether the engine can quote it. Check from outside your own network, because your office IP is often allowlisted:

for ua in "GPTBot/1.4" "OAI-SearchBot/1.4" "ClaudeBot" "Claude-User" "PerplexityBot/1.0"; do
  printf '%-22s %s\n' "$ua" \
    "$(curl -sI -A "$ua" -o /dev/null -w '%{http_code}' https://yourbrand.com/products/hero)"
done

Anything other than 200 on a page you expect to be quoted is the finding. A 403 is a bot rule. A 200 that returns an empty shell is a rendering problem: these engines do not reliably run your JavaScript, so a specification that appears only after hydration is a specification they do not have.

Two honest limits on this step. The live fetchers sit outside robots.txt by design — OpenAI documents that its rules may not apply to ChatGPT-User because a person initiated the request, and Perplexity says Perplexity-User generally ignores it. And allowing the retrieval agent is necessary rather than sufficient: in the study above, canary pages with server logs showed OAI-SearchBot nearly silent while the index kept serving results, so what fills that index is not established. Watch your own logs rather than assuming the documented agent is the one that came.

Where do you start, and what do you measure first?

Fix a set of questions before you change anything, and record what comes back today. Twenty to forty questions a buyer would ask without knowing your name — a prompt panel — run on each engine you care about, with the date and model version noted.

Do this first because it is the only step that cannot be done later. Once you have shipped a change the before is gone, and every claim you make afterwards about what worked is a story rather than a measurement. Searchd publishes its own at /results — zero, at launch, which is what a new domain should expect.

The method, including why a single check tells you almost nothing, is in how to check if ChatGPT mentions your brand.

Does your brand name resolve to one thing?

An engine has to decide that the home-market company name, the trading name on the invoice, the Instagram handle and the romanised spelling on a retailer page are one brand. That decision is entity resolution, and it is the quietest of the six failures. When it cannot, the evidence for you is divided among several thin identities instead of pooled behind one.

Five broken pieces of one grey stone bar stand apart with gaps between them, their fracture faces still matching; beside them the same bar whole and unbroken in deep blue. The five stand for one brand written five ways — Seolwha, SEOLWHA Co., Ltd., 설화, Sulhwa and @seolwha_official — with the evidence for it divided five ways so each looks thin. The whole bar is that same evidence pooled behind a single name.
One brand written five ways — Seolwha, SEOLWHA Co., Ltd.,설화, Sulhwa, @seolwha_official — divides the evidence between five identities, and each one then looks too thin to cite. Names are placeholders. What resolution changes is not how much is written about a brand, it is how much of it counts as being about the same brand. Consistent naming, one canonical page and matching structured data are how the resolved state is reached.

This is ground a cross-border brand loses that a domestic competitor never had to defend. A US brand has one spelling. You have a Hangul or Kana original, at least one romanisation, a legal entity that differs from the trading name, and a distributor who chose a third form on a marketplace listing.

Pick one string. Use it byte-identically on your own site, in your schema, on every marketplace listing and in every directory record. Then link the profiles that prove it is you — sameAs in your organisation schema exists for exactly this, and it is an ownership assertion rather than a guess, so only put profiles in it that are actually yours.

A one-hour job that everything else compounds on, which is why it sits above the expensive work rather than below it.

Is there a sentence on your page worth quoting?

A model does not quote a page. It quotes a passage. The properties that decide whether a paragraph survives being pulled out of its page — its citability — are mechanical.

sample

Lifted out of the page, this survives — the four that decide extraction

The Haruwell Daily Fluid is ratedSPF 50+, tested at the full2 mg/cm² application rate to the2019 edition of ISO 24444.

This does not

As mentioned above, it performs well against the standard discussed, and this is what makes it such a strong choice for the shopper described.
  • Names the brandno unresolved “we” or “the company”
  • Carries a numbera value that can be checked
  • Carries a dateso staleness is visible
  • Stands aloneno “as mentioned above”
Sample copy for a fictional brand, written twice about the same product. The second version loses every property that lets a model quote it — so a passage that reads fine in place becomes unusable the moment it is pulled out of the page.

“Lightweight daily protection with a comfortable finish” is grammatical, accurate and useless. “SPF 50+ PA++++, tested at the full 2 mg/cm² application rate to ISO 24444, chemical filters, no white cast” is a sentence ChatGPT can lift and attribute. The second is not better writing. It is writeable only because someone went and got the numbers.

The move is the same in every category. An air purifier states its CADR figure and the room size that figure was measured against. An appliance states its capacity in both the unit its home market uses and the unit the buyer uses, and the voltage it ships with. A device states the standard, its edition, and the test conditions.

For a brand whose English pages were translated rather than written, this is usually the step with the most to gain. Translation optimises for fidelity to the original sentence. It does not introduce the checkable claim that was never in the original either.

Who owns your product facts — you, or a marketplace listing?

When ChatGPT needs your specification it uses whichever source it retrieved, not whichever source you would prefer. For a cross-border brand that is frequently a marketplace listing written by a distributor two revisions ago.

Your product page

Weight limit
3.5–15 kg
Newborn insert
Included
Last updated
This season

Distributor’s marketplace listing cited

Weight limit
3.5–15 lb
Newborn insert
Sold separately
Last updated
Two versions ago

What the assistant then states“It holds 3.5 to 15 pounds, and the newborn insert is a separate purchase.”

The listing here is invented, with the kilogram figures carrying the wrong unit, which is the error it fails on. The point is the mechanism: whichever page an assistant treats as the source of record is the one your buyer hears back, correct or not. Which page yours is, is checkable — ask the question and read the sources.

The failure has a recognisable shape: units converted wrongly, an age or size range from a superseded revision, an accessory listed as included when it is now sold separately, a reformulation missing from an ingredient list. In a regulated category, an engine stating your product’s rating wrongly is a compliance-shaped problem with your name on it, and the correction is not a website edit — it is a request to a retailer with its own queue.

Find them first: run your panel questions, read the sources the engine cites, and open every one that is not yours.

Why do competitors keep appearing instead of you?

Because the answer is assembled from a reading list — the corpus — and most of it is not yours to write.

Seven surfaces one answer is assembled fromsample

  • Your siteyours to write
  • Amazon listingyours to correct
  • Retailer pagesyours to correct
  • Review round-upssomeone else's
  • Reddit threadssomeone else's
  • YouTube reviewssomeone else's
  • Korean reviewswritten, wrong language

How many you write

1/7is a page you controlThe other six decide the shortlist.
You write one of these seven. Two more you can get corrected. Four are written by other people, and an English answer is assembled almost entirely from English sources — which is why the last row matters: a decade of reviews already exists in Korean, and the answer your US buyer reads was written without any of it. Which rows apply to you is what an audit establishes. The last four rows are the work.

For a brand whose home market is not English-speaking, this is where the home advantage inverts. The round-ups, comparison threads and long-form reviews that anchor English answers often exist — in Korean, or Japanese, or Chinese. The engine answering an American buyer in English is not reading any of it. The English record is thinner than the brand’s actual reputation, and closing that gap depends on other people publishing, which is why it takes months while the steps above take days.

Does schema markup or an llms.txt file help?

The most expensive part of this work is the part sold most confidently and supported least.

Raw before and after, treated pages only

+43%

Same pages, after matched control pages

+2.4%

Change in citations 30 days after JSON‑LD was added, per engine, after controls
Google AI Overviews−4.6%significant, about 1 in 2,500 by chance
Google AI Mode+2.4%not distinguishable from zero
ChatGPT+2.2%not distinguishable from zero
1,885 pages that added JSON‑LD schema, each matched against three control pages on other domains with similar prior citation levels, tracked August 2025 to March 2026. Source: Ahrefs, 11 May 2026. We did not run this study. We read it, and it is the reason we do not sell schema as a citation lever.

Schema markup did not lift citations in the only controlled test we could find. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against matched controls and reported roughly +2.4% on AI Mode and +2.2% on ChatGPT — indistinguishable from noise — and −4.6% on AI Overviews (Ahrefs, 11 May 2026). One caveat that vendors on both sides skip: every page in that study already carried 100+ AI Overview citations, so it shows schema will not lift a page that is already visible. It does not show what schema does for a page an engine has never retrieved. Keep structured data for what it demonstrably does, and do not buy it as a citation lever.

llms.txt is ignored by Google Search. Its guidance states you do not need “new machine readable files, AI text files, markup, or Markdown to appear in Google Search,” and that keeping such files “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them” (Google, updated 10 July 2026). The file is cheap, so keep one if another system reads it. Do not pay for one as a visibility deliverable.

Guaranteed placements cannot be delivered by anyone who does not run the engine. Google’s guidance says so in its own words: “no third-party tool has access to our internal ranking or AI systems.”

More of what we checked and rejected, with the sources, is on our evidence page.

How will you know whether any of it worked?

Re-run the same panel, and compare against pages you deliberately left alone. Engine answers vary between runs, so a single before-and-after on a single page is a coin flip presented as a result. A control group is what separates “we changed something and the number moved” from “we changed something.”

Expect the first weeks to read as zero. A new or newly-corrected corpus takes time to be retrieved, and treating early noise as failure is how teams abandon the cheap steps before they have had a chance to matter.

Who does this work — your team, or an agency?

You can do all of it yourself. The first four steps need an afternoon each and a person willing to read robots.txt. The fifth is a queue at someone else’s company, so it needs someone who will keep asking. The sixth needs outreach sustained for months, and it is the one that quietly slips.

Searchd does this as an engagement — the panel, the index and access audit, the entity work, the passage rewriting, the marketplace corrections and the third-party outreach — and publishes its own number while doing it. What that engagement guarantees, what it only commits to, and what it refuses to promise are set out on the audit page and the retainer page. The short version is that we guarantee the layer we control and measure the rest, because the alternative is selling a mechanism nobody outside the engine can see.

Related questions

Do we need a page for every question an AI might ask?

No, and Google names that pattern as a spam policy violation. Its guidance says building pages for every search variation, mainly to manipulate rankings or AI responses, falls under scaled content abuse and is ineffective as a strategy. The version that works is one page that genuinely covers the subtopics a question decomposes into, which is why this page is organised as sections rather than split into nine thin ones.

If we already rank on Google, are we not covered?

Only for Google's own surfaces. Google's index feeds AI Overviews and AI Mode. Claude's web search runs on a different index, ChatGPT serves results from pipelines including one of its own, and Perplexity operates its own crawler. Search Console and Bing Webmaster Tools report on neither of the last three, so a brand can be indexed everywhere it checks and absent from the index ChatGPT or Claude is actually reading.

Should we block AI crawlers to protect our content?

That is a legitimate business decision, but make it per agent rather than as one rule. Training crawlers and retrieval agents are separate tokens at every engine, so a blanket block written against training also removes you from the index that decides whether the engine can quote you. If you want out of training but still quotable, block the training agent and allow the retrieval one.

How long before anything changes?

The first four steps are days of work each, and their effect shows in the next panel run. Correcting a retailer's listing runs at that retailer's pace, usually weeks. The last step is months, because it depends on other people publishing. Anyone quoting you a date for a citation is quoting a date for someone else's behaviour.

Is this just SEO with a new name?

For Google's surfaces, Google says it is: its own guidance states that optimizing for generative AI search is optimizing for the search experience, and thus still SEO. That settles one surface. ChatGPT, Claude and Perplexity retrieval are not covered by that guidance, and the steps here that touch those engines — which index they read, their separate retrieval agents, the entity they resolve you to — are where the work stops looking like an SEO deliverable.

AEO vs SEO: what actually changes

The technical foundations are shared. The content shape, the off-site work and the definition of success are not. A practical breakdown of where the two disciplines diverge.

Find out what AI says about your brand.

Send us your brand and category. We run 20 buyer questions across five engines by hand and send you the transcript within 48 hours.

What arrives
20 questions, 5 engines, every answer quoted in full
Who sends it
Snow Lee, by hand, from Seoul
Your email
Used to send the report and reply about it. Privacy.
Free report · your named ratesample
Buyer questionEngineNamed
best Korean sunscreen for oily skinChatGPTnot named
K-beauty sunscreen that does not pillClaudenamed
Korean skincare brands sold at UltaAI Overviewsnot named
…17 more questions, with the verbatim answer and every source cited
Three rows from what the free report looks like. Yours is built from your own category and market.