Corpus
When an assistant answers “best Korean sunscreen for oily skin”, it is not consulting a ranked index of pages. It is drawing on the material it can retrieve and has been trained on for that category, and producing a summary of what those sources collectively say.
Seven surfaces one answer is assembled fromillustrative
- Your siteyours to write
- Amazon listingyours to correct
- Retailer pagesyours to correct
- Review round-upssomeone else's
- Reddit threadssomeone else's
- YouTube reviewssomeone else's
- Korean reviewswritten, wrong language
How many you write
That material is the corpus. For a consumer category it can include:
- The brand’s own product and specification pages
- Marketplace listings, which may have been written by a distributor rather than the brand
- Retailer category and product pages
- Review publications and comparison round-ups
- Forum threads, Reddit prominent among them in English-language markets
- Video transcripts from review and comparison content
Two properties of the corpus are what we look at first when a brand is not being named. Presence — does the brand appear in it at all, in the language of the market. Agreement — do several independent sources state the same specification, so a model can assert it without relying on the brand’s own word.
Both are where a cross-border brand can come up short. A Korean brand can have a rich Korean-language corpus and a near-empty English one, and the assistant answering an American shopper reads only the second.
Related: entity resolution, citability, named rate.