AEO fundamentals
How AEO agencies prove results, and what a monthly report has to show

Every AEO pitch arrives at the same question, and most of them arrive at it badly. You are being asked to pay for a number that the vendor also produces. If the same party sets the questions, runs them, and reports the score, the report is not a measurement — it is a marketing asset with a table in it.
So the useful thing to establish before you sign is not whether an agency is confident. It is which of its claims you could check yourself, today, without its cooperation.
How do AEO agencies prove their results?
Three ways, and they are not equally checkable. SearchD sorts vendor evidence by a single test: how much of the claim can you verify before you pay for it?
| Evidence | What it looks like | Checkable before you sign? |
|---|---|---|
| Client case study | “Brand X went from 4% to 31% named rate in five months” | Only by asking the client. Get the introduction, then ask which panel was used and whether it was frozen beforehand |
| Third-party tool score | A dashboard screenshot with a visibility percentage | Partly. You can buy the same dashboard. It reports the score, not what caused it |
| The vendor’s own published number | A dated figure on a question set fixed in advance | Yes. Open the page and read it, today |
The first row is the strongest of the three when the introduction comes with it, and SearchD tells buyers to ask for it first — of every vendor, itself included. What decides whether it is evidence is the baseline behind it. A named rate means nothing except against its own before-state, and that is the one thing nobody can reconstruct later: once a change has shipped, the before is gone. So the two questions to put to a reference are which panel was used and whether it was frozen before the work began. A case study that survives both is worth more than anything else on this page. One with no client you can reach is a claim wearing a number.
That same logic is why the third row has to be published before the selling starts rather than after a good month.
Google’s own guidance is the reason to be strict here. It states plainly that no third-party tool has insight into its ranking or AI systems (Google, AI features and your website). Any vendor promising you a placement inside an answer is promising a mechanism nobody outside the engine can see.
What should a monthly AI visibility report show?
Six sections, on one page, and a report missing any of them is hiding something structural rather than saving you reading time. This is the shape SearchD publishes, including the sections that make it look worse.
| # | Section | What it contains |
|---|---|---|
| 1 | Named rate by engine | Questions where the engine named you, over questions asked — per engine, this month against last |
| 2 | Deltas against a noise floor | Every change marked inside or outside the floor, applied symmetrically in both directions |
| 3 | Quotes, and who was named instead | Verbatim answer text and the competitor names that appeared in your place |
| 4 | Factual errors found | Wrong specification, price or availability, with the URL the engine took it from |
| 5 | What shipped | The changes made that month, dated, so cause and effect are at least alignable |
| 6 | Assistant referral traffic | Sessions arriving from ChatGPT, Perplexity and the rest, in absolute numbers |
Section 1 is per engine for a reason. The engines do not read the same index — Claude’s web search runs on Brave’s, Google’s surfaces read Google’s, and ChatGPT serves from pipelines including one of its own, identified by instrumenting 1,200 answers where only 1.5% of the URLs appeared in Bing’s top 20 for the same queries (Search Engine Land, 17 August 2026). A single blended percentage averages four separate facts into one that describes none of them.
Section 5 is where most reports go quiet. If the report does not say what was shipped and when, no delta in it can be attributed to anything, and the agency has reserved the right to claim every rise. Model updates move these numbers too, in both directions, and nobody can fully separate the two — SearchD records the model version on every observation and re-runs within a week of a major release, and where a jump is plainly a model update the report says so.
Why does an AI visibility report need a noise floor?
Because the same question asked twice does not reliably return the same answer, so a small movement is not evidence of anything until you know how big the instrument’s own wobble is. SearchD measures that floor from a stability run — the same panel, repeated with nothing changed — and reports every delta against it.
The discipline that matters is symmetry. A +2 and a −1 inside the same ±3 band are both noise, and an agency that marks the +2 as progress while calling the −1 noise has told you the floor is decorative. SearchD reports named rate per engine at n=40, which is the basis the floor is derived from; the arithmetic behind reading such a number is in what a good named rate is.
Is there an agency that publishes its own AI visibility numbers before selling AEO?
SearchD does, and the number is currently zero. In the pilot reading of 8 September 2026 — 17 buyer questions asked on ChatGPT, Claude and Gemini, 51 observations — no engine named SearchD on any of them. The figure is on /results and goes up monthly, including the months it goes the wrong way.
A new domain with no corpus should be at zero, so a launch number that is anything else is the thing worth interrogating. The monthly panel SearchD reports against is 40 questions across five engines — ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews — frozen in September 2026 and unchanged from here.
The same standard is what SearchD asks buyers to hold every vendor to, itself included. The full list of questions to put to an AEO vendor, with SearchD’s own answers filled in including the three it loses, is on the evidence page.
Which proof claims should make you walk away?
Five, and each one fails against a source you can read rather than against SearchD’s opinion. These are claims SearchD checked before deciding what it would sell, and declined to sell.
| The claim | What the evidence says |
|---|---|
| “Schema markup will get you cited” | After matched controls across 1,885 pages: +2.4% on AI Mode, +2.2% on ChatGPT, −4.6% on AI Overviews (Ahrefs, 11 May 2026) |
| “An llms.txt file, so the models can read you” | Google states Search ignores these files, and that keeping one neither harms nor helps (Google) |
| “A guaranteed citation, or a top-five placement” | Nobody outside the engine can see the mechanism being guaranteed, and Google says so in its own documentation |
| “Your buyers ask ChatGPT this question N times a month” | No engine publishes prompt frequency. Every figure on the market is reconstructed from an opt-in panel or extrapolated from Google volume |
| “We rank you for these keywords, so you’ll be in the answers” | 95% of the queries an engine generates for itself have no monthly search volume in conventional keyword tools (AirOps, March 2026) |
The last row is the one that most often survives a procurement review unchallenged, because it sounds like the SEO deliverable everyone already knows how to buy. It fails because the engine does not search the sentence your buyer typed — it generates several searches of its own and answers from what those return. What ChatGPT actually searches before it names a brand covers that mechanism and the order of work that follows from it.
How do you check an agency’s proof in ten minutes?
Four steps, none of which require the agency’s cooperation, and all of which you can run before the second call.
- Ask for a client you can call, and for the agency’s own number. Both, not either. If there is no reference, ask what the agency’s own named rate was last month and on what date — an agency selling measurement that does not measure itself has answered the question.
- Ask who holds the question list. If the panel is proprietary, the denominator is hidden and the number cannot be audited. You should receive the questions, and they should be frozen before any work begins.
- Ask what the noise floor is and how it was measured. Then look at last month’s report and check whether a negative delta inside that band was also called noise.
- Ask what happens if it does not work. The answer should name something the agency controls and point at a clause in the agreement, not at a page on its website.
A vendor that answers all four cleanly may still not move your number — engine behaviour is nobody’s to promise. But you will at least be buying a measurement rather than a narrative, and you will know on which month to stop.
What that measurement is worth on your own brand depends on whether there is a quotable passage for an engine to lift once it can reach you, which is a separate piece of work: what makes a passage get cited. The monthly loop SearchD runs against the panel above is described on the retainer page, and where a monitoring tool is the better purchase is set out on alternatives.
Related questions
Is a client case study not proof?
A case study with a client who will take your call is the strongest evidence an AEO agency can offer, and you should ask every vendor for one — SearchD included. What it needs to survive is two questions put to that client: which question panel was used, and was it frozen before the work started. A case study nobody will put you in touch with is a claim, not a reference. Ask SearchD the same question; the answer today is its own published panel number rather than a client reference.
Can I not just buy a tracking tool and check myself?
You can, and for some buyers that is the right purchase. A tool reports the score; it does not tell you which of the causes behind the score is yours, and it will not correct a wrong specification on a retailer's listing. SearchD's own position is on its alternatives page: if you have a team to act on a score, buy the tool.
What if an agency says its panel is proprietary?
Then the number it reports to you cannot be audited, because the denominator is hidden. A panel is only a measurement instrument if you hold the question list. SearchD sets each panel from a week of live readings, agrees it with the client in a session they sit in, hands over the list, and leaves it unchanged for twelve months — adding questions later moves the number without moving the business.
How soon should a monthly report show anything?
Access and entity fixes show up in the next panel run, because they change what an engine can fetch and which brand it resolves you to. Corrections on a retailer's listing run at that retailer's pace, usually weeks. Corpus work is months, because it depends on other people publishing. An agency quoting you a date for a citation is quoting a date for someone else's behaviour.
Does SearchD guarantee a result?
SearchD guarantees the things it controls, not engine behaviour. The audit invoice is cancelled if it surfaces no cause you could not reasonably have found yourself, and the retainer stops billing at day 60 if your named rate has not moved beyond your noise floor. Both clauses are written into the engagement agreement rather than stated on a page.
