How do you measure AEO?
You measure AEO by running a fixed set of queries across a fixed set of answer engines on a schedule, in a logged-out session, and recording what each one returns. There is no dashboard that reports this for you, because the engines do not publish it. The measurement is a protocol, not a tool, and its value comes from being repeated unchanged.
Answer engines do not report impressions, positions or citations the way a search console does. What can be observed is the answer itself: the text a person actually receives. So the unit of measurement is one observation, defined as one query, on one engine, in one locale, on one date. Everything else is built from those.
The five numbers worth separating
Most reporting on this collapses into a single "visibility" figure. That hides the part that matters, because these are different events with different costs to the engine.
| Measure | What it means | Why it is separate |
|---|---|---|
| Observations made | How many of the planned runs produced an answer | Coverage is never total. Engines block, gate and stall, and the shortfall is information |
| Brand named | The answer wrote the name | Costs the engine nothing and commits it to nothing |
| Any link cited | The answer attributed something to a source | Shows the engine was willing to attribute at all for that query |
| Our link cited | The attributed source was ours | The only measure that reflects a decision about your pages |
| Entity resolved | The engine treated the name as a known thing | Distinguishes being known from having your string repeated back |
An engine that names you without citing you has not chosen your page. An engine that cites a link but not yours has chosen someone else's. Reporting those as one number turns a loss into a gain on paper.
What does not become a row
This is the rule that decides whether a log is worth anything. An engine that requires an account, returns a block page, shows a CAPTCHA or simply stops responding has produced no observation. It must not be recorded as "no name returned", because that reads later as a result when it was an absence of access.
The citation log published here keeps those absences visible instead of silently dropping them. Every run states how many of the planned observations were made, so a thin week looks thin rather than looking like bad performance.
Fixed queries, and why they must not drift
The queries are chosen once and then left alone. Changing them between runs makes the series incomparable, and the change usually happens by accident: someone rephrases a prompt, the numbers move, and the movement gets read as a result. A query list is a commitment, and editing it should be a decision with a date attached, not a typing error.
Good queries are the ones a real person would type, phrased neutrally, and they do not contain the brand name. A query that already names the answer will return the answer, and measures nothing.
Cadence
One run per week, on the same day. A single reading is an anecdote: these systems are non-deterministic, and the same prompt can return different sources minutes apart, in different countries, and on different accounts. Three runs begin to show direction. Six answer whether engines are converging on the same sources or diverging.
What this method cannot tell you
It cannot tell you why an engine chose a source, because none of them explain it. It cannot promise that a change you make will produce a citation. And it cannot be generalised from one locale to another: the same query, on the same day, returns different names in different countries, which is documented with screenshots on the main claim page and is the reason locale is recorded on every row.
What it does give you is an honest series: what was asked, what came back, what was missing and when. That is a low bar, and almost nobody in this field clears it. The reason the title itself is contested is that most of the claims around it were announced rather than measured.