How Broadcastwell measures AI visibility
This is the documented method behind every client figure we report. It states how the questions are built, which engines are asked, how many times, how an answer is scored, and how much uncertainty sits on the result. It is versioned and dated, and a change to the method is published here before it is used. Version 1.0, 25 August 2026.
Scope
What is in scope: the AI Visibility Diagnostic, and the ongoing measurement carried out inside the monthly program. Both use the rules on this page.
What is out of scope: the published research volumes. Each of those states its own method in full, on its own page, and each held one engine constant across its runs so that differences between questions were not confused with differences between engines. Client measurement queries four engines and answers a different question, so the two methods are stated separately and their numbers are never combined.
The free mini-audit is a third instrument with its own limits, described in section 04.
What is measured
No composite visibility score is produced, and that is a deliberate choice rather than an omission. Mentions, citations and position move for different reasons and at different speeds. Adding them together produces one number that goes up, and hides which of the three actually moved, which is the only part you can act on. Visibility figures are also reported separately from referral traffic, self-reported discovery and pipeline correlation, for the same reason.
What is not measured: sentiment, or what an answer says about a brand as distinct from whether it names it. The measurement records presence, not tone. If an engine names a brand while describing it poorly, that is scored as a mention, and the answer text is delivered in full so the description can be read directly.
The question set
The questions are the ones a buyer types when they are choosing between vendors and do not yet know who the vendors are. They are written for the category, agreed with the client before the run, and fixed in wording.
Three classes of question exist. They are defined here because the difference between them decides whether a figure means anything, and because they are reported separately and never blended into one headline number.
A question that already contains a name hands the answer that name, so an answer repeating it measures the question rather than the engine's choice. Our own published research measured this effect directly and found that name-seeded questions materially inflated raw mention counts. That is why a Diagnostic runs only the first class, and why, when the other two are used, their results are reported as separate figures with the class stated.
The question set is frozen between runs. The same wording is sent every time, because a reworded question is a different question and the two results are not comparable. If a question set changes, that opens a new baseline. It does not continue the old one, and the comparison to earlier runs is dropped rather than presented with a note.
Engines
The four are ChatGPT, Claude, Perplexity and Google AI Overviews. Every question is sent to all four, in the same words, in the same run.
The free mini-audit is a different instrument. It queries one engine, Perplexity, returns a position on the absence ladder, and is not a Diagnostic. A mini-audit result and a Diagnostic result are not comparable, and a mini-audit is never used as a baseline.
Model version strings, the automation that runs the queries, and any third-party service involved in collection are implementation rather than method. They are not published. What is being bought is the engine products named above, the rules on this page, and the answers underneath the figures.
Repetition
Engines are not deterministic. Ask one the same question an hour apart and the vendor list can change without anything in the world having changed. A method that asks once and reports the result cannot tell the difference between a real position and a coin landing one way.
The current configuration is ten buyer questions, four engines, and three documented repeat runs of every question on every engine. That is 120 observed answers per Diagnostic. Every observation is retained, including the ones that disagree with the others, and the run-to-run spread is reported alongside the rate rather than averaged away.
Ten questions, four engines, three repeats. 120 observed answers per Diagnostic. This is the configuration as of version 1.0. Any change to it is published in the change log in section 09 before it is used.
Scoring rules
Uncertainty
We do not observe every answer an engine could ever give. We observe a sample of them, under a documented protocol, and estimate a rate from that sample. A different sample of the same size would land near the same figure but not exactly on it, and how near is something that can be calculated rather than guessed.
Every headline figure is therefore published with its sample size and a 95 percent confidence interval, calculated using the Wilson score interval for a binomial proportion. Wilson is used rather than the textbook normal approximation because it stays sensible at small samples and at rates near zero, which is exactly where a first baseline usually sits.
The practical consequence is a rule we hold ourselves to: a change between two runs that sits inside both intervals is not reported as movement. It is reported as a figure that did not separate from the previous one, with both intervals shown, and the run-to-run spread published beside it.
This is also why the full four-engine re-score is quarterly rather than monthly. A monthly headline re-score on a sample this size would spend most of its life reporting movement that cannot be distinguished from noise. Leading indicators that do move on a monthly timescale, such as mentions, citations and which sources the engines are pulling from, are reported every month instead.
What is published and what is held
The line between the two is method against implementation. The method is on this page in enough detail that a technically literate reader could reproduce its shape and check the arithmetic. The implementation is how we happen to run it this quarter.
Version and change log
| Version | Date | What changed |
|---|---|---|
| v1.0 | 25 August 2026 | First published. |
A change to the method opens a new baseline. Results produced under one version are not compared against results produced under another, because the comparison would measure the method change rather than the client's position. The version in force when a figure was produced is stated on the report that carries it.
Measure your category under this method
Ten buyer questions for your category, one named engine, your position on the absence ladder. Free, under a minute, no call required.
Run the free mini-audit