How Broadcastwell measures AI visibility
This is the documented method behind every client figure we report. It states how the questions are built, which engines are asked, how many times, how an answer is scored, and how much uncertainty sits on the result. It is versioned and dated, and a change to the method is published here before it is used. Every Broadcastwell Diagnostic asks 35 buyer questions of five AI engines, three times each, and prints the confidence interval on every number. Version 1.1, 8 September 2026.
Scope
What is in scope: the AI Visibility Diagnostic, and the ongoing measurement carried out inside the monthly program. Both use the rules on this page.
What is out of scope: the published research volumes. Each states its own method, dates and limitations. Volumes I and II use one engine; Volume III compares four engines. Research and client measurements are reported separately, and their numbers are never combined.
The free 10-question check is a third instrument with its own limits, described in section 04.
Instrument scope. The measurement is single turn, API based, logged out, US English, gl=us and hl=en. It is a controlled instrument, not a replica of any one buyer's screen.
What is measured
No composite visibility score is produced, and that is a deliberate choice rather than an omission. Mentions, citations and position move for different reasons and at different speeds. Adding them together produces one number that goes up, and hides which of the three actually moved, which is the only part you can act on. Visibility figures are also reported separately from referral traffic, self-reported discovery and pipeline correlation, for the same reason.
What is not measured: sentiment, or what an answer says about a brand as distinct from whether it names it. The measurement records presence, not tone. If an engine names a brand while describing it poorly, that is scored as a mention, and the answer text is delivered in full so the description can be read directly.
The question set
The questions are the ones a buyer types when they are choosing between vendors and do not yet know who the vendors are. They are written for the category, agreed with the client before the run, and fixed in wording.
Every question in the set is a generic buyer intent question, containing no brand name at all. These are the discovery questions and they are the whole measure. When an engine names a brand in answer to one of these, it chose that brand.
A question that already contains a name hands the answer that name, so an answer repeating it measures the question rather than the engine's choice. Our own published research measured this effect directly and found that name-seeded questions materially inflated raw mention counts. That is why no name appears in any question we ask. Where a client asks for a branded or unbranded comparison, those figures are recorded and reported as their own set. Branded and unbranded figures are never blended into one headline.
The question set is frozen between runs. The same wording is sent every time, because a reworded question is a different question and the two results are not comparable. If a question set changes, that opens a new baseline. It does not continue the old one, and the comparison to earlier runs is dropped rather than presented with a note.
Engines
The five engines are ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Every question is sent to all five, in the same words, in the same run. Google AI Overviews and Google AI Mode are separate surfaces and are recorded separately, because they cite different sources from each other most of the time.
The free 10-question check is a different instrument. It queries one engine, Perplexity, returns a position on the absence ladder, and is not a Diagnostic. A Free 10-question check result and a Diagnostic result are not comparable, and a Free 10-question check is never used as a baseline.
The automation that runs the queries and any third-party service involved in collection are implementation rather than method, and are not published. Model version strings appear in the published research. Client reports name the engine, not the model string. What is being bought is the engine products named above, the rules on this page, and the answers underneath the figures.
Repetition
Engines are not deterministic. Ask one the same question an hour apart and the vendor list can change without anything in the world having changed. A method that asks once and reports the result cannot tell the difference between a real position and a coin landing one way.
The current configuration is 35 buyer questions, five engines, and three scheduled runs of every question on every engine. That is 525 scheduled observed answers per Diagnostic. This is the scheduled count, not an upper limit. Scheduled and adaptive observations are reported separately, and the headline comparison uses only the scheduled panel. Where the three verdicts for a question and engine pair disagree, we run that pair up to twice more, so the extra observation goes where the disagreement is rather than everywhere at once. Runs beyond three add almost nothing where the three agree. Every observation is retained, including the ones that disagree, and the run-to-run spread is reported alongside the rate rather than averaged away.
35 questions, five engines, three scheduled runs, plus up to two further runs on any disagreeing pair. 525 scheduled observed answers per Diagnostic. This is the configuration as of version 1.1. Any change to it is published in the change log in section 09 before it is used.
The measurement grid: 35 questions, five engines, three scheduled runs
A table of four question groups against the five engines: ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Each group states its question count, the engines, the three scheduled runs for every question and engine pair, and the scheduled answers that result, adding to 525.
| Question group | Questions | Engines | Scheduled runs | Scheduled answers |
|---|---|---|---|---|
| Shortlist | 10 | 5 | 3 | 150 |
| Buyer role | 10 | 5 | 3 | 150 |
| Use case and problem | 10 | 5 | 3 | 150 |
| Evaluation and pricing | 5 | 5 | 3 | 75 |
| Total | 35 | 5 | 3 | 525 |
35 questions x 5 engines x 3 scheduled runs = 525 scheduled observed answers, plus up to two further runs on any question and engine pair whose three verdicts disagreed.
In words: 35 buyer questions are asked on five engines, ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode, and every question and engine pair is run three times on schedule. Thirty-five questions multiplied by five engines multiplied by three scheduled runs is 525 scheduled observed answers. Any pair whose three verdicts disagreed is run up to twice more.
Scoring rules
Uncertainty
We do not observe every answer an engine could ever give. We observe a sample of them, under a documented protocol, and estimate a rate from that sample. A different sample of the same size would land near the same figure but not exactly on it, and how near is something that can be calculated rather than guessed.
Every headline figure is therefore published with its sample size and a 95 percent confidence interval, calculated using the Wilson score interval for a binomial proportion. Wilson is used rather than the textbook normal approximation because it stays sensible at small samples and at rates near zero, which is exactly where a first baseline usually sits.
The comparison is between two proportions on the same fixed panel, a repeated-measures design: questions repeat across engines and runs, so 525 answers are not 525 independent draws. The figures estimate change within that fixed panel, reported per engine and question group, not across every possible buyer question. Wilson intervals describe the observed proportions; they do not by themselves establish a significant change between baselines. A significance claim requires a documented paired comparison that accounts for this dependence. No such test is specified by this worked illustration, so it makes no significance claim.
This is also why the full re-score is quarterly rather than monthly. A monthly headline re-score would spend most of its life reporting movement that cannot be distinguished from noise. Brand mentions, citations and the sources engines use are reported every month, and the 10 shortlist questions carry monthly leading indicators. The controlled five-engine re-score at the full configuration remains quarterly.
What is published and what is held
The line between the two is method against implementation. The method is on this page in enough detail that a technically literate reader could reproduce its shape and check the arithmetic. The implementation is how we happen to run it this quarter.
Version and change log
| Version | Date | What changed |
|---|---|---|
| v1.1 erratum | 11 September 2026 | The worked Wilson intervals overlap from 16.8 to 18.4 percent. The former claim of non-overlap and the inference of real movement were incorrect. This correction clarifies dependence, denominators and scheduled versus adaptive observations; the method version remains v1.1. |
| v1.1 | 8 September 2026 | 8 September 2026. Method v1.1. Questions per client 10 to 35, in four groups. Engines four to five, adding Google AI Mode. Repeat runs three, plus up to two further runs on any question and engine pair whose verdicts disagreed. Confidence intervals unchanged. Exclusion rule extended to ChatGPT answers that did not search the live web. Why: 35 questions cover four buyer-question groups within a fixed repeated-measures panel; runs beyond three add almost nothing except where the engines disagree, so that is where the extra runs go; Google AI Mode is a distinct surface that cites different sources from AI Overviews most of the time. Engagements that began under v1.0 stay on v1.0 and are never compared across versions. |
| v1.0 | 25 August 2026 | First published. |
A change to the method opens a new baseline. Results produced under one version are not compared against results produced under another, because the comparison would measure the method change rather than the client's position. The version in force when a figure was produced is stated on the report that carries it. Comparability. Figures are compared only within one method version and one question set. Engagements that began under v1.0 are re-scored under v1.0. Nothing is compared across versions.
Start with a measured baseline
See which buyer questions leave you out, the sources visible in those answers, and the three changes worth testing first. The Category Audit covers ten questions; the Diagnostic extends to 35 questions and 525 scheduled observed answers, adaptive runs extra. AI answers vary, so repeated observations and disagreements are reported.
Start with the $490 Category AuditResults appear in your account; a sign-in link arrives by email. Credits in full against the Diagnostic within 30 days.
Get the Diagnostic, $990Results appear in your account; a sign-in link arrives by email. Credits in full against the program.