Skip to content

How Broadcastwell measures AI visibility

This is the documented method behind every client figure we report. It states how the questions are built, which engines are asked, how many times, how an answer is scored, and how much uncertainty sits on the result. It is versioned and dated, and a change to the method is published here before it is used. Every Broadcastwell Diagnostic asks 35 buyer questions of five AI engines, three times each, and prints the confidence interval on every number. Version 1.1, 8 September 2026.

This document covers client measurement. It does not cover the published research.

What is in scope: the AI Visibility Diagnostic, and the ongoing measurement carried out inside the monthly program. Both use the rules on this page.

What is out of scope: the published research volumes. Each states its own method, dates and limitations. Volumes I and II use one engine; Volume III compares four engines. Research and client measurements are reported separately, and their numbers are never combined.

The free 10-question check is a third instrument with its own limits, described in section 04.

Instrument scope. The measurement is single turn, API based, logged out, US English, gl=us and hl=en. It is a controlled instrument, not a replica of any one buyer's screen.

Four outcomes are recorded for every observed answer. They are reported separately and are never combined into a single score.
Mention. The brand is named in the answer text. An answer counts once for a brand it names, however many times it names it.
Citation. The brand's own domain appears among the sources the answer cited. A citation and a mention are separate outcomes. Supplying the material an engine reads and being named in the answer it writes are different results, and an answer can do either without the other.
Recommendation position. Where the brand falls when an answer returns an ordered or grouped shortlist. Recorded when the answer has a discernible order. Otherwise the report says: no discernible order in this answer. Position is never inferred from the order in which names happen to appear in a paragraph.
Competitor named. A rival from the list agreed before the run appears in the answer text. Scored by the same rule as a mention.

No composite visibility score is produced, and that is a deliberate choice rather than an omission. Mentions, citations and position move for different reasons and at different speeds. Adding them together produces one number that goes up, and hides which of the three actually moved, which is the only part you can act on. Visibility figures are also reported separately from referral traffic, self-reported discovery and pipeline correlation, for the same reason.

What is not measured: sentiment, or what an answer says about a brand as distinct from whether it names it. The measurement records presence, not tone. If an engine names a brand while describing it poorly, that is scored as a mention, and the answer text is delivered in full so the description can be read directly.

A Diagnostic asks 35 questions, written for the category in real buyer language. No question names any brand: not the client's, not a competitor's. The 35 sit in four groups: 10 shortlist questions, 10 questions phrased for different buyer roles, 10 use case and problem questions, and 5 evaluation and pricing questions. The 10 shortlist questions are the ones tracked between full re-scores. Findings and re-measurement report presence for each of the four groups separately, named in findings as shortlist, buyer role, use case and evaluation, so a rise in citations on use case questions is never presented as presence on the shortlist. The numbers here do not prove causation, do not predict traffic, and describe the observation window only.

The questions are the ones a buyer types when they are choosing between vendors and do not yet know who the vendors are. They are written for the category, agreed with the client before the run, and fixed in wording.

Every question in the set is a generic buyer intent question, containing no brand name at all. These are the discovery questions and they are the whole measure. When an engine names a brand in answer to one of these, it chose that brand.

A question that already contains a name hands the answer that name, so an answer repeating it measures the question rather than the engine's choice. Our own published research measured this effect directly and found that name-seeded questions materially inflated raw mention counts. That is why no name appears in any question we ask. Where a client asks for a branded or unbranded comparison, those figures are recorded and reported as their own set. Branded and unbranded figures are never blended into one headline.

The question set is frozen between runs. The same wording is sent every time, because a reworded question is a different question and the two results are not comparable. If a question set changes, that opens a new baseline. It does not continue the old one, and the comparison to earlier runs is dropped rather than presented with a note.

Client measurement queries five engine products, each asked the identical wording.

The five engines are ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Every question is sent to all five, in the same words, in the same run. Google AI Overviews and Google AI Mode are separate surfaces and are recorded separately, because they cite different sources from each other most of the time.

The free 10-question check is a different instrument. It queries one engine, Perplexity, returns a position on the absence ladder, and is not a Diagnostic. A Free 10-question check result and a Diagnostic result are not comparable, and a Free 10-question check is never used as a baseline.

The automation that runs the queries and any third-party service involved in collection are implementation rather than method, and are not published. Model version strings appear in the published research. Client reports name the engine, not the model string. What is being bought is the engine products named above, the rules on this page, and the answers underneath the figures.

A single run is not evidence. The same question asked twice can return different answers.

Engines are not deterministic. Ask one the same question an hour apart and the vendor list can change without anything in the world having changed. A method that asks once and reports the result cannot tell the difference between a real position and a coin landing one way.

The current configuration is 35 buyer questions, five engines, and three scheduled runs of every question on every engine. That is 525 scheduled observed answers per Diagnostic. This is the scheduled count, not an upper limit. Scheduled and adaptive observations are reported separately, and the headline comparison uses only the scheduled panel. Where the three verdicts for a question and engine pair disagree, we run that pair up to twice more, so the extra observation goes where the disagreement is rather than everywhere at once. Runs beyond three add almost nothing where the three agree. Every observation is retained, including the ones that disagree, and the run-to-run spread is reported alongside the rate rather than averaged away.

35 questions, five engines, three scheduled runs, plus up to two further runs on any disagreeing pair. 525 scheduled observed answers per Diagnostic. This is the configuration as of version 1.1. Any change to it is published in the change log in section 09 before it is used.

The measurement grid: 35 questions, five engines, three scheduled runs

A table of four question groups against the five engines: ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Each group states its question count, the engines, the three scheduled runs for every question and engine pair, and the scheduled answers that result, adding to 525.

The measurement grid: 35 questions, five engines, three scheduled runs
Question groupQuestionsEnginesScheduled runsScheduled answers
Shortlist1053150
Buyer role1053150
Use case and problem1053150
Evaluation and pricing55375
Total3553525

35 questions x 5 engines x 3 scheduled runs = 525 scheduled observed answers, plus up to two further runs on any question and engine pair whose three verdicts disagreed.

In words: 35 buyer questions are asked on five engines, ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode, and every question and engine pair is run three times on schedule. Thirty-five questions multiplied by five engines multiplied by three scheduled runs is 525 scheduled observed answers. Any pair whose three verdicts disagreed is run up to twice more.

The rules below decide what counts. They are applied the same way to a win and to a loss.
Brand matching is on word boundaries. A brand counts as named when the answer contains the brand as a whole word, case-sensitive for a single plain word. A longer word that happens to contain the brand inside it does not count.
Domain matching compares hostnames. A citation counts when the cited source resolves to the brand's hostname. Subdomains of that hostname count. Near-miss domains that merely resemble it do not.
Failed responses leave the denominator. When an engine returns nothing, errors, refuses, or, in the case of ChatGPT, answers without searching the live web, that observation is excluded from both the named count and the denominator. Missing answers and engine errors are listed with the denominator for each engine and question group. It is never scored as a miss. This cuts both ways and that is the point: counting an outage as an absence would understate a client's position, and quietly dropping only the unfavourable failures would overstate it. Every exclusion is listed in full, with its reason, in the delivered data.
A question-and-engine cell has one verdict. For the proof gate, a cell has one named verdict: unanimous across its three scheduled runs, or the majority after the adaptive reruns. Each cell counts once. Extra adaptive answers do not enlarge the frozen panel. Cells lost to engine errors are excluded from both baseline and re-score and reported with the common denominator. The ten shortlist questions form 50 cells; the current v1.1 contractual gate covers all 35 agreed questions and 175 cells, as specified in the terms. Presence is measured, tone is not.
Floors are labelled as floors. Where a figure is a minimum rather than an exact count, because the underlying source is incomplete or an engine truncated its own source list, the figure is labelled a minimum. It is not presented as an exact number with a caveat somewhere else on the page.
Every rate we report is a sample, not a census. A rate published without a range is not evidence, and this section is the reason this page exists.

We do not observe every answer an engine could ever give. We observe a sample of them, under a documented protocol, and estimate a rate from that sample. A different sample of the same size would land near the same figure but not exactly on it, and how near is something that can be calculated rather than guessed.

Every headline figure is therefore published with its sample size and a 95 percent confidence interval, calculated using the Wilson score interval for a binomial proportion. Wilson is used rather than the textbook normal approximation because it stays sensible at small samples and at rates near zero, which is exactly where a first baseline usually sits.

A worked illustration. Suppose a brand is named in 79 of 525 scheduled observed answers. That is a rate of 15.0 percent, and its 95 percent confidence interval runs from 12.2 percent to 18.4 percent. Now suppose the next quarterly re-score names it in 105 of 525, a rate of 20.0 percent, whose interval runs from 16.8 percent to 23.6 percent. The intervals overlap from 16.8 to 18.4 percent. Interval overlap does not by itself decide whether a change is real. These figures illustrate the arithmetic. They are not a Broadcastwell result and not a client result.

The comparison is between two proportions on the same fixed panel, a repeated-measures design: questions repeat across engines and runs, so 525 answers are not 525 independent draws. The figures estimate change within that fixed panel, reported per engine and question group, not across every possible buyer question. Wilson intervals describe the observed proportions; they do not by themselves establish a significant change between baselines. A significance claim requires a documented paired comparison that accounts for this dependence. No such test is specified by this worked illustration, so it makes no significance claim.

This is also why the full re-score is quarterly rather than monthly. A monthly headline re-score would spend most of its life reporting movement that cannot be distinguished from noise. Brand mentions, citations and the sources engines use are reported every month, and the 10 shortlist questions carry monthly leading indicators. The controlled five-engine re-score at the full configuration remains quarterly.

The client receives everything underneath their own figures. Their figures stay theirs.
Delivered to the client in full. Every observed answer, as the engine returned it, with the engine that returned it, the date and time it ran, and every source it cited. Wins and losses alike, including every observation excluded from a denominator and the reason it was excluded. No figure we report is one a client cannot trace back to the answers underneath it.
Never published without written permission. Client baselines, findings, question sets and outcomes. A client story appears on this site only with explicit written approval. Reviews are optional, candid and never tied to a discount, payment or preferential treatment.
Not published at all. The operational implementation: the automation that runs the queries, model version strings, endpoints, third-party services, and what any of it costs to run. None of that changes whether a figure is true.

The line between the two is method against implementation. The method is on this page in enough detail that a technically literate reader could reproduce its shape and check the arithmetic. The implementation is how we happen to run it this quarter.

This method is versioned. Changes are published here.
Data table: Version, Date, What changed
VersionDateWhat changed
v1.1 erratum11 September 2026The worked Wilson intervals overlap from 16.8 to 18.4 percent. The former claim of non-overlap and the inference of real movement were incorrect. This correction clarifies dependence, denominators and scheduled versus adaptive observations; the method version remains v1.1.
v1.18 September 20268 September 2026. Method v1.1. Questions per client 10 to 35, in four groups. Engines four to five, adding Google AI Mode. Repeat runs three, plus up to two further runs on any question and engine pair whose verdicts disagreed. Confidence intervals unchanged. Exclusion rule extended to ChatGPT answers that did not search the live web. Why: 35 questions cover four buyer-question groups within a fixed repeated-measures panel; runs beyond three add almost nothing except where the engines disagree, so that is where the extra runs go; Google AI Mode is a distinct surface that cites different sources from AI Overviews most of the time. Engagements that began under v1.0 stay on v1.0 and are never compared across versions.
v1.025 August 2026First published.

A change to the method opens a new baseline. Results produced under one version are not compared against results produced under another, because the comparison would measure the method change rather than the client's position. The version in force when a figure was produced is stated on the report that carries it. Comparability. Figures are compared only within one method version and one question set. Engagements that began under v1.0 are re-scored under v1.0. Nothing is compared across versions.

Start with a measured baseline

See which buyer questions leave you out, the sources visible in those answers, and the three changes worth testing first. The Category Audit covers ten questions; the Diagnostic extends to 35 questions and 525 scheduled observed answers, adaptive runs extra. AI answers vary, so repeated observations and disagreements are reported.

Start with the $490 Category Audit

Results appear in your account; a sign-in link arrives by email. Credits in full against the Diagnostic within 30 days.

Get the Diagnostic, $990

Results appear in your account; a sign-in link arrives by email. Credits in full against the program.

Every figure comes from the published method, v1.1.