How Broadcastwell measures AI visibility

This is the documented method behind every client figure we report. It states how the questions are built, which engines are asked, how many times, how an answer is scored, and how much uncertainty sits on the result. It is versioned and dated, and a change to the method is published here before it is used. Version 1.0, 25 August 2026.

01

Scope

This document covers client measurement. It does not cover the published research.

What is in scope: the AI Visibility Diagnostic, and the ongoing measurement carried out inside the monthly program. Both use the rules on this page.

What is out of scope: the published research volumes. Each of those states its own method in full, on its own page, and each held one engine constant across its runs so that differences between questions were not confused with differences between engines. Client measurement queries four engines and answers a different question, so the two methods are stated separately and their numbers are never combined.

The free mini-audit is a third instrument with its own limits, described in section 04.

02

What is measured

Four outcomes are recorded for every observed answer. They are reported separately and are never combined into a single score.
Mention. The brand is named in the answer text. An answer counts once for a brand it names, however many times it names it.
Citation. The brand's own domain appears among the sources the answer cited. A citation and a mention are separate outcomes. Supplying the material an engine reads and being named in the answer it writes are different results, and an answer can do either without the other.
Recommendation position. Where the brand falls when an answer returns an ordered or grouped shortlist. Recorded when the answer has a discernible order, and left blank when it does not, rather than inferred from the order in which names happen to appear in a paragraph.
Competitor named. A rival from the list agreed before the run appears in the answer text. Scored by the same rule as a mention.

No composite visibility score is produced, and that is a deliberate choice rather than an omission. Mentions, citations and position move for different reasons and at different speeds. Adding them together produces one number that goes up, and hides which of the three actually moved, which is the only part you can act on. Visibility figures are also reported separately from referral traffic, self-reported discovery and pipeline correlation, for the same reason.

What is not measured: sentiment, or what an answer says about a brand as distinct from whether it names it. The measurement records presence, not tone. If an engine names a brand while describing it poorly, that is scored as a mention, and the answer text is delivered in full so the description can be read directly.

03

The question set

A Diagnostic asks ten questions, written for the category, and none of them contains a brand name.

The questions are the ones a buyer types when they are choosing between vendors and do not yet know who the vendors are. They are written for the category, agreed with the client before the run, and fixed in wording.

Three classes of question exist. They are defined here because the difference between them decides whether a figure means anything, and because they are reported separately and never blended into one headline number.

Generic buyer intent. Questions containing no brand name at all. These are the discovery questions and the primary measure. Every question in a Diagnostic is of this class. When an engine names a brand in answer to one of these, it chose that brand.
Competitor-named. Questions that name a rival, in the shape of alternatives to a named vendor, or one vendor against another.
Own-brand-named. Questions that name the client.

A question that already contains a name hands the answer that name, so an answer repeating it measures the question rather than the engine's choice. Our own published research measured this effect directly and found that name-seeded questions materially inflated raw mention counts. That is why a Diagnostic runs only the first class, and why, when the other two are used, their results are reported as separate figures with the class stated.

The question set is frozen between runs. The same wording is sent every time, because a reworded question is a different question and the two results are not comparable. If a question set changes, that opens a new baseline. It does not continue the old one, and the comparison to earlier runs is dropped rather than presented with a note.

04

Engines

Client measurement queries four engine products, each asked the identical wording.

The four are ChatGPT, Claude, Perplexity and Google AI Overviews. Every question is sent to all four, in the same words, in the same run.

The free mini-audit is a different instrument. It queries one engine, Perplexity, returns a position on the absence ladder, and is not a Diagnostic. A mini-audit result and a Diagnostic result are not comparable, and a mini-audit is never used as a baseline.

Model version strings, the automation that runs the queries, and any third-party service involved in collection are implementation rather than method. They are not published. What is being bought is the engine products named above, the rules on this page, and the answers underneath the figures.

05

Repetition

A single run is not evidence. The same question asked twice can return different answers.

Engines are not deterministic. Ask one the same question an hour apart and the vendor list can change without anything in the world having changed. A method that asks once and reports the result cannot tell the difference between a real position and a coin landing one way.

The current configuration is ten buyer questions, four engines, and three documented repeat runs of every question on every engine. That is 120 observed answers per Diagnostic. Every observation is retained, including the ones that disagree with the others, and the run-to-run spread is reported alongside the rate rather than averaged away.

Ten questions, four engines, three repeats. 120 observed answers per Diagnostic. This is the configuration as of version 1.0. Any change to it is published in the change log in section 09 before it is used.

06

Scoring rules

The rules below decide what counts. They are applied the same way to a win and to a loss.
Brand matching is on word boundaries. A brand counts as named when the answer contains the brand as a whole word, case-sensitive for a single plain word. A longer word that happens to contain the brand inside it does not count.
Domain matching compares hostnames. A citation counts when the cited source resolves to the brand's hostname. Subdomains of that hostname count. Near-miss domains that merely resemble it do not.
Failed responses leave the denominator. When an engine returns nothing, errors, or refuses, that observation is excluded from the denominator. It is never scored as a miss. This cuts both ways and that is the point: counting an outage as an absence would understate a client's position, and quietly dropping only the unfavourable failures would overstate it. Every exclusion is listed with its reason in the delivered data.
Floors are labelled as floors. Where a figure is a minimum rather than an exact count, because the underlying source is incomplete or an engine truncated its own source list, the figure is labelled a minimum. It is not presented as an exact number with a caveat somewhere else on the page.
07

Uncertainty

Every rate we report is a sample, not a census. A rate published without a range is not evidence, and this section is the reason this page exists.

We do not observe every answer an engine could ever give. We observe a sample of them, under a documented protocol, and estimate a rate from that sample. A different sample of the same size would land near the same figure but not exactly on it, and how near is something that can be calculated rather than guessed.

Every headline figure is therefore published with its sample size and a 95 percent confidence interval, calculated using the Wilson score interval for a binomial proportion. Wilson is used rather than the textbook normal approximation because it stays sensible at small samples and at rates near zero, which is exactly where a first baseline usually sits.

A worked illustration. Suppose a brand is named in 18 of 120 observed answers. That is a rate of 15.0 percent, and its 95 percent confidence interval runs from 9.7 percent to 22.5 percent. Now suppose the next quarterly re-score names it in 24 of 120, a rate of 20.0 percent, whose interval runs from 13.8 percent to 28.0 percent. The headline looks like a jump from 15 percent to 20 percent. It is not one. Each rate sits inside the other's interval, so the honest reading is that the position may not have changed at all. These figures illustrate the arithmetic. They are not a Broadcastwell result and not a client result.

The practical consequence is a rule we hold ourselves to: a change between two runs that sits inside both intervals is not reported as movement. It is reported as a figure that did not separate from the previous one, with both intervals shown, and the run-to-run spread published beside it.

This is also why the full four-engine re-score is quarterly rather than monthly. A monthly headline re-score on a sample this size would spend most of its life reporting movement that cannot be distinguished from noise. Leading indicators that do move on a monthly timescale, such as mentions, citations and which sources the engines are pulling from, are reported every month instead.

08

What is published and what is held

The client receives everything underneath their own figures. Their figures stay theirs.
Delivered to the client in full. Every observed answer, as the engine returned it, with the engine that returned it, the date and time it ran, and every source it cited. Wins and losses alike, including every observation excluded from a denominator and the reason it was excluded. No figure we report is one a client cannot trace back to the answers underneath it.
Never published without written permission. Client baselines, findings, question sets and outcomes. A client story appears on this site only with explicit written approval, and reviews are optional, never required and never incentivized.
Not published at all. The operational implementation: the automation that runs the queries, model version strings, endpoints, third-party services, and what any of it costs to run. None of that changes whether a figure is true.

The line between the two is method against implementation. The method is on this page in enough detail that a technically literate reader could reproduce its shape and check the arithmetic. The implementation is how we happen to run it this quarter.

09

Version and change log

This method is versioned. Changes are published here.
VersionDateWhat changed
v1.025 August 2026First published.

A change to the method opens a new baseline. Results produced under one version are not compared against results produced under another, because the comparison would measure the method change rather than the client's position. The version in force when a figure was produced is stated on the report that carries it.

Measure your category under this method

Ten buyer questions for your category, one named engine, your position on the absence ladder. Free, under a minute, no call required.

Run the free mini-audit