Skip to content

How Broadcastwell measures AI visibility

This is the documented method behind every client figure we report. It states how the questions are built, which engines are asked, how many times, how an answer is scored, and how much uncertainty sits on the result. It is versioned and dated, and a change to the method is published here before it is used. Every full Broadcastwell baseline asks 35 buyer questions of five AI engines, three times each, and prints the confidence interval on every number. Version 1.1, 8 September 2026.

This document covers client measurement. It does not cover the published research.

What is in scope: the Category Audit, the 35-question baseline, and the ongoing measurement carried out inside the monthly program. All use the rules on this page.

What is out of scope: the published research volumes. Each states its own method, dates and limitations. Volumes I and II use one engine; Volume III compares ChatGPT, Claude, Perplexity and Google AI Overviews. Research and client measurements are reported separately, and their numbers are never combined.

The free 10-question check is a third instrument with its own limits, described in section 04.

Instrument scope. The measurement is single turn, API based, logged out, US English, gl=us and hl=en. It is a controlled instrument, not a replica of any one buyer's screen.

Four outcomes are recorded for every observed answer. They are reported separately and are never combined into a single score.
Mention. The brand is named in the answer text. An answer counts once for a brand it names, however many times it names it.
Citation. The brand's own domain appears among the sources the answer cited. A citation and a mention are separate outcomes. Supplying the material an engine reads and being named in the answer it writes are different results, and an answer can do either without the other.
Recommendation position. Where the brand falls when an answer returns an ordered or grouped shortlist. Recorded when the answer has a discernible order. Otherwise the report says: no discernible order in this answer. Position is never inferred from the order in which names happen to appear in a paragraph.
Competitor named. A rival from the list agreed before the run appears in the answer text. Scored by the same rule as a mention.

No composite visibility score is produced, and that is a deliberate choice rather than an omission. Mentions, citations and position move for different reasons and at different speeds. Adding them together produces one number that goes up, and hides which of the three actually moved, which is the only part you can act on. Visibility figures are also reported separately from referral traffic, self-reported discovery and pipeline correlation, for the same reason.

What is not measured: sentiment, or what an answer says about a brand as distinct from whether it names it. The measurement records presence, not tone. If an engine names a brand while describing it poorly, that is scored as a mention, and the answer text is delivered in full so the description can be read directly.

A full baseline asks 35 questions, written for the category in real buyer language. No question names any brand: not the client's, not a competitor's. The 35 sit in four groups: 10 shortlist questions, 10 questions phrased for different buyer roles, 10 use case and problem questions, and 5 evaluation and pricing questions. The 10 shortlist questions are the ones tracked between full re-scores. Findings and re-measurement report presence for each of the four groups separately, named in findings as shortlist, buyer role, use case and evaluation, so a rise in citations on use case questions is never presented as presence on the shortlist. The numbers here do not prove causation, do not predict traffic, and describe the observation window only.

The questions are the ones a buyer types when they are choosing between vendors and do not yet know who the vendors are. They are written for the category, agreed with the client before the run, and fixed in wording.

Every question in the set is a generic buyer intent question, containing no brand name at all. These are the discovery questions and they are the whole measure. When an engine names a brand in answer to one of these, it chose that brand.

A question that already contains a name hands the answer that name, so an answer repeating it measures the question rather than the engine's choice. Our own published research measured this effect directly and found that name-seeded questions materially inflated raw mention counts. That is why no name appears in any question we ask. Where a client asks for a branded or unbranded comparison, those figures are recorded and reported as their own set. Branded and unbranded figures are never blended into one headline.

The question set is frozen between runs. The same wording is sent every time, because a reworded question is a different question and the two results are not comparable. If a question set changes, that opens a new baseline. It does not continue the old one, and the comparison to earlier runs is dropped rather than presented with a note.

Complete question sets are public in the Absence Index release files and in the public sample. Each client's own question set is agreed before the run and stays private. To see which GEO agencies publish theirs, use the GEO agency RFP: 25 questions and who answers them.

Client measurement queries five engine products, each asked the identical wording.

The five engines are ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Every question is sent to all five, in the same words, in the same run. Google AI Overviews and Google AI Mode are separate surfaces and are recorded separately, because they cite different sources from each other most of the time.

The Free 10-question check is a different instrument. The visitor types a website, the check reads that site to propose a category, and one click queries one engine, Perplexity, once per question. It returns a shortlist wheel and a position on the absence ladder, and it is not a paid measurement. A Free 10-question check result and a paid measurement result are not comparable, and a Free 10-question check is never used as a baseline.

The automation that runs the queries and any third-party service involved in collection are implementation rather than method, and are not published. Model version strings appear in the published research. Client reports name the engine, not the model string. What is being bought is the engine products named above, the rules on this page, and the answers underneath the figures.

A single run is not evidence. The same question asked twice can return different answers.

Engines are not deterministic. Ask one the same question an hour apart and the vendor list can change without anything in the world having changed. A method that asks once and reports the result cannot tell the difference between a real position and a coin landing one way.

The current configuration is 35 buyer questions, five engines, and three scheduled runs of every question on every engine. That is 525 scheduled observed answers per full baseline. This is the scheduled count, not an upper limit. Scheduled and adaptive observations are reported separately, and the headline comparison uses only the scheduled panel. Where the three verdicts for a question and engine pair disagree, we run that pair up to twice more, so the extra observation goes where the disagreement is rather than everywhere at once. Runs beyond three add almost nothing where the three agree. Every observation is retained, including the ones that disagree, and the run-to-run spread is reported alongside the rate rather than averaged away.

35 questions, five engines, three scheduled runs, plus up to two further runs on any disagreeing pair. 525 scheduled observed answers per full baseline. This is the configuration as of version 1.1. Any change to it is published in the change log in section 09 before it is used.

The measurement grid: 35 questions, five engines, three scheduled runs

A table of four question groups against the five engines: ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Each group states its question count, the engines, the three scheduled runs for every question and engine pair, and the scheduled answers that result, adding to 525.

The measurement grid: 35 questions, five engines, three scheduled runs
Question groupQuestionsEnginesScheduled runsScheduled answers
Shortlist1053150
Buyer role1053150
Use case and problem1053150
Evaluation and pricing55375
Total3553525

35 questions x 5 engines x 3 scheduled runs = 525 scheduled observed answers, plus up to two further runs on any question and engine pair whose three verdicts disagreed.

In words: 35 buyer questions are asked on five engines, ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode, and every question and engine pair is run three times on schedule. Thirty-five questions multiplied by five engines multiplied by three scheduled runs is 525 scheduled observed answers. Any pair whose three verdicts disagreed is run up to twice more.

The rules below decide what counts. They are applied the same way to a win and to a loss.
Brand matching is on word boundaries. A brand counts as named when the answer contains the brand as a whole word, case-sensitive for a single plain word. A longer word that happens to contain the brand inside it does not count.
Domain matching compares hostnames. A citation counts when the cited source resolves to the brand's hostname. Subdomains of that hostname count. Near-miss domains that merely resemble it do not.
Failed responses leave the denominator. When an engine returns nothing, errors, refuses, or, in the case of ChatGPT, answers without searching the live web, that observation is excluded from both the named count and the denominator. Missing answers and engine errors are listed with the denominator for each engine and question group. It is never scored as a miss. This cuts both ways and that is the point: counting an outage as an absence would understate a client's position, and quietly dropping only the unfavorable failures would overstate it. Every exclusion is listed in full, with its reason, in the delivered data.
A question-and-engine cell has one verdict. For the proof gate, a cell has one named verdict: unanimous across its three scheduled runs, or the majority after the adaptive reruns. Each cell counts once. Extra adaptive answers do not enlarge the frozen panel. Cells lost to engine errors are excluded from both baseline and re-score and reported with the common denominator. The ten shortlist questions form 50 cells; the current v1.1 contractual gate covers all 35 agreed questions and 175 cells, as specified in the terms. Presence is measured, tone is not.
Floors are labeled as floors. Where a figure is a minimum rather than an exact count, because the underlying source is incomplete or an engine truncated its own source list, the figure is labeled a minimum. It is not presented as an exact number with a caveat somewhere else on the page.
Every rate we report is a sample, not a census. A rate published without a range is not evidence, and this section is the reason this page exists.

We do not observe every answer an engine could ever give. We observe a sample of them, under a documented protocol, and estimate a rate from that sample. A different sample of the same size would land near the same figure but not exactly on it, and how near is something that can be calculated rather than guessed.

Every headline figure is therefore published with its sample size and a 95 percent confidence interval, calculated using the Wilson score interval for a binomial proportion. Wilson is used rather than the textbook normal approximation because it stays sensible at small samples and at rates near zero, which is exactly where a first baseline usually sits.

A worked illustration. Suppose a brand is named in 79 of 525 scheduled observed answers. That is a rate of 15.0 percent, and its 95 percent confidence interval runs from 12.2 percent to 18.4 percent. Now suppose the next quarterly re-score names it in 105 of 525, a rate of 20.0 percent, whose interval runs from 16.8 percent to 23.6 percent. The intervals overlap from 16.8 to 18.4 percent. Interval overlap does not by itself decide whether a change is real. These figures illustrate the arithmetic. They are not a Broadcastwell result and not a client result.

The comparison is between two proportions on the same fixed panel, a repeated-measures design: questions repeat across engines and runs, so 525 answers are not 525 independent draws. The figures estimate change within that fixed panel, reported per engine and question group, not across every possible buyer question. Wilson intervals describe the observed proportions; they do not by themselves establish a significant change between baselines. A significance claim requires a documented paired comparison that accounts for this dependence. No such test is specified by this worked illustration, so it makes no significance claim.

This is also why the full re-score is quarterly rather than monthly. A monthly headline re-score would spend most of its life reporting movement that cannot be distinguished from noise. Brand mentions, citations and the sources engines use are reported every month, and the 10 shortlist questions carry monthly leading indicators. The controlled five-engine re-score at the full configuration remains quarterly.

The client receives everything underneath their own figures. Their figures stay theirs.
Delivered to the client in full. Every observed answer, as the engine returned it, with the engine that returned it, the date and time it ran, and every source it cited. Wins and losses alike, including every observation excluded from a denominator and the reason it was excluded. No figure we report is one a client cannot trace back to the answers underneath it.
Never published without written permission. Client baselines, findings, question sets and outcomes. A client story appears on this site only with explicit written approval. Reviews are optional, candid and never tied to a discount, payment or preferential treatment.
Not published at all. The operational implementation: the automation that runs the queries, model version strings, endpoints, third-party services, and what any of it costs to run. None of that changes whether a figure is true.

The line between the two is method against implementation. The method is on this page in enough detail that a technically literate reader could reproduce its shape and check the arithmetic. The implementation is how we happen to run it this quarter.

Did the change we shipped cause the movement, or did the category move anyway?

A single before and after on the questions we worked on cannot answer that, because engines drift on their own. The controlled lift addendum answers it by measuring a second group of questions that we deliberately did not work on, and reporting the difference between the two changes. This addendum is labeled v1.2. The method version remains v1.1 and no figure published under v1.1 changes because of it.

Two groups from one bank. Target questions and control questions are drawn from the same published question bank for the category, so the two groups are asked in the same words and scored under the same rules. The control questions receive no work during the experiment.
Assignment is locked first. Which questions are target and which are control is fixed before the first follow up run, and the lock time is printed with the result. An assignment that could be changed after seeing the answers would not be a control.
Scheduled runs only. Only scheduled runs enter a controlled lift result. Adaptive runs, taken where verdicts disagreed, are reported separately and never added in.
Four counts. Target before, target after, control before, control after. Each is a count over its own base, and each carries a Wilson 95 percent interval.
Two changes, then the difference. The change within each group is reported with a Newcombe 95 percent interval for the difference of two proportions. Controlled lift is the difference of those two changes: what the target questions did, minus what the untouched control questions did over the same window.
Three verdicts, and only three. Lift shown, when the lower bound of the controlled lift is above zero and the target questions rose. Gain on target questions, not separable from drift, when the target questions rose but the interval on the controlled lift includes zero. No gain observed, when the target questions did not rise. There is no fourth verdict and no score.
The limit, stated. Repeat runs of one question are correlated with each other, so the true uncertainty on a controlled lift is wider than the interval printed beside it. The interval is a floor on the uncertainty, not a ceiling.

A worked example, on a fictional sample company, is open at app.broadcastwell.com/sample/proof, including an experiment that reads No gain observed and stays published.

Each answer in a Pre-Call Brief is captured by a person, in a signed in seat of the named product, in a fresh chat or window with no prior context: ChatGPT in a temporary chat with personalization off, Claude in a new incognito chat with no project and no connectors, Perplexity in a new search thread with no Space and memory off, Google AI Overviews and Google AI Mode in a fresh window. One question, one answer, no follow-up, United States, English. The answer is copied word for word. Where the product offers a share link, the link is saved; otherwise a full screenshot is saved. The time, the product name and the plan shown are recorded, with every address the answer cited. All 25 answers are captured on one calendar day. Nothing is typed except the question. No person's name or contact details ever appear in a question. The prospect is described by industry, size and region, and named only when the buyer asks for that.

For each answer we record whether the vendor is named, its position among the named vendors, whether the rival is named, the answer's own verdict (the vendor, the rival, both, neither), the price the answer states and whether it matches the vendor's own pricing page read the same day, every wrong statement about the vendor with the cited page that carries it, and every address cited. A brief prints counts out of 25, never a rate.

This brief records what five named AI products answered on one date, to one set of questions, from a fresh seat. Answers change between days and between seats, and a signed in buyer with history can see a different answer. A verdict is not a ranking, an endorsement or a judgment of product quality. The brief does not promise a shortlist, a position, a win against the rival, a lead or revenue. Counts are out of 25 answers and are not rates.

Every answer is captured by a member of our team, by hand, in a signed-in seat, one question per fresh chat, with no custom instructions, memory and personalisation off where the product allows, web search on where the product offers it, in US English, using the exact question text and nothing else. We record the engine, the model label as displayed, the seat type, the UTC time, the full answer text (for Perplexity, the 60-word excerpt that carries the statements and the share link), every visible citation, and a screenshot. Google AI Overviews is recorded as present or absent before capture; an absent Overview is a valid record. Google AI Mode is captured in the AI Mode tab. Nothing is captured by a script, an extension or an automation; the engines' terms forbid it, and so do we.

A statement is one factual claim an answer makes about the company. Every statement in the 25 answers is listed and classed by our reviewer against the company's fact sheet and its own public pages: true when it matches; wrong when it contradicts the published fact on the capture date; stale when it was true before a dated change the company documented; unverifiable when neither the fact sheet nor the public pages settle it within five minutes. Unverifiable counts as neither right nor wrong. Each wrong or stale statement is traced to its causing page: the citation the engine attached if it carries the statement, else the page found by searching the statement's distinctive words, else "cause not found". Causing pages are typed: own page, directory, review site, press, forum, other. Time to AI is the number of days from a launch date to the first capture in which an engine states the launch correctly, or "not yet". Time to correction is the number of days from a fix shipping to the first capture in which an engine states the fact correctly. Counts are reported with their bases; answers with at least one wrong or stale statement are reported as n of 25 with a 95 percent interval.

The Access Check reads the company's robots.txt, home page and pricing page once each with the documented user agents of ChatGPT, Claude, Perplexity and Google (the agent traffic guide lists them with their sources), and checks for an llms.txt file and a Vendor Facts file. For each agent it prints allowed or disallowed by robots.txt and the status the server returned; a 403 or a challenge page is printed as "blocked at the edge". Nothing else is fetched, nothing is retried more than once, and the check runs only for a company that bought or requested it. A company the engines cannot read cannot be corrected by content; the Correction List starts there.

The Record records what five AI engines stated about a company in answer to five questions on stated dates, captured by people and verified against the company's own published facts. It does not prove what any buyer was told, that a correction will be learned, or when. "Wrong" means the statement contradicted the published fact on that date; "stale" means it was true before a documented change; a causing page is the page that carried the statement, not proof that the engine read it. Time to AI and time to correction are observations on capture dates. Before and after figures are observations with intervals, not a lift verdict; the controlled-lift verdict belongs to the shortlist measurement under our proof method. The Record promises no correction, no position, no leads and no revenue.

Read Set the Record for the Record Check and the free Correction Desk.

The Proof Page prints four lines, each with its source and date. Measured presence: the named rate from the latest Category Audit or Absence Index record, with its 95 percent interval. AI-referred demand: sessions in the AI referred class by engine, and leads whose self-reported source is an AI option, with the questions they said they asked, de-identified. AI-sourced pipeline: opportunities, pipeline amount and won amount from the company's own CRM for leads whose source is an AI option, as recorded. Lift: the controlled-lift verdict from our proof method, printed only where a fix shipped inside a locked target and control assignment and a follow-up run exists; otherwise "no fix window yet". A missing source prints "not measured" and the reason. Self-reported values are labelled self-reported. Google AI Overviews and AI Mode traffic is not separable by referrer and appears only through the field. Nothing on the page is modelled, weighted or forecast.

The Proof Page shows what the company's analytics, forms and CRM recorded, beside what we measured, on stated dates. A buyer who read an AI answer and typed the address arrives as direct traffic and is counted only if they say so. A referrer can be stripped by a browser or a link. A self-reported source is what the person chose. An amount is what the CRM holds. Presence and pipeline moving together is an association; the only causal statement we print is the controlled-lift verdict, under its published rule. The page promises no leads, no revenue and no position.

AI Pipeline Proof. AI Source Standard v1.

A criterion is a capability, feature, property or proof an answer tells a buyer to require, check or compare. A person builds the canonical list from the 25 captured answers. Two phrasings merge only if a buyer would write them as one line in an RFP; both phrasings and the merge note are retained. A criterion no answer stated does not enter the report.

Each answer is marked present or absent for each criterion. Frequency is n of 25 with a Wilson 95 percent interval, using the same interval as the Index. The table lists the engines that stated it, the source attached to it or no source shown, and the vendor named beside it exactly as the engine said it. Rival-favoured means a named client rival appeared beside the criterion in two or more of the 25 answers; it is an engine observation, not our judgement of that rival.

The reviewer maps each of the client's five differentiators, in the client's words, to criteria describing its capability in words a buyer would recognise. The mapping is written down; a doubtful match is absent. Coverage counts each answer once even if several mapped criteria appear. Absent means the captured answers did not tell the buyer to look for the capability, not that the product lacks it.

For each absent differentiator and rival-favoured criterion, the Fix List records the page types the engines cited, what the client lacks, the fix and whether it is a Fix Sprint target. Day-30 and monthly capture repeats use the same questions. Coverage before and after is a dated observation with intervals; it does not establish causation or promise inclusion, position, leads or revenue.

People capture in signed-in seats, one exact question per fresh chat, in US English, with no custom instructions, memory and personalisation off where available, and web search on where offered. Records include engine, displayed model label, seat, UTC time, visible citations, capturer initials and a screenshot. Perplexity retains the scored excerpt of at most 60 words and its share link. An absent Google AI Overview is a valid record with no criteria stated; Google AI Mode uses its own tab. We prohibit automated engine capture. Priya Nair, Management Consultant, reviews before release. We work on published evidence; we do not control engines or sell placement in their answers.

This method is versioned. Changes are published here.
Data table: Version, Date, What changed
VersionDateWhat changed
v1.2 addendum, controlled lift21 September 2026Adds controlled lift: target and control questions drawn from the same bank, assignment locked before the first follow up run with the lock time printed, scheduled runs only, four counts each with a Wilson 95 percent interval, the change in each group with a Newcombe interval, and the controlled lift as the difference of the two changes. Exactly three verdicts: Lift shown; Gain on target questions, not separable from drift; No gain observed. Repeat runs of one question are correlated, so true uncertainty is wider than shown. The method version remains v1.1 and no published figure changes.
v1.1 erratum11 September 2026The worked Wilson intervals overlap from 16.8 to 18.4 percent. The former claim of non-overlap and the inference of real movement were incorrect. This correction clarifies dependence, denominators and scheduled versus adaptive observations; the method version remains v1.1.
v1.18 September 20268 September 2026. Method v1.1. Questions per client 10 to 35, in four groups. Engines four to five, adding Google AI Mode. Repeat runs three, plus up to two further runs on any question and engine pair whose verdicts disagreed. Confidence intervals unchanged. Exclusion rule extended to ChatGPT answers that did not search the live web. Why: 35 questions cover four buyer-question groups within a fixed repeated-measures panel; runs beyond three add almost nothing except where the engines disagree, so that is where the extra runs go; Google AI Mode is a distinct surface that cites different sources from AI Overviews most of the time. Engagements that began under v1.0 stay on v1.0 and are never compared across versions.
v1.025 August 2026First published.

A change to the method opens a new baseline. Results produced under one version are not compared against results produced under another, because the comparison would measure the method change rather than the client's position. The version in force when a figure was produced is stated on the report that carries it. Comparability. Figures are compared only within one method version and one question set. Engagements that began under v1.0 are re-scored under v1.0. Nothing is compared across versions.

Fix Sprint, pay when named

Pay-when-named re-measure: day 30 after Sprint delivery, same 10 questions, same five engines (ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode), three runs, 150 answers, named count compared with the Audit baseline, recorded as an observation under method v1.1; one no-cost re-run on request.

Nothing is charged on day 30. Our day 30 result email gives you five business days to request one no-cost re-run. If you do not request it, we charge on business day six only if the named count is greater than the baseline. If you request it, the re-run decides and no charge occurs before that result or the end of the request window. Business days are Monday through Friday, in Bloomington, Indiana time. The email and your account show the exact deadline.

Read the payment condition

Agent Test method

Agent used as of 6 October 2026: ChatGPT Work with Cloud browser, two runs per task. Product name, dates and tools actually used are recorded in each receipt. A web search or content fetch is identified separately from a rendered browser visit.

Each task is run by ChatGPT Work with Cloud browser twice, at least two hours apart. For every session we record four things. Found: the agent identified <Vendor> as the vendor in question. Correct: the agent stated no false fact about <Vendor> when checked against <Vendor>'s own site read the same day. Reached: the page the task needed loaded for the agent with no block, challenge page or empty render. Position: where <Vendor> stood in any shortlist or comparison the agent produced. A task passes when Found, Correct and Reached are all yes in both runs. The pass count is the number of passed tasks out of ten. Blocked sessions are reported with the reason the agent itself gave and a screenshot. Nothing is typed into any form. Every transcript is delivered in full.

What the test does not prove

The Agent Test records what ChatGPT Work with Cloud browser did on stated dates with stated prompts. The agent changes over time, answers vary between runs, and a pass today is not a promise of a pass tomorrow. A pass is not a ranking, an endorsement or a judgment of product quality. The test does not measure how many buyers use agents, and it does not promise a shortlist, a position, a lead or revenue.

The Agent Gap is a separate published study of 70 vendors, three tasks, one agent and one primary run per task, with 30 repeat pairs. A tested vendor may request one re-run of its three tasks at no cost by writing to index@broadcastwell.com. Both results are published. Read The Agent Gap.

Start with a measured baseline

See which buyer questions leave you out, the sources visible in those answers, and the three changes worth testing first. The Category Audit covers ten buyer questions on five engines, three runs each. AI answers vary, so repeated observations and disagreements are reported.

Get the Category Audit, $490

Findings within 48 hours of your category confirmation. Refundable in full within 30 days of delivery.

One time price. No annual contract. All five engines included. The $490 credits once against the $2,900 Fix Sprint within 30 days of delivery, so the Sprint is $2,410.

Every figure comes from the published method, v1.1.

The monthly letter of the Absence Index

The Absence Report

Who AI engines name, what the evidence shows and what it means for B2B software teams.

Subscribe on Substack

Free. Monthly. Unsubscribe any time. Record Watch is a separate subscription.