How Broadcastwell measures AI visibility
This is the documented method behind every client figure we report. It states how the questions are built, which engines are asked, how many times, how an answer is scored, and how much uncertainty sits on the result. It is versioned and dated, and a change to the method is published here before it is used. Every full Broadcastwell baseline asks 35 buyer questions of five AI engines, three times each, and prints the confidence interval on every number. Version 1.1, 8 September 2026.
Scope
What is in scope: the Category Audit, the 35-question baseline, and the ongoing measurement carried out inside the monthly program. All use the rules on this page.
What is out of scope: the published research volumes. Each states its own method, dates and limitations. Volumes I and II use one engine; Volume III compares ChatGPT, Claude, Perplexity and Google AI Overviews. Research and client measurements are reported separately, and their numbers are never combined.
The free 10-question check is a third instrument with its own limits, described in section 04.
Instrument scope. The measurement is single turn, API based, logged out, US English, gl=us and hl=en. It is a controlled instrument, not a replica of any one buyer's screen.
What is measured
No composite visibility score is produced, and that is a deliberate choice rather than an omission. Mentions, citations and position move for different reasons and at different speeds. Adding them together produces one number that goes up, and hides which of the three actually moved, which is the only part you can act on. Visibility figures are also reported separately from referral traffic, self-reported discovery and pipeline correlation, for the same reason.
What is not measured: sentiment, or what an answer says about a brand as distinct from whether it names it. The measurement records presence, not tone. If an engine names a brand while describing it poorly, that is scored as a mention, and the answer text is delivered in full so the description can be read directly.
The question set
The questions are the ones a buyer types when they are choosing between vendors and do not yet know who the vendors are. They are written for the category, agreed with the client before the run, and fixed in wording.
Every question in the set is a generic buyer intent question, containing no brand name at all. These are the discovery questions and they are the whole measure. When an engine names a brand in answer to one of these, it chose that brand.
A question that already contains a name hands the answer that name, so an answer repeating it measures the question rather than the engine's choice. Our own published research measured this effect directly and found that name-seeded questions materially inflated raw mention counts. That is why no name appears in any question we ask. Where a client asks for a branded or unbranded comparison, those figures are recorded and reported as their own set. Branded and unbranded figures are never blended into one headline.
The question set is frozen between runs. The same wording is sent every time, because a reworded question is a different question and the two results are not comparable. If a question set changes, that opens a new baseline. It does not continue the old one, and the comparison to earlier runs is dropped rather than presented with a note.
Complete question sets are public in the Absence Index release files and in the public sample. Each client's own question set is agreed before the run and stays private. To see which GEO agencies publish theirs, use the GEO agency RFP: 25 questions and who answers them.
Engines
The five engines are ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Every question is sent to all five, in the same words, in the same run. Google AI Overviews and Google AI Mode are separate surfaces and are recorded separately, because they cite different sources from each other most of the time.
The Free 10-question check is a different instrument. The visitor types a website, the check reads that site to propose a category, and one click queries one engine, Perplexity, once per question. It returns a shortlist wheel and a position on the absence ladder, and it is not a paid measurement. A Free 10-question check result and a paid measurement result are not comparable, and a Free 10-question check is never used as a baseline.
The automation that runs the queries and any third-party service involved in collection are implementation rather than method, and are not published. Model version strings appear in the published research. Client reports name the engine, not the model string. What is being bought is the engine products named above, the rules on this page, and the answers underneath the figures.
Repetition
Engines are not deterministic. Ask one the same question an hour apart and the vendor list can change without anything in the world having changed. A method that asks once and reports the result cannot tell the difference between a real position and a coin landing one way.
The current configuration is 35 buyer questions, five engines, and three scheduled runs of every question on every engine. That is 525 scheduled observed answers per full baseline. This is the scheduled count, not an upper limit. Scheduled and adaptive observations are reported separately, and the headline comparison uses only the scheduled panel. Where the three verdicts for a question and engine pair disagree, we run that pair up to twice more, so the extra observation goes where the disagreement is rather than everywhere at once. Runs beyond three add almost nothing where the three agree. Every observation is retained, including the ones that disagree, and the run-to-run spread is reported alongside the rate rather than averaged away.
35 questions, five engines, three scheduled runs, plus up to two further runs on any disagreeing pair. 525 scheduled observed answers per full baseline. This is the configuration as of version 1.1. Any change to it is published in the change log in section 09 before it is used.
The measurement grid: 35 questions, five engines, three scheduled runs
A table of four question groups against the five engines: ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode. Each group states its question count, the engines, the three scheduled runs for every question and engine pair, and the scheduled answers that result, adding to 525.
| Question group | Questions | Engines | Scheduled runs | Scheduled answers |
|---|---|---|---|---|
| Shortlist | 10 | 5 | 3 | 150 |
| Buyer role | 10 | 5 | 3 | 150 |
| Use case and problem | 10 | 5 | 3 | 150 |
| Evaluation and pricing | 5 | 5 | 3 | 75 |
| Total | 35 | 5 | 3 | 525 |
35 questions x 5 engines x 3 scheduled runs = 525 scheduled observed answers, plus up to two further runs on any question and engine pair whose three verdicts disagreed.
In words: 35 buyer questions are asked on five engines, ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode, and every question and engine pair is run three times on schedule. Thirty-five questions multiplied by five engines multiplied by three scheduled runs is 525 scheduled observed answers. Any pair whose three verdicts disagreed is run up to twice more.
Scoring rules
Uncertainty
We do not observe every answer an engine could ever give. We observe a sample of them, under a documented protocol, and estimate a rate from that sample. A different sample of the same size would land near the same figure but not exactly on it, and how near is something that can be calculated rather than guessed.
Every headline figure is therefore published with its sample size and a 95 percent confidence interval, calculated using the Wilson score interval for a binomial proportion. Wilson is used rather than the textbook normal approximation because it stays sensible at small samples and at rates near zero, which is exactly where a first baseline usually sits.
The comparison is between two proportions on the same fixed panel, a repeated-measures design: questions repeat across engines and runs, so 525 answers are not 525 independent draws. The figures estimate change within that fixed panel, reported per engine and question group, not across every possible buyer question. Wilson intervals describe the observed proportions; they do not by themselves establish a significant change between baselines. A significance claim requires a documented paired comparison that accounts for this dependence. No such test is specified by this worked illustration, so it makes no significance claim.
This is also why the full re-score is quarterly rather than monthly. A monthly headline re-score would spend most of its life reporting movement that cannot be distinguished from noise. Brand mentions, citations and the sources engines use are reported every month, and the 10 shortlist questions carry monthly leading indicators. The controlled five-engine re-score at the full configuration remains quarterly.
What is published and what is held
The line between the two is method against implementation. The method is on this page in enough detail that a technically literate reader could reproduce its shape and check the arithmetic. The implementation is how we happen to run it this quarter.
Controlled lift, addendum v1.2
A single before and after on the questions we worked on cannot answer that, because engines drift on their own. The controlled lift addendum answers it by measuring a second group of questions that we deliberately did not work on, and reporting the difference between the two changes. This addendum is labeled v1.2. The method version remains v1.1 and no figure published under v1.1 changes because of it.
A worked example, on a fictional sample company, is open at app.broadcastwell.com/sample/proof, including an experiment that reads No gain observed and stays published.
Pre-Call Brief capture
Each answer in a Pre-Call Brief is captured by a person, in a signed in seat of the named product, in a fresh chat or window with no prior context: ChatGPT in a temporary chat with personalization off, Claude in a new incognito chat with no project and no connectors, Perplexity in a new search thread with no Space and memory off, Google AI Overviews and Google AI Mode in a fresh window. One question, one answer, no follow-up, United States, English. The answer is copied word for word. Where the product offers a share link, the link is saved; otherwise a full screenshot is saved. The time, the product name and the plan shown are recorded, with every address the answer cited. All 25 answers are captured on one calendar day. Nothing is typed except the question. No person's name or contact details ever appear in a question. The prospect is described by industry, size and region, and named only when the buyer asks for that.
For each answer we record whether the vendor is named, its position among the named vendors, whether the rival is named, the answer's own verdict (the vendor, the rival, both, neither), the price the answer states and whether it matches the vendor's own pricing page read the same day, every wrong statement about the vendor with the cited page that carries it, and every address cited. A brief prints counts out of 25, never a rate.
This brief records what five named AI products answered on one date, to one set of questions, from a fresh seat. Answers change between days and between seats, and a signed in buyer with history can see a different answer. A verdict is not a ranking, an endorsement or a judgment of product quality. The brief does not promise a shortlist, a position, a win against the rival, a lead or revenue. Counts are out of 25 answers and are not rates.
Sponsored Watch capture
Every capture is made by a person, in a ChatGPT seat on the Free or Go plan that the person owns, with personalization off, in a fresh chat, never a Temporary Chat, United States, one question, one answer. The person records whether a sponsored card appears beneath the answer, the advertiser, the card's title and landing domain, a screenshot and the time. For Google AI Mode the same question is asked in a fresh window and any ad unit is recorded the same way. Nothing is clicked. A capture reports what one seat saw on one date; a card absent in one capture may appear in the next. We print counts out of the checks made, never a rate.
Sponsored Watch is a separate observation. It never changes an Absence Index release figure. Read Absence Ads for the managed Sprint.
Record rules
Every answer is captured by a member of our team, by hand, in a signed-in seat, one question per fresh chat, with no custom instructions, memory and personalisation off where the product allows, web search on where the product offers it, in US English, using the exact question text and nothing else. We record the engine, the model label as displayed, the seat type, the UTC time, the full answer text (for Perplexity, the 60-word excerpt that carries the statements and the share link), every visible citation, and a screenshot. Google AI Overviews is recorded as present or absent before capture; an absent Overview is a valid record. Google AI Mode is captured in the AI Mode tab. Nothing is captured by a script, an extension or an automation; the engines' terms forbid it, and so do we.
A statement is one factual claim an answer makes about the company. Every statement in the 25 answers is listed and classed by our reviewer against the company's fact sheet and its own public pages: true when it matches; wrong when it contradicts the published fact on the capture date; stale when it was true before a dated change the company documented; unverifiable when neither the fact sheet nor the public pages settle it within five minutes. Unverifiable counts as neither right nor wrong. Each wrong or stale statement is traced to its causing page: the citation the engine attached if it carries the statement, else the page found by searching the statement's distinctive words, else "cause not found". Causing pages are typed: own page, directory, review site, press, forum, other. Time to AI is the number of days from a launch date to the first capture in which an engine states the launch correctly, or "not yet". Time to correction is the number of days from a fix shipping to the first capture in which an engine states the fact correctly. Counts are reported with their bases; answers with at least one wrong or stale statement are reported as n of 25 with a 95 percent interval.
The Access Check reads the company's robots.txt, home page and pricing page once each with the documented user agents of ChatGPT, Claude, Perplexity and Google (the agent traffic guide lists them with their sources), and checks for an llms.txt file and a Vendor Facts file. For each agent it prints allowed or disallowed by robots.txt and the status the server returned; a 403 or a challenge page is printed as "blocked at the edge". Nothing else is fetched, nothing is retried more than once, and the check runs only for a company that bought or requested it. A company the engines cannot read cannot be corrected by content; the Correction List starts there.
The Record records what five AI engines stated about a company in answer to five questions on stated dates, captured by people and verified against the company's own published facts. It does not prove what any buyer was told, that a correction will be learned, or when. "Wrong" means the statement contradicted the published fact on that date; "stale" means it was true before a documented change; a causing page is the page that carried the statement, not proof that the engine read it. Time to AI and time to correction are observations on capture dates. Before and after figures are observations with intervals, not a lift verdict; the controlled-lift verdict belongs to the shortlist measurement under our proof method. The Record promises no correction, no position, no leads and no revenue.
Read Set the Record for the Record Check and the free Correction Desk.
Proof Page rules
The Proof Page prints four lines, each with its source and date. Measured presence: the named rate from the latest Category Audit or Absence Index record, with its 95 percent interval. AI-referred demand: sessions in the AI referred class by engine, and leads whose self-reported source is an AI option, with the questions they said they asked, de-identified. AI-sourced pipeline: opportunities, pipeline amount and won amount from the company's own CRM for leads whose source is an AI option, as recorded. Lift: the controlled-lift verdict from our proof method, printed only where a fix shipped inside a locked target and control assignment and a follow-up run exists; otherwise "no fix window yet". A missing source prints "not measured" and the reason. Self-reported values are labelled self-reported. Google AI Overviews and AI Mode traffic is not separable by referrer and appears only through the field. Nothing on the page is modelled, weighted or forecast.
The Proof Page shows what the company's analytics, forms and CRM recorded, beside what we measured, on stated dates. A buyer who read an AI answer and typed the address arrives as direct traffic and is counted only if they say so. A referrer can be stripped by a browser or a link. A self-reported source is what the person chose. An amount is what the CRM holds. Presence and pipeline moving together is an association; the only causal statement we print is the controlled-lift verdict, under its published rule. The page promises no leads, no revenue and no position.
Checklist rules
A criterion is a capability, feature, property or proof an answer tells a buyer to require, check or compare. A person builds the canonical list from the 25 captured answers. Two phrasings merge only if a buyer would write them as one line in an RFP; both phrasings and the merge note are retained. A criterion no answer stated does not enter the report.
Each answer is marked present or absent for each criterion. Frequency is n of 25 with a Wilson 95 percent interval, using the same interval as the Index. The table lists the engines that stated it, the source attached to it or no source shown, and the vendor named beside it exactly as the engine said it. Rival-favoured means a named client rival appeared beside the criterion in two or more of the 25 answers; it is an engine observation, not our judgement of that rival.
The reviewer maps each of the client's five differentiators, in the client's words, to criteria describing its capability in words a buyer would recognise. The mapping is written down; a doubtful match is absent. Coverage counts each answer once even if several mapped criteria appear. Absent means the captured answers did not tell the buyer to look for the capability, not that the product lacks it.
For each absent differentiator and rival-favoured criterion, the Fix List records the page types the engines cited, what the client lacks, the fix and whether it is a Fix Sprint target. Day-30 and monthly capture repeats use the same questions. Coverage before and after is a dated observation with intervals; it does not establish causation or promise inclusion, position, leads or revenue.
People capture in signed-in seats, one exact question per fresh chat, in US English, with no custom instructions, memory and personalisation off where available, and web search on where offered. Records include engine, displayed model label, seat, UTC time, visible citations, capturer initials and a screenshot. Perplexity retains the scored excerpt of at most 60 words and its share link. An absent Google AI Overview is a valid record with no criteria stated; Google AI Mode uses its own tab. We prohibit automated engine capture. Priya Nair, Management Consultant, reviews before release. We work on published evidence; we do not control engines or sell placement in their answers.
Version and change log
| Version | Date | What changed |
|---|---|---|
| v1.2 addendum, controlled lift | 21 September 2026 | Adds controlled lift: target and control questions drawn from the same bank, assignment locked before the first follow up run with the lock time printed, scheduled runs only, four counts each with a Wilson 95 percent interval, the change in each group with a Newcombe interval, and the controlled lift as the difference of the two changes. Exactly three verdicts: Lift shown; Gain on target questions, not separable from drift; No gain observed. Repeat runs of one question are correlated, so true uncertainty is wider than shown. The method version remains v1.1 and no published figure changes. |
| v1.1 erratum | 11 September 2026 | The worked Wilson intervals overlap from 16.8 to 18.4 percent. The former claim of non-overlap and the inference of real movement were incorrect. This correction clarifies dependence, denominators and scheduled versus adaptive observations; the method version remains v1.1. |
| v1.1 | 8 September 2026 | 8 September 2026. Method v1.1. Questions per client 10 to 35, in four groups. Engines four to five, adding Google AI Mode. Repeat runs three, plus up to two further runs on any question and engine pair whose verdicts disagreed. Confidence intervals unchanged. Exclusion rule extended to ChatGPT answers that did not search the live web. Why: 35 questions cover four buyer-question groups within a fixed repeated-measures panel; runs beyond three add almost nothing except where the engines disagree, so that is where the extra runs go; Google AI Mode is a distinct surface that cites different sources from AI Overviews most of the time. Engagements that began under v1.0 stay on v1.0 and are never compared across versions. |
| v1.0 | 25 August 2026 | First published. |
A change to the method opens a new baseline. Results produced under one version are not compared against results produced under another, because the comparison would measure the method change rather than the client's position. The version in force when a figure was produced is stated on the report that carries it. Comparability. Figures are compared only within one method version and one question set. Engagements that began under v1.0 are re-scored under v1.0. Nothing is compared across versions.
Fix Sprint, pay when named
Pay-when-named re-measure: day 30 after Sprint delivery, same 10 questions, same five engines (ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode), three runs, 150 answers, named count compared with the Audit baseline, recorded as an observation under method v1.1; one no-cost re-run on request.
Nothing is charged on day 30. Our day 30 result email gives you five business days to request one no-cost re-run. If you do not request it, we charge on business day six only if the named count is greater than the baseline. If you request it, the re-run decides and no charge occurs before that result or the end of the request window. Business days are Monday through Friday, in Bloomington, Indiana time. The email and your account show the exact deadline.
Read the payment conditionAgent Ready
Agent Test method
Agent used as of 6 October 2026: ChatGPT Work with Cloud browser, two runs per task. Product name, dates and tools actually used are recorded in each receipt. A web search or content fetch is identified separately from a rendered browser visit.
Each task is run by ChatGPT Work with Cloud browser twice, at least two hours apart. For every session we record four things. Found: the agent identified <Vendor> as the vendor in question. Correct: the agent stated no false fact about <Vendor> when checked against <Vendor>'s own site read the same day. Reached: the page the task needed loaded for the agent with no block, challenge page or empty render. Position: where <Vendor> stood in any shortlist or comparison the agent produced. A task passes when Found, Correct and Reached are all yes in both runs. The pass count is the number of passed tasks out of ten. Blocked sessions are reported with the reason the agent itself gave and a screenshot. Nothing is typed into any form. Every transcript is delivered in full.
What the test does not prove
The Agent Test records what ChatGPT Work with Cloud browser did on stated dates with stated prompts. The agent changes over time, answers vary between runs, and a pass today is not a promise of a pass tomorrow. A pass is not a ranking, an endorsement or a judgment of product quality. The test does not measure how many buyers use agents, and it does not promise a shortlist, a position, a lead or revenue.
The Agent Gap is a separate published study of 70 vendors, three tasks, one agent and one primary run per task, with 30 repeat pairs. A tested vendor may request one re-run of its three tasks at no cost by writing to index@broadcastwell.com. Both results are published. Read The Agent Gap.
Start with a measured baseline
See which buyer questions leave you out, the sources visible in those answers, and the three changes worth testing first. The Category Audit covers ten buyer questions on five engines, three runs each. AI answers vary, so repeated observations and disagreements are reported.
Get the Category Audit, $490Findings within 48 hours of your category confirmation. Refundable in full within 30 days of delivery.
One time price. No annual contract. All five engines included. The $490 credits once against the $2,900 Fix Sprint within 30 days of delivery, so the Sprint is $2,410.