Run-to-run noise: what one AI answer tells you and what three do
Dek. Broadcastwell Research ran 140 buyer questions three times each on ChatGPT, Claude, Perplexity, Google AI Overviews and Google AI Mode on 22 to 23 Sep 2026 (UTC), then asked how often one run would have told you something the other two did not. Study 3 of the September 2026 series. Data and report CC BY 4.0.
Key figures
22.2% of single-run vendor namings were not repeated by both of the other two runs of the same question on the same engine.
Base: 8,055 namings in 14,192 question x engine x vendor cells with three valid runs. Strict variant: 9.1% of namings were contradicted by both other runs (733 of 8,055). Cell variant: 37.6% of cells that named a vendor at all did not do so in all three runs (1,259 of 3,349).
What we found
- The three runs agreed on 91.1% of all cells, but on only 62.4% of the cells where the vendor was named at all (2,090 of 3,349).
- Google AI Overviews and Google AI Mode disagreed with themselves on 50.4% and 48.9% of active cells; ChatGPT 27.4%, Claude 31.2%, Perplexity 30.6%.
- A single run called 57 to 59 of 286 vendors never named across all questions and engines; three runs called 46. Per engine, one run called on average 16 to 31 more vendors absent than three runs did.
- 28 vendors (9.8%) were absent in some runs and present in others. Ranked from one run, a vendor's rank within its category differed from its three-run rank 49.1% of the time (competition ranking; with ties broken in the three-run order the floor is 28.3%); the top vendor changed in 3 of 42 single-run category rankings.
- The mean Wilson 95% interval around a vendor's mention rate was 16.9 points wide with one run and 9.45 with three.
Chart 1: was a single run's naming repeated?

On ChatGPT 15.6% of namings were not repeated by both other runs, on Claude 17.8%, on Perplexity 17.2%, on Google AI Overviews 32.6% and on Google AI Mode 30.6%.
Chart 2: absence is overstated by a single run

A single run calls on average 16 to 31 more vendors absent per engine than three runs do. On ChatGPT 115 vendors were never named across three runs against 131.3 in one run on average; on Google AI Mode 101 against 131.7.
Chart 3: three runs cut the interval almost in half

The widest case was Jira in project management: named in 16, 22 and 25 of 50 answers across the three runs, a single-run mention rate from 32.0% to 50.0%. Three runs put it at 42.0% (34.4 to 50.0).
Why it matters
A single screenshot is a sample of one. A vendor missing from one answer is a weaker signal than a vendor present in one: 77.8% of namings were repeated by both other runs, while one run overstated the count of never-named vendors by 23.9% to 28.3%. Google's two surfaces varied most in this data; this study does not say why.
Method
Source: The Absence Index, release 2026-09 (Broadcastwell), DOI 10.5281/zenodo.22907695, CC BY 4.0. 14 categories, 286 checked vendors, 140 questions, five engines, three scheduled repeats per engine and question, captured 22 to 23 Sep 2026 (UTC); 2,095 valid target answers. Matching rule as published: names and listed aliases match case-insensitive word boundaries; plain-word vendors require a listed qualifier within five words in the same sentence or the canonical domain in the cited URLs; Wilson 95% intervals. Broadcastwell's published method (broadcastwell.com/methodology, version 1.1, 8 September 2026) states: "The measurement is single turn, API based, logged out, US English, gl=us and hl=en" and "It is a controlled instrument, not a replica of any one buyer's screen." Every figure was recomputed from the raw answers by a Python implementation of the published matching rule, shipped with the dataset, and cross-checked against the Index site's own matching code: 0 of 14,300 cells differ. A cell is one question x engine x vendor; 14,192 cells have three valid runs and 108 that lost a run to an excluded answer are left out. This study measures within-day repeat variation, not change over time.
Limitations
Three runs is a small sample per cell. Same-day runs understate longer-horizon variation. API based and logged out, not a consumer interface. The question set is fixed. Single-run ties are common among rarely named vendors; the rank-change figure is 49.1% under competition ranking and 28.3% with ties broken in the three-run order. Checked vendor lists come from buyer directories. Plain-word vendors need a qualifier or domain to count. Excluded answers drop 108 cells. No causal claim about why engines vary. Engines change their models and retrieval without notice.
Dataset
Report, charts and CSVs: https://doi.org/10.5281/zenodo.22982432 (CC BY 4.0). Files: run_cells_2026-09.csv (14,192 rows), vendors_by_run_2026-09.csv (286 rows), README, scripts.
How to cite
Broadcastwell Research. (2026). Run-to-run noise in AI search answers: three runs of 140 buyer questions on five engines (Version 1.0) [Data set and report]. Broadcastwell. https://doi.org/10.5281/zenodo.22982432
A second measurement of a result you were shown is available as Broadcastwell's Second Opinion: https://broadcastwell.com/second-opinion