Do AI search engines return the same vendors? We measured it against their own noise.
ChatGPT, Claude, Perplexity and Google AI Overviews are four different products from three different companies. If they return different shortlists for the same buyer question, a single visibility score is a measurement of one product wearing the clothes of a category. We put 280 B2B software buyer questions across 40 categories to all four on 4 August 2026, and re-ran 25 of them three times so engine-to-engine difference could be read against each engine's own run-to-run noise.
The result, per engine
A pooled figure would be dominated by whichever engine contributed the most repeat pairs, so every number here names its engine and its sample size.
| Engine | Repeat pairs | Agrees with itself | Agrees with other engines | Gap (95% CI) |
|---|---|---|---|---|
| Google AI Overviews | 62 | 0.499 | 0.240 | +0.258 (0.178 to 0.338) |
| Claude | 74 | 0.442 | 0.277 | +0.165 (0.082 to 0.242) |
Jaccard similarity of the vendor set, 25 stability questions, bootstrap 95% confidence intervals. Source: The 2026 State of Generative Engine Optimization v3.0, DOI 10.5281/zenodo.21789120

Figure 3: within-engine versus between-engine agreement, per engine, under three measures. Source: DOI 10.5281/zenodo.21789120
Then we controlled for list length, and it changed the answer
Engines name very different numbers of vendors per answer: Google AI Overviews about 4.8, Claude 8.8, Perplexity 10.7. Jaccard between two similar-length lists runs mechanically higher than between two lists of very different length. And an engine compared with itself is length-matched by construction, while two different engines are not. So we reran the identical comparison two more ways.
| Engine | Raw Jaccard | Truncated to first 5 named | Overlap coefficient |
|---|---|---|---|
| Google AI Overviews | +0.258 holds | +0.167 holds | +0.193 holds |
| Claude | +0.165 holds | +0.075 not significant | +0.002 not significant |
On Google AI Overviews the separation survives every control. On Claude it does not. The length-normalised gap is +0.002. On Claude, what looks like cross-engine divergence is not distinguishable from the two engines naming different-length lists.
Source: DOI 10.5281/zenodo.21789120
What holds regardless
An engine agrees with itself about half the time. Even the strongest case is 0.499 across reruns of the identical question. A single-run visibility score is a coin flip reported to one decimal place.
Source: DOI 10.5281/zenodo.21789120
Mean pairwise Jaccard across all engine pairs is 0.313 (95% CI 0.295 to 0.330), 590 pair observations.
Source: DOI 10.5281/zenodo.21789120

Figure 1: 4x4 Jaccard matrix. Source: DOI 10.5281/zenodo.21789120
Across 142 questions three engines all answered, 63.3% of distinct vendor mentions came from exactly one engine and 17.5% from all three.
Source: DOI 10.5281/zenodo.21789120

Figure 2: consensus distribution for the two engines that collected in full. Of 2,664 vendor mentions across the 258 questions Claude and Google AI Overviews both answered, 69.8% came from one engine only and 30.2% from both. The 63.3% and 17.5% above are the three-engine reading of the same measure. Source: DOI 10.5281/zenodo.21789120
Of 32 companies testable on more than one engine, 14 were visible on some and invisible on others.
Source: DOI 10.5281/zenodo.21789120

Figure 6: v1.0 baseline against v3.0 multi-engine visibility, 32 companies, Pearson r 0.70. Source: DOI 10.5281/zenodo.21789120
Google returned an AI Overview for 92.1% of these buyer questions. Where it did, it named 4.8 vendors in an average answer of about 1,300 characters.
Source: DOI 10.5281/zenodo.21789120

Figure 4: answers, non-answers and collection attrition per engine. Source: DOI 10.5281/zenodo.21789120

Figure 5: citation source mix per engine. Source: DOI 10.5281/zenodo.21789120
Method and limits
280 questions, 40 categories, four production AI search products, collected 05:28 to 16:29 UTC on 4 August 2026. Models as reported by each API: gpt-5.5-2026-04-23, claude-sonnet-5, sonar-pro, and Google AI Overviews via SerpApi. ChatGPT had to be forced to search: with tool_choice auto it returned zero citations on every pilot question. Three of four API accounts ran out of credit during collection and one was topped up and finished, so per-engine sample sizes differ: 280 answers on Google AI Overviews and Claude, 175 on Perplexity, 18 on ChatGPT. The four-engine intersection is 18 questions and never carries a headline here.
This is an independent industry study published as an open dataset with the analysis code that produced every figure in it. It has not been through academic peer review.
Data and analysis code: github.com/Broadcastwell/state-of-geo-2026. DOI 10.5281/zenodo.21789120
Our own score
Broadcastwell ran the measurement and sells in this category, so it is excluded from every ranking on this site. Broadcastwell was named in 0 of the 200 answers and cited 0 times among the 663 citations.
Our standing scorecard is at broadcastwell.com/visibility.
Cite this page
Sivakumar, S. (2026). Divergence Survives a Length Control on Google AI Overviews and Vanishes on Claude. Measuring Vendor-Set Agreement Across Four AI Search Engines. The 2026 State of Generative Engine Optimization, v3.0. Zenodo. https://doi.org/10.5281/zenodo.21789120
Prior volumes: The 2026 State of GEO, Volume I, DOI 10.5281/zenodo.21537014, and The Absence Ladder, State of GEO Volume II, DOI 10.5281/zenodo.21586091.
Last updated: 4 August 2026.
Run the same check on your own category
Five buyer questions through two engines, results by email in about ten minutes. No call required.
Run your free instant check