Do AI search engines return the same vendors?
ChatGPT, Claude, Perplexity and Google AI Overviews are four different products from three different companies. If they return different shortlists for the same buyer question, a single visibility score is a measurement of one product wearing the clothes of a category. We put 280 B2B software buyer questions across 40 categories to all four on 4 August 2026, and re-ran 25 of them three times so engine-to-engine difference could be read against each engine's own run-to-run noise.
The result, per engine
One question on the left divides into four separate paths on the right, each ending at a different height, showing that engines answer the same question differently.
- One buyer question, asked of every engine
- Each engine returns its own answer set
- The difference between them is measured, not averaged away
A pooled figure would be dominated by whichever engine contributed the most repeat pairs, so every number here names its engine and its sample size.
| Engine | Repeat pairs | Agrees with itself | Agrees with other engines | Gap (95% CI) |
|---|---|---|---|---|
| Google AI Overviews | 62 | 0.499 | 0.240 | +0.258 (0.178 to 0.338) |
| Claude | 74 | 0.442 | 0.277 | +0.165 (0.082 to 0.242) |
Jaccard similarity of the vendor set, 25 stability questions, bootstrap 95% confidence intervals. Source: The 2026 State of Generative Engine Optimization v3.0, DOI 10.5281/zenodo.21789120

Figure 3: within-engine versus between-engine agreement, per engine, under three measures. Source: DOI 10.5281/zenodo.21789120
Then we controlled for list length, and it changed the answer
Engines name very different numbers of vendors per answer: Google AI Overviews about 4.8, Claude 8.8, Perplexity 10.7. Jaccard between two similar-length lists runs mechanically higher than between two lists of very different length. And an engine compared with itself is length-matched by construction, while two different engines are not. So we reran the identical comparison two more ways.
| Engine | Raw Jaccard | Truncated to first 5 named | Overlap coefficient |
|---|---|---|---|
| Google AI Overviews | +0.258 holds | +0.167 holds | +0.193 holds |
| Claude | +0.165 holds | +0.075 not significant | +0.002 not significant |
On Google AI Overviews the separation survives every control. On Claude it does not. The length-normalised gap is +0.002. On Claude, what looks like cross-engine divergence is not distinguishable from the two engines naming different-length lists.
Source: DOI 10.5281/zenodo.21789120
What holds regardless
An engine agrees with itself about half the time. Even the strongest case is 0.499 across reruns of the identical question. A single-run visibility score is a coin flip reported to one decimal place.
Source: DOI 10.5281/zenodo.21789120
Mean pairwise Jaccard across all engine pairs is 0.313 (95% CI 0.295 to 0.330), 590 pair observations.
Source: DOI 10.5281/zenodo.21789120

Figure 1: 4x4 Jaccard matrix. Source: DOI 10.5281/zenodo.21789120
Across 142 questions three engines all answered, 63.3% of distinct vendor mentions came from exactly one engine and 17.5% from all three. Vendor-mention shares: 1,416 of 2,238 from one engine and 392 of 2,238 from all three, across 142 questions. Interval not reported in the published source.
Source: DOI 10.5281/zenodo.21789120

Figure 2: consensus distribution for the two engines that collected in full. Of 2,664 vendor mentions across the 258 questions Claude and Google AI Overviews both answered, 69.8% came from one engine only and 30.2% from both. The 63.3% and 17.5% above are the three-engine reading of the same measure. Vendor-mention shares: 1,860 of 2,664 from one engine and 804 of 2,664 from both, across 258 questions. Interval not reported in the published source. Vendor-mention shares: 1,416 of 2,238 from one engine and 392 of 2,238 from all three, across 142 questions. Interval not reported in the published source. Source: DOI 10.5281/zenodo.21789120
Of 32 companies testable on more than one engine, 14 were visible on some and invisible on others.
Source: DOI 10.5281/zenodo.21789120

Figure 6: v1.0 baseline against v3.0 multi-engine visibility, 32 companies, Pearson r 0.70. Source: DOI 10.5281/zenodo.21789120
Google returned an AI Overview for 92.1% of these buyer questions. Observed question proportion: 258 of 280 questions. Interval not reported in the published source. Where it did, it named 4.8 vendors in an average answer of about 1,300 characters.
Source: DOI 10.5281/zenodo.21789120

Figure 4: answers, non-answers and collection attrition per engine. Source: DOI 10.5281/zenodo.21789120

Figure 5: citation source mix per engine. Source: DOI 10.5281/zenodo.21789120
Method and limits
280 questions, 40 categories, four production AI search products, collected 05:28 to 16:29 UTC on 4 August 2026. Models as reported by each API: gpt-5.5-2026-04-23, claude-sonnet-5, sonar-pro, and Google AI Overviews via SerpApi. ChatGPT had to be forced to search: with tool_choice auto it returned zero citations on every pilot question. Collection completed with 280 answers each on Google AI Overviews and Claude, 175 on Perplexity and 18 on ChatGPT. The four-engine intersection is 18 questions and never carries a headline here.
This is an independent industry study published as an open dataset with the analysis code that produced every figure in it. Published open access under CC BY 4.0. Not peer reviewed.
Data and analysis code: github.com/Broadcastwell/state-of-geo-2026. DOI 10.5281/zenodo.21789120
Publisher disclosure
Broadcastwell publishes this page and provides AI search visibility services. The dated research findings are presented separately from current commercial claims and client evidence.
The current client measurement protocol is at broadcastwell.com/methodology.
Cite this page
Sivakumar, S. (2026). Divergence Survives a Length Control on Google AI Overviews and Vanishes on Claude. Measuring Vendor-Set Agreement Across Four AI Search Engines. The 2026 State of Generative Engine Optimization, v3.0. Zenodo. https://doi.org/10.5281/zenodo.21789120
Prior volumes: The 2026 State of GEO, Volume I, DOI 10.5281/zenodo.21537014, and The Absence Ladder, State of GEO Volume II, DOI 10.5281/zenodo.21586091.
Last updated: 4 August 2026.
Start with a measured baseline
See which buyer questions leave you out, the sources visible in those answers, and the three changes worth testing first. The Category Audit covers ten questions; the Diagnostic extends to 35 questions and 525 scheduled observed answers, adaptive runs extra. AI answers vary, so repeated observations and disagreements are reported.
Start with the $490 Category AuditResults appear in your account; a sign-in link arrives by email. Credits in full against the Diagnostic within 30 days.
Get the Diagnostic, $990Results appear in your account; a sign-in link arrives by email. Credits in full against the program.