Do AI search engines return the same vendors? We measured it against their own noise.

ChatGPT, Claude, Perplexity and Google AI Overviews are four different products from three different companies. If they return different shortlists for the same buyer question, a single visibility score is a measurement of one product wearing the clothes of a category. We put 280 B2B software buyer questions across 40 categories to all four on 4 August 2026, and re-ran 25 of them three times so engine-to-engine difference could be read against each engine's own run-to-run noise.

0.499
Google AI Overviews agrees with itself
0.240
and with other engines
+0.002
Claude's gap once list length is controlled
92.1%
of these buyer questions returned an AI Overview
01

The result, per engine

A pooled figure would be dominated by whichever engine contributed the most repeat pairs, so every number here names its engine and its sample size.

EngineRepeat pairsAgrees with itselfAgrees with other enginesGap (95% CI)
Google AI Overviews620.4990.240+0.258 (0.178 to 0.338)
Claude740.4420.277+0.165 (0.082 to 0.242)

Jaccard similarity of the vendor set, 25 stability questions, bootstrap 95% confidence intervals. Source: The 2026 State of Generative Engine Optimization v3.0, DOI 10.5281/zenodo.21789120

Per-engine chart of within-engine versus between-engine vendor-set agreement. On Google AI Overviews the gap stays positive and significant under raw Jaccard, under top-five truncation and under the overlap coefficient. On Claude the gap falls to +0.002 and is no longer significant once list length is controlled.

Figure 3: within-engine versus between-engine agreement, per engine, under three measures. Source: DOI 10.5281/zenodo.21789120

02

Then we controlled for list length, and it changed the answer

Engines name very different numbers of vendors per answer: Google AI Overviews about 4.8, Claude 8.8, Perplexity 10.7. Jaccard between two similar-length lists runs mechanically higher than between two lists of very different length. And an engine compared with itself is length-matched by construction, while two different engines are not. So we reran the identical comparison two more ways.

EngineRaw JaccardTruncated to first 5 namedOverlap coefficient
Google AI Overviews+0.258 holds+0.167 holds+0.193 holds
Claude+0.165 holds+0.075 not significant+0.002 not significant

On Google AI Overviews the separation survives every control. On Claude it does not. The length-normalised gap is +0.002. On Claude, what looks like cross-engine divergence is not distinguishable from the two engines naming different-length lists.

Source: DOI 10.5281/zenodo.21789120

03

What holds regardless

An engine agrees with itself about half the time. Even the strongest case is 0.499 across reruns of the identical question. A single-run visibility score is a coin flip reported to one decimal place.

Source: DOI 10.5281/zenodo.21789120

Mean pairwise Jaccard across all engine pairs is 0.313 (95% CI 0.295 to 0.330), 590 pair observations.

Source: DOI 10.5281/zenodo.21789120

Four-by-four matrix of mean pairwise Jaccard similarity between the four engines. The most alike pair is Claude and Perplexity at 0.327 and the least alike is ChatGPT and Google AI Overviews at 0.223.

Figure 1: 4x4 Jaccard matrix. Source: DOI 10.5281/zenodo.21789120

Across 142 questions three engines all answered, 63.3% of distinct vendor mentions came from exactly one engine and 17.5% from all three.

Source: DOI 10.5281/zenodo.21789120

Consensus distribution for Claude and Google AI Overviews, the two engines that completed collection. Of 2,664 distinct vendor mentions across the 258 questions both answered, 69.8 percent were named by one engine only and 30.2 percent by both.

Figure 2: consensus distribution for the two engines that collected in full. Of 2,664 vendor mentions across the 258 questions Claude and Google AI Overviews both answered, 69.8% came from one engine only and 30.2% from both. The 63.3% and 17.5% above are the three-engine reading of the same measure. Source: DOI 10.5281/zenodo.21789120

Of 32 companies testable on more than one engine, 14 were visible on some and invisible on others.

Source: DOI 10.5281/zenodo.21789120

Scatter plot of each company's Volume I single-engine visibility against its Volume III multi-engine visibility, Pearson r of 0.70 across 32 companies.

Figure 6: v1.0 baseline against v3.0 multi-engine visibility, 32 companies, Pearson r 0.70. Source: DOI 10.5281/zenodo.21789120

Google returned an AI Overview for 92.1% of these buyer questions. Where it did, it named 4.8 vendors in an average answer of about 1,300 characters.

Source: DOI 10.5281/zenodo.21789120

Stacked bars of questions attempted per engine. Google returned no AI Overview on 22 of 280 questions, a 92.1 percent answer rate, and the differing bar heights show that three of the four API accounts ran out of credit during collection.

Figure 4: answers, non-answers and collection attrition per engine. Source: DOI 10.5281/zenodo.21789120

Citation source mix per engine. Vendor-owned domains are 85.6 percent of ChatGPT's citations, 44.0 percent of Claude's, 39.2 percent of Perplexity's and 38.3 percent of Google AI Overviews'.

Figure 5: citation source mix per engine. Source: DOI 10.5281/zenodo.21789120

04

Method and limits

280 questions, 40 categories, four production AI search products, collected 05:28 to 16:29 UTC on 4 August 2026. Models as reported by each API: gpt-5.5-2026-04-23, claude-sonnet-5, sonar-pro, and Google AI Overviews via SerpApi. ChatGPT had to be forced to search: with tool_choice auto it returned zero citations on every pilot question. Three of four API accounts ran out of credit during collection and one was topped up and finished, so per-engine sample sizes differ: 280 answers on Google AI Overviews and Claude, 175 on Perplexity, 18 on ChatGPT. The four-engine intersection is 18 questions and never carries a headline here.

This is an independent industry study published as an open dataset with the analysis code that produced every figure in it. It has not been through academic peer review.

Data and analysis code: github.com/Broadcastwell/state-of-geo-2026. DOI 10.5281/zenodo.21789120

05

Our own score

Broadcastwell ran the measurement and sells in this category, so it is excluded from every ranking on this site. Broadcastwell was named in 0 of the 200 answers and cited 0 times among the 663 citations.

Our standing scorecard is at broadcastwell.com/visibility.

06

Cite this page

Sivakumar, S. (2026). Divergence Survives a Length Control on Google AI Overviews and Vanishes on Claude. Measuring Vendor-Set Agreement Across Four AI Search Engines. The 2026 State of Generative Engine Optimization, v3.0. Zenodo. https://doi.org/10.5281/zenodo.21789120

Prior volumes: The 2026 State of GEO, Volume I, DOI 10.5281/zenodo.21537014, and The Absence Ladder, State of GEO Volume II, DOI 10.5281/zenodo.21586091.

Last updated: 4 August 2026.

Run the same check on your own category

Five buyer questions through two engines, results by email in about ten minutes. No call required.

Run your free instant check