Skip to content

Do AI search engines return the same vendors?

ChatGPT, Claude, Perplexity and Google AI Overviews are four different products from three different companies. If they return different shortlists for the same buyer question, a single visibility score is a measurement of one product wearing the clothes of a category. We put 280 B2B software buyer questions across 40 categories to all four on 4 August 2026, and re-ran 25 of them three times so engine-to-engine difference could be read against each engine's own run-to-run noise.

Jaccard: Google AI Overviews agrees with itself
Jaccard: and with other engines
Claude's gap once list length is controlled
258 of 280 questions triggered an AI Overview. Interval not reported in the published source.

One question on the left divides into four separate paths on the right, each ending at a different height, showing that engines answer the same question differently.

  1. One buyer question, asked of every engine
  2. Each engine returns its own answer set
  3. The difference between them is measured, not averaged away
One question in. Separate answer sets out. The spread is the finding.

A pooled figure would be dominated by whichever engine contributed the most repeat pairs, so every number here names its engine and its sample size.

Data table: Engine, Repeat pairs, Agrees with itself, Agrees with other engines, Gap (95% CI)
EngineRepeat pairsAgrees with itselfAgrees with other enginesGap (95% CI)
Google AI Overviews620.4990.240+0.258 (0.178 to 0.338)
Claude740.4420.277+0.165 (0.082 to 0.242)

Jaccard similarity of the vendor set, 25 stability questions, bootstrap 95% confidence intervals. Source: The 2026 State of Generative Engine Optimization v3.0, DOI 10.5281/zenodo.21789120

Per-engine chart of within-engine versus between-engine vendor-set agreement. On Google AI Overviews the gap stays positive and significant under raw Jaccard, under top-five truncation and under the overlap coefficient. On Claude the gap falls to +0.002 and is no longer significant once list length is controlled.

Figure 3: within-engine versus between-engine agreement, per engine, under three measures. Source: DOI 10.5281/zenodo.21789120

View full-size figure

Engines name very different numbers of vendors per answer: Google AI Overviews about 4.8, Claude 8.8, Perplexity 10.7. Jaccard between two similar-length lists runs mechanically higher than between two lists of very different length. And an engine compared with itself is length-matched by construction, while two different engines are not. So we reran the identical comparison two more ways.

Data table: Engine, Raw Jaccard, Truncated to first 5 named, Overlap coefficient
EngineRaw JaccardTruncated to first 5 namedOverlap coefficient
Google AI Overviews+0.258 holds+0.167 holds+0.193 holds
Claude+0.165 holds+0.075 not significant+0.002 not significant

On Google AI Overviews the separation survives every control. On Claude it does not. The length-normalised gap is +0.002. On Claude, what looks like cross-engine divergence is not distinguishable from the two engines naming different-length lists.

Source: DOI 10.5281/zenodo.21789120

An engine agrees with itself about half the time. Even the strongest case is 0.499 across reruns of the identical question. A single-run visibility score is a coin flip reported to one decimal place.

Source: DOI 10.5281/zenodo.21789120

Mean pairwise Jaccard across all engine pairs is 0.313 (95% CI 0.295 to 0.330), 590 pair observations.

Source: DOI 10.5281/zenodo.21789120

Four-by-four matrix of mean pairwise Jaccard similarity between the four engines. The most alike pair is Claude and Perplexity at 0.327 and the least alike is ChatGPT and Google AI Overviews at 0.223.

Figure 1: 4x4 Jaccard matrix. Source: DOI 10.5281/zenodo.21789120

View full-size figure

Across 142 questions three engines all answered, 63.3% of distinct vendor mentions came from exactly one engine and 17.5% from all three. Vendor-mention shares: 1,416 of 2,238 from one engine and 392 of 2,238 from all three, across 142 questions. Interval not reported in the published source.

Source: DOI 10.5281/zenodo.21789120

Consensus distribution for Claude and Google AI Overviews, the two engines that completed collection. Of 2,664 distinct vendor mentions across the 258 questions both answered, 69.8 percent were named by one engine only and 30.2 percent by both.

Figure 2: consensus distribution for the two engines that collected in full. Of 2,664 vendor mentions across the 258 questions Claude and Google AI Overviews both answered, 69.8% came from one engine only and 30.2% from both. The 63.3% and 17.5% above are the three-engine reading of the same measure. Vendor-mention shares: 1,860 of 2,664 from one engine and 804 of 2,664 from both, across 258 questions. Interval not reported in the published source. Vendor-mention shares: 1,416 of 2,238 from one engine and 392 of 2,238 from all three, across 142 questions. Interval not reported in the published source. Source: DOI 10.5281/zenodo.21789120

View full-size figure

Of 32 companies testable on more than one engine, 14 were visible on some and invisible on others.

Source: DOI 10.5281/zenodo.21789120

Scatter plot of each company's Volume I single-engine visibility against its Volume III multi-engine visibility, Pearson r of 0.70 across 32 companies.

Figure 6: v1.0 baseline against v3.0 multi-engine visibility, 32 companies, Pearson r 0.70. Source: DOI 10.5281/zenodo.21789120

View full-size figure

Google returned an AI Overview for 92.1% of these buyer questions. Observed question proportion: 258 of 280 questions. Interval not reported in the published source. Where it did, it named 4.8 vendors in an average answer of about 1,300 characters.

Source: DOI 10.5281/zenodo.21789120

Stacked bars of questions attempted per engine. Google returned no AI Overview on 22 of 280 questions, a 92.1 percent answer rate, and the differing bar heights show the per-engine answer counts: 280 each on Google AI Overviews and Claude, 175 on Perplexity and 18 on ChatGPT.

Figure 4: answers, non-answers and collection attrition per engine. Source: DOI 10.5281/zenodo.21789120

View full-size figure

Citation source mix per engine. Vendor-owned domains are 85.6 percent of ChatGPT's citations, 44.0 percent of Claude's, 39.2 percent of Perplexity's and 38.3 percent of Google AI Overviews'.

Figure 5: citation source mix per engine. Source: DOI 10.5281/zenodo.21789120

View full-size figure

280 questions, 40 categories, four production AI search products, collected 05:28 to 16:29 UTC on 4 August 2026. Models as reported by each API: gpt-5.5-2026-04-23, claude-sonnet-5, sonar-pro, and Google AI Overviews via SerpApi. ChatGPT had to be forced to search: with tool_choice auto it returned zero citations on every pilot question. Collection completed with 280 answers each on Google AI Overviews and Claude, 175 on Perplexity and 18 on ChatGPT. The four-engine intersection is 18 questions and never carries a headline here.

This is an independent industry study published as an open dataset with the analysis code that produced every figure in it. Published open access under CC BY 4.0. Not peer reviewed.

Data and analysis code: github.com/Broadcastwell/state-of-geo-2026. DOI 10.5281/zenodo.21789120

Broadcastwell publishes this page and provides AI search visibility services. The dated research findings are presented separately from current commercial claims and client evidence.

The current client measurement protocol is at broadcastwell.com/methodology.

Sivakumar, S. (2026). Divergence Survives a Length Control on Google AI Overviews and Vanishes on Claude. Measuring Vendor-Set Agreement Across Four AI Search Engines. The 2026 State of Generative Engine Optimization, v3.0. Zenodo. https://doi.org/10.5281/zenodo.21789120

Prior volumes: The 2026 State of GEO, Volume I, DOI 10.5281/zenodo.21537014, and The Absence Ladder, State of GEO Volume II, DOI 10.5281/zenodo.21586091.

Last updated: 4 August 2026.

Start with a measured baseline

See which buyer questions leave you out, the sources visible in those answers, and the three changes worth testing first. The Category Audit covers ten questions; the Diagnostic extends to 35 questions and 525 scheduled observed answers, adaptive runs extra. AI answers vary, so repeated observations and disagreements are reported.

Start with the $490 Category Audit

Results appear in your account; a sign-in link arrives by email. Credits in full against the Diagnostic within 30 days.

Get the Diagnostic, $990

Results appear in your account; a sign-in link arrives by email. Credits in full against the program.

Every figure comes from the published method, v1.1.