05
Methodology
Engine. All answers were generated with a frontier large language model (Anthropic Claude Sonnet) with live web search enabled, so every answer reflects real-time retrieval, not training data alone. This edition holds the engine constant across all 860 runs so results stay directly comparable across categories; a four-engine study of this size would require 3,440 runs and introduce cross-engine variance we could not yet control for. The full audit does run all four engines (ChatGPT, Perplexity, Claude and Google AI Overviews) and compares them; the single-engine constraint applies to this study only. The Q4 2026 edition will extend the multi-engine comparison to the full research sample.
Prompts. For each company we generated 10 standardized buyer questions in its exact category: a mix of best-of, direct comparison, and use-case questions, phrased the way a real buyer types them.
Scoring. Each answer was scored on two binary outcomes: whether the measured company was named, and whether its domain appeared among the cited sources. Every cited URL was logged, producing 5,160 citations across 1,753 domains. One company was swept twice, returning 20 answers rather than 10, which is why 85 companies produce 860 scored answers. Volume II's per-company analysis uses a single score per company and therefore reports 850.
Sample. 85 B2B software companies across 61 categories, skewing toward challenger brands (the companies most likely to need this measurement). Category leaders enter the data through mentions, not through selection.
Classification. The 100 most-cited domains (36% of all citations) were classified by hand into vendor-authored, review platform, independent media, analyst, and community. Separately, all 5,160 cited URLs were string-matched against reddit.com, quora.com and stackoverflow.com; none appeared in any answer. The community figure therefore covers the full citation set, not only the classified sample.
Limitations. Single engine, single run per prompt. AI answers vary run to run: repeated runs of the same question can return different leaders, so all figures are point-in-time estimates, not fixed rankings. The sample skews toward challengers, which lowers average named rates relative to a random sample of all vendors. Percentages are rounded.