Comparison prompts that name no brand write up to 92 percent more words, and branded comparisons barely change length
Two earlier posts found that comparison prompts pull longer answers on Gemini and Perplexity and more citations on Gemini and Claude. A third post then showed that every branded prompt in our data is a comparison, so the "comparison" group mixes two different kinds of question. We went back and split length and citation count the same way.
Word count: the whole effect sits in the unbranded comparison
We took every real, non-error response across our 270 live audit panels (project management, customer support, and HR software; ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search), 8,664 in all. Each was joined by promptId to its panel's prompt list. Recommendation and best-of prompts are all organic in this dataset. Comparison prompts come in two kinds: branded ("Gladly vs Five9: which is better?") and organic ones that name no brand ("What are the pros and cons of different customer support software?"). Figures are average words per response, with the change against the recommendation baseline.
| Engine | Recommendation (n) | Comparison, branded (n) | Comparison, organic (n) |
|---|---|---|---|
| Gemini | 599.9 (1,399) | 595.3 (547), -0.8% | 1,101.6 (285), +83.6% |
| Perplexity | 221.0 (1,429) | 248.8 (543), +12.6% | 424.7 (292), +92.2% |
| ChatGPT | 704.0 (767) | 623.7 (327), -11.4% | 782.1 (153), +11.1% |
| Claude | 467.0 (1,608) | 466.5 (631), -0.1% | 457.0 (331), -2.1% |
A Gemini answer to an open comparison question runs nearly twice as long as its answer to a recommendation ask, and a Perplexity answer runs nearly twice as long as well. The branded comparison, where the buyer names two products, is almost indistinguishable from a recommendation ask on Gemini and only 13 percent longer on Perplexity. The earlier finding that comparison prompts run 29 percent longer on Gemini and 40 percent longer on Perplexity was an average of a flat group and a very long one. Weighting the two rows above by their response counts gives back those published figures.
The length effect holds in every category
| Engine / category | Branded comparison vs recommendation (n) | Organic comparison vs recommendation (n) |
|---|---|---|
| Gemini, project management | -8.9% (376) | +86.2% (202) |
| Gemini, customer support | +20.5% (118) | +77.9% (52) |
| Gemini, HR software | +8.1% (53) | +77.3% (31) |
| Perplexity, project management | +9.9% (392) | +92.6% (210) |
| Perplexity, customer support | +10.5% (84) | +75.9% (40) |
| Perplexity, HR software | +28.5% (67) | +107.4% (42) |
The organic comparison is between 76 and 107 percent longer than the recommendation baseline in all six cells. The branded comparison moves between -9 and +29 percent with no steady direction. Several organic cells rest on 31 to 52 responses, so read the exact percentages loosely. The gap between the two columns is large enough that the ordering does not depend on them. ChatGPT is milder and less consistent: its organic comparison is 11 percent longer than a recommendation ask while its branded comparison is 11 percent shorter. Claude does not change length with any intent.
Citations: Gemini and Claude both lean on the unbranded comparison
The same split on average citations per response:
| Engine | Recommendation | Comparison, branded | Comparison, organic |
|---|---|---|---|
| Gemini | 12.11 | 15.31, +26.5% | 19.76, +63.2% |
| Claude | 5.45 | 6.68, +22.5% | 7.32, +34.3% |
| ChatGPT | 6.42 | 6.40, -0.3% | 6.38, -0.7% |
| Perplexity | 8.78 | 8.85, +0.8% | 8.77, -0.1% |
Here the branded comparison does carry a real effect on Gemini and Claude, roughly +22 to +27 percent, and the unbranded comparison adds more on top: Gemini cites 29 percent more on an organic comparison than on a branded one, Claude 10 percent more. ChatGPT and Perplexity are flat on both, which fits what earlier posts found: Perplexity returns a fixed window of 5 to 10 sources whatever the prompt, and ChatGPT's citation count did not move with intent in any split.
By category, both Gemini and Claude show a positive gap for both comparison types in all six of their cells. The organic comparison sits above the branded one in five of those six. The exception is Claude on HR software, where the branded comparison is +26.3 percent and the organic one +15.1 percent (45 responses). Gemini customer support is nearly a tie, +30.2 against +31.3 percent.
What this changes about the earlier posts
The earlier word-count post found Gemini and Perplexity swing hard on intent and Claude does not. That holds. What it could not say is which comparison prompts cause the swing. It is the ones that name no brand. The earlier citation post found comparison prompts pull more sources on Gemini and Claude. That holds too, and it has two layers: a moderate rise for any comparison, and a larger one when the buyer asks for a survey of the category. Together with the response time finding, the unbranded comparison is the heaviest request in our panels on time, on length where the engine writes more, and on citations where the engine cites more. The branded head-to-head is the light version of the same question on length and time, and only moderately heavier than a recommendation ask on citations.
The three measures still do not move together on every engine. Gemini moves on all three. Perplexity moves on length and time while its citation count stays inside its window. Claude moves on citations and time while its word count stays flat. ChatGPT is the quiet one on citations and on time. Nothing in this data says a longer or more heavily cited answer is more likely to name your brand.
Consistent with the known baseline
Adding the four groups back together gives 8,664 responses, with per-engine counts of 2,680 for Claude, 2,326 for Gemini, 1,300 for ChatGPT, and 2,358 for Perplexity. These match every earlier post in the series. Blended average word counts (465.7 Claude, 658.3 Gemini, 692.1 ChatGPT, 251.4 Perplexity) and citation counts (6.02, 13.77, 6.48, 8.80) match the published figures, so this is the same dataset read along a finer field. Three Gemini responses carry an empty text field but no error, and they are counted as zero words and zero citations, the same as in the earlier posts.
What to take from it
When you read any length or citation benchmark for an AI engine, ask which kind of comparison it includes. A panel heavy on open-ended category comparisons will show longer answers and more sources on Gemini, and longer answers on Perplexity, than a panel built from head-to-head prompts, even with the same brands and the same engines. A panel that mixes both without saying so produces an average that describes neither.
Find out how the AI engines your buyers use answer both kinds of question about your category: the open comparison and the head-to-head with a named competitor.
Get an AI Visibility Audit, $490Methodology
Data drawn from 270 live, search-grounded audit panels (project management, customer support, and HR software brands), each run across four AI engines. We re-read every non-error response from results.jsonl fresh this session, 8,664 in total, and joined each response's promptId to that panel's panel.json to recover its intent (recommendation, comparison, or best-of) and promptType (branded or organic). Two legacy panels lack a promptType field and were classified by whether the prompt text contains the panel's brand name, which affects 20 prompts. Word count is a whitespace split on the response text. Citation count is the length of the response's citations array, raw and not deduplicated, the same measure as the earlier citation posts. Figures are averages. Best-of prompts are not shown in the tables because they rest on 53 to 110 responses per engine; they sit near the recommendation baseline on words and citations except on ChatGPT and Claude citations (+24.6 and +22.0 percent). ChatGPT contributes 1,300 non-error responses out of 5,225 calls because of quota errors covered in earlier posts, so its figures describe the calls that succeeded. Differences are correlations from an unrandomized prompt set. No development-rail or fixture data is included; all responses came from live engine calls. Data collected June-August 2026.