CitedWell

Comparison prompts run 16 to 31 percent slower on three engines, and the slowest ones name no brand

Earlier posts on this blog found that response time is mostly a trait of the engine, and that branded prompts run slower than organic ones on three of four engines. Prompt intent, the finer field that separates a recommendation ask from a comparison ask from a best-of list, had not been checked against time. We checked it, and the result also changes how the earlier branded-prompt finding should be read.

Comparison prompts are slower on three engines, flat on the fourth

We pulled the response time (latencyMs) on every real, non-error response across our 270 live audit panels (project management, customer support, and HR software; ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search). That is 8,664 responses. Each one was joined by promptId to its panel's prompt list to recover the prompt's intent. The live panel generator only produces three intents in this dataset: recommendation, comparison, and best-of. All figures below are medians, the same measure the earlier latency posts use.

EngineBest-of: median (n)Comparison: median (n)Recommendation: median (n)Comparison vs recommendation
Claude22.0s (110)25.6s (962)22.1s (1,608)+15.7%
Gemini9.1s (95)12.8s (832)9.8s (1,399)+31.1%
ChatGPT10.6s (53)11.0s (480)10.8s (767)+1.7%
Perplexity3.0s (94)4.3s (835)3.4s (1,429)+26.2%

Percentages are computed on unrounded medians. On Claude, Gemini, and Perplexity a comparison prompt takes noticeably longer than a recommendation prompt, and the best-of prompts sit at or below the recommendation figure. ChatGPT is the exception again: its three medians sit within 0.4 seconds of each other. The best-of column rests on 53 to 110 responses per engine, so we read it as a direction and do not break it down further.

On three of the four engines, asking for a comparison costs 3 to 4 extra seconds on Claude and Gemini and about one extra second on Perplexity, which is a 16 to 31 percent slowdown against a plain recommendation ask. ChatGPT does not show it.

The direction holds in every category on three engines

We re-split each engine by category to check whether one crowded vertical was driving the averages.

Engine / categoryRecommendation: median (n)Comparison: median (n)Gap
Claude, project management22.4s (1,040)25.6s (624)+14.5%
Claude, customer support22.4s (330)25.6s (215)+14.1%
Claude, HR software21.4s (238)25.8s (123)+20.7%
Gemini, project management9.5s (968)12.1s (578)+26.5%
Gemini, customer support9.9s (263)14.1s (170)+41.3%
Gemini, HR software10.5s (168)14.0s (84)+34.3%
ChatGPT, project management10.7s (329)9.8s (220)-8.4%
ChatGPT, customer support11.0s (336)11.9s (219)+8.4%
ChatGPT, HR software10.8s (102)11.7s (41)+9.1%
Perplexity, project management3.5s (1,016)4.4s (602)+26.9%
Perplexity, customer support3.4s (196)4.1s (124)+20.6%
Perplexity, HR software3.1s (217)4.1s (109)+33.4%

Claude, Gemini, and Perplexity are slower on comparison prompts in all nine of their category cells, with gaps between 14% and 41%. ChatGPT's category cells do not agree with each other: project management runs 8.4% faster on comparison prompts while customer support and HR run 8.4% and 9.1% slower. That is why its blended figure looks like nothing. It is a small effect pointing in different directions.

The slowest comparison prompts name no brand

Here the earlier branded-prompt post needs a correction. That post found branded prompts run slower than organic ones on Claude, Gemini, and Perplexity, and guessed the cause was the engine reconciling two named products. Splitting the intent field by that same promptType field shows a confound. In this dataset every branded prompt is a comparison ("Gladly vs Five9: which is better?"), and every recommendation and best-of prompt is organic. So the organic side of that earlier split was a blend of fast recommendation prompts and slow organic comparison prompts, and the branded side was comparison prompts only.

Organic comparison prompts are the ones like "What are the pros and cons of different customer support software?" or "What are the differences between customer support software providers?". They name no brand at all. Splitting the comparison group into branded and organic gives this:

EngineRecommendation, organic (n)Comparison, branded (n)Comparison, organic (n)Organic vs branded comparison
Claude22.1s (1,608)25.1s (631)26.3s (331)+4.7%
Gemini9.8s (1,399)11.8s (547)14.8s (285)+25.5%
ChatGPT10.8s (767)10.5s (327)12.4s (153)+18.8%
Perplexity3.4s (1,429)3.8s (543)5.4s (292)+43.2%

On all four engines the organic comparison prompt is the slowest of the three groups. It beats the recommendation baseline by 15% on ChatGPT, 19% on Claude, 52% on Gemini, and 59% on Perplexity. By category, organic comparison prompts are slower than recommendation prompts in all 12 engine-by-category cells and slower than branded comparison prompts in 11 of 12. The one exception is Gemini on customer support, 13.4s organic against 15.0s branded, where the organic cell has only 52 responses.

This also explains ChatGPT. Its blended comparison figure hides two opposite pieces: branded comparison prompts are 3.3% faster than recommendation prompts, and organic comparison prompts are 14.9% slower. Its branded comparisons run about as fast as a recommendation ask and its unbranded comparisons do not, which is why it was the lone engine that looked faster on branded prompts in the earlier post.

Taken together, the data does not support the earlier explanation that naming two brands is what costs time. On every engine, an open-ended comparison with no brand named takes longer than a comparison between two named products. What separates the slow prompts is the request to survey a whole category and weigh options, which is a broader synthesis task than checking one product against another. The branded slowdown reported earlier was real, but it came from the comparison, and the brand names were riding along.

Latency, length, and citations move on different engines

Two earlier posts split other measures by the same intent field. Putting comparison-versus-recommendation results side by side shows that no single mechanism explains the slowdown:

EngineResponse timeWord countCitation count
Claude+15.7%-0.8%+26.5%
Gemini+31.1%+28.6%+39.1%
ChatGPT+1.7%-4.2%-0.4%
Perplexity+26.2%+40.4%+0.5%

Response time is a median comparison from this post, while word count and citation count are the average-based figures from the earlier posts, so read the table by direction and rough size. Gemini moves on all three measures. Claude gets slower and pulls more sources without writing more, which is consistent with extra retrieval work. Perplexity gets slower and writes more, but its citation count stays inside its fixed 5 to 10 source window, which is consistent with the extra time going into the text. ChatGPT changes on none of them when the two comparison types are pooled. A slower comparison answer is a different piece of extra work on each engine.

Consistent with the known baseline

Blending all three intents back together per engine gives median response times of 23.3s for Claude, 10.5s for Gemini, 10.8s for ChatGPT, and 3.6s for Perplexity across 8,664 responses. Those match the response-latency-by-engine post's published figures exactly, along with its per-engine response counts, so this is the same dataset read along a new field. The branded and organic medians in the table above (for example Claude 25.1s branded, Gemini 11.8s branded) also match the earlier branded-versus-organic latency post.

What to take from it

Response time is not a visibility measure, and nothing here says a slower response is more likely to name your brand. It does say something about which buyer questions make an engine work hardest. A prompt asking an engine to survey a category and weigh the options is the heaviest request in our panels, and that is the kind of question a buyer asks while building a shortlist. A panel that overweights that style will run longer, and any latency figure quoted without saying which prompt types were in the mix is hard to compare.

Find out how the AI engines your buyers use answer both kinds of question about your category: the open comparison and the head-to-head with a named competitor.

Get an AI Visibility Audit, $490

Methodology

Data drawn from 270 live, search-grounded audit panels (project management, customer support, and HR software brands), each run across four AI engines: ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search. We re-read every real (non-error) response from results.jsonl fresh this session, 8,664 in total, and joined each response's promptId to that panel's panel.json prompt list to recover its recorded intent (recommendation, comparison, or best-of; the live panel generator does not produce review, alternative, or research intents) and its promptType (branded or organic). Two legacy panels lack a promptType field and were classified by whether the prompt text contains the panel's brand name, which affects 20 prompts. Response time is the wall-clock latencyMs recorded by our runner, reported as a median. Latency includes provider load at the time of each call and was not randomized across time of day, so small gaps should be treated as directional. ChatGPT contributes 1,300 non-error responses out of 5,225 calls because of quota errors covered in earlier posts, so its figures describe the calls that succeeded. No development-rail or fixture data is included; all responses came from live engine calls. Data collected June-August 2026.