CitedWell

Three quarters of uncited AI answers come from six engine and prompt-wording pairs

An earlier post on this blog found that 3.4% of AI responses carry no citations at all, and that those answers run shorter. It did not ask which prompts produce them. We checked whether a prompt that names the brand behaves differently from one that does not. It does, but the split turns out to be a proxy for something narrower: the exact wording of the question, on the engines that answer some wordings without any sources.

Branded prompts almost always come back with sources

We took every real, non-error response across our 270 live audit panels (project management, customer support, and HR software; ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search). That is 8,664 responses. Each was joined by promptId to its panel's prompt list to recover whether the prompt named the brand (branded) or not (organic), and we counted the responses whose citation list was empty.

EngineOrganic prompts: no citations (n)Branded prompts: no citations (n)
ChatGPT12.2% (119 of 973)0.6% (2 of 327)
Claude6.3% (128 of 2,049)0.8% (5 of 631)
Gemini2.0% (35 of 1,779)0.7% (4 of 547)
Perplexity0.0% (0 of 1,815)0.0% (0 of 543)
All engines4.3% (282 of 6,616)0.5% (11 of 2,048)

On the face of it, naming a brand in the prompt makes the engine go and read something. Perplexity gives no signal here, since it never returned fewer than 5 citations in this dataset.

Branding turns out to be a proxy

An earlier post found that in these panels every branded prompt is a head-to-head comparison ("Brand A vs Brand B: which is better?"), while every recommendation and best-of prompt is organic. So the organic column above mixes three different kinds of ask. Splitting by intent, and splitting comparison prompts by whether they name a brand, gives a less tidy picture:

EngineRecommendationBest-ofComparison, organicComparison, branded
ChatGPT11.1% (85 of 767)0.0% (0 of 53)22.2% (34 of 153)0.6% (2 of 327)
Claude8.0% (128 of 1,608)0.0% (0 of 110)0.0% (0 of 331)0.8% (5 of 631)
Gemini2.3% (32 of 1,399)0.0% (0 of 95)1.1% (3 of 285)0.7% (4 of 547)
Perplexity0.0% (0 of 1,429)0.0% (0 of 94)0.0% (0 of 292)0.0% (0 of 543)
All engines4.7% (245 of 5,203)0.0% (0 of 352)3.5% (37 of 1,061)0.5% (11 of 2,048)

The intent field does not explain it either. ChatGPT goes uncited most often on organic comparison prompts (22.2%), the very group where Claude never does (0 of 331). Claude's uncited answers are almost all in recommendation prompts (128 of 133). Best-of prompts never go uncited on any engine. Neither "branded" nor "intent" is the variable that separates the answers with sources from the ones without. Something finer is.

The variable is the wording of the question

Our panel generator uses a fixed set of prompt templates, with the category name filled in. Grouping responses by template, restricted to the versions with no use-case qualifier tacked on ("for agencies", "for remote teams"), gives nine wordings. This table shows the share of responses with no citations for each, by engine:

Prompt wordingChatGPTClaudeGeminiPerplexity
What [category] do you recommend?59.3% (32/54)45.3% (53/117)24.0% (23/96)0% (0/104)
What [category] should I use?65.0% (26/40)0% (0/92)1.3% (1/80)0% (0/83)
Can you recommend a good [category]?0% (0/60)44.5% (53/119)1.9% (2/104)0% (0/102)
What are the differences between [category] providers?71.1% (32/45)0% (0/101)1.1% (1/92)0% (0/88)
What are the pros and cons of different [category]?4.3% (2/46)0% (0/108)1.1% (1/90)0% (0/96)
Compare the top [category] options0% (0/62)0% (0/122)1.0% (1/103)0% (0/108)
Which [category] is best for small businesses?0% (0/54)0% (0/113)0% (0/102)0% (0/102)
Who offers the best [category] service?0% (0/47)0% (0/99)0% (0/89)0% (0/86)
Best [category]0% (0/53)0% (0/110)0% (0/95)0% (0/94)

Six cells stand out: ChatGPT on three wordings, Claude on two, and Gemini on one. Everywhere else the rate sits at or near zero. The contrast inside a single engine is sharper than any prompt-type split. ChatGPT returns no sources for 71.1% of "What are the differences between [category] providers?" responses and 4.3% of "What are the pros and cons of different [category]?" responses. Both are organic comparison asks, worded a few words apart. Claude returns no sources for 44.5% of "Can you recommend a good [category]?" and none at all for "What [category] should I use?".

Those six engine-and-wording cells cover 471 responses, 5.4% of the dataset, and hold 219 of the 293 uncited responses (74.7%). That includes 106 of Claude's 133, 90 of ChatGPT's 121, and 23 of Gemini's 39.

For Claude, it is also a category effect

The two Claude wordings look alike, so we split them by category. All 106 of Claude's uncited answers on those two wordings are in project management: 53 of 75 for "What project management software do you recommend?" and 53 of 71 for "Can you recommend a good project management software?", 72.6% combined. In customer support and HR software the same two wordings produced 0 uncited answers in 90 responses. ChatGPT's pattern is less category-bound: its three heavy wordings are uncited in both project management and customer support (for example 12 of 17 and 14 of 21 for "should I use"), with HR too thin to read at 2 to 9 responses per cell.

This also ties back to an earlier finding that Claude's project management null responses, where none of the tracked brands were named, come almost entirely from the "for agencies" qualifier. On that qualifier, Claude returned citations on every one of 328 responses. The same engine that answers a plain project management recommendation ask with no sources 30.3% of the time (106 of 350) returns sources almost every time once the question is qualified (20 of 690 uncited, 2.9%).

Adding a qualifier changes the behavior

The same two recommendation wordings also appear with a use-case qualifier added, such as "for agencies" or "for startups". Those versions come back without sources far less often than the plain versions.

Recommendation wordingPlain: no citations (n)With use-case qualifier: no citations (n)
What [category] do you recommend?, all engines29.1% (108 of 371)3.7% (27 of 737)
Can you recommend a good [category]?, all engines14.3% (55 of 385)2.8% (18 of 635)
Can you recommend a good [category]?, Claude only44.5% (53 of 119)8.7% (17 of 196)
What [category] do you recommend?, ChatGPT only59.3% (32 of 54)20.6% (22 of 107)

The direction is the same for both engines and both wordings. It is consistent with the engine deciding whether a question needs current sources, and a bare "what do you recommend" reading as answerable from what the model already knows.

Consistent with the baseline

Across all 8,664 responses, 293 carried no citations (3.4%): ChatGPT 121 (9.3%), Claude 133 (5.0%), Gemini 39 (1.7%), Perplexity 0. Those match the counts and rates in the zero-citation-responses post exactly, as do the per-engine response totals (ChatGPT 1,300, Claude 2,680, Gemini 2,326, Perplexity 2,358). The intent-group counts in the second table sum back to 8,664 and the zero-citation counts to 293, so this is the same dataset read along the prompt template rather than a new sample.

What to take from it

An empty citation list does not mean an engine failed. On these panels it means the engine answered a certain kind of question without retrieving anything we can see. On such an answer, nothing on the web is cited, so a brand missing from it cannot be diagnosed as a gap in third-party coverage, and a fix that works by getting onto the pages an engine reads has nothing to act on. It has to be read as a different kind of result, from a different kind of question.

It also means the mix of wordings in a prompt panel changes what a measurement can show. A panel built mostly from short, generic recommendation asks will look very different on Claude in project management than one built from qualified or head-to-head questions. A scan that leaned on one or two of the six wordings above would see roughly 45% to 71% uncited answers on some engines and near zero on others, depending only on which wordings it picked.

Find out which of your buyers' questions the AI engines answer with sources and which they answer without, and what each one says about your brand.

Get an AI Visibility Audit, $490

Methodology

Data drawn from 270 live, search-grounded audit panels (173 project management, 58 customer support, and 39 HR software brands), each run across four AI engines: ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search. We re-read every real (non-error) response from results.jsonl fresh this session, 8,664 in total, and joined each response's promptId to that panel's panel.json prompt list to recover its intent (recommendation, comparison, or best-of), its promptType (branded or organic), its use-case qualifier if any, and its text. Two legacy panels lack a promptType field and were classified by whether the prompt text contains the panel's brand name, which affects 43 responses. Prompt wording was identified by replacing the category name, brand, competitor names, and use-case phrase in each prompt's text with placeholders and grouping the results; the wording table covers only prompts with no use-case qualifier. A response counts as uncited when its citation list is empty. That reflects what the API returned; we cannot see inside the call to know whether a search was attempted. ChatGPT contributes 1,300 non-error responses out of 5,225 calls because of quota errors covered in earlier posts, so its per-wording counts (40 to 62 responses) are small and describe the calls that succeeded. Cell sizes below 60 responses should be read as direction, not precision. No development-rail or fixture data is included; all responses came from live engine calls. Data collected June-August 2026.