CitedWell

Perplexity fails more on branded prompts than organic ones, in every category. The other three engines do not agree.

Two earlier posts established provider API error rate by category and by failure type, but neither checked the split most audits care about: does naming the brand in the prompt itself change how often the call fails? We joined every response line, including errored ones, to each prompt's branded-or-organic label across the same 270 live panels. The blended answer is almost nothing, a gap under 5 points on every engine. Broken out by category, that flatness turns out to be two different things depending on the engine.

The blended gap is small everywhere

We read every line of results.jsonl across the 270 live panels this session, including errored calls, and joined each one to its prompt's branded/organic label from panel.json (two legacy panels without the field, teaser-gladly-live and teaser-kustomer-live, were classified the same way the scorer itself falls back: brand name present in the prompt text means branded). Totals match every prior error post on this dataset exactly: OpenAI 5,225 calls, Gemini 2,737, Perplexity 3,571, Claude 2,680.

EngineBrandedOrganicGap
Claude0.0% (0/631)0.0% (0/2,049)0.0 pts
Gemini15.3% (99/646)14.9% (312/2,091)0.4 pts
OpenAI73.0% (885/1,212)75.8% (3,040/4,013)2.8 pts
Perplexity37.2% (321/864)33.0% (892/2,707)4.2 pts

Claude stays at zero on both. The other three engines all show a gap under 5 percentage points, small enough on its own to call the branded/organic split irrelevant to error rate. It is not, once you split by category.

Perplexity's gap holds in all three categories. Gemini and OpenAI's does not.

EnginePM softwareHR softwareCustomer support
Perplexity, branded minus organic+0.9 pts (3.4% vs 2.5%)+3.6 pts (23.0% vs 19.4%)+1.4 pts (77.4% vs 76.0%)
Gemini, branded minus organic+0.3 pts (7.4% vs 7.1%)+6.6 pts (41.8% vs 35.2%)-2.5 pts (20.8% vs 23.3%)
OpenAI, branded minus organic-3.0 pts (80.9% vs 83.9%)+2.4 pts (89.9% vs 87.5%)-0.5 pts (0.0% vs 0.5%)

Perplexity is the only engine where the sign does not change: branded prompts error out more than organic ones in project management, HR, and customer support software alike, every time, by 0.9 to 3.6 points. The margins are modest, but three category cells all pointing the same direction is a real pattern, not noise landing on the same side three times by chance.

Gemini and OpenAI look flat in the blended table for a different reason. Gemini's category gaps range from -2.5 points in customer support to +6.6 points in HR software, crossing zero in the middle. OpenAI's range from -3.0 points in project management to +2.4 points in HR software, also crossing zero. Their small blended numbers (0.4 and 2.8 points) are not a small effect held steady across categories, they are a real per-category effect that flips sign and cancels out in the average. Reporting either engine's "branded vs organic error rate" as one number hides that the direction itself is not fixed.

If you are debugging why one of your brand's audit runs shows more provider errors than another, prompt framing is not a reliable lever to point at on three of the four engines: whether branded or organic prompts fail more depends on which category you are in, not on branded vs organic as a rule. Perplexity is the exception. There, a batch heavy on branded head-to-head prompts should be expected to run a somewhat higher error rate than one built from open recommendation prompts, in any of the three categories we track.

Claude still does not move

Across 2,680 real Claude calls split every way we have checked so far, by category, by prompt type, now by both together, the error rate stays at exactly zero. That is the same result the category-only and error-type posts already found; this split gave it one more chance to show a crack and it did not.

See your brand's real visibility on all four engines, with every error logged and excluded from your score, not averaged over.

Get an AI Visibility Audit, $490

Methodology

Data drawn from 270 live, search-grounded audit panels (project management, customer support, and HR software brands), each run across four AI engines: ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search. We read every line of results.jsonl for each panel this session, including errored calls, and joined each response's promptId to its prompt's promptType field in panel.json (branded or organic), using the scorer's own fallback rule for two legacy panels that predate the field. A response counts as errored if the provider's API returned a non-empty error field. Per-engine total call counts (OpenAI 5,225, Gemini 2,737, Perplexity 3,571, Claude 2,680) and blended error rates match error-rate-by-category and error-type-breakdown-by-engine exactly, confirming this is the same dataset re-joined on a new field, not a different sample. Distinct from error-rate-by-category (split by category only) and error-type-breakdown-by-engine (split by failure type, not prompt framing). No development-rail or fixture data is included; all responses came from live engine calls. Data collected June-July 2026.