Prompt intent barely moves provider error rate. Category moves it by 80 points.
The branded-versus-organic split showed a small gap on every engine, but it lumped every recommendation, best-of, and comparison ask into one organic bucket. We re-split every call, including errored ones, into four intents across the same 270 live panels. Whether you ask for a recommendation or a comparison changes the error rate by a few points at most. Which category you are asking about changes it by tens of points.
Four intents, one flat result per engine
We read every line of results.jsonl across the 270 live panels and joined each call to its prompt's intent and promptType in panel.json. Branded comparison means a head-to-head naming the brand, organic comparison means a comparison of category options without the brand. Totals match the earlier error posts exactly: OpenAI 5,225 calls, Gemini 2,737, Perplexity 3,571, Claude 2,680.
| Engine | Recommendation | Best-of | Branded comparison | Organic comparison |
|---|---|---|---|---|
| Claude | 0.0% (0/1,608) | 0.0% (0/110) | 0.0% (0/631) | 0.0% (0/331) |
| Gemini | 14.8% (243/1,642) | 15.2% (17/112) | 15.6% (101/648) | 14.9% (50/335) |
| OpenAI | 75.4% (2,357/3,124) | 77.7% (185/238) | 72.9% (885/1,214) | 76.7% (498/649) |
| Perplexity | 32.8% (698/2,127) | 37.7% (57/151) | 37.2% (321/864) | 31.9% (137/429) |
Gemini sits between 14.8% and 15.6% in all four columns. OpenAI spans 72.9% to 77.7%. Perplexity has the widest spread, 31.9% to 37.7%, and its two highest cells are the ones with the fewest calls (151 best-of, 864 branded). No engine shows a comparison ask failing much more than a recommendation ask, or the reverse.
Inside a category, intent still barely matters
A blended table can hide a category effect, so we repeated the split within each category (best-of left out, since it is only 611 calls across all engines and too thin per cell).
| Engine and category | Recommendation | Branded comparison | Organic comparison |
|---|---|---|---|
| OpenAI, PM software | 83.8% | 80.9% | 83.4% |
| OpenAI, HR software | 86.5% | 89.9% | 91.8% |
| OpenAI, customer support | 0.3% | 0.0% | 1.4% |
| Perplexity, PM software | 2.3% | 3.4% | 3.7% |
| Perplexity, HR software | 19.6% | 23.0% | 14.3% |
| Perplexity, customer support | 76.0% | 77.4% | 75.3% |
| Gemini, PM software | 6.9% | 7.4% | 7.3% |
| Gemini, HR software | 35.9% | 41.8% | 35.4% |
| Gemini, customer support | 22.6% | 21.9% | 24.6% |
Across these nine engine and category rows, the gap between the highest and lowest intent is never more than 8.7 points, and in six of the nine it is under 4. The widest, Perplexity in HR software, rests on 49 organic comparison calls, so treat it as thin. Compare that with the category effect on the same engine: OpenAI errors on 0.3% of customer support recommendation asks and 86.5% of HR software ones, and Perplexity runs from 2.3% in project management to 76.0% in customer support. The category sets the error rate. The intent of the question rides on top of it by a few points.
Claude, again
Claude returned zero errors across all 2,680 calls in every intent, as it did in the category and prompt-type splits. A fourth cut did not move it either.
See your brand's real visibility on all four engines, with every error logged and excluded from your score, not averaged over.
Get an AI Visibility Audit, $490Methodology
Data drawn from 270 live, search-grounded audit panels (project management, customer support, and HR software brands), each run across four AI engines: ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search. We read every line of results.jsonl for each panel this session, including errored calls, and joined each response's promptId to the prompt's intent field in panel.json. Comparison prompts were split into branded and organic using promptType, with the scorer's own fallback (brand name present in the prompt text means branded) for two legacy panels that predate the field. A response counts as errored if the provider's API returned a non-empty error field. Per-engine call counts and blended error rates (OpenAI 75.1%, Perplexity 34.0%, Gemini 15.0%, Claude 0.0%) match error-rate-by-category and error-type-breakdown-by-engine, confirming this is the same dataset re-joined on a new field. Best-of is omitted from the category table because of its small per-cell counts. Distinct from error-rate-by-prompt-type, which split only branded against organic. No development-rail or fixture data is included; all responses came from live engine calls. Data collected June-July 2026.