State of AI Search 2026: We Asked 5 AI Engines to Recommend Software
We took 10 ordinary buying questions — “what’s the best CRM?”, “what’s the best password manager?” — and ran each one three times through five grounded AI engines: ChatGPT, Claude, Gemini, Perplexity, and Grok. That’s 150 real answers. Then we counted every brand each engine named. The result is a clear, and slightly unnerving, picture of who AI recommends — and who it has quietly erased.
Three Findings, Up Front
If you read nothing else, read these:
- Every category has a locked shortlist. In all 10 categories, at least one brand was named in all 15 runs (five engines, three times each). The engines converge hard on the same handful of leaders. On average, the top three brands captured 66% of all mentions in a category.
- The leaders are stable; the challengers flicker. Overall, 25% of brand appearances were inconsistent across three identical runs — but that noise is concentrated in the middle tier. The top brands showed up every single time. Ask once and you might miss a challenger; ask three times and the leaders never move.
- Well-known brands can be completely invisible. Eleven established products were never named once across 150 answers. Notion did not appear for engineering project management. NetSuite did not appear for accounting. Carrd did not appear for website builders. Reputation in the market is not the same as visibility in AI.
Who AI Recommends, by Category
Here is the full board. “Always named” means the brand appeared in all 15 runs. “Sometimes named” shows the share of the 15 runs a brand appeared in. “Never named” means zero mentions across every engine and every run.
| Category | Always named (15/15) | Sometimes named | Never named |
|---|---|---|---|
| CRM Best CRM for a B2B SaaS company | Salesforce, HubSpot, Pipedrive | Attio (73%), Zoho CRM (60%) | Freshsales, Insightly, Nutshell |
| Project management Best PM tool for a software engineering team | Jira | Linear (87%), ClickUp (67%), Asana (53%), Monday.com (47%), Trello (13%), Wrike (7%) | Notion, Basecamp, Smartsheet |
| Email marketing Best email platform for a small business | Brevo | Klaviyo (87%), Mailchimp (80%), ActiveCampaign (60%), HubSpot (40%), ConvertKit (33%), Constant Contact (33%), Beehiiv (7%) | — |
| Help desk Best customer support help desk | Zendesk, Freshdesk | Help Scout (73%), Zoho Desk (67%), Intercom (47%), HappyFox (33%), Gorgias (27%) | Kustomer |
| Web analytics Best privacy-friendly analytics tool | Plausible, Matomo, Umami | Fathom (80%), Google Analytics (67%), Simple Analytics (27%), PostHog (20%) | Cloudflare Web Analytics |
| Password manager Best password manager for a small team | 1Password, Bitwarden, Keeper | NordPass (73%), LastPass (33%), Dashlane (27%), Proton Pass (13%) | — |
| Accounting Best accounting software for a small business | QuickBooks, Xero, FreshBooks, Zoho Books | Sage (20%), Wave Accounting (20%) | NetSuite |
| E-signature Best e-signature software | DocuSign | PandaDoc (93%), Adobe Acrobat Sign (93%), SignNow (80%), Dropbox Sign (67%) | Signeasy |
| Website builder Best website builder for a small business | Squarespace, Wix, Shopify | WordPress (60%), Webflow (47%), GoDaddy (7%), Framer (7%) | Carrd |
| Video conferencing Best video conferencing for remote teams | Zoom, Google Meet, Microsoft Teams | Webex (87%), Whereby (20%), Zoho Meeting (20%), GoTo Meeting (7%) | — |
The Consensus Is Real — and It’s a Moat
The striking thing isn’t that AI has opinions. It’s how much the five engines agree. Ask any of them for the best CRM and you get Salesforce, HubSpot, and Pipedrive — every time, from every model. Best help desk? Zendesk and Freshdesk, universally. Best video conferencing? Zoom, Google Meet, Microsoft Teams, no exceptions.
This convergence is a moat for whoever is inside it. When five different AI systems, trained by five different companies, all reach for the same two or three names, that shortlist becomes self-reinforcing: they cite the sources that mention those brands, which makes those brands the default answer, which generates more content mentioning them. If you’re on the list, you compound. If you’re not, you have to break in from outside.
Being Famous Isn’t Being Visible
The most useful column in that table is the last one. These are not obscure products — they are brands with real customers and real market share that AI simply did not surface:
- Notion — zero mentions for “best project management tool for a software engineering team.” The engines read “engineering” and reached for Jira and Linear; Notion’s enormous general-purpose brand didn’t register for the specific query.
- NetSuite, Kustomer, Signeasy, Carrd, Cloudflare Web Analytics, Freshsales, Insightly, Nutshell, Basecamp, Smartsheet — all named zero times in their categories.
And plenty of household names were only partly visible. Mailchimp, arguably the most recognized email brand on earth, was named in only 80% of runs — it missed one in five. LastPass appeared in just a third. GoDaddy, despite blanketing the world in advertising, showed up in 7% of website-builder answers. The lesson is blunt: the marketing budget that wins mindshare with humans does not automatically win a citation from a model.
Is your brand on the shortlist or in the void?
This study used generic categories. openllmrank runs the exact same engine on your brand and your competitors — five grounded providers, your real buying questions, your citation rate with the evidence and a plan to close the gaps.
Get my report — $29.99Why You Can’t Trust a Single Check
A quarter of all brand appearances were inconsistent between identical runs. Same question, same model, same day — different answer. If you had checked ChatGPT once and seen your brand, you might have relaxed; if you’d checked once and missed it, you might have panicked. Both would be wrong, because one query is a coin flip on the challengers.
The signal only appears with repetition. Run each question several times across several engines and a stable picture emerges: the locked leaders, the flickering middle, the invisible tail. That’s the difference between an anecdote and a measurement, and it’s the whole reason a real visibility check runs prompts in triplicate rather than once.
What This Means for Your AEO Strategy
Four takeaways if you want to move from the invisible tail toward the locked shortlist:
- Know which bucket you’re in. Always-named, sometimes-named, or never-named is the only diagnosis that matters, and you can only know it by measuring across engines and repeated runs.
- Specificity beats fame. Notion lost the engineering query it could plausibly win because the model associates the category with Jira and Linear. Win the specific phrasing your buyers use, not just general awareness.
- Get into the sources the engines read. Every category leader shares a trait: it’s everywhere in the review sites and listicles models cite. We break down exactly which domains in which sources AI engines cite.
- Re-measure after you ship. Because answers move, you’ll only know your content worked by re-running the same prompts and watching the rate climb. See how to get mentioned in ChatGPT for the tactics.
Methodology
We want this to be reproducible, so here’s exactly what we did and where it’s limited.
- Scope: 10 software categories, one buying question each, run in July 2026.
- Engines: five grounded (web-search-enabled) models — OpenAI gpt-5.4-mini, Anthropic claude-haiku-4-5, Google gemini-3.5-flash, Perplexity sonar, xAI grok-4.3.
- Runs: three samples per question per engine = 150 answers. 149 of 150 returned grounded web sources.
- Counting: a brand counts as “named” when its name or a known alias appears in the answer text. We deliberately excluded candidate brands whose names are common English words (to avoid false matches) and used full product names where needed.
- Limits: we queried grounded provider APIs, not the consumer apps, which can use different model versions and personalization. Ten categories and three runs is a directional sample, not a census. Substring matching can miss an oblique reference or, rarely, over-count. Treat the numbers as a reproducible benchmark, not gospel.
- Tooling and cost: the whole study ran on the open-source openllmrank CLI and cost $6.01 in API calls.
Frequently Asked Questions
Which AI engines were tested?
Five grounded (web-search-enabled) models, in July 2026: OpenAI gpt-5.4-mini, Anthropic claude-haiku-4-5, Google gemini-3.5-flash, Perplexity sonar, and xAI grok-4.3. Each of 10 buying questions was run three times per engine — 150 answers total. We used the providers' grounded APIs, not the consumer apps, so treat results as a reproducible directional benchmark rather than a promise of what any single user sees.
Do different AI engines recommend different brands?
Less than you'd expect. In every one of the 10 categories, at least one brand was named in all 15 runs across all five engines. The engines converge on the same short list of leaders and largely draw recommendations from it. The variation shows up in the middle tier — challenger brands that some engines name and others skip.
How consistent are AI recommendations run to run?
The leaders are rock-solid; the challengers flicker. Overall, 25% of brand appearances were inconsistent across the three identical runs — a brand named in one run but not the next. But that inconsistency is concentrated in mid-tier brands. The top brands in each category appeared in all three runs, every time. A single query is still a sample size of one, which is why measuring visibility requires repeated runs.
Can a well-known brand be invisible in AI search?
Yes, and that's the most important finding. Eleven established brands were never named once across 150 answers — including Notion (for engineering project management), NetSuite, and Carrd. Others were named far less than their reputation implies: Mailchimp missed one run in five, LastPass appeared in only a third of runs. Brand awareness in the market does not equal visibility in AI answers.
How much did this study cost to run?
About six dollars in API calls. The entire benchmark — 150 grounded queries across five providers — was produced with the open-source openllmrank CLI for $6.01. That's the same workflow the hosted $29.99 report runs for a single brand and its competitors.
Find out what AI says about your brand
One emailed report, five grounded providers, your citation rate versus competitors, and what to do about it. $29.99, delivered in about fifteen minutes.
Get my report — $29.99