CiteWorks Studio

How AI Search Is Recommending Chatbots

Mark HuntleyBy Mark HuntleyFounder and CEO
14 minutes read

On this report

Key Takeaways

  • LiveChat, Tidio, and Zendesk capture most recommendation value, showing that AI-led buyer discovery is concentrating around a small group of chatbot platforms.
  • Raw mention frequency does not predict shortlist success: Drift, Ada, and Landbot appear in AI responses but earn little recommendation credit or Top 3 placement.
  • Top 3 rate, Top 1 rate, and average recommended rank are stronger indicators of competitive position than visibility alone in AI-generated comparisons.
  • Pricing and comparison prompts carry the highest commercial weight, and performance varies by platform, creating brand-specific openings across ChatGPT, Copilot, Gemini, Google AI, and Perplexity.

Buyer discovery in the chatbot and customer messaging market is shifting. Instead of starting with Google searches and visiting vendor websites, buyers are asking AI systems to compare platforms, explain pricing, and recommend shortlists. The AI response has become the first filter, and being named in that response is no longer enough. Being ranked and recommended is what determines whether a brand makes the buyer's shortlist.

The LLM Authority Index benchmark for June 2026 reveals a clear pattern in this category. Recommendation power is concentrating around a small set of platforms, while several well-known brands appear frequently in AI responses but fail to convert that presence into ranked shortlist positions. LiveChat leads the category in recommendation value, followed by Tidio and Zendesk. Meanwhile, brands like Drift and Ada are visible but rarely recommended. CiteWorks Studio interprets this benchmark to show what the data means for competitive positioning in AI-led discovery.

Methodology

1. Market studied: Chatbots and customer messaging platforms, including AI-powered chatbot solutions, live chat software, and customer support automation platforms.

2. Brands/entities included: Ada, Chatfuel, Drift, Freshchat, Intercom, Landbot, LiveChat, ManyChat, Tidio, and Zendesk. This universe covers major platforms but is not a full market census.

3. Data collection date/window: June 2026, snapshot taken on June 18, 2026.

4. AI platforms tested: ChatGPT, Copilot, Gemini, Google AI Mode, Google AI Overviews, and Perplexity.

5. Number of prompts tested: Prompt count was not provided. A total of 1,083 observations were analyzed across three public high-intent clusters.

6. Prompt categories: Discovery (awareness-stage), Comparison (consideration-stage), and Pricing (decision-stage) prompts representing three stages of the buyer journey.

7. Definition of a mention: A mention means the company appeared in an AI-generated response, regardless of sentiment or ranking.

8. Definition of a valid recommendation: A valid recommendation is a positive, shortlist-quality recommendation or ranked recommendation that earns recommendation credit. This is the key CiteWorks distinction: visibility is not the same as recommendation credit.

9. Ranking/scoring metrics used: Valid recommendation coverage, Top 3 rate, Top 1 rate, average recommended rank, net sentiment score, monthly AI Authority Value, monthly AI Recommendation Value, monthly AI Visibility Assist Value, and captured share of AI opportunity.

10. Limitations: This is a point-in-time benchmark. AI outputs can change with model updates and source changes. Modeled values are estimates based on commercial intent modeling and are not revenue. This report is not a full audit or full market census. The public version covers 3 of 10 total prompt clusters.

Key Findings

Recommendation power is concentrated in three brands. LiveChat, Tidio, and Zendesk capture the majority of recommendation value in the chatbot category. LiveChat leads with a monthly AI Authority Value of $586,476, followed by Tidio at $410,264 and Zendesk at $362,103. These three brands combined capture 9.5% of the total modeled monthly AI opportunity of $14.27 million. The remaining seven brands compete for a substantially smaller share.

Visibility does not equal recommendation credit. Drift appears in 10.0% of all AI responses but earns valid recommendations in only 2.3% of observations. Ada appears in 4.9% of responses but earns recommendations in 2.8%. Landbot appears in 4.8% of responses but earns virtually no recommendation credit. These brands are being seen by AI systems but are not being selected for buyer shortlists.

Top 3 placement is the decisive competitive metric. LiveChat achieves a Top 3 rate of 9.1% and an average recommended rank of 1.95, the strongest combination in the category. Zendesk achieves a Top 3 rate of 15.2% and a Top 1 rate of 10.0%. Tidio achieves a Top 3 rate of 12.6%. Brands with low Top 3 rates, including Drift at 0.7% and Ada at 0.8%, are effectively absent at the decision moment.

No single brand dominates all six platforms. LiveChat holds its strongest position on Copilot, with an AI Authority Value of $322,901, and on Google AI Mode with $106,710. Tidio leads on ChatGPT with $97,249 and on Perplexity with $66,633. Zendesk leads on Gemini with $41,809. No platform-wide leader exists, which means competitive openings are platform-specific and brand-specific.

Pricing and comparison clusters carry the highest commercial stakes. The Pricing cluster carries a buyer-stage intent multiplier of 1.5, and Tidio leads this cluster with an AI Authority Value of $131,598. The Comparison cluster carries a multiplier of 1.25, and Tidio also leads there with $164,641. Brands that underperform in these high-intent clusters lose recommendation influence precisely where purchase decisions are forming.

What Changed in the Market

Buyers evaluating chatbot platforms are no longer moving exclusively from Google search results to brand websites. They are asking AI systems to compare providers, explain reputation, summarize pricing, surface alternatives, and recommend shortlists. The AI response now functions as a pre-filter, and the difference between being mentioned and being recommended carries real commercial weight.

For B2B software categories like chatbots, this shift changes how procurement confidence is built. Buyers expect AI systems to surface platforms that are well-documented, widely reviewed, and consistently compared across multiple public sources. Brands with strong official content, broad review coverage, active community discussions, and consistent comparison page presence earn higher recommendation rates. Brands with thinner public footprints earn mentions but not shortlist positions.

The benchmark data shows that AI systems do not simply list every brand they have encountered. They select, rank, and recommend based on the strength and consistency of available public evidence. Traditional brand visibility, measured by advertising spend or website traffic, does not automatically translate into AI recommendation power. A brand can be widely discussed online yet rarely earn a Top 3 position in an AI-generated response.

This dynamic is especially pronounced in a category like chatbots, where buyers face a crowded field of platforms with overlapping feature sets. AI systems are being asked to make sense of that crowded field on the buyer's behalf. The brands that earn shortlist placement are those that have given AI systems clear, consistent, and credible material to synthesize. The gap between mention and recommendation is the central competitive dynamic in this market right now.

What the Benchmark Found

LiveChat leads the chatbot category with a monthly AI Authority Value of $586,476. The benchmark shows it appears in 29.6% of all AI responses and earns valid recommendations in 11.7% of observations. Its average recommended rank of 1.95 is the strongest in the category, and it achieves a Top 1 rate of 7.5%. LiveChat captures 4.1% of the total modeled monthly AI opportunity. It is the most consistently recommended platform across AI systems and holds the strongest combination of recommendation rate and average rank.

Tidio holds second place with a monthly AI Authority Value of $410,264. The analysis found it appears in 40.5% of all AI responses, the highest raw mention presence in the category. It earns valid recommendations in 17.0% of observations and achieves a Top 3 rate of 12.6%, with an average recommended rank of 2.53. Tidio captures 2.9% of the total modeled monthly AI opportunity and leads both the Pricing and Comparison clusters. It has the broadest visibility across AI platforms but converts that presence into recommendation credit at a lower rate than LiveChat, indicating a visibility-to-recommendation gap worth investigating.

Zendesk ranks third with a monthly AI Authority Value of $362,103. The dataset marked it as appearing in 44.0% of all AI responses and earning valid recommendations in 19.7% of observations, the highest recommendation coverage rate in the category. It achieves a Top 1 rate of 10.0% and an average recommended rank of 2.14. Zendesk captures 2.5% of the total modeled monthly AI opportunity. Its recommendation coverage rate and Top 1 rate are both category-leading figures, suggesting AI systems surface it frequently as a primary option.

ManyChat holds fourth place with a monthly AI Authority Value of $272,123. It appears in 14.8% of all AI responses and earns valid recommendations in 4.8% of observations, with an average recommended rank of 3.04. ManyChat captures 1.9% of the total modeled monthly AI opportunity, driven largely by visibility assist value rather than direct recommendation credit. Its modeled rank position suggests it earns shortlist placement but typically not at the top.

Intercom ranks fifth with a monthly AI Authority Value of $167,999. It appears in 37.1% of all AI responses and earns valid recommendations in 19.9% of observations, the second highest recommendation coverage rate in the category. Its average recommended rank of 3.21 means it is frequently included in AI shortlists but positioned further down the list than its coverage rate might suggest. Intercom is a recommendation presence leader by coverage, but a rank-weighted underperformer relative to its visibility.

Drift presents the most significant warning pattern in the dataset. It appears in 10.0% of all AI responses but earns valid recommendations in only 2.3% of observations. Its Top 3 rate is 0.7%, and its average recommended rank is 4.28. Drift carries a net sentiment score of 0.37, the second lowest among the ten tracked brands. The combination of moderate visibility, very low recommendation credit, and weak framing score identifies Drift as visible but commercially weak in AI-led discovery.

Ada appears in 4.9% of responses and earns valid recommendations in 2.8% of observations. Its Top 3 rate is 0.8%, and its average recommended rank is 4.69. Ada's net sentiment score of 0.64 indicates positive framing when mentioned, but the low recommendation rate means it is rarely positioned as a top option. The positive framing without shortlist placement is a structural gap: Ada's source material may support favorable descriptions without generating the ranked recommendation signal that shortlist eligibility requires.

Landbot appears in 4.8% of responses but earns recommendation credit in only 0.3% of observations. It achieves a Top 3 rate of 0.0%. Landbot's net sentiment score of 0.10 is the lowest in the category, suggesting that when AI systems mention Landbot, they do so in neutral or mixed contexts without endorsing it. This framing pattern is the most commercially concerning among the lower-visibility brands.

Chatfuel appears in 4.7% of responses and earns valid recommendations in 1.3% of observations. Its Top 3 rate is 0.9%, and its average recommended rank is 2.85. Chatfuel captures 0.07% of the total modeled monthly AI opportunity. Its average rank is relatively strong when it does earn recommendation credit, but the frequency of that credit is very low.

Freshchat appears in 8.9% of responses and earns valid recommendations in 1.7% of observations. Its Top 3 rate is 0.3%, and its average recommended rank is 5.07, the lowest in the category. Freshchat captures 0.15% of the total modeled monthly AI opportunity. Its rank position when recommended suggests it is appearing at the tail of shortlists rather than near the top.

Why Visibility Is Not Enough

A brand can appear in AI answers and still fail to win the buyer shortlist. This is the central commercial distinction the benchmark reveals, and it separates brands that benefit from AI-led discovery from brands that are merely present in it.

Raw mention presence measures how often a company appears in AI responses. It does not measure whether the company is recommended, how it is framed, or where it ranks. Drift appears in 10.0% of responses but earns a Top 3 recommendation only 0.7% of the time. LiveChat appears in 29.6% of responses and earns a Top 3 recommendation 9.1% of the time. Both brands have presence. Only one is consistently shortlisted.

Top 3 placement and rank-one placement tell different stories about competitive position. Zendesk has the highest Top 1 rate at 10.0%, meaning it is frequently the first brand named. LiveChat has the strongest average recommended rank at 1.95, meaning it consistently lands near the top when it appears. Tidio has the highest Top 3 rate at 12.6%, meaning it is the most frequently included in shortlists. Each metric reflects a distinct competitive advantage. None of them can be inferred from raw mention frequency alone.

Neutral or cautionary mentions do not count as recommendations. Landbot appears in 4.8% of responses but earns recommendation credit in only 0.3% of observations, and its net sentiment score of 0.10 indicates that most of those mentions are neutral rather than endorsing. Being named without being recommended is not a neutral outcome. It creates the appearance of competitive presence without the commercial benefit of shortlist eligibility.

Modeled benchmark value is not revenue. The monthly AI Authority Value figures in this report represent modeled recommendation influence, not bookings, pipeline, or return on investment. They are calibrated estimates of the commercial weight attached to recommendation positions, and they are most useful as a relative measure of competitive standing across the ten tracked brands.

Ahrefs-based search visibility, where present in the supporting data, reflects traditional organic search footprint. Strong search presence may contribute to the public evidence layer that AI systems draw from, but it does not independently produce AI recommendation credit. The relationship between search visibility and AI recommendation strength is indirect and mediated by source quality, framing, and consistency.

The Citation Layer

AI recommendation patterns reflect the public evidence that AI systems can retrieve and synthesize. The concentration of recommendation value around LiveChat, Tidio, and Zendesk is consistent with the strength of their public source footprints. These three brands are likely supported by broad editorial coverage, strong review platform presence, active comparison page visibility, and consistent community discussion across public forums and industry publications.

The source types that appear to shape AI answers in the chatbot category include official brand websites, editorial reviews on platforms such as G2, Capterra, and TrustRadius, comparison pages on software review sites, industry publication articles covering customer service technology, community discussions on Reddit and similar forums, YouTube and video content, and integration-focused content from partner ecosystems. Brands with more retrievable, consistent, and credible material across these source types appear to earn higher recommendation rates.

Drift and Ada, which have strong market recognition but weak recommendation credit, may have public evidence architectures that are misaligned with how AI systems construct recommendations in this category. This does not necessarily mean their content is thin overall. It may mean that the specific source categories AI systems prioritize for chatbot recommendations are underdeveloped, inconsistently structured, or framed in ways that produce mention signals without shortlist signals.

Landbot's net sentiment score of 0.10 is the most notable framing risk in the dataset. When AI systems do encounter source material about Landbot, the aggregate framing appears neutral rather than positive. This pattern may indicate that the available public evidence is sparse, mixed in tone, or concentrated in source types that do not carry strong recommendation weight in this category.

Traditional search visibility contributes to the public evidence layer. Brands that rank well for comparison queries, pricing searches, and category-level keywords create more retrievable source material for AI systems. But search ranking alone does not determine recommendation framing. The quality, authority, and consistency of the source pages that rank also matter. A brand can have strong organic search presence and still produce weak AI recommendation outcomes if the ranked pages do not provide the kind of structured, authoritative, comparison-ready content that AI systems draw on.

What Brands Need to Fix

Weak valid recommendation coverage. Drift, Ada, Landbot, Chatfuel, and Freshchat all earn valid recommendations in fewer than 3% of observations. These brands need to diagnose why AI systems mention them without recommending them. The gap usually traces to source quality, framing consistency, or the absence of credible third-party validation in the specific source categories AI systems prioritize.

Low Top 3 and Top 1 presence. Even brands with moderate recommendation coverage often appear at the bottom of shortlists. Freshchat's average recommended rank of 5.07 and Drift's rank of 4.28 mean these brands are listed but not positioned as primary options. Moving up the ranked shortlist requires stronger source material that supports top-position framing, not simply more mentions.

Poor prompt-cluster coverage at the decision stage. The Pricing and Comparison clusters carry the highest commercial intent multipliers in this benchmark. Brands that underperform in these clusters lose recommendation influence at the moment buyers are closest to making a decision. Addressing pricing transparency, comparison page presence, and feature-level content is a strategic priority for brands with low coverage in these clusters.

Neutral or cautionary framing. Landbot's net sentiment score of 0.10 and Freshchat's score of 0.25 indicate that AI systems are not generating endorsing language when these brands are mentioned. Improving framing quality requires addressing the public source material that shapes how AI systems describe the brand, including review platform profiles, editorial summaries, and community discussions.

Thin or misaligned source footprint. Brands with low recommendation rates may lack the public evidence that AI systems need to construct confident, ranked recommendations. Building a stronger citation architecture means ensuring that official content, third-party reviews, comparison pages, and community discussions collectively support the brand's positioning as a credible shortlist candidate.

Inconsistent entity information. Brands with inconsistent naming, product descriptions, feature claims, or pricing information across public sources create confusion in AI synthesis. Consistent entity information across the full public evidence layer is a prerequisite for reliable AI recommendation framing.

How CiteWorks Studio Helps

1. Map AI recommendation visibility. Track prompts, platforms, company presence, valid recommendations, Top 3 and Top 1 performance, framing quality, and citation sources across the chatbot category to establish a clear competitive baseline.

2. Identify the sources shaping AI answers. Find the editorial, review, forum, directory, owned, and search-visible sources that are influencing how AI systems describe and rank each brand in generated responses.

3. Build the citation architecture plan. Strengthen the public evidence layer so AI systems have more accurate, consistent, and persuasive source material to synthesize when constructing recommendations.

Commercial Takeaway

AI-led discovery is changing where buyer shortlists are formed in the chatbot category. The benchmark shows that recommendation power is concentrating around three brands, while several visible brands are losing the recommendation-stage competition. Brands that appear in AI responses but fail to earn ranked recommendations are at risk of being displaced by competitors with stronger public evidence architectures, even when those competitors have lower raw mention presence.

The gap between mention and recommendation is not fixed. It can widen if brands with weak recommendation credit do not address the source layers that AI systems rely on. It can also close for brands that invest in the citation architecture, framing quality, and prompt-cluster coverage that recommendation-stage visibility requires. Traditional search and source visibility remain important because they contribute to the retrievable public evidence layer, but they are not sufficient on their own.

The total modeled monthly AI opportunity value for this category is $14.27 million. That figure is a modeled benchmark estimate, not revenue. But the pattern it reflects is real: recommendation power is concentrated, the gap between mention and recommendation is wide for several major brands, and the competitive stakes in high-intent prompt clusters are growing as AI adoption increases among buyers evaluating chatbot platforms.

The benchmark shows where the category stands, but every brand has a different profile. Some brands are visible but not recommended. Others are recommended on certain platforms but absent on others. Some have reasonable framing in discovery-stage prompts but weak coverage in pricing and comparison clusters where purchase decisions form.

CiteWorks Studio can show where your brand appears, where competitors are recommended instead, which prompt clusters carry the most commercial risk for your position, which sources are shaping AI answers about your category, and what needs to change to improve recommendation-stage visibility. Request an AI Visibility Audit, an AI Company Discovery Report, or a Citation Architecture Review to understand your brand's standing in AI-led discovery before that standing is set by your competitors.

Benchmark Source

This analysis is based on the 2026 AI Market Discovery Index for Chatbots, published by LLM Authority Index. The benchmark dataset and public industry report were supplied for this category. Visit LLM Authority Index to read the full benchmark report.

/ Take the next step

Want to Understand Your AI Citation Footprint?

We start every engagement with a full audit of how AI systems reference your brand today.

Measurable, Repeatable Programme

Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge

Citation Architecture Review

Identify which high-authority community sources are and aren't working in your favour across AI platforms.

AI Visibility Audit

Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.

/ Learn More

Understanding AI search visibility.

AI search experiences create answers by pulling information from many places online and summarizing it into a single response.

About The Author

Mark Huntley

Mark Huntley

Founder and CEO

Mark Huntley, J.D. is founder of CiteWorks Studio, a strategic advisory focused on visibility, authority, and recommendation presence in AI-shaped search environments. His work centers on embedding-level GEO, vector optimization, and cosine gap engineering — helping brands align their digital presence with the retrieval systems that increasingly shape discovery, interpretation, and choice.

VIEW ALL CASE STUDIESREQUEST AN AI VISIBILITY AUDIT