CiteWorks Studio

How AI Search Is Recommending AI Chatbots: Monthly Trends

Mark HuntleyBy Mark HuntleyFounder and CEO
10 minutes read

Key Takeaways

  • WATI remained the top recommended chatbot in September 2026 with 22.1% valid recommendation coverage, widening its lead over Yellow.ai to 15.2 points.
  • Yellow.ai and Interakt both lost recommendation coverage despite stable mention presence, indicating weaker conversion from visibility into recommendation slots.
  • Gallabox, Engati, and Haptik fell to 0.0% valid recommendation coverage in September, while Geta.ai showed no qualified presence across the tracked period.
  • The category was broadly stable month to month, but the valid recommendation shortlist share declined from 38.2% in July to 22.6% in September.

Executive Summary

WATI remains the clear leader in AI chatbot recommendations, holding 22.1% valid recommendation coverage in September 2026, down from 26.2% in the July baseline. Yellow.ai follows at 6.9% coverage (down from 11.1% in July), with Interakt third at 6.0% (down from 9.3% in July). The gap between WATI and the next brand stands at 15.2 points in September, compared with 10.1 points in August.

The category recorded no significant movement on the primary coverage metric this month. All eight tracked brands are classified as stable, and the benchmark describes September as a quiet month overall. Within that quiet baseline, the most visible pattern is a narrowing of the field immediately behind the leader: Yellow.ai and Interakt both moved down from their August readings, compressing the spread between second, third, and fourth place even as the gap to the leader widened.

Two brands showed zero valid recommendation coverage this month. Engati moved from 0.4% coverage in both July and August to 0.0% in September, and Gallabox moved from 1.3% in both prior months to 0.0%. Haptik declined from 1.3% in July to 0.4% in August and then to 0.0% in September. Geta.ai remained at 0.0% across all three tracked months, with no recorded valid recommendations in the benchmark's tracked period.

Each monthly run begins with a collection of prompt-surface observations across the benchmark's defined AI/search surface universe. September 2026 began from 454 total prompts (305 unique questions). Of those, 445 mentioned a tracked brand or competitor; 340 were relevant and 105 were irrelevant. The public metrics use 217 qualified observations that survive both qualification stages. The July baseline began from 397 prompts (308 unique questions) and produced 225 qualified observations; August began from 491 prompts (385 unique questions) and produced 236 qualified observations.

AI recommendation trend

valid recommendation coverage, Jul 2026 to Sep 2026

0%10%20%30%40%Jul 2026Aug 2026Sep 2026
  • WATI22.1%
  • Yellow.ai6.9%
  • Interakt6.0%
  • Gupshup2.3%
  • Engati0.0%
  • Gallabox0.0%
  • Geta.ai0.0%
  • Haptik0.0%

Key Findings

Signal

September 2026 finding

Category leader

WATI holds 22.1% valid recommendation coverage, down from 26.2% in July

Leader gap to second place

15.2 points ahead of Yellow.ai (6.9%), up from a 10.1-point gap in August

Largest point movement

Interakt fell from 9.8% in August to 6.0% in September, a 3.8-point move

Two brands at zero

Engati and Gallabox both moved from positive coverage to 0.0% this month

Multi-month pattern

Yellow.ai has declined for two consecutive months, from 11.1% to 10.2% to 6.9%

Category status

Quiet month; no brand's movement exceeded normal month-to-month variation on the primary coverage metric

Benchmark Context

The report separates the raw collection universe from the qualified analysis set. Brand-level recommendation percentages are calculated within the qualified benchmark set.

Research stage

Jul 2026

Sep 2026

What it represents

Source prompt-surface observations collected

397

454

Total prompt-surface observations gathered

Unique questions

308

305

Distinct questions after deduplication

Brand / competitor mentions

359

445

Prompts mentioning a tracked brand or competitor

Relevant prompts

269

340

Prompts relevant to the vertical

Irrelevant prompts

90

105

Prompts not relevant to the vertical

Qualified benchmark observations

225

217

Public denominator after qualification

Qualified surface breadth

6

6

AI surface families with qualified observations

Benchmark-Level Metrics

The metrics below describe the qualified observation set as a whole, before breaking results out by brand.

Metric

Jul 2026

Sep 2026

Change

Qualified observations

225

217

Down 8

Companies tracked

8

8

Flat

Recommendation-shaped answer share

13.8%

19.8%

Up 6.0 points

Valid recommendation shortlist share

38.2%

22.6%

Down 15.6 points

Category leader by coverage

WATI

WATI

Stable

August 2026 sat between the baseline and current months with 236 qualified observations, a recommendation-shaped answer share of 18.2%, and a valid recommendation shortlist share of 28.8%. The shortlist share moved downward across each of the three months, from 38.2% in July to 22.6% in September.

AI Recommendation Trend

The Leader's Position Holds as the Field Behind It Narrows

Brand

Jul 2026

Sep 2026

Movement

Sep 2026 rank

WATI

26.2%

22.1%

Down 4.1 points

1st

Yellow.ai

11.1%

6.9%

Down 4.2 points

2nd

Interakt

9.3%

6.0%

Down 3.3 points

3rd

Gupshup

4.9%

2.3%

Down 2.6 points

4th

Gallabox

1.3%

0.0%

Down 1.3 points

5th

Engati

0.4%

0.0%

Down 0.4 points

6th

Haptik

1.3%

0.0%

Down 1.3 points

7th

Geta.ai

0.0%

0.0%

Flat

8th

No brand's movement on the primary coverage metric exceeded the benchmark's normal month-to-month variation this month. The category-level change reflects the combination of several smaller movements, with the most concentrated pressure on the brands directly behind the leader. WATI's decline of 4.1 points from the July baseline was partly offset by a 1.8-point move up from August, leaving the leader's September position well above the rest of the field while four of the seven other tracked brands registered coverage at or near zero.

What Changed This Month

WATI: Leader by a Widening Margin

WATI's valid recommendation coverage moved from 26.2% in July to 22.1% in September, a net decline across the full window, though the current month sits above the August reading of 20.3%.

WATI's raw mention presence rose to 64.5% in September, up from 61.8% in the July baseline, meaning the brand appeared in nearly two-thirds of qualified observations. Its rank-one recommendation rate fell from 13.3% in July to 5.5% in September, a decline of 7.8 points. Its top-three rate moved from 17.8% to 12.4%.

The distinction to notice: WATI is more visible in September than it was in the July baseline, but it is being placed at the top of recommendations less often. The brand still converts presence into valid recommendations at 22.1%, and with the rest of the field declining, that conversion leaves it further ahead of second place.

Highest-priority diagnostic: Which prompts shifted WATI away from the rank-one position, and which brands are appearing in its place within the top three?

Yellow.ai and Interakt: The Middle Narrows Behind the Leader

Yellow.ai's valid recommendation coverage fell from 11.1% in July to 6.9% in September, a decline of 4.2 points, and has now declined for two consecutive months (11.1% to 10.2% to 6.9%). Interakt fell from 9.3% to 6.0% over the same window, down 3.3 points, including a 3.8-point move from August to September alone.

Both brands maintained their raw mention presence across the period. Yellow.ai's presence was 25.8% in September, nearly flat against its 25.3% July baseline, while Interakt's presence rose from 25.8% to 26.7%. The decline sits in recommendation conversion rather than in visibility: both brands appear in roughly a quarter of qualified observations but are recommended less often.

The distinction to notice: Yellow.ai's top-three rate fell from 7.1% in July to 2.3% in September, while Interakt's top-three rate moved from 6.2% to 3.7%. Both brands remain present in the conversation but are losing placement in the recommendation slots that matter most.

Highest-priority diagnostic: Which competitor is capturing the recommendation slots Yellow.ai and Interakt held in July, and in which prompt categories?

Gupshup: Gradual Decline Across Three Months

Gupshup's valid recommendation coverage fell from 4.9% in July to 2.3% in September, a decline of 2.6 points across the window, with August at 3.4% between the two. The brand has now declined for two consecutive months.

Gupshup's raw mention presence held relatively steady, moving from 10.7% in July to 9.7% in September. Its rank-one rate moved from 0.0% in July to 0.5% in September (1 rank-one mention out of 217 observations), while its top-three rate declined from 1.8% to 0.9%.

The distinction to notice: Gupshup operates at small counts where a handful of observations move percentages materially. Its 2.3% coverage represents 5 valid recommendations in September, down from 11 in July, so the decline should be read as a directional signal rather than an established trend.

Highest-priority diagnostic: Which prompts stopped returning Gupshup in a recommendation position, and is the decline concentrated on particular AI surfaces?

Engati, Gallabox, Haptik: At Zero Coverage This Month

Three brands recorded 0.0% valid recommendation coverage in September. Engati fell from 0.4% in both July and August to 0.0%, with 2 mentions in September and no valid recommendations. Gallabox moved from 1.3% in July and August to 0.0%, with 4 mentions and no valid recommendations. Haptik declined from 1.3% in July to 0.4% in August to 0.0%, with 4 mentions and no valid recommendations in September.

All three brands retain some raw presence, meaning AI systems still mention them, but none of those mentions converted into a recommendation this month. The distinction to notice: presence without recommendation is a distinct signal from absence. Engati, Gallabox, and Haptik are still surfacing in answers but not being placed in recommendation positions.

Highest-priority diagnostic: For each brand, are the remaining mentions contextual citations rather than recommendation placements, and what would be required to convert presence into recommendation?

Geta.ai: Zero Coverage and No Presence Across the Tracked Period

Geta.ai recorded 0.0% valid recommendation coverage in September 2026, unchanged from 0.0% in August and 0.0% in July. Unlike Engati, Gallabox, and Haptik, which retained some raw mention presence even after falling to zero coverage, Geta.ai had no recorded presence in the September qualified observation set, matching the same zero-presence reading in August.

WATI records 22.1% valid recommendation coverage in September, while Geta.ai stands at 0.0%. The more important issue is what that gap means for Geta.ai's recommendation position: every other tracked brand in this benchmark, including the three brands that fell to zero coverage this month, still appeared in at least some qualified observations, while Geta.ai has now shown no presence for two consecutive tracked months. This is not a case of visibility failing to convert into recommendation — it is a case of no visibility at all, against a leader capturing better than one in five qualified observations and a field of seven other brands that all registered at least some presence at some point in the tracked window.

Highest-priority diagnostic: Why does Geta.ai have zero presence in AI-generated answers across two consecutive tracked months, and what would be required to establish even a baseline presence in the prompts where competitors are currently being surfaced?

Buyer-Intent Interpretation

Buyer-intent cluster

What it captures

Strategic question

Brand Recommendation

Queries where a specific chatbot brand is recommended

Which brands win the recommendation when a direct ask is made?

Pricing & Value

Queries about cost, plans, and value comparison

How do AI systems position brands on price and value?

Multi-Brand Comparison

Queries comparing two or more brands head-to-head

Which brand is favored when options are weighed side by side?

In September 2026, all 217 qualified observations fell into the Brand Recommendation cluster. No observations were recorded in the Pricing & Value or Multi-Brand Comparison clusters, meaning the public benchmark cannot yet answer questions about how AI systems position brands on price, value, or direct head-to-head comparison. The evidence captures which brand is recommended, not the reasoning or trade-offs behind that recommendation.

Brand Opportunity Summary

The table below distills the current per-brand signal down to the single highest-priority diagnostic question for each brand.

Brand

Sep 2026 coverage

Current signal

Highest-priority diagnostic

WATI

22.1%

Leader; rank-one rate down 7.8 points from July

Which prompts lost the top recommendation, and to whom?

Yellow.ai

6.9%

Second place; two-month decline from 11.1%

Which competitor is capturing Yellow.ai's lost recommendation slots?

Interakt

6.0%

Third place; down 3.8 points from August despite stable presence

Why did presence stop converting into recommendations?

Gupshup

2.3%

Two-month decline from 4.9%; small counts

Which prompts stopped returning Gupshup as a recommendation?

Gallabox

0.0%

Down from 1.3%; presence remains at 1.8%

Are remaining mentions contextual rather than recommendations?

Engati

0.0%

Down from 0.4%; presence at 0.9%

What would convert Engati's presence into recommendation?

Haptik

0.0%

Three-month decline from 1.3% to zero

Which prompts stopped returning Haptik entirely?

Geta.ai

0.0%

No presence and no recommendations in any tracked month

Why is Geta.ai absent from AI answers entirely?

The benchmark identifies where attention is warranted; a company-level analysis is needed to explain why.

Evidence Behind the Benchmark

The aggregate metrics are built from prompt-level observations (query, surface, recommendation outcome, rank, sentiment, and citations where exposed). Company-level analysis can go deeper into prompt, competitor, surface, and evidence patterns. Source presence is not automatically treated as proof of causation.

About This Benchmark

This report is part of the LLM Authority Index AI Market Discovery research program.

Report-Specific Interpretation Notes

  • Small-count movements: brands such as Engati, Gallabox, Gupshup, and Haptik operate at counts where a single observation changes percentages meaningfully. Read these as directional signals, not established trends.
  • Qualified denominator: all brand-level percentages are calculated against the qualified observation count (217 in September), not the raw collection size (454 prompts). This separates the analysis set from the broader collection universe.
  • Month-over-month movement identifies changes worth investigating. It does not by itself establish the cause of those changes, and this month's category-level pattern was quiet, with no brand exceeding the benchmark's normal month-to-month variation on the primary coverage metric.

Next Step

The Public Benchmark Shows Where a Brand Is Winning or Losing. A Company-Level Audit Shows Why.

Beneath the aggregate percentages sit the questions that matter: which high-intent prompts are won, which competitor takes the recommendation when a brand loses, what attributes AI systems associate with each option, and which external sources shape those answers. The public benchmark identifies the surface-level movements; it does not expose the underlying drivers.

A company-specific AI visibility audit maps those prompt, surface, competitor, ranking, sentiment, and evidence-source patterns into a prioritized visibility strategy.

Request an AI visibility audit

/ Take the next step

Want to Understand Your AI Citation Footprint?

We start every engagement with a full audit of how AI systems reference your brand today.

Measurable, Repeatable Programme

Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge

Citation Architecture Review

Identify which high-authority community sources are and aren't working in your favour across AI platforms.

AI Visibility Audit

Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.

/ Learn More

Understanding AI search visibility.

AI search experiences create answers by pulling information from many places online and summarizing it into a single response.

What Is AI Citation Intelligence?
AI citation intelligence is the process of measuring where AI platforms source their information and how frequently a brand is mentioned or referenced in AI-generated responses. Because LLMs synthesize across multiple sources, the sites and brands that appear repeatedly tend to influence how a topic or company is framed. This practice focuses on identifying which sources shape AI outputs and tracking brand visibility across different AI systems.
What Is Citation Architecture?
Citation architecture describes the set of sources that consistently inform how AI systems talk about a brand, product, or topic. LLMs draw from websites, articles, forums, and public discussion, and the sources they rely on most often become the backbone of their answers. Building strong citation architecture means ensuring that accurate, credible, high authority sources are the ones most likely to shape the way AI tools summarize and recommend a brand.
What Is Generative Engine Optimization?
Generative engine optimization (GEO) is the practice of improving the chances that AI systems use and cite your brand or content when generating answers. While traditional SEO is centered on ranking pages in search results, GEO focuses on how LLMs retrieve, interpret, and combine information when responding to a question. The objective is to strengthen the content and sources AI systems rely on, so your brand is treated as a trusted reference in AI responses.
What Is AI Share of Voice?
AI share of voice tracks how often a brand appears in AI-generated answers compared with competitors in the same category. It reflects visibility across AI platforms such as ChatGPT, Gemini, Claude, and Perplexity. Monitoring AI share of voice helps organizations see whether AI systems consistently include and recommend their brand for key queries or whether competitor brands are showing up more often.

About The Author

Mark Huntley

Mark Huntley

Founder and CEO

Mark Huntley, J.D. is founder of CiteWorks Studio, a strategic advisory focused on visibility, authority, and recommendation presence in AI-shaped search environments. His work centers on embedding-level GEO, vector optimization, and cosine gap engineering — helping brands align their digital presence with the retrieval systems that increasingly shape discovery, interpretation, and choice.

VIEW ALL CASE STUDIESREQUEST AN AI VISIBILITY AUDIT