How AI Search Is Recommending AI Chatbots: Monthly Trends

Mark HuntleyBy Mark HuntleyFounder and CEO
13 minutes read

Key Takeaways

  • WATI remained the category leader in October 2026, rising to 27.9% valid recommendation coverage and extending a two-month upward pattern.
  • Yellow.ai recovered to 12.1% after a September dip, while its presence stayed broadly flat and its gains came from better placement.
  • Interakt was the only brand with a clear multi-month decline, falling to 5.8% and showing the largest drop in raw mention presence.
  • The benchmark found five high-severity factual inconsistencies across four AI platforms, including conflicting claims about WATI’s headquarters and Haptik’s pricing model.

Executive Summary

October 2026 is a quiet month across the AI chatbot benchmark: no tracked brand's valid recommendation coverage moved beyond the range the benchmark treats as normal month-to-month variation, and all eight tracked brands are classified stable. WATI remains the category leader at 27.9% valid recommendation coverage, up from a July baseline of 26.2% and extending an upward pattern that is now in its second consecutive month, from 22.1% in September to 27.9% in October. Yellow.ai holds second place at 12.1%, above its July baseline of 11.1%, following a one-month recovery from 6.9% in September. The gap between WATI and Yellow.ai stands at 15.8 points in October, compared with 15.2 points in September.

Interakt is the only brand showing a multi-month downward pattern, moving down for two consecutive months from 6.0% in September to 5.8% in October and sitting 3.5 points below its July baseline of 9.3%. Its raw mention presence fell from 25.8% in July to 17.9% in October, the largest presence-rate move recorded for any tracked brand across the window, though the move itself falls within the benchmark's normal range for a single reporting period.

Measured against the July baseline, the gap between WATI and Interakt has widened from 16.9 points to 22.1 points, though the gap did not widen in every intervening month. The remaining brands — Gupshup, Gallabox, Engati, Haptik, and Geta.ai — are all classified stable this month, with movements concentrated at small observation counts that do not establish a trend on their own. For Geta.ai specifically, that stability reflects a continued absence of qualified presence rather than steady performance, and the gap to the category leader is addressed directly in its company section below.

Each monthly run begins with prompt-surface observations across the benchmark's defined AI/search surface universe. October 2026 began from 427 total prompts (284 unique questions). Of those, 411 mentioned a tracked brand or competitor; 313 were relevant and 98 were irrelevant. The public metrics use 190 qualified observations that survive both qualification stages. The July baseline began from 397 prompts (308 unique questions) and produced 225 qualified observations; August began from 491 prompts (385 unique questions) and produced 236 qualified observations; September began from 454 prompts (305 unique questions) and produced 217 qualified observations.

AI recommendation trend

valid recommendation coverage, Jul 2026 to Oct 2026

0%10%20%30%40%Jul 2026Aug 2026Sep 2026Oct 2026
  • WATI27.9%
  • Yellow.ai12.1%
  • Interakt5.8%
  • Gupshup4.2%
  • Gallabox1.6%
  • Engati0.5%
  • Geta.ai0.0%
  • Haptik0.0%

Key Findings

Signal

October 2026 finding

Category status

Quiet month; no brand's coverage movement exceeded the benchmark's normal month-to-month variation

Category leader

WATI holds 27.9% valid recommendation coverage, up from 26.2% in the July baseline

Leader gap to second place

15.8 points ahead of Yellow.ai (12.1%), compared with 15.2 points in September

Multi-month pattern

WATI has moved up for two consecutive months; Interakt has moved down for two consecutive months

Largest presence-rate change

Interakt's raw mention presence fell from 25.8% (July) to 17.9% (October)

WATI–Interakt gap

Widened from 16.9 points (July) to 22.1 points (October); did not widen in every month

AI Response Inconsistency Alerts

Questions This Section Answers

  • What factual inconsistencies did AI platforms report about WATI this month?
  • Where did AI platforms conflict on Haptik's pricing model?

The benchmark detected 5 critical or high-severity factual inconsistencies across 4 AI platforms this month. All flagged inconsistencies carry high confidence, and each involves conflicting claims from two different AI platforms answering the same question.

WATI

AI platforms provided conflicting information about WATI's headquarters location four times this month. When asked "Is Wati an Indian company?", ChatGPT stated the company is "Headquartered in Kuala Lumpur, Malaysia," citing a LinkedIn company page, while Copilot stated it is "Headquartered in Hong Kong," citing yespress.io, Venture Intelligence, and a Tracxn legal-entity page. The flagged sources show the conflict explicitly: a LinkedIn excerpt reading "Headquarters Kuala Lumpur" on one side and a yespress.io excerpt reading "Est. 2020, Hong Kong, Remote-First" on the other.

A second conflict on the same question saw Gemini state "Headquartered in Hong Kong," citing a Google Cloud customer case study, while ChatGPT stated "Headquartered in Kuala Lumpur, Malaysia," citing the LinkedIn company page. The flagged Google Cloud excerpt reads "Founded in 2016 and headquartered in Hong Kong," against the LinkedIn "Headquarters Kuala Lumpur" excerpt.

A third conflict again paired Gemini, stating "Headquartered in Hong Kong" and citing the Google Cloud case study, against Perplexity, stating "Headquartered in Kuala Lumpur" and citing LinkedIn, the company's own About Us page, and an Economic Times article. The flagged sources include a LinkedIn post excerpt stating "24% of Hong kong headquartered Wati's site traffic comes from Japan" and an Economic Times excerpt describing "Hong Kong-based customer engagement platform WATI."

A fourth conflict paired Perplexity, stating "Headquartered in Kuala Lumpur," against Copilot, stating "Headquartered in Hong Kong," each citing the source sets described above.

Haptik

AI platforms provided conflicting information about Haptik's pricing model once this month. When asked "What is the difference between yellow AI and Haptik?", Copilot stated "Structured annual pricing (~$5,000/year reported)," citing bestremotetools.com, G2, and Software Advice, while Gemini stated the company "Operates strictly on custom enterprise pricing models, requiring a demo and sales discussion," citing SoftwareWorld, Slashdot, and G2. Responses differed on whether Haptik publishes a fixed annual price or requires a sales conversation to obtain any price at all.

Benchmark Context

Questions This Section Answers

  • How does the qualified benchmark set differ from the raw prompt collection?
  • How have the qualified observation count and recommendation-shaped answer share shifted between July and October?

The report separates the raw collection universe from the qualified analysis set. Brand-level recommendation percentages are calculated within the qualified benchmark set.

Research stage

Jul 2026

Oct 2026

What it represents

Source prompt-surface observations collected

397

427

Total prompt-surface observations gathered

Unique questions

308

284

Distinct questions after deduplication

Brand / competitor mentions

359

411

Prompts mentioning a tracked brand or competitor

Relevant prompts

269

313

Prompts relevant to the vertical

Irrelevant prompts

90

98

Prompts not relevant to the vertical

Qualified benchmark observations

225

190

Public denominator after qualification

Qualified surface breadth

6

6

AI surface families with qualified observations

August 2026 produced 236 qualified observations from 491 source prompts, and September 2026 produced 217 qualified observations from 454 source prompts. The qualified observation count has fallen across the window even as the raw collection grew in August and September, so the October public denominator of 190 is the smallest in the tracked series.

Benchmark-Level Metrics

Metric

Jul 2026

Oct 2026

Change

Qualified observations

225

190

Down 35

Companies tracked

8

8

Flat

Recommendation-shaped answer share

13.8%

40.0%

Up 26.2 points

Valid recommendation shortlist share

38.2%

43.7%

Up 5.5 points

Category leader by coverage

WATI

WATI

Stable

The shortlist share moved down across July, August, and September before reversing in October, from 38.2% to 28.8% to 22.6% to 43.7%. The recommendation-shaped answer share shows a similar reversal, from 13.8% to 18.2% to 19.8% to 40.0%, with the October reading more than double the August and September levels.

AI Recommendation Trend

Questions This Section Answers

  • Which brands lead the October 2026 AI chatbot recommendation rankings?
  • Which brand showed the largest presence-rate change, and does it establish a trend?

WATI Leads a Field the Benchmark Classifies as Stable This Month

Brand

Jul 2026

Oct 2026

Movement

Oct 2026 rank

WATI

26.2%

27.9%

Up 1.7 points

1st

Yellow.ai

11.1%

12.1%

Up 1.0 points

2nd

Interakt

9.3%

5.8%

Down 3.5 points

3rd

Gupshup

4.9%

4.2%

Down 0.7 points

4th

Gallabox

1.3%

1.6%

Up 0.3 points

5th

Engati

0.4%

0.5%

Up 0.1 points

6th

Haptik

1.3%

0.0%

Down 1.3 points

7th

Geta.ai

0.0%

0.0%

Flat

8th

No brand's movement on the primary coverage metric exceeded the benchmark's normal month-to-month variation this period, and the category-level pattern reflects a combination of small movements rather than one dominant shift. WATI and Yellow.ai hold the top two positions, Interakt's presence decline is the period's most notable move among supporting metrics, and the remaining five brands sit within normal range at small observation counts. Geta.ai sits at the bottom of the table at 0.0%, a full 27.9 points behind category leader WATI — the widest gap recorded between any tracked brand and the leader this month.

What Changed This Month

Questions This Section Answers

  • What distinguishes WATI's October gains from Yellow.ai's recovery?
  • Why does Interakt's decline differ from brands whose placement softened but presence held?
  • Which brands are at 0.0% coverage, and what does that signal?

WATI: Leading Position Holds

WATI's valid recommendation coverage stands at 27.9% in October, up from a July baseline of 26.2%, a gain of 1.7 points across the window. Set against the immediately prior month, WATI rose from 22.1% in September to 27.9% in October — a 5.8-point move that sits within the benchmark's normal range for a single month. The brand has now moved up for two consecutive months.

Its top-three rate moved from 17.8% in July to 22.1% in October, and its rank-one rate moved from 13.3% to 15.3%. Raw mention presence rose from 61.8% to 64.2%, meaning WATI appeared in close to two-thirds of qualified observations.

The distinction to notice: WATI's gains touch both visibility and placement. Its October rank-one count stands at 29 top-position mentions out of 190 qualified observations, and its coverage lead over second-place Yellow.ai is 15.8 points in October, compared with 15.2 points in September.

Highest-priority diagnostic: Which prompts placed WATI in the top recommendation position this month, and on which AI surfaces did that placement concentrate?

Yellow.ai: A One-Month Recovery Within Normal Range

Yellow.ai's valid recommendation coverage stands at 12.1% in October, up from a July baseline of 11.1%, a gain of 1.0 point across the window. Against the immediately prior month, Yellow.ai rose from 6.9% in September to 12.1% in October — its largest single-month move in the tracked series, though one the benchmark classifies as normal month-to-month variation rather than a significant shift.

Its top-three rate moved from 7.1% in July to 7.4% in October, after falling to 2.3% in September, and its rank-one rate moved from 2.7% to 2.6%. Raw mention presence slipped slightly, from 25.3% in July to 24.2% in October.

The distinction to notice: Yellow.ai's October coverage gain came with essentially flat presence and a rebound in top-three placement rather than an expansion of visibility. The brand is being placed in recommendation slots at a higher rate while appearing in a similar share of answers.

Highest-priority diagnostic: Which prompts returned Yellow.ai to top-three placement in October after its September dip, and were those placements concentrated on a narrow set of surfaces?

Interakt: A Second Consecutive Month of Decline

Interakt's valid recommendation coverage stands at 5.8% in October, down from a July baseline of 9.3%, a decline of 3.5 points across the window. The brand has now moved down for two consecutive months, from 6.0% in September to 5.8% in October.

Interakt's raw mention presence fell from 25.8% in July to 17.9% in October, a 7.9-point decline that is the largest presence-rate move recorded for any tracked brand this period. Its top-three rate moved from 6.2% to 3.2%, and its rank-one rate moved from 3.6% to 1.1%. Its October coverage represents 11 valid recommendations out of 190 qualified observations, down from 21 in July.

The distinction to notice: Interakt's decline sits in visibility as well as placement. The brand is appearing in fewer answers and holding fewer recommendation positions, which separates its pattern from brands whose presence held steady while placement softened.

Highest-priority diagnostic: Which prompt categories account for the reduction in Interakt's raw mention presence, and which brands appear in those answers instead?

Gupshup: Partial Recovery From a Series Low

Gupshup's valid recommendation coverage stands at 4.2% in October, down from a July baseline of 4.9%, a decline of 0.7 points across the window. Against the immediately prior month, Gupshup rose from 2.3% in September — the low point of its tracked series — to 4.2% in October, a 1.9-point recovery.

Its top-three rate moved from 1.8% in July to 1.6% in October, and its rank-one rate moved from 0.0% to 1.1%. Raw mention presence rose from 10.7% to 11.6%. Its October coverage represents 8 valid recommendations out of 190 qualified observations.

The distinction to notice: Gupshup operates at small counts where a handful of observations moves percentages materially. Its top-three count of 3 and rank-one count of 2 (each out of 190 observations) are directional signals, not an established pattern.

Highest-priority diagnostic: Which prompts produced Gupshup's October rank-one placements, and is that placement durable or concentrated in a small number of queries?

Haptik: Presence Without Recommendation

Haptik's valid recommendation coverage stands at 0.0% in October, down from a July baseline of 1.3%, and unchanged from its September reading of 0.0%. The brand registered 5 raw mentions in October with no valid recommendations.

Its top-three rate moved from 0.9% in July to 0.0% in October, and its rank-one rate moved from 0.4% to 0.0%. Raw mention presence fell from 4.0% to 2.6%.

The distinction to notice: Haptik is still surfaced in answers but is not being placed in recommendation positions. Presence without recommendation is a distinct signal from absence, and the brand's October position sits at that boundary.

Highest-priority diagnostic: Are Haptik's remaining mentions contextual citations rather than recommendation placements, and which prompts retain them?

Engati and Gallabox: Low-Count Positions

Engati's valid recommendation coverage moved from 0.4% in July to 0.5% in October, a 0.1-point gain, recovering from a 0.0% reading in September. Its October coverage represents 1 valid recommendation out of 190 qualified observations, and its raw mention presence stands at 0.5%, down from 1.3% in July.

Gallabox's valid recommendation coverage moved from 1.3% in July to 1.6% in October, a 0.3-point gain, recovering from a 0.0% reading in September. Its October coverage represents 3 valid recommendations out of 190 qualified observations, and its top-three rate moved from 0.9% to 1.1%. Raw mention presence moved from 3.1% to 2.6%.

The distinction to notice: both brands operate at counts where a single observation changes percentages materially. Engati's October reading rests on 1 valid recommendation and Gallabox's on 3 — directional signals rather than established patterns.

Highest-priority diagnostic: For Engati and Gallabox, which prompts produced their October recommendations, and is that activity likely to repeat?

Geta.ai: No Qualified Presence in the Current Period

Geta.ai recorded 0.0% valid recommendation coverage in October, unchanged from 0.0% in July, August, and September. The brand registered no presence at all in the October qualified observation set, compared with a single presence observation in the July baseline — meaning Geta.ai has gone from a single recorded mention to none across the tracked window.

WATI records 27.9% valid recommendation coverage in October, while Geta.ai stands at 0.0%. The more important issue is what that gap means for Geta.ai's recommendation position: unlike Haptik, which still registers raw mentions (2.6% presence) without converting them into recommendations, Geta.ai currently has no presence at all to convert. Every other tracked brand, including the lowest-count brands Engati (1 valid recommendation) and Gallabox (3 valid recommendations), registered at least one qualified observation this month; Geta.ai did not. That places Geta.ai outside the recommendation conversation entirely in the current period, not merely behind on placement within it.

Highest-priority diagnostic: Does any qualified prompt in the current benchmark surface Geta.ai at all, and what would need to change in prompt coverage or AI-surface exposure for the brand to register even a single observation in the next reporting period?

Buyer-Intent Interpretation

Questions This Section Answers

  • Which buyer-intent clusters did the October qualified observations fall into?
  • Why can't the benchmark answer how AI systems position chatbot brands on price or head-to-head comparison?

Buyer-intent cluster

What it captures

Strategic question

Brand Recommendation

Queries where a specific chatbot brand is recommended

Which brands win the recommendation when a direct ask is made?

Pricing & Value

Queries about cost, plans, and value comparison

How do AI systems position brands on price and value?

Multi-Brand Comparison

Queries comparing two or more brands head-to-head

Which brand is favored when options are weighed side by side?

In October 2026, all 190 qualified observations fell into the Brand Recommendation cluster. No observations were recorded in the Pricing & Value or Multi-Brand Comparison clusters, so the public benchmark cannot yet answer how AI systems position brands on price, value, or direct head-to-head comparison. The evidence captures which brand is recommended, not the trade-offs behind that recommendation.

Brand Opportunity Summary

Questions This Section Answers

  • Which brand-specific diagnostics should each tracked chatbot brand investigate first?

Brand

Oct 2026 coverage

Current signal

Highest-priority diagnostic

WATI

27.9%

Leader; up 1.7 points from July and 5.8 points from September

Which prompts produced the October top-position gains?

Yellow.ai

12.1%

Second place; recovered 5.2 points from September

Which prompts returned the brand to top-three placement?

Interakt

5.8%

Third place; down 3.5 points from July with the period's largest presence-rate decline

Which prompt categories drove the fall in raw mention presence?

Gupshup

4.2%

Up 1.9 points from September; small counts

Are the October rank-one placements durable or concentrated?

Gallabox

1.6%

Recovered from 0.0% in September; 3 valid recommendations

Which prompts produced the October recommendations?

Engati

0.5%

Recovered from 0.0% in September; 1 valid recommendation

Which single prompt produced the October recommendation?

Haptik

0.0%

Presence at 2.6% with no valid recommendations

Are remaining mentions contextual rather than recommendations?

Geta.ai

0.0%

No current-period presence; a single presence observation in the July baseline

Does any qualified prompt surface the brand at all?

The benchmark identifies where attention is warranted; a company-level analysis is needed to explain why.

Evidence Behind the Benchmark

Questions This Section Answers

  • What data layers build the aggregate recommendation metrics?
  • Does source presence in AI answers prove causation?

The aggregate metrics are built from prompt-level observations (query, surface, recommendation outcome, rank, sentiment, and citations where exposed). Company-level analysis can go deeper into prompt, competitor, surface, and evidence patterns. Source presence is not automatically treated as proof of causation.

About This Benchmark

This report is part of the CiteWorks Studio AI Visibility Industry Market research program.

Report-Specific Interpretation Notes

  • Small-count movements: brands such as Engati, Gallabox, Gupshup, and Haptik operate at counts where a single observation changes percentages meaningfully. Read these as directional signals, not established trends. Engati's October coverage rests on 1 valid recommendation and Gallabox's on 3.
  • Qualified denominator: all brand-level percentages are calculated against the qualified observation count (190 in October), not the raw collection size (427 prompts). This separates the analysis set from the broader collection universe.
  • Directional analysis: month-over-month movement identifies changes worth investigating. It does not by itself establish the cause of those changes, and this month's category-level pattern was quiet, with no brand exceeding the benchmark's normal month-to-month variation on the primary coverage metric.

Next Step

The Public Benchmark Shows Where a Brand Is Winning or Losing. A Company-Level Audit Shows Why.

Beneath the aggregate percentage sit the questions that matter: which high-intent prompts are won, which competitor takes the recommendation when a brand loses, what attributes AI systems associate with each option, and which external sources shape those answers. The public benchmark identifies the surface-level movements; it does not expose the underlying drivers.

A company-specific AI visibility audit maps those prompt, surface, competitor, ranking, sentiment, and evidence-source patterns into a prioritized visibility strategy.

Request an AI visibility audit

/ Take the next step

Want to Understand Your AI Citation Footprint?

We start every engagement with a full audit of how AI systems reference your brand today.

Measurable, Repeatable Programme

Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge

Citation Architecture Review

Identify which high-authority community sources are and aren't working in your favour across AI platforms.

AI Visibility Audit

Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.

/ Learn More

Understanding AI search visibility.

AI search experiences create answers by pulling information from many places online and summarizing it into a single response.

What Is AI Citation Intelligence?
AI citation intelligence is the process of measuring where AI platforms source their information and how frequently a brand is mentioned or referenced in AI-generated responses. Because LLMs synthesize across multiple sources, the sites and brands that appear repeatedly tend to influence how a topic or company is framed. This practice focuses on identifying which sources shape AI outputs and tracking brand visibility across different AI systems.
What Is Citation Architecture?
Citation architecture describes the set of sources that consistently inform how AI systems talk about a brand, product, or topic. LLMs draw from websites, articles, forums, and public discussion, and the sources they rely on most often become the backbone of their answers. Building strong citation architecture means ensuring that accurate, credible, high authority sources are the ones most likely to shape the way AI tools summarize and recommend a brand.
What Is Generative Engine Optimization?
Generative engine optimization (GEO) is the practice of improving the chances that AI systems use and cite your brand or content when generating answers. While traditional SEO is centered on ranking pages in search results, GEO focuses on how LLMs retrieve, interpret, and combine information when responding to a question. The objective is to strengthen the content and sources AI systems rely on, so your brand is treated as a trusted reference in AI responses.
What Is AI Share of Voice?
AI share of voice tracks how often a brand appears in AI-generated answers compared with competitors in the same category. It reflects visibility across AI platforms such as ChatGPT, Gemini, Claude, and Perplexity. Monitoring AI share of voice helps organizations see whether AI systems consistently include and recommend their brand for key queries or whether competitor brands are showing up more often.

About The Author

Mark Huntley

Mark Huntley

Founder and CEO

Mark Huntley, J.D. is founder of CiteWorks Studio, a strategic advisory focused on visibility, authority, and recommendation presence in AI-shaped search environments. His work centers on embedding-level GEO, vector optimization, and cosine gap engineering — helping brands align their digital presence with the retrieval systems that increasingly shape discovery, interpretation, and choice.

VIEW ALL CASE STUDIESREQUEST AN AI VISIBILITY AUDIT