How AI Search Is Recommending AI Chatbots: Monthly Trends
This analysis is based on the source benchmark: AI Chatbots: 2026 AI Visibility Market Discovery Index
Key Takeaways
- WATI remained the category leader in October 2026, rising to 27.9% valid recommendation coverage and extending a two-month upward pattern.
- Yellow.ai recovered to 12.1% after a September dip, while its presence stayed broadly flat and its gains came from better placement.
- Interakt was the only brand with a clear multi-month decline, falling to 5.8% and showing the largest drop in raw mention presence.
- The benchmark found five high-severity factual inconsistencies across four AI platforms, including conflicting claims about WATI’s headquarters and Haptik’s pricing model.
Executive Summary
October 2026 is a quiet month across the AI chatbot benchmark: no tracked brand's valid recommendation coverage moved beyond the range the benchmark treats as normal month-to-month variation, and all eight tracked brands are classified stable. WATI remains the category leader at 27.9% valid recommendation coverage, up from a July baseline of 26.2% and extending an upward pattern that is now in its second consecutive month, from 22.1% in September to 27.9% in October. Yellow.ai holds second place at 12.1%, above its July baseline of 11.1%, following a one-month recovery from 6.9% in September. The gap between WATI and Yellow.ai stands at 15.8 points in October, compared with 15.2 points in September.
Interakt is the only brand showing a multi-month downward pattern, moving down for two consecutive months from 6.0% in September to 5.8% in October and sitting 3.5 points below its July baseline of 9.3%. Its raw mention presence fell from 25.8% in July to 17.9% in October, the largest presence-rate move recorded for any tracked brand across the window, though the move itself falls within the benchmark's normal range for a single reporting period.
Measured against the July baseline, the gap between WATI and Interakt has widened from 16.9 points to 22.1 points, though the gap did not widen in every intervening month. The remaining brands — Gupshup, Gallabox, Engati, Haptik, and Geta.ai — are all classified stable this month, with movements concentrated at small observation counts that do not establish a trend on their own. For Geta.ai specifically, that stability reflects a continued absence of qualified presence rather than steady performance, and the gap to the category leader is addressed directly in its company section below.
Each monthly run begins with prompt-surface observations across the benchmark's defined AI/search surface universe. October 2026 began from 427 total prompts (284 unique questions). Of those, 411 mentioned a tracked brand or competitor; 313 were relevant and 98 were irrelevant. The public metrics use 190 qualified observations that survive both qualification stages. The July baseline began from 397 prompts (308 unique questions) and produced 225 qualified observations; August began from 491 prompts (385 unique questions) and produced 236 qualified observations; September began from 454 prompts (305 unique questions) and produced 217 qualified observations.
AI recommendation trend
valid recommendation coverage, Jul 2026 to Oct 2026
- WATI27.9%
- Yellow.ai12.1%
- Interakt5.8%
- Gupshup4.2%
- Gallabox1.6%
- Engati0.5%
- Geta.ai0.0%
- Haptik0.0%
Key Findings
Signal | October 2026 finding |
|---|---|
Category status | Quiet month; no brand's coverage movement exceeded the benchmark's normal month-to-month variation |
Category leader | WATI holds 27.9% valid recommendation coverage, up from 26.2% in the July baseline |
Leader gap to second place | 15.8 points ahead of Yellow.ai (12.1%), compared with 15.2 points in September |
Multi-month pattern | WATI has moved up for two consecutive months; Interakt has moved down for two consecutive months |
Largest presence-rate change | Interakt's raw mention presence fell from 25.8% (July) to 17.9% (October) |
WATI–Interakt gap | Widened from 16.9 points (July) to 22.1 points (October); did not widen in every month |
AI Response Inconsistency Alerts
Questions This Section Answers
- What factual inconsistencies did AI platforms report about WATI this month?
- Where did AI platforms conflict on Haptik's pricing model?
The benchmark detected 5 critical or high-severity factual inconsistencies across 4 AI platforms this month. All flagged inconsistencies carry high confidence, and each involves conflicting claims from two different AI platforms answering the same question.
WATI
AI platforms provided conflicting information about WATI's headquarters location four times this month. When asked "Is Wati an Indian company?", ChatGPT stated the company is "Headquartered in Kuala Lumpur, Malaysia," citing a LinkedIn company page, while Copilot stated it is "Headquartered in Hong Kong," citing yespress.io, Venture Intelligence, and a Tracxn legal-entity page. The flagged sources show the conflict explicitly: a LinkedIn excerpt reading "Headquarters Kuala Lumpur" on one side and a yespress.io excerpt reading "Est. 2020, Hong Kong, Remote-First" on the other.
A second conflict on the same question saw Gemini state "Headquartered in Hong Kong," citing a Google Cloud customer case study, while ChatGPT stated "Headquartered in Kuala Lumpur, Malaysia," citing the LinkedIn company page. The flagged Google Cloud excerpt reads "Founded in 2016 and headquartered in Hong Kong," against the LinkedIn "Headquarters Kuala Lumpur" excerpt.
A third conflict again paired Gemini, stating "Headquartered in Hong Kong" and citing the Google Cloud case study, against Perplexity, stating "Headquartered in Kuala Lumpur" and citing LinkedIn, the company's own About Us page, and an Economic Times article. The flagged sources include a LinkedIn post excerpt stating "24% of Hong kong headquartered Wati's site traffic comes from Japan" and an Economic Times excerpt describing "Hong Kong-based customer engagement platform WATI."
A fourth conflict paired Perplexity, stating "Headquartered in Kuala Lumpur," against Copilot, stating "Headquartered in Hong Kong," each citing the source sets described above.
Haptik
AI platforms provided conflicting information about Haptik's pricing model once this month. When asked "What is the difference between yellow AI and Haptik?", Copilot stated "Structured annual pricing (~$5,000/year reported)," citing bestremotetools.com, G2, and Software Advice, while Gemini stated the company "Operates strictly on custom enterprise pricing models, requiring a demo and sales discussion," citing SoftwareWorld, Slashdot, and G2. Responses differed on whether Haptik publishes a fixed annual price or requires a sales conversation to obtain any price at all.
Benchmark Context
Questions This Section Answers
- How does the qualified benchmark set differ from the raw prompt collection?
- How have the qualified observation count and recommendation-shaped answer share shifted between July and October?
The report separates the raw collection universe from the qualified analysis set. Brand-level recommendation percentages are calculated within the qualified benchmark set.
Research stage | Jul 2026 | Oct 2026 | What it represents |
|---|---|---|---|
Source prompt-surface observations collected | 397 | 427 | Total prompt-surface observations gathered |
Unique questions | 308 | 284 | Distinct questions after deduplication |
Brand / competitor mentions | 359 | 411 | Prompts mentioning a tracked brand or competitor |
Relevant prompts | 269 | 313 | Prompts relevant to the vertical |
Irrelevant prompts | 90 | 98 | Prompts not relevant to the vertical |
Qualified benchmark observations | 225 | 190 | Public denominator after qualification |
Qualified surface breadth | 6 | 6 | AI surface families with qualified observations |
August 2026 produced 236 qualified observations from 491 source prompts, and September 2026 produced 217 qualified observations from 454 source prompts. The qualified observation count has fallen across the window even as the raw collection grew in August and September, so the October public denominator of 190 is the smallest in the tracked series.
Benchmark-Level Metrics
Metric | Jul 2026 | Oct 2026 | Change |
|---|---|---|---|
Qualified observations | 225 | 190 | Down 35 |
Companies tracked | 8 | 8 | Flat |
Recommendation-shaped answer share | 13.8% | 40.0% | Up 26.2 points |
Valid recommendation shortlist share | 38.2% | 43.7% | Up 5.5 points |
Category leader by coverage | WATI | WATI | Stable |
The shortlist share moved down across July, August, and September before reversing in October, from 38.2% to 28.8% to 22.6% to 43.7%. The recommendation-shaped answer share shows a similar reversal, from 13.8% to 18.2% to 19.8% to 40.0%, with the October reading more than double the August and September levels.
AI Recommendation Trend
Questions This Section Answers
- Which brands lead the October 2026 AI chatbot recommendation rankings?
- Which brand showed the largest presence-rate change, and does it establish a trend?
WATI Leads a Field the Benchmark Classifies as Stable This Month
Brand | Jul 2026 | Oct 2026 | Movement | Oct 2026 rank |
|---|---|---|---|---|
WATI | 26.2% | 27.9% | Up 1.7 points | 1st |
Yellow.ai | 11.1% | 12.1% | Up 1.0 points | 2nd |
Interakt | 9.3% | 5.8% | Down 3.5 points | 3rd |
Gupshup | 4.9% | 4.2% | Down 0.7 points | 4th |
Gallabox | 1.3% | 1.6% | Up 0.3 points | 5th |
Engati | 0.4% | 0.5% | Up 0.1 points | 6th |
Haptik | 1.3% | 0.0% | Down 1.3 points | 7th |
Geta.ai | 0.0% | 0.0% | Flat | 8th |
No brand's movement on the primary coverage metric exceeded the benchmark's normal month-to-month variation this period, and the category-level pattern reflects a combination of small movements rather than one dominant shift. WATI and Yellow.ai hold the top two positions, Interakt's presence decline is the period's most notable move among supporting metrics, and the remaining five brands sit within normal range at small observation counts. Geta.ai sits at the bottom of the table at 0.0%, a full 27.9 points behind category leader WATI — the widest gap recorded between any tracked brand and the leader this month.
What Changed This Month
Questions This Section Answers
- What distinguishes WATI's October gains from Yellow.ai's recovery?
- Why does Interakt's decline differ from brands whose placement softened but presence held?
- Which brands are at 0.0% coverage, and what does that signal?
WATI: Leading Position Holds
WATI's valid recommendation coverage stands at 27.9% in October, up from a July baseline of 26.2%, a gain of 1.7 points across the window. Set against the immediately prior month, WATI rose from 22.1% in September to 27.9% in October — a 5.8-point move that sits within the benchmark's normal range for a single month. The brand has now moved up for two consecutive months.
Its top-three rate moved from 17.8% in July to 22.1% in October, and its rank-one rate moved from 13.3% to 15.3%. Raw mention presence rose from 61.8% to 64.2%, meaning WATI appeared in close to two-thirds of qualified observations.
The distinction to notice: WATI's gains touch both visibility and placement. Its October rank-one count stands at 29 top-position mentions out of 190 qualified observations, and its coverage lead over second-place Yellow.ai is 15.8 points in October, compared with 15.2 points in September.
Highest-priority diagnostic: Which prompts placed WATI in the top recommendation position this month, and on which AI surfaces did that placement concentrate?
Yellow.ai: A One-Month Recovery Within Normal Range
Yellow.ai's valid recommendation coverage stands at 12.1% in October, up from a July baseline of 11.1%, a gain of 1.0 point across the window. Against the immediately prior month, Yellow.ai rose from 6.9% in September to 12.1% in October — its largest single-month move in the tracked series, though one the benchmark classifies as normal month-to-month variation rather than a significant shift.
Its top-three rate moved from 7.1% in July to 7.4% in October, after falling to 2.3% in September, and its rank-one rate moved from 2.7% to 2.6%. Raw mention presence slipped slightly, from 25.3% in July to 24.2% in October.
The distinction to notice: Yellow.ai's October coverage gain came with essentially flat presence and a rebound in top-three placement rather than an expansion of visibility. The brand is being placed in recommendation slots at a higher rate while appearing in a similar share of answers.
Highest-priority diagnostic: Which prompts returned Yellow.ai to top-three placement in October after its September dip, and were those placements concentrated on a narrow set of surfaces?
Interakt: A Second Consecutive Month of Decline
Interakt's valid recommendation coverage stands at 5.8% in October, down from a July baseline of 9.3%, a decline of 3.5 points across the window. The brand has now moved down for two consecutive months, from 6.0% in September to 5.8% in October.
Interakt's raw mention presence fell from 25.8% in July to 17.9% in October, a 7.9-point decline that is the largest presence-rate move recorded for any tracked brand this period. Its top-three rate moved from 6.2% to 3.2%, and its rank-one rate moved from 3.6% to 1.1%. Its October coverage represents 11 valid recommendations out of 190 qualified observations, down from 21 in July.
The distinction to notice: Interakt's decline sits in visibility as well as placement. The brand is appearing in fewer answers and holding fewer recommendation positions, which separates its pattern from brands whose presence held steady while placement softened.
Highest-priority diagnostic: Which prompt categories account for the reduction in Interakt's raw mention presence, and which brands appear in those answers instead?
Gupshup: Partial Recovery From a Series Low
Gupshup's valid recommendation coverage stands at 4.2% in October, down from a July baseline of 4.9%, a decline of 0.7 points across the window. Against the immediately prior month, Gupshup rose from 2.3% in September — the low point of its tracked series — to 4.2% in October, a 1.9-point recovery.
Its top-three rate moved from 1.8% in July to 1.6% in October, and its rank-one rate moved from 0.0% to 1.1%. Raw mention presence rose from 10.7% to 11.6%. Its October coverage represents 8 valid recommendations out of 190 qualified observations.
The distinction to notice: Gupshup operates at small counts where a handful of observations moves percentages materially. Its top-three count of 3 and rank-one count of 2 (each out of 190 observations) are directional signals, not an established pattern.
Highest-priority diagnostic: Which prompts produced Gupshup's October rank-one placements, and is that placement durable or concentrated in a small number of queries?
Haptik: Presence Without Recommendation
Haptik's valid recommendation coverage stands at 0.0% in October, down from a July baseline of 1.3%, and unchanged from its September reading of 0.0%. The brand registered 5 raw mentions in October with no valid recommendations.
Its top-three rate moved from 0.9% in July to 0.0% in October, and its rank-one rate moved from 0.4% to 0.0%. Raw mention presence fell from 4.0% to 2.6%.
The distinction to notice: Haptik is still surfaced in answers but is not being placed in recommendation positions. Presence without recommendation is a distinct signal from absence, and the brand's October position sits at that boundary.
Highest-priority diagnostic: Are Haptik's remaining mentions contextual citations rather than recommendation placements, and which prompts retain them?
Engati and Gallabox: Low-Count Positions
Engati's valid recommendation coverage moved from 0.4% in July to 0.5% in October, a 0.1-point gain, recovering from a 0.0% reading in September. Its October coverage represents 1 valid recommendation out of 190 qualified observations, and its raw mention presence stands at 0.5%, down from 1.3% in July.
Gallabox's valid recommendation coverage moved from 1.3% in July to 1.6% in October, a 0.3-point gain, recovering from a 0.0% reading in September. Its October coverage represents 3 valid recommendations out of 190 qualified observations, and its top-three rate moved from 0.9% to 1.1%. Raw mention presence moved from 3.1% to 2.6%.
The distinction to notice: both brands operate at counts where a single observation changes percentages materially. Engati's October reading rests on 1 valid recommendation and Gallabox's on 3 — directional signals rather than established patterns.
Highest-priority diagnostic: For Engati and Gallabox, which prompts produced their October recommendations, and is that activity likely to repeat?
Geta.ai: No Qualified Presence in the Current Period
Geta.ai recorded 0.0% valid recommendation coverage in October, unchanged from 0.0% in July, August, and September. The brand registered no presence at all in the October qualified observation set, compared with a single presence observation in the July baseline — meaning Geta.ai has gone from a single recorded mention to none across the tracked window.
WATI records 27.9% valid recommendation coverage in October, while Geta.ai stands at 0.0%. The more important issue is what that gap means for Geta.ai's recommendation position: unlike Haptik, which still registers raw mentions (2.6% presence) without converting them into recommendations, Geta.ai currently has no presence at all to convert. Every other tracked brand, including the lowest-count brands Engati (1 valid recommendation) and Gallabox (3 valid recommendations), registered at least one qualified observation this month; Geta.ai did not. That places Geta.ai outside the recommendation conversation entirely in the current period, not merely behind on placement within it.
Highest-priority diagnostic: Does any qualified prompt in the current benchmark surface Geta.ai at all, and what would need to change in prompt coverage or AI-surface exposure for the brand to register even a single observation in the next reporting period?
Buyer-Intent Interpretation
Questions This Section Answers
- Which buyer-intent clusters did the October qualified observations fall into?
- Why can't the benchmark answer how AI systems position chatbot brands on price or head-to-head comparison?
Buyer-intent cluster | What it captures | Strategic question |
|---|---|---|
Brand Recommendation | Queries where a specific chatbot brand is recommended | Which brands win the recommendation when a direct ask is made? |
Pricing & Value | Queries about cost, plans, and value comparison | How do AI systems position brands on price and value? |
Multi-Brand Comparison | Queries comparing two or more brands head-to-head | Which brand is favored when options are weighed side by side? |
In October 2026, all 190 qualified observations fell into the Brand Recommendation cluster. No observations were recorded in the Pricing & Value or Multi-Brand Comparison clusters, so the public benchmark cannot yet answer how AI systems position brands on price, value, or direct head-to-head comparison. The evidence captures which brand is recommended, not the trade-offs behind that recommendation.
Brand Opportunity Summary
Questions This Section Answers
- Which brand-specific diagnostics should each tracked chatbot brand investigate first?
Brand | Oct 2026 coverage | Current signal | Highest-priority diagnostic |
|---|---|---|---|
WATI | 27.9% | Leader; up 1.7 points from July and 5.8 points from September | Which prompts produced the October top-position gains? |
Yellow.ai | 12.1% | Second place; recovered 5.2 points from September | Which prompts returned the brand to top-three placement? |
Interakt | 5.8% | Third place; down 3.5 points from July with the period's largest presence-rate decline | Which prompt categories drove the fall in raw mention presence? |
Gupshup | 4.2% | Up 1.9 points from September; small counts | Are the October rank-one placements durable or concentrated? |
Gallabox | 1.6% | Recovered from 0.0% in September; 3 valid recommendations | Which prompts produced the October recommendations? |
Engati | 0.5% | Recovered from 0.0% in September; 1 valid recommendation | Which single prompt produced the October recommendation? |
Haptik | 0.0% | Presence at 2.6% with no valid recommendations | Are remaining mentions contextual rather than recommendations? |
Geta.ai | 0.0% | No current-period presence; a single presence observation in the July baseline | Does any qualified prompt surface the brand at all? |
The benchmark identifies where attention is warranted; a company-level analysis is needed to explain why.
Evidence Behind the Benchmark
Questions This Section Answers
- What data layers build the aggregate recommendation metrics?
- Does source presence in AI answers prove causation?
The aggregate metrics are built from prompt-level observations (query, surface, recommendation outcome, rank, sentiment, and citations where exposed). Company-level analysis can go deeper into prompt, competitor, surface, and evidence patterns. Source presence is not automatically treated as proof of causation.
About This Benchmark
This report is part of the CiteWorks Studio AI Visibility Industry Market research program.
- AI Visibility Industry Market Methodology
- AI Visibility Industry Market Metrics
- AI Visibility Industry Market Standards
Report-Specific Interpretation Notes
- Small-count movements: brands such as Engati, Gallabox, Gupshup, and Haptik operate at counts where a single observation changes percentages meaningfully. Read these as directional signals, not established trends. Engati's October coverage rests on 1 valid recommendation and Gallabox's on 3.
- Qualified denominator: all brand-level percentages are calculated against the qualified observation count (190 in October), not the raw collection size (427 prompts). This separates the analysis set from the broader collection universe.
- Directional analysis: month-over-month movement identifies changes worth investigating. It does not by itself establish the cause of those changes, and this month's category-level pattern was quiet, with no brand exceeding the benchmark's normal month-to-month variation on the primary coverage metric.
Next Step
The Public Benchmark Shows Where a Brand Is Winning or Losing. A Company-Level Audit Shows Why.
Beneath the aggregate percentage sit the questions that matter: which high-intent prompts are won, which competitor takes the recommendation when a brand loses, what attributes AI systems associate with each option, and which external sources shape those answers. The public benchmark identifies the surface-level movements; it does not expose the underlying drivers.
A company-specific AI visibility audit maps those prompt, surface, competitor, ranking, sentiment, and evidence-source patterns into a prioritized visibility strategy.
/ Take the next step
Want to Understand Your AI Citation Footprint?
We start every engagement with a full audit of how AI systems reference your brand today.
Measurable, Repeatable Programme
Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge
Citation Architecture Review
Identify which high-authority community sources are and aren't working in your favour across AI platforms.
AI Visibility Audit
Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.
/ Learn More
Understanding AI search visibility.
AI search experiences create answers by pulling information from many places online and summarizing it into a single response.


