CiteWorks Studio

How AI Search Is Recommending Language Learning Software: Monthly Trends

Mark HuntleyBy Mark HuntleyFounder and CEO
9 minutes read

Key Takeaways

  • Duolingo led the category in September 2026 with 83.0% valid recommendation coverage, ahead of Babbel at 76.9%.
  • Babbel strengthened its second-place position, while Pimsleur extended a two-month rise to 66.8% coverage.
  • Rosetta Stone recorded the most significant quarter-over-quarter drop, falling 11.9 points from July to 36.6% in September.
  • September month-over-month movement was muted, with no brand exceeding the significance threshold despite smaller gains and declines across the field.

Executive Summary

Duolingo remains the coverage leader in Language Learning Software with valid recommendation coverage of 83.0% in September 2026, holding a 6.1-point edge over Babbel's 76.9%. Babbel holds the clear second position, up 2.8 points from August's 74.1% and up 2.3 points from the July baseline of 74.6%. The top two brands are stable and separated from the rest of the field.

The strongest upward mover this month was Lingoda, rising 2.9 points from 18.2% in August to 21.1% in September, though this remains within normal variation. Pimsleur continues a two-month upward streak, reaching 66.8% from a July baseline of 63.1%. The sharpest decliner was italki, falling 3.7 points from 49.4% in August to 45.7% in September, though the most significant story of the quarter remains Rosetta Stone's 11.9-point decline from 48.5% in July to 36.6% in September.

Against the prior month, the category was quiet: no brand exceeded the significance threshold for month-over-month movement in September. The most consequential baseline-to-current finding is Rosetta Stone's significant decline, which widened its gap with Pimsleur and Babbel in every month of the series.

Each monthly run begins with 800 prompt-surface observations across the benchmark's defined AI/search surface universe. Unique questions totaled 464 in July, 525 in August, and 518 in September. All 800 observations each month mentioned a tracked brand or competitor. Relevant prompts totaled 800 in July, 798 in August, and 792 in September; irrelevant prompts totaled 0, 2, and 8 respectively. The public metrics use 520 qualified observations in July, 522 in August, and 558 in September that survive both qualification stages.

AI recommendation trend

valid recommendation coverage, Jul 2026 to Sep 2026

0%25%50%75%100%Jul 2026Aug 2026Sep 2026
  • Duolingo83.0%
  • Babbel76.9%
  • Pimsleur66.8%
  • italki45.7%
  • Busuu44.6%
  • Rosetta Stone36.6%
  • Memrise31.0%
  • Lingoda21.1%
  • Mango Languages7.5%
  • Mondly6.8%

Key Findings

Signal

September 2026 finding

Category leader

Duolingo leads with 83.0% valid recommendation coverage, holding a 6.1-point edge over Babbel's 76.9%

Largest riser

Lingoda rose 2.9 points to 21.1% coverage from August's 18.2%

Largest decliner

italki fell 3.7 points to 45.7% coverage from August's 49.4%

Significant decliner

Rosetta Stone is down 11.9 points to 36.6% from July's 48.5%, exceeding the significance threshold

Widening gap

Pimsleur's lead over Rosetta Stone widened from 14.6 points in July to 30.2 points in September

Top-3 placement

Babbel leads top-3 placement at 51.8%, Duolingo follows at 50.5%

Benchmark Context

The report separates the raw collection universe from the qualified analysis set. Brand-level recommendation percentages are calculated within the qualified benchmark set.

Research stage

July 2026

September 2026

What it represents

Source prompt-surface observations collected

800

800

Total prompts collected across AI/search surfaces

Unique questions

464

518

Distinct questions after deduplication

Brand / competitor mentions

800

800

Prompts mentioning a tracked brand or competitor

Relevant prompts

800

792

Prompts relevant to the category

Irrelevant prompts

0

8

Prompts not relevant to the category

Qualified benchmark observations

520

558

Public denominator for all brand-level metrics

Qualified surface breadth

6

6

Canonical AI surface families with at least one observation

These qualified counts anchor every brand-level percentage below; the raw collection volume is shown for context only and is not used as a denominator anywhere in the report.

Benchmark-Level Metrics

Metric

July 2026

September 2026

Change

Qualified observations

520

558

Up 38

Companies tracked

10

10

No change

Recommendation-shaped answer share

64.6%

55.4%

Down 9.2 points

Valid recommendation shortlist share

84.6%

81.4%

Down 3.2 points

Category leader by coverage

Duolingo (81.7%)

Duolingo (83.0%)

Stable

August 2026 is the intermediate month in this series, with 522 qualified observations and a recommendation-shaped answer share of 66.9%.

AI Recommendation Trend

The top tier is stable and wide; the competitive battle is for third through sixth

September 2026 shows the same commercial structure as July: Duolingo and Babbel occupy a tier of their own, while the middle of the category continues to reorder. Duolingo's 83.0% and Babbel's 76.9% leave Pimsleur at 66.8% as the clear third-place challenger, 10.1 points behind Babbel and 21.1 points ahead of italki.

Brand

July 2026

September 2026

Movement

September 2026 rank

Duolingo

81.7%

83.0%

Up 1.3 points

1st

Babbel

74.6%

76.9%

Up 2.3 points

2nd

Pimsleur

63.1%

66.8%

Up 3.7 points

3rd

italki

49.0%

45.7%

Down 3.3 points

4th

Busuu

40.8%

44.6%

Up 3.8 points

5th

Rosetta Stone

48.5%

36.6%

Down 11.9 points

6th

Memrise

32.5%

31.0%

Down 1.5 points

7th

Lingoda

20.0%

21.1%

Up 1.1 points

8th

Mango Languages

10.6%

7.5%

Down 3.1 points

9th

Mondly

6.2%

6.8%

Up 0.6 points

10th

The category-level change in September came primarily from the combination of several smaller movements rather than any single significant shift. No brand exceeded the significance threshold for month-over-month movement. The most material movement in the full series is Rosetta Stone's 11.9-point baseline-to-current decline.

What Changed This Month

Rosetta Stone: Significant decliner across the quarter

Rosetta Stone's valid recommendation coverage fell 11.9 points from 48.5% in July to 36.6% in September, a significant baseline-to-current decline exceeding the threshold for normal variation. The steepest single-month movement was 12.3 points between July and August; September was essentially flat against August at 36.6% versus 36.2%.

The decline is accompanied by a drop in raw mention presence from 58.3% to 46.2%, down 12.1 points. Top-3 placement fell from 10.4% to 5.6%, and rank-one presence declined from 1.5% to 0.7%. Rosetta Stone's valid recommendation count fell from 252 in July to 189 in August, then recovered slightly to 204 in September.

The distinction to notice: Rosetta Stone's shift shows up primarily in visibility, not sentiment. When it is mentioned, sentiment remains broadly positive at 0.8 net sentiment, but it appears in far fewer responses than at the start of the quarter. The benchmark cannot distinguish platform behavior from measurement effects.

Highest-priority diagnostic: Which surfaces or prompt categories reduced their mention of Rosetta Stone across the quarter, and which evidence sources are those AI systems relying on instead?

Babbel: Stable leader gaining ground

Babbel's valid recommendation coverage rose 2.8 points from 74.1% in August to 76.9% in September, its highest level in the tracked series and up 2.3 points from the July baseline of 74.6%. This is the first month of an upward streak, though it remains within normal variation.

Raw mention presence rose from 84.6% in July to 85.8% in September, while top-3 placement declined from 56.4% to 51.8% over the same period. Rank-one presence fell from 21.3% to 16.7%. Babbel's valid recommendation count rose from 388 in July to 429 in September.

The distinction to notice: Babbel is being recommended in more responses overall, but its placement quality is shifting. It is appearing in more shortlists while being named first or in the top three less often than at baseline.

Highest-priority diagnostic: Which surfaces are adding Babbel to shortlists at higher rates, and where is it losing top-3 and rank-one placement?

italki: Softening from its mid-quarter peak

italki's valid recommendation coverage fell 3.7 points from 49.4% in August to 45.7% in September, erasing most of its modest mid-quarter gain. Against the July baseline of 49.0%, the decline is 3.3 points, within normal variation but notable for its direction.

Raw mention presence declined from 57.9% in July to 53.4% in September, down 4.5 points. Top-3 placement fell from 16.2% to 12.0%, a decline of 4.2 points. Rank-one presence rose slightly from 1.7% to 2.0%. italki's valid recommendation count held steady at 255 in September versus 258 in August.

The distinction to notice: italki's presence is holding in absolute counts, but its top-3 placement has eroded. It remains in shortlists but is appearing higher in those lists less often.

Highest-priority diagnostic: Which prompt types shifted italki out of top-3 positions, and which competitors are capturing those placements?

Mango Languages: Two-month decline

Mango Languages fell 3.1 points from 10.6% in July to 7.5% in September, marking a two-month downward streak. The decline is within normal variation but consistent across the quarter, with coverage of 8.6% in August between the two endpoints.

Raw mention presence declined from 11.9% to 10.2%, while top-3 placement held steady at 2.1%. Rank-one presence remained flat at 0.2%. Mango Languages' valid recommendation count fell from 55 in July to 45 in August, then to 42 in September.

The distinction to notice: Mango Languages' decline is driven by presence rather than placement. When it appears, it is recommended at roughly the same rate, but it appears in fewer responses across the quarter.

Highest-priority diagnostic: Which niche prompts still include Mango Languages in recommendations, and which surfaces reduced their mentions most?

Buyer-Intent Interpretation

Buyer-intent cluster

What it captures

Strategic question

Brand Recommendation

Prompts seeking a recommended language learning option

Which brands win the direct recommendation when a buyer asks for a single best option?

Pricing & Value

Prompts exploring cost, plans, or value comparisons

How do AI systems position brands on price and value?

Multi-Brand Comparison

Prompts comparing multiple brands head-to-head

Which brands are included and favored in structured comparisons?

In September, all 558 qualified observations fell into the Brand Recommendation cluster. There were no qualified observations in the Pricing & Value or Multi-Brand Comparison clusters, meaning the public benchmark cannot yet answer questions about how AI systems position these brands on price, value, or head-to-head comparison. The current evidence is confined to direct brand recommendation behavior.

Brand Opportunity Summary

Brand

September 2026 coverage

Current signal

Highest-priority diagnostic

Duolingo

83.0%

Stable leader with dominant rank-one presence (33.3%)

Which high-intent prompts does Duolingo win that its closest challengers do not?

Babbel

76.9%

Stable second position with rising coverage

Where does Babbel regain top-3 and rank-one placement against Duolingo?

Pimsleur

66.8%

Two-month upward streak with improving top-3 rate (32.4%)

Which surfaces are driving Pimsleur's sustained top-3 gains?

italki

45.7%

Softening with declining top-3 placement

Which prompt types moved italki out of top-3 positions?

Busuu

44.6%

Stable with coverage above baseline

Which prompts drove Busuu's mid-quarter rise, and can it be sustained?

Rosetta Stone

36.6%

Significant decliner across the quarter

Which surfaces reduced Rosetta Stone's mentions most sharply?

Memrise

31.0%

Stable with slight coverage decline

What distinguishes prompts where Memrise still earns placement?

Lingoda

21.1%

Largest riser this month with coverage above baseline

Which use cases or language prompts are driving Lingoda's September gain?

Mango Languages

7.5%

Two-month decline with steady placement when present

Which niche prompts retain Mango Languages in recommendations?

Mondly

6.8%

Stable with meaningful rise in rank-one rate (0.7%)

Which prompts produce Mondly's small but improving placements?

The benchmark identifies where attention is warranted; a company-level analysis is needed to explain why.

Evidence Behind the Benchmark

The aggregate metrics are built from prompt-level observations (query, surface, recommendation outcome, rank, sentiment, and citations where exposed). Company-level analysis can go deeper into prompt, competitor, surface, and evidence patterns. Source presence is not automatically treated as proof of causation.

About This Benchmark

This report is part of the CiteWorks Studio AI Industry Market Discovery research program.

Report-Specific Interpretation Notes

  • Qualified observations (558 in September) form the public denominator; raw collection volume was 800 prompts. Brand-level percentages are calculated only within the qualified set.
  • Small absolute counts matter. For example, Mango Languages' 7.5% coverage represents 42 qualified recommendations, and Mondly's 6.8% represents 38. Movement in such cases is less reliable than for brands with larger counts.
  • Significant movement is identified when the baseline-to-current change exceeds the threshold for normal month-to-month variation. It indicates a change worth investigating, not a proven cause.
  • This analysis identifies where attention is warranted; it does not establish why a brand gained or lost recommendation credit.

Next Step

The Public Benchmark Shows Where a Brand Is Winning or Losing. A Company-Level Audit Shows Why.

Beneath the aggregate percentages sit the questions that matter: which high-intent prompts a brand wins, which competitor takes the recommendation when a brand loses, what attributes AI systems associate with each option, and which external sources shape those answers. The September data shows Rosetta Stone's significant quarter-long decline and Babbel's steady rise, but it does not show which prompts changed or which evidence sources drove the shifts.

A company-specific AI visibility audit maps those prompt, surface, competitor, ranking, sentiment, and evidence-source patterns into a prioritized visibility strategy.

Request an AI visibility audit

/ Take the next step

Want to Understand Your AI Citation Footprint?

We start every engagement with a full audit of how AI systems reference your brand today.

Measurable, Repeatable Programme

Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge

Citation Architecture Review

Identify which high-authority community sources are and aren't working in your favour across AI platforms.

AI Visibility Audit

Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.

/ Learn More

Understanding AI search visibility.

AI search experiences create answers by pulling information from many places online and summarizing it into a single response.

What Is AI Citation Intelligence?
AI citation intelligence is the process of measuring where AI platforms source their information and how frequently a brand is mentioned or referenced in AI-generated responses. Because LLMs synthesize across multiple sources, the sites and brands that appear repeatedly tend to influence how a topic or company is framed. This practice focuses on identifying which sources shape AI outputs and tracking brand visibility across different AI systems.
What Is Citation Architecture?
Citation architecture describes the set of sources that consistently inform how AI systems talk about a brand, product, or topic. LLMs draw from websites, articles, forums, and public discussion, and the sources they rely on most often become the backbone of their answers. Building strong citation architecture means ensuring that accurate, credible, high authority sources are the ones most likely to shape the way AI tools summarize and recommend a brand.
What Is Generative Engine Optimization?
Generative engine optimization (GEO) is the practice of improving the chances that AI systems use and cite your brand or content when generating answers. While traditional SEO is centered on ranking pages in search results, GEO focuses on how LLMs retrieve, interpret, and combine information when responding to a question. The objective is to strengthen the content and sources AI systems rely on, so your brand is treated as a trusted reference in AI responses.
What Is AI Share of Voice?
AI share of voice tracks how often a brand appears in AI-generated answers compared with competitors in the same category. It reflects visibility across AI platforms such as ChatGPT, Gemini, Claude, and Perplexity. Monitoring AI share of voice helps organizations see whether AI systems consistently include and recommend their brand for key queries or whether competitor brands are showing up more often.

About The Author

Mark Huntley

Mark Huntley

Founder and CEO

Mark Huntley, J.D. is founder of CiteWorks Studio, a strategic advisory focused on visibility, authority, and recommendation presence in AI-shaped search environments. His work centers on embedding-level GEO, vector optimization, and cosine gap engineering — helping brands align their digital presence with the retrieval systems that increasingly shape discovery, interpretation, and choice.

VIEW ALL CASE STUDIESREQUEST AN AI VISIBILITY AUDIT