CiteWorks Studio

How AI Search Is Recommending Language Learning Software

Mark HuntleyBy Mark HuntleyFounder and CEO
13 minutes read

On this report

Key Takeaways

  • Duolingo leads overall AI presence and recommendation volume, but Babbel captures more modeled recommendation value in higher-intent comparison and pricing prompts.
  • Rosetta Stone shows the clearest gap between visibility and recommendation, appearing often in AI answers while rarely being advanced as a top choice.
  • Pricing and plans prompts carry the highest modeled commercial value, making decision-stage recommendation coverage especially important in this market.
  • Several brands, including Mango Languages and Mondly, are mentioned in AI responses but receive no meaningful recommendation credit, suggesting factual presence without shortlist status.

AI platforms are reshaping how language learners discover and evaluate software. When a prospective learner asks ChatGPT, Gemini, or Perplexity for the best app to learn Spanish or a comparison of Duolingo versus Babbel, the AI-generated response often determines which brands enter the buyer's consideration set before a single website is visited. Traditional search rankings and ad impressions no longer capture where buyer shortlists are actually formed in this category.

The July 2026 LLM Authority Index benchmark for language learning software reveals a market where appearing in AI answers is not the same as being recommended. Duolingo dominates raw visibility at 38.1% of all AI responses, yet the gap between appearing and being advanced as a top choice is significant across the category. Several well-known brands appear frequently in AI answers without earning consistent recommendation credit. CiteWorks Studio interprets this benchmark to show which brands are winning recommendation-stage visibility, which are visible but not recommended, and what the commercial consequences are for companies competing in this space.

Methodology

  1. Market studied: Language Learning Software, including mobile apps, web platforms, and subscription-based learning services targeting consumer and self-directed learners.
  2. Brands/entities included: Duolingo, Babbel, Busuu, italki, Lingoda, Mango Languages, Memrise, Mondly, Pimsleur, and Rosetta Stone. This is not a complete market census. Other language learning platforms not included in the benchmark universe may be present in the broader market.
  3. Data collection date/window: July 2026, snapshot-based measurement. All metrics reflect a single point-in-time observation window.
  4. AI platforms tested: ChatGPT, Gemini, Copilot, Perplexity, Google AI Mode, and Google AI Overviews.
  5. Number of prompts tested: Prompt count was not provided in the available dataset. The analysis is based on 845 observations distributed across three buyer-stage prompt clusters.
  6. Prompt categories: Three clusters were measured: Consideration (Best Language Learning Apps and Platforms), Evaluation (Language Learning App Comparisons), and Decision (Language Learning App Pricing and Plans).
  7. Definition of a mention: A mention means the company appeared in an AI-generated response, regardless of framing, sentiment, or ranking position.
  8. Definition of a valid recommendation: A valid recommendation is a positive, shortlist-quality recommendation or ranked recommendation that earns recommendation credit. Neutral references, cautionary mentions, comparison anchors, and factual listings do not qualify as valid recommendations. This distinction is central to the benchmark methodology.
  9. Ranking/scoring metrics used: Valid recommendation coverage, top-three rate, rank-one rate, average rank, net sentiment score, monthly AI authority value, monthly AI recommendation value, monthly AI visibility assist value, and captured share of AI opportunity.
  10. Limitations: This is a point-in-time benchmark. AI outputs can change across platforms and prompt variations. Modeled values are estimates based on prompt volume, commercial intent, buyer stage multipliers, and platform weights. They are not revenue, pipeline, or booked demand. This report is not a full audit or a full market census.

Key Findings

Duolingo leads recommendation volume but Babbel leads on recommendation value. Duolingo earns 38 valid recommendations across 845 observations, producing a 4.5% recommendation coverage rate and an average rank of 1.29 when recommended. Babbel earns 19 valid recommendations with a 2.25% coverage rate and an average rank of 1.5. Despite lower recommendation volume, Babbel captures $5,131 in modeled monthly AI recommendation value compared to Duolingo's $2,241. The benchmark analysis suggests Babbel is winning recommendation credit in higher-value buyer contexts, particularly comparison and decision-stage prompts where AI answers carry more commercial weight.

Rosetta Stone carries the sharpest visibility-to-recommendation gap in the dataset. Rosetta Stone appears in 10.8% of all AI responses but converts those appearances into only 9 valid recommendations, a 1.07% recommendation coverage rate. On ChatGPT specifically, Rosetta Stone appears in 14.1% of responses but earns zero valid recommendations. The brand is present in AI conversations as a recognized name but is not being advanced as a top choice. This pattern, high presence with low recommendation conversion, represents the clearest cautionary case in the benchmark.

The Decision-stage pricing cluster concentrates the highest commercial value. The Language Learning App Pricing and Plans cluster carries a 1.5x buyer stage multiplier and represents $340,260 in total modeled monthly opportunity. Duolingo leads with 25 valid recommendations and a 5.6% coverage rate in this cluster. Babbel follows with 7. Pimsleur earns 5, and Rosetta Stone earns 4. Several brands in the dataset earn no recommendations at all in this cluster, meaning they are absent at the moment when buyer intent is highest and AI answers are most likely to shape final decisions.

Several brands have measurable presence but zero modeled recommendation value. Mango Languages and Mondly appear in AI responses but earn no valid recommendations and no modeled recommendation value. Busuu appears in 8.2% of observations and earns 6 valid recommendations but captures zero recommendation value in the model. italki appears in 10.7% of observations with 4 valid recommendations and $8.50 in modeled recommendation value. These brands are being named in AI answers, but the evidence suggests they are functioning as factual references rather than buyer recommendations.

Platform performance varies significantly and creates uneven exposure. Duolingo achieves its strongest recommendation coverage on Google AI Mode at 6.29% and on Gemini at 4.93%. Babbel performs best on Google AI Mode at 4.2% and on Google AI Overviews at 2.1%. Rosetta Stone earns recommendation credit only on Google AI Mode and Google AI Overviews, with zero valid recommendations on ChatGPT and Perplexity. Brands that perform well on one platform may be effectively absent from buyer consideration on others, a risk that aggregate visibility metrics do not capture.

What Changed in the Market

Language learners are no longer only moving from Google search results to brand websites. They are increasingly asking AI systems to compare providers, explain teaching methods, summarize pricing, surface alternatives, and recommend shortlists. This shift relocates where buyer consideration begins. A learner who receives a ranked AI answer may never see a brand that ranks well in organic search but is absent from that AI shortlist.

For language learning software, a category where trust, pedagogy, and price transparency are critical purchase factors, AI systems are becoming the first filter in the evaluation process. When a learner asks which app is best for complete beginners or how Babbel compares to Rosetta Stone in price, the AI response often determines which two or three brands are evaluated in depth and which are excluded from consideration entirely.

The benchmark shows that AI systems are selective in ways that raw visibility does not reveal. Across 845 observations, only 38 of Duolingo's appearances resulted in a valid recommendation. For Babbel, 19 of its appearances earned recommendation credit. For most other brands in the dataset, the conversion ratio is lower still. This selectivity means brand recognition alone does not guarantee shortlist placement. AI systems appear to reward brands with structured, authoritative, and positively framed public evidence, not simply brands that are well known.

The buyer stage at which a prompt is asked also changes the competitive picture. Consideration-stage prompts are more likely to surface broad lists that include more brands. Evaluation and decision-stage prompts, which ask AI systems to compare or price specific options, are more concentrated. The brands that perform well across all three stages are the ones most likely to capture buyers regardless of where in the journey an AI conversation begins.

What the Benchmark Found

Duolingo is the visibility leader and the recommendation volume leader in the benchmark. It appears in 38.1% of all 845 AI responses and earns 38 valid recommendations, producing a 4.5% recommendation coverage rate. Duolingo achieves a 4.0% rank-one rate, meaning it is the first recommendation in approximately 34 observations. Its average rank of 1.29 confirms that when Duolingo receives a valid recommendation, it is almost always positioned at the top of the list. The modeled monthly AI authority value is $16,893. Duolingo wins on volume, top-of-list placement, and decision-stage concentration, but its modeled recommendation value per appearance is lower than Babbel's, which points to Babbel outperforming in the contexts where AI recommendations carry more buyer weight.

Babbel is the recommendation value leader and the strongest challenger to Duolingo at the commercial level. It appears in 18.8% of AI responses and earns 19 valid recommendations with a 2.25% recommendation coverage rate. Babbel achieves a 1.4% rank-one rate with 12 first-place recommendations. Its average rank of 1.5 is strong. Babbel captures $5,131 in modeled monthly AI recommendation value, more than double Duolingo's $2,241, indicating that Babbel's recommendation credit concentrates in higher-value prompt contexts. The modeled monthly AI authority value is $12,374. Babbel's profile is that of a brand that earns fewer mentions but converts them more efficiently.

Pimsleur is a solid but limited performer. It appears in 10.1% of AI responses and earns 14 valid recommendations with a 1.66% recommendation coverage rate. Its average rank of 2.86 is lower than the top two, and it has 4 rank-one placements. Pimsleur captures $2,215 in modeled monthly recommendation value. The modeled monthly AI authority value is $4,107. Pimsleur has consistent recommendation coverage but is typically positioned below Duolingo and Babbel, which limits its top-of-list commercial influence in competitive prompts.

Rosetta Stone represents the most commercially significant warning in the dataset. It appears in 10.8% of AI responses but earns only 9 valid recommendations with a 1.07% recommendation coverage rate. Its average rank is 2.63 when recommended, and it has 4 rank-one placements. Rosetta Stone captures $787 in modeled monthly recommendation value. The modeled monthly AI authority value is $3,839. On ChatGPT, Rosetta Stone appears in 14.1% of responses but earns zero valid recommendations. This is the definition of visible but under-recommended: the brand is present in AI conversations but is not being advanced as a buyer choice on the platform with the largest share of AI-led discovery activity.

Memrise is a modest but relatively efficient performer. It appears in 5.9% of AI responses and earns 6 valid recommendations with a 0.71% recommendation coverage rate. Its average rank of 2.0 is competitive relative to its presence rate. Memrise captures $378 in modeled monthly recommendation value. The modeled monthly AI authority value is $1,654. Memrise demonstrates that smaller presence shares can still produce proportionate recommendation efficiency when framing and source support are aligned.

Busuu, italki, Lingoda, Mondly, and Mango Languages form the exposed group in the benchmark. These brands appear in AI responses at varying rates but rarely earn meaningful recommendation credit. Busuu appears in 8.2% of observations with 6 valid recommendations and zero modeled recommendation value. italki appears in 10.7% of observations with 4 valid recommendations and $8.50 in modeled recommendation value. Lingoda appears in 10.1% of observations with 2 valid recommendations and $4.68 in modeled recommendation value. Mondly and Mango Languages register no valid recommendations. These brands are being named in AI answers, but the evidence suggests they are functioning primarily as factual references rather than shortlist-quality recommendations.

Why Visibility Is Not Enough

A brand can appear in AI answers and still fail to win the buyer shortlist. This is the central commercial distinction the benchmark surfaces, and it is the reason raw mention counts are an insufficient measure of AI search performance.

Raw mention presence measures how often a company is named in an AI response. Valid recommendation coverage measures how often a company is actually recommended or shortlisted in a meaningful, positive way. These are not the same signal, and conflating them produces misleading conclusions about competitive positioning. A brand named as a historical product, a comparison anchor, or a neutral listing is not receiving the same commercial benefit as a brand positioned first in a ranked shortlist.

Top-three placement and rank-one placement matter more commercially than total mention counts. A brand that appears first in an AI recommendation list captures buyer attention at a qualitatively different level than a brand that appears fourth or fifth. Duolingo's rank-one rate of 4.0% reflects meaningful top-of-list control. Rosetta Stone's rank-one rate of 0.47% across 845 observations means it is rarely the first name a buyer receives, even when it is mentioned.

Neutral and cautionary framing does not equal a positive recommendation. Busuu carries a 7.2% neutral visibility rate. italki carries a 10.2% neutral rate. Lingoda carries a 9.7% neutral rate. These brands are being named, but the framing quality is not producing recommendation credit. AI systems appear to distinguish between naming a brand as context and advancing it as a choice.

Modeled benchmark value is not revenue. The monthly AI authority values and recommendation values in this analysis are modeled estimates. They are directional indicators of competitive positioning based on prompt volume, commercial intent multipliers, buyer stage weights, and platform factors. They should be read as signals of where recommendation-stage advantage is concentrating, not as revenue projections or pipeline forecasts.

The Citation Layer

AI systems build recommendations from publicly available evidence, not from brand recognition alone. The sources that appear to shape AI answers in the language learning software category include official brand websites, editorial reviews, app store content, comparison articles, learning methodology explainers, pricing pages, forum discussions, and community-generated content across platforms such as Reddit and language learning communities.

Brands with structured, authoritative, and consistently positive public content are more likely to be retrieved reliably and recommended accurately. Duolingo's scale generates substantial review content, community discussions, press coverage, and comparison articles across a wide range of sources, giving AI systems extensive material to synthesize when forming recommendations. Babbel's structured pricing and subscription content appear to make it particularly retrievable in decision-stage prompts, which may help explain its concentration of modeled recommendation value.

Brands with thinner source footprints, inconsistent entity information, or limited third-party validation face a structural challenge in AI-led discovery. Lingoda, Mondly, and Mango Languages appear in AI responses at low rates and earn minimal recommendation credit. Whether this reflects weaker source footprints, less structured pricing and comparison content, or lower third-party coverage is a question that a full source audit would need to answer, but the pattern is consistent with brands that are less retrievable in competitive prompt contexts.

Sentiment signals across the public evidence layer also appear to influence recommendation placement. Duolingo and Babbel maintain positive net sentiment scores above 0.15 in the dataset. Lingoda registers 0.035, and italki registers 0.044. Brands with lower net sentiment in the public evidence layer appear less likely to be advanced as top recommendations, even when they appear in AI responses at comparable rates.

The search-visible evidence layer that supports AI retrievability includes the kinds of pages that appear in traditional organic search: comparison articles, methodology reviews, expert assessments, pricing summaries, and authoritative third-party coverage. Brands that have built strong search-visible presence across these page types are creating source material that AI systems can access when forming answers. This does not guarantee AI recommendation credit, but a stronger and more consistently positive source footprint gives AI systems more accurate and persuasive material to draw from.

What Brands Need to Fix

Weak valid recommendation coverage. Rosetta Stone converts fewer than 1 in 10 AI appearances into a recommendation. Busuu converts roughly 1 in 14. Brands in this position need to understand why AI systems are naming them without recommending them, which typically points to framing quality, source authority, or content structure gaps.

Low top-three and rank-one presence. Pimsleur averages a rank of 2.86 when it earns recommendation credit. Rosetta Stone earns only 4 rank-one placements across 845 observations. Brands that appear in AI answers but not at the top of ranked lists are losing buyer attention to competitors at the critical shortlist moment.

Absent or thin decision-stage coverage. The pricing and plans cluster carries the highest buyer intent multiplier in the benchmark and the highest total modeled opportunity. Brands that are absent from this cluster, or that earn zero recommendation credit in it, are missing the moments when AI answers are most likely to shape final decisions.

Neutral or non-endorsing framing. Busuu, italki, and Lingoda carry high neutral visibility rates and very low positive visibility. These brands are being named in AI answers but not endorsed. The quality of framing in AI-generated responses is a distinct signal from mention frequency, and it requires attention to the evidence layer that AI systems draw from when forming characterizations.

Thin or inconsistent source footprint. Brands with limited editorial coverage, weak review ecosystems, unstructured pricing information, or inconsistent entity definitions across the public web face a structural disadvantage in AI-led discovery. Building a stronger, more consistent, and more authoritative source footprint is a prerequisite for improving recommendation-stage visibility.

Platform-specific gaps. Rosetta Stone earns zero valid recommendations on ChatGPT and Perplexity. Other brands have similar platform-specific blind spots. Brands that depend on recommendation visibility from a single platform are exposed to significant risk if that platform's behavior changes or if buyers shift their AI usage patterns.

How CiteWorks Studio Helps

  1. Map AI recommendation visibility. Track prompts, platforms, company presence, valid recommendations, top-three and rank-one performance, framing, and citation sources across the language learning software category and within your specific competitive set.
  2. Identify the sources shaping AI answers. Find the editorial, review, forum, directory, owned, and search-visible sources that appear to influence brand framing and recommendation eligibility in AI-generated responses.
  3. Build the citation architecture plan. Strengthen the public evidence layer so AI systems have more accurate, consistent, and persuasive source material to synthesize when generating recommendations across buyer-stage prompt clusters.

Commercial Takeaway

AI-led discovery is changing where buyer shortlists are formed in the language learning software market. The July 2026 LLM Authority Index benchmark shows that Duolingo and Babbel have established a two-tier leadership structure. Duolingo wins on volume, top-of-list placement, and decision-stage concentration. Babbel wins on modeled recommendation value and high-intent buyer contexts. The distance between these two leaders and the rest of the competitive set is significant and is likely to compound as AI-led discovery becomes a larger share of how buyers find and evaluate language learning options.

Brands below this tier face a choice. Several well-known and well-funded brands in this benchmark, including Rosetta Stone, Busuu, italki, and Lingoda, are present in AI answers at meaningful rates but are not consistently advanced as top choices. Competitors are intercepting demand in high-intent prompt clusters while these brands remain visible but under-recommended. The risk is not just losing a mention. It is being excluded from the shortlist at the moment a buyer is ready to decide.

Traditional search and source visibility still matter because they contribute to the public evidence layer that AI systems draw from when forming answers. The opportunity for brands in this category is not to chase AI mentions but to build the recommendation-stage visibility that turns appearances into shortlist placement. The brands that will compete effectively in AI-led discovery are those that treat citation architecture, source footprint quality, and framing consistency as strategic priorities alongside their traditional search and content investments.

The benchmark shows the market shape. A company-specific analysis shows where your brand stands, where competitors are being recommended instead, and what the evidence gap looks like from an AI system's perspective.

CiteWorks Studio can show where your brand appears across AI platforms, which prompts carry the most commercial risk for your category, which sources appear to be shaping competitor recommendations, and what needs to change to improve your recommendation-stage visibility.

Request an AI Visibility Audit, an AI Company Discovery Report, or a Citation Architecture Review to understand your brand's current position in AI-generated language learning recommendations.

Benchmark Source

This analysis is based on the July 2026 AI Market Discovery Index for Language Learning Software, published by LLM Authority Index. Read the full benchmark report at the LLM Authority Index website.

/ Take the next step

Want to Understand Your AI Citation Footprint?

We start every engagement with a full audit of how AI systems reference your brand today.

Measurable, Repeatable Programme

Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge

Citation Architecture Review

Identify which high-authority community sources are and aren't working in your favour across AI platforms.

AI Visibility Audit

Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.

/ Learn More

Understanding AI search visibility.

AI search experiences create answers by pulling information from many places online and summarizing it into a single response.

About The Author

Mark Huntley

Mark Huntley

Founder and CEO

Mark Huntley, J.D. is founder of CiteWorks Studio, a strategic advisory focused on visibility, authority, and recommendation presence in AI-shaped search environments. His work centers on embedding-level GEO, vector optimization, and cosine gap engineering — helping brands align their digital presence with the retrieval systems that increasingly shape discovery, interpretation, and choice.

VIEW ALL CASE STUDIESREQUEST AN AI VISIBILITY AUDIT