CiteWorks Studio

Tilly & Wilbur AI Market Strategy Report - Kids and Family Graphic Apparel

Mark HuntleyBy Mark HuntleyFounder and CEO
10 minutes read

Key Takeaways

  • Tilly & Wilbur increased valid recommendation coverage from 28.0% in July 2026 to 48.0% in September 2026, moving into the category lead.
  • The brand’s strongest advantage is rank-one placement, reaching 32.9% versus 5.3% for Pixie and Elf.
  • Google AI Overviews shows the strongest scaled platform performance, while Copilot shows presence without any valid recommendations.
  • The main measurement gap is outside discovery prompts, with no qualified observations yet in pricing, value, or multi-brand comparison queries.

Answer Capsule

Tilly & Wilbur is now the clear recommendation leader in the Kids and Family Graphic Apparel category, with valid recommendation coverage of 48.0% in September 2026, up from 28.0% in July 2026. The brand overtook Pixie and Elf, which fell from 38.7% to 28.9% over the same period, reversing a 10.7-point deficit into a 19.1-point lead. The clearest win is rank-one placement, where Tilly & Wilbur holds 32.9% against Pixie and Elf's 5.3%. The clearest weakness is the absence of qualified observations in pricing, value, and multi-brand comparison clusters, leaving the brand's performance in those buyer-intent classes unmeasured. The biggest opportunity is converting its strong discovery-stage recommendation position into a comparable presence across comparison and price-sensitive prompts.

Who This Report Is For

This report is for marketing, brand, and ecommerce leaders at Tilly & Wilbur and other Kids and Family Graphic Apparel brands tracking how AI and search surfaces recommend brands during buyer discovery.

Report Card

Field

Value

Report type

AI Company Market Strategy Report

Target company

Tilly & Wilbur

Category / market studied

Kids and Family Graphic Apparel

Reporting month

September 2026

AI platforms tracked

5 (ChatGPT, Copilot, Gemini, Google AI Mode, Google AI Overviews)

Public high-intent clusters

1 (Brand Recommendation)

AI observations analyzed

152

Competitors tracked

6

Executive Summary

Tilly & Wilbur holds the strongest recommendation position in the Kids and Family Graphic Apparel category, with valid recommendation coverage of 48.0% in September 2026, up from 28.0% in July 2026. The brand appears in 107 of 152 qualified observations, a raw mention presence rate of 70.4%, and converts that presence into 73 valid recommendations. This is a broad-based gain across presence, coverage, top-three placement, and rank-one placement, not a case of rising mentions without rising recommendations.

The strongest cluster is Best Apparel and Gifts Discovery, which accounts for all 152 qualified observations in the current public series. Within that cluster, Tilly & Wilbur records a top-three rate of 38.2% and a rank-one rate of 32.9%, with an average recommended rank of 1.32. The brand holds 84 positive mentions against 23 neutral mentions and no negative mentions, producing a net sentiment score of 0.785.

The clearest platform signal is Google AI Overviews, where Tilly & Wilbur reaches 61.67% valid recommendation coverage and a 46.67% rank-one rate across 60 observations. ChatGPT shows a perfect 100% rank-one rate, though on a much smaller sample of 2 observations. The clearest gap is the complete absence of qualified observations in pricing, value, and multi-brand comparison clusters, meaning the benchmark does not yet measure how AI systems frame Tilly & Wilbur when price or head-to-head comparison enters the query.

The benchmark classifies Tilly & Wilbur as the sole significant riser from July 2026 to September 2026, with a 20.0-point gain in valid recommendation coverage that exceeds normal variation. The brand moved upward in each of the two months since baseline, from 28.0% in July to 35.6% in August to 48.0% in September.

What Tilly & Wilbur Is Winning

Questions This Section Answers

  • Which recommendation metrics does Tilly & Wilbur lead the category on?
  • How does the brand's rank-one placement compare with its closest competitor?
  • What does the absence of negative mentions signal about how the brand is framed?

Tilly & Wilbur holds the category leadership position in valid recommendation coverage at 48.0%, a 19.1-point lead over Pixie and Elf. The brand is the only significant riser in the benchmark, with a 20.0-point gain from July 2026 to September 2026 that exceeds normal variation.

Rank-one placement is the sharpest win. Tilly & Wilbur holds rank one in 32.9% of qualified observations, up from 12.0% in July 2026, while its closest competitor fell from 29.3% to 5.3%. The brand also improved average recommended rank from 1.62 in July 2026 to 1.32 in September 2026, meaning that when Tilly & Wilbur is recommended, it tends to appear at or near the top of the list.

The brand shows no negative framing in the current series, with 84 positive mentions and 23 neutral mentions across 152 qualified observations. Growth is broad-based across presence, coverage, top-three, and rank-one, which is a stronger signal than a brand that gains mentions without converting them into recommendations.

Where Tilly & Wilbur Has the Clearest AI Visibility Gaps

Questions This Section Answers

  • Why is the lack of observations in pricing and comparison clusters a measurement gap rather than a performance one?
  • Which platform shows Tilly & Wilbur present but never recommended?

The most significant gap is structural rather than competitive. All 152 qualified observations in September 2026 fall into the Brand Recommendation cluster, with zero qualified observations in Pricing and Value or Multi-Brand Comparison. The public benchmark measures which brand AI systems recommend in response to discovery prompts, but it does not yet show how Tilly & Wilbur is framed when buyers ask about cost, value, or explicit head-to-head comparison. This is a measurement gap, not necessarily a performance gap, but it leaves an important part of the buyer journey unverified.

Within the measured cluster, Tilly & Wilbur still loses the recommendation in 52.0% of qualified observations. The brand is present in 70.4% of observations but converts that presence into a valid recommendation only 48.0% of the time. The gap between presence and recommendation suggests there are prompts where Tilly & Wilbur appears in the answer but is not the brand being recommended.

Platform coverage is uneven. Copilot shows Tilly & Wilbur present in only 2 of 9 observations, all neutral, with no valid recommendations. The brand has no presence in Copilot's recommendation output at all, despite appearing as a neutral reference. This is the clearest platform-level gap in the current data.

Biggest Opportunity

Questions This Section Answers

  • Which buyer-intent clusters should Tilly & Wilbur target next to extend its discovery-stage lead?

The clearest opportunity is converting Tilly & Wilbur's strong discovery-stage recommendation position into comparable strength in comparison and price-sensitive prompts. The current benchmark measures only the Brand Recommendation cluster, where Tilly & Wilbur leads. The next frontier is ensuring the brand is also the recommended choice when buyers ask for alternatives, compare multiple brands head to head, or ask about value and pricing. These clusters carry higher buyer-stage multipliers in the underlying methodology, meaning they represent later-stage, higher-intent moments where recommendation decisions are more commercially decisive.

Competitive Landscape

Questions This Section Answers

  • How does Tilly & Wilbur's recommendation position compare with Pixie and Elf's?
  • Where does the competitive gap actually lie between the two leading brands?

Tilly & Wilbur holds the strongest recommendation-stage position in the category, with Pixie and Elf as the only other brand showing meaningful coverage. The remaining four tracked brands hold single-digit coverage or no valid recommendations.

Brand

Top-3 rate

Rank-1 rate

Avg recommended rank

Sentiment

Tilly & Wilbur

38.16%

32.89%

1.32

0.785

Pixie and Elf

21.05%

5.26%

2.11

0.791

Triple One Designs

1.97%

0.66%

2.80

0.857

Three2Tango Tees

1.97%

1.97%

1.00

1.000

Babies2Infinity

0.00%

0.00%

0.000

House of Hide

0.00%

0.00%

0.000

Average recommended rank covers rank-eligible recommendations only.

The table shows Tilly & Wilbur leading on every placement metric while Pixie and Elf holds a nearly identical sentiment score. The competitive gap is not in how the two brands are framed, but in how often each is selected and at what position. Pixie and Elf appears in 44.1% of observations but converts to rank one only 5.3% of the time, while Tilly & Wilbur converts 70.4% presence into a 32.9% rank-one rate.

Prompt Evidence

Google AI Overviews / Best Apparel and Gifts Discovery Prompt: "dad shirt" Result: Tilly & Wilbur appears as a rank-one recommendation, consistent with the platform's 46.67% rank-one rate across the cluster.

Google AI Mode / Best Apparel and Gifts Discovery Prompt: "custom polo shirts australia" Result: Tilly & Wilbur is recommended within the top three, contributing to the platform's 32.91% top-three rate across 79 observations.

ChatGPT / Best Apparel and Gifts Discovery Prompt: "graphic tees women" Result: Tilly & Wilbur holds rank one in both ChatGPT observations, showing a perfect but small-sample recommendation record on this platform.

Copilot / Best Apparel and Gifts Discovery Prompt: "big brother shirt" Result: Tilly & Wilbur appears as a neutral reference in 2 of 9 Copilot observations but receives no valid recommendation, showing presence without recommendation conversion.

What CiteWorks Studio Would Do Next

Phase 1: AI Market Discovery Audit Map which specific prompts drive Tilly & Wilbur's 73 valid recommendations and identify the prompt patterns behind the 20.0-point coverage gain since July 2026.

Phase 2: Recommendation Readiness Plan Close the gap between 70.4% presence and 48.0% recommendation coverage by identifying the prompts where Tilly & Wilbur appears but is not selected.

Phase 3: Owned Answer Layer Buildout Strengthen owned content around comparison, pricing, and value topics so the brand is positioned for the buyer-intent clusters the current benchmark does not yet measure.

Phase 4: Citation / Authority Layer Development Build the external evidence layer that supports Tilly & Wilbur's recommendation strength on Google AI Overviews and Google AI Mode, where the brand already leads.

Phase 5: Monthly AI Visibility and Recommendation Tracking Track rank-one rate and platform-specific coverage monthly to confirm whether the September 2026 leadership position holds or reverts toward normal variation.

Why This Matters

AI presence alone is not enough. Tilly & Wilbur appears in 70.4% of qualified observations, but the commercial difference comes from being the first recommendation 32.9% of the time. Pixie and Elf proves the risk: the brand still appears in 44.1% of observations but has lost the conversion of that presence into top recommendations, and a competitor has taken the position it previously held.

The next move for Tilly & Wilbur is targeted correction of the prompt, page, and citation layers to protect the rank-one position it now holds, extend recommendation strength into comparison and pricing prompts, and close the Copilot gap where the brand is present but never recommended.

Core Metrics

Metric

Value

Mentions

107

Valid recommendations

73

Top 3 recommendation count

58

Rank #1 recommendation count

50

Average recommended rank

1.32

Positive mentions

84

Neutral mentions

23

Negative mentions

0

Raw mention presence rate

70.39%

Valid recommendation coverage

48.03%

Top 3 recommendation rate

38.16%

Rank #1 recommendation rate

32.89%

Net sentiment score

0.785

Strongest cluster by recommendation behavior

Best Apparel and Gifts Discovery

Strongest platform by recommendation behavior

Google AI Overviews

Sentiment Score

Sentiment Score = (positive mentions × 1 + neutral mentions × 0 + negative mentions × -1) / total mentions

For Tilly & Wilbur, this is (84 × 1 + 23 × 0 + 0 × -1) / 107 = 0.785.

This score matters because unclassified mention counts are misleading. Share of voice is a diagnostic metric, not a business KPI. A positive recommendation, neutral reference, cautionary mention, and competitor-displaced mention are not equal, and counting all mentions as wins is bad measurement. Classified sentiment is required before interpreting AI visibility, because a brand can appear frequently while being framed neutrally or negatively, and that framing changes the commercial meaning of its presence.

Sentiment by Platform

Platform

Mentions

Positive

Neutral

Negative

Sentiment Score

Readout

ChatGPT

2

2

0

0

1.00

Strongest public recommendation signal

Copilot

2

0

2

0

0.00

Present as context, not recommendation

Gemini

1

1

0

0

1.00

Positive, but sample too small

Google AI Mode

53

39

14

0

0.736

Present, but not recommendation-led

Google AI Overviews

49

42

7

0

0.857

Strongest public recommendation signal

Methodology

  1. Report orientation: This is a benchmark-based analysis of how AI and search surfaces recommend brands in the Kids and Family Graphic Apparel category, produced from the LLM Authority Index AI Market Discovery Index public dataset. It is not a client implementation case study.
  2. Reporting window: Data covers September 2026, with baseline comparisons to July 2026 and August 2026 where available.
  3. Platforms tracked: ChatGPT, Copilot, Gemini, Google AI Mode, and Google AI Overviews, representing five qualified surface families in September 2026.
  4. Observation count: 152 qualified benchmark observations in September 2026, derived from 197 source prompt-surface observations and 187 unique questions.
  5. Competitor universe: Six tracked brands: Tilly & Wilbur, Pixie and Elf, Triple One Designs, Three2Tango Tees, Babies2Infinity, and House of Hide.
  6. Public clusters used: All 152 qualified observations fell into the Brand Recommendation cluster (Best Apparel and Gifts Discovery). No qualified observations were recorded in Pricing and Value or Multi-Brand Comparison clusters.
  7. Stage 0 role: Raw prompt-surface observations were collected and passed through relevance and qualification stages before becoming the public denominator. Brand-level percentages use the 152 qualified observations, not the 197 raw observations.
  8. Definition of a mention: A brand mention is any appearance of the brand in a qualified observation, regardless of whether the brand was recommended.
  9. Definition of a valid recommendation: A valid recommendation requires the brand to be positively recommended within the answer, with rank-eligible placement. Neutral references, cautionary mentions, and comparison anchors are not counted as valid recommendations.
  10. Limitations: The public benchmark measures the Brand Recommendation cluster only. Pricing, value, and multi-brand comparison behavior is not yet measured. Small-count movement affects brands with fewer than 10 qualified observations. Month-to-month movement identifies changes worth investigating but does not by itself establish cause. Source presence in AI responses is evidence about the information environment, not proof that the source caused the recommendation.
  11. Unique prompt count: The public version does not expose the full unique prompt list. Prompt examples in this report are drawn from the cluster-level examples provided in the company index packet.
  12. Metric interpretation: Raw mention presence, valid recommendation coverage, top-three rate, rank-one rate, and sentiment are separate signals and should not be collapsed into a single visibility metric.

See How AI Is Recommending Your Brand

The public benchmark shows where Tilly & Wilbur stands in AI-generated recommendations for Kids and Family Graphic Apparel. A company-level AI visibility audit goes deeper, mapping the specific prompts, competitor displacement patterns, platform gaps, and external evidence sources behind the aggregate numbers. Understanding why AI systems recommend your brand at rank one in one surface but not another is the first step to protecting and extending the position this benchmark reveals.

/ Take the next step

Want to Understand Your AI Citation Footprint?

We start every engagement with a full audit of how AI systems reference your brand today.

Measurable, Repeatable Programme

Build a durable foundation of credible citations that compounds over time and continues to influence AI answers as new queries emerge

Citation Architecture Review

Identify which high-authority community sources are and aren't working in your favour across AI platforms.

AI Visibility Audit

Understand exactly how LLMs are referencing your brand today and which sources are shaping those answers.

/ Learn More

Understanding AI search visibility.

AI search experiences create answers by pulling information from many places online and summarizing it into a single response.

What Is AI Citation Intelligence?
AI citation intelligence is the process of measuring where AI platforms source their information and how frequently a brand is mentioned or referenced in AI-generated responses. Because LLMs synthesize across multiple sources, the sites and brands that appear repeatedly tend to influence how a topic or company is framed. This practice focuses on identifying which sources shape AI outputs and tracking brand visibility across different AI systems.
What Is Citation Architecture?
Citation architecture describes the set of sources that consistently inform how AI systems talk about a brand, product, or topic. LLMs draw from websites, articles, forums, and public discussion, and the sources they rely on most often become the backbone of their answers. Building strong citation architecture means ensuring that accurate, credible, high authority sources are the ones most likely to shape the way AI tools summarize and recommend a brand.
What Is Generative Engine Optimization?
Generative engine optimization (GEO) is the practice of improving the chances that AI systems use and cite your brand or content when generating answers. While traditional SEO is centered on ranking pages in search results, GEO focuses on how LLMs retrieve, interpret, and combine information when responding to a question. The objective is to strengthen the content and sources AI systems rely on, so your brand is treated as a trusted reference in AI responses.
What Is AI Share of Voice?
AI share of voice tracks how often a brand appears in AI-generated answers compared with competitors in the same category. It reflects visibility across AI platforms such as ChatGPT, Gemini, Claude, and Perplexity. Monitoring AI share of voice helps organizations see whether AI systems consistently include and recommend their brand for key queries or whether competitor brands are showing up more often.

About The Author

Mark Huntley

Mark Huntley

Founder and CEO

Mark Huntley, J.D. is founder of CiteWorks Studio, a strategic advisory focused on visibility, authority, and recommendation presence in AI-shaped search environments. His work centers on embedding-level GEO, vector optimization, and cosine gap engineering — helping brands align their digital presence with the retrieval systems that increasingly shape discovery, interpretation, and choice.

VIEW ALL CASE STUDIESREQUEST AN AI VISIBILITY AUDIT