Skip to content

AI Market Discovery Methodology

Overview

CiteWorks Studio's AI Market Discovery research is designed to measure how brands appear in commercially relevant AI-assisted discovery, evaluation, comparison, and recommendation moments.

The methodology deliberately separates the raw collection universe from the qualified benchmark set used for public brand-level metrics. This prevents incidental brand mentions, citation-only appearances, or low-relevance answers from receiving the same weight as genuine commercial recommendation moments.

Individual category reports contain their own category-specific data and findings. This page is the authoritative shared reference for the research process, so the same methodology does not need to be duplicated across hundreds of report pages.

> In short: collect broadly, qualify consistently, classify commercial intent, code recommendation outcomes, and publish the resulting metrics with transparent denominators.


1. Build the Category Query Universe

Each category begins with a set of high-demand queries selected using Ahrefs search-volume data as a demand proxy.

The demand signal is used to prioritize questions that represent meaningful consumer interest. It should not be interpreted as measured prompt volume inside ChatGPT, Gemini, Copilot, Google AI Mode, Google AI Overviews, or any other AI system.

For each category, the research pipeline should retain at minimum:

  • exact query text;
  • search-demand signal used for selection;
  • category assignment;
  • collection month;
  • geography or market, where relevant;
  • language;
  • any query-clustering or normalization metadata; and
  • whether the query belongs to the fixed recurring panel or a refreshed discovery panel, if that distinction is used.

Why search-demand-based selection is used

AI platforms do not generally expose reliable public prompt-volume datasets comparable to search keyword tools. Search-demand data therefore functions as a practical external proxy for prioritizing high-interest consumer questions while maintaining a repeatable benchmark design.


2. Run Queries Across the AI/Search Surface Universe

The selected query set is evaluated across the AI/search surfaces included in the benchmark.

A prompt-surface observation is one source query tested on one AI/search surface together with the resulting answer and collection metadata.

For the current category-report implementation, a monthly run typically begins with approximately 800 prompt-surface observations before qualification.

The benchmark may include environments such as:

  • ChatGPT;
  • Gemini;
  • Microsoft Copilot;
  • Google AI Mode;
  • Google AI Overviews; and
  • other AI/search surfaces defined in the active benchmark configuration.

The authoritative surface list should be maintained centrally. If a surface is added, removed, or materially changes collection behavior, the methodology version should be updated rather than silently changing the longitudinal series.

Collection metadata

Where available, the underlying dataset should preserve:

  • surface name;
  • model/version when exposed;
  • exact collection date and time;
  • market or geography;
  • language;
  • signed-in/logged-out state;
  • personalization controls;
  • prompt text;
  • raw answer text;
  • citations or linked sources; and
  • any other environment setting that could materially affect reproducibility.

3. Apply Tracked-Brand Relevance Qualification

The raw collection universe intentionally contains more observations than the final benchmark requires. The first qualification stage determines whether an answer is relevant to the tracked competitive set.

An observation proceeds when at least one tracked brand is meaningfully surfaced in the answer context under the benchmark's inclusion rules.

Examples of signals that may qualify include:

  • a tracked brand being recommended;
  • a tracked brand being explicitly compared with alternatives;
  • a tracked brand being discussed in a pricing or value context; or
  • another explicit commercial mention that satisfies the benchmark rules.

Examples that should not automatically qualify include:

  • a tracked company appearing only in an unrelated citation;
  • a domain being cited without the brand being surfaced in the relevant answer context;
  • incidental references that do not contribute to the consumer decision being analyzed; or
  • answers with no meaningful connection to the tracked competitive set.

The objective is not to maximize the final sample size. It is to ensure that downstream metrics describe commercially relevant brand discovery, not arbitrary brand presence.


4. Classify Commercial Buyer Intent

Brand-relevant observations are then evaluated for buyer intent. To enter the public AI Market Discovery benchmark, an observation must fit one of the defined commercial-intent clusters.

Brand Recommendation

Discovery questions in which the user asks AI to suggest brands, products, stores, or options that fit a need.

Typical questions include:

  • Which brands enter the consideration set?
  • Which brands appear in the top three?
  • Which brand is recommended first?
  • What attributes does AI use when explaining the recommendation?

Pricing & Value

Questions where price, affordability, discounts, value, or budget suitability materially affects the recommendation or evaluation.

Typical questions include:

  • Which brands remain visible when a budget constraint is introduced?
  • Which brands are framed as value-oriented or premium?
  • Does recommendation rank change when price becomes explicit?

Multi-Brand Comparison

Questions in which AI compares two or more brands, products, or alternatives directly.

Typical questions include:

  • Which brand is preferred in a head-to-head comparison?
  • Which attributes are used to differentiate competitors?
  • Which alternative is presented as stronger for a particular use case?

Observations that pass brand relevance but do not fit the benchmark's commercial buyer-intent framework are excluded from the final public denominator.


5. Code Recommendation Outcomes

For each qualified observation, the benchmark records the variables required to describe both whether a brand appeared and how strongly it was positioned.

Depending on the observation and available data, coded fields can include:

  • valid recommendation status;
  • recommendation rank;
  • top-three recommendation inclusion;
  • rank-one recommendation inclusion;
  • raw mention presence;
  • sentiment;
  • buyer-intent cluster;
  • AI/search surface;
  • cited or attributable evidence source; and
  • other structured fields required by the benchmark.

The same coding rules should be applied consistently across categories and measurement periods.

For the public definitions and denominator rules associated with these fields, see AI Visibility Metric Definitions.


6. Build the Qualified Benchmark Set

The observations that survive both qualification stages form the qualified benchmark set.

This is the denominator used for the public recommendation metrics unless a metric explicitly defines a different applicable observation set.

This distinction is important:

> A report with 87 qualified observations does not mean only 87 AI queries were run. It means 87 observations from the much larger raw collection universe satisfied the benchmark's tracked-brand relevance and commercial buyer-intent rules.

Category reports should state both the upstream collection context and the final qualified denominator so readers can interpret rates correctly.


7. Qualified Surface Breadth

The benchmark can also report qualified surface breadth: the number of tested AI/search surfaces that contributed at least one qualified benchmark observation during the measurement period.

Qualified surface breadth is an output of the filtering process, not the number of surfaces tested.

For example, if the benchmark tests the same full surface universe in two consecutive months but qualified observations appear on three surfaces in one month and seven in another, the collection methodology has not necessarily changed. What changed is how broadly qualified commercial observations were distributed across the tested environments.

This metric can help identify whether a category or brand's discoverability is concentrated in a small number of AI environments or distributed more broadly.


8. Calculate and Publish Metrics

Metrics are calculated from stored coded observations using the denominator rules defined in the AI Visibility Metric Definitions.

Publication should follow four principles:

  1. Name the denominator. Do not imply that a rate calculated on qualified observations represents all raw AI responses.
  2. Pair rates with counts where practical. This is particularly important when one or two observations can materially change a percentage.
  3. Use exact stored numerators and denominators. Do not reverse-engineer production counts from rounded public percentages when source counts are available.
  4. Describe percentage movement correctly. When a rate changes from 27.3% to 44.8%, the movement is 17.5 percentage points. In compact tables, this may be displayed as Up 17.5 points.

9. Source and Citation Analysis

When an AI/search surface exposes citations, links, or attributable evidence sources, those sources can be analyzed as an additional research layer.

Source analysis may capture:

  • source domain or URL;
  • associated tracked brands;
  • buyer-intent cluster;
  • surface using the source;
  • frequency of appearance;
  • relationship to the recommendation outcome; and
  • whether the source is first-party, retailer, editorial, review, comparison, marketplace, or another source type.

Source presence should not automatically be interpreted as causal. It is evidence about the information environment surrounding an AI answer and should be described accordingly.


10. Quality Assurance

Before publication, benchmark data should pass a defined QA process.

At minimum, the production workflow should verify that:

  • query counts and surface counts match the configured research run;
  • qualification rules were applied consistently;
  • every published percentage resolves to a stored numerator and denominator;
  • brand names and entity mappings are normalized consistently;
  • recommendation ranks do not contain impossible or duplicate assignments;
  • sentiment values fall within the approved coding scale;
  • charts match their underlying tables;
  • report dates match the current benchmark refresh;
  • illustrative or test data has been removed; and
  • public copy reflects the current methodology version.

Where automated classification is used, high-impact exceptions and ambiguous cases should be reviewable against the original answer text.


11. Longitudinal Updates

Category reports use an evergreen URL. The same report page is updated as new monthly measurements become available rather than creating a new URL for every month.

The ongoing report should preserve meaningful historical context, including:

  • baseline values;
  • current values;
  • major historical highs or lows;
  • material reversals;
  • persistent leadership changes;
  • unusually large movements; and
  • methodological annotations when relevant.

This allows the page to become a cumulative research asset rather than a collection of increasingly outdated monthly pages.

If a methodology, surface universe, query framework, or calculation rule materially changes, that change should be documented under Research Standards & Interpretation.


How This Methodology Relates to Individual Reports

Individual category reports should contain only enough methodology to interpret the local data correctly:

  • raw research scale;
  • qualified denominator;
  • category-specific scope;
  • a concise explanation of qualification; and
  • links to this methodology and the shared metric glossary.

They should not reproduce this entire page or paraphrase it hundreds of different ways merely to create textual variation.

The unique value of each report should come from its category-specific observations, historical movements, competitive analysis, prompt evidence, and source environment.