Skip to content

Research Standards & Interpretation

Overview

This page defines the universal data-quality, interpretation, versioning, and publication standards that apply across CiteWorks Studio's AI Market Discovery research.

Category reports are designed to contain primarily unique market evidence and analysis. Universal limitations and research controls are maintained here so they do not need to be duplicated across hundreds of evergreen report pages.

For the research workflow, see the AI Market Discovery Methodology. For calculation terminology, see AI Visibility Metric Definitions.


What the Benchmark Is Designed to Show

AI Market Discovery research is designed to measure directional patterns within a defined competitive and query framework.

It can show:

  • recommendation coverage among tracked brands;
  • changes in recommendation position over time;
  • top-three and rank-one preference patterns;
  • brand mention presence;
  • coded sentiment movement;
  • differences by buyer-intent cluster;
  • differences by prompt and AI/search surface;
  • qualified surface breadth;
  • source and citation patterns where evidence is observable; and
  • historical inflection points as the same evergreen report accumulates additional monthly data.

These signals are intended to identify where brands are gaining, losing, or concentrating AI-mediated discovery visibility.


What the Benchmark Does Not Establish by Itself

The benchmark is not designed to prove:

  • total category market share;
  • revenue share;
  • search-engine market share;
  • total AI prompt share;
  • attributable revenue or sales impact;
  • every answer a user could receive from an AI system; or
  • causality from a metric increase or decline alone.

A movement in recommendation coverage can identify a change worth investigating. It does not, by itself, establish why that change occurred.

Causal interpretation requires deeper review of the underlying prompts, recommendation ranks, cited sources, brand content, competitor changes, and collection conditions.


AI Outputs Are Dynamic

AI-generated answers can vary because of factors including:

  • model or system updates;
  • collection date and time;
  • geography;
  • language;
  • personalization;
  • conversation history;
  • logged-in state;
  • interface or product surface;
  • retrieval/citation changes; and
  • stochastic model behavior.

A benchmark should therefore be read as a structured measurement of defined observations, not a claim that every end user will receive exactly the same answer.

Repeated monthly measurement is valuable because it helps distinguish one-off observations from persistent competitive patterns.


Small-Count Interpretation

For lower-visibility brands, one or two observations can materially change a percentage.

Category reports should therefore pair rates with absolute counts where practical and add a report-specific interpretation note when a headline movement is driven by a small numerator.

Example: A decline from seven valid recommendations to one is strategically relevant, but the percentage movement should be read alongside the absolute counts rather than presented as a large rate change without context.

Small counts do not make an observation meaningless. They change the level of confidence and the type of conclusion that can reasonably be drawn from it.


Raw Collection Volume vs. Qualified Sample Size

The benchmark intentionally begins with a larger raw collection universe and narrows it through relevance and buyer-intent qualification.

A final qualified set of 87 observations does not mean only 87 AI interactions were tested. It means 87 observations from the broader collection satisfied the benchmark's criteria for the public commercial analysis.

Reports should state this distinction clearly near the top of the page and link to the Methodology for the full qualification logic.


Evergreen Report Policy

Each industry/category report should use one evergreen URL that is updated as new monthly measurements become available.

The report should accumulate historical context rather than create a separate indexable page for each month.

The evergreen page should retain meaningful historical information such as:

  • original baseline;
  • current measurement;
  • notable historical peaks or lows;
  • major leadership changes;
  • persistent gains or declines;
  • reversals in direction;
  • important ranking changes; and
  • any methodology annotations that materially affect interpretation.

The objective is to create a progressively stronger longitudinal research asset rather than a sequence of increasingly outdated monthly pages.


Longitudinal Integrity

Month-over-month comparisons are only useful when the reader can distinguish market movement from methodology movement.

The research program should maintain an internal change log for:

  • AI/search surfaces added or removed;
  • query-universe changes;
  • geography changes;
  • collection-setting changes;
  • entity/brand-set changes;
  • buyer-intent classification changes;
  • recommendation-validity rule changes;
  • sentiment model changes;
  • formula changes; and
  • pipeline or QA changes that could affect the series.

When a change materially affects comparability, the relevant report should include a concise annotation.

Where feasible, historical measurements should be recalculated under the new rule. Where that is not possible, the report should make the break in comparability explicit.


Qualified Surface Breadth

The number of AI/search surfaces represented in the final qualified set can change even when the benchmark tests the same surface universe.

A surface may contribute no qualified observations in one month and several in another.

This is why reports should distinguish:

  • surfaces tested, which are part of the benchmark design; and
  • qualified surface breadth, which is the number of tested surfaces contributing at least one qualified observation after filtering.

A change in qualified surface breadth should not automatically be described as a change in platform coverage methodology.


Source and Citation Interpretation

Where AI systems expose citations or attributable evidence, those sources can provide valuable context about the information environment surrounding a recommendation.

However:

  • citation presence does not prove that the source caused the recommendation;
  • the same source can support multiple brands;
  • different surfaces may retrieve different evidence for the same prompt; and
  • source frequency should be interpreted with the underlying prompt and recommendation outcome.

Public reports should use language such as “associated with,” “appeared around,” “was cited in,” or “coincided with” unless causal evidence exists.


Reporting Movement in Rates

Changes between percentage rates are reported in percentage points.

For example, a change from 27.3% to 44.8% is an increase of 17.5 percentage points.

For readability:

  • prose should use “percentage points”;
  • compact tables and charts should use “Up 17.5 points” or “Down 8.0 points”; and
  • public-facing reports should use points or percentage points rather than abbreviated technical shorthand.

A percentage-point movement should not be labeled as a simple percent change, because readers may interpret that as relative growth rather than the difference between two rates.


Report-Specific Limitations

Individual category reports should not repeat all universal limitations on this page.

Instead, they should add only limitations that materially affect the specific report, such as:

  • unusually small brand counts;
  • sparse observations within one buyer-intent cluster;
  • missing citation/source data on particular surfaces;
  • the first month of a new historical series;
  • a material methodology change;
  • an entity entering or leaving the competitive set; or
  • another concrete constraint that changes how the local results should be interpreted.

This keeps category pages concise while preserving the context required for responsible analysis.


Publication Quality Standards

Before an AI Market Discovery report is published or refreshed, the following checks should be completed.

Data integrity

  • Exact numerators and denominators are pulled from stored benchmark data.
  • Percentages reconcile with their underlying counts.
  • Duplicate or malformed observations are removed or resolved.
  • Brand/entity mappings are normalized.
  • Historical values are not accidentally overwritten.

Editorial integrity

  • The executive summary reflects the actual current data.
  • Claims distinguish observation from interpretation.
  • Causal language is avoided unless supported.
  • Small-count movements are contextualized.
  • Percentage-point movement is labeled correctly.
  • Generic methodology is linked rather than repeated in full.
  • Illustrative or prototype data is removed.

Human readability

  • The most important findings appear near the top.
  • Tables use explicit labels and units.
  • Acronyms are defined or avoided where unnecessary.
  • Charts have textual or tabular equivalents.
  • Readers can understand what the denominator represents without reading a separate technical document first.

Machine readability

  • Important facts are present in HTML text, not only images.
  • Headings describe the content beneath them accurately.
  • Tables use real header cells in the rendered page.
  • Internal links connect the report to methodology, metrics, standards, and the research hub.
  • Structured data, where used, matches visible page content.
  • Updated/published dates accurately reflect the research page.

Corrections and Updates

Material corrections should be reflected in the report's modified date and retained in an internal audit trail.

Examples include:

  • a coding error affecting published rates;
  • an incorrectly mapped brand;
  • a duplicate observation;
  • an incorrect denominator;
  • a source-attribution error; or
  • a material methodology description error.

Minor copy edits that do not change the research result do not require the same level of correction disclosure.


Versioning Principle

The shared Methodology, Metric Definitions, and Research Standards pages should use stable URLs and be updated centrally as the research program evolves.

Do not create multiple near-identical versions of these pages for different industries. Do not paraphrase the same universal methodology across hundreds of category pages purely to create textual variation.

The differentiation of the report library should come from unique category evidence, longitudinal movement, company implications, prompt observations, and source patterns.