AI Visibility

    How to Analyse Sources Cited by AI Models

    This guide describes a measurement method. Automated AI prompt, mention and citation monitoring is not available in the current public BeKnow product. Collect the observations with an external service or a documented manual sample.

    Marco Salvo
    Marco SalvoFounder of BeKnow · SEO & AI
    Updated September 5, 2026
    6 min read
    How to Analyse Sources Cited by AI Models

    This guide describes a measurement method. Automated AI prompt, mention and citation monitoring is not available in the current public BeKnow product. Collect the observations with an external service or a documented manual sample.

    A list of URLs cited by an AI model is not yet an AI visibility analysis. The useful work begins when those URLs are connected to the prompt, the claim they support, the type of source selected and the brand that benefits. Without that context, teams tend to count domains, imitate competitors and publish more content without knowing why one source was chosen.

    Source analysis should answer a practical question: what evidence environment does the model assemble for this topic, and where can the brand make a legitimate contribution to it? The answer may involve an owned page, independent corroboration, clearer technical documentation or no new content at all. The method must therefore combine quantitative patterns with close reading.

    Preserve the complete observation

    Every citation belongs to an answer, and every answer belongs to a prompt, model, date, language and configuration. Store the complete observation before extracting domains. A bare URL loses whether it supported a definition, product recommendation, statistic, objection or warning. It also loses whether browsing was enabled and whether the citation was visibly attached to a particular sentence.

    Record the exact prompt, generated answer, model or interface, observation time, source URL, page title and cited passage where available. Resolve redirects and normalize URLs carefully, but keep the original address for auditability. Two tracking parameters should not create two sources; two distinct articles on the same domain should not be collapsed into one observation.

    Repeated runs matter because generative answers vary. A source appearing once is different from one selected consistently across models and dates. The prompt portfolio should also be stable and intentional. The guide to choosing prompts for AI visibility monitoring explains how to avoid a sample dominated by branded or artificially favourable questions.

    Classify the source before judging it

    Start with ownership: owned, competitor-owned, independent editorial, institutional, community, marketplace or unknown. Then identify the page type: homepage, product page, documentation, research, comparison, review, news article, profile or user-generated discussion. These dimensions explain more than a league table of domains.

    Next classify the citation role. A source can establish a fact, define a concept, provide a firsthand statement, compare alternatives, validate reputation or document a limitation. The same domain may play different roles in different answers. A competitor’s documentation cited for a technical fact is not the same strategic signal as its product page being recommended.

    Finally assess freshness, specificity and inspectability. Does the page state the claim directly? Is the evidence dated? Can a reader identify the author or methodology? Does the page remain accessible without scripts, login or geographic restriction? These properties do not prove why a model selected the source, but they produce testable hypotheses rather than folklore.

    Separate occurrence, persistence and prominence

    Occurrence measures whether a source appeared. Persistence measures how often it reappears across repeated, comparable observations. Prominence describes where and how it appears: first citation, central support, peripheral reference or long list. All three matter.

    A domain cited in 30% of answers may look dominant, yet most appearances could be peripheral. Another source may occur less often but support the decisive recommendation. Report raw counts and rates beside qualitative role. Do not merge languages, markets and models until the underlying distributions have been inspected; an aggregate can hide that a source dominates only one interface.

    This distinction complements mention rate and citation rate. Those metrics reveal whether a brand or its sources enter answers. Source analysis explains which pages were chosen, for what purpose and alongside whom.

    Compare competitor citations without copying competitors

    When competitors are cited and the brand is not, the first question should not be “how do we reproduce their article?” Examine the query intent and the evidence role. The competing source may contain original data, a concise definition, product documentation, independent recognition or a comparison structure that makes verification easy.

    Create a source-gap record containing the prompt family, recurring cited pages, role, supported claims and the brand’s current coverage. Then decide whether the brand has a genuine basis to add better evidence. If it owns relevant data, publishing methodology and limitations may be useful. If the missing signal is independent reputation, another self-authored landing page will not substitute for external corroboration. If the competitor is cited because its product actually offers a capability the brand lacks, content cannot repair the product gap.

    This restraint is important. Analysis should not become an instruction to manufacture claims, mimic wording or pursue citations at any cost. The article on why AI cites competitors instead of your brand develops the broader causes behind these gaps.

    Turn patterns into an evidence backlog

    Each proposed action should retain the observation that justified it. A useful backlog item states the affected prompt family, recurring source pattern, hypothesis, owner and success check. “Write more GEO content” is not actionable. “Create an inspectable methodology page because comparison prompts repeatedly cite competitors’ research pages for this claim” is testable.

    Possible interventions include clarifying an existing page, consolidating conflicting facts, adding dated methodology, improving crawlable documentation, strengthening entity consistency or earning legitimate independent coverage. Some findings belong to product, PR, technical SEO or customer support rather than the editorial team. Routing the action correctly is part of the analysis.

    Validate what the cited page actually says

    Automated extraction helps at scale, but a human should inspect important sources. Confirm that the page is live, the cited claim exists and the answer has not exaggerated it. Look for dates, primary references, conflicts of interest and whether the AI cited a secondary summary instead of the original evidence.

    Also inspect your own cited pages. A citation is not automatically a success if the destination is obsolete, contradictory or unable to convert an interested reader. Ensure that the page has a clear canonical, accessible content, consistent entity naming and a useful next step. AI visibility and human experience cannot be separated once a citation sends someone to the site.

    Avoid false causal claims

    No external observer can normally prove that a particular on-page change caused a model to cite a URL. Model training, retrieval, browsing indexes, interface rules and source availability are partly opaque. A before-and-after increase is evidence of association under documented conditions, not proof of a universal ranking factor.

    Use controlled comparisons where possible. Keep the prompt cohort stable, record changes, repeat observations and compare source roles. Avoid public promises such as “this schema guarantees citations.” The broader AI visibility and GEO framework treats visibility as measurable behaviour, not a guaranteed optimization formula.

    Frequently asked questions

    Should every cited URL count equally?

    No. Count occurrence transparently, then distinguish role, persistence and prominence. A central source supporting a recommendation carries a different meaning from a peripheral link in a long list.

    Are citations proof that a page was used to generate the answer?

    They are visible evidence presented with the answer, but interfaces differ. Avoid claiming access to hidden training or reasoning. Analyse what can be observed and document the system and mode.

    How often should cited sources be reviewed?

    Use a cadence that matches the market and observation volume. Recheck important prompt families after material site changes, model changes or product launches, while retaining a stable recurring sample for comparison.

    Is domain authority enough to explain citations?

    No. Reputation may matter, but relevance, specificity, source type, freshness, accessibility and the model’s retrieval environment can all influence the observed set. Source analysis should not reduce selection to one third-party metric.

    Build a map of evidence, not just a domain ranking

    The goal is not to collect the most citations in isolation. It is to understand which sources support which claims, how consistently they appear and where the brand can contribute accurate, accessible evidence. That map produces better editorial decisions and more honest AI visibility reporting.

    What BeKnow keeps

    Keep interventions, hypotheses, assets, observations and decisions in a searchable workspace. Record what is known and what still needs checking.

    A change in performance after an intervention does not prove that the intervention caused it. Cross-platform attribution and a complete analytics dashboard are not available today.

    Project memory and data imports do not require an AI model key. A compatible external AI client may have its own costs. BYOK applies only to available functions that actually call an external provider.

    Next step

    Start with one project, one documented change and the evidence needed to review it. Source connections.

    Record your first intervention. How it works.

    About the author

    Marco Salvo is the founder of BeKnow. With more than 20 years in SEO, he created BeKnow to connect project changes with real-world results and turn that history into knowledge people and AI can use.

    Record your first intervention

    Start with one project, one documented change and the evidence needed to review it.

    Create a free workspace