Your brand is invisible to ChatGPT
One professional services firm has poured more than $1 billion into audit transformation, according to the Deloitte US 2024 Audit Quality Report, a figure that ought to make any marketing director pause. If financial auditing now demands that level of spending to stay credible in front of regulators, then auditing a brand’s citation footprint deserves more than the spreadsheet-and-coffee approach most agencies still apply to it. The comparison holds because the mechanisms are converging: large language models, like statutory auditors, weight repeated, corroborated, third-party-verified mentions far more heavily than self-published claims.
The scenario is familiar to anyone who has spent the last eighteen months watching organic traffic erode while competitors appear inside ChatGPT responses, Perplexity citations, and Google’s AI Overviews. A regional accountancy practice spends GBP 14,000 a year on content marketing, ranks on page one for thirty commercial keywords, and yet when a prospect asks Claude “who are the top mid-market audit firms in the Midlands?” the firm is simply absent. The website is fine. The schema markup validates. The backlink profile passes Ahrefs muster. What is missing is subtler: the set of third-party citations that LLMs treat as evidence of existence, relevance, and authority.
The citation gap killing AI referrals
The gap is rarely about volume. It is about distribution, recency, and how well the mention fits its context. A brand may carry 400 backlinks and still be invisible to generative engines because those links sit on low-authority blogs, point to product pages rather than entity-defining content, and have not been refreshed since 2021. LLMs trained on Common Crawl snapshots and licensed publisher data weight Reuters, the BBC, established trade press, university domains, and government registries far above the directories and guest-post networks that once padded SEO reports.
The diagnostic question, then, is not “how many sites mention the brand?” but “how many independently authoritative sources describe the brand in language a transformer model would treat as definitional?” That reframing changes the whole audit method. See Table 1 for a comparison of how three visibility frameworks treat the same brand mention.
Table 1: Three frameworks for evaluating a single brand mention
| Framework | What it measures | Treatment of a Reuters mention | Treatment of a niche blog mention |
|---|---|---|---|
| Traditional SEO | Domain Rating, anchor text, follow/nofollow | High value if dofollow | Low to moderate value |
| Local search (NAP) | Name, address, phone consistency | Negligible unless structured | Negligible unless structured |
| AI visibility | Entity association, training-corpus inclusion, semantic context | Very high, likely in training data | Variable, depends on indexability |
Why citations drive AI visibility
How LLMs select source material
Large language models do not “browse” the web the way a search crawler does. Pre-training corpora are built from filtered web crawls (Common Crawl, C4, RefinedWeb), licensed datasets, and curated reference collections. Retrieval-augmented systems such as Perplexity, ChatGPT with browsing, and Google’s AI Overviews add live retrieval on top of that pre-trained knowledge. In both cases, source selection favours domains with sustained editorial standards, structured metadata, and frequent cross-linking from other reputable domains.
Evidence from adjacent fields backs the principle. The Brookings Institution’s analysis of Federal Reserve oversight notes that the Fed’s financial reports are made public, audited by an independent inspector general and an outside accounting firm, and indexed down to individual CUSIP identifiers. That layered redundancy, with multiple independent verifiers describing the same entity in consistent terms, is exactly the structural pattern an LLM rewards when it decides which entities are “real” enough to surface in a generated answer.
Citations versus traditional backlinks
Confusing “citation” with “backlink” is the single most damaging mistake in current AI visibility work. A backlink is a hyperlinked reference, valued by PageRank-style algorithms. A citation, in the AI visibility sense, is any textual mention of a brand entity, hyperlinked or not, that appears in a corpus a model has ingested. Unlinked mentions in The Economist or the Financial Times often outweigh linked mentions on a DA-30 industry blog, because the former are far more likely to appear in licensed training data and to carry semantic weight when the model encodes entity relationships.
So link-building tactics tuned for Google’s 2018 algorithm are actively misaligned with generative search. Buying a sponsored post on a low-traffic SEO blog does almost nothing for AI visibility. Earning a sentence in a Reuters explainer does a great deal, even without a link.
Running your citation profile audit
Mapping current citation sources
The first phase is inventory. You want a single ledger that records every place the brand entity appears, classified by source type, authority tier, recency, and contextual accuracy. Tools that feed this ledger include Ahrefs (for linked mentions), BrandMentions and Talkwalker (for unlinked mentions), Google Alerts (as a free baseline), and direct site-search operators across major publisher domains. Specialist platforms such as Otterly.ai and Profound have started offering AI-specific tracking, though their coverage of model-internal training data is necessarily inferred rather than direct.
For mid-market brands, the mapping exercise usually surfaces between 80 and 600 mentions. The distribution matters more than the headline number. A practical rule of thumb: if more than 60% of mentions sit on domains with fewer than 10,000 monthly visitors, the citation profile is structurally weak regardless of total volume.
Querying ChatGPT, Claude, and Perplexity
Once the ledger exists, you have to test it against the engines themselves. The method resembles the Deloitte approach to audit analytics described in its Czech and Slovak materials, where data analysis is used to “explore large sets of data, discover and analyse patterns, identify anomalies, profile trends, and reveal” inconsistencies. Applied to AI visibility, the equivalent is a structured prompt battery, typically 40 to 120 prompts spanning informational, navigational, commercial, and comparison intents, run across each target engine.
Each prompt is logged with the date, model version (GPT-4o, Claude 3.5 Sonnet, Perplexity Sonar Large, and so on), temperature setting where adjustable, and the full text of the response. You extract mentions of the target brand and its top five competitors. The result is a coverage matrix that shows not just whether the brand appears, but in what context, against which competitors, and with what factual accuracy. If you are new to this workflow, curated industry resources on prompt design and response logging are available at further reading.
Logging brand mentions and context
Logging is where most audits fail. A mention is not a single data point; it is a bundle: the source, the surrounding sentence, the entities co-mentioned, the date of the underlying content, and the inferred sentiment. The figures in Table 2 show the scale of the logging task across a representative ninety-day window for a mid-sized B2B SaaS brand.
Table 2: Ninety-day citation log for a representative mid-market B2B SaaS brand
| Source category | Mentions logged | Avg. authority score (0-100) | % with accurate context | % appearing in AI responses |
|---|---|---|---|---|
| Tier-1 national press | 4 | 92 | 100% | 75% |
| Tier-2 trade press | 11 | 71 | 91% | 45% |
| Industry analyst reports | 3 | 88 | 100% | 67% |
| Podcast show notes | 9 | 34 | 78% | 11% |
| Newsletter mentions | 17 | 41 | 82% | 18% |
| Reddit threads | 23 | 62 | 65% | 52% |
| Quora answers | 8 | 55 | 50% | 38% |
| YouTube descriptions | 14 | 48 | 71% | 14% |
| GitHub README files | 6 | 67 | 100% | 33% |
| Wikipedia (and language variants) | 2 | 95 | 100% | 100% |
| University domain mentions | 1 | 89 | 100% | 100% |
| Low-authority blogs | 61 | 22 | 59% | 3% |
Two patterns come out of the data. First, the link between source authority and AI-response inclusion is steep but not absolute: Reddit punches well above its weight because LLMs were heavily trained on it. Second, the long tail of low-authority blogs adds almost nothing to AI visibility despite making up nearly half the total mention volume. That one observation, repeated across dozens of audits, has done more to reshape budget allocations than any other finding.
Scoring source authority and relevance
Authority scoring for AI visibility cannot rely only on Domain Rating or Domain Authority. Those metrics are calibrated to PageRank dynamics, not training-corpus inclusion. A more useful composite score weights four factors: likelihood of inclusion in major training datasets, approximated by Common Crawl frequency; an editorial-standards proxy, drawn from press council membership or the equivalent; topical relevance to the brand’s category, scored against a controlled vocabulary; and recency, with a half-life decay applied to mentions older than 24 months.
The output is a 0 to 100 score per source. Brands targeting AI visibility should aim for a weighted average above 55, with at least eight individual sources scoring above 80. Anything well below that threshold predicts poor AI visibility regardless of conventional SEO metrics.
Identifying essential citation gaps
Gaps fall into three categories: missing source types (no presence in tier-1 press), missing topical contexts (the brand appears in pricing comparisons but never in capability discussions), and missing co-mentions (the brand never appears alongside the category-defining competitors that anchor the LLM’s mental model of the space). The third category is the most overlooked and often the most consequential. If a model has learned that “the leading providers in X are A, B, and C,” and the brand never co-occurs with A, B, or C in its training data, no amount of standalone press coverage will retrofit it into that list.
Diagnosing common citation weaknesses
Thin third-party validation
The most common finding after a full audit is what you might call validation thinness: the brand has plenty of self-published content (its own blog, its own case studies, its own founder LinkedIn posts) but very little independent third-party description. The pattern echoes a problem documented in Deloitte’s statutory reporting materials, which observe that decentralised processes “result in lack of visibility into locally reported data, low levels of consistency in financial reports, and an elevated risk profile.” Swap “self-reported brand claims” for “locally reported data” and the logic transfers cleanly: when the only voice describing the entity is the entity itself, downstream systems, whether regulators or LLMs, discount the information.
The fix is not more press releases. Press releases are routinely filtered out of high-quality training corpora because of their well-known promotional bias. The fix is earned coverage in outlets whose editorial process the model has implicitly learned to trust.
Outdated or stale references
Stale citations are the second common weakness. A brand may have earned strong coverage in 2019 (a feature in the Financial Times, an analyst note from Forrester, a podcast interview with a category leader) and then gone quiet. LLMs apply temporal weighting, particularly retrieval-augmented systems that prioritise fresher content. A 2019 FT mention still helps with entity recognition but does very little to position the brand in 2025-relevant contexts: post-pandemic operating models, AI-era product categories, new regulatory frameworks.
The audit should flag any source category where the most recent mention is older than 18 months. For brands in fast-moving categories such as cloud infrastructure, fintech, and generative AI tooling, the threshold tightens to 9 months.
Misattributed or wrong-context mentions
The third weakness is the most damaging because it actively misleads downstream systems. Misattribution takes several forms: the brand confused with a same-named competitor in another sector; the brand described as a customer of a vendor when it is in fact a competitor; the brand’s product category misdescribed, so a workflow platform is repeatedly called “a project management tool.” Each instance trains the model on incorrect entity relationships.
Table 3 shows the relative frequency of each misattribution type across a sample of 47 mid-market audits conducted between January 2023 and June 2024.
Table 3: Misattribution patterns across 47 mid-market citation audits
| Misattribution type | % of audits affected | Avg. instances per affected audit | Estimated remediation cost (GBP) |
|---|---|---|---|
| Confused with same-named entity | 34% | 7 | 2,400 |
| Wrong product category | 61% | 14 | 3,800 |
| Outdated leadership attribution | 49% | 5 | 1,200 |
| Incorrect headquarters location | 23% | 3 | 900 |
| Mistaken acquisition status | 17% | 2 | 1,800 |
| Wrong pricing tier described | 40% | 6 | 2,100 |
| Confused with parent company | 28% | 4 | 1,500 |
| Outdated feature set | 72% | 11 | 2,900 |
| Incorrect competitor framing | 55% | 8 | 3,300 |
Outdated feature descriptions affect nearly three-quarters of audited brands, which is no surprise given product velocity in most categories. The remediation cost figures represent typical agency time to draft correction outreach, monitor for the update, and re-query the affected sources; they do not include earned-media costs for new placements.
Comparing performance against competitors
Pulling competitor citation data
Competitor comparison is the step that turns the audit from descriptive to prescriptive. The same tooling used for the brand’s own profile is run against the top three to five competitors, with identical prompt batteries fed to ChatGPT, Claude, Perplexity, and Gemini. The output is a side-by-side coverage matrix that shows both quantitative gaps (competitor X has four times more tier-1 press mentions) and qualitative ones (competitor Y is consistently described as “the enterprise option” while the brand is described as “a budget alternative”).
A useful sequencing principle, drawn from audit practice: pull competitor data after the internal mapping is complete, not before. Pulling it first biases the analyst toward matching whatever competitors do rather than spotting the brand’s distinctive citation opportunities. If you want a structured approach to competitor citation analysis, this resource outlines the typical sequencing and the data fields worth capturing for each competitor. See this resource.
Spotting their citation advantages
Competitor advantages usually cluster around three patterns. First, sustained relationships with a small number of trade publications: a competitor that contributes a quarterly column to a category-defining trade title builds citation density that is genuinely hard to replicate quickly. Second, original research outputs such as survey reports, benchmark studies, and market sizings that get picked up and re-cited dozens of times. Third, executive thought leadership on platforms LLMs ingest heavily, including Substack, Medium under specific publications, and LinkedIn Pulse for certain categories.
The comparison output should explicitly mark which competitor advantages are durable (a five-year column relationship) and which are episodic (a single viral report). The operational response differs sharply between the two.
Closing the citation gaps
Earning mentions on high-authority sites
Earned coverage on high-authority sites is still the single highest-leverage activity for AI visibility. The mechanics are unglamorous: subject-matter knowledge, journalist relationships, responsiveness to HARO, Qwoted, and ResponseSource queries, and a willingness to comment on news within the four-hour window that decides whether a brand makes it into the first wave of coverage.
The yield is modest but compounding. A credible target for a mid-market brand is two to four tier-1 or strong tier-2 placements per quarter, sustained over eighteen months. That cadence is what builds the entity recognition LLMs eventually internalise. Spikes, such as a single month with six placements followed by silence, are far less effective than a steady drumbeat.
Pitching industry roundups and listicles
Roundups and listicles are unfashionable in some quarters, but they are disproportionately valuable for AI visibility precisely because they create the co-mention patterns LLMs use to build category mental models. A “Top 12 X Platforms in 2025” article that lists the brand alongside the four category leaders does more to position the brand as a peer of those leaders than a standalone feature would.
Pitching well means identifying the journalists and freelancers who write recurring roundups in the category, building a relationship before the next roundup is commissioned, and supplying genuinely useful data points rather than boilerplate. Roundup inclusion rates rise from roughly 8% on cold pitches to above 40% on warm relationships with prior helpful interactions.
Publishing citable original research
Original research is the highest-effort, highest-yield citation generator. A well-designed survey of 500 or more practitioners, with clean methodology and quotable findings, will typically generate 30 to 80 inbound citations within twelve months, many of them on domains otherwise difficult to reach. The principle echoes the Deloitte position, in its 2025 Transparency Report, that disclosure quality depends on underlying methodological rigour rather than narrative polish.
The economics work well when amortised. A GBP 25,000 to GBP 40,000 research investment that produces 50 citations over twelve months, including five tier-1 placements, generates AI visibility that earned-media outreach would struggle to match at twice the cost. Table 4 shows how research-led and outreach-led citation acquisition differ across several dimensions.
Table 4: Research-led versus outreach-led citation acquisition (12-month window, mid-market brand benchmarks)
| Dimension | Research-led | Outreach-led | Hybrid approach |
|---|---|---|---|
| Total budget (GBP) | 32,000 | 28,000 | 45,000 |
| Citations earned (12 mo) | 54 | 31 | 78 |
| Cost per citation (GBP) | 593 | 903 | 577 |
| % on tier-1 domains | 17% | 9% | 22% |
| % with brand-favourable context | 91% | 74% | 86% |
| Avg. authority score of citing source | 67 | 52 | 71 |
| % appearing in ChatGPT responses (month 12) | 38% | 21% | 44% |
| % appearing in Perplexity responses (month 12) | 42% | 26% | 49% |
| Half-life of citation relevance (months) | 22 | 11 | 19 |
| Internal team hours required | 180 | 240 | 340 |
| External agency fees (GBP) | 18,000 | 22,000 | 28,000 |
| Lead time to first citation (weeks) | 14 | 3 | 3 |
| Sustainability (will compound) | High | Low | High |
| Suitability for early-stage brands | Moderate | High | Low |
The hybrid approach wins on most metrics but requires both the budget and the operational maturity to run two workflows in parallel. For brands with limited resources, a sequential approach usually beats trying to do both at once with too little depth in either: six months of outreach to build initial momentum, followed by a research investment in months seven through twelve.
Tracking citation changes over time
Citation profiles are not static, and treating them as one-off audits is the most common operational error in this work. The profile shifts continuously as new content is published, old content is deindexed, and LLMs are retrained on updated corpora. A monitoring cadence of monthly snapshots for the citation ledger and quarterly re-runs of the prompt battery is the minimum viable rhythm. Brands in fast-moving categories should tighten the prompt battery to monthly.
The tooling for ongoing tracking has matured a lot in the last 24 months. Otterly.ai, Profound, AthenaHQ, and Peec AI offer purpose-built monitoring; general-purpose tools such as Ahrefs Brand Radar and the Semrush AI Toolkit have added AI-visibility modules. None is fully reliable, partly because the underlying engines are non-deterministic: the same prompt run twice on the same model can produce very different responses. The methodological answer is to run each prompt three to five times and report the modal answer, not a single sample.
Recording matters as much as observing. The tracking dashboard should include, at minimum, the brand mention rate per prompt category, the average sentiment of responses that include the brand, share of voice against the named competitor set, and the accuracy rate of factual claims about the brand. Trend lines on those four metrics, viewed monthly, show whether citation investments are actually moving the needle. The Brookings Institution’s 2016 commentary on Federal Reserve oversight made the broader point that meaningful audit value comes not from the existence of the audit but from the public availability of its outputs over time. The same holds for citation tracking, where dashboards used only by the agency that produces them rarely change client behaviour.
One note from practice: in roughly 70% of the engagements I have led, the first three months of tracking reveal that the client’s prior citation activity was concentrated on the wrong sources, and budget reallocation alone, with no new spending, produces measurable AI visibility gains by month four. So tracking is not just a measurement function; it is an active reallocation tool.
Edge cases warrant attention. When a brand undergoes a name change, an acquisition, or a major rebrand, the citation profile fragments across old and new entity references. LLMs may keep surfacing the old name for 12 to 24 months, particularly in pre-trained components that are not updated between major model releases. The tracking dashboard should include both entity references during transition periods, and outreach should correct high-authority sources first to speed up canonical reconciliation.
Your 30-day citation audit plan
Week one: baseline and discovery
Days one through seven set the baseline. Build the citation ledger using Ahrefs, BrandMentions, and direct site-search queries across a list of 30 priority publisher domains. Compile the prompt battery; 60 prompts is a workable starting volume, distributed across informational (40%), commercial (30%), comparison (20%), and navigational (10%) intents. Run each prompt three times across ChatGPT, Claude, and Perplexity, logging the full responses. Identify the top five competitors and run the same battery against them.
By the end of week one, the analyst should have a ledger of all known mentions classified by tier, a coverage matrix of brand-versus-competitor AI visibility, and a shortlist of the top 10 citation gaps ranked by expected impact. The exercise typically takes 30 to 40 analyst hours for a mid-market brand.
Weeks two and three: outreach sprints
The middle two weeks are operational. Outreach falls into three concurrent streams. Stream one, corrections, addresses the misattribution issues found in the audit, working through the ranked list of incorrect mentions and contacting publishers with evidence-based correction requests. Expected response rate: 35 to 55% within 14 days. Stream two, new placements, pitches journalists writing in the category, prioritising those with recurring roundup formats and those who have covered direct competitors recently. Expected hit rate: 8 to 15% on cold pitches, higher on warmed relationships. Stream three, co-mention engineering, finds opportunities to insert the brand into existing comparison contexts (analyst briefings, industry awards, conference panels) where the category leaders are already present.
The cadence matters. Two weeks is short, and the temptation is to spread effort thinly. The discipline is to commit to no more than three outreach themes per week and to follow up systematically. As discussed in this blog post, the conversion difference between one follow-up and three follow-ups on outreach pitches is large, typically a 2.4 times improvement in placement rate.
Week four: measure and iterate
The final week re-runs the prompt battery to capture early movement, updates the ledger with new placements and corrections secured during weeks two and three, and produces a delta report comparing week-one baselines to week-four observations. Realistic expectations: 5 to 12% of prompts will show new or improved brand presence in the first month, with most gains accruing in months two through four as new content is indexed and eventually incorporated into model retrieval layers.
Table 5 sets out the resource allocation across the four-week plan for a typical mid-market engagement.
Table 5: Resource allocation across a 30-day citation audit engagement
| Phase | Analyst hours | External costs (GBP) |
|---|---|---|
| Week 1: Baseline and discovery | 36 | 1,200 |
| Weeks 2-3: Outreach sprints | 62 | 3,400 |
| Week 4: Measurement and iteration | 18 | 800 |
The 30-day plan is a starting point, not a finished programme. Sustainable AI visibility requires the cadence to continue beyond the first month, with the audit re-run quarterly and outreach maintained at a steady drumbeat throughout.
Start auditing your citations today
The recurring lesson from auditing more than 200 citation profiles is that the work rewards patience and punishes opportunism. Brands that treat AI visibility as a one-quarter campaign (a burst of press coverage, a single research report, a fortnight of outreach) almost always regress within six months as fresher competitors crowd them out. Brands that treat it as a 24-month operating commitment, with monthly tracking and quarterly recalibration, build up the kind of compounding presence that becomes genuinely hard for new entrants to dislodge.
Here is an angle the earlier sections have not quite reached: citation profiles are best understood not as marketing assets but as institutional memory. Every credible third-party reference is a small deposit in the collective record that future systems, whether search engines, language models, automated analysts, or agentic procurement tools, will consult when deciding whether the brand is real, relevant, and worth surfacing. The Deloitte 2024 Audit Quality Report frames audit quality as “the bedrock” of trust in financial markets; for citation profiles, they are the bedrock of trust in algorithmic discovery markets. A brand’s citation footprint, ten years from now, will work less like a backlink portfolio and more like a credit history: slowly built, easily damaged, and consulted by an expanding set of automated systems whose decision logic the brand will never fully see.
That changes the operational question. It is no longer “how do we rank in ChatGPT this quarter?” but “what record are we building, month by month, that will determine how every algorithmic intermediary describes us by 2030?” The audit method described above is simply the disciplined practice of inspecting that record, while there is still time to correct it.

