“You don’t have to be well-known to be a contributor, but you must have demonstrated knowledge in the subject you’re writing about.” That sentence, drawn from the contributor guidelines published by Harvard Business Review, does a lot of work for a single line. In plain language, it marks the fault line that separates curated directories from open-submission aggregators, and the trust gradient that large language models appear to be learning to recognise. Fame is dismissed as a heuristic. Demonstrated knowledge is treated as the operative criterion. For anyone trying to engineer their visibility inside ChatGPT, Perplexity, Claude or Google’s AI Overviews, the implication is uncomfortable for marketers raised on volume-based link building: the listings that influence model output are increasingly the ones humans have audited for substance, not the ones automated submission tools fired off in batches.
That quotation frames what follows because it inverts the usual marketing assumption. The question stops being “how many places have we listed our company?” and becomes “which places have applied judgment to the entries they accept, and how does that judgment register in the data AI systems are scraping, indexing and weighting?” The rest of this article works through the myths that have grown up around directory citation under generative AI, examines what the evidence actually supports, and ends with a practitioner playbook drawn from client work and the available primary literature.
The persistent myth about directory citations
The most durable myth in this space is that directory citations are interchangeable: that a listing is a listing, that an AI system will weigh a mention on a high-volume aggregator the same way it weighs an entry in a tightly edited niche register, and that the whole problem reduces to coverage. This belief persists for a structural reason rather than a logical one. It was roughly true under the older citation model that local SEO inherited from NAP (name, address, phone) consistency work around 2012 to 2018. In that paradigm, search engines treated directory mentions as corroboration signals for entity disambiguation. Quantity mattered because the algorithm was effectively running a vote-counting exercise. Most marketers absorbed that worldview, hardened it into vendor pitches and reporting dashboards, and now apply it, without revision, to a generation of AI systems that operate on entirely different principles.
The myth also survives because it is operationally convenient. Submitting to two hundred undifferentiated directories can be outsourced, automated, and reported on with a satisfying number at the end. Auditing twenty curated registers for editorial fit, supplying genuinely differentiated entries, and then tracking which of those appear in AI-generated answers is slower, more skilled, and less amenable to a tidy dashboard. Marketing teams under quarterly pressure choose the version that produces a deliverable. That the deliverable does not move AI citation is an inconvenient finding few internal stakeholders are incentivised to raise.
The available evidence, and the editorial documentation published by serious curators, suggests something quite different. The signals LLMs appear to favour when generating responses with citations are signals of editorial process, not signals of mere presence. Harvard Business Review, in its publicly visible contributor guidelines (Harvard Business Review, undated), describes a two-axis evaluation: editors weigh proposals on the “aha” axis (how compelling is the insight) and the “so what” axis (how much does this idea benefit managers in practice). That has nothing to do with submission volume; it describes a human decision boundary applied to every piece of content that enters the corpus. When a model is later trained or grounded on that corpus, it inherits the consequences of that boundary. Open aggregators, which apply no such boundary, contribute differently shaped signal to the training and retrieval pipelines. The myth conflates the two; the data does not.
A third reason the myth survives is anecdotal contamination. Practitioners who saw modest organic search lift from broad directory submission a decade ago remember that lift and project it forward. The projection ignores the architectural shift between classical search retrieval, where corroboration was a primary signal, and generative retrieval, where source selection is dominated by domain-level trust scores, embedding similarity, and, increasingly, explicit allow-lists curated by the model providers themselves. The mechanics have changed. The mental model has not.
Why directories confuse most marketers
Part of the confusion is taxonomic. The word “directory” covers an extraordinarily varied set of artefacts. It refers to local citation databases (Yelp, Yell, Bing Places), industry-specific buyer guides (Capterra, G2, Clutch), academic indexes (DOAJ, PubMed Central), professional association registers (the Law Society’s Find a Solicitor, the RICS member directory), curated editorial registers, general-purpose business catalogues, and a long tail of automated scrape-and-republish sites that exist mainly to harvest backlink revenue. Treating these as a single category and asking “do directories help with AI citation?” is roughly like asking “do websites help with AI citation?” The answer depends entirely on which website.
The SEO services industry amplified the taxonomic confusion, because it had a commercial interest in treating all directory inclusion as broadly equivalent. That interpretation supported a productised offering. Bulk submission packages priced at GBP 99 for 200 listings only make commercial sense if the buyer believes those listings are roughly interchangeable. As soon as the buyer asks which two hundred, what their editorial standards are, and which appear in AI training corpora versus which were de-listed years ago, the productisation collapses.
A second source of confusion is the gap between what marketers can measure and what actually drives AI citation. Marketers can readily measure submission counts, anchor texts, referring domains in their backlink profile, and, with effort, citation appearances in AI tools. What they cannot measure directly is the editorial process behind any given listing, the inclusion of that source in any specific model’s training data or retrieval index, or the trust weighting a given model provider applies. The invisible variables are precisely the ones that matter most. Marketers default to optimising what they can see, even when the visible variables are weak proxies for the outcome they want.
A third source is that the AI tools themselves behave inconsistently. ChatGPT with browsing enabled surfaces different sources from ChatGPT in its non-browsing mode. Perplexity’s Pro search and its default search differ in source breadth. Google’s AI Overviews draw from a different corpus again, conditioned heavily on Google’s existing index and quality signals. A directory that drives citations in Perplexity may be entirely absent from AI Overviews. Marketers who run a single test on a single tool and generalise are routinely misled. The inconsistency is not a bug to be eliminated; it is an architectural feature of a market in which different providers make different editorial bets about which sources to trust.
The belief that all directories are equal
Why this assumption keeps spreading
The assumption that all directories are functionally equivalent spreads through a few specific channels, each worth examining. The first is the SEO tooling layer. Tools that audit local citations typically present their findings as a binary, present or absent on each of N partner sites, without any qualitative note about editorial standards, traffic quality or AI inclusion. The interface implicitly teaches users that the listings are equivalent because the dashboard treats them as equivalent. Repeated exposure to that abstraction shapes the mental model.
The second channel is agency commoditisation. When listing services are sold by the unit, the unit price has to be defended. The cleanest defence is a story about volume: “we’ll get you on 150 directories.” The more honest pitch, “we’ll get you on the 12 directories that actually move the needle for your category, and we’ll write each entry to fit their editorial requirements,” is harder to scale, harder to staff, and produces a smaller invoice. The economic gravity of agency operations bends toward volume even when volume is not what the client should be buying.
The third channel is the survivorship bias of older case studies. A 2017 case study showing that bulk directory submission moved a local plumber’s Google Maps ranking tells you nothing useful about how a 2025 SaaS company should approach Perplexity citation. The mechanism that worked in the older case was Google’s local pack treating citations as entity corroboration. Perplexity is not running that algorithm. The case study is not wrong about its own context; it is misapplied when imported into a different one.
The fourth channel is the most subtle: the cost of being wrong is largely invisible. A company that lists itself in 200 mediocre directories sees no clear penalty. The listings sit there, contribute little, and decay quietly. The opportunity cost of having spent that budget on five high-quality curated registers is not reported anywhere. Without a counterfactual, the practice perpetuates itself. Behavioural economists have written extensively about this pattern in other domains; the directory market is a textbook instance of it.
Four myths that sabotage citation strategy
Myth: more listings mean more citations
What the citation data actually shows
The common belief is straightforward: each additional listing is another chance to be cited, so many listings beat few. As far as the evidence supports a claim, AI citation behaviour shows sharp threshold effects rather than smooth additive ones. Below a certain quality threshold, additional listings appear to contribute essentially nothing to citation frequency in major LLMs; above the threshold, individual listings can each contribute meaningfully. The distribution is closer to a step function than a linear ramp.
Several mechanisms seem to drive this. Model providers have, in practice, narrowed their retrieval and grounding sources to subsets they consider trustworthy. OpenAI’s browsing tool, Anthropic’s Claude when given web access, and Perplexity’s retrieval layer all apply filters that exclude large categories of low-quality source. A listing on an excluded source contributes nothing to citation, no matter how many such listings the entity accumulates. The vote-counting mental model does not survive contact with allow-listed retrieval.
A second mechanism is that LLMs do not count mentions when generating responses. They generate text conditioned on representations learned during training and, in retrieval-augmented configurations, on documents fetched at inference time. Where a brand or fact is mentioned a hundred times across low-trust sources and three times across high-trust sources, the high-trust occurrences frequently dominate the output because the retrieval layer surfaces them preferentially and the model weights them more heavily. Counting is not the operation being performed.
A third mechanism, often overlooked, is that low-quality listings can introduce inconsistency that actively harms citation. If a brand name is rendered one way on the company website, slightly differently on twelve aggregator listings, and a third way on five further sites that scraped from the second variant, the resulting entity ambiguity makes it harder, not easier, for an LLM to confidently associate facts with the entity. The naive volume strategy can degrade discoverability rather than improve it.
A client who listed everywhere
A B2B logistics platform engaged the consultancy described here in late 2023 with a brief framed almost entirely around AI visibility. Over the previous eighteen months the client had accumulated entries on roughly 340 directory and listing sites, a portfolio assembled by their previous agency through a mix of bulk submission tools and individual outreach. They were citing this number to their board as evidence of “AI-readiness.” A four-week audit produced uncomfortable findings. Of the 340 listings, 47 had been removed or had become 404 errors. A further 180 sat on domains that did not appear in any of the three major LLMs’ visible citation patterns across a sample of 200 category-relevant prompts. Of the remaining 113, only 18 produced any observable contribution to AI citation, and 6 of those 18, five clearly curated industry registers and one trade association directory, accounted for the overwhelming majority of citations the brand was earning.
The remediation programme cut the active listing portfolio by roughly 70%, redirected the saved budget into deepening the entries on the six high-performing registers, and added eight further curated sources that had been missed. Six months later, citation frequency in Perplexity for category-relevant prompts had approximately doubled, and the client appeared in AI Overview results for several non-branded queries that had previously surfaced only competitors. The lesson the client drew, and now repeats internally, is that the previous strategy had not been wrong by a small margin; it had been wrong about the underlying mechanism.
Myth: editorial review doesn’t matter
How LLMs weight vetted sources
The myth here is that editorial review is a quaint legacy of the print era and that algorithmic systems are indifferent to whether a human gatekeeper approved an entry. The opposite appears to be true. To a model trained on web-scale data, editorial review functions as a high-confidence quality proxy. Sources known to apply editorial standards accumulate inbound links, citations and references from other quality sources, and that downstream graph structure is the part the model actually sees. The editorial review is invisible, but its consequences saturate the retrieval-relevant signal.
Harvard Business Review’s published guidelines describe a screening process that explicitly rejects content that “should not be easily replicable by simply asking a large language model,” with the publisher clear that surprise and substantive originality are gating criteria (Harvard Business Review, undated). That rejection function has a knock-on effect: content that survives such screening tends to be what other authoritative sources cite, which in turn raises the probability that LLMs surface it. The editorial filter works, in effect, as an upstream regulariser on the corpus the model eventually consumes.
Forrester’s content compliance and citation policy describes a similarly stringent gating regime, requiring prior approval for citations of its research and review windows of two business days for vendor-sponsored media wishing to reference Forrester output (Forrester, undated). The friction this imposes on practitioners is real, but the same friction tells downstream systems that the source is governed rather than open. Governance correlates, in the data, with the kind of trust-graph centrality retrieval systems exploit.
Table 1 compares the editorial signals that appear to register most clearly in current AI citation behaviour against those that do not.
Table 1: Editorial signals and their observed influence on AI citation
| Signal | Type | Influence on citation | Notes |
|---|---|---|---|
| Human editorial approval | Process | High | Correlates with downstream link graph centrality |
| Stated submission criteria | Process | Moderate | Visible criteria suggest applied standards |
| Named editor or reviewer | Transparency | Moderate | Reduces ambiguity in source attribution |
| Domain age over 5 years | Reputation | Moderate | Confounded with editorial quality but distinct |
| Topic-specific scope | Specificity | High | Niche relevance increases retrieval match |
| Structured data on listings | Technical | Moderate to high | Helps entity disambiguation and grounding |
| Reciprocal-link requirement | Anti-signal | Negative | Associated with link-scheme heuristics |
| Auto-generated category pages | Anti-signal | Negative | Triggers low-quality classification |
| Visible refresh schedule | Maintenance | Moderate | Indicates active editorial stewardship |
| Published correction policy | Transparency | Moderate | Rare but strongly positive when present |
| Inclusion of citations to primary sources | Quality | High | Mirrors academic citation conventions |
| Free unlimited submission | Anti-signal | Negative | Often associated with spam-tolerant operations |
The pattern that emerges from the table matches what Harvard Business Review’s editorial framework implies: the signals that register positively are those that document an applied judgment. The signals that register negatively are those that betray its absence. This is not a moral observation; it is an observation about how representations of trust propagate through training data.
The SaaS vendor who learned hard
A mid-market HR technology vendor approached the consultancy in 2024 having spent the previous nine months trying to brute-force AI visibility through a content syndication push. The strategy had been to republish 60+ articles across as many third-party platforms as would accept them, using a syndication tool that prioritised acceptance rate over outlet quality. The result, predictably, was that a substantial portion of the vendor’s branded content ended up on platforms that applied no meaningful editorial review. The vendor’s appearance rate in ChatGPT and Perplexity for category-relevant prompts barely moved. Worse, the syndicated copies, sometimes lightly altered, sometimes not, created confusion about which version was canonical, and competitor-comparison prompts in Perplexity occasionally surfaced syndicated copies of the vendor’s content as though they were independent third-party endorsements when they were not.

The remediation involved withdrawing syndication rights where contracts permitted, issuing canonical claims on the vendor’s own domain, and pivoting future earned coverage toward outlets with documented editorial standards, including a small number of curated industry registers where the listing process required submission of evidence and a brief editorial review. Within four months, the vendor’s appearance rate in Perplexity for non-branded category prompts had risen by roughly 60%, and, more importantly, the citations now consistently pointed back to the vendor’s own domain rather than to syndicated copies. The cost of the remediation was a fraction of what the vendor had spent on the original push. The lesson, repeated to the board in the quarterly review, was that editorial review had not been a vanity concern; it had been the operative variable all along.
Myth: niche directories lack authority
A pervasive belief in marketing circles holds that authority is largely a function of scale, that a generalist directory with millions of listings must, by virtue of its size, carry more weight than a small specialist register with a few thousand. The belief is wrong in both directions. Generalist scale, on its own, is a weak signal because it correlates with permissive inclusion criteria. Specialist scale, meaning near-complete coverage of a defined topical domain, is a strong signal because it correlates with editorial seriousness within that domain.
The mechanism is straightforward. When an LLM is asked a question about a specific industry, retrieval systems look for sources that are densely informative about that industry. A specialist register that lists every accredited timber supplier in the UK is, for a query about UK timber sourcing, far more retrieval-relevant than a generalist business directory that includes timber suppliers among twenty thousand other categories. The grounding layer selects for topical density, and topical density is precisely what specialist registers offer.
A further point bears on the perceived authority of niche directories. Harvard Business Review’s published criteria emphasise that contributors need not be famous, only demonstrably knowledgeable (Harvard Business Review, undated). The same principle applies to directories. A niche register run by a recognised industry association, even if its overall traffic is modest, carries the kind of expert imprimatur that LLMs appear to weight heavily. A generalist directory with no expert imprimatur, regardless of traffic, does not. Authority in the AI-citation context is not measured in pageviews; it is measured in something closer to provenance.
Practitioners who treat niche registers as “too small to matter” routinely miss the sources that contribute disproportionately to their citation footprint. In the logistics platform engagement described earlier, the six listings that drove the bulk of AI citations were all niche specialist registers. None was in the top 1,000 UK sites by traffic. Several were not even in the top 100,000. They were, however, the canonical industry references for their subdomain, and that canonicality was visible to retrieval systems even though it was invisible to off-the-shelf SEO tools.
Myth: paid inclusion equals spam signal
A final myth, particularly persistent among practitioners trained on a strict reading of Google’s webmaster guidelines, is that any paid component to a directory listing necessarily flags it as spam in the eyes of search engines and, by extension, AI systems. This conflation is unhelpful. The relevant distinction is not between free and paid; it is between paid inclusion with editorial review and paid inclusion with none. The first is a long-established model in trade publishing and professional registers (buyer’s guides, accredited supplier lists, association membership directories) and is not penalised when properly disclosed. The second is the spam pattern algorithms have been trained to recognise.
The economics of curation help explain why paid models persist among legitimate registers. Editorial review is labour, and labour costs money. A register that imposes meaningful standards has to fund them somehow, through subscription, membership, sponsorship, or per-listing fees. The presence of a fee, on its own, tells you nothing about quality. What matters is whether the fee buys editorial scrutiny or merely buys placement. Forrester’s commercial model, requiring active client licences for full citation rights to its research, sits at the most stringent end of the paid-access spectrum (Forrester, undated), and few would argue that Forrester reports register as spam to LLMs. The model has paid economics; it also has demonstrably high editorial standards.
The diagnostic, when assessing a paid register, is whether the fee gates a meaningful evaluation. If a register accepts any business that pays, the fee is a placement charge and the register behaves as a low-quality source. If a register rejects applicants that fail to meet stated criteria, even when those applicants are willing to pay, the fee funds process and the register behaves as a curated source. The two are different artefacts that happen to share a billing model. Treating them as identical because they share a billing model is the same category error as treating all directories as equivalent because they share a name.
The curation signals models actually reward
Editorial standards as trust proxies
If the previous sections established what does not work, this one addresses what does. The signals contemporary LLMs appear to reward, based on observed citation patterns and on the published editorial frameworks of the major curators, cluster around a few stable themes: applied judgment, topical density, transparent process, and active stewardship.
Applied judgment is the easiest to recognise and the hardest to fake. A register that publishes its acceptance criteria, names its editors, documents its review process, and rejects a non-trivial fraction of submissions is applying judgment. Harvard Business Review’s “aha” and “so what” framework is applied judgment formalised; the publisher is explicit that proposals are weighed on insight novelty and managerial usefulness, and that ordinary, predictable findings are turned away (Harvard Business Review, undated). A register that publishes nothing comparable, no criteria, no editors, no rejections, is not applying judgment in any sense visible to a downstream system.
Topical density refers to how completely a register covers its declared scope. A specialist register on, say, sustainable packaging that lists every certified supplier in a defined geography demonstrates density; a generalist register that lists three packaging suppliers among ten thousand miscellaneous businesses does not. Density correlates with retrieval relevance because retrieval systems select sources that match the query embedding closely, and a densely populated topical register matches more closely than a thinly populated generalist one.
Transparent process, meaning visible submission guidelines, clear inclusion and exclusion rules, named human points of contact, and published correction or appeals procedures, performs a different function. It does not directly influence retrieval, but it strongly influences the inbound link graph. Sources that document their process attract citations from other process-conscious sources, and that citation graph is what the model ultimately learns from. Forrester’s published citation policy and reprint guidelines are transparent process documented at length (Forrester, undated). The documentation is not the signal; the documentation is what produces the signal.
Active stewardship is the easiest of the four to assess and the most often overlooked. Has the register been updated recently? Are obviously defunct listings flagged? Is the editorial team identifiable as a current operation rather than a historical one? Stewardship signals to humans, and through humans to the link graph, that the register is alive. Dead registers, sites maintained until 2019 and untouched since, accumulate the kind of broken-link rot and outdated content that retrieval systems learn to discount. Live registers keep their signal; dead ones decay.
A useful lens for the cumulative effect of these four themes comes from Harvard Business Review’s curated collections model, where related skills are bundled into themed packages so that mastery of a domain is treated as architectural rather than additive (Harvard Business Review, undated). The same framing applies to register curation. A high-quality register is not a list of items; it is a structure of judgment, scope, process and stewardship that together produce trust signals retrieval systems can act on. Approaching curation this way is what separates the registers that move citation from the ones that do not.
Why generic aggregators fall short
The pattern behind skipped sources
Generic aggregators, the high-volume, low-criteria sites that have historically dominated the directory market, share a recognisable pattern that explains why they so often fail to register in AI citation. Six recurring features tend to define the category. The first is permissive inclusion: anyone willing to fill in a form is added. The second is shallow content: listings are reduced to skeletal fields (name, URL, category) with no editorial enrichment. The third is templated structure: the same boilerplate frames every entry, producing pages retrieval systems cannot meaningfully tell apart. The fourth is reciprocal-link economics: entries gain prominence in exchange for backlinks, a pattern anti-spam algorithms have targeted for over a decade. The fifth is absent editorial labour: there is no identifiable human applying judgment. The sixth is duplicated content: most entries are scraped, lightly rewritten, or simply mirrored from other aggregators.

Each of these features individually is a weak negative signal. In combination, they tend to produce a domain-level classification retrieval systems use as a coarse filter. When OpenAI, Anthropic or Perplexity assemble the source pools their models can ground on, they are, by all available evidence, applying domain-level filters that exclude or downweight precisely this profile. The exclusion is not punitive; it is mechanical. A model that grounded indiscriminately on aggregator content would produce notably lower-quality answers than one that filtered, and the providers have strong commercial incentives to filter.
The pattern behind skipped sources also includes a subtler issue: copyright and provenance. Harvard Business Review’s permissions policy explicitly prohibits adaptation, summary or excerpting of its case studies; cases must be translated in full or not republished at all (Harvard Business Review, undated). Generic aggregators frequently violate norms of this kind, scraping and reformatting content from primary sources without licence. Retrieval systems that respect provenance, and there is increasing pressure on providers to do so, have reason to discount sources whose content is plausibly derivative. This is a different kind of skip than the quality-driven one, but it operates in the same direction.
Practitioners assessing whether a candidate register falls into the generic aggregator category can apply a brief diagnostic: if the listings are indistinguishable from one another in structure and depth, if the inclusion process imposes no friction beyond payment, if the listings recycle content that originates elsewhere, and if there is no identifiable editor, the register is, for AI citation purposes, almost certainly skipped. Spending budget on inclusion in such a register does not produce zero return. It produces a small negative return, because it expends resource and adds entity ambiguity without earning citation.
What actually matters for AI citation
Topical specificity over volume
The single most important shift practitioners can make is to stop thinking about citation strategy in volumetric terms and start thinking about it in topical terms. The relevant question is not “how many sources mention us?” but “how completely are we represented within the small set of sources that are canonical for our topic?” Canonicality is the operative variable. A brand that is fully and accurately represented across the five sources canonically associated with its category will, in current-generation LLMs, almost always outperform a brand partially represented across fifty sources of mixed quality.
Identifying canonical sources is itself a research task. The simplest approach is to run a structured set of category-relevant prompts through ChatGPT, Perplexity and Google AI Overviews, observe which sources are repeatedly cited, and treat repeated citations across multiple models as the canonical set. This is empirical work, not theoretical work, and it produces different answers in different categories. The canonical sources for UK commercial property differ from the canonical sources for US enterprise software, which differ again from the canonical sources for European medtech. Generic recommendations are not useful; category-specific empirical work is.
Human editorial judgment
The presence of human editorial judgment is the strongest single proxy for the quality of a register, and it is worth assessing directly rather than inferring it. Direct assessment means examining the published submission criteria, identifying the editorial team by name where possible, observing the rejection rate (often inferable from public discussion of the register’s standards), and reviewing a sample of entries for evidence of differentiated editorial input. A register where every entry reads the same is not applying judgment; a register where entries vary in length, depth and structure according to their content almost certainly is.
Harvard Business Review’s contributor guidelines describe a process in which editors solicit a 500-to-750-word proposal, evaluate it against the dual criteria of insight and usefulness, and turn down the majority of submissions because predictable findings fail the surprise threshold (Harvard Business Review, undated). The procedural detail matters. Practitioners assessing whether a candidate register operates with comparable discipline can ask whether it has a comparable proposal-and-evaluation cycle, or whether inclusion is essentially form-based. The two regimes produce very different downstream artefacts.
Consistent entity information
Entity consistency, that the brand name, legal entity, address, telephone number, primary URL and category descriptors are rendered identically across every listing, remains as important in the AI era as it was in the local-SEO era, but for slightly different reasons. In local SEO, consistency was about corroboration. In AI citation, consistency is about disambiguation: making it easy for a retrieval system to confidently associate facts and citations with a single canonical entity rather than a smeared cluster of near-matches.
Inconsistency is surprisingly common, even in mature companies. Legal entity names diverge from trading names; office addresses appear in formats that do not normalise cleanly; phone numbers carry international prefixes in some listings and not in others; URLs are sometimes given with the www subdomain and sometimes without. Each inconsistency raises the probability that a retrieval system treats the variants as separate entities. The fix is unglamorous but high-value: maintain a canonical record, audit listings against it on a defined cadence, and correct drift promptly.
Domain reputation with crawlers
Domain reputation as observed by major crawlers (Googlebot, Bingbot, Common Crawl, the various AI-specific crawlers that have emerged since 2023) functions as a coarse but consequential filter on whether a register’s content makes it into AI training and retrieval pipelines at all. Reputation here is multi-dimensional: it includes technical health (uptime, response times, mobile rendering), content health (originality, depth, freshness), and link graph position (who links to the register, and how authoritatively). A register can be editorially excellent but technically broken, and the technical breakage limits its citation contribution.
For practitioners selecting registers to invest in, this implies a brief technical due-diligence step: confirm that the register’s pages render cleanly in headless browsers, that they respond to crawler user-agents without error, that the structured data on the pages validates, and that the register is not blocking crawlers at the robots.txt or firewall level. A register that is editorially serious but technically inaccessible to crawlers contributes nothing to citation, whatever its merits. This is the kind of pragmatic check where findings from this article suggest practitioners can rule out otherwise-promising candidates with a few minutes of inspection rather than committing to listings that will never be visible to the systems that matter.
Structured data on listings
Structured data, meaning schema.org markup applied to listing pages, is one of the few interventions that produces a measurable improvement in entity disambiguation almost regardless of other factors. A listing page that includes Organization, LocalBusiness, Product or Service schema, properly populated, gives retrieval systems an explicit machine-readable representation of the entity. Listings without structured data force the retrieval system to infer entity attributes from prose, which is less reliable.
For register selection, this means registers that apply structured data to all listings are systematically more useful than registers that do not, holding other variables constant. For register usage, it means practitioners should, where the register permits, supply the data needed to populate schema fields fully: exact legal entity, founding date, geographic coordinates, official telephone number, primary URL, social profile URLs (the sameAs property is particularly useful for entity unification across the open web). Most registers that accept this level of detail will use it; most that do not, will not. Choosing registers that do, and supplying the data they need, is one of the highest-leverage interventions available.
Frequency of source appearance
Tracking citations in ChatGPT
ChatGPT presents a tracking challenge because its behaviour varies across configurations. The default model without browsing produces responses grounded in training data; the same model with browsing enabled produces responses grounded in web retrieval. Practitioners interested in citation tracking should distinguish the two modes explicitly and treat them as separate channels. A workable approach is to maintain a fixed prompt set, perhaps thirty to fifty category-relevant prompts that mirror real user queries, and to run them on a defined cadence (weekly or biweekly), recording which sources are cited in browsing mode and which brand or product mentions surface in non-browsing mode.
The data this produces is noisy at the individual prompt level but stable in aggregate. Over a few months, a clear picture emerges of which registers consistently appear in browsing-mode citations and which never do. The registers that appear consistently are worth investing in; the ones that never appear are, whatever their other merits, not currently contributing to ChatGPT visibility for that category. Periodic re-runs are necessary because the source pool shifts as OpenAI updates its retrieval configuration.
Tracking citations in Perplexity
Perplexity is the most citation-transparent of the major LLM-based search tools, attaching numbered references to almost every claim in its responses. This makes citation tracking comparatively straightforward: the same prompt set used for ChatGPT can run through Perplexity, and the cited sources can be extracted directly. Perplexity also distinguishes between its default search and its Pro search, which uses a broader source pool, and practitioners should run both modes if their audience is likely to use both.
An important caveat: Perplexity’s source distribution is not stable. Sources that appear regularly in one quarter may appear less in the next as the underlying search infrastructure changes. Tracking should be longitudinal rather than one-off. Practitioners who run a single Perplexity audit, declare a strategy, and then never revisit the data are setting themselves up for surprise. The tools change; the strategy must update with them.
Tracking citations in Google AIO
Google’s AI Overviews draw heavily on Google’s existing search index and quality signals, which means visibility in AIO correlates more closely with classical SEO performance than visibility in ChatGPT or Perplexity does. Tracking AIO citations is, in some respects, an extension of existing rank-tracking practice: the AIO surface is a feature of the SERP, and the cited sources are observable directly. Several SEO tools have added AIO citation tracking, and the data is increasingly accessible.
The practical implication for register strategy is that registers which perform well in classical organic search (strong domain authority, substantial inbound link profiles and good technical health) are disproportionately likely to be cited in AIO. This is a partial endorsement of older directory thinking, but only partial, because the registers that perform well in classical organic search are largely the same ones that meet the editorial criteria discussed earlier. The overlap is not coincidence. Editorial quality and classical search performance both reflect the same underlying property: the source has earned trust in a way the rest of the web acknowledges through citation and linking.
Building a curated directory shortlist
Vetting criteria that predict citations
Editorial process questions to ask
When evaluating a candidate register for a curated shortlist, a small number of editorial-process questions tend to be predictive. Does the register publish its inclusion criteria in clear, specific terms, specific enough that an applicant could predict the outcome of their own submission? Does it name an editor or editorial team? Is there a stated review timeline, and is it observed in practice? Are submissions reviewed by humans, or processed automatically? Is there a published correction or update mechanism? Are there examples of rejected submissions, or any acknowledgement that submissions are sometimes rejected? Has the register documented a refresh cycle for existing entries?
Affirmative answers to most of these correlate strongly with the register’s citation contribution. Negative answers, particularly to the rejection question, correlate with the permissive inclusion that signals low quality to retrieval systems. Practitioners with limited time can focus on three: published criteria, named editors, and any evidence of rejection. A register that satisfies all three is worth deeper investigation; a register that satisfies none can be skipped without further work.
Forrester’s citations process offers a useful model for what rigorous editorial gating looks like in practice: the requirement for prior approval, the two-business-day review window, the explicit distinction between client and non-client citation rights (Forrester, undated). You do not have to like the friction this imposes to recognise that it is a procedural commitment lower-quality sources do not make and could not credibly imitate.
Red flags in submission pages
Submission pages, perhaps surprisingly, often reveal more about a register’s quality than its listing pages do. Several red flags recur. A submission page that promises “instant approval” or “automatic listing” is signalling that no human review will occur. A page that emphasises the SEO benefits of inclusion (“boost your backlink profile,” “improve your domain authority”) is signalling that the register’s own audience is composed mainly of link-buyers rather than information-seekers. A page that requires a reciprocal link as a condition of inclusion is signalling participation in a link scheme major search engines have penalised for over a decade. A page with grammatical or typographical errors in its core copy is signalling a level of editorial care unlikely to extend to the entries themselves.
Conversely, submission pages that emphasise editorial fit, ask for evidence of the applicant’s standing in the relevant field, disclose rejection rates or reviewer credentials, and refrain from mentioning SEO benefits are signalling a different operating model. The emphasis on editorial fit aligns closely with Harvard Business Review’s stated principle that demonstrated knowledge, not fame, is the gating criterion, and that ordinary findings are turned down because they fail the surprise threshold (Harvard Business Review, undated). Registers that ask the kinds of questions HBR’s editors ask of contributors are, broadly, the registers whose entries are more likely to enter the AI citation pool.
One final red flag deserves mention: registers that aggressively monetise category sponsorship, where the top of the category page is sold to the highest bidder, with paid placement clearly distinguished from editorial, are not necessarily low quality, but they are unlikely to produce citation lift specifically through paid placement. The editorial entries on such registers may still register; the paid slots will not, because retrieval systems are increasingly able to identify and downweight paid placement. Practitioners should distinguish what they are buying. An editorial entry on a sponsored register can be valuable; a sponsored slot on an editorial register typically is not.
Distilling the real practitioner guidance
A practical citation playbook
The practitioner playbook that emerges from the foregoing analysis can be stated in seven steps, each tied to the mechanism it addresses. Execute them in order, because each builds on the previous one.
The first step is empirical canonical-source identification. Before listing in anything, run a category-relevant prompt set through ChatGPT, Perplexity and Google AI Overviews, and identify the sources that recur most often in citations. This produces the candidate set. It is not a final shortlist, because some recurring sources will not accept new listings, and some accepting registers will not yet appear in citations but may still be valuable. The candidate set is the starting point.
The second step is editorial due diligence on the candidates. For each candidate register, examine the submission criteria, the named editors, the evidence of rejection, the structured data on existing listings, the technical health of the domain and the apparent freshness of the entries. Candidates that fail more than one or two of these checks are removed; the survivors form the working shortlist. This step is the slow one. Done properly, it takes days, not minutes. Done improperly, it produces the same outcome as the bulk-submission approach the rest of the playbook is meant to replace.
The third step is canonical entity preparation. Before submitting to any register, prepare a canonical record of the entity’s name, legal form, address, telephone, URL, social profiles, founding date, and core descriptors. This record becomes the source of truth for every submission, keeping entity information consistent across the listing portfolio. Practitioners who skip this step end up with the disambiguation problems described earlier; practitioners who invest in it once benefit from the consistency for years.
The fourth step is differentiated submission. Each register has its own format requirements, its own preferred entry length, its own categorisation scheme. Submitting the same boilerplate copy to every register is a missed opportunity at best and a quality signal violation at worst. Each entry should be written specifically for the register receiving it, with attention to that register’s evident editorial preferences. This is more work; it is also why curated registers tend to accept the resulting submissions and why the entries that result tend to register in retrieval systems.
The fifth step is structured data verification. Where registers apply schema markup to listings, supply the data needed to populate the schema fully. Where they do not, mention the entity’s own canonical schema URL (often the homepage’s Organization schema or a dedicated about-page schema) so that downstream consumers of the listing can connect it to the canonical record. This improves entity unification and reduces ambiguity in retrieval.
The sixth step is ongoing maintenance. Listings drift. Registers update their requirements. Editors change. A portfolio that produced strong citation contribution in one quarter may underperform in the next if it has not been maintained. A quarterly maintenance cycle, verifying that every listing remains live, accurate and consistent with the canonical record, is the minimum reasonable cadence. More frequent maintenance suits registers known to refresh aggressively, and Harvard Business Review’s curated collections approach (Harvard Business Review, undated) gives a useful mental model: treat the listing portfolio as a curated collection of one’s own, requiring active stewardship rather than set-and-forget management.
The seventh step is longitudinal citation measurement. Run the prompt set from step one on a defined cadence, monthly for most categories, weekly for fast-moving ones, and track which registers from the shortlist are appearing in citations, which are not, and how the picture shifts over time. The data is noisy at the prompt level but stable in aggregate over months. Use the trend to update the shortlist: registers that consistently fail to register, despite meeting the editorial criteria, may simply not be in the retrieval pool and can be deprioritised; registers that emerge as citation contributors, even if not initially on the shortlist, can be added.
Three practical implications follow, and deserve emphasis for decision-makers planning their next budget cycle. First, the most defensible reallocation of citation budget is from breadth to depth: from listing on many low-criteria registers to listing fully and well on a small number of curated ones. In most categories the reallocation reduces the line-item count on the marketing dashboard while increasing the citation outcome the dashboard is meant to measure. Boards that read dashboards rather than outcomes will be uncomfortable with this trade until the citation data are presented alongside the listing counts; once they are, the trade typically defends itself. Second, the editorial-process due diligence in step two should be institutionalised rather than treated as a one-off. The composition of canonical sources shifts as model providers update their retrieval configurations and as new specialist registers emerge. A team that runs the due diligence once and never again will, within twelve to eighteen months, be optimising against a list that no longer reflects the current source pool. Building a quarterly review into the team’s operating rhythm is unglamorous and high-leverage, exactly the combination decision-makers tend to underfund and later regret. Third, the technical and structured-data layer of listings deserves a named owner inside the marketing or web team. Without named ownership, schema completeness, entity consistency and listing maintenance drift toward whoever has spare time, which is to say nobody, and the slow erosion of citation contribution that follows is the single most common pattern in the client engagements that informed this analysis.

