The 62% citation surprise
When you ask a generative answer engine a local commercial question, about 62% of the time it will cite at least one structured listing source among its first three references. That number cuts against the common assumption that large language models have made curated directories obsolete. It stands out because the story since late 2022 has been one of imminent extinction: if a model can synthesize the open web, why would it lean on Yelp, Yellow Pages, or any of the thousands of vertical listing sites that spread during the early 2000s?
The answer has less to do with the directories themselves and more to do with what retrieval-augmented generation systems actually need. They need verifiable, structured, frequently updated entities with consistent attributes: names, addresses, phone numbers, hours, categories, and credible third-party signals. That is, almost by definition, what a business directory provides. Models do not cite directories because the directories are charming. They cite them because directories solve a grounding problem that unstructured web pages create.
This article works through that dynamic with the available evidence. It looks at what changed between the pre-LLM era and the post-LLM era of search, what the data say about which directories now appear in AI outputs, where the evidence is solid and where it stays limited, and what practitioners should do differently in response. The framing draws on enterprise research from Forrester, market data aggregated by Statista, and the wider literature on emerging-technology adoption documented in Harvard Business Review. Read the headline statistic above, and several related figures throughout, as directional rather than definitive. The measurement methods vary widely, and the picture is moving fast enough that any quarterly snapshot is partial by design.
What the data do support is a more careful conclusion than either the doom or the boosterism around directories tends to allow. Directories have not died. They have not thrived uniformly either. A split is underway: a small number of high-authority, structured, well-verified directories are gaining citation share at the expense of a long tail of thin, unverified, scraped or duplicative properties. The middle is hollowing out. To see why, you have to look at how AI search works, what information it needs, and how those needs meet the way directories have historically been built.
How AI search reshaped directory traffic
Click-through decline since 2022
The most visible change practitioners report is a drop in referral clicks from search results pages, especially for informational queries that AI summaries can answer without a click-through. This trend is well documented across the wider publishing ecosystem and has come up repeatedly in coverage from Harvard Business Review, where commentary on generative AI’s effect on content economics has picked up since the launch of ChatGPT in late 2022. Directories with traffic models built around informational queries, such as “what time does X open” or “is Y open on Sundays”, have absorbed an outsized share of this loss, because those queries are exactly the ones AI Overviews resolve in-line.
Transactional and high-intent queries tell a different story. Searches that imply imminent commercial action, like booking, calling, or requesting a quote, still tend to send users to a destination, because models are tuned to defer the final commercial step rather than complete it on their own. A directory whose listings act as conversion endpoints, with booking widgets, click-to-call, and lead forms, has fared much better than one whose pages work only as reference material. The decline is not uniform. It clusters at the informational end of the funnel.
Citation volume in LLM responses
Against the click-decline picture, citation volume, meaning the number of times a directory’s domain is named or linked within an AI-generated response, has risen for the better-structured listing sites. The mechanism is simple: retrieval-augmented systems sample candidate sources during inference, score them for relevance and authority, and embed citations to support factual claims. Directories that publish clean structured data, keep listings fresh, and show real depth of user reviews tend to clear those scoring thresholds again and again. Forrester’s own materials describe surveying over 500,000 consumers, executives, and technology leaders annually, and that research places grounded retrieval at the center of credible enterprise AI deployments.
Citation is not the same as traffic, and that distinction matters. A model can name a directory in its answer without sending the user there. The relationship between citations and downstream visits is shaped by interface design choices the directory operator does not control. This split, visibility without click-through, is one of the hardest measurement problems in the current environment, and the section on attribution comes back to it in detail.
Directory domains in ChatGPT outputs
Sampling ChatGPT outputs across local commercial queries shows a pattern that has settled over the last several quarters: a small number of directory domains account for most citations. The concentration recalls organic search results before 2022, where high-authority sites dominated head terms. The names involved will surprise nobody who has worked in local SEO: Yelp, TripAdvisor, the major aggregators, the Better Business Bureau, and a handful of vertical leaders in legal, medical, and home services. Past that head, citation drops off steeply.
For newer or smaller directories, the implication is uncomfortable but clear: aggregate authority signals matter, and they compound. A directory that has spent a decade building up reviews, structured data, and inbound links benefits in AI search for the same reason it benefited in classical search. The underlying ranking systems weight similar signals, even when the surface presentation has changed.
Perplexity source frequency data
Because Perplexity surfaces its sources prominently, it offers the cleanest observational dataset for citation behavior. Source-frequency analyses across local and B2B queries consistently show directory domains appearing in the top five sources at rates above their organic search visibility. One reasonable reading is that Perplexity’s retrieval layer prefers structured, entity-rich pages because they are easier to parse for the disambiguation tasks the model performs. A second reading, which fits with the first, is that directories’ aggregation of third-party reviews gives the kind of multi-source corroboration retrieval systems treat as a quality signal.
Both readings match the wider principle in Harvard Business Review’s ongoing coverage of AI adoption: production AI systems work best when they cross-check human-generated signals rather than relying on any single source. Directories, seen this way, are cross-checking infrastructure.
Google AI Overview inclusion rates
Google’s AI Overview is a more complicated case. Inclusion rates for directory content vary a lot by query category and by whether the Local Pack appears in the same SERP. When the Local Pack is present, AI Overview tends to defer to Google’s own entity graph. When it is absent, third-party directory citations rise. This substitution pattern is what you would expect if Google preferred its first-party data where available and back-filled with directory citations where its own coverage was thin.
For practitioners, the takeaway is that competing for AI Overview inclusion is partly about finding the queries Google’s own data does not yet cover well, and making sure directory presence in those gaps is strong. Cross-referencing Table 1 shows how those category gaps map to citation behavior across the major AI surfaces.
Table 1: Estimated directory citation share by AI search surface and query category
| Query category | ChatGPT citation share | Perplexity citation share | Google AI Overview share | Local Pack overlap |
|---|---|---|---|---|
| Local restaurant discovery | High | High | Moderate | High |
| Home services (plumbing, HVAC) | High | High | Moderate | High |
| Legal services | Moderate | High | Low | Moderate |
| Medical and dental | Moderate | Moderate | Low | High |
| B2B SaaS evaluation | Moderate | High | Low | Low |
| Travel and hospitality | High | High | Moderate | Moderate |
| Financial services | Low | Moderate | Low | Moderate |
| Niche professional (e.g. translation) | Moderate | High | Low | Low |
| Generic retail | Low | Low | Low | High |
The pattern in the table is worth sitting with. Where Google has dense first-party data, such as generic retail and mainstream medical, third-party directory citation share is low across surfaces. Where the entity graph is sparser, such as niche professional services and B2B SaaS, third-party directories carry outsized weight, especially on Perplexity. This is the inversion practitioners need to absorb: directory investment pays off most exactly where the platforms have not yet built strong first-party coverage.
Defining the modern business directory
Structured data requirements
A modern business directory, in the technical sense that matters for AI search, is no longer a list of hyperlinked names. It is a structured database whose individual records expose enough machine-readable attributes, such as name, address, geographic coordinates, category, opening hours, payment methods, accreditations, photographs, reviews, and identifiers cross-referenced with other authoritative graphs, that a retrieval system can confidently tie a query to a specific entity. The underlying schema is what decides whether a directory can be cited reliably, not the visual presentation of its pages.
This shift mirrors wider patterns in enterprise data infrastructure, where the value of intangibles, data and technology has, as Deloitte notes in its legal services materials, grown as software, ideas and intellectual property come to dominate balance sheets. A directory’s database works the same way: it is the directory’s main asset, and the discipline applied to maintaining it shapes the directory’s commercial trajectory.
Schema.org markup adoption rates
Schema.org markup adoption among the top-tier directories is now near-universal for core types like LocalBusiness, Organization, Review, and AggregateRating, and it increasingly extends to vertical-specific subtypes such as MedicalBusiness, LegalService, and FinancialProduct. Among mid-tier and long-tail directories, adoption is uneven. Many still rely on legacy templates that produce visually competent listings without machine-readable attribute markup. The gap matters. In practice, AI retrieval systems do not parse unstructured HTML with the same confidence they extract entities from typed graphs.
Adoption is also moving from descriptive to prescriptive. Where directories once treated schema markup as an SEO bonus, the leading operators now treat it as the production interface to AI search. Listings are designed schema-first, with the human-readable presentation generated downstream from the structured record rather than the other way around.
Verification signals that matter
Verification, the process of confirming that a listed entity is what it claims to be, has become a primary differentiator. The verification signals AI systems appear to weight include domain age and consistency, phone-number verification, address geocoding accuracy, review velocity and pattern (reviews concentrated suspiciously in time tend to depress trust), licence and accreditation lookups (especially in regulated verticals), and cross-reference with authoritative third-party graphs.
The reason, again, is grounding. A model that hallucinates a clinic’s address creates immediate harm. A model that defers to a directory whose verification process detects address mismatches against postal data creates far less. Operators who have invested in verification infrastructure now earn back that investment through citation share, even when the verification work is invisible to end users.
Category taxonomy depth metrics
Taxonomy depth, meaning the granularity of category trees and the precision of category assignments, has become unexpectedly important. Shallow directories that classify a business as merely “Restaurant” do worse in AI citations than directories that classify it as “Restaurant > Italian > Neapolitan-style pizzeria > wood-fired > gluten-free options”. Retrieval queries rarely match on broad categories. They match on specific attributes. A taxonomy designed to capture those attributes in a structured way generates more matches and therefore more citations.
Operators who treat taxonomy as an editorial discipline, refreshed annually, informed by query-log analysis, and mapped against external reference taxonomies, outperform those who treat it as a one-time setup task. The practice resembles, at a smaller scale, the ongoing work catalogued in the World Bank’s Business Ready data programme, where sustained classification rigor underwrites the comparability of cross-country commercial information.
Review volume thresholds for citation
Review volume works as a citation threshold rather than a linear lift. Listings below a minimum review count (the exact figure varies by vertical, but is usually in the range of 5 to 25 reviews) are much less likely to be cited than those above it. Past the threshold, extra reviews yield diminishing returns. The pattern fits what you would expect if AI systems used review count as a confidence filter: enough reviews to cross-check, but no preference for inflated counts that may point to manipulation.
This cuts against the volume-maximizing instincts that dominated review acquisition strategy a decade ago. In practice, spreading review acquisition effort across many listings to clear the threshold is more productive than piling it onto a few flagship listings to maximize count.
Which directories AI models cite most
Yelp citation share by query type
Yelp’s citation share stays substantial in food, hospitality, and personal services, where its review depth and entity coverage are hard for newer competitors to match. In professional services and B2B contexts, Yelp’s share drops sharply, pushed aside by vertical-specific directories whose taxonomies and verification regimes are better tuned to those domains. The pattern reinforces the wider point that horizontal directories gain advantage in consumer-discretionary categories, while vertical directories gain advantage in considered or regulated categories.
The sensitivity to query type is sharper than many practitioners assume. Yelp can dominate citations for “best brunch in [neighbourhood]” while disappearing entirely from citations for “commercial litigation attorney in [city]”. A citation strategy built on a single horizontal directory is brittle by design.
Industry-specific directory performance
Industry-specific directories, such as Avvo in legal, Healthgrades in medical, Houzz in home design, and G2 and Capterra in B2B software, have, on the available evidence, done comparatively well in the LLM era. Their structured attribute sets are tuned to the questions users in those verticals actually ask, their verification regimes fit the regulatory context, and their review communities are usually less polluted by promotional or reciprocal activity.
Coverage from Harvard Business Review on enterprise AI adoption has repeatedly stressed that domain-specific tooling outperforms general-purpose tooling once users move past early experimentation. The same dynamic seems to govern directory citation: as AI search matures, models lean more on sources whose vocabularies match the domain of the query.
Local Pack versus AI citation overlap
The overlap between Google’s Local Pack and AI citations is high but not total. Listings that appear in the Local Pack are more likely than baseline to be cited in AI Overviews, and the reverse holds at lower magnitude. The two ranking systems share enough underlying signals (proximity, relevance, prominence) that visibility in one tends to predict visibility in the other. Practitioners who have historically optimized only for Local Pack visibility will find that much of that work transfers, though not all of it.
The non-overlapping part is interesting. Some businesses appear in AI citations without ranking in the Local Pack, and vice versa. The gap tends to track the strength of third-party directory citations: businesses with deep, structured presence across multiple authoritative directories gain AI visibility on their own, independent of their Local Pack performance. That is a real signal that off-Google directory work still matters, perhaps more than it did when the Local Pack was the only game in town.
Niche directory citation lift
Niche directories, meaning those covering tightly defined verticals or geographies, have shown a citation lift that is, on the available evidence, larger relative to their size than generalist directories. The mechanism appears to be specificity: when a query is narrow, a directory whose entire content set addresses that narrowness signals more relevance per unit of authority than a generalist directory whose treatment of the topic is shallow. Curated platforms such as an in-depth piece of editorially reviewed listing infrastructure can build authority out of proportion to their absolute size, as long as the editorial standards hold over time.
The lift is not automatic. Niche directories that stop staying fresh, or whose verification practices lapse, lose citation share quickly. The advantage of specificity depends on the structural discipline that makes the specificity legible to retrieval systems.
Geographic variation in source selection
Geographic variation in citation behavior is large and often underappreciated. In markets where Google’s first-party data is dense, such as most of the United States, Western Europe, and urban centers in developed economies, third-party directory citation share is lower. In markets where Google’s coverage is thinner, such as secondary cities, emerging economies, and language regions with fewer resources, third-party directories carry more weight.
This has implications for international strategy. A directory operator expanding into a new geography may find that AI citation share rises faster than organic traffic share, because the gap in first-party data is wider than the gap in conventional search infrastructure. The reverse holds too: directories operating in saturated markets face a citation ceiling that no amount of optimization will fully lift.
Strong evidence versus weak signals
It is worth stopping to separate where the evidence here is strong from where it is genuinely limited. The strong evidence concerns the existence and shape of citation behavior: that AI systems cite directories, that they cite some directories more than others, and that citation patterns vary by query type and surface. These findings replicate across observational studies run by independent practitioners and are stable enough across quarters to treat as structural rather than passing.

The weaker evidence concerns the exact mechanisms. In most cases we have no direct access to the retrieval and ranking algorithms AI systems use to select citations. The available evidence is inferential: we observe inputs and outputs and deduce what the system must be weighting. That is a respectable method, but it is not the same as reading the source code. Claims about specific weighting (for example, “review velocity is weighted at X% of the citation score”) deserve skepticism unless they come with methodology that explains how the figure was derived.
The thinnest evidence concerns causation and durability. Even where we see that directories with property X are cited more often than directories without property X, we cannot always tell whether X causes the citation or whether X correlates with some hidden variable (domain age, brand strength, editorial discipline) that does the causal work. Practitioners who treat correlational findings as causal levers tend to be disappointed when their interventions fail to produce the predicted lift. Deloitte’s wider transformation work stresses this distinction repeatedly: what looks like a rule in one quarter may be an artifact of a temporary configuration in the underlying systems.
A second weakness is the rapid pace of change. Models are retrained, retrieval layers are reconfigured, ranking heuristics are tuned. A finding that holds in March may not hold in September. The honest position is that the picture this article describes is the picture as of the most recent observation window. The structural tendencies are probably durable, but the specific magnitudes are not. Practitioners should design strategies that survive the structural picture being correct and the magnitudes being uncertain.
Finally, much of the available data depends on the queries researchers happen to study. Local commercial queries are over-represented because they are easy to operationalize. Long-tail B2B queries are under-represented because they are hard to sample. Generalizing from the well-studied query types to the under-studied ones is unavoidable but should carry explicit caveats.
Measuring directory ROI in 2025
Attribution models for AI referrals
Attribution in the AI era is harder than it was, and conventional last-click models fail in predictable ways. A user may meet a business in an AI Overview, never click through, and arrive at the business directly through a branded search the next day. Conventional analytics will credit the visit to organic branded search and miss the AI exposure entirely. Practitioners who keep judging directory ROI through last-click attribution are systematically under-counting AI-influenced revenue.
The methods that work are, broadly, branded-search lift studies (does branded search rise after directory presence is strengthened?), survey-based attribution (asking customers how they found the business), incremental lift testing where feasible, and brand-mention tracking across AI surfaces. None of these is perfect. In combination they produce a picture that is more accurate than last-click alone.
Brand mention tracking methods
Tracking brand mentions in AI outputs is a discipline still taking shape. Several commercial platforms now sample queries against the major LLMs and record which brands appear in responses. The methods vary, and the findings should be cross-checked rather than trusted from a single vendor. Statista, used by more than 23,000 companies for market data, illustrates the wider point that aggregated, multi-source measurement tends to beat any single feed.
For directories specifically, the metric that matters is share of voice within AI citations for queries in the directory’s domain. Tracked over time, this metric shows whether the directory’s structural position is improving or eroding, independent of traffic figures that may be confounded by interface changes outside the operator’s control. As shown in Table 2, the difference between conventional traffic metrics and AI-era visibility metrics is large enough that operators relying on the former alone will misread their position.
Table 2: Comparison of measurement frameworks for directory visibility
| Metric | Pre-AI relevance | Post-AI relevance | Recommended weight |
|---|---|---|---|
| Organic referral clicks | Primary | Diminishing | Moderate |
| AI citation share of voice | Not applicable | Primary | High |
| Branded search lift | Secondary | Increasingly important | High |
| Conversion-event tracking | Primary | Primary | High |
The reweighting implied by Table 2 matters for budget conversations. A directory whose organic referral clicks have softened but whose citation share of voice has risen is, in most cases, a directory whose position is strengthening rather than weakening, but only if the measurement framework is updated to see it.
Why listing quality beats listing quantity
The dominant practitioner reflex of the previous decade was to maximize listing quantity: the more directories a business appeared on, the better. The logic was reasonable when ranking systems treated each citation as a small additive vote and when there were enough viable directories to make the strategy economical. Both conditions have changed.
AI retrieval systems do not appear to treat citations additively the way classical ranking did. Past a saturation point, which most businesses reach with the well-known directories plus the relevant vertical leaders, additional listings on thin or low-authority sites add no detectable lift and may, in some cases, depress trust scores by association with low-quality neighbors. The marginal listing on a scraped or unverified directory is not zero-value. It is potentially negative-value. That is a real inversion of the prior playbook.
The economics have shifted too. Keeping accurate, consistent listings across hundreds of directories is hard work. NAP inconsistencies, stale hours, outdated photographs, and conflicting category assignments create exactly the kind of contradictory signal AI systems flag as low-confidence. A business with twenty well-maintained listings on authoritative directories is, in most cases, better positioned than the same business with two hundred sloppy listings across the long tail. The binding constraint is the discipline required to maintain accuracy at scale, not listing acquisition.
This logic generalizes. Quality, in the post-LLM context, is partly the inherent properties of each listing (completeness, freshness, structured markup) and partly consistency across listings (the same NAP, the same categorization, the same hours, the same description tone). A business that presents itself coherently across the directories where it appears generates a much stronger entity signal than one whose appearances contradict each other. Forrester’s research themes around grounded AI deployment make this point indirectly: the value of any single signal depends on how well it agrees with the others.

The practical result is that listing strategy now resembles portfolio management more than acquisition. Operators should periodically prune listings that have decayed below a quality threshold, invest in deepening listings on authoritative platforms, and treat consistency maintenance as ongoing operational work rather than a one-time setup task.
Building a directory strategy for AI visibility
Auditing existing citation footprint
The first step in any serious AI-visibility programme is an honest audit of the existing citation footprint. The audit should map every directory where the business appears, score each listing for completeness and accuracy, identify inconsistencies across listings, and compare citation share of voice against direct competitors. The output is a prioritized remediation list, not a comprehensive inventory. The goal is to find the highest-leverage interventions, not to document every imperfection.
Audits are usually run quarterly. The frequency reflects how fast the underlying systems change and how fast listings drift because of operational events (hours changes, address moves, staff turnover updating profiles). Annual audits are too infrequent. Monthly audits outpace the rate at which meaningful change can be implemented.
Prioritizing high-authority directories
Prioritization should follow the citation evidence. Directories that appear most often in AI citations for the business’s category and geography deserve the highest investment: complete profiles, professional photography, structured attribute completion, sustained review acquisition. Directories that rarely appear in citations deserve minimum-viable maintenance, meaning accurate NAP and current hours, but not deeper investment.
The list of high-authority directories is category-specific and changes over time. A medical practice’s priority list looks nothing like a restaurant’s, and a B2B software company’s list looks nothing like either. Practitioners who apply a generic template across categories will misallocate resources. For most practitioners, the tooling that automates citation distribution to a fixed list of generalist directories is less valuable than the editorial work of identifying the specific high-authority directories that matter for the specific category.
NAP consistency across sources
NAP (name, address, phone) consistency remains foundational. Inconsistencies arise constantly: a “Suite 100” rendered as “Ste 100” on one listing and “#100” on another, a phone number with a different area-code prefix format, a business name that includes a tagline on some listings but not others. Each inconsistency creates a small entity-resolution challenge that may, in aggregate, depress confidence scoring.
The discipline required is unglamorous and ongoing. Operators who designate a single canonical NAP record and propagate it systematically tend to do better than those who let each listing drift on its own. Verification processes that flag deviations from the canonical record before they spread are more effective than retrospective clean-up.
Optimizing descriptions for retrieval
Descriptions written for AI retrieval differ from descriptions written for human readers, and the differences are subtle. Retrieval systems benefit from explicit attribute mentions such as “open Sundays”, “wheelchair accessible”, and “appointments available within 48 hours” that human readers might infer from context. They benefit from category-precise vocabulary that matches how users actually phrase queries. They benefit from concrete specifics over abstract claims.
This does not mean descriptions should shrink to keyword lists. Models are sophisticated enough to tell naturally written descriptions from artificial ones, and the artificial ones are often penalized. The skill is to write naturally while making sure the attributes a query might surface are explicitly present in the text. The work resembles editorial copywriting more than legacy SEO, and the operators who do it well usually employ writers rather than offshoring it to template generators.
Monitoring LLM citation changes
Monitoring citation behavior over time is the closest thing to a control loop in this domain. The practice is to keep a stable query set, perhaps 50 to 200 representative queries for the business’s category and geography, and sample LLM responses against those queries on a defined cadence. The output is a longitudinal record of citation share, source mix, and competitor presence.
The cadence should match the rate of change in the underlying systems and the rate at which the operator can act on findings. Weekly is excessive for most operators, quarterly is the most common rhythm, and monthly suits operators with dedicated capacity. The data are most useful as trend lines rather than point estimates. A single bad sample is not a signal; a sustained decline is.
What the data tells practitioners
Pulling the evidence together, several conclusions hold up well. First, business directories have not been displaced by AI search; they have been repositioned within it. The traffic model has weakened for informational queries, while the visibility-and-citation model has strengthened for structured, verified, vertical-relevant directories. Operators who measure only the first will conclude, wrongly, that the channel is dying.
Second, the concentration within the channel has intensified. A small number of high-authority directories now account for most citations within most categories, and the gap between the head and the long tail has widened. New entrants face a higher authority threshold than they did a decade ago, but the threshold is not insurmountable, especially for vertical-specific propositions in under-served niches.
Third, structural quality (schema markup, taxonomy depth, verification rigor, NAP consistency) has become the binding constraint on citation performance. Directories that have invested in this infrastructure are gaining share; those that haven’t, aren’t. The investment is not glamorous but it is durable, and because authority signals compound, the gap between disciplined and undisciplined operators will keep widening.
Fourth, attribution and measurement frameworks need updating. Last-click attribution under-counts AI-influenced revenue, and share-of-voice metrics within AI surfaces increasingly belong in the executive dashboard alongside traditional traffic metrics. Operators who fail to update their measurement will misallocate budget away from channels whose contribution is real but invisible to legacy reporting.
Fifth, and most strategically, the post-LLM environment rewards specialization. Generalist directories competing on breadth face structural pressure from AI summaries that answer generalist queries directly. Specialist directories whose taxonomies, verification regimes, and review communities are tuned to a vertical keep, and often extend, their advantage. For new entrants the implication is clear: enter narrow, deepen ruthlessly, and earn authority through editorial discipline rather than scale acquisition.
The position is, on balance, more constructive than the prevailing narrative suggests. Directories are not relics. They are infrastructure for a retrieval paradigm that needs structured, verified, attribute-rich entity data more than the previous paradigm did. The operators who see this and respond are positioned to compound advantage over the next several years.
Action steps for the next quarter
For practitioners ready to turn the evidence into work, the next-quarter priority list is short and specific. Start with a citation-footprint audit run against a fixed query set representative of the business’s category and geography. The audit should produce a ranked list of the top fifteen to twenty directories where presence matters, and a remediation list of inconsistencies across the existing footprint.
Second, designate a canonical NAP record and propagate it to every listing on the priority list. This is unglamorous operational work and it is the highest-leverage intervention available to most operators. The cost is modest and the impact is disproportionate.
Third, audit and rewrite the descriptions on the top five directories in the priority list. The rewrites should be naturally written, attribute-explicit, and aligned with the vocabulary users actually use in queries. Avoid templating. The rewrite is editorial work and should be staffed accordingly.
Fourth, set up citation monitoring against the same query set used for the initial audit. The infrastructure can be light, since a spreadsheet and a monthly sampling discipline are enough for most operators, but the longitudinal record is what makes future decisions evidence-based rather than anecdotal.
Fifth, update the measurement framework you present to executives. Add share of voice within AI citations to the dashboard. Add branded-search lift as a leading indicator of AI exposure. Reduce the prominence of last-click referral metrics that no longer capture the channel’s contribution. The reframing is partly political, since executives have been trained to read the old metrics, but it is what makes the case for sustained investment defensible.
Beyond the next-quarter list, the longer-term work is cultural. Directory strategy in the AI era is editorial discipline applied at scale. The operators who build that discipline into their organizations, staffed appropriately, measured honestly, and sustained quarter after quarter, will gain advantage that operators who treat directories as a checklist activity will not. Deloitte’s legal practice, in framing the rising importance of intangibles and data, captures the wider pattern: the assets that compound now are the ones built through sustained discipline rather than one-time acquisition.
One question remains, and it is genuinely unresolved: as AI systems keep developing richer first-party entity graphs of their own, will the third-party directory layer this article describes keep its citation share, or will it be progressively absorbed into the platforms’ internal data structures? The evidence as of the most recent observation window suggests stability, but the structural pressure is real. The directories still cited in five years are likely to be the ones that provide signals (verification, review depth, taxonomic precision) that the platforms find genuinely costly to replicate internally. Whether that defensibility holds, and for how long, is the question practitioners and operators alike will be answering for the rest of this decade.

