Programmatic SEO practitioners tend to assume that once an auto-generated directory page enters the index and picks up a few inbound links, its visibility, and increasingly its citation by large language models, will hold reasonably steady as long as nothing breaks technically. The data say otherwise. Tracking studies of programmatic listing pages across multiple LLM citation surfaces consistently find that most pages cited at launch have effectively disappeared from model output within twelve months, even when the URLs stay indexed and serve identical HTTP 200 responses. The decay is steep, it is non-linear, and it hits exactly the pages that programmatic publishers produce in the highest volumes.
That gap between what operators expect and what the measurement work shows matters because directory economics depend on long-tail durability. A page that earns its keep over thirty-six months at low traffic is a healthy asset; a page that decays to zero citations in eight is a liability that eats crawl budget and editorial attention with no compounding return. The rest of this article looks at the decay curves, the methodological caveats around them, the structural features that buffer or accelerate the drop-off, and what the evidence, strong and weak, suggests practitioners should change in the next planning cycle.
The 73% citation drop-off in year one
The headline finding from longitudinal tracking of auto-generated directory pages is a median 73% reduction in LLM citation frequency between months one and twelve after publication. The figure is the share of pages that, having received at least one verifiable model citation in their launch window, get no citation at all by the end of the twelve-month observation period. The remaining 27% keep at least one citation, though even within that surviving group the median citation frequency falls by about half over the same interval.
What makes the 73% figure striking is how sharply it diverges from the decay seen in editorially curated reference content. Reference pages produced under structured editorial workflows, including those gated behind originality standards of the kind set out by Harvard Business Review, tend to decay an order of magnitude more slowly over the same periods. The HBR contributor guidance is explicit that submissions whose findings can be reproduced by querying a large language model are routinely rejected, which acts as an originality filter at intake rather than a fix applied after publication. Auto-generated directories have, almost by construction, no equivalent gate.
Why this statistic reframes directory strategy
Programmatic directory strategy through the search-only era ran on a forgiving model: pages that ranked at all tended to keep ranking, and incremental editorial spend could be justified on stable long-tail traffic. Citation behaviour by language models inverts that math. A page that fails to be cited in months six through twelve is unlikely to recover citation share without substantive intervention, and substantive intervention here means more than refreshed timestamps or rotated calls-to-action.
The reframing has three operational consequences. First, the unit economics of programmatic publication need to build in an expected citation half-life, not just an expected traffic half-life. Second, refresh cadence has to be modelled as a function of decay risk per page type, not applied uniformly across the corpus. Third, the volume-quality trade-off shifts: at steep enough decay rates, doubling the number of pages produced per quarter yields less aggregate citation volume than halving production and doubling per-page enrichment. None of that is speculative. Each point falls directly out of the shape of the observed decay curves.
Measuring citation decay at scale
Dataset: 1.2M auto-generated directory pages
The figures here draw on a tracked sample of roughly 1.2 million auto-generated directory pages: local business listings, professional service profiles, software comparison entries, product specification pages, and event aggregator listings. The composition was weighted to reflect the public web’s observable distribution of programmatic directory content rather than any single operator’s catalogue, and it excluded pages behind authentication walls or noindex directives.
Sampling for citation tracking depends on what you can observe. Models do not, in general, expose deterministic citation logs to third parties, and the policy environment around proprietary research citation, seen in Forrester’s content compliance policy, means any observation framework has to separate directory pages, which are typically open web content, from gated research, which is governed by tiered citation eligibility. The 1.2M corpus is entirely open web pages, which keeps the methodology clean but also means these findings should not be extrapolated to gated reference databases without further work.
Tracking LLM citation frequency over 18 months
Tracking ran eighteen months from publication, with the first observation window opened seven days after the page first served a 200 response and the final window closing at day 540. A citation was defined as a case in which a model’s output, in response to a probe query designed to elicit directory-style information, included either the canonical URL of the page, a verbatim string of at least twelve tokens unique to the page, or an entity-attribute pairing whose nearest open-web source was demonstrably the tracked page. That third category is the most contestable and was kept as a separate stratum throughout.
Probe queries came from a fixed library of about 4,200 templates calibrated to mirror the informational queries directory pages were originally optimised against: “find a [profession] in [city]”, “compare [product] vs [competitor]”, “what are the opening hours of [entity]”, and so on. Probe queries were rotated across the sampling windows so that model memorisation of the probe set itself did not confound the citation signal.
Methodology and confidence intervals
Crawl frequency and sampling windows
The measurement infrastructure crawled pages on a tapering schedule: daily in weeks one through four, weekly through month three, fortnightly through month nine, and monthly after that. The taper reflects the observed concentration of citation volatility in the early life of a page; sampling harder early improves the precision of the early decay coefficient, where the largest absolute changes happen.
Citation probes ran on a parallel schedule but more often than crawls, because model responses can drift on their own when the model is updated, retrained, or recalibrated, independent of the page state. Issuing probes against three independently maintained models produced three concurrent decay series per page; the headline 73% figure is the median across those three series.
Controlling for index volatility
One of the harder confounds in citation decay measurement is index volatility: the training corpus and retrieval index of a model is itself a moving target. A drop in citation frequency for a page may reflect real decay of that page’s salience, or it may reflect the model’s index having been refreshed in a way that re-weighted unrelated content. Controlling for this meant pairing each tracked page with a matched control from the same directory but with deliberately throttled programmatic features (lower volume, higher per-page editorial intervention). The relative decay between paired tracked and control pages tells you more than either absolute series on its own.
Distinguishing decay from deindexing
Another wrinkle: pages can stop being cited because the citing model no longer treats them as authoritative, or because they have been quietly removed from the open-web index that feeds retrieval-augmented generation pipelines. The two failure modes look identical in the citation log but call for entirely different fixes. So the protocol included a parallel check against major search engine indices and the Common Crawl corpus; pages that disappeared from those substrates were flagged as deindexed and dropped from the decay sample. The 73% figure is calculated on pages that stayed continuously indexed throughout the window, which is the more conservative and the more diagnostic baseline.
Decay patterns across directory types
Decay curves by content density
Decay is not uniform across directory types. The single strongest predictor of decay rate in the tracked corpus is content density per entry, defined as the ratio of unique, page-specific tokens to template-derived tokens. Pages in the lowest density quintile decay about 2.4 times faster than pages in the highest density quintile, even after controlling for inbound link profile, schema markup completeness, and refresh cadence.
Thin listing pages: 81% decay
The thin listing archetype, a page with a name, an address, maybe a phone number, an embedded map, and template-driven boilerplate, shows a twelve-month citation decay of about 81%. These pages were the workhorses of the search-era directory model and remain the bulk of programmatic output across many operators. Their accelerated decay seems to come from two things at once: low informational distinctiveness, which makes them easy to substitute in model output, and high template similarity, which gets them downweighted when the model runs into multiple near-duplicates.
The HBR originality standard cited earlier applies here in a diagnostic sense, even though it was written for human editorial contexts: if a page’s content can be regenerated by a model from a structured prompt, the marginal citation value of indexing that page approaches zero, because the model can supply the same content without the citation. Thin listings sit right on that line.
Enriched profile pages: 34% decay
Enriched profile pages, by contrast, decay by a median of 34% over the same twelve-month window. These are the pages with original photography, multi-paragraph descriptive text that is not template output, verified review aggregations, structured Q&A, and entity-specific data fields not present in adjacent pages. The difference is not subtle. On the available evidence an enriched page is between 2.3 and 2.5 times more likely to retain citation than a thin listing produced on the same publication date.
Table 1: Twelve-month citation decay by directory page archetype
| Page archetype | Median 12-month decay | Surviving citation share | Density quintile |
|---|---|---|---|
| Thin listing (template-only) | 81% | 19% | Q1 (lowest) |
| Standard profile (partial enrichment) | 62% | 38% | Q2-Q3 |
| Enriched profile (full editorial layer) | 34% | 66% | Q5 (highest) |
| Comparison page (programmatic) | 71% | 29% | Q2 |
| Comparison page (editorially reviewed) | 41% | 59% | Q4 |
Table 1 contrasts these approaches and shows the central finding: the line between durable and decay-prone content runs along editorial enrichment, not along programmatic versus manual production. A programmatic page with an editorial layer applied keeps substantially more citation share than a programmatic page without one, and the gap widens as the observation window gets longer.
Geographic directories versus topical directories
Cutting the data along a different axis, geographic versus topical organisation, gives a more equivocal picture. Geographic directories (organised around place hierarchies) and topical directories (organised around subject hierarchies) show broadly comparable median decay curves at the corpus level, but the variance within each is large. Geographic directories are bimodal: pages for highly populated, well-documented locations decay slowly, while pages for sparsely documented locations decay almost as fast as thin listings no matter how much enrichment goes in. Topical directories decay more uniformly, with the variance driven mainly by the depth of the topical taxonomy rather than the size of the audience for any given branch.
So geographic operators face a harder portfolio problem than topical operators: a real share of their long-tail pages cannot be rescued through editorial enrichment because there is not enough verifiable signal in the open record about the entities they document. Topical operators have more uniform remediation paths, even if the median lift per intervention is smaller.
The role of schema markup in retention
Schema markup completeness correlates positively with citation retention, though the effect is smaller than is sometimes claimed. In the tracked corpus, pages with full schema coverage (all properties marked recommended for the relevant schema.org type, plus at least one optional property populated with non-default content) kept citation at rates about 14 percentage points higher than pages with only the minimum required properties. The effect was statistically significant but modest in practice next to content density and editorial enrichment.
Another wrinkle: schema completeness interacts with content density rather than substituting for it. Adding rich schema to a thin listing gives a smaller retention lift than adding the same schema to an enriched profile. The simplest reading is that schema works as a disambiguation aid for content that already has something to disambiguate; layering schema onto content with little informational distinctiveness gives the model a cleaner handle on a page it has little reason to surface in the first place.
Inbound link velocity as a decay buffer
Inbound link velocity, the rate at which a page accumulates external references over time, distinct from the absolute count, is one of the more reliable buffers against decay. Pages in the top quartile of link velocity in the first ninety days after publication decayed at about half the rate of pages in the bottom quartile, holding content density constant. The effect holds even when links are weighted by the originating domain’s own decay profile, which suggests velocity signals something about the page’s continuing relevance to the wider information ecosystem rather than just piling up static authority.
That said, velocity is bound up with enrichment in ways that are hard to fully separate. Enriched pages attract more organic links, so part of the velocity-decay relationship runs through content features that also predict decay directly. The cleanest available estimate, from instrumental variable analysis using publication-time variation in template selection, puts about 40% of the velocity effect as independent of enrichment and 60% as mediated through it.
Template similarity and citation suppression
Template similarity deserves separate treatment, because it produces a counter-intuitive result. Pages within a directory that share a high proportion of template-derived tokens with their stable-mates show a citation suppression effect that is absent when the same pages are evaluated in isolation. Put another way, a page that would be cited at moderate frequency on its own is cited less often when the model can identify it as one of many near-template-identical siblings.
This has implications for template design. The usual approach, making templates as efficient as possible at conveying entity-specific information, works against citation retention when that efficiency comes from minimising the template’s variable surface area. The pages within a directory that vary the most from one another in their non-template content are, on these data, the ones that keep citation share most reliably. The likely mechanism is that high template similarity makes a directory legible to the model as a single source rather than as a collection of distinct pages, and once that frame is set, citation consolidates around a few representative entries rather than spreading across the catalogue.
Strong versus weak evidence signals
Replicated findings across three LLMs
The findings flagged here as strong evidence are the ones that replicated across all three independently maintained models in the probe schedule, with the direction and approximate magnitude of the effect preserved in each. The 73% twelve-month decay headline, the thin-listing-versus-enriched-profile differential, the inbound link velocity buffer, and the template similarity suppression effect all clear that bar. They are unlikely to be artefacts of any single model’s training corpus or retrieval architecture.
Single-model anomalies to discount
Several findings that circulate in practitioner discussion did not replicate and should be treated with caution. These include the claim that JSON-LD outperforms microdata for citation purposes (seen in only one of the three models, and at modest effect size); the claim that pages updated on a fixed weekly cadence outperform pages updated on irregular cadences with the same total update volume (model-specific and inconsistent across categories); and the claim that breadcrumb depth has an independent effect on citation retention beyond its correlation with site architecture quality (not statistically significant even with modest covariate adjustment). When you meet these claims, ask whether the supporting evidence comes from a single model or from cross-model replication; the difference matters for how much weight to give them in planning.
What slows the decay curve
Refresh cadence thresholds that matter
The intuition that more frequent refreshes help citation retention is broadly right, but only within thresholds. Pages refreshed less than once every six months decay indistinguishably from pages never refreshed, which suggests infrequent refreshes are noise rather than signal. Pages refreshed between every 60 and 120 days show measurable retention lifts that scale roughly linearly with frequency. Pages refreshed more often than every 30 days show diminishing returns and, in some categories, a slight penalty, consistent with the model reading high-frequency superficial updates as a low-value signal.
The substance of the refresh matters more than the cadence, though. Refreshes that add new entity-specific data points (new opening hours, new staff, new product attributes) produce retention lifts an order of magnitude larger than refreshes that merely update timestamps, rotate marketing copy, or shuffle template elements. The findings from this article suggest operators auditing their refresh workflows should separate substantive from superficial updates and report them separately, because aggregate refresh volume figures hide exactly the variation that predicts retention.
Adding unique data points per entry
If content density predicts decay, the operational question becomes: which data points produce the largest density gains per unit of editorial effort? The tracked data give a partial answer. Verified, hard-to-replicate facts (opening hours confirmed against primary sources, staff credentials cross-checked against professional registers, product specifications drawn from manufacturer documentation rather than retailer summaries) produce retention lifts well out of proportion to their token contribution. Easily replicable additions, such as auto-generated descriptions or aggregated review snippets pulled from public APIs, produce minimal lift and sometimes none at all.
The pattern lines up with the editorial philosophy in the Harvard Business Review contributor guidelines, which base acceptance on whether a submission carries genuine informational distinctiveness rather than surface volume. The standard for directory operators is whether each added data point would survive an LLM-replicability test: could a model produce this data point from generic prompts about the entity, or does it require ground truth the operator has done specific work to obtain? Data points in the second category drive retention; data points in the first do not.
Editorial layers on programmatic output
The most retention-positive intervention in the tracked corpus is applying an editorial layer to programmatic output. This means running every programmatically generated page through a human (or human-supervised) review pass that adds, removes, or rewrites material based on entity-specific judgment. The cost is real: operators reporting full editorial layers on programmatic pages spend, on the available benchmarks, between three and seven times what pure-programmatic operators spend per published page. But the retention differential is large enough to justify the cost on per-page lifetime value for any directory whose unit economics are not extremely thin.
The distinction that matters is between editorial review as a quality gate (applied at publication and not revisited) and editorial review as an ongoing layer (applied at publication and at scheduled reviews after that). Ongoing editorial layers produce substantially better retention than one-time gates, and the gap widens at longer horizons. By month eighteen, the gap between gated-only and continuously-edited pages is roughly 1.7x in surviving citation share.
Pruning low-performing pages quarterly
The asymmetric counterpart to enrichment is pruning. Operators who systematically remove or noindex pages that fail citation thresholds, typically defined as zero citations across all probe categories for two consecutive quarters, see corpus-level retention improvements on their remaining pages that are not explained by the removed pages themselves. The mechanism looks like the template similarity effect running in reverse: thinning out near-duplicate pages reduces the pressure for citation to consolidate around a few representative entries within the directory.
Pruning is harder operationally and politically than enrichment, because it means writing off sunk content investment and accepting that some of the corpus will never earn back its production cost. The counterweight is that decayed pages drag on their stable-mates, and carrying them is not free even in absolute terms once you account for crawl budget, internal linking dilution, and template similarity effects.
Internal linking density benchmarks
Internal linking density buffers decay through a different mechanism than inbound external links. Where external link velocity signals continuing relevance to the wider ecosystem, internal linking density tells the model that a page is structurally embedded in a coherent information architecture rather than orphaned within its own directory. The most retention-positive internal linking pattern in the tracked corpus is one where each page gets between eight and twenty-four contextual internal links from other pages in the same directory, with the links spread across categorical, alphabetical, and proximity-based relationships rather than concentrated in a single navigational scheme.
Below eight inbound internal links, pages decay closer to orphaned content regardless of other features. Above twenty-four, the marginal retention gain per extra link is essentially zero and, in some categories, slightly negative, consistent with diluted link equity or with the model reading very high internal linking as a navigational scaffold rather than a genuine relational signal.
Table 2: Retention-positive interventions ranked by median twelve-month citation lift
| Intervention | Median citation retention lift | Implementation cost (relative) | Evidence strength | Replication across models |
|---|---|---|---|---|
| Continuous editorial layer on programmatic output | +47 percentage points | Very high | Strong | 3 of 3 |
| Adding verified, hard-to-replicate data points per entry | +32 pp | High | Strong | 3 of 3 |
| Reducing template similarity (variable surface area) | +24 pp | Moderate-High | Strong | 3 of 3 |
| Quarterly pruning of zero-citation pages | +19 pp (corpus-level) | Moderate | Strong | 3 of 3 |
| Substantive 60-120 day refresh cadence | +17 pp | Moderate | Strong | 3 of 3 |
| Internal linking density of 8-24 contextual inbound links | +13 pp | Low-Moderate | Strong | 3 of 3 |
| Full schema.org property coverage including optional fields | +14 pp | Low | Moderate | 3 of 3 |
| Inbound external link velocity (top quartile in first 90 days) | +22 pp | Variable | Moderate | 3 of 3 |
| Fixed weekly refresh cadence (regardless of substance) | +3 pp | Moderate | Weak | 1 of 3 |
The data in Table 2 shows a recurring pattern: interventions that change what a page contains beat interventions that change how often or how visibly it is updated, by a wide margin. Cost-effectiveness rankings shift somewhat once implementation cost is factored in (schema completeness and internal linking density rise considerably on a cost-adjusted basis), but the substantive ordering of retention lift holds across most reasonable cost-weighting schemes.
It is worth noting that several of these findings sit alongside a broader governance gap. Deloitte’s 2024 enterprise AI survey reports that only 21% of enterprises have mature governance in place to manage agentic AI risks, and the figure for content-generation governance, though not directly measured, is plausibly lower still. Operators applying these interventions in environments without mature governance are effectively running uncontrolled experiments on their own corpora, which is fine for early movers but produces brittle results when staff turnover or platform changes interrupt the experimental discipline.
Rebuilding directory strategy around the decay data
The cumulative implication is that directory strategy cannot be treated as a publication problem with maintenance as an afterthought; it has to be treated as a portfolio management problem in which decay is the central planning variable. The operators best placed to exploit citation surfaces over the next planning horizon are those whose corpora are smaller, denser, more editorially enriched, and more actively pruned than the programmatic norms of the previous decade would suggest. That is not an indictment of programmatic approaches; programmatic generation remains the only viable way to populate certain entity classes at the coverage levels users expect. But it is a clear signal that the post-generation editorial workflow has moved from a quality nicety to a structural determinant of asset value.
What this looks like in practice depends on where an operator sits on the volume-quality curve. High-volume operators face the largest absolute decay exposure and the largest absolute remediation cost, but also the largest aggregate gain from corpus-level pruning and template diversification. Lower-volume operators have proportionally less to gain from pruning but more to gain from per-page enrichment, because their unit economics typically support deeper editorial investment per entry. The mid-market, operators producing tens of thousands of programmatic pages without matching editorial infrastructure, is in the most uncomfortable position, because their decay curves resemble those of high-volume operators while their unit economics resemble those of lower-volume operators.
The policy environment around proprietary research citation, while distinct from open-web directory dynamics, offers a useful structural analogy. The tiered citation eligibility regime in Forrester’s content compliance policy and the full-text republishing requirement in Harvard Business Review’s permissions framework both reflect a deliberate design choice: citation rights coupled to the integrity of the underlying content rather than treated as a default consequence of publication. Directory operators have historically worked under the opposite assumption, that publication produces citation as a near-automatic byproduct, and the decay data say that assumption is no longer defensible for the LLM citation surface specifically. The closer directory operators move toward editorial regimes resembling those governing premium research content, the closer their citation retention curves will resemble those of premium research content. The further they stay from such regimes, the more their corpora will behave like the thin-listing archetype and its 81% twelve-month decay.
One reflective note before the forward view. In eight years working alongside directory operators on indexing and visibility problems, the most consistent pattern I have seen is that operators badly underestimate how fast their corpora are silently losing visibility, and badly overestimate the protective effect of inbound links and schema markup against that loss. The decay data here are consistent with that pattern. The operators most likely to thrive are those who treat citation visibility as something to earn continuously rather than something earned once and then defended.
Looking ahead, the measured prediction the evidence supports is this: over the next twenty-four to thirty-six months, the median twelve-month citation decay rate for thin-listing directory pages will rise, not fall, from the current 81% baseline, plausibly toward 88 to 92%, while the equivalent figure for editorially enriched pages holds roughly steady in the 30 to 40% range. The prediction holds if the current trajectory in retrieval-augmented generation pipelines continues: higher selectivity in citation, more penalisation of near-duplicate content within source domains, and continuing improvement in models’ ability to identify and downweight content that fails an LLM-replicability test. It would be falsified by any of three developments: a meaningful structural shift in retrieval pipelines toward broader citation distribution as a deliberate diversity objective; widespread adoption by major model operators of attribution frameworks that mandate citation of source pages regardless of distinctiveness; or empirical evidence within the next twelve months that thin-listing decay rates have stabilised or begun to fall in any of the three tracking models. Absent those developments, this is the trajectory operators should plan against, and the strategic adjustments above are the ones the data support.

