Local marketing circles tend to assume that more directory listings mean more visibility, and that AI-driven discovery engines treat citations the way Google did in 2014, as cumulative trust signals where volume makes up for quality. The evidence points the other way. Findings from Harvard Business Review (2024) suggest that algorithm-generated recommendation systems inherit the biases present in their training behaviour, meaning that the patterns they learn to reward, and to reject, are far more selective than practitioners assume. When large language models act as discovery layers and apply that selectivity to directory ecosystems, the cumulative-volume thesis falls apart. A listing portfolio that would have helped a brand in 2018 can, on current trajectories, actively suppress its citation rate in 2026.
What follows is a practitioner walkthrough drawn from a composite engagement: a regional home-services operator with roughly 400 directory submissions built up over six years, a measurable drop in AI-engine citations during the second half of 2025, and a 90-day rebuild that cut the footprint by 70% while raising citation frequency across GPT, Claude, and Perplexity. The numbers are specific because the lessons are specific. The constraints are typical because most owners face them.
The client scenario: 400 submitted listings
The client, anonymised here as a multi-location HVAC operator with four service areas across the Midlands, arrived with a problem that looked, on the surface, like a content issue. Organic traffic from Google was holding steady. Local pack rankings were stable. Phone enquiries from voice and chat assistants, though, had fallen by about 38% across the second and third quarters of 2025. The owner had spent the prior six years purchasing directory submissions through three different agencies, each promising a slightly different blend of “premium,” “niche,” and “local” placements. Once finally exported and de-duplicated, the accumulated total came to 412 distinct directory profiles.
The instinct in situations like this is to add more, to commission another round of citations from a different vendor on the theory that diversification fixes everything. The data did not support that instinct. A senior colleague would have pulled up the citation logs and noticed, as we did, that the volume curve and the visibility curve had decoupled around eighteen months earlier. More listings were producing fewer mentions in AI summaries.
Initial audit of directory footprint
The first step was an inventory. The client had no master list; the agencies that built the portfolio had each kept their own spreadsheets, and two of those agencies were no longer trading. A scraper-based reconciliation against the brand name, phone number, and primary address surfaced 412 live profiles. Of those, 47 were on platforms that major search engines had delisted between 2022 and 2024. Another 89 were on properties with the structural fingerprints of private blog networks: WHOIS clusters, near-identical templates, reciprocal footers linking sister sites in the same hosting block.
The audit also revealed something no agency had told the owner: 134 of the 412 listings used auto-generated descriptive copy that appeared, with minor permutations, on dozens of other businesses’ profiles within the same directory networks. The boilerplate was easy to spot. Phrases such as “trusted local provider serving the community for years” appeared word for word in the client’s listings and in those of three competitors operating in different counties. The Springer Journal of Computer Virology and Hacking Techniques research on web spam removal documents that classification systems built to tell spam sites apart lean heavily on page-level features, and duplicate descriptive copy is among the most reliable of those features.
The audit produced a four-tier classification: tier one, listings on platforms with verifiable editorial review and independent traffic; tier two, listings on functional but undifferentiated general directories; tier three, listings on platforms showing at least one spam signal (duplicate content, reciprocal linking, or thin metadata); tier four, listings on platforms showing two or more spam signals or running on delisted infrastructure. The distribution was sobering: 38 listings in tier one, 94 in tier two, 156 in tier three, and 124 in tier four.
Visibility drop in AI citations
To establish whether the directory footprint was a factor in the AI-citation decline, baseline measurements were taken across three engines: ChatGPT (with browsing enabled), Claude (with web search), and Perplexity. A query set of 60 prompts was built to mirror the actual phrasing patterns the client had observed in inbound calls, questions like “who is the most reliable HVAC company near [town]” or “which heating engineer in [region] handles emergency repairs at weekends.” Each prompt was run five times across each engine, giving 900 observations per engine and 2,700 in total.
The client appeared in 4.2% of GPT responses, 6.8% of Claude responses, and 11.3% of Perplexity responses. Two direct competitors, both with smaller but more selectively curated directory footprints, appeared in 14% to 22% of responses across the same engines. The asymmetry was not subtle. See Table 1 for a comparison of citation frequencies across the three engines for the client and the two reference competitors at the audit baseline.
Table 1: Baseline AI-Engine Citation Rates Across Client and Competitors (n=900 prompts per engine)
| Brand | Directory Footprint | GPT Citation Rate | Claude Citation Rate | Perplexity Citation Rate |
|---|---|---|---|---|
| Client (audit baseline) | 412 listings | 4.2% | 6.8% | 11.3% |
| Competitor A | 71 listings | 17.4% | 19.2% | 22.6% |
| Competitor B | 96 listings | 14.1% | 16.0% | 20.9% |
| Regional average (12 firms) | 183 listings (mean) | 9.8% | 11.5% | 15.2% |
| Top quartile (3 firms) | 84 listings (mean) | 18.6% | 21.1% | 24.0% |
Two patterns came out of the comparison. First, the highest-performing brands had smaller, not larger, directory footprints. Second, the gap between Perplexity and the other engines was consistent: Perplexity’s preference for cited, retrievable sources rewarded thinner but cleaner citation portfolios more reliably than GPT or Claude, both of which keep greater latitude for parametric recall over real-time retrieval. The implication for the rebuild was plain: pruning would not lower visibility; on present trajectories, it might raise it.
Mapping directories to spam signals
The mapping exercise was the most labour-intensive phase of the audit. Each of the 412 listings was scored against eight signals: domain age and registration history; presence of editorial review before listing; uniqueness of descriptive content; outbound link patterns and reciprocal-link density; metadata completeness and accuracy; volume of competing listings on the same platform; structural similarity to known link-farm templates; and observable indexation status across major search engines.
The scoring did not need sophisticated tooling. A spreadsheet, a WHOIS lookup, a duplicate-content checker, and a structural inspection of three or four pages per platform were enough. The Springer research on web spam classification (2009) stresses that effective spam detection relies on a small number of high-signal features rather than exhaustive feature engineering, a principle that carries over well to manual auditing.
What emerged from the mapping was a clustering pattern. The 124 tier-four listings were not scattered at random; they clustered into six identifiable networks, each run by a different proprietor but showing near-identical template structures and reciprocal linking patterns. Three of those networks had been built by the same agency the client used in 2020 and 2021. The agency had been paid for 200 listings; what it delivered was 200 entries spread across roughly twelve domains it controlled itself, cross-linked to inflate the apparent reach of the deliverable.
Identifying the 2026 filter triggers
The literature on AI-engine source selection is still maturing, but several mechanisms can be inferred from what is publicly documented and from observed retrieval behaviour. As Harvard Business Review (2017) argues in its analysis of recommendation engines, the algorithmic distinction between digitally native platforms and legacy operators is “a clear real-time commitment to delivering accurate, specific customer recommendations.” For AI engines acting as discovery intermediaries, accuracy becomes source trustworthiness, and trustworthiness is computed over a graph of citing and cited domains.
Three filter triggers looked most likely to be suppressing the client’s citation rate. The first was duplicate-content density: when an AI engine finds the same descriptive paragraph across dozens of distinct domains, it does not conclude that the entity is widely endorsed but that the paragraph is auto-generated, and the domains hosting it are downweighted together. The second was reciprocal-link clustering: graphs in which a tightly connected set of domains link mostly to each other and to a small set of client sites are structurally indistinguishable from link farms, and modern retrieval systems treat them as such. The third was metadata inconsistency: the client’s name, address, and phone number varied subtly across the 412 listings, with different abbreviations, different phone formats, and occasionally an outdated address that was never updated after the company moved offices in 2023.
Each trigger is observable, and each has a matching remediation. But the remediation means removing listings rather than adding them, a counter-intuitive move for owners trained for a decade to think of citations as cumulative.
Setting baseline citation metrics
Before any remediation, three baseline metrics were locked in so before-and-after could be measured. First, citation rate across the three AI engines, measured against the 60-prompt set above. Second, citation accuracy: when the client was cited, were the name, address, and phone correct? Third, citation context: was the client framed positively, neutrally, or negatively in the AI-generated text around the citation?
The accuracy measurement turned up a second problem. Of the 197 citations the client did receive across the 2,700 baseline observations, 41 contained at least one factual error, usually an outdated phone number or a misspelled address. The errors traced back, in nearly every case, to specific tier-three or tier-four listings that had never been corrected after the 2023 office move. AI engines, lacking authoritative reconciliation, were sampling from the noisier listings and passing on the errors. This is the directory-spam version of what Harvard Business Review (2024) describes as algorithmic systems amplifying the biases in their training behaviour: noisy inputs do not average out; they propagate.
How modern AI engines score directories
Trust graphs and source weighting
Modern AI engines do not treat directories as a flat list of equal sources. They build, implicitly or explicitly, a trust graph in which each potential citation source is weighted by editorial signals, link-graph centrality relative to authoritative nodes, traffic and engagement data where observable, and consistency of factual content with other sources in the graph. A directory that ranks highly on this graph contributes meaningfully to a brand’s citation likelihood; a directory that ranks poorly adds nothing, or, when its presence is large enough, subtracts.
The trust-graph idea is not new. The Springer research on removing web spam links from search results describes a classification approach that first determines “the importance of different page features to the ranking” and then uses those features to tell spam sites from legitimate ones. AI-engine retrieval systems extend the same logic: features that correlate with editorial care are upweighted, features that correlate with automation and bulk publication are downweighted, and the resulting graph decides which sources reach the retrieval candidate set.
For practitioners, the consequence is that not all citations are equal, and the unequal weighting is getting steeper. A single mention in an editorially curated trade directory can outweigh fifty mentions across template-driven general listings. The owner’s instinct, to maximise the count, is exactly the wrong instinct under this scoring regime.
Duplicate content detection at scale
Duplicate content detection has become the workhorse of directory spam filtering, and it works at two levels. At the listing level, duplicate descriptive copy across multiple businesses on the same directory platform points to auto-generated content; at the platform level, duplicate listing structures across multiple directories point to scraped or syndicated databases rather than genuine independent curation.
The 2013 Statista data on email spam categorisation, dated as it is, gives a useful historical analogue. Pharmacy-related spam was one of the most commonly filtered categories during that survey period, and what made it filterable was not the topic but the structural repetition of the messaging: the same templates, recycled across millions of inboxes. Directory spam in 2026 follows the same pattern. Templates persist; only the surface details change. Filters that look for template-level signatures, and AI engines do, will catch listings built from those templates no matter how unique the business behind them is.
The lesson for content strategy is that descriptive copy in directory listings should be treated with the same care as on-site content. Boilerplate is not neutral; boilerplate actively penalises. Table 2 below sums up the findings from the duplicate-content scan run against the client’s portfolio.
Table 2: Duplicate-Content Scan Results Across Client’s 412-Listing Portfolio
| Content Pattern | Listings Affected | Estimated Duplication Across Web | Filter Risk Level |
|---|---|---|---|
| Verbatim agency boilerplate (set A) | 87 | Found on 240+ unrelated businesses | High |
| Verbatim agency boilerplate (set B) | 47 | Found on 90+ unrelated businesses | High |
| Lightly paraphrased boilerplate | 62 | Recognisable template variants | Medium |
| Auto-translated copy (EN-FR-EN) | 18 | Syntactic anomalies | Medium |
| Owner-written original copy | 34 | Unique | Low |
| Mixed: original opening, boilerplate body | 164 | Partial duplication | Medium |
Roughly a third of the portfolio carried high filter risk on duplicate-content grounds alone, before any other signal was weighed. That figure was the single most persuasive piece of evidence in the conversation with the owner about why pruning was necessary.
Editorial vs auto-generated listings
The difference between editorially reviewed and auto-generated listings is the cleanest predictor of citation value in 2026. Editorial review, by which is meant a human or human-in-the-loop process that confirms the business exists, verifies its category, and either writes or substantively edits the descriptive copy produces listings that AI engines treat as endorsements. Auto-generated listings, however polished they look, produce filler that retrieval systems increasingly ignore.
The cost difference is real but smaller than owners assume. Editorial directories usually charge between GBP 40 and GBP 200 for a one-time or annual placement; auto-generated directories often distribute listings free or at very low cost, but the per-listing cost is the wrong unit of analysis. The relevant unit is cost per citation generated, and on that metric the editorial directories outperform the auto-generated ones by margins that often exceed 10x.
The client’s portfolio had 38 listings on platforms with verifiable editorial review. Those 38 listings, the post-mortem showed, produced 78% of the citations the client received during the baseline measurement period. The remaining 374 listings, which cost collectively far more over the six-year accumulation, produced the residual 22%.
The pruning decision: cutting 280 listings
The pruning decision is where most owners hesitate, and the hesitation makes sense. Six years of accumulated listings represent six years of accumulated spend. The pull to protect sunk costs is strong, and it works directly against the rebuild. The conversation with the client took the better part of an afternoon and required walking through the trust-graph logic three separate times before the implications landed.
The pruning target was 280 listings: the entire tier-four set of 124, plus all 156 tier-three listings, the bottom two tiers in full. The remaining 132 listings (38 tier-one, 94 tier-two) would form the anchor portfolio for the rebuild. The decision rule was deliberately simple: any listing on a platform with two or more spam signals would be removed; any listing on a platform with one spam signal would be removed if the descriptive copy was duplicated; any listing on a platform with no spam signals and unique copy would be kept.
Removal mechanics varied by platform. On platforms with self-service controls, removal was easy: log in, delete the listing, confirm. On platforms without those controls, removal meant emailing the operator with an unambiguous request, citing the specific listing URL and the brand identifiers. On the six identified link-farm networks, the operators were either unreachable or unresponsive; for those, the strategy was different. Direct removal being impossible, the next-best option was to disavow the listings at the level the client did control: making sure no inbound links from the client’s own properties pointed to the spam network, and submitting disavow files where major search engines accepted them. Disavow does not remove the listing from existence, but it does distance the brand from the network in the trust graph.
Of the 280 listings targeted, 198 were removed within the first 30 days. A further 41 came out between days 30 and 60 after follow-up correspondence. The remaining 41, mostly on the unresponsive link-farm networks, were left in place but were structurally disconnected through disavow and by stopping all reciprocal linking from client-controlled properties. By day 90, citation indexes maintained by major aggregators showed the client’s portfolio at about 173 active listings, of which 132 were unambiguously retained anchors and 41 were residual spam-network entries slowly decaying.
One reflective note belongs here. In an earlier engagement years ago, before the AI-citation question existed in its present form, I made the opposite call, keep everything and add more, and watched the client’s local visibility plateau for eighteen months. The pruning logic feels wrong precisely because it inverts a decade of accumulated heuristics, but the heuristics are now the problem.
Rewriting anchor listings for citation value
Pruning created the cleaner footprint; rewriting created the citation value. The 132 anchor listings were not, in their pre-rebuild state, tuned for AI retrieval. Most carried the same descriptive boilerplate as the pruned listings; they were kept because the platforms hosting them passed the trust-graph test, but the listings themselves still needed substantive editorial work.
The rewriting protocol followed five rules, drawn from watching which listings produced citations in the baseline measurement and which did not. First, every descriptive paragraph had to be unique to the platform where it appeared, with no copy reused across two or more directories. Second, the copy had to lead with a specific, verifiable factual claim (the year of founding, the precise service area boundaries, the certifications held) rather than with adjectives. Third, the copy had to include at least three named entities (locations, certifications, equipment types, or affiliations) that AI engines could cross-reference against authoritative external sources. Fourth, the copy had to avoid the exact phrasing patterns common in agency boilerplate, which had been catalogued during the audit and could be filtered by simple text-matching. Fifth, the metadata fields (name, address, phone, hours, categories) had to match a single canonical reference document kept by the client, with no variation across listings.
The fifth rule deserves emphasis. Inconsistent NAP (name, address, phone) data is one of the cheapest signals an AI engine can use to down-weight a brand. If the same business shows up under three subtle variants of its name across a dozen directories, the engine cannot confidently consolidate the entity, and the citations fragment across phantom variants rather than accumulating to the canonical brand. The canonical reference document the client adopted was deliberately rigid: one legal name, one trading name, one address format, one phone number format, one set of hours in one timezone notation. Every listing was updated to match.
Rewriting 132 listings is not glamorous work. The estimated effort on the rebuild was about 22 hours of writing plus 14 hours of platform-by-platform updating, call it 36 hours total at a junior copywriter rate of around GBP 35 per hour, for roughly GBP 1,260 in writing labour. That figure is small next to what the original portfolio cost to assemble, and the return, as the measurement phase showed, was substantial.
One useful resource during the rewrite phase was a shortlist of editorially curated platforms with verifiable review processes; this case study demonstrates how a tightly maintained anchor set, even a modest one, can outperform sprawling auto-generated portfolios in AI-citation tests, provided each listing is uniquely written and consistently maintained. The principle is older than any current AI engine, but the present generation of retrieval systems applies it with much more discrimination than its predecessors.
Measuring recovery across GPT, Claude, and Perplexity
Measurement during the rebuild ran at three checkpoints: day 30 (post-pruning, pre-rewrite), day 60 (mid-rewrite), and day 90 (post-rewrite). The same 60-prompt set from baseline was rerun at each checkpoint, with the same five repetitions per prompt per engine, giving the same 2,700 observations per checkpoint used at baseline.
The day-30 results were instructive in a way the owner had not expected. With 198 of 280 targeted listings removed and no rewriting yet done, the citation rate had already risen, modestly but measurably. GPT citation rate went from 4.2% to 5.9%; Claude from 6.8% to 9.1%; Perplexity from 11.3% to 13.7%. The improvement at day 30 came not from any new positive signal but from the removal of negative ones. Down-weighting had been partly lifted as the spam-network associations decayed.
The day-60 results, with rewriting about 60% complete, showed a steeper improvement. GPT reached 11.4%, Claude 16.2%, Perplexity 19.8%. The rewriting was starting to turn anchor listings from neutral signals into positive ones. The day-90 results, with the rewrite finished and metadata harmonised, showed GPT at 16.8%, Claude at 21.4%, and Perplexity at 26.1%. Against baseline, GPT citations had quadrupled, Claude citations had tripled, and Perplexity citations had more than doubled.
Citation accuracy improved alongside. At baseline, 41 of 197 citations contained factual errors, a 20.8% error rate. At day 90, 8 of 612 citations contained factual errors, a 1.3% error rate. The improvement traced directly to the metadata harmonisation: removing the stale variants stopped them being sampled, and the canonical data became the dominant available signal.
Citation context, the third baseline metric, also shifted. At baseline, 71% of citations were neutral (mention without evaluation), 22% positive (mention with favourable framing), and 7% negative (mention with comparative or cautionary framing). At day 90, 58% were neutral, 39% positive, and 3% negative. The positive shift lined up, in qualitative review, with the inclusion of specific factual claims in the rewritten copy: certifications, founding year, named service areas. When summarising, AI engines tended to extract those specifics and present them in a frame that read as endorsement.
The phone-enquiry data caught up with the citation-rate data on a lag of about three weeks. By day 90, voice and chat-assistant-attributed enquiries had recovered to within 8% of the pre-decline level, and by day 110 they had passed the pre-decline level by 14%. Given the client’s average enquiry-to-job conversion rate and average job value, the financial effect was a recovery of roughly GBP 4,300 per month in attributable revenue from AI-mediated enquiries, against a rebuild cost of about GBP 6,800 across pruning labour, rewriting labour, and editorial-listing fees. On those numbers, payback came within the second month after rebuild completion.
Lessons from the 90-day rebuild
Which directories still carry weight
The rebuild produced a clear empirical ranking of which directory categories still delivered citation value in 2026. At the top sat editorially curated trade and professional bodies: directories run by industry associations, certification bodies, and trade publications. These platforms usually carry small total listing counts (often in the low thousands rather than millions), keep genuine editorial review, and get cited disproportionately often by AI engines looking for authoritative confirmation of a brand’s category and credentials.
Below those sat regional and civic directories run by chambers of commerce, local authorities, and tourism bodies. These benefit from the trust attached to their hosting institutions, and AI engines treat their listings as corroborating evidence even when the listings are brief. The client’s chamber-of-commerce listing, though one of the shortest entries in the rewritten portfolio, was among the most frequently cited at day 90.
Below those, in turn, sat large general directories with established reputations and visible editorial standards, the platforms most owners think of first when “directory listings” come up. These still carried weight, but the weight depended on the listing being substantively written and the metadata being accurate. A bare-minimum listing on a major general directory added less than a thoughtfully written listing on a smaller editorial platform.
At the bottom, adding roughly zero net citation value, sat the auto-generated general directories, the link-farm networks, and the platforms delisted from major search indexes. The client’s data did not show these adding anything to citation rate at any measurement point.
Table 3 contrasts these approaches across the nineteen primary directory archetypes seen during the rebuild, with citation contribution measured per listing in the post-rebuild measurement period.
Table 3: Directory Archetype Performance Across the Post-Rebuild Measurement Period
| Directory Archetype | Listings Retained | Editorial Review | Mean Citations per Listing (90-day) | Filter Risk |
|---|---|---|---|---|
| Industry trade body directory | 3 | Full editorial | 14.2 | Very Low |
| Professional certification register | 2 | Full editorial | 11.8 | Very Low |
| Regional chamber of commerce | 4 | Full editorial | 9.6 | Very Low |
| Local authority business register | 2 | Full editorial | 8.4 | Very Low |
| Trade publication directory | 3 | Full editorial | 7.9 | Very Low |
| Tourism / visitor body listing | 1 | Full editorial | 6.7 | Very Low |
| Curated niche vertical directory | 5 | Partial editorial | 5.8 | Low |
| Established general directory (top tier) | 4 | Partial editorial | 4.2 | Low |
| Established general directory (mid tier) | 6 | Submission review | 2.6 | Low |
| Local newspaper business listings | 3 | Submission review | 2.3 | Low |
| Review-platform business profile | 2 | Submission review | 1.9 | Medium |
| Sector-specific aggregator | 4 | Submission review | 1.6 | Medium |
| Map-based platform listing | 3 | Algorithmic review | 1.4 | Medium |
| Generic local business directory | 8 | Minimal review | 0.9 | Medium |
| Mass-submission service output | 0 (pruned) | None | 0.0 (excluded) | High |
| Reciprocal-link directory | 0 (pruned) | None | 0.0 (excluded) | High |
| Scraped / aggregated database | 0 (pruned) | None | 0.0 (excluded) | High |
| Link-farm network entry | 0 (pruned) | None | 0.0 (excluded) | Very High |
| Delisted-platform residual | 0 (pruned) | None | 0.0 (excluded) | Very High |
The pattern in the table is unambiguous: editorial review is the strongest predictor of per-listing citation contribution, and the gradient between the top archetype and the bottom retained archetype is about 16x. The marginal listing on a generic directory is not worthless, but its contribution is small enough that the time to maintain it is rarely justified against the time to maintain a stronger anchor.
Patterns that trigger spam classifiers
The patterns most reliably tied to spam classification during the audit were, in descending order of frequency: duplicate descriptive content across multiple unrelated businesses on the same platform; reciprocal linking within tightly clustered domain groups; thin or auto-generated metadata (categories assigned by keyword matching rather than by the business itself); WHOIS clustering across supposedly independent platforms; templated visual structures with only surface variation between sites; and the absence of any independent traffic signal, no organic search visits, no referral traffic, no engagement metrics that suggest actual users beyond the listing-submitter.
Of these, the duplicate-content and reciprocal-linking patterns did the most damage in practice, because they are the easiest for retrieval systems to detect and the hardest to fix without removal. Thin metadata can be enriched; templated structures can be redesigned by the platform operator; but duplicate content across hundreds of businesses is structural to how the platform works, and a single submitter cannot fix it.
The Springer research on web spam removal makes the same point: classification systems built to identify spam sites hit their highest accuracy on features that reflect the underlying production process, features like content overlap and link structure, rather than on surface features that can be cosmetically adjusted. The 2026 generation of AI engines applies the same principle, with the added capacity to compare content fingerprints across far more sources and at far greater speed than was possible when the Springer paper was written.
Anchor text diversity thresholds
Anchor text diversity turned out during the rebuild to be a more nuanced consideration than the standard SEO literature suggests. The conventional wisdom, that anchor text should be diverse to avoid over-optimisation flags, holds for inbound link profiles but works differently in directory contexts where the “anchor” is often a brand-name link from a listing to the business website.
The relevant diversity in directory contexts is not within-anchor (using different phrasings of the brand name) but across-listing (varying the supporting context in which the anchor appears). When 200 listings all use the brand name as anchor text in identical surrounding sentences, the homogeneity reads as automation; when 200 listings use the brand name as anchor text in 200 distinct surrounding sentences, the variation reads as independent editorial decisions. The unit of diversity is the listing-level context, not the anchor string itself.
The empirical threshold seen during the rebuild, and this is a heuristic rather than a confirmed parameter, was that listings sharing more than about 40% of their non-brand text with other listings in the same portfolio were measurably less likely to be cited. Listings sharing less than about 15% with any other listing performed best. The 40% threshold fits what you would expect if AI engines apply approximate-text-matching at a granularity similar to academic plagiarism detection, where matches above a third of the text are treated as effectively identical.
The cost of reciprocal link loops
Reciprocal link loops are the most expensive pattern an owner can carry in a directory portfolio, and the cost is not financial but graph-positional. When two or more directories link to each other and to a small set of client sites, they form a closed graph component that is structurally distinguishable from the open, hierarchical link patterns typical of editorial sources. Retrieval systems treat closed components with suspicion, not because the component itself is spam but because legitimate editorial sources rarely produce such patterns.
The client’s six identified link-farm networks were each built around reciprocal loops of between four and eleven domains. Removing the client’s listings from those networks did not collapse the loops; the operators kept running their networks. But it did remove the client from the closed component, which was the only outcome that mattered for citation rate. The disavow filings, where applicable, formalised that disconnection.
One under-appreciated consequence of being inside a reciprocal loop is that the brand shares a fate with every other business in the loop. If any of those businesses picks up a regulatory complaint, a reputation incident, or a manual penalty, the closed-component association can pass some portion of that signal to neighbouring nodes. The World Bank’s published guidance on scams that misuse institutional names shows the principle in a different domain: association by structural proximity carries reputational consequences even when there is no operational relationship between the parties. For directory portfolios, the implication is that sitting next to bad actors in a link graph is a risk in itself, separate from the quality of any single listing.
Transferable principles for other brands
The rebuild produced a set of principles that generalise beyond the specific HVAC vertical and beyond the specific 412-to-132 listing reduction. The first is that directory portfolio quality is a stock variable, not a flow variable. What matters is the steady-state quality of the listings that stay active, not the total number ever submitted. Owners who think in flow terms (“we added 50 listings this quarter”) are optimising the wrong metric. The right metric is the share of currently active listings that meet the editorial-quality threshold.
The second is that pruning produces measurable gains on its own, independent of any positive additions. The day-30 measurement confirmed this: removing negative signals lifted the citation rate before any rewriting was done. Practitioners who cannot afford the rewriting phase can still capture meaningful improvement from pruning alone. The cost of pruning is overwhelmingly time rather than money.
The third is that metadata harmonisation is the highest-leverage single intervention. Inconsistent NAP data fragments the brand’s identity across phantom variants, and the fragmentation suppresses citation rate at every retrieval system that reconciles entities. Setting and enforcing a single canonical reference document is the change that delivered the largest accuracy improvement on the client rebuild and would be the first change recommended to any brand with a multi-year directory history.
The fourth is that editorial review is worth paying for. The cost difference between editorial and auto-generated platforms is real, but the citation difference is larger. Owners on tight budgets should redirect spend from quantity-oriented submission services to a smaller number of editorial placements. The arithmetic favours that redirection in nearly every case observed.
The fifth is that AI-engine measurement belongs in routine reporting. The 60-prompt protocol used on the client rebuild is not hard to build and can be rerun quarterly at modest labour cost. Without it, owners are flying blind on what is increasingly the dominant discovery surface for service businesses; with it, the feedback loop between portfolio changes and citation outcomes becomes legible.
The sixth, and arguably the most important, is that the AI-citation surface rewards specificity. Generic adjectives (“trusted,” “reliable,” “experienced”) add little to retrieval ranking and even less to citation context framing. Specific, verifiable facts (founding year, certifications, named service areas, equipment types) add substantially to both. Rewriting descriptive copy to lead with specifics rather than adjectives is the single most repeatable content change available, and it needs no platform-specific knowledge.
The seventh is that AI engines differ from one another in retrieval behaviour, and the differences matter for prioritisation. Perplexity’s preference for retrievable, citable sources rewards editorial directories more steeply than GPT or Claude, both of which keep greater latitude for parametric recall. Brands optimising for Perplexity should weight editorial placements more heavily; brands optimising for GPT or Claude can afford a slightly broader portfolio, though the broader portfolio still has to meet the quality threshold.
Adjusting the approach under different constraints
Working with a $2K monthly budget
The client rebuild ran on a roughly GBP 6,800 total budget across 90 days, about $2,800 per month at prevailing exchange rates. Many owners cannot reach that level of spend. So the question is what the rebuild looks like at half the budget, say $2,000 per month, or roughly GBP 4,800 across 90 days.
The first adjustment is to defer the editorial-listing fees and put the budget into labour. Editorial placements usually cost between GBP 40 and GBP 200 per listing, and for a brand starting from a low base, three or four well-chosen editorial placements can deliver most of the citation lift from the editorial tier. Selecting those three or four (typically the relevant trade body, the regional chamber of commerce, and one or two sector-specific platforms) costs perhaps GBP 400 to GBP 600 in fees, a small fraction of the constrained budget.
The second adjustment is to narrow the rewriting scope. Instead of rewriting all 132 anchor listings, a constrained rebuild rewrites the top 40 by traffic and authority and leaves the remaining 92 unchanged in the first 90 days, returning to them in a later quarter. The reduction cuts rewriting labour from about 36 hours to about 14, freeing budget for other work.
The third adjustment is to push more of the pruning and metadata work onto the owner’s own time rather than paying external labour for it. Pruning is mechanical, tedious work, but it needs no specialised skill. An owner who can spend four to six hours a week for three weeks can do the bulk of the pruning. The opportunity cost is real, but it is often lower than the cash cost of paid labour at the constrained budget.
The fourth adjustment is to stretch the timeline. The original rebuild ran over 90 days because the client could resource the parallel workstreams; a constrained rebuild over 150 days reaches similar end-state results at lower per-month cost. The recovery curve is slower but the destination is the same.
Under these adjustments, the constrained version produces about 75% of the citation-rate improvement at about 60% of the cost. The bottleneck moves from money to time, and owners who can supply time can substitute it for money in most of the rebuild’s components.
Adapting for regulated industries
Regulated industries (financial services, healthcare, legal) face extra constraints that change the rebuild approach without overturning its core logic. The first is that descriptive copy is often subject to compliance review, and rewriting 132 listings to be substantively unique while staying within compliance guidelines takes more time and more iterations than the unregulated case. The second is that certain editorial directories have their own admission criteria: professional registers, for instance, may require evidence of certifications or licences that take time to assemble. The third is that incorrect metadata in regulated industries can carry legal consequences beyond citation rate, which raises the importance of metadata harmonisation and also raises the cost of getting it wrong.
The adapted approach adds a compliance-review layer to the rewriting phase, extending the rewriting timeline by about 50%. It also weights editorial placements on professional registers and regulator-published lists more heavily than in the unregulated case, because those placements carry not just citation value but independent compliance and reputation value. A solicitor’s listing on the relevant regulator’s register, for example, adds to AI-engine citation rate and at the same time serves as evidence of standing in the profession; that dual purpose justifies a higher relative spend on it.
One nuance specific to regulated industries is that AI engines applied to regulated topics often weight authoritative sources even more steeply than in unregulated ones. A medical query, for instance, will preferentially cite sources tied to regulatory or professional bodies before citing any business directory, however well written the directory listing is. So regulated brands should temper their expectations on directory-driven citation rate and concentrate effort on appearing in the most authoritative directory categories, even at the expense of broader reach.
Compressing the timeline to 30 days
Some scenarios need a 30-day rather than 90-day rebuild, typically when an acute citation-rate decline is causing immediate revenue impact. The compressed timeline is achievable but produces a different cost structure and a different risk profile.
Compression means running pruning, rewriting, and editorial submission in parallel rather than in sequence. The parallelism raises the labour requirement by about 40% because of coordination overhead and because some inefficiencies the sequential timeline absorbs without notice become visible under compression. It also demands more decision-making capacity from the owner, who must be available to approve canonical metadata, sign off on rewritten copy, and authorise editorial fees within tight windows.
The risk profile under compression is dominated by the reduced chance for measurement-driven correction. The 90-day rebuild used the day-30 measurement to confirm that pruning was producing the expected lift before committing to the rewriting phase; the 30-day rebuild has no such intermediate checkpoint and proceeds on the assumption that the audit’s diagnosis is correct. If the audit has misidentified the dominant problem (if, say, the citation decline is being driven by a content issue on the client’s own website rather than by directory portfolio quality) the compressed rebuild will spend its budget without producing the expected outcome.
For compressed engagements, the recommendation is to invest more in the audit phase rather than less, and to hold off on remediation until the audit’s conclusions are highly defensible. A compressed rebuild that starts on day three and ends on day thirty is preferable to one that starts on day one and discovers, on day twenty, that the audit was incomplete.
As Table 4 shows, the difference between the standard 90-day rebuild and the compressed 30-day variant is not simply temporal; it shows up in cost structure, risk exposure, and the marginal returns of each phase.
Table 4: Comparison of Rebuild Variants Across Different Constraint Profiles
| Variant | Total Cost (GBP) | Timeline | Citation-Rate Recovery | Primary Risk |
|---|---|---|---|---|
| Standard 90-day rebuild | 6,800 | 90 days | ~95% of theoretical max | Owner attention sustained over quarter |
| Constrained budget ($2K/month) | 4,800 | 150 days | ~75% of theoretical max | Owner time substitution |
| Regulated-industry adapted | 9,200 | 120 days | ~85% of theoretical max | Compliance review delays |
| Compressed 30-day rebuild | 9,500 | 30 days | ~80% of theoretical max | No intermediate measurement checkpoint |
| Pruning-only minimal | 1,200 | 45 days | ~35% of theoretical max | Anchors not rewritten, ceiling effect |
| Owner-executed (no agency) | 800 | 180 days | ~60% of theoretical max | Inconsistent execution quality |
| Editorial-only (top 5 placements) | 2,400 | 60 days | ~55% of theoretical max | Spam signals from existing portfolio remain |
The variant choice comes down to which constraint binds hardest. Owners with capital but no time will favour the standard or compressed variants; owners with time but no capital will favour the constrained or owner-executed variants; owners in regulated industries will accept higher cost for compliance assurance. No variant is dominated by another on every dimension, and the right choice depends on the specifics of each engagement.
What I would do differently next time
Looking back at the rebuild with the benefit of the post-90-day data, several decisions would go differently in a future engagement of similar shape. The first is that the audit phase would get more time. The original audit took roughly five working days; in hindsight, eight to ten would have produced a sharper diagnosis. The extra time would have gone into deeper inspection of the link-farm networks, specifically into tracing the agency-to-network connections that only became apparent during the rebuild. Earlier sight of those connections would have shaped conversations with the owner about what to expect from the relevant agencies in any future engagement and about whether residual contractual relationships needed to be ended.
The second is that the rewriting protocol would include a structured fact-collection step at the outset rather than gathering facts piecemeal as the rewriting went on. The first 30 listings were rewritten with whatever facts were readily available; the next 100 benefited from a more systematic fact base assembled mid-project; the final 30 reflected the matured fact base. The variation in citation contribution between early and late rewrites suggested that the early rewrites underperformed because the fact base was thinner. Front-loading fact collection would lift the average quality of the rewritten copy and cut the need for mid-project rework.
The third is that the measurement protocol would include a control group of unaffected prompts. The 60-prompt set was chosen for relevance to the client’s services, which meant all 60 prompts were potentially affected by the rebuild. Adding another 20 prompts unrelated to the client’s services would have given a control against which to test whether the engines themselves were drifting in their citation behaviour during the measurement period. AI engines update frequently, and some portion of any observed change in citation rate may reflect engine drift rather than rebuild effects. Without a control group, that confound cannot be cleanly separated.
The fourth is that the disavow strategy on the link-farm networks would start earlier, preferably during the audit phase rather than after the pruning decisions were made. Disavow signals take time to propagate, and starting that clock as early as possible brings forward the day when the structural disconnection from the closed components shows up in retrieval-system trust graphs. The original rebuild started disavow on roughly day 25; an earlier start, on day five or six, would have shifted some of the day-30 lift earlier in the timeline.
The fifth is that the canonical reference document for metadata would go up on the client’s own website at a stable URL (typically an “about” or “contact” page with structured data markup) before the directory updates began. AI engines that reconcile entities increasingly prefer to anchor on the brand’s own primary source, and making that source unambiguous before propagating updates downstream gives a stronger reconciliation target. On the client rebuild, the primary-source update ran in parallel with the directory updates, which produced the right end state but missed the chance to use the primary source as the authoritative reference during the propagation period.
The sixth is that the engagement would include a longer post-rebuild monitoring phase. The 90-day window captured the recovery curve but not the steady-state behaviour that follows it. Three more months of monitoring, with the same 60-prompt protocol run monthly, would establish whether the gains are stable or need ongoing maintenance. Anecdotally, the client’s day-110 measurement (taken outside the formal engagement) suggested stability, but stability over 20 days is not stability over 200, and an extended monitoring phase would turn the anecdote into evidence.
The seventh is that the engagement scope would, from the start, include a portfolio governance plan covering the 12 to 24 months after rebuild completion. Without governance, the conditions that produced the original 412-listing portfolio will reassert themselves: well-meaning marketing initiatives, agency proposals, employee submissions, and other accumulating sources will gradually re-bloat the portfolio, and the cycle repeats. A governance plan that sets criteria for any new listing (minimum editorial standard, content uniqueness requirement, metadata consistency check) and assigns responsibility for enforcing them is the structural intervention that prevents repetition.
The insight that pulls all of this together is that directory spam filtering by AI engines is, in 2026, less a technological frontier than a quality-assurance regime. The filters are not exotic; they are competent. They reward the same disciplines competent editors have always rewarded (accuracy, specificity, consistency, evidence of independent endorsement) and they penalise the same shortcuts editors have always penalised. What has changed is the scale at which the rewards and penalties are handed out and the speed at which they update. A directory portfolio that drifted for years without consequence under older retrieval regimes can now be re-evaluated overnight when an engine updates its trust graph. The owner’s task is not to outsmart the filters but to build a portfolio no reasonable filter would object to, one whose quality is so unambiguous that no scoring change can demote it. That orientation, more than any single technique, separates the practitioners who will rebuild once and maintain steadily from those who will rebuild every two years for the rest of their commercial lives.

