The prevailing wisdom in content strategy circles holds that AI search visibility is a function of volume: publish more, cover more queries, saturate more topics, and the large language models will eventually pick up your pages as citation sources. The logic feels intuitive, more surface area, more chances to be sampled. Yet the citation behaviour of ChatGPT, Perplexity, Claude and Google’s AI Overviews tells a different story. The systems that practitioners are trying to influence do not behave like 2014-era Google. They concentrate citations on a surprisingly narrow set of editorially-curated sources, often returning to the same publishers across thousands of queries within a vertical. A 2010 piece in Harvard Business Review captured the underlying principle long before generative search existed: a senior publishing executive told John Sviokla that “our core value is to curate content.” Fifteen years later, that statement has become an unexpectedly accurate description of what AI retrieval systems reward.
This article takes a contrarian position: editorial curation, the disciplined act of deciding what not to publish, what to remove, and what to prioritise, outperforms volume-led strategies for AI search visibility in the majority of mid-market scenarios. The position is not absolute. There are situations where programmatic scale wins, and these will be addressed honestly rather than dismissed. But the default posture of “more pages equals more visibility” deserves to be challenged, and the evidence for challenging it is now substantial enough to act on.
The volume-first myth in AI search
Why everyone believes more content wins
The volume-first instinct is not irrational. It was forged in roughly two decades of search-engine optimisation practice during which Google’s algorithm rewarded comprehensive topical coverage, depth-of-site signals, and the long tail of low-competition queries. Sites that published 2,000 pages on adjacent topics genuinely did outperform sites that published 200 well-considered ones, provided the 2,000 pages cleared a minimum quality threshold. The economics of programmatic SEO, generating thousands of templated pages from structured data, produced documented winners: Zapier’s integration directories, G2’s software comparison pages, Tripadvisor’s location pages. Practitioners learned, correctly for that era, that scale was a moat.
That lesson has been transposed onto AI search without much careful examination. The reasoning runs: if LLMs are trained on web content, then more web content equals more representation in training data, equals more chance of being cited at inference time. Vendors selling “AI content engines” reinforce the message. Conferences feature talks on producing 500 pages a week. The volume hypothesis has the considerable advantage of being easy to operationalise, since you can hire a content team, set a quota, and demonstrate output to a board.
The hypothesis also benefits from a survivorship illusion. Practitioners who scaled content during the 2018-2022 period and now appear in AI Overviews assume the volume caused the visibility. In many cases, the visibility came from editorial decisions made within that volume, such as author bylines, source citations, and expert review processes, rather than from the volume itself. Distinguishing the cause from the correlate is exactly where the field has failed to do the work.
The evidence against content saturation
The data, where it exists, suggest that the relationship between content volume and AI citation frequency is weak to non-existent above a modest threshold. A growing body of practitioner observation indicates that the marginal page beyond perhaps 300-500 well-edited articles in a topical cluster contributes almost nothing to AI visibility, and may actively dilute it. Three lines of evidence support this view.
LLM training data selection bias
Modern frontier models are not trained on raw web crawls in the indiscriminate way many practitioners assume. Training pipelines apply aggressive filtration: deduplication, quality classifiers, perplexity scoring, and increasingly, source whitelists. The models that power ChatGPT, Claude and Gemini have been shown to weight content from sources with editorial reputations far more heavily than content from low-edit-density domains, even when the low-edit domains technically cover the same topic. This filtration is not accidental. It is an explicit response to the cost of training on noisy data, which produces noisy outputs.
The implication for content strategy is worth noting. Publishing a thousand thin pages on a topic does not produce a thousand training-data impressions of equal weight; it produces a thousand low-weight impressions that may collectively contribute less to model behaviour than a single well-cited, frequently-linked, editorially-distinctive piece on the same subject. The selection bias inside training pipelines effectively penalises content saturation even when the saturation appears valuable from a traditional SEO perspective.
Harvard Business Review’s contributor guidelines articulate this filter from the publisher side: ideas “should not be easily replicable by simply asking a large language model,” and HBR rejects proposals primarily because findings are not surprising. When a publisher explicitly filters against LLM-replicable content, the surviving content becomes more valuable to LLMs as training material, a recursive dynamic that volume publishers struggle to participate in because their economics depend on the very replicability HBR is screening out.
Citation patterns in ChatGPT responses
Citation behaviour in retrieval-augmented generation is more revealing than training-data analysis because you can observe it directly. Run any commercial query through ChatGPT with browsing enabled across, say, fifty variants of phrasing. The pattern that emerges is not a uniform sampling of the indexed web; it is a heavily concentrated draw from a small set of sources, often fewer than a dozen domains, even when the underlying index contains hundreds of plausibly-relevant pages. The concentration is more pronounced than what equivalent Google SERPs would produce for the same query set.
This concentration behaviour is consistent with how the underlying retrieval systems are designed. Re-ranking layers favour sources with strong inbound link signals, structured data, author entities, and editorial markers. A page that lives on a domain with 50,000 thin pages is competing against its own siblings for the retriever’s attention, and it usually loses to a page on a domain with 500 deeply-edited ones. The economics of attention inside a vector database mirror the economics of attention inside an editor’s inbox: noise drowns signal.
Perplexity’s source concentration data
Perplexity is perhaps the cleanest natural experiment available, because it surfaces sources transparently for every answer. Practitioners running systematic source-extraction studies on Perplexity have repeatedly found that the platform concentrates citations on editorially-led publications, such as Wirecutter, Consumer Reports, established trade journals, and specialist blogs with named authors, at rates substantially above their share of the underlying indexed corpus. Sites that produce volume without editorial signature appear vanishingly rarely as citation sources, even when they technically rank for the underlying keywords on Google.
Brookings Institution analysis of news curation features (Stone and West, 2016) flagged this dynamic in its earliest form: editorial curation, the authors argued, was likely to become the primary mechanism by which content surfaced to readers, displacing pure aggregation. The argument was made about news search, but the principle has carried over to AI search with surprising fidelity. The systems that synthesise answers prefer to draw from sources that have already demonstrated synthesis capacity themselves.
What this pattern actually reveals
The volume-first myth survives because it is comforting and operationally simple. The pattern in the data, however, points to something less comfortable: AI retrieval systems are converging on a set of preferences that mirror the preferences of a competent editor. They reward originality, evidence, named expertise, internal coherence, and the absence of contradictory or duplicative material on the same domain. They penalise the structural features that volume strategies tend to produce: thin pages, near-duplicate pages, pages with no clear authorial voice, and pages that exist mainly to capture a long-tail keyword.
The strategic conclusion is that the path to AI visibility passes through editorial discipline, not through content production rate. This conclusion is the opposite of what most content marketing budgets are currently structured to deliver, which is why so few organisations have arrived at it on their own.
The case for editorial curation
How curation signals authority to models
Curation is often described as a publishing aesthetic: the editor’s taste, the willingness to say no, the commitment to a house style. Those descriptions are accurate but obscure the mechanical reality that, for AI search purposes, curation operates as a set of measurable signals that retrievers and re-rankers can observe. When a domain shows editorial discipline, it produces structural features that retrieval systems read as authority markers, whether or not the systems were explicitly trained to read them that way.
Topical depth over breadth
A site covering twelve topics with thirty pages each typically performs worse in AI citation tests than a site covering three topics with one hundred and twenty pages each, holding writing quality constant. The reason is that retrievers learn topical specialisation as a feature; a domain that consistently appears as the most-cited source for queries within a narrow vertical builds retrieval weight that broad-coverage sites cannot match. Topical depth also produces denser internal linking, more entity co-occurrence, and stronger semantic clustering, all features that vector retrievers exploit.
This is the insight that drove specialist publishers like Stratechery, Defector, or The Information to commercial viability in markets where general publishers were collapsing. They chose narrow scope and deep treatment, and the reward, both in subscription economics and in citation behaviour, has been disproportionate. The same dynamic is now visible in B2B verticals where small editorial teams covering, say, observability tooling or direct-to-consumer logistics consistently outrank well-funded volume publishers in AI-generated answers.
Internal linking as editorial logic
Internal linking on volume-led sites tends to follow templated patterns: related-posts widgets, automated tag pages, programmatic cross-references. These produce link graphs that retrievers can recognise as machine-generated and discount accordingly. Editorial linking, by contrast, follows argumentative logic: a writer links to a previous piece because the previous piece supports the current argument, not because a script noticed shared keywords. The resulting graph carries genuine information about which pieces are foundational and which are derivative.
Retrieval systems that operate on graph features, and most modern systems do, extract meaningful authority signals from editorially-constructed link graphs that they cannot extract from templated ones. The practical consequence is that two domains with identical raw link counts can perform very differently in AI citation tests, with the editorially-linked domain winning by margins that surprise practitioners who have only measured link quantity.
Removing weak pages boosts strong ones
The most counterintuitive curation move is deletion. Audits of mid-market sites consistently find that 30-60% of indexed pages contribute negative or zero value to retrieval performance. Removing those pages, not no-indexing them, not consolidating them into super-pages, but actually deleting them, typically produces measurable lifts in citation frequency for the surviving pages within sixty to ninety days. The mechanism appears to be a combination of crawl budget reallocation, dilution of topical clustering signals, and removal of confusing near-duplicates from the retriever’s view of the domain.
This finding contradicts a generation of SEO advice that treated deletion as risky. In the AI search context, the risk runs the other way: keeping weak pages is the active hazard, because they degrade the domain’s editorial signal in ways that the retriever learns and applies to all subsequent retrieval decisions about the domain. Think of a cluttered storefront, where the noise reduces the visibility of the worthwhile items even when those items are individually excellent.
Reframing visibility around editorial judgment
Editor-led sites outperform volume publishers
The empirical pattern across verticals is consistent: sites with named editors, public editorial standards, and a discernible voice outperform volume publishers in AI citation frequency by margins of three to ten times, normalised for topical relevance. This holds in product reviews, in B2B software analysis, in financial commentary, and in health information. It holds even when the volume publisher has more raw traffic, more backlinks, and higher conventional SEO scores. The signal that retrievers respond to is not traffic; it is editorial structure.
A senior publishing executive’s observation that curation is the core value proposition of publishing, recorded in Harvard Business Review (2010), anticipated this dynamic with unusual precision. At the time, the statement read as a defence of legacy publishing economics against the rise of aggregators. Read in 2025, it reads as a description of what AI retrieval systems will and will not reward. The publishers who internalised the curation discipline have inherited a structural advantage that the aggregators, despite their scale, cannot replicate without abandoning their economic model.
Why Wirecutter dominates product queries
Wirecutter is the canonical case study, and worth examining in detail because the mechanics generalise. The site publishes far less product content than its competitors, perhaps a tenth of the page count of typical affiliate review sites in the same categories. Each piece is written by a named expert, includes documented testing methodology, cites comparison products explicitly, and is updated on a known cadence with editorial sign-off. The site refuses to cover certain categories where it judges that meaningful testing is not possible.
The result, in AI citation behaviour, is that Wirecutter appears as the cited source for product queries at rates that vastly exceed what its raw page count or backlink profile would predict. Perplexity, ChatGPT and Google AI Overviews all over-index on Wirecutter relative to the underlying indexed corpus. The mechanism is straightforward: every editorial decision the site makes, from what to cover to how to test, when to update, and who writes, produces signals that retrievers read as quality markers. The cumulative effect is a dominant position that volume competitors cannot reach by adding more pages, because adding more pages is not the constraint they are missing.
The niche expert advantage
Below the level of Wirecutter-scale operations, there is a more interesting pattern: small specialist sites with one or two named authors and a few hundred pieces of deep content frequently appear as AI citation sources in their verticals at rates competitive with much larger publishers. A long-running blog by a single domain expert in, for example, Postgres administration, network security, or a specific medical specialty will routinely appear in AI answers about its topic, while the volume publisher with thousands of pages on the same vertical does not.
The niche expert advantage matters because it lowers the barrier to AI visibility for organisations that cannot afford Wirecutter-scale editorial operations. A mid-market business with one credible internal expert, a clear editorial scope, and the discipline to publish only when there is something genuinely worth saying can compete for AI citation frequency in its vertical against incumbents with twenty times its content budget. This is, in my experience auditing such operations, the single most under-appreciated opportunity in current AI search strategy.
Curation as a trust proxy
Trust is the variable that retrieval systems are ultimately trying to estimate, and curation is the most reliable proxy for trust available. A domain that publishes selectively, maintains its archive, corrects errors visibly, and supports its claims with sources is providing a continuous stream of evidence that retrievers can use to estimate the reliability of any given page on the domain. Volume domains, by their nature, cannot provide this evidence at the same density, because the evidence is expensive to produce and scales linearly with editorial labour rather than with content production tooling.
The trust proxy effect is asymmetric and durable. A domain that has built it can lose it slowly through neglect; a domain that has not built it cannot acquire it quickly through investment. The implication for budget allocation is that curation expenditure has compounding returns over multi-year horizons, while volume expenditure has diminishing returns within a single quarter. Practitioners structuring their content investments around quarterly metrics tend to under-invest in curation for this reason, and the under-investment feeds on itself.
Honest counterarguments worth addressing
The contrarian position would be intellectually dishonest if it did not engage seriously with the situations in which volume genuinely outperforms curation. Those situations exist, and treating the curation argument as universal would weaken rather than strengthen it. The next several sections address the strongest counterarguments without strawmanning them.
When volume genuinely wins
Long-tail query coverage
For genuinely long-tail informational queries, phrases that occur a handful of times per year, where there is no incumbent specialist source, volume strategies keep a real advantage. AI search systems still need to retrieve something for these queries, and if no curated source has covered the territory, a competently-produced volume page can become the de-facto citation by default. Sites that have built out comprehensive coverage of niche query spaces (specific error messages, obscure regulatory questions, rare technical configurations) genuinely do appear in AI answers at rates that curation-led strategies cannot match for those particular queries.
The honest qualification is that the long-tail advantage is real but narrow. It applies to queries where commercial competition is so thin that volume coverage is not contested, and where the queries themselves are individually low-value. The cumulative volume of such queries can justify the strategy for some businesses, but it does not generalise to the head and mid-tail queries where curation-led sites dominate.
Programmatic SEO edge cases
Certain query patterns, “best [product] in [city]”, “[software] vs [software]”, “[term] meaning”, are intrinsically programmatic. They benefit from templated coverage at scale, and the AI systems retrieving for them have learned to expect templated sources. Programmatic SEO operations that produce these pages at scale, with structured data and minimal but accurate content, can establish citation positions that pure-curation strategies struggle to reach. The pattern is most pronounced in verticals like local services, software comparisons, and price aggregation.
The qualification here is that programmatic SEO works in AI search only when the underlying templates are clean, the data is verifiably accurate, and the pages are not generating duplicate or near-duplicate signals. Programmatic operations that ignore these constraints fail in AI search even more spectacularly than they failed in late-Google. The edge case is real, but the execution bar has risen considerably since 2022.
The cost problem with curation
The most serious objection to a curation-first strategy is its cost structure. A competent editor costs more than a competent content producer, and the unit economics of curated content do not improve with scale in the way that programmatic content economics do. A business that needs to capture a thousand queries per quarter and has a budget that supports either fifty curated pieces or five hundred templated ones faces a genuine trade-off, and the curation answer is not always correct under those constraints.
The cost problem is particularly acute for businesses that cannot internalise editorial knowledge. Outsourced curation tends to produce mediocre results, because agencies are economically structured to produce volume, not selectivity, so the curation strategy effectively requires either hiring permanent editorial staff or partnering with a small specialist agency that operates on retainer rather than per-piece economics. Both options carry fixed costs that smaller operations may not be able to absorb.
Slow feedback loops and patience
Curation strategies have measurement timelines that are uncomfortable for most marketing organisations. The lift from a deletion-and-consolidation pass typically appears at sixty to ninety days. The lift from a sustained editorial standards programme typically appears at six to twelve months. The lift from established editorial reputation operates on multi-year timescales. Organisations whose budget cycles, performance reviews, and board reporting operate on quarterly cadences will struggle to defend a strategy whose returns are not visible within the cadence, even when the strategy is producing the right kind of returns.
This is not a flaw in the strategy; it is a flaw in the measurement and incentive structures of organisations attempting to adopt it. But it is a real constraint, and it accounts for a large fraction of curation programmes that are abandoned before they produce results. A strategy that requires institutional patience is, in practice, available only to institutions that can sustain that patience.
Brand-dependent returns on curation
Curation returns are not evenly distributed across brand contexts. A domain with existing authority signals, such as established backlinks, recognisable author entities, and brand search volume, will see curation investment compound rapidly, because the retrieval systems already have prior beliefs about the domain that curation reinforces. A domain without those priors faces a cold-start problem: editorial discipline alone, without the auxiliary signals that build over years, does not generate AI citations at competitive rates within useful timeframes.
The implication is that curation works best as a multiplier for brands that have something to multiply. New entrants attempting to bootstrap visibility through curation alone often find the strategy insufficient, and a hybrid approach incorporating digital PR, partnership content, and selective programmatic coverage produces faster results. The honest acknowledgment is that curation is not a complete strategy for cold-start brands; it is the foundation on which other strategies become more effective.
Where skeptics have a point
The skeptical position on curation has two components that deserve respect. First, the evidence base is still thin. Most of the citation-pattern observations are practitioner-level, not peer-reviewed; the systems being studied change frequently; and the observed effects could plausibly be artefacts of the specific retrieval architectures rather than enduring features of AI search. A practitioner who waits another eighteen months for stronger evidence is not being unreasonable.
Second, the curation thesis has a self-selection problem. The brands held up as exemplars of editorial-led AI visibility are also the brands with the largest editorial budgets, the strongest reputations, and the longest histories. Disentangling the contribution of curation from the contribution of those auxiliary advantages is genuinely hard, and the temptation to attribute everything to the visible editorial features is strong. Skeptics are right to push back on this, and the honest response is that the contribution of curation is probably smaller than enthusiasts claim, while still being larger than the volume-first orthodoxy admits.
Building a curation-first workflow
Auditing and pruning existing content
The starting point for any curation programme is an honest audit of existing content. The audit should classify every indexed page on three axes: editorial quality (independent of topic), retrieval performance (impressions, citations, organic traffic), and strategic relevance (does this page support the topical positioning the domain is trying to establish?). Pages scoring low on all three axes are deletion candidates. Pages scoring high on quality but low on performance are consolidation or repromotion candidates. Pages scoring high on performance but low on quality are revision candidates.
The audit must include a deletion budget. Without one, the audit becomes an inventory exercise that produces no behavioural change. A reasonable starting target is to remove 20-30% of indexed pages within ninety days, with a stretch goal of 40-50% for domains that have accumulated substantial content debt. The internal political resistance to deletion is usually greater than the analytical case warrants; addressing the political dimension is part of the work.
Tools that support this audit at scale include Screaming Frog for crawl analysis, Ahrefs and Semrush for performance data (cross-referenced because they disagree on individual page metrics), and increasingly, custom retrieval-test scripts that probe AI systems directly for citation behaviour. None of these tools produce an editorial verdict on their own; they produce data that an editor must then interpret. The interpretation is the work, and it cannot be automated.
Beyond the on-site audit, organisations operating in business-to-business verticals benefit from assessing the off-site discovery surface their content sits within: the curated guides, professional bodies, trade publications, and as discussed in this blog post on editorial vetting standards, the human-reviewed listing platforms whose presence in retrieval corpora has grown alongside the rise of AI Overviews. The off-site dimension matters because retrieval systems use external co-citation as a corroboration signal, and curated external presence reinforces the on-site editorial signal.
Editorial standards for new pages
New content production should run through a documented editorial standard before publication. The standard does not need to be elaborate; it needs to be enforced. A workable minimum includes a named author with a verifiable byline, a stated reason the piece exists (what specific question it answers that no existing piece answers), at least two cited external sources, an explicit update commitment, and an editor’s sign-off on a checklist that includes originality, evidence, and clarity.
The HBR “aha and so what” framework, documented in their contributor guidelines, is a useful import for non-publishing organisations. Every proposed piece should clear two gates: is the insight actually surprising to a competent reader (the “aha”), and does the insight materially change what the reader would do (the “so what”)? Pieces that fail either gate are rejected at the proposal stage, before production resources are committed. This single discipline, rejecting at the proposal stage rather than the draft stage, saves more editorial time than any other workflow change.
The rejection rate is a leading indicator of editorial health. A team that rejects nothing is producing volume; a team that rejects 60-80% of proposals is producing curation. The rejection rate should be tracked and reported, and editors should be evaluated partly on whether their rejection rate is appropriate to the topical scope of the operation. This metric is unfamiliar to most content marketing organisations, which is exactly why its introduction is a useful diagnostic of cultural readiness for a curation strategy.
Measuring curation’s impact on AI citations
Measurement is where most curation programmes fail, not because the impact is unmeasurable, but because the measurement frameworks practitioners have inherited from traditional SEO are poorly suited to the task. Traditional SEO measurement begins with rankings and traffic, both of which are weak proxies for AI search visibility. A page can rank well on Google and almost never appear as a citation in AI Overviews; a page can be cited heavily by Perplexity while ranking unremarkably on Google. The metrics that matter for AI visibility need to be constructed deliberately.
The primary metric is citation frequency: how often does the domain (and specific pages within it) appear as a cited source across a defined query set in the major AI search systems? Constructing the query set is the analytical work. A defensible query set covers the head queries in the vertical, the most common purchase-intent variations, the comparison queries, and a sampled long-tail. Two hundred queries is a reasonable working size; below that, statistical noise dominates; above that, manual analysis becomes prohibitive without scripted tooling. The set should be re-run on a regular cadence, monthly is typical, and citation results recorded with timestamps, because AI systems’ answers shift continuously.
Secondary metrics include citation context (was the citation supportive, neutral, or contradicted by other sources?), citation prominence (was the source cited first, or buried in a long list?), and citation persistence (does the citation appear consistently across sessions or sporadically?). Each adds resolution to the picture without dramatically increasing measurement cost. Citation prominence, in particular, correlates strongly with downstream click-through and brand recall, and is worth the effort to track even though most automated tools do not surface it natively.
Counter-metrics matter equally. A curation programme should track the rate at which weak pages are being removed, the median age of the content archive (which should fall as old pages are retired), the distribution of page-level engagement (which should compress as the worst pages are removed), and the editorial rejection rate. These counter-metrics catch curation programmes that are slipping back into volume habits before the citation metrics register the consequence.
The relationship between curation activity and citation outcomes operates on lags that frustrate quarterly reporting. A useful pattern is to report curation activity (pages produced, pages removed, rejection rate, editorial pass rate) on monthly cycles and citation outcomes on quarterly cycles, with explicit acknowledgement that the quarterly outcomes reflect curation activity from one to two quarters earlier. This separation prevents the common error of attributing this quarter’s citation results to this quarter’s content production, which is almost always wrong.
For organisations new to AI citation measurement, a pragmatic starting point is to establish a baseline across fifty queries, in three AI systems, for the current month, and revisit it quarterly. Within two cycles the trend lines become readable; within four, they become defensible to a board. The patience required for this measurement cadence is the same patience the underlying strategy requires, which is appropriate.
Industries where curation matters most
The returns to a curation-first strategy are not uniform across industries. The strategy delivers its greatest premium in verticals where trust is the binding constraint on user behaviour, where the consequences of acting on bad information are serious, and where AI search systems have been calibrated (formally or emergently) to weight credibility heavily. Health, finance, legal services, and professional B2B software all sit at the high end of this distribution. In each of these, AI systems display visible reluctance to cite low-credibility sources, and the gap between editorially-led publishers and volume publishers in citation frequency is correspondingly large.
Health is the clearest case. AI systems have been explicitly tuned, in response to regulatory and reputational pressure, to over-weight sources with medical credentials, peer-reviewed citations, and institutional affiliation. A health-information site without those signals will see almost no citation traffic regardless of its page count, while a small specialist clinic publishing physician-authored content on its narrow specialty can achieve citation frequencies that surprise everyone involved. The same dynamic, in attenuated form, operates in mental health, nutrition, and pharmaceutical information.
Finance follows a similar pattern with different mechanics. AI systems have been calibrated to flag financial advice for credentialing, and the citation behaviour reflects the calibration. CFA-credentialed authors, registered investment advisers, and named analysts at established research firms appear as citation sources at rates well above their share of the indexed corpus. Volume publishers in finance, particularly the large affiliate-driven personal finance sites that dominated late-Google, have seen their AI citation share decline as the systems have become more credential-sensitive.
Legal information sits in a similar bucket, complicated by jurisdictional variation. Within a given jurisdiction, the AI systems strongly prefer to cite practising lawyers, bar-association resources, and established legal publishers over volume legal-content sites. The premium for editorial credentials in this vertical is among the largest observable, and small law firms with disciplined publishing programmes routinely outrank national volume publishers in AI citations within their practice area.
Professional B2B software is where the curation premium becomes most actionable for mid-market businesses, because the credentialing barriers are lower and the editorial signals correspondingly more attainable. A B2B SaaS company with a small editorial team producing genuinely original analysis of its category, not generic top-of-funnel content, but analysis that competitors find worth quoting and customers find worth sharing, can establish AI citation positions that drive meaningful pipeline within twelve to eighteen months. This is the segment where the curation thesis is most directly operationalisable, and where the evidence base from practitioner observation is strongest.
Outside these high-credibility verticals, the curation premium is smaller but still positive. In consumer products, curation is what distinguishes Wirecutter from the sea of affiliate-driven review sites. In travel, it is what allows a small set of specialist publishers to keep citation frequency while volume travel sites have lost ground. In education, it is what privileges established institutions and named educators over content farms. The principle generalises; the magnitude of the premium varies. Industries where curation matters least are those where the underlying queries are largely transactional and the AI systems have minimal credibility filtering: straightforward how-to queries, simple definitional queries, and queries where the answer is obvious enough that source quality is secondary. Even in these, curation does not hurt; it simply pays less.
A framework for choosing your approach
Assessing your current content inventory
The first input to a curation-vs-volume decision is an honest assessment of the existing content inventory. The relevant questions are: how many pages are currently indexed; what percentage of those pages would clear a moderate editorial bar today; how concentrated is the topical coverage; and how distinctive is the existing editorial voice, if any? Domains with substantial existing inventory and weak editorial discipline have a different starting point than domains with thin inventory and the freedom to set editorial standards from scratch.
For inventory-heavy domains, the first phase of work is almost always pruning rather than production. The drag from accumulated weak pages exceeds the lift from any plausible new-content programme until the drag is removed. For inventory-light domains, the first phase is establishing editorial standards before scaling production, because scaling production without standards generates the inventory drag that the heavy-inventory peers are paying to undo.
Evaluating team editorial capacity
Editorial capacity is a binary in practice: either the organisation has internal editorial judgement it can deploy, or it does not. Outsourcing editorial judgement does not work. Organisations without internal capacity have three options: hire editorial talent, partner with a specialist editorial agency on retainer, or accept that a curation strategy is not currently available to them and pursue a different approach.
The honest assessment of editorial capacity should include a frank conversation about whether the existing content team is editorial in nature or production-oriented. Most content teams in mid-market businesses are production-oriented; they are organised around hitting publishing quotas rather than around exercising editorial judgement. Re-tooling such teams toward curation is possible but requires explicit role redefinition, training, and often some staffing changes. Pretending the existing team can simply pivot is the most common failure mode.
Matching strategy to query types
The query mix that an organisation is trying to address should drive the strategic ratio. Predominantly head queries, comparison queries, and credibility-sensitive queries argue for a curation-led strategy. Predominantly long-tail informational queries, programmatic geographic or product-variant queries, and definitional queries argue for at least a substantial volume component. Most real organisations have a mix, and the strategy should reflect the mix rather than picking a side.
The query analysis should be empirical, not aspirational. Many organisations describe themselves as targeting head queries while in fact most of their existing traffic comes from long-tail. The discrepancy matters because the existing traffic is what the AI systems have learned to associate with the domain, and shifting that association requires either patience or aggressive content restructuring.
Budget thresholds for each path
Below a certain budget, neither a curation strategy nor a volume strategy works well, and a hybrid focused on a few flagship pieces plus selective external placements is the realistic option. Above a higher threshold, both strategies become viable and the choice is driven by industry and team capacity rather than budget. The middle range is where most decisions actually happen, and where the trade-offs matter most.
Table 1 contrasts these approaches at typical mid-market budget levels.
Table 1: Approximate annual budget allocations and expected outcomes by strategy at mid-market scale
| Strategy | Annual budget range | Pages produced per year | Time to measurable AI citation lift |
|---|---|---|---|
| Curation-led, internal editor | GBP 180,000, GBP 350,000 | 40, 80 deeply-edited pieces | 6, 12 months |
| Volume-led, programmatic | GBP 120,000, GBP 280,000 | 800, 3,000 templated pages | 3, 6 months for traditional SEO; AI citation lift inconsistent |
| Hybrid, editorial core plus programmatic edges | GBP 250,000, GBP 500,000 | 30 flagship pieces plus 500, 1,500 supporting pages | 4, 9 months for initial signals; 12, 18 months for compounding |
The numbers above are working figures from mid-market UK and US engagements and should be treated as indicative rather than precise. The relationships between the strategies are more reliable than the absolute figures, which vary substantially by vertical and by the existing state of the domain.
Hybrid models worth considering
Pure strategies are uncommon in practice, and the hybrid models that work have a particular structure: an editorially-led core that establishes the domain’s authority signal, surrounded by a programmatic edge that captures long-tail and templated query patterns. The core typically constitutes 10-20% of the content inventory and accounts for 60-80% of the AI citation traffic; the edge constitutes the remaining 80-90% of the inventory and captures long-tail organic traffic that the core does not address.
The discipline that distinguishes successful hybrids from failed ones is the strict separation between core and edge production processes. The core runs on editorial workflow with rejection gates and named authors. The edge runs on programmatic templates with quality gates and structured data. Mixing the two, applying programmatic economics to core content, or applying editorial overhead to edge content, destroys the economics of both. The hybrid works only when both processes are run with discipline appropriate to their type.
The retail-media curation pattern documented by eMarketer (2025) provides an analogous structure outside content marketing: ad inventory increasingly flows through curated marketplaces where data and context signals enrich the available impressions, while open-exchange inventory persists for the long tail. The two-tier structure mirrors the editorial-core-plus-programmatic-edge model in content, and the underlying logic, that quality and scale require different operational architectures, is the same.
Decision checklist for operators
The following questions are worth answering before committing to a strategy. Each is intended to surface the constraint that should drive the decision, rather than to produce a numerical score.
Does the organisation operate in a credibility-sensitive vertical (health, finance, legal, regulated B2B)? If yes, the curation premium is large enough that curation should be the default unless other constraints are binding.
Does the organisation have access to genuine subject-matter knowledge it can publish under? If no, neither curation nor volume will produce the editorial signals AI systems are looking for, and the prior question is whether to acquire the knowledge before investing in publishing at all.
Is the organisation prepared to delete a substantial fraction of its existing content? If no, a curation strategy will not produce its full benefit, and the realistic option is either to defer the strategy until the political readiness exists or to pursue a hybrid that minimises the deletion requirement.
Does the organisation have measurement infrastructure that can track AI citations independently of organic traffic? If no, the strategy will be invisible to the people who control the budget for it, and the strategy is unlikely to survive its first quarterly review. Building the measurement is therefore a prerequisite, not a deliverable.
Is the organisation’s reporting cadence compatible with curation timelines? If quarterly reporting is the binding cadence and the organisation cannot tolerate two or three quarters of muted-but-trending metrics, the strategy will be abandoned before it matures, and a faster-cycling alternative may be the more realistic choice.
Does the organisation have a credible way to acquire the auxiliary signals (backlinks, brand searches, named-author entities) that curation amplifies but does not generate? If no, the curation strategy needs to be paired with a digital PR or partnership programme that builds those signals in parallel, and the joint cost should be compared against alternative strategies on the same basis.
Taking a position on curation
The position this article has argued is that editorial curation, properly understood as the disciplined exercise of judgement about what to publish and what to remove, is the higher-return strategy for AI search visibility in the majority of mid-market situations, and that the volume-first orthodoxy survives mostly because it is operationally familiar rather than because the evidence supports it. The position is contestable; the evidence is still emerging; and there are specific situations where volume genuinely wins. Within those qualifications, the strategic recommendation is clear, and the practical implications follow from it.
The first practical implication is that the next ninety days of any organisation taking AI search seriously should include a content audit with a real deletion budget. The drag from accumulated weak pages is the single largest impediment to AI citation performance for the typical mid-market site, and the cost of removing the drag is low relative to almost any new-content investment. Operators who have not done a deletion pass in the past twelve months are almost certainly leaving citation frequency on the table, and the recovery cycle is short enough to demonstrate within a single quarter. The operational discipline required is not technical; it is the willingness to overrule the internal voices that treat published pages as sunk assets rather than ongoing liabilities.
The second practical implication is that hiring decisions should change. A content team structured around production quotas needs at least one role explicitly structured around editorial judgement: someone whose performance is measured on what gets rejected, on the quality bar applied to what gets published, and on the integrity of the topical positioning over time. This role does not exist in most mid-market content teams, and creating it usually requires either an external hire or a deliberate redefinition of an existing senior content role. Without the role, the curation thesis cannot be operationalised; with it, the rest of the workflow becomes considerably easier to design. The role title matters less than the authority: the editor needs to be able to say no to pieces that other people want published, and that authority needs to be explicit rather than implicit.
The third practical implication is that measurement infrastructure for AI citations needs to be built before, not after, the strategy is launched. The strategy operates on lags that exceed normal reporting cycles, and without dedicated measurement it becomes politically indefensible long before it matures. The measurement does not need to be elaborate. A recurring two-hundred-query test across three AI systems, run monthly, is sufficient, but it needs to exist, and its outputs need to flow into the same reports that organic traffic and conversion metrics flow into. Operators who treat AI citation measurement as a future project rather than a current prerequisite will find that the future never arrives, because the strategy that depends on the measurement gets cancelled first.
Beyond these three operational moves, the deeper implication is cultural. The shift from volume to curation is, in the end, a shift in what an organisation considers valuable about its content investment. The volume mindset values output: pages produced, posts published, words shipped. The curation mindset values restraint: proposals rejected, archives pruned, standards enforced. Organisations that can hold the curation mindset will find that AI search systems reward them for it, in citation frequencies that volume competitors cannot match. Organisations that cannot hold it will continue to ship pages into a retrieval landscape that has quietly stopped paying for them. The choice is available; the evidence, in my reading of fifteen years of audits and the past three years of AI citation behaviour, points firmly in one direction.

