On July 20, 2026, an arXiv preprint numbered 2607.18045, titled “The Shared Discovery Paradox,” put formal numbers on a failure mode that AEO practitioners have been circling anecdotally for a year: when a search or discovery system converges on a single shared answer, improving the underlying information pool can make individual search outcomes worse, not better. The practical consequence for anyone optimizing content for AI-assistant citations is that the standard playbook, publish the best answer and wait to be cited, inverts under one-answer consolidation. The paper is a theoretical model, so the direction of the effect is actionable while its specific magnitudes are not, and that distinction shapes everything below.
What does the Shared Discovery Paradox actually model?
The paper models a discovery game in which many agents hunt for answers and can either follow their own private clues or copy a pooled, shared recommendation, and it proves that copying the pooled recommendation can raise the accuracy of the single best answer while collapsing the rate at which the group as a whole finds anything. The setup is deliberately stylized: agents hold partial information, a pooling mechanism aggregates that information into a shared report, and each agent chooses between acting on its own signal or repeating the consensus. That choice structure maps directly onto how AI search products behave today. An assistant that retrieves, synthesizes, and returns one synthesized answer is a pooling mechanism, and every user who accepts that answer without clicking through is an agent repeating the shared recommendation.
The paradox in the title is the gap between two things that intuition says should move together. Pooling information should help, because aggregation averages out individual noise. And it does help the pooled artifact itself: in the paper’s benchmark, the accuracy of the best single recommendation rises from 0.20 to 0.3835 once information is pooled, according to the preprint. The catch is what agents do with it. When everyone repeats the same recommendation, the group stops exploring, and discovery, the probability that someone in the group actually finds the target, craters. Better shared information, worse collective search. That is the whole paradox in one sentence, and it is worth pausing on because it runs against the default assumption baked into most content strategy: that making the canonical answer better is strictly good for the ecosystem that produced it.
The paper is careful about scope. It does not claim consolidation is always harmful. It identifies regimes where a shared answer is exactly what you want: consensus formation on settled questions, and query efficiency when the cost of redundant search outweighs the value of diversity. If ten thousand people ask the same factual question and nine thousand nine hundred of them get the right answer from one shared response, that is a win, not a monoculture. The failure mode appears specifically when the task is discovery, when the goal is to surface something new, under-covered, or contested, rather than to confirm something known. Much of what publishers and documentation teams actually care about lives in the discovery regime, which is why the model lands uncomfortably close to home.
What happens to discovery rates when everyone repeats the best answer?
Group discovery falls from 0.8322 under decentralized clue-following to 0.3835 when every agent repeats the pooled recommendation, a collapse of more than half, according to the benchmark results in arXiv:2607.18045. Read those two numbers against each other. Decentralized search, where each agent follows its own imperfect private signal, discovers the target about 83 percent of the time. Full consolidation on the pooled answer, an answer that is itself nearly twice as accurate as any individual’s best guess, discovers it about 38 percent of the time. The group traded a diverse portfolio of mediocre searches for a monoculture of one decent search, and the monoculture lost.
The mechanism is correlation. Decentralized agents fail independently; their errors partially cancel and their successes cover different regions of the search space. Agents repeating a shared recommendation fail together, because they have all placed the same bet. When the pooled recommendation is wrong about where to look, everyone is wrong simultaneously, and there is no residual diversity left to pick up the slack. This is the same reason a portfolio of uncorrelated assets beats a concentrated position at equal expected return, transplanted into information search. It is also, not coincidentally, the reason a search results page with ten distinct sources historically surfaced more total correct information than a single synthesized paragraph that averages them.
For the publishing side of the market, the implication is sharper than it first looks. Under consolidation, the citation becomes winner-take-all: the assistant picks one source to ground its answer, and every equally correct alternative source is suppressed. The 0.3835 discovery figure is what happens to the ecosystem. What happens to the individual publisher is a positional tournament in which being marginally better than the consensus clone yields near-zero marginal reward, because you either win the single slot or you are invisible. Investing heavily to be 5 percent better than the next equally correct explainer is, under this model, roughly the same expected citation outcome as doing nothing, while collectively degrading the diversity that makes the discovery layer useful at all.
Where is the tipping point between consolidation and diversity?
In the paper’s canonical environment, a consolidated market overtakes decentralized report-following at a copying probability of c = 0.788462, per the analysis in arXiv:2607.18045. Below that threshold, enough agents are still following their own signals that consolidation’s accuracy gains outweigh its diversity losses. Above it, the monoculture effect dominates and the shared-answer regime produces worse group outcomes than everyone fending for themselves. The number itself is precise to six decimal places, which should be read as a property of the math rather than a prediction about any real product; nobody has measured the effective copying probability of a production assistant, and the paper does not pretend otherwise.
What matters is that a threshold exists and that the system’s position relative to it is what determines whether “improve the pooled answer” helps or hurts. This reframes a question the industry currently argues about with vibes. Debates about whether AI Overviews are “good for the open web” are usually conducted at the level of sentiment. The model says the honest version of that question is quantitative: what fraction of users accept the synthesized answer and stop, and where does that fraction sit relative to the regime boundary? An assistant whose users mostly click through behaves like decentralized search with a nice summary on top. An assistant whose users mostly stop at the answer is operating above the tipping point, and every improvement to its synthesis quality deepens the discovery collapse.
The paper also computes the cost of uncoordinated behavior directly. In the canonical equal-split game, the anonymous symmetric equilibrium, which the paper characterizes as a water-filling rule where agents spread effort across actions until marginal returns equalize, achieves a value of 0.5991, and the exact mixed price of anarchy is 2 - 1/N, meaning equilibrium losses relative to the coordinated optimum shrink toward zero as the number of players N grows, per the preprint. That second result cuts in an interesting direction. With many independent agents, the gap between what selfish optimization achieves and what coordination would achieve narrows. With few agents, or equivalently with many agents whose behavior is correlated through a shared recommendation, the gap is wide, and consolidation sits exactly in the correlated regime.
Does LLM consensus mean correctness?
No, and a companion audit posted this month quantifies the gap: in arXiv:2607.08065, the most consistent frontier model reached agreement of 0.8 or higher on 77 percent of GPQA case-result entries, yet 48 percent of those high-agreement entries were wrong. Consensus among language models, or between a model and itself across repeated sampling, is a calibration signal of the weakest kind. It measures how confidently the system repeats itself, not how often repetition tracks the truth. Nearly half the time in that audit, high confidence and high agreement coexisted with a wrong answer.
This finding matters to the Shared Discovery Paradox as more than a footnote. One defense of one-answer search is that synthesis quality keeps improving, so the single answer is increasingly likely to be right, which softens the diversity loss. The agreement audit undercuts the version of that defense that runs on self-consistency: if you are using the assistant’s own confidence, or agreement between assistants trained on overlapping corpora, as evidence that the shared answer is reliable, you are leaning on a signal that failed on roughly half of high-agreement cases in a graduate-level science benchmark. Correlated confidence is what you would expect from models sharing training data and fine-tuning regimes, which is exactly the correlation structure the paradox warns about.
There is a second-order consequence for evaluators and procurement teams. “The assistant consensus must be right” is an assumption that shows up quietly in how organizations deploy these systems: if two or three assistants agree, the answer ships without verification. The audit suggests that agreement between models is evidence of shared training distribution more than shared correctness, and the discovery model suggests the same correlation failure at the ecosystem level. Both point the same way. Independence is the scarce resource, in search results and in model ensembles alike.
What does this mean for AEO strategy?
The model’s sharpest strategic implication is that query intent splits into two regimes with opposite optimal strategies, and treating them the same is the expensive mistake. Under winner-take-all intent, where the assistant converges on one answer and cites one source, competing to be a marginally better version of the consensus answer yields diminishing returns and suppresses equally correct alternatives. Under multi-answer intent, where assistants still surface several sources, lists, comparisons, contested questions, regional or situational variance, the old rules apply and quality improvements still convert into citations. The paper’s practical verdict, as summarized in the research behind this piece, is to stop treating “publish the single best answer” as a universal strategy and start segmenting by which regime a query type actually lives in.
That segmentation is observable in practice even without the paper’s formalism. Factual lookups with settled answers behave winner-take-all; there is one correct boiling point for water at sea level and the assistant will say it once. Comparison queries, tool-selection queries, and anything with legitimate tradeoff structure still surface multiple answers, because collapsing them to one answer would be visibly wrong to the user. The strategic move is to concentrate depth investment in the second category: queries where multiple correct framings coexist, where a differentiated angle earns its own citation slot rather than competing for the single slot against the consensus clone. The alternative, spending to out-answer the incumbent canonical answer by a margin users cannot perceive, is precisely the behavior the model shows producing near-zero marginal citation reward under consolidation.
The paper also names the vendor dynamic to watch. AEO tooling that sells “optimize to be THE answer” without distinguishing intent types is selling a playbook the model shows failing exactly when consolidation is strongest. That does not make such tooling useless; in fragmented, multi-answer query spaces, positional optimization still works the way its marketing claims. It makes the tooling mispriced for consolidated intent, where the marginal citation probability of being the second-best correct answer rounds to zero. Buyers should ask any AEO vendor a simple question: does your optimization target winner-take-all intent or multi-answer intent, and how do you tell the difference? A vendor without an answer is optimizing you for a tournament where second place is indistinguishable from last.
Does coordinated differentiation beat solo optimization?
Yes, in the model’s cleanest result: a coordinated eight-action portfolio using the same pooled reports reaches a discovery rate of 0.8594, above both the 0.8322 decentralized baseline and far above the 0.3835 pure-consolidation outcome, according to arXiv:2607.18045. This is the counter-evidence that keeps the paper from being a simple argument against pooling. Pooling is not the problem. Repetition is. When agents share information and then deliberately cover different actions instead of all copying the single best recommendation, they get the accuracy benefits of aggregation and the coverage benefits of diversity simultaneously. Coordinated diversity on pooled data is the best of the three regimes the paper tests.
Translated to publishing: the countermove to consolidation is not to abandon shared knowledge or to write worse content on principle. It is to coordinate coverage so that the ecosystem, whether that means a company’s own content portfolio, a partner network, or a documentation program, occupies differentiated positions rather than stacking duplicates on the consensus answer. One canonical page, plus deliberately angled coverage of adjacent framings, edge cases, and query variants, dominates eight pages each trying to be the single best answer to the same question. The model formalizes something good content strategists already practice under the name of topical authority, but it adds a warning: the strategy works through differentiation, and it fails the moment the portfolio collapses into near-duplicates chasing the same citation slot.
There is also a coordination problem hiding here, and the paper’s price-of-anarchy result (2 - 1/N, equilibrium value 0.5991 in the canonical game) quantifies it. Solo optimizers, each picking what looks individually best, do not reach the coordinated portfolio outcome on their own; they converge toward the water-filling equilibrium instead, which is better than monoculture but short of the optimum. In practice that means differentiated coverage has to be a deliberate allocation decision, not an emergent property of many writers each following the same keyword tools. Keyword tools are, structurally, pooling mechanisms that recommend the same targets to everyone who subscribes. Treating their top recommendation as your top priority is the publishing equivalent of every agent copying the shared report.
What are the limitations of the model?
The strongest limitation is that the Shared Discovery Paradox is a formal game-theoretic model, not an empirical study of production AI search, so its thresholds describe a stylized environment and nobody knows where real assistants sit relative to them. The copying probability tipping point of 0.788462, the discovery figures 0.8322 and 0.3835, and the portfolio figure 0.8594 are properties of the paper’s constructed benchmark. Whether ChatGPT, Perplexity, or Google’s AI surfaces operate above or below the effective tipping point is unmeasured. The direction of the effect, consolidation degrades discovery once correlation gets high enough, is robust within the model and consistent with independent results like the LLM agreement audit. The magnitudes are not measurements and should not be quoted as if they were.
The model also compresses a lot of real-world mess into clean parameters. Real users do not face a binary copy-or-search choice; they skim, partially verify, reformulate, and occasionally click through in patterns no single copying probability captures. Real assistants mix behaviors by query type, sometimes returning one synthesized answer and sometimes a citation wall, sometimes in the same session. And the publisher side has its own complications the model abstracts away: citation selection is not purely meritocratic, grounding sources are chosen by retrieval pipelines with their own biases, and “equally correct” is doing heavy lifting anywhere the correct answer is contested. None of these gaps invalidate the framework. They do mean the framework is a lens for reasoning about incentives, not a calculator for predicting traffic.
Context on the venue is worth one paragraph. The paper appears on arXiv, which as of July 1, 2026 operates as an independent nonprofit after separating from Cornell University, a transition announced in March 2026. The platform was handling roughly 24,000 submissions per month as of November 2024, run by approximately 27 staff on a $6 million annual budget. Two things follow. Preprints on arXiv are not peer-reviewed, and this one should be read with the same discount applied to any fresh theoretical result. And the volume figure is itself a small illustration of the paper’s subject: 24,000 papers a month competing for discovery through interfaces that increasingly return one answer at a time. The discovery problem the paper models is not abstract for the platform hosting it.
Where should content effort go now?
The practical verdict is to stop treating “publish the single best answer” as a universal strategy, segment query targets by intent regime, and concentrate differentiation where assistants still surface multiple answers, while treating any coordinated coverage of pooled intelligence as a portfolio allocation problem rather than a race to clone consensus. Concretely, that means three reallocations. First, shift depth investment away from settled factual queries, where one-answer consolidation already dominates and marginal quality buys no marginal citation, toward comparison, selection, and contested-tradeoff queries where multiple framings still earn slots. Second, within any content portfolio, audit for near-duplicate pages chasing the same consolidated answer and differentiate or consolidate them; the model says the portfolio wins through coverage, not through stacking. Third, discount any AEO tooling or playbook that optimizes for THE answer without an intent-type segmentation behind it.
Carry the skeptical counterweight alongside the strategy. The evidence for the direction of the consolidation effect is now twofold, a formal model showing the mechanism and an independent audit showing that model agreement fails as a correctness proxy on 48 percent of high-agreement cases, but neither tells you where your specific query space sits relative to the tipping point. That measurement problem is open, and it is the right next question: someone will build the tooling that classifies live query classes by consolidation behavior, and teams that have already segmented their content by intent regime will be positioned to act on it. Until then, the defensible position is the boring one. Bet on differentiated coverage, treat consensus-cloning as a depreciating asset, and do not mistake the assistant’s confidence for the world’s agreement.
Frequently Asked Questions
How does the water-filling equilibrium differ from the coordinated portfolio outcome?
The water-filling equilibrium is an uncoordinated strategy where agents spread effort until marginal returns equalize, achieving a value of 0.5991 in the canonical game. This falls short of the coordinated portfolio outcome of 0.8594 because solo optimizers cannot internalize the diversity benefits their actions provide to the group. The gap between these two states is quantified by the price of anarchy, which shrinks as the number of players increases.
What is the practical impact of arXiv’s 24,000 monthly submissions on discovery?
The volume of 24,000 monthly preprints creates a dense information environment where traditional search surfaces diminishing returns. As assistants converge on single answers to manage this volume, the risk of the Shared Discovery Paradox increases because the system prioritizes the most statistically probable consensus over novel, uncorrelated findings. This dynamic forces publishers to compete for a single citation slot against a backdrop of massive, uncoordinated output.
Why is model self-consistency a poor confidence signal for search systems?
An audit of frontier models found that 48 percent of entries with high self-agreement were factually wrong, despite 77 percent of case results showing agreement above 0.8. This indicates that self-consistency measures the stability of a model’s internal distribution rather than its alignment with external truth. Search systems relying on this signal risk amplifying confident errors, as the model repeats its own biases without external verification.
How does the price of anarchy change as the number of agents grows?
The mixed price of anarchy in the equal-split game is exactly 2 - 1/N, meaning the relative loss from uncoordinated behavior decreases as the number of players N increases. In large populations, the gap between the selfish water-filling equilibrium and the coordinated optimum narrows significantly. However, in smaller or highly correlated groups, the inefficiency of uncoordinated discovery remains substantial.