Imagine a local reporter spends two weeks investigating a zoning scandal. In the old internet arrangement, curious readers would search for the story, click through to her newspaper’s website, see ads, maybe subscribe, and leave clicks and links that told search engines her article was worth finding. Now imagine that same reader typing the question into a chatbot, getting a tidy summary that draws on her reporting, and never visiting the site at all. The reader is satisfied. The reporter’s paper gets nothing.
That scenario is the starting point for a new NBER working paper by Alex Chan of Harvard Business School, who asks what happens to the economics of the open web when generative AI systems answer users’ questions directly instead of routing them to the sources that produced the underlying information. His conclusion, developed through a formal model, is that even when every AI answer is accurate and every user acts rationally, the arrangement that has financed online content for two decades can quietly stop working.
The two things a visit used to do
Chan begins by pointing out that a visit to a website has historically done two jobs at once. First, it created a “revenue event”: ads shown, subscriptions converted, affiliate links clicked. Second, it created a “measurement event”: data on which articles readers returned to, linked to, bookmarked, or corrected. Search engines used those signals to figure out which sources were worth surfacing in the future.
Generative AI answer engines, Chan argues, can absorb both events. When a chatbot synthesizes information from a publisher’s article without sending the user to the page, the publisher loses the ad revenue and the signal that would have told ranking systems the article was useful. Chan frames this as a “missing market” problem rather than a story about bad AI or misinformation. Even with truthful publishers, accurate answers, and reasonable users, he writes, the platform “can appropriate current answer value without fully paying for the future reproduction value and source-level information value of publisher traffic.”
A participation constraint and a reproduction number
The mathematical spine of the paper is what Chan calls a participation constraint. If producing a piece of costly content, say, investigative reporting, expert review, or product testing, requires some minimum expense, and if the share of paid attention reaching the publisher shrinks toward zero, then the publisher’s rational choice eventually collapses to producing nothing. Chan proves this without assuming a specific demand curve, a smooth cost function, or a single dominant publisher. All he needs is that positive content has positive avoidable cost and that revenue per visit is finite.
He then extends this into a dynamic model. Content today generates immediate traffic and durable “attention capital”: subscribers, backlinks, bookmarks, reputation, search authority. That capital produces future visits, which finance future content. Chan derives what he calls an open-web reproduction number, borrowing the mathematics epidemiologists use to determine whether an infection sustains itself. If each generation of costly content produces less than one generation of replacement content, the system is “subcritical,” and the stock of costly human information decays toward zero.
The number depends on how much monetizable attention publishers retain. AI diversion lowers it. So does anything that makes conventional search less good at surfacing genuine sources. When the reproduction number drops below one, Chan shows, every path within his no-transfer upper bound converges to zero over time.
Why the AI platform’s math and society’s math differ
The paper’s central strategic result concerns the choice facing an AI platform: how much traffic to send back to publishers. Chan compares this choice to what a social planner would pick. Both parties value future content quality to some degree, but the platform only captures part of the continuation value. It does not fully internalize publisher revenue, the public-good value of verified information, the diversity of topic coverage, or the source-level measurements that also help rival systems function.
Using tools from lattice theory, Chan shows that whenever the platform places less weight on continuation value than the social planner does, the platform will retain weakly less referral traffic. If the platform’s chosen level of diversion pushes the reproduction number below one and a feasible policy exists that would have kept costly information alive at higher social value, the private outcome is inefficient. Chan is careful to note the result is one-sided: proving the system is above the survival threshold is not the same as proving survival, which also requires reinvestment, entry, and functioning ranking systems.
More pages, less information
One counterintuitive prediction is what Chan calls the “abundance paradox.” Because generative AI also makes it cheap to produce web pages, the raw count of URLs can rise even as the stock of costly human information falls. Cheap synthetic pages can flood the graph that link-based ranking systems like PageRank rely on. Chan shows that if synthetic pages dominate the pool of indexed content and random walks rarely escape synthetic clusters, PageRank mass on genuine pages can approach zero, unless the search engine deliberately reserves exposure for verified sources through what Chan calls “trust anchors.”
He treats visits and links not just as payments but as statistical experiments. Drawing on Blackwell’s classical comparison of experiments, Chan shows that higher AI diversion makes the old web’s quality signals less informative. Answer-level feedback, a thumbs-up on a chatbot response, cannot generally identify which source was original, which was stale, or which corrected an error, because the same answer can be produced from different underlying source qualities.
A feedback loop
Chan describes a tipping condition under which the situation can reinforce itself. If conventional search becomes less useful because genuine sources are harder to find, users have more reason to rely on AI answers. Greater AI reliance removes further visits and links, which further weakens conventional search. The paper models this as a contraction: once search quality drops into a certain region, the loop drives effective publisher monetization and source-level measurement toward zero together.
What a repair mechanism would need to do
Chan devotes a substantial section to what he calls “market design” for the problem. He is explicit that he is not proposing a fully worked-out institution, only listing what any workable one would have to do. His five elements are: measure how much traffic AI answers actually displace; pay a “visitor-replacement royalty” that restores what publishers would have earned; audit the provenance of information used in AI answers; verify that paid-for content is genuinely costly human information rather than cheap synthetic imitation; and direct compensation toward “keystone topics” whose reproduction is most important to the wider network.
The audit requirement matters because Chan proves a screening result: if a mechanism cannot statistically distinguish costly human sources from cheap synthetic mimics, paying for “quality” attracts imitators. He derives a likelihood-ratio condition an audit must meet to overcome the cost advantage that synthetic content enjoys.
The keystone-topics idea pushes back against the intuition that payment should follow current clicks. Using the Perron eigenvector of his reproduction operator, Chan argues that scarce compensation is best directed to topics that sit in central positions in the network of content reproduction, which may not be the topics with the most current traffic. Local news, technical maintenance, and minority-language content, he notes, tend to have high fixed costs and low direct-click monetization, precisely the profile that AI diversion selects against.
The scope of the claim
Chan is deliberate about what his paper does not say. It does not claim every website will disappear, that all AI reduces content diversity, or that modern search must fail. Search engines with hard verified-source anchors, provenance rules, or protected exposure for trusted publishers can avoid the collapse channel he describes. AI systems whose retrieval policies preserve heterogeneous interests need not concentrate source attention. And where content is produced for reasons other than commercial revenue, hobby sites, government pages, philanthropically funded outlets, the participation constraint does not bind in the same way.
The paper’s empirical predictions, which Chan leaves for future work, include declines in expensive content types in categories most exposed to AI answers, concentrated exit among high-fixed-cost topics, falling entropy in source attention, and weakening predictive power of link-based ranking signals over time.




