The first phase of the AI-data conflict was extraction: models were trained on the open web, and publishers discovered it after the fact. The second phase was legal: lawsuits from the New York Times, authors, artists and stock-image libraries, alongside a fair-use doctrine suddenly carrying billions of dollars of weight.
The third phase, now underway, is commercial. Data has a price, contracts have been signed, and the economics of publishing are adjusting around a customer that reads everything and links to nothing.
Why it matters
The open web's implicit bargain — content in exchange for traffic — breaks when the reader is a model that answers instead of referring. Licensing is the candidate replacement: if AI companies pay for the data that makes their answers good, quality publishing retains a revenue line. If the market fails, the incentive structure points toward paywalls, blocking, and a web that gets steadily worse for everyone, including the models.
How it works
Deals fall into two shapes. Training licenses grant access to archives for model training — the Associated Press, Axel Springer, News Corp, the Financial Times and others have signed these with OpenAI and peers. Retrieval deals cover real-time access: the AI product queries the publisher's content at answer time, with attribution, turning the publisher into a cited source rather than training exhaust. Reddit and Stack Overflow license their corpora on both theories, having discovered their user-generated archives are among the most valuable training assets in existence.
The technical layer is catching up to the commercial one: robots.txt remains the blunt instrument, but proposals for machine-readable licensing signals and per-crawler pricing are moving from blog posts to standards discussions.
Evidence
The signed deals are public: News Corp's agreement with OpenAI was reported at over $250 million across five years; Reddit's data licensing generates tens of millions per quarter and appears in its SEC filings as a distinct revenue line; Stack Overflow, Shutterstock, and a long list of news publishers have announced agreements. The litigation track is equally public, with courts beginning to rule on the fair-use questions in 2025 — early decisions gave AI companies qualified wins on training while leaving the market for licensing intact.
The competing read
Publishers frame licensing as existential: without it, AI products free-ride on journalism until there is no journalism left to ride. AI companies counter that training on public data is transformative fair use, that licensing at scale is impractical across the whole web, and that retrieval-with-citation delivers value back. The courts will settle the doctrine; the market is already settling the practice — the largest players are paying, which is its own signal about how confident their lawyers are.
What happens next
Watch the long tail: the deals so far cover the biggest archives, but the web's value is its breadth. If collective licensing or marketplace mechanisms emerge for smaller publishers, the model scales; if not, the two-tier web hardens. Also watch the retrieval tier — per-answer citation economics may prove larger than the training tier, because it recurs forever.
