The Data That Trained on Itself
Licensing deals worth hundreds of millions a year are flowing to platforms, not the people who wrote the words.
Google pays Reddit roughly $60 million a year for access to its forums. OpenAI pays Reddit around $70 million for the same thing, and separately signed a five-year, $250 million deal with News Corp. None of the people who wrote the posts, comments and articles now changing hands for that money were party to any of these deals. They wrote for free, on platforms that were free to join, and the platforms sold what they wrote.
That should be the fact that stops you, before any argument gets built on top of it. Live-access licensing deals like these have grown from two in 2023 to a projected thirty-four in 2026 — a market that didn’t exist three years ago, already moving hundreds of millions of dollars a year, for something the AI industry used to scrape for nothing.
Why the market suddenly needs “real” data
The reason is a 2024 finding, published in Nature, that AI models degrade when they’re trained on their own output — not goodwill toward journalism or forum culture. Ilia Shumailov, a computer scientist then at Oxford, and his co-authors showed that models fed a diet of AI-generated text suffer what they called model collapse: an early drift away from the true distribution of language, followed by the permanent loss of rare, low-frequency information — the unusual phrasing, the minority view, the edge case that only shows up once in a million real sentences. Train a model on text that was itself written by a model, repeatedly, and each generation grows narrower and worse.
The supply of text that hasn’t been through that loop is shrinking fast. By April 2025, an estimated 74.2% of new webpages already contained AI-generated text. Epoch AI has projected that the internet’s stock of fresh, verifiably human-written material could be exhausted sometime between 2026 and 2032. Put those two facts together and the licensing boom stops looking like opportunism and starts looking like the market pricing in a resource it only just discovered was finite.
What buying real data gets you
For a lab that can afford it, a licensing deal buys something specific: a continuously refreshed pipeline of text a real, currently living human being wrote, with a paper trail behind it. That’s a genuine practical advantage: slower degradation, a defensible answer to “what did you train this on,” and a supply of language that keeps including the rare, awkward, human sentence that a model trained on its own past output would quietly lose.
Regulation has started reaching toward this problem, but not from the direction that matters here. The EU AI Act’s labelling mandate, which takes effect in August 2026, will require AI-generated images, video and audio to carry machine-readable labels, with fines of up to €15 million for non-compliance. Text is explicitly exempted — the European Commission’s own reasoning is that reliable detection of AI-written text doesn’t exist yet. So the one piece of law built to tag AI output has nothing to say about the market this article is describing: hundreds of millions of dollars a year, changing hands over exactly the kind of text the mandate can’t touch.
What it leaves out, and for whom
The money flows to the platform, not the person. Reddit and News Corp get paid because they own the pipe the content flows through; the individual who posted the comment or wrote the article does not see a share of either deal, and has no legal or technical route to ask for one. Nothing built so far reaches past the platform to the creator.
Smaller and open-source developers are shut out from a different direction. A verified, licensed data pipeline costs tens or hundreds of millions of dollars to assemble. A developer who can’t pay that price is left training on the same open web that’s now more than 74% synthetic, with no way to check what’s clean and what isn’t. The gap this opens is concrete: a well-resourced lab’s model gets steadily better data, a smaller developer’s model gets steadily worse data, and neither the developer nor their users are told which side of that split they’re on.
Government hasn’t closed the gap either. The UK’s March 2026 Report on Copyright and Artificial Intelligence chose to keep the existing rules in place, favouring voluntary licensing and transparency over reform, even though 88% of the more than 10,000 people who responded through the consultation’s Citizen Space platform backed the option that would have required licences for AI training — the strongest protection on the table. The market has already answered the question of what this content is worth — hundreds of millions of dollars a year and rising. Nobody in government has yet answered the question of who, if anyone, is owed a share of it.
My Opinion
I write a lot both here and in my published books, so I read this as someone whose words could end up in the same pipeline. The question of paying for what someone else created isn’t new — publishers have negotiated royalties for a century. What’s new is the scale: this isn’t licensing a book or a photograph, it’s the wholesale absorption of everything a person has ever written, folded into a model with no accounting for whose sentences did the work. Some form of monetary recognition has to follow the creator, not stop at the platform that happened to host them.
The AI industry has just told everyone, in real money, exactly what their unpaid words are worth. It has not told anyone how, or whether, they will ever see any of it.
You’re reading The Next Evolution by Neil Catton, articles that explore the human world and the intersection of technology, they try and ask difficult questions - not to scare - but to inform. If someone forwarded this to you, you can subscribe free at neilcatton.substack.com.
Neil Catton is the author of The Next Evolution, The Cognitive Crucible and The Shadow System - available on Amazon, and writes at the intersection of technology, ethics, and human purpose.


