Pollution with a second life¶
The complaint about generated material is usually a complaint about taste, which is why it goes nowhere. Taste arguments cannot be won, and the people making the material have a serviceable answer: nobody is obliged to read it. The property doing the damage is not quality. It is that the material does not stay where it lands.
Four destinations for a generated page¶
A person reads the page and decides whether it is any good. Ordinary, familiar, and where most accounts of slop stop.
A crawler collects it and it enters a training set. The substrate is contaminated, and the contamination is discovered, if at all, a generation later.
A model reads it while scoring another model’s answer, or it ends up inside a benchmark. The scorer is contaminated, which is worse and much quieter, because a broken scorer reports success.
Code is written from it, or the page was code, and it enters the tooling that produces the next system. Recursion arrives in the development loop.
Only the first is a taste problem. The other three change what the next system is made of, and none of them requires that anybody read the page.
The collapse result is narrower than its reputation¶
Model collapse is the term the argument reaches for, usually as a prophecy: the machines eat their own output and degrade. The published result is narrower and more useful than the prophecy. Shumailov and colleagues, in The Curse of Recursion, found that training recursively on model-generated content causes irreversible defects, with the tails of the original distribution disappearing. Gerstgrasser and colleagues then ran the distinction the prophecy skips: if generated data replaces the real data at each round, test error increases with every iteration, and if generated data accumulates alongside the real data, the error has a finite upper bound and collapse does not occur.
What decides the outcome is therefore not the presence of synthetic material. It is whether anything keeps the originals and filters the additions. Synthetic data is ordinary practice, and the practice is inseparable from the filtering: the phi-4 report describes a model that “strategically incorporates synthetic data throughout the training process” on “a training recipe that is centrally focused on data quality”. The filter is doing the work the data gets credit for.
And the filter is increasingly a model. That is the same recursion arriving one level up, at the point where it is least visible.
Volume is not readership¶
Graphite, examining 43,000 English-language URLs from Common Crawl published between January 2020 and May 2025, found that in November 2024 the quantity of AI-generated articles being published passed the quantity of human-written ones, and that the proportion has been relatively stable since. The same analysis found that these articles “largely do not appear in Google and ChatGPT”, and its authors say they suspect the articles are not viewed in proportion to their number.
Ahrefs, sampling 900,000 new English-language pages in April 2025, one per domain, found 74.2 per cent containing some AI-generated content, with 2.5 per cent classed as pure AI, 25.8 per cent as pure human, and 71.7 per cent a mix of the two.
Two measurements of two different things, and neither of them measures quality. Read together they describe something more specific than a flood: a very large quantity of material that few people ever see, accumulating in the places downstream systems collect from. Whether it reaches the next training set, the next benchmark or the next retrieval index is a question about how those are assembled, and nothing in the way this material was produced argues against it.
The pollution metaphor holds for the first destination and misleads for the rest. This is not litter. It is feedstock.
The clerk’s brief¶
From the clerks, for the Patrician’s eyes
Compiled August 2026. Newest first; settled items sink into the assessment at the end. The clerks note that counts of published material are easy to obtain and counts of read material are not, and have filed accordingly.
April 2025: Most new pages are neither one thing nor the other¶
Ahrefs’ sample of 900,000 pages found pure machine authorship at 2.5 per cent and pure human authorship at 25.8 per cent, with the remaining 71.7 per cent a mixture. The clerks observe that a labelling scheme, a detection tool and a policy on machine-generated content all assume a category boundary which most of the material does not respect.
November 2024: More articles published by machine than by person¶
Graphite’s analysis of Common Crawl put the crossover here, and reported the proportion broadly flat across the twelve months to May 2025, where its sample ends. The crossover is a publishing statistic. The finding sitting underneath it, that the articles largely do not reach readers through search or assistants, is the one the clerks consider worth the Patrician’s minute: the market for this material is not attention, and something else is paying for it.
2024: The curse of recursion, and its condition¶
The collapse literature reached the public as a prophecy and the laboratories as an engineering constraint. Shumailov and colleagues established the defect; Gerstgrasser and colleagues established the condition, which is data replacement rather than data presence. The clerks record that the response in practice is to keep the archives and filter the additions, with data quality treated as the recipe rather than a precaution, that the filters are themselves increasingly machine-run, which leaves the clerks with the less fashionable question of who audits the filters.
Filed under feedstock¶
The volume figures are the least interesting part of the file. What they establish is that a large and growing share of the material available for collection was produced at negligible cost, without requiring a reader, by systems whose successors can be built from it. The clerks’ standing assessment is that the pollution framing has been a distraction, because pollution is something a population is exposed to, and this is something the next generation of machines will be built out of.