The loop’s own paperwork (a resource collection)¶
Almost every primary source below was produced by the field examining itself: laboratories publishing on the limits of laboratories, benchmark authors on the failure of benchmarks, a research system evaluated by researchers. The field is both the subject and the maker of most of the instruments used to examine it, which is not a reason to discount the material. It is a reason to note who held the instrument. Sources from outside that production system, a European institution counting electricity and journalists examining labour and litigation, are identified where they supply a different vantage point.
Measuring the substrate¶
More articles are now created by AI than humans, Graphite, 2025. The crossover study: 43,000 English-language URLs from Common Crawl, published between January 2020 and May 2025, with machine-written articles passing human-written ones in November 2024 and the proportion broadly flat since. The finding under the headline is the more useful one, where the authors record that these articles “largely do not appear in Google and ChatGPT” and that they suspect readership does not follow publication.
What percentage of new content is AI-generated, Ahrefs, 2025. A different question and a different answer: 900,000 new English pages sampled in April 2025, one per domain, with 74.2 per cent containing some machine-written text, 2.5 per cent classed as pure machine, 25.8 per cent as pure human, and the remaining 71.7 per cent a mixture. The mixture figure is the useful one, since it is the category a binary human-or-machine distinction cannot represent. Produced by a company selling search tooling, using its own detector.
2025 Bad Bot Report, Imperva. Automated traffic at 51 per cent of web traffic in 2024, above human activity for the first time in a decade, with bad bots at 37 per cent. A vendor report, counting requests rather than readers, and useful here as a measure of how much web activity is now automated rather than of how much of it is generated.
Collapse, and the condition attached to it¶
The Curse of Recursion: Training on Generated Data Makes Models Forget, Shumailov, Shumaylov, Zhao, Gal, Papernot and Anderson, 2023. The paper the whole model collapse argument descends from, showing that replacing real training data with model-generated content can produce irreversible defects, with the tails of the distribution going first.
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data, Gerstgrasser and colleagues, April 2024. The condition that rarely travels with the prophecy: replacing real data with generated data sends test error up with every iteration, while accumulating generated data alongside the real leaves the error with a finite upper bound. Anyone citing collapse without this paper is citing half the result.
Machines designing machines¶
Neural Architecture Search with Reinforcement Learning, Zoph and Le, 2016. An early and influential example of machines proposing designs for machines, scored by validation accuracy. A decade on it reads as ordinary technique rather than as a threshold being crossed.
Discovering faster matrix multiplication algorithms with reinforcement learning, Fawzi and colleagues, Nature, October 2022. AlphaTensor, including the 47-multiplication algorithm for four by four matrices in a finite field against Strassen’s 49. The setting travels badly: the headline case is modular arithmetic, and the practical speedups reported are hardware-specific.
Faster sorting algorithms discovered using deep reinforcement learning, Mankowitz and colleagues, Nature, June 2023. AlphaDev, and perhaps the strongest of these cases, because the routines went into the LLVM standard C++ sort library and stayed there. Improvements of up to 70 per cent for sequences of length five and roughly 1.7 per cent above 250,000 elements. Both Nature links may bounce through a cookie gate before settling.
Evaluating Sakana’s AI Scientist: Bold Claims, Mixed Results, and a Promising Future?, Beel, Kan and Baumgart, February 2025. The independent read of an automated research system: 42 per cent of experiments failing on coding errors, established concepts reported as novel, a median of five citations mostly predating 2020, and some hallucinated numerical results, at a reported six to fifteen dollars per paper and around three and a half hours of human involvement.
Measuring AI Ability to Complete Long Tasks, METR, March 2025. The time horizon measure, with a doubling of roughly seven months across six years of models. The page itself now records that the figure is out of date and superseded, which is a better advertisement for the organisation than the graph everybody reposts.
Phi-4 Technical Report, Abdin and colleagues, December 2024. Useful for what it says about method rather than about the model: synthetic data incorporated “throughout the training process” on “a training recipe that is centrally focused on data quality”, and a system reported to surpass the teacher it was built from. A laboratory describing its own recipe, so read as an account of practice rather than an independent finding.
The instruments¶
Reflections on Trusting Trust, Ken Thompson, the 1983 Turing Award lecture as published in Communications of the ACM in August 1984. Three pages, a scanned PDF, and still the clearest statement of what a closed loop can hide: a compiler that reinserts its own alteration with nothing in the source to show for it, and the moral that inspection cannot settle the question from inside.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Zheng and colleagues, 2023. The paper that helped establish LLM-as-judge as a practical evaluation method, reporting over 80 per cent agreement with human preferences, roughly the level of agreement between humans. It names position bias, verbosity bias and self-enhancement bias in the same document, qualifications that sit rather less comfortably in the headline version of the result.
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges, Chehbouni, Haddou, Cheung and Farnadi, August 2025. Measurement theory turned on the practice, reporting that adoption “has outpaced rigorous scrutiny of their reliability and validity as evaluators”. The most direct statement available of the gap between how fast the instrument spread and how slowly it was checked.
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework, Li and colleagues, April 2026. 11,727 studies screened, 49 included, risk of bias testing absent in 36 of them and one study at production. The number to carry away is not the adoption rate but the validation rate.
Benchmarks as Microscopes: A Call for Model Metrology, Saxon, Holtzman, West, Wang and Saphra, July 2024. The argument from inside the field that static benchmarks saturate without telling anybody what a system will do once deployed, and that broad claims about reasoning rest on weak metrics.
Computing Power and the Governance of Artificial Intelligence, Sastry, Heim, Belfield, Anderljung, Brundage and colleagues, February 2024. The argument behind every compute threshold: compute is “detectable, excludable, and quantifiable, and is produced via an extremely concentrated supply chain”, which is what makes it a lever where capability is not. Written largely by people who work on AI policy, which is the vantage point rather than a disqualification.
The physical loop¶
Data centres: an energy-hungry challenge, European Commission, November 2025. Around 415 terawatt hours globally in 2024, some 1.5 per cent of world electricity, projected to more than double by 2030, with EU consumption at about 70 terawatt hours heading for 115, and the observation that new demand tends to meet a lack of available grid connection capacity.
[AI and the energy sector](https://www.europarl.europa.eu/RegData/etudes/BRIE/2025/775859/EPRS_BRI(2025) 775859_EN.pdf), European Parliamentary Research Service, July 2025. The same problem in local units: about 3 per cent of EU electricity demand, above 20 per cent in Ireland, geographic clustering of AI facilities, and the Cloud and AI Development Act’s intention to triple EU data centre capacity within five to seven years.
The statute itself¶
Regulation (EU) 2024/1689, the AI Act, in the Official Journal text. Article 51 sets the presumption of high impact capabilities at cumulative training compute above 10^25 floating point operations, adjustable by delegated act. Article 55 requires providers of such models to perform their own evaluation to state-of-the-art protocols, including documented adversarial testing. Article 92 lets the AI Office evaluate such a model itself and demand access through interfaces and source code. Read together, the three mark out what a regulator may reach for and what it still takes on trust.
Who does the checking¶
OpenAI Used Kenyan Workers on Less Than $2 Per Hour, TIME, January 2023. The outside view, and the only source here written by someone with no stake in the field’s self-assessment. Take-home pay of around $1.32 to $2 an hour for labelling descriptions of abuse, torture and worse, against $12.50 an hour billed to the client, with the work cancelled in February 2022, eight months early.
Meta and Sama lawsuit, on working conditions and human trafficking, Business and Human Rights Resource Centre. A maintained timeline of the Kenyan litigation, filed May 2022 and still short of a final determination, with the rulings and the parties’ statements attached.
Jurisdiction over Meta Inc. in Kenyan courts, Toussaint Nothias, conflictoflaws.net, March 2026. Three separate suits laid out side by side, which is what it takes to keep them apart: the whistleblower claim, the moderators’ dismissal claim, and the Ethiopian hateful content claim. The private international law framing is the useful part, since jurisdiction is where all three have spent their time.
Provenance, added afterwards¶
How we’re helping creators disclose altered or synthetic content, YouTube, 18 March 2024, and Meta’s approach to labelling AI-generated content and manipulated media, April 2024. Both schemes begin at upload, resting on creator disclosure supported by detection of industry-standard indicators, and therefore say little about provenance before material reaches a platform.