Where the lever reaches

The argument about regulating these systems assumes the noun names one thing a lever can be attached to. It does not. Models, training runs, compute, published research, research assisted by machines, research designed by machines, and the tooling that produces all of it are seven different layers, in one loop, with entirely different degrees of exposure to anybody holding a rulebook.

The question that decides whether a rule does anything is which layer it can actually touch.

The law attaches to a product

European law reaches these systems through the architecture it uses for kettles and medical devices: obligations attach to placing something on the market or putting it into service, and to the providers who do so. This works well for products. It works well enough for models, which are placed on markets and have providers with addresses.

The loop’s more consequential behaviour happens upstream of any placement. Internal tooling is not placed on a market. A training pipeline is not put into service. An evaluation harness has no provider, no CE mark and no customer, and a research system used inside a laboratory to design the next research system is not a product in any sense the framework recognises. The Regulation says so itself: it “does not apply to any research, testing or development activity regarding AI systems or AI models prior to their being placed on the market or put into service”. The layers where recursion actually runs are the ones with no obvious commercial transaction to hang a duty on, and the one place it runs hardest is written out of scope on purpose.

A threshold made of arithmetic

The AI Act does reach for a physical quantity. A general-purpose model is presumed to have high impact capabilities, and is therefore classified as carrying systemic risk, when the cumulative compute used for its training exceeds 10^25 floating point operations. That is a good surface: countable, expensive, hard to hide, and attached to a small number of organisations.

It is also the quantity the loop is busy reducing. Efficiency is what recursive engineering produces first and most reliably, in kernels, in architectures, in training procedures, in data selection. A boundary defined by training compute gets eroded by exactly the improvement it was drawn to watch, and the same capability arrives later under the line. The Commission can adjust the threshold by delegated act, which is the mechanism the Act provides, and which moves at the speed of committees while the erosion moves at the speed of publication.

The regulator’s instrument problem

For models above that line, the Act requires providers to “perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks”.

The provider performs it. That is not a drafting failure. Independent evaluation capacity does exist, and the UK’s AI Security Institute is the clearest case, running its own pre-deployment tests of frontier models and publishing the framework it uses. It is also small next to the parties it examines, and its access to unreleased models rests on agreements rather than powers. The Act does better on paper. It lets the AI Office evaluate a general-purpose model itself, demand access through interfaces and technical means including source code, and appoint independent experts to carry out the work. What no statute can legislate is the instrument. An evaluation run by the AI Office and one run by the provider reach for the same instruments, which carry the blind spots of the thing they are pointed at, and which are increasingly generated too.

A regime built on documented evaluation inherits every property of the evaluation. The law can require the document. It cannot require the document to be measuring anything.

The surface that can be counted

What remains reachable is physical. Fabrication plants are few, enormous and take years to build. Data centres are large and fixed. Grid connections are granted by named bodies that keep records. Electricity is metered by people whose job is metering it.

This is why the governance literature keeps returning to compute. The standard survey of the question sets out why, in four properties: relative to data and algorithms, compute is “detectable, excludable, and quantifiable, and is produced via an extremely concentrated supply chain”. An agreement about capability has nothing to inspect. An agreement about power stations has inspectors already. Countability is not agreement, and nobody has solved the geopolitics. But a treaty needs an object, and buildings make better objects than intentions.

Where a duty can attach

A duty attaches in three places now. The physical one, where things are large and located. The commercial one, where something is placed on a market and a provider takes on duties. The professional one, where a named person signs to say a system was checked.

Three places it does not reach. The internal research loop, which produces no product. The generated infrastructure, which nobody is looking at because everybody is looking at the number it prints. And the instruments of evaluation, which everything else depends on and which are being automated fastest of all.

A rule aimed at the model layer, while the capability is produced several layers upstream, arrives at the right building on the wrong floor.

The clerk’s brief

From the clerks, for the Patrician’s eyes

Compiled August 2026, with the AI Act’s general-purpose provisions in application since August 2025. Newest first; settled items sink into the assessment at the end. These entries concern where a rule can be made to bite, rather than what the rules say, which is filed at greater length elsewhere in the observatory.

2024: The systemic risk line, drawn in arithmetic

Article 51 of the AI Act provides that a general-purpose model is presumed to have high impact capabilities when the cumulative training compute exceeds 10^25 floating point operations, with the Commission able to move the figure by delegated act. The clerks admire the choice of a countable quantity over a contested adjective, and note the awkwardness that follows: the industry’s most dependable achievement is doing more with less, so the line moves towards the capability whether or not anybody amends it.

2024: The examination is set by the candidate

Article 55 requires providers of models with systemic risk to perform model evaluation to state-of-the-art protocols, including documented adversarial testing, and to report serious incidents. Article 92 gives the AI Office power to evaluate such a model itself and to demand access through interfaces and source code, which is a real power and a recent one. The clerks observe that independent testing capacity is nonetheless small next to the providers it would examine, that the AI Security Institute in London evaluates frontier models before release by arrangement rather than by right, and that a regime resting largely on self-administered examinations has a long history in other trades.

The reachable edges

Governance of these systems reaches the layers where something is bought, built or signed for, and stops at the layers where the loop actually runs. The clerks’ standing assessment is that the useful question for any proposed rule is not whether it is strict but which edge of the loop it can be enforced against, and that on present evidence the answer is the physical edge, the commercial edge, and little further in.