The check that cannot be generated¶
The question everyone asks is whether a machine can design a machine. The question that decides the answer is whether anything left in the arrangement can say no and be right about it.
Three columns, and one of them emptying¶
Take any part of the pipeline and there are three things to record: what the system produces, what a machine checks, and what a person checks. Code, tests, benchmarks, research, architecture. Until recently the third column had an entry in every one.
In the systems now being assembled it has been emptying row by row, and each removal arrived with a decent argument attached. A person read the code, then generated tests began doing it, because reading every diff does not scale. A person wrote the tests against a requirement, then the tests were generated too, because writing them is tedious and the model is good at tedium. A person built the benchmark, then benchmark items were generated, because a static benchmark runs out. Peer review was people, and is now assisted by the systems being reviewed, because reviewers are scarce and unpaid. Architecture is the row where least has gone: machines have proposed designs for a decade, and the ones that ship are still, for the most part, chosen by people who can say why they chose them.
No step in that sequence is unreasonable on its own terms, and each was taken by somebody solving a real problem in front of them. What is left at the end is a pipeline that produces its own evidence about itself.
The question to ask a loop¶
What observation cannot be generated by the system being evaluated?
Some answers hold up well. Execution is one: the test passes or it does not, the proof checks or it does not, and no amount of confident prose alters the exit status. Physical measurement is another: the part holds at the clock speed or fails, and a fabrication plant is unimpressed by argument. Time is a third, because a prediction about next quarter cannot be produced in advance by the thing making it. Time arrives late and says nothing about why, which does not stop it being an answer no system can manufacture in advance. Money is a fourth and a cruder one: somebody paid, or nobody did.
Then there is a person’s judgement, which is the only answer that covers taste, novelty, relevance and whether the question was worth asking. It is also the expensive one, which is why it is the entry disappearing from the third column, and why the cost of keeping it gets pushed somewhere less visible rather than paid.
Cheap and correct, against cheap and gameable¶
A check that is cheap and correct is the best thing in engineering, and the reason the compiler loop and the chip loop compounded for decades while alarming almost nobody, Thompson excepted.
A check that is cheap and gameable is worse than no check, because it produces confidence. Benchmarks drift into this category by ageing: they saturate, the answers circulate, items reach the training set, and a score that once measured a capability comes to measure exposure to the test. The measurement does not announce the moment it stops working. It keeps returning numbers, and the numbers keep going up.
The substitution now in wide use is the model as judge, adopted for the ordinary reason that human evaluation cannot keep pace. It works better than it has any right to: agreement with human preference in the region of eighty per cent, which in the same study is about what the human raters managed with each other. The same work that established the agreement also named the biases, including a preference for longer answers, sensitivity to which answer came first, and a tendency to favour the judge’s own output.
That last bias is the one that does the damage. A judge inclined towards its own work is not lying and has nothing to hide. It is wrong in a direction that nothing else in the loop is positioned to notice, because everything else in the loop shares the inclination.
Three years after the practice was established, a group working from measurement theory put the position plainly: adoption of these judges “has outpaced rigorous scrutiny of their reliability and validity as evaluators”.
Coherence is not grounding¶
Deception gets the attention and is the easier problem, because deception implies something that knows better. The failure in a recursive loop is quieter than that. A system evaluated by instruments that share its assumptions becomes more internally consistent over time and no better attached to anything outside itself. Every reading is green. Nothing inside reports a fault, because a fault of this kind has no inward-facing symptom.
Institutions do this without any machinery at all. A firm that reads only its own dashboards, a literature that cites its way to consensus, a service that collects the material confirming what it already believed. What changes with a generative system in the loop is that the cost of producing agreement falls much faster than the cost of producing disagreement.
When the answer is a person¶
Any of these systems can be examined with one question, and the answer is short: which observation can this loop not produce for itself? For a compiler, a program whose correct output was established without it. For a chip, the measurement of a fabricated part. For a coding agent, a failing test written against a requirement invented outside the loop. For a model trained on generated text, the answer is the awkward one: the real data that was kept is a reference rather than a verdict, and a reference can be thin, stale or contaminated in its own right without ever announcing it.
Where the answer is a person, there is a second question, and it is not a technical one. It is what that person is paid, how long they are given, and whether anybody has arranged for them to be able to say no. The first two are budget. The third is not, and no wage on its own produces it.
These are different kinds of outside, and they are easily mistaken for each other. A fabricated part is causally outside: the system cannot manufacture the failed measurement. An external evaluator can be institutionally outside the provider while sitting inside the same epistemic system, working from its benchmarks, its frameworks and access it was granted rather than holds. A person at a checkpoint can be outside only by personnel, able to see the output and not to alter the framing, reject the result or stop the deployment. All three look identical on an architecture diagram. They are not equivalent checks.
The clerk’s brief¶
From the clerks, for the Patrician’s eyes
Compiled August 2026. Newest first; settled items sink into the assessment at the end. These entries concern the instruments rather than the systems, on the principle that a report is only as good as the scales it was weighed on.
April 2026: The instrument reaches the clinic before the checks do¶
A scoping review of LLM-as-a-judge in healthcare screened 11,727 studies and included 49. Risk of bias testing was absent in 36 of them, 73.5 per cent, and exactly one study, 2 per cent, had reached production. The clerks note the ordering. The instrument is being taken up in a domain where a wrong answer has a patient attached to it, and the work of establishing whether the instrument is any good is mostly not being done.
August 2025: Adoption outpaced scrutiny¶
Chehbouni and colleagues applied measurement theory from the social sciences to four assumptions behind LLMs as judges and reported that “their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators”. The clerks record that this is the field auditing its own instrument, and that the finding has the shape of everything else in the file: the measurement was adopted because it was cheap and available, and the question of whether it measures anything arrived afterwards.
2024: The instruments were called into question by their own field¶
Saxon and colleagues argued in a call for model metrology that “static benchmarks inevitably saturate without providing confidence in the deployment tolerances of LM-based systems”, and that broad claims about reasoning and understanding rest on flawed metrics. The clerks note that this objection came from inside the discipline rather than from its critics, that the proposed remedy is dynamic assessment generated for the occasion, and that generating the assessment is precisely the step that keeps turning up on the wrong side of the loop.
2023: The judge became a model¶
Zheng and colleagues established the practice now in wide use, reporting in Judging LLM-as-a-Judge over 80 per cent agreement between model judges and human preferences, matching the agreement humans reach with each other. The same paper named position bias, verbosity bias and self-enhancement bias, the last being a judge’s inflated assessment of its own output. The clerks record that the agreement figure is the one cited when the substitution is defended, that it is a good figure, and that the biases were published in the same document and have been considerably less quoted.
Made in the same works¶
Nothing in this file describes a machine deceiving anybody. It describes measuring instruments being replaced, one at a time and for good local reasons, by instruments manufactured in the same works as the thing being measured. The clerks’ standing assessment is that the useful audit question for any of these arrangements is not how capable the system is, nor how carefully it was tested, but which of the tests could have come back negative, and what would have happened next if one had.