The words somebody chose

A dispute produces a few million documents. Nobody reads a few million documents. Somebody writes a query, the query returns a subset, and a team of reviewers works through the subset deciding which parts of it the other side is entitled to see.

The reviewers are careful. Their judgement is real, and it is exercised on everything in front of them. What is in front of them was settled earlier, by whoever chose the words.

Complete and proportional at the same time

The rules ask for two things at once. In the United States, Rule 37(a)(4) treats an evasive or incomplete response as a failure to respond at all, while Rule 26(b)(2)(C)(iii) requires a court to limit discovery where “the burden or expense of the proposed discovery outweighs its likely benefit”. Grossman and Cormack, writing in 2011, put the position plainly: the rules together reflect a tension between completeness on one hand, and burden and cost on the other, one that exists in every electronic discovery exercise.

Neither rule says how to resolve it, because neither can. The tension is resolved anyway, several times a day, by people deciding which terms to run. A query is where a procedural contradiction gets settled without anybody recording that a decision was taken.

What a reviewer is given

In a technology-assisted process, a human examines and codes a small part of the collection, and the computer extends that coding to the rest. Grossman and Cormack describe the human’s share as “a tiny fraction of the entire collection”.

So there are two selections here, not one. The first decides what a person sees. The second takes the person’s judgement about that fraction and applies it to everything they did not see. The judgement is genuine at the point it is made and it travels considerably further than the person who made it.

The uncomfortable finding

The obvious moral would be that reading everything is better. The evidence does not support it.

Grossman and Cormack analysed the 2009 TREC Legal Track, run by the National Institute of Standards and Technology, and found that two technology-assisted processes achieved results exceeding those the official assessors would have reached had they conducted a manual review of the entire collection. The assessors were law students and lawyers employed by professional document-review companies. Reading everything, carefully, by people paid to do it, came second.

Four years earlier the Sedona Conference had put the assumption under examination, warning of “a myth that manual review by humans of large amounts of information is as accurate and complete as possible” and that this was treated as the standard against which searches were measured.

That result cuts against the easy version of the argument, and it belongs here for that reason. Selection is not the enemy of accuracy. It can be the source of it.

Which makes the query more interesting, not less

If choosing what a person looks at can beat looking at everything, then the choosing is not an administrative preliminary. It is the part of the process that determines the outcome, and it is the part that leaves the lightest record.

An exhaustive review can be described afterwards: this many documents, these reviewers, this many hours. A query can be described too, and usually is, as a string. What cannot be described from the string is what it did not return, because the documents it did not return were never collected into anything, never counted, and never assigned to a reviewer who might have noticed their absence.

The clerk’s brief

From the clerks, for the Patrician’s eyes

Compiled August 2026. Newest first; settled items pass into the note at the foot. These entries concern a selection step that no regulator has named, examined instead by a standards institute, a conference of lawyers and a working group of judges.

May 2026: The judges look at the machinery

On 8 May 2026 the Courts and Tribunals Judiciary reported that its Disclosure Review Working Group was considering simplification of Practice Direction 57AD, which governs disclosure in the Business and Property Courts, and had been convened “to examine the operation of PD 57AD … and, in that context, the use of Technology Assisted Review (TAR) and Artificial Intelligence (AI)”. Its survey asked whether “developments in technology (including the use of AI) are impacting the process of disclosure”. The clerks note that the question is put as one about impact on a process, and that the technology in question is the process.

2011: Reading everything came second

Grossman and Cormack published an analysis of the TREC 2009 Legal Track Interactive Task in the Richmond Journal of Law and Technology, demonstrating that the performance of two technology-assisted processes exceeded what the official TREC assessors would have achieved by manually reviewing the entire document collection. Both authors were coordinators of the track. The clerks record the finding without qualification, and note that it makes the unexamined step the important one rather than the harmless one.

2009: A gold standard, built on purpose

The TREC Legal Track, sponsored by the National Institute of Standards and Technology, ran an Interactive Task whose results were published as NIST Special Publication SP 500-278. Its method was to construct a gold standard for a document collection and measure retrieval methods against it. The clerks observe that this is the only entry in the file where somebody established what the correct answer was before asking how well anybody found it, and that it required a national standards body to do so.

2007: The myth named

The Sedona Conference’s commentary on search and information retrieval methods warned of “a myth that manual review by humans of large amounts of information is as accurate and complete as possible”, treated as the standard by which searches were measured. The clerks note the date: the assumption was identified as an assumption two years before anybody tested it, and fifteen years before a working group of judges began asking what the technology was doing to the process.

What the file cannot contain

Every entry here concerns methods of finding, and every one is assessed on what it returned. A search is measured against a gold standard where one exists, and where one does not it is measured against the confidence of the person who wrote it. The clerks’ standing assessment is that the record of a disclosure exercise is a record of what was reviewed, that the reviewing is the visible and expensive part, and that the words somebody chose beforehand are recorded as a string and assessed as a formality.