The twin pillars of ML disappointment¶
In the great and creaking city of Ankh-Morpork, success depends on luck, a tolerance for interesting smells, and the ability to keep a straight face while the whole system declares that everything is going exactly to plan. Our model has none of these. It has two pillars instead, Data and Requirements, both as reliable as a guild oath sworn after midnight, each holding the thing up with the enthusiasm of a wet cardboard box.
Between them they ensure the model will be flawed if it is lucky, scandalous if it is not, and profoundly misunderstood by the people who commissioned it either way.
Garbage in, gospel out¶
In theory, data is the fuel that powers the clever contraptions. In reality it behaves more like the damp kindling the Night Watch uses when attempting to start the annual solstice bonfire. Everyone gathers round, insists it will work, swears they have done this before, and then watches the thing smoulder resentfully while producing nothing but smoke and disappointment.
The eternal whinge¶
Data scientists have a single truly reliable setting. It is the one where they explain, with the desperation of a man who has seen the edge of the world and found it insufficiently documented, that there is not enough data. Ten million samples? A mere whiff. More data is always needed. More data is the answer to every question. More data is the oxygen of their existence, and there is always someone, apparently, standing on the hose.
There are, to be fair, moments of genuine famine. Training a model on three JPEGs and an optimistic shrug is unlikely to succeed unless the target domain is children’s fridge art. More often the problem is not a lack of data but a lack of useful data.
A warehouse full of inputs can still amount to a drought. Does the dataset represent humanity in all its chaotic, contradictory glory, or does it look more like a damp pile of Tesco receipts and mislabelled cat photos? Ten thousand shots of the same blurry stapler teach a model very little beyond the unfortunate truth about the procurement department. A billion rows of missing values and timestamps set to the Unix epoch teach it only despair.
Quantity counts, but not as much as the people chanting for more seem to believe, especially when the last time they inspected a histogram was sometime before The Patrician outlawed unlicensed fortune telling.
Data quality, where hope goes to die¶
Data quality is the noble, shimmering ideal that everyone mentions and no one recognises. Ask for a definition and the answers range from “not horrible” to “the thing that did not crash the pipeline last week”.
It fails in three reliably different ways.
Consistency¶
Consistency is the hope that data will behave sensibly, like a decently trained dog. What usually turns up is a cat that has learned the art of irony. One column says “United Kingdom”. Another says “UK”. A third says “U.K.”. A fourth offers “Europe, sort of”, as though the field itself were having an existential crisis. Attempt a join across that lot and watch a soul leave the building.
Correctness¶
Correct data is meant to reflect reality. Instead it often reflects an intern’s dreams during the final hour of a shift. “Customer Age: 217” turns up, and the choice is between an immortal customer and a mug landing on a keyboard. Meanwhile the fraud model is busy awarding a youth savings account to a lich.
Completeness¶
Completeness is the cheerful fiction that a dataset contains information rather than an art installation composed mainly of absence. Eighty per cent missing values is fine, says someone. Just impute the mean. Now everyone has exactly average income, average height, and the same number of cats. Useful for modelling Lego figures. Less so for people.
Cleaning data is tedious, thankless, and absolutely necessary. Much like sweeping the streets of Ankh-Morpork: nobody praises the work, and skipping it makes everything unpleasant in ways that accumulate quietly and smell alarming.
The art of lowering expectations¶
If gathering requirements feels like group therapy held in a broom cupboard, that is because it is. Stakeholders describe an AI system that is ethical, transparent, unbiased, fast, accurate, and cheap. Two of those are available. Two is generous.
Requirements documents begin life filled with hope, then slowly become graveyards of abandoned ambition. There is always a line that reads “must handle edge cases”. It will not. There is always “must be explainable”. It will not be, unless the explanation involves three wizards and a chalk circle. There is usually something about avoiding bias. The details follow, along with a stiff drink.
Once in a while someone produces a 73-page PDF. Nobody reads it until six months after deployment, when the system has done something outrageous and everyone is trying to work out who approved it.
Success lies not in fulfilling the vision but in managing expectations early enough that the outrage arrives in smaller, more manageable bursts.
Neither pillar keeps to itself. What arrives downstream is a model that discriminates without ever being told to and cannot account for itself when asked, in a world that has since written laws about both.
Ethics and bias: the minefield¶
Every team insists it does not discriminate. Meanwhile the model is redlining half the city with the efficiency of a Guild accountant on deadline. Bias is not hiding in the dataset. It is the dataset. It arrives baked in, like raisins in a fruitcake nobody wanted but everybody receives, year after year.
Common practice is to remove the obviously sensitive columns and consider the matter closed. Race, gender, the usual suspects. Never mind that fields such as postcode and forename are quietly carrying the same information with the determination of a mule that has not been fed in three days.
A system that subjects “Jamal from East Ham” to a lengthy security review while fast-tracking “Sebastian from Surrey” into preferential customer support has something amiss with it, and it is not just the ambient fog.
The traditional remedy is a fairness audit, which in many organisations means skimming a Medium article titled “Ethical AI for Busy People” shortly before the quarterly meeting. Bias is a mirror, and few care for the reflection.
Explainability: the corporate fig leaf¶
Explainable AI is the respectable cloak a model wears to meet regulators. It looks sincere, it sounds thoughtful, and it conceals the fact that nobody truly knows what the thing is doing when left to its own devices.
Feature 147 and its deviation from baseline can be discussed at length. Users nod politely, then go looking for legal advice. SHAP values and LIME plots are genuinely clever, but they answer the wrong question. The applicant refused a mortgage wants to know what to do differently next time, and receives a ranked list of contributions from variables they have never heard of.
In meetings, executives pretend to understand the diagrams, engineers pretend to enjoy producing them, and regulators pretend to find them reassuring. It is all terribly civil and almost completely pointless.
An explanation requiring a whiteboard, three hours, and a sworn oath from the High Priest of Quantitative Mysticism is not an explanation. It is performance art.
Regulations: the compliance theatre¶
Onto the stage shuffle the great chorus of acronyms. GDPR, HIPAA, the AI Act, and the rest. Their purpose is to ensure that when something important gets broken by machine learning, it is at least broken in a documented fashion.
There will be a debate about whether log files count as personal data. They do. There will be an argument about whether storing them in the cloud is legal. It depends. There will be the consoling thought that nobody audits these things until a newspaper becomes involved.
Compliance is not about being ethical. It is about being demonstrably correct on paper. It resembles keeping a diary, not for posterity, but in case the Thieves’ Guild asks questions.
The golden rule is simple. Document everything. Not because it improves the system, but because the regulator arrives eventually, and a sufficiently thick paper trail will soak up a surprising quantity of tears.
The EU AI Act entered into force on 1 August 2024, which means it is no longer a threat on the horizon but a fact of life, like damp and unreliable contractors. It operates on a risk-based tiered system, the logic being that not all machine learning is equally capable of ruining someone’s afternoon.
At the top sit the prohibited practices: social scoring, whether the scoring is done by a government or by a company, real-time remote biometric identification in publicly accessible spaces, where the ban falls on law enforcement uses and carries narrow exceptions that lawyers are already widening, and systems designed to manipulate people through subliminal techniques. These are banned outright, which has not stopped several organisations from discovering that what they built falls uncomfortably close to the line.
Below that are high-risk systems, which arrive by two routes. One is a list of standalone uses: credit scoring, recruitment, law enforcement, education, critical infrastructure. The other covers AI built into products already regulated for safety, medical devices among them. The distinction is dry and it decides which deadline applies. A model deciding who gets a loan, who gets an interview, or who gets flagged at the border is in this category whether anyone planned it that way or not. The obligations run to transparency, human oversight, data governance, accuracy and robustness, documented to a standard that satisfies auditors who were not present when any of the decisions were made and will not be impressed by a notebook.
Fines for prohibited practices reach 35 million euros or seven per cent of global annual turnover, whichever is greater, with the lower of the two applying to smaller companies. For most other violations the ceiling is fifteen million euros or three per cent. The European AI Office oversees the general-purpose model rules, while enforcement of the rest falls to national authorities, staffed throughout by people who have read the regulation rather more carefully than the average practitioner.
The compliance timeline has been phased, and then re-phased. Prohibitions applied from February 2025 and the obligations on general-purpose models from August 2025. The high-risk deadlines proved less durable: the Digital Omnibus, agreed in May 2026, adopted in June and signed in July, moves standalone high-risk systems back to December 2027 and those embedded in regulated products to August 2028. It is still waiting on publication in the Official Journal, so at the time of writing none of it is law. This has produced the familiar spectacle of organisations simultaneously insisting they are fully compliant and quietly reclassifying their systems to avoid the high-risk category. Consulting firms have never been busier. The paperwork has never been thicker.
Between bad data and impossible promises, the model is already fragile before anyone lays a finger on it. Which is unfortunate, because a good many people are about to.