The Root Cause That Fits Every Incident
Open most incident reports to the last page and you find the same few words. Operator error. Failure to follow procedure. Complacency. They read like causes. They aren't, and there's a precise way to see why.
The physicist David Deutsch offers a test for telling a real explanation from a hollow one. A good explanation is hard to vary. Its parts are nailed so tightly to the thing it explains that you can't change them and still account for what happened.
His example is the seasons. For most of history people explained summer and winter with myth. The Greeks had Demeter, grieving for her stolen daughter half the year and freezing the world while she mourned. The problem isn't that the story is false. The problem is that it bends. Tell the Greeks that Australia has summer while Greece has winter and they would have said Demeter sends the cold to the far side of the earth when she grieves. The myth absorbs any fact you throw at it, because you can always adjust it.
The real explanation, the tilt of the earth's axis, cannot be adjusted. Change one piece of it and it stops producing the seasons we actually get. A good explanation has nowhere to hide.
"Operator error" is the Demeter explanation. You can fasten it to any incident ever recorded. Lift it off one release, set it on an unrelated one, and it reads just as well. It fits everything, which is exactly why it explains nothing.
None of this is news to the field. Sidney Dekker's new view of human error has been orthodoxy for years: human error is a symptom, not a cause, and an investigation that stops at the operator stopped too soon. True, and widely accepted. But it's a warning, not a tool. It tells you to look past the human. It doesn't tell you when you've reached the cause. Hard to vary is the test the warning never came with.
In 2010 a refrigeration plant in Theodore, Alabama released 32,000 pounds of ammonia. The hard-to-vary account: the control system let an operator restart refrigeration on a coil still in defrost, putting hot gas and cold liquid in the same pipe, and a single valve group commanded four coils, so the surge hit all four at once.
Try to slip "the operator was careless" into that sentence and it falls apart. Change one real piece of it, give the control logic a lockout that refuses the restart, and the release doesn't happen. The explanation is bolted to the machine.
That difference isn't academic, because of what a hard-to-vary explanation can do that the other can't. It reaches past the incident in front of you. "Operator error" protects no one, because there's nothing inside it to fix. The control-logic account protects every plant with the same gap, including the ones whose incident hasn't happened yet. Deutsch calls this reach. A good explanation explains more than the case you built it for.
The hollow explanation has the opposite habit. It doesn't reach outward to the next plant. It climbs. Look past the operator, as the field finally learned to, and the empty explanation doesn't disappear. It moves up a level and changes clothes. "Weak safety culture." "Production pressure." "Inadequate management of change." Point at the system, blame no one, pass the audit. Each of them is still pinnable to any incident on record, which means each of them fails the same test "operator error" did. The blame got more sophisticated. The explanation got no better.
It climbs one more rung, into entire schools of thought. Human and organizational performance (HOP), Safety-II, behavior-based safety, critical control management, safety culture, and the AI camp now forming. Each comes with a vocabulary that can re-describe any incident in its own terms. The test of a framework is not whether it can explain what you saw, because each of them can. The test is whether it forbids anything. An idea compatible with every outcome predicts none of them, and an investigation run in a language that fits all incidents carries the same defect as a root cause that fits all incidents, raised to the level of a worldview. A diagnosis that fits every patient cures none. We'd laugh it out of a clinic and sign it in an incident review.
The Theodore explanation was systemic too. Control logic, valve grouping, no operator to blame. It held anyway, because it was built from that one release and couldn't survive being moved to another. The level you investigate at was never the test. Hard to vary is.
What makes the hollow explanation dangerous is what happens after the file closes, which is nothing. A real cause tells you what to change. An explanation that fits everything tells you to change nothing in particular, so the same class of incident comes back wearing a new date. An organization running on explanations like that isn't learning from incidents. It's filing them and growing more confident with every folder, while the exposure underneath sits untouched and compounds. That is the quiet way a safety program turns fragile. Not by missing the hazard in front of it, but by explaining the hazard in language that hides the fact it was never addressed. Hydraulic shock still happens, more than a decade after the industry supposedly explained it.
The test costs nothing, and you can run it on the next report that crosses your desk. Take the cause it lands on, whether it names a person, a system or a school of thought, and try to set it on an unrelated incident. If it fits, you've found a label. Keep going until you reach something that could only be true of this one. The same test works on the fix. If the corrective action is more training, more management attention or a stronger safety culture, you haven't reached the cause, because a remedy that suits any incident came from an explanation that fits them all.
We taught the field to stop blaming the operator, which was right. Then we moved the blame onto the system and finally into the language of whole disciplines. Each time, we mistook the new vocabulary for understanding. The blame moved. The explanation never did.