DERIVEREFRIGERATIONJournal
2026-06-04

Compliant and Fragile

Every five years, a tech climbs to the top of a vessel, pulls a relief valve that has never lifted or made a sound, and drops it in a bucket. There may be nothing wrong with it. But a valve that has never lifted has proven nothing about whether it will, and we don't let its silence fool us into trusting it.

Everywhere else, we do.

A clean audit file, a current binder, an RMP five-year accident history with nothing in it — every signal that a plant is safe is a record of something that hasn't gone wrong yet, taken as proof that it never will.

A compliance audit checks whether your paperwork matches the standard. That five-year history is narrower still: it logs only releases that killed or injured someone, forced an evacuation, or did real property damage. A contained spill, a release that hurt no one. None of it has to appear. Both look backward, and an empty record doesn't mean the system ran clean. It means nothing crossed the reporting threshold. The record measures catastrophe, not operation.

Even a true record of quiet would miss the thing you need to know. Most days at an ammonia plant are the same day: it runs, the logs fill in, nobody gets hurt. But that pattern says nothing about the day that breaks it — the release that hospitalizes a crew, the rupture that empties a neighborhood. That's a fat tail: the risk lives in the extreme event that almost never happens and is catastrophic when it does. Years without an incident tell you nothing about the size of the one you haven't had yet.

Nassim Taleb studies exactly this. Systems that look stable and well-behaved right up until the one event they can't absorb. His word for them is *fragile*. Not failing, just quietly accumulating risk that doesn't show up in any performance metric until the day it shows up all at once. The danger isn't in the data you have. It's in the data you don't.

He draws a hierarchy that's useful here. A fragile system breaks under stress. A robust system survives it. An antifragile system gets stronger from it. It learns from small shocks so it's better prepared for large ones. Most ammonia facilities are aiming for compliance. Compliance, in Taleb's framework, is fragile. It's a system optimized to look orderly, not to absorb disorder.

The question is what keeps it there. The answer is mostly structural.

The records stay clean for a simpler reason than anyone wants to admit: the work behind them has become the work of signing them. Every year the procedures are certified current and accurate. Most years nobody walks the line to check, because nothing went wrong, so the signature lands and the procedure drifts from the plant it describes. The five-year PHA revalidation becomes a re-stamp of the last one. Nothing went wrong, so there's no reason to re-examine the nodes or revisit the risk rankings. The assumptions carry forward unchallenged, whether or not the plant still matches them.

That's what a recurring task becomes when what gets measured is whether it happened, not whether it worked. A hollow program produces paper identical to a living one. Same certifications, same dates, same signatures. Only the bad day tells them apart.

The drift compounds it. Change accumulates without getting written down: a setpoint nudged to kill a nuisance trip, a red-line on a P&ID after someone caught a mislabeled valve, a temporary bypass that quietly became permanent. Each seems small and sensible in the moment, none feels like a management-of-change event, and after enough of them the binder describes a plant that no longer exists — robust to the auditor, brittle to the failure mode that lives in the gap between the drawing and the pipe.

And slack looks like waste. The redundant pump, the maintenance window you don't strictly need this quarter, the second operator on a shift that runs fine with one. On a budget review, each is a line to question. Trim them and nothing happens. Nothing happens for years. Then the buffers that were there to absorb an abnormal day are gone, and you find out on the abnormal day.

This is a system that is fragile by design — not because anyone decided to make it fragile, but because everything that would make it robust gets optimized away. The signatures replace the walkthroughs. The drift goes unrecorded. The margins get cut. And at every step, the paperwork still looks fine.

Taleb's answer to fragility isn't more documentation. It's learning from small shocks instead of hiding them.

The cheapest lesson a plant gets is the release that hurt no one. A near-miss is a free data point about a failure mode that almost happened, exactly the kind of information a fragile system needs and never captures. Treat it as a confession, treat the investigation as a hunt for who to fire, and people report the minimum and bury the rest. A blame culture doesn't just stay fragile. It actively destroys the one feedback mechanism that doesn't require a catastrophe to function.

A plant that treats near-misses as intelligence, that investigates without hunting for a throat to choke, converts disorder into knowledge. That's antifragility: getting stronger from the shocks instead of suppressing them. It's also the thing compliance culture is least equipped to produce, because reporting a near-miss makes the audit trail messier, and a clean audit trail is the whole point of the exercise.

Robustness itself is mostly subtraction. You take fragility out rather than documenting over it. Less ammonia in the system, not because it drops you under a regulatory threshold, but because a smaller inventory is a smaller tail. The hazard doesn't change at 9,999 pounds. It changes when there's less ammonia to release. Redundancy and maintenance margins defended as the cost of surviving a bad day, not waste to trim when the budget gets tight.

The most fragile asset in any facility isn't the equipment. It's the thing the twenty-five-year operator carries in his head — which valve sticks, which reading to trust, what the plant sounds like before it misbehaves. That knowledge is undocumented, untransferable, and invisible to any safety management system. It walks out at retirement, and no audit ever recorded it was there. Writing it down while the person is still here to write it isn't a compliance task. It's the difference between a plant that knows itself and one that has to relearn its own failure modes from scratch.

The work isn't making the paperwork match the standard. It's making the documentation match the plant — the real one, the one that drifted — so that what the plant knows outlives the people who know it. Do that and the audit takes care of itself. Aim at the audit and you get paper that passes and a plant that doesn't.

We already know how to think this way. Every five years we throw away a working part because we won't let "it never failed" mean "it works." We've quarantined that one piece of clear thinking to a single valve. The silence everywhere else is making the same promise, and we keep believing it.

We write about what we see and what were building. No fluff.

New articles delivered to your inbox.