I’m old enough to remember watching Space Shuttle Challenger take off on January 28th 1986. Tragically it broke apart shortly after takeoff with all seven crew losing their lives.
As a team we were recently reflecting on explicitly positioning a Structured Recovery offering to address ERP and WMS implementation projects that are going off-course or failing to perform to expectation. In many of the instances where we’ve been involved in over the years the same idea comes up repeatedly.
Premise: An isolated technical fault or limitation that is either not understood, or in as in the case of the Space Shuttle, simply not communicated through middle-management to the executive level causes a whole project to either underperform or fail catastrophically.
Fast-forward to 2009. It’s my first consulting job and it’s my first week. I’m wearing a suit from the South Yarra end of Chapel Street and a nice tie. This isn’t a project engagement, this is a customer whose warehouse management system is taking two minutes per order to confirm shipments with the ERP system and they’re not a small operation so it’s a big problem. The usual rumblings are happening “the system’s not built for this demand, the partner doesn’t know what they’re doing” - we all know those conversations from both sides of the fence.
Long story short. At the time I walked in, I didn’t know that SQL Server had something called “statistics” let alone that for one reason or another those statistics can be wrong, and when they’re wrong the database can take a very long time to do something that ought to take it milliseconds. Forty-eight hours later, after a crash course in SQL Server internals and the relative merits of “FULLSCAN”, the two-minute transaction was back to taking milliseconds. There was nothing fundamentally wrong with the system after all.
Much more recently, a large multi-site operation called us when they were at their wits’ end with the partner who had implemented and “supported” the system for five years.
The biggest single problem? It took a full minute after scanning a product barcode to return a list of valid batches on the handheld.
A long way from the lone green consultant in the sharp suit, we were now somewhat greying guys in jeans and hi-vis vests. We were simultaneously delighted and depressed to discover that, after taking a backup and some other precautions, we could remove a massive amount of redundant data and cut the response time from a minute to one second.
That wasted minute, repeated across every stock movement, had been costing them hundreds of hours a year.
It’s not all down at the database level. The CEO of one of our long-term clients is a motoring enthusiast. We needed to explain the apparent contradiction between the long-standing advice to restart the application every night and our advice to leave it running and manage it properly.
A memory leak is like a fuel leak in your car. Restarting the application is like telling someone with a leaking fuel tank to visit the service station every morning, regardless of how far they’ve driven. It replaces what has leaked out, but does nothing to fix the leak. Worse, every restart carries a non-zero chance that the application won’t come back up cleanly.
The accepted workaround was hiding one fault while introducing another opportunity for failure.
So back to Challenger.
It came to mind because the shuttle was lost to an isolated technical problem with the O-ring seals. The engineers knew about it, but the risk was not carried through management with enough force or clarity to stop the launch.
President Reagan established the Rogers Commission to investigate the disaster. I knew of it because Richard Feynman, something of a household name in my family growing up, wrote about his involvement in the enquiry. There is a famous clip of him demonstrating the behaviour of the O-ring in ice water.
Perhaps because of his fame, Feynman is sometimes credited with discovering the O-ring problem. The engineers had known about it long before. His contribution was to make the problem clear, public and impossible to explain away.
Feynman approached the Commission with characteristic irreverence—this was, after all, a man who had taken up safe-cracking as a hobby during his work on the Manhattan Project. His written contribution was a scathing critique of weak safety standards and the way good engineering practice had been defeated by poor corporate culture.
His closing remark applies to consulting engagements and systems work as well as it does, one assumes, to aerospace:
“For a successful technology, reality must take precedence over public relations, for nature cannot be fooled.”
My take? When your project or your system is in trouble, you don’t need more of the same. You need people who can get close to the problem, learn from the people on the ground, get across technology they may not know and communicate the technical reality without fear.
The US government did not abandon NASA or its contractors, and the findings were not buried. The booster joints were redesigned, safety and management arrangements were changed, and the Space Shuttle returned to flight 32 months later.
That is Structured Recovery. Not pretending the failure is small. Not declaring the whole system unfit for purpose. Finding the fault, understanding how it was allowed to persist and fixing both.
— Jonathan Embrey