Useful Failure – Reading IT News Critically Part 2

Introduction
Welcome to the 2nd Part of our three-part discussion of how to read IT News critically. For those only joining us now: in Part 1 we discussed various forms of illusionary success that make for great reads but won’t transfer to fix your projects’ problems. We specifically mentioned selection bias, regression to the mean – Coach’s Fallacy, anyone? – and the insufficiently sized sampling window.
In Part I, I only briefly alluded to the famous survivorship bias, and I promised a fuller discussion of that issue. The reason I wanted to take the extra time is because there is a broader picture of focussing on success that is limiting what we could learn. There is a psychological preference for success that is worth unpacking. This is the point of the post here.
Survivorship Bias
The Wikipedia article on Survivorship Bias comes with an excellent illustration. It is excellent because it is both evocative and manipulative. Looking at that picture is like playing out survivorship bias in real time on yourself. The drawing is by Martin Grandjean (vector), McGeddon (picture), US Air Force (hit plot concept) - Own work, CC BY-SA 4.0, File:Survivorship-bias.svg
The context is World War 2, specifically determining which parts of an aircraft needed to be strengthened to increase the safety of the bomber crews. Since strengthening has a weight penalty, this is a tough problem.
The use of the color red – which also looks a bit like blood, though airplanes do not bleed – immediately draws your attention to the areas where the plane was hit by anti-aircraft fire. The picture collects all known areas of damage. As an engineer, seeing this picture, my first instinct is, “We gotta strengthen these parts!”
The idea of survivorship bias is of course that planes that were hit and made it back to be analyzed give misleading information about vulnerabilities. What this picture wants to illustrate is that planes that were hit in the white parts did not make it back – and are therefore absent from the sample. Even though the red screams attention, it is the undotted parts that need the most reinforcement. That’s survivorship bias in a nutshell.
Abraham Wald, the statistician on whose 1943 research these patterns are based, knew that. Wald was part of the Statistical Research Group at Columbia University during World War 2. He was not the first to notice survivorship bias. His contribution was developing statistical estimators that could counteract the bias in the sample. He wanted to give those in charge a method to compute where to expend their precious reinforcement weight to increase survival. Here is a screenshot of the 1980s reprint of Wald’s report A Method of Estimating Plane Vulnerability based on Damage of Survivors (which was classified during the War years), p.65.
In short, Wald was reconstructing the failures’ distribution, because the failures were the most informative.
The Preference for Success
There are two common justifications why people focus on success rather than failure.
Pragmatically, we are engineers and we need to get stuff done. While a failure is a useful warning of what to stay away from, it does not give much guidance on what to do instead. Hence we consider it best practice to keep lists of best practices – to support our doing. However, this overlooks the lesson Wald was teaching us since 1943 – even a reconstructed failure distribution will tell you clearly what needs to be done next!
The other aspect is that failure can stem from a whole slew of causes. Or, to put it in the terms of Leo Tolstoy’s famous opening line from Anna Karenina: “All happy families are alike; each unhappy family is unhappy in its own way.”
I have been arguing since the previous blog post on critical reading of IT news that this is too simple. Success is just as complicated as failure and has just as many causes. Some causes for success are so specific that we cannot let the success inspire our best practice – think of the dream team selection bias. Because we are interested in success that can guide our own actions, we have to cross the success-failure distinction with the axis of informative vs uninformative.
Informative success is what drives best practices and invites emulation. But success that we cannot replicate, because the evidence is not actually clear enough or the preconditions do not match, is uninformative. That was the argument from Part I.
Admittedly, failure is often just as uninformative as irreplicable success is. But there is a category of informative failure that is worth chasing down.
Anti-Patterns and Post-Mortems
As an industry, software engineering has become better at organizing failure experience. The Pattern movement in the 1980s gave rise to the inverse movement of collecting anti-patterns as well. So that’s good, right? Is that not what I am after? Oddly, not really. Anti-patterns are very pedagogical, and it makes sense to learn the patterns and the anti-patterns together. They can illuminate each other.
But the social result of that pedagogical usage makes applying an anti-pattern look like a beginner’s move. It’s the kind of thing the wise experienced engineer points out to the summer intern. It is hard enough to get people to talk about failures; try getting them to admit to having used an anti-pattern! Plus, we already know that one, right? We even named it!
What about post-mortems? Are those not precisely the type of informative failure reporting I am arguing for? Funny that you should ask. There clearly are post-mortems published with the intent to educate the discipline about lessons to learn. Fred Brooks’s classical The Mythical Man Month absolutely belongs in the category; that is why it is still studied in software engineering classes, though the events it describes are over fifty years in the past.
But the kind of post mortems that get issued when a key software provider almost brings the world to a halt have a purposefully uninformative shape. Many post mortems are exercises in damage control. Their whole point is to reassure not just the industry, but even the wider public, that this was a one-in-a-million edge case. To paraphrase Tolstoy, the post mortem intends to convince us that this problem was unhappy in its very, very own way.
And if we are honest about it, useful research like the analysis of the recent ChainDrop supply chain attack is an IT success story of the security lab that publishes it. It therefore participates in all the complexities of success stories that we have been mulling over for two posts.
Across the Fence
As a community, we have come to value findings that are novel and eye-catching over those that are true, incentivising a host of questionable practices that twist the evidence to suit the narrative. –Chris Chambers, 2013
Fortunately, the IT industry is not the only endeavor grappling with the absence of documented informative failure. So a peek across into the other gardens may well suggest some ideas to draw inspiration from.
That is perhaps not surprising – success as the criterion for social advancement is a common occurrence across science and technology. With the reward structures skewed toward “publish or perish” in academia, the mere appearance of success will make it into the literature. In a much-discussed essay Why Most Published Research Findings Are False from 2005, John Ioannidis sounded the alarm on exactly this trend in medicine.
In the related field of psychology, Theodore Sterling introduced his 1959 empirical survey article (cf JASA 54(285):30) with an observation that readily extends into a thought-experiment: Assume an appealing research question whose true answer is null. Any psychologist who performs the experiment is then likely to not publish it. Consequently, neither a positive finding nor any indication of the failed trials marks the spot in the research landscape. Periodically, another scientist will find that time sink and try again. Eventually – due to error bars, small sample size or similar chance – the experiment wanders into spurious significance – a Type I error. Only now will it finally get published and become common mythology – at a possibly staggering expense of wasted efforts to the psychology research community.
The situation had hardly improved some 50-plus years later, when Chris Chambers in 2013 concluded that “the simple fact that editorial decisions are often based on the results“ was at the root of all this evil. It was the motivation for evidence twisting, p-hacking, and selective reporting. Chambers scoffed at the absurd training experience of incoming researchers: As undergraduates they had to state their theories first. As graduates and post-graduates, when publishing suddenly mattered, the research practice turned out to be recovering theories from collected facts.
For the scientific journal Cortex, where Chambers is on the editorial team, they instituted a pre-registration process for the experimental design called Registered Reports. Chambers has argued that experiments should be accepted based on their design, not their outcomes. As long as the study proposals are “judged to be methodologically valid, detailed, replicable, and which address an important scientific question” they deserve to be published, regardless of outcome.
Because: If a quality design fails to show an effect, that is – in our diction – informative failure.
Transfer Learning
Like many software companies, Posedio has a rigorous methodology that tracks both the good and the bad. Each iteration in the Deming-Shewhart Cycle of Plan, Do, Check, Act (PDCA) identifies what went well, what did not go well, and what could be improved. Design is done up front, and then checked against the implementation, to improve both the design process and the implementation considerations.
Sadly, the practices of the software industry do not amount to research registration in the style of the Registered Reports of the journal Cortex. Though all the pieces seem to be there – those learnings that PDCA identifies remain internal. The introspection does not translate into cross-industry sharable know-how. No registry or repository administers the informative failures. Those spots on the industry’s architectural map remain white. Apparently unexplored.
The reason is often legal, not only cultural. We neither own the computational context – network, stack, hardware, configurations – nor the data that is running on it. Our task is to make things work inside the setups that our clients have constructed. That setup is part of their IP, under their control, expressing their vision. Any software system design is underspecified without that construction.
But that valid concern about IP and competitive advantage leaves us open to the trap Sterling described above: Individually lured by a seemingly promising architecture, we try and fail collectively, over and over, because no one can speak about the pain.
Outlook
Every project manager is looking for that one software component, data management system or algorithm that will remove the stumbling block and unlock the potential of their project. We dig through the stories that speak of success and recommendation to find that missing link.
In the first blog I argued that success is more complicated than that. Some stories are uninformative – they won’t transfer, they won’t scale, they didn’t show that they meant to show. Biases. Mirages. Sampling windows. Red flags.
In this blog we dug down on the failures, again pulling out the informative cases. Abraham Wald reminded us that we need to hunt for the full distribution. A review of failure tracking methods like anti-patterns and post-mortems shows that we are off to a good start, but have a ways to go. Peeking at kindred spirits in medicine and psychology cautioned us that talking about failure requires infrastructure, conventions and venues that make identifying and sharing failure possible. PDCA cannot give us industry-wide benefits if the explorations are firewalled off.
This was a tough hike, and I appreciate that. But we are engineers. We work by trading off advantages and disadvantages. Only a complete picture of success and failure, of informative and uninformative makes that balancing act possible.
So I am going to head over to finish Part 3, where we look at some specimens of success stories from the wild. This is where we bring it all together to make it real.
And if you are disappointed that I hawked no ready-made solution to sell you on: read the blog one more time 😉
PS: If you are interested in an analysis of the estimation work of Abraham Wald, there is an often-cited explication by Marc Mangel and Francisco Samaniego, Abraham Wald's Work on Aircraft Survivability, June 1984, Journal of the American Statistical Association 79(386):259-267. There is a copy of the PDF on ResearchGate.
Appendix
Purposeful Omissions
Here is information that I left out because I felt it distracted from the overall argument. However, one could argue that this is relevant to the overall shape of the discussion.
Like his hero Ronald Fisher, Theodor David Sterling began working for the tobacco companies in the mid-1960s as an expert witness, detracting from the mounting evidence of tobacco use and health concerns. In the late 1980s he was finally outed as a paid consultant of the tobacco industry and thereafter not taken seriously as a witness in the tobacco discussion. However, this does not invalidate his 1959 paper that we are citing here.
The 2013 initiative of Registered Reports was instrumental in the reboot of the way psychological research conducted in the Western tradition. By 2017, when Chris Chambers wrote his research critique, The Seven Deadly Sins of Psychology, [Google Books], 40 journals were already offering Registered Reports. By the time the 2019 paperback of Seven Deadly Sins was published, that number had grown to 150 journals (Preface, Chambers 2019). Thus, the proposal was not a singular event, but the beginning of a movement.
Are you ready to bring the good, the bad and the improvables about your own software projects to the surface?
Don’t let the informative failures within your organization remain buried. Give Posedio a call.



