Skip to content
Back to Blog

Can you believe that? Reading IT News Critically

12 min read
Can you believe that? Reading IT News Critically

Introduction

Eat jelly doughnuts and lose twenty pounds a day

Hear the story of the man born without a head

And top psychics all agree // That the telephone company

Will have a brand new service that lets you talk to the dead

— Weird Al Yankovic, Midnight Star (1984)

Boy, 1984 seems light years away. Back when US comedian and singer Weird Al Yankovic wrote this song, you had to go to a news kiosk or supermarket checkout lane to get the trashy publications that contained these stunning revelations. Many countries were still under the dominance of their national telephone companies, which were struggling to maintain their monopolies with customer-responsive add-on services.

Fast-forward to 2026, and the media landscape has changed considerably. Social media like Facebook and X (formerly Twitter) or subscription blogging sites like Medium and substack are now the common go-to places for interesting news and news that are just … interesting.

What has not changed is the breathless tone of the headlines – now called click-bait, due to the differences in monetization. And the IT industry is no better, frankly, than other knowledge domains are.

Which is understandable. Every IT project manager, systems architect or software developer is looking for that “silver bullet” that is going to make their tasks complete on time, on budget, while delivering customer satisfaction. Simply plopping in database X or rewriting in language Y or trying library Z … it is tempting! Perhaps not as tempting as eating jelly doughnuts, but pretty close.

The choice is clear. If the tale is true, then waiting to switch is a waste of time and resources. If the tale is false, the switch will be a waste of time and resources, pushing a troubled project into Titanic mode. The validity of the claims is decisive.

But how to tell?

Coverage

Let’s circle back to the jelly doughnuts for a moment. 20 pounds a day is probably “a typo”, but 2 pounds a week would be a healthy rate of weight loss. What evidence would we like to see to make that move? Probably a medical study. Something like a randomized control trial. Since there are no placebo doughnuts (now there is a market!), it is not going to be double-blind. A good mix of subjects, in both groups, definitely some people in our own age and weight and life-style bracket. 

What this thought experiment is pointing out is that we are concerned with repeatability. Someone did X and they achieved (amazing) result Y. I won’t want to go through the effort until I know that I have a chance of achieving that same outcome. The positive effect needs to repeat itself for me.

My trust in that repeatability goes up if this was a fair experiment, that lasted long enough and took into consideration a mix of circumstances. If the situation sounds too different from my own, it becomes a roll of the dice whether the change will help or hinder. 

By definition, software projects are unique. Each has its own set of circumstances, at minimum due to the participants, location and time frame. Repeating the project is a ridiculous idea. The river is different, each time you step into it, as the Ionian nature philosopher Heraclitus pointed out. 

What is repeatable is the constituent parts: the databases, the development frameworks, the programming language, the algorithms, even the domain models. That is the part one might reasonably learn about and transfer to one’s own tool box. 

Except, as with the medical study, we would want a good mix of projects that use a particular constituent successfully. This is akin to the good mix of doughnut eaters in our medical study. How many of them? Simplifying the central limit theorem considerably, about thirty. So thirty independent narratives about projects successfully delivering with database X should do it?

Which pushes the question to the story of the success. We want to use database X because it was instrumental to the success. That is the justification for extending our toolkit. But for that, we have to make sure that the database – and not something else — contributed materially to the success. The success is after all what we wish to replicate.

Illusions of Success

Sometimes a simple success story is just an illusion. The story tells the success one way, but there were other causes – not mentioned – that contributed to the outcome. This is important because you have to replicate all causes, not just the ones the story fronts … otherwise too different; and then the results may not hold. Such a story works like a mirage in the desert – everything looks great until you get closer and realize there is no there there. Which may well leave you with a parched feeling!

Pointing out an illusion of success does not accuse anyone of lying. Calling it a mirage just means that the causal story is more complicated than presented. That there were other aspects to the story; aspects that did not receive the weight they possibly had. That there are alternate explanations; explanations that may not involve database X at all.

One famous illusion of success is due to survivor-bias. All the companies that went out of business picking database X won’t be there to blog, present at industry conferences or make online videos. That does not just skew the results; that is downright unfortunate, because their lessons learned might have been especially informative. However, the story around survivor-bias is more complex than that and deserves fuller treatment in a later blog; so just name-checking that bias here.

The IT industry is particularly partial to some illusions of success. Looking at them will make that very clear – and alert us as to what to look out for.

Selection Bias – The Atypical will not Generalize

This is the “We adopted X and improved Y” type of story. This story is appealing because everyone wants to improve Y. The story is illusory if it was not X that improved Y, but something else entirely.

If that adoption of X was implemented by an ordinary team in an ordinary situation within an ordinary organization, this is about as close to the random medical trial as we can expect. But if this is a select organization, situation or team …?

Let’s think these through in turn.

Our first concern: Perhaps the organization was exceptional. It sounds like they had a management that supports engineering culture and is willing to drop resources on adoption experiments. Notice that they signed off on the budget without knowing that Y would improve at all. That’s hindsight.

Or: Perhaps the situation was exceptional. Perhaps the team was freed up from the usual distractions of bug crawls, deadlines and meetings.

Alternatively: Perhaps the team was exceptional. Perhaps the best architects, lead developers and most talented interns were pooled. The kind of people that could have improved Y by rewriting everything in Object-Oriented Cobol.

Perhaps the team size was exceptional. Hyper-Scalers can field team sizes that the usual companies could not afford.

In order to make “adopting X improves Y” work for your team, those adopters have to be like your team. Otherwise you can never shake the suspicion that their “dream team” status, situation or corporate support overwhelmingly affected the success.

Underestimating Natural Variation

Every situation is another roll of the dice. Some are amazing (green), some are awful (red), and most are just average (yellow). That’s what natural variation means.

Because people like to think that they are in control, they tend to underestimate how much of what happens is just such natural variation. The return to the average after the exception is called Regression to the Mean. Put differently, because the ends of the distribution are rare, they can trip up attribution when looking at sequences of situations.

Regressing toward the mean from the bad end is the Coach’s Fallacy. Some coaches like to think that their yelling at the team, after a bad performance, is what makes the next game better. But statistics say that the team was bound to play average anyway. In our drawing, that happens when a red-arrow situation is followed by a yellow-arrow situation. This is eventually unavoidable, because the red-arrow is unusual and therefore rare. Since yellow-arrow is not as bad as red-arrow, the effect is an improvement. Even without doing anything at all, the situation would have gone back to “average” on its own.

In IT land, this is the “After bad event X we ….” type of story. Notice how the red-arrow is the hook of the whole narrative. Anything that comes after the “we” will most likely have restored the average state.

The opposite direction is called the Winner’s Curse, a situation where the exceptionally good – the green-arrow – happens. Alas, because green-arrow is also rare and yellow-arrow the normal, the success turns out to be irreproducible. A one-hit wonder, to put it in popular music terms.

In IT land, this is the “For client X, we …” type of story. There was a single glorious constellation where all the pieces came together just right. But will that generalize for anyone else, ever?

Insufficient Sampling vs Experience

In order to understand the next mirage, we need to look at a more involved example. The following is the depiction of a stream of events – say, purchases on a website – over time, from the point of view of the team hosting the website. Time moves left to right in this chart.

At the top, visitors / requests /queries are coming at an average clip. Then there are times of unusual traffic; some of these have names in the business world, e.g. the end of the fiscal quarter or Black Friday. Some are cultural and do not have specific names in the IT world – Super Bowl, religious holidays or the start of the school year. And then there are the typical outages and devops disasters.

All of these together influence whether the website is delivering the intended performance or not. A true understanding of the way the website behaves is highly dependent on which slice of time you study. There are multiple slices of various widths in this event stream where everything looks great. There are multiple times when unusual things happen and things might not be going smoothly at all.

In order to want to believe whatever story they are telling, you want the sampling window to be as wide as possible. If the web site braved outages, social cycles and cultural events, there is a lot to learn from a success story. But if the success is due to the happy slice of nothing out of the ordinary happening that could challenge the system … you may want to keep looking.

Notice that as these events roll by – the ups and downs in traffic, the outages and security updates – not only does the team somehow deal with that; they gain experience. They become seasoned. By the end they have morphed into the atypical team that we were worried about when talking about selection bias above.

All other things being equal, the success story is most impressive if the window is earlier rather than later. But the window must be sizeable – to dodge the Winner’s Curse!

More things to Watch out For

There can be other warning flags in a narrative that can contribute to the illusion of success. One should well remain hesitant to change tooling due to such a tale.

One flag is when the focus on new metrics only began with the change. This is the “we started gathering metrics” story. Ideally you want them to have focused on key metrics before they started making the change. Not only are there complex psychological issues – the so-called Hawthorne Effect – to worry about; but the lack of a baseline to measure the improvement is unhelpful.

Another flag is when the connection of the star component to the other elements of the tech stack is not spelled out. Clearly the other components were present; clearly they influenced the outcome. Perhaps the way they tuned the edge cache allowed database X to shine. No edge cache, no shine for you!

Another flag is the claim of improvements without any counterfactual. If the “resulting system was more Y” then we would like to know: compared to what? What system exhibited the “less Y”? And was it a comparable system? Ideally, there would be competitive prototyping or adversarial testing to firm up that claim.

Finally, keep an eye out for regulatory frameworks. There are substantial differences between countries in terms of data protection laws, privacy, financial regulations or even worker rights. Skipping privacy guards can make software faster and cheaper to build. Implementing data regulations costs resources. US vs EU on GDPR issues – that’s comparing apples to oranges.

Outlook

All writers want to grab their audience’s attention. The question is how to distinguish those that want to explain from those that want to monetize eyeballs. Learning something new or useful seems like a fair trade for someone pocketing the ad revenue. But for that, the story has to work out – which is why I read with a critical eye.

In this blog post we discussed how to accumulate helpful information by getting the proper coverage of examples for a technology, programming language, algorithm or library.

I looked at illusory stories of success that read great but will not help your team: Marred by selection bias, regression to the mean – whether as Coach’s Fallacy or Winner’s Curse, and the insufficiently sized sampling window, they beckon like the desert mirage but will leave you stranded.

I mentioned specific warning flags to watch out for as well – such as changing metrics, hazy descriptions and improvements without a proper counterfactual. Just like regulatory frameworks, any of these can keep what worked for them from working for you.

Some of you might feel that all of this was a bit theoretical or too abstract. Well, stick around – in Part Three we are going to look at some case studies! Part Three? Yes! Because I still owe you the promised analysis of survivor-bias. Hope to see you in Part Two, where we discuss the problem of Useful Failure!

Do it NOW.

Interested in working together on something like this?

Get in touch
Robert C. Kahlert

Robert C. Kahlert is a Senior Software Researcher at Posedio. He comes from a background in symbolic AI research and has contributed to projects and products in the fields of pharma, defense, medicine, natural language processing, and resource exploration. His passions include software engineering, compiler construction, cloud computing, knowledge representation, and databases. In his free time, he enjoys Vienna's diverse museum scene.

View all posts by this author

Similar Posts

Become Part of the Community!

Sign up for our newsletter and never miss an event, talk, or update.