The soft underbelly of hard data
Deming, John Graunt and the data rush: why interpreting, measuring and experimenting in digital product is more slippery than it looks.
Digital product management is no stranger to the growing tendency to grant special relevance to data analysis. The most influential technology companies have normalized the expression "data-driven companies", recasting the spirit of what William Edwards Deming used to say, a statistician and a pioneer in the study of quality management in organizations: "In God we trust; all others must bring data." Deming had a substantial influence across countries and industries with what came to be known as his "14 principles". The book in which he set them out could have been published this very year: Out of the Crisis.
Neither the social sciences nor public opinion escape this rebirth of data. In the former, you can see how quantitative information has revived the drive to get closer to the "hard sciences", trying to make them more empirical and better tested against reality. In the latter, the Covid crisis turned public opinion into a constant conversation about figures for tests, patients and deaths. We are all a little like John Graunt, a forerunner of statistics and considered the first demographer, who worked from weekly death records looking for how to make better decisions in order to survive the plague.
That said, the social sciences, public opinion and anyone working on software-based products all have to face, once the data high wears off, a series of obstacles that make the whole thing genuinely interesting and far more slippery than we would like.
In the first place, data has to be interpreted. And so there is what has sometimes been called the "soft underbelly" of hard data. It is obvious, for instance, that in economics the analysis of the available evidence is complex and always has many edges: establishing causal relationships is a real art, when it is not simply a web of statistical correlations built on certain hypotheses, theories or intuitions. What happens to public opinion when it cheerfully shares infection percentages runs along the same lines: the relationships between factors run in many directions. The "soft underbelly" of someone working on a digital product also has to take in things like rumors -"I've been told WhatsApp is down"-, gossip and impressions, without being able to ignore them, because sometimes they really are the canary in the mine.
In the second place, there is the common mistake of illegitimately using inadequate data. It is fairly frequent that, by not paying attention to where the data comes from and how it was produced, you end up using figures that are incorrect or even manipulated. But it is no rarer to end up handling well-built data for which you do not have the right unit of measurement. Neither the sociologist comparing two provinces nor the product person comparing an acquisition campaign against a competitor is safe from doing it without proportion.
Then it is very natural to confuse statistical association with causality: we shipped the update on accounts and the customers we have from the agreement with that other company started complaining, therefore.... Front pages are full of bright ideas where coincidence dresses up as causation. And in companies in need of fuel and investment there is a fairly frequent, atavistic pull towards this kind of magical thinking.
Besides, data does not travel alone. In the social sciences there is a current called the "credibility revolution", which advocates the use of experiments as happens, again, in the "hard sciences", in search of conclusive empirical evidence. This current has popularized field experiments and "Randomized Controlled Trials" (RCT), widespread in medicine and drug development, comparing randomly chosen groups: one that receives a given measure (intervention) and another that serves as a placebo (control). None of this will sound entirely foreign to someone moving in the terrain of software-based products either, because the use of experiments, A/B tests and similar mechanisms has been preached there too.
These new proposals are not free of obstacles either. In many cases some of the ones we mentioned about data come back, along with new ones related to the impossibility of reproducing certain scenarios as experiments, whether for reasons of time, resources or even ethics.
So we have to live with data and with experiments, but without losing our heads, because when we lose them we do not get better decisions, or more truthfulness, or more credibility. Along with a serious effort not to fall into greater confusion and opacity, which get in the way of transparency and of agreed action.