Digital product and experimentation
From scurvy to streptomycin: what centuries of clinical trials teach product teams about efficacy, relative improvement and the real cost of experimenting.

At Instituto Tramontana we ran regular events to introduce the product management and culture programme to prospective candidates. Rather than describe the content, I preferred to reproduce it: half a tour of what was happening in the current edition, half an abbreviated simulation of an actual class, so that the dynamic of debate and reflection could be experienced rather than summarised.
One of those sessions dealt with digital product and experimentation. What follows is roughly the thread we followed.
Pyramids and medicine
The most complete synthesis available for anyone entering the world of digital product experiments remains Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, by Ron Kohavi, Diane Tang and Ya Xu. The use of the word trustworthy is significant, because in the literature it has become the third link in a chain that began with cause, continued with correlation, and finally arrived at trustworthiness.
Two references appear in the authors' main argument for situating the framework of experimentation, and we used both as the thread of the session.

The hierarchy of evidence reaches digital product work on loan from medicine.
The evidence pyramid has become the standard reference when ranking the sources that serve to make decisions. One of the reflections we worked through was what it would look like to carry something similar into digital product territory, imagining that we could classify initiatives on different rungs.

At the base of the pyramid, expert opinion is labelled HiPPO.
The exercise is uncomfortable in a useful way. Most product decisions, honestly classified, sit near the bottom of that pyramid: expert opinion and case reports. That is not a scandal — medicine spent centuries there — but it does change how you talk about a decision when you know which rung it is standing on.
Efficacy and the relative
The reference to the medical literature led us straight to clinical trials. They are a comparative experiment which, like so many simple ideas, conceals several powerful points.

Comparing requires treating the differences inside each group as negligible.
The first is that it assumes individuals can be grouped on the basis of similarities sufficient to disregard their differences. The idea is to apply different treatments — including none at all, if we want a null control group — to different groups in order to evaluate the results. The second strength connects to that evaluation: efficacy becomes the central concept, made operational through some criterion. Treatment C is better than treatment B because it shortens sick leave by two days, by allowing faster recovery.

Efficacy only exists once someone fixes the criterion it is measured against.
Efficacy, of course, is always relative. And here we find another strength that suits thinking about digital product rather well, given how prone that thinking is to absolute criteria: we're going to redesign the whole site to fix all the problems, when will we stop having payment issues with Apple?, let's give this time so we can clear all the incidents.
The absolute framing is not merely optimistic. It is unmeasurable, which is its real defect. If the target is all, there is no arithmetic that can tell you whether you got closer.
From scurvy to tuberculosis
The history of clinical trials is a history of centuries. That suits us very well, because once again we learn that we ride on the shoulders of giants. It is also, as it happens, a history full of good anecdotes. What is considered the first proto-clinical trial involves ships, scurvy, and a naval surgeon, James Lind, who set about testing — experimenting with — different treatments applied to different groups, which is to say segments. The citrus-based treatment turned out to be the most effective.

James Lind hands different treatments to different groups, on board.
How effective? In the eighteenth century they still could not answer that question. Various experiments followed one another over time and combined with the principal source of decision in prescription: clinical judgment. Based on experience, on the observation of cases, on familiarity with variation, expert judgment will not stop being the protagonist of decisions, even in our own day. Its main objection is the one that accompanies us permanently as a biological population: bias.

Clinical judgment remains the protagonist of the decision, biases included.
But in the first decades of the twentieth century the development of statistics as a discipline, and its expansive use across different areas, did begin to answer that question of how much, the one tied to efficacy. It is almost at mid-century that we get what is considered the reference clinical trial.

Statistics finally makes the «how much» question answerable.
It is also a small work of art of human knowledge, for the way it combines humility in its framing with ambition in its results. This time the subject was not scurvy but the treatment of tuberculosis. Bradford Hill synthesised in a single table the percentages of efficacy from a streptomycin treatment group and a control group, together totalling 107 patients.

The streptomycin clinical trial, published in 1948.
Alongside several dimensions, the one considered most critical — deaths — began to display a quantitative criterion: 7% against 27%. Efficacy does not seek to abolish the problem. The deaths are accepted as part of the equation. What it seeks is relative improvement.
That sentence is the one I would take out of the session and pin to a wall.
What to spend on
An instrumentation of efficacy like this would not, on its own, have had to become the standard we work with today. A series of social factors, in play at the same time, gradually led to the current situation.
The first was the role of the state as an economic agent and, in particular, its prominence in creating social insurance and subsidising treatments. A logical question begins to surface, easy enough to understand from inside the digital product industry even if we rarely see it day to day: what should we spend the money on? Or, put differently, how can we have confidence that we are putting money into the problems and the solutions that genuinely produce results?

Whoever pays for the treatments ends up asking for proof that they work.
The second factor has to do with the other protagonist of the story: the pharmaceutical industry. How do we correct incentives that can falsify a supposed efficacy? There is a search for a degree of impartiality, and this is the reason instruments like the double blind — neither the patient nor the professional should know which treatment is being applied — also end up becoming a standard.

The double blind exists because incentives distort results.
The parallel is worth sitting with. Every product organisation has a party with an interest in the result. The person who proposed the initiative usually designs the measurement, interprets the outcome and presents it. Medicine took a century to conclude that this arrangement does not work, and built machinery to prevent it. We have retrospectives.
Careful about playing at laboratories
This brief tour, besides encouraging debate and an exchange of experiences, helped us widen our perspective: beyond modifying a few strings or buttons in a signup flow, beyond changing a design in a transactional email.
Widening the view helps you take a more critical perspective on what you are doing. You now know that when you pronounce the word experiment there is a historical and methodological weight behind it, and that can help you pronounce it in lower case rather than capitals. It can also help you think about what it costs: the scaffolding needed to instrument all of this is not trivial. Success or failure often has to do with the dose.
James Flier's argument about irreproducibility in bioscience is the necessary counterweight to any enthusiasm here. The method does not guarantee the result. A field with vastly more rigour than ours, and considerably higher stakes, still produces a large volume of findings that do not survive replication. An A/B test run on insufficient traffic does not produce knowledge; it produces a confident number, which is worse than no number at all.
What we take away is that the spirit resonating through all of this can help us professionalise our organisations and teams. Complex environments — and we already know that the oil stain that is software makes everything especially complex — require a reflection on how they decide what they decide, what opportunity cost their decisions carry, and finally what efficacy they end up having. Whether in the form of experiments, of bets built out of qualified judgment, or of plain intuitions.
Peter Medawar's phrase for this was the art of the soluble: the researcher's real skill lies less in answering hard questions than in recognising which questions are currently answerable. That skill transfers directly, and it is the one no framework can install for you.