Digital product and experimentation

From scurvy to streptomycin: what centuries of clinical trials teach product teams about efficacy, relative improvement and the real cost of experimenting.

November 20, 2024
Digital product and experimentation

At Instituto Tramontana we ran regular events to introduce the product management and culture programme to prospective candidates. Rather than describe the content, I preferred to reproduce it: half a tour of what was happening in the current edition, half an abbreviated simulation of an actual class, so that the dynamic of debate and reflection could be experienced rather than summarised.

One of those sessions dealt with digital product and experimentation. What follows is roughly the thread we followed.

Pyramids and medicine

The most complete synthesis available for anyone entering the world of digital product experiments remains Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, by Ron Kohavi, Diane Tang and Ya Xu. The use of the word trustworthy is significant, because in the literature it has become the third link in a chain that began with cause, continued with correlation, and finally arrived at trustworthiness.

Two references appear in the authors' main argument for situating the framework of experimentation, and we used both as the thread of the session.

A book paragraph on Guyatt's hierarchy of evidence, with «evidence» and «medical literature» circled in red

The hierarchy of evidence reaches digital product work on loan from medicine.

The evidence pyramid has become the standard reference when ranking the sources that serve to make decisions. One of the reflections we worked through was what it would look like to carry something similar into digital product territory, imagining that we could classify initiatives on different rungs.

The evidence pyramid, with meta-analyses at the apex and expert opinion at the base

At the base of the pyramid, expert opinion is labelled HiPPO.

The exercise is uncomfortable in a useful way. Most product decisions, honestly classified, sit near the bottom of that pyramid: expert opinion and case reports. That is not a scandal — medicine spent centuries there — but it does change how you talk about a decision when you know which rung it is standing on.

Efficacy and the relative

The reference to the medical literature led us straight to clinical trials. They are a comparative experiment which, like so many simple ideas, conceals several powerful points.

Two blocks of identical figures, one orange and one green, under the label «comparative experiment»

Comparing requires treating the differences inside each group as negligible.

The first is that it assumes individuals can be grouped on the basis of similarities sufficient to disregard their differences. The idea is to apply different treatments — including none at all, if we want a null control group — to different groups in order to evaluate the results. The second strength connects to that evaluation: efficacy becomes the central concept, made operational through some criterion. Treatment C is better than treatment B because it shortens sick leave by two days, by allowing faster recovery.

The word «efficacy» handwritten and underlined in orange

Efficacy only exists once someone fixes the criterion it is measured against.

Efficacy, of course, is always relative. And here we find another strength that suits thinking about digital product rather well, given how prone that thinking is to absolute criteria: we're going to redesign the whole site to fix all the problems, when will we stop having payment issues with Apple?, let's give this time so we can clear all the incidents.

The absolute framing is not merely optimistic. It is unmeasurable, which is its real defect. If the target is all, there is no arithmetic that can tell you whether you got closer.

From scurvy to tuberculosis

The history of clinical trials is a history of centuries. That suits us very well, because once again we learn that we ride on the shoulders of giants. It is also, as it happens, a history full of good anecdotes. What is considered the first proto-clinical trial involves ships, scurvy, and a naval surgeon, James Lind, who set about testing — experimenting with — different treatments applied to different groups, which is to say segments. The citrus-based treatment turned out to be the most effective.

A painting of an eighteenth-century naval surgeon giving citrus to sailors sick with scurvy below deck

James Lind hands different treatments to different groups, on board.

How effective? In the eighteenth century they still could not answer that question. Various experiments followed one another over time and combined with the principal source of decision in prescription: clinical judgment. Based on experience, on the observation of cases, on familiarity with variation, expert judgment will not stop being the protagonist of decisions, even in our own day. Its main objection is the one that accompanies us permanently as a biological population: bias.

A clinician examining an X-ray beside a patient's hospital bed

Clinical judgment remains the protagonist of the decision, biases included.

But in the first decades of the twentieth century the development of statistics as a discipline, and its expansive use across different areas, did begin to answer that question of how much, the one tied to efficacy. It is almost at mid-century that we get what is considered the reference clinical trial.

A normal curve with the right tail shaded beyond a cut-off point, under the label «statistically significant difference»

Statistics finally makes the «how much» question answerable.

It is also a small work of art of human knowledge, for the way it combines humility in its framing with ambition in its results. This time the subject was not scurvy but the treatment of tuberculosis. Bradford Hill synthesised in a single table the percentages of efficacy from a streptomycin treatment group and a control group, together totalling 107 patients.

A 1948 British Medical Journal table comparing the streptomycin group with the control group, showing 7 per cent and 27 per cent deaths

The streptomycin clinical trial, published in 1948.

Alongside several dimensions, the one considered most critical — deaths — began to display a quantitative criterion: 7% against 27%. Efficacy does not seek to abolish the problem. The deaths are accepted as part of the equation. What it seeks is relative improvement.

That sentence is the one I would take out of the session and pin to a wall.

What to spend on

An instrumentation of efficacy like this would not, on its own, have had to become the standard we work with today. A series of social factors, in play at the same time, gradually led to the current situation.

The first was the role of the state as an economic agent and, in particular, its prominence in creating social insurance and subsidising treatments. A logical question begins to surface, easy enough to understand from inside the digital product industry even if we rarely see it day to day: what should we spend the money on? Or, put differently, how can we have confidence that we are putting money into the problems and the solutions that genuinely produce results?

A handwritten label: «the state as investor» and, below it, «where to spend?»

Whoever pays for the treatments ends up asking for proof that they work.

The second factor has to do with the other protagonist of the story: the pharmaceutical industry. How do we correct incentives that can falsify a supposed efficacy? There is a search for a degree of impartiality, and this is the reason instruments like the double blind — neither the patient nor the professional should know which treatment is being applied — also end up becoming a standard.

A handwritten label: «impartiality» and, below it, «double blind»

The double blind exists because incentives distort results.

The parallel is worth sitting with. Every product organisation has a party with an interest in the result. The person who proposed the initiative usually designs the measurement, interprets the outcome and presents it. Medicine took a century to conclude that this arrangement does not work, and built machinery to prevent it. We have retrospectives.

Careful about playing at laboratories

This brief tour, besides encouraging debate and an exchange of experiences, helped us widen our perspective: beyond modifying a few strings or buttons in a signup flow, beyond changing a design in a transactional email.

Widening the view helps you take a more critical perspective on what you are doing. You now know that when you pronounce the word experiment there is a historical and methodological weight behind it, and that can help you pronounce it in lower case rather than capitals. It can also help you think about what it costs: the scaffolding needed to instrument all of this is not trivial. Success or failure often has to do with the dose.

James Flier's argument about irreproducibility in bioscience is the necessary counterweight to any enthusiasm here. The method does not guarantee the result. A field with vastly more rigour than ours, and considerably higher stakes, still produces a large volume of findings that do not survive replication. An A/B test run on insufficient traffic does not produce knowledge; it produces a confident number, which is worse than no number at all.

What we take away is that the spirit resonating through all of this can help us professionalise our organisations and teams. Complex environments — and we already know that the oil stain that is software makes everything especially complex — require a reflection on how they decide what they decide, what opportunity cost their decisions carry, and finally what efficacy they end up having. Whether in the form of experiments, of bets built out of qualified judgment, or of plain intuitions.

Peter Medawar's phrase for this was the art of the soluble: the researcher's real skill lies less in answering hard questions than in recognising which questions are currently answerable. That skill transfers directly, and it is the one no framework can install for you.

2026 © Íñigo Medina