A statistics demonstration

The Garden of Forking Paths

Below is a randomised trial of a study-skills app, run on 240 students. There is no treatment effect in this data. None. The outcomes were generated before anyone was assigned to a group, so the app cannot possibly have done anything.

Your job is to find a significant result anyway. Pick an outcome, decide who to exclude, choose what to adjust for, look at a subgroup. Every one of those is a choice a careful analyst has to make, and none of them is cheating. Keep going until the p-value drops below 0.05. It will.

The trial

240 students, half assigned to the app and half to business as usual by shuffling a balanced list. Three outcomes were measured at the end: a final exam score, a homework completion rate, and a self-reported confidence index. Four baseline variables were recorded before assignment.

The pre-registered analysis

Written down before the data arrived: compare the final exam score between the two arms with a two-sample t-test, on everybody, adjusting for nothing. One number, decided in advance. This is the only result on this page that means what it says.

Your analysis

Six decisions. Change any of them and watch the answer move.

p-value
waiting

Paths you have walked

0 analyses run.

    The multiverse

    There are 1,920 combinations of those six choices. Once you have found your result, look at the other 1,919.

    What just happened

    Nothing in the data changed while you were clicking. What changed was how many chances you gave yourself. A single pre-specified test on null data returns p < 0.05 about 5% of the time, which is the deal you signed up for. Search across all 1,920 defensible specifications and at least one of them clears the bar 87.3% of the time. That figure is measured rather than asserted. The project's test suite sweeps 150 null cohorts on every run and 131 of them hand over at least one significant result, against 6.0% for the single pre-specified test on the very same cohorts.

    The uncomfortable part is that nobody has to run all 1,920 for this to happen. Gelman and Loken's argument is that a researcher who runs exactly one analysis can still be caught by it, because the analysis they ran was chosen after seeing the data. Had the numbers come out differently they would have reasonably chosen differently, and that whole set of would-have-beens is what sets the real error rate. The single path actually walked says very little about it. This is why they framed it as a problem for honest researchers. No fabrication is involved and nobody needs to know they are doing it.

    The defences are unglamorous. Pre-register the analysis, as the box above does. Report the whole multiverse rather than the corner of it that worked, which is what Steegen and colleagues proposed. Disclose every choice considered, which is what Simmons, Nelson and Simonsohn's disclosure standard asks for.

    Sources

    • Gelman, A. and Loken, E. (2013). The garden of forking paths: Why multiple comparisons can be a problem, even when there is no "fishing expedition" or "p-hacking" and the research hypothesis was posited ahead of time. Department of Statistics, Columbia University.
    • Simmons, J. P., Nelson, L. D. and Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science 22(11), 1359–1366.
    • Steegen, S., Tuerlinckx, F., Gelman, A. and Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science 11(5), 702–712.
    • Simonsohn, U., Simmons, J. P. and Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour 4, 1208–1214. The two-panel figure above follows their layout.