A Power Primer
Jacob Cohen
Gives working effect-size conventions and a lookup table for statistical power, arguing most published psychology studies were underpowered to detect the effects they claimed to be testing for.
link checked 17 Sept 2026Structuring an experiment so the answer means something.
9 topics · 13 curated works
No prior grounding assumed.
A Power Primer
Jacob Cohen · 1992
Gives working effect-size conventions and a lookup table for statistical power, arguing most published psychology studies were underpowered to detect…
+2 more at this level
Assumes you know the vocabulary.
Controlled Experiments on the Web: Survey and Practical Guide
Ron Kohavi, Roger Longbotham, Dan Sommerfield & Randal M. Henne · 2009
Argues that online controlled experiments are cheap enough to run constantly, and catalogues the practical failure modes — from clock skew to novelty…
+1 more at this level
Primary sources and full treatments.
The Arrangement of Field Experiments
Ronald A. Fisher · 1926
Introduces randomisation itself as the device that justifies treating the difference between plots as evidence about the treatment, rather than about…
+7 more at this level
12 of 13 works
Jacob Cohen
Gives working effect-size conventions and a lookup table for statistical power, arguing most published psychology studies were underpowered to detect the effects they claimed to be testing for.
link checked 17 Sept 2026Veronica Czitrom
Shows with worked examples that changing one factor at a time while holding others fixed misses interactions a factorial design catches for the same or fewer runs, so OFAT is not even the cheap option it is chosen for.
link checked 17 Sept 2026Elise Whitley & Jonathan Ball
Walks through the four quantities every sample-size calculation trades off against each other — effect size, variability, significance level and power — so a stated sample size can be audited rather than taken on faith.
link checked 17 Sept 2026Ron Kohavi, Roger Longbotham, Dan Sommerfield & Randal M. Henne
Argues that online controlled experiments are cheap enough to run constantly, and catalogues the practical failure modes — from clock skew to novelty effects — that make a technically correct A/B test still mislead.
link checked 17 Sept 2026Kenneth F. Schulz, Douglas G. Altman & David Moher
Sets a minimum checklist for reporting a randomised trial, arguing that incomplete reporting of allocation and blinding makes trials that were actually well designed look indistinguishable from ones that were not.
link checked 17 Sept 2026Ronald A. Fisher
Introduces randomisation itself as the device that justifies treating the difference between plots as evidence about the treatment, rather than about whatever else varied between them.
link checked 17 Sept 2026Frank Yates
Works out how to vary several factors at once in a single experiment and still separate their effects, including deliberately confounding higher-order interactions with blocks to keep the design tractable.
link checked 17 Sept 2026Abraham Wald
Derives the sequential probability ratio test, showing that letting sample size depend on the data collected so far reaches a decision with fewer observations on average than any fixed-size test of the same reliability.
link checked 17 Sept 2026Herbert Robbins
Poses the multi-armed bandit problem formally and shows the sequential-design question is not just when to stop but which arm to pull next, founding adaptive experimentation as a distinct topic from fixed sequential testing.
link checked 17 Sept 2026Bruce D. Meyer
Sets out what a quasi-experiment can and cannot license once randomisation is gone, and catalogues the extra checks — placebo tests, pre-trends, robustness to functional form — that stand in for it.
link checked 17 Sept 2026Kari Lock Morgan & Donald B. Rubin
Proposes discarding and redrawing a randomisation whenever it leaves covariates badly imbalanced, and shows this can be done without the usual randomisation-based inference breaking.
link checked 17 Sept 2026Ron Kohavi, Alex Deng, Brian Frasca, Toby Walker, Ya Xu & Nils Pohlmann
Describes the infrastructure and statistical discipline needed to run thousands of A/B tests at once, including why most proposed changes fail to beat the control and why that is the expected, useful outcome.
link checked 17 Sept 2026