Tidy Data
Hadley Wickham
Defines a tidy dataset as one where every variable is a column, every observation a row and every value a cell, and shows most data-cleaning pain comes from a fixable violation of this.
link checked 17 Sept 2026Doing statistics at a scale where the arithmetic matters.
7 topics · 10 curated works
No prior grounding assumed.
Tidy Data
Hadley Wickham · 2014
Defines a tidy dataset as one where every variable is a column, every observation a row and every value a cell, and shows most data-cleaning pain…
+1 more at this level
Assumes you know the vocabulary.
The Monte Carlo Method
Nicholas Metropolis & Stanislaw Ulam · 1949
Proposes solving otherwise-intractable integrals and diffusion problems by simulating random samples on a computer instead of solving the underlying…
+5 more at this level
Primary sources and full treatments.
Maximum Likelihood from Incomplete Data via the EM Algorithm
Arthur P. Dempster, Nan M. Laird & Donald B. Rubin · 1977
Formalises the EM algorithm as a general way to maximise a likelihood when data are incomplete, alternating between guessing the missing part and…
+1 more at this level
10 works
Hadley Wickham
Defines a tidy dataset as one where every variable is a column, every observation a row and every value a cell, and shows most data-cleaning pain comes from a fixable violation of this.
link checked 17 Sept 2026Hadley Wickham & Garrett Grolemund
Teaches the whole analysis cycle — import, tidy, transform, visualise, model — as one workflow rather than as separate tools.
link checked 17 Sept 2026Nicholas Metropolis & Stanislaw Ulam
Proposes solving otherwise-intractable integrals and diffusion problems by simulating random samples on a computer instead of solving the underlying equations directly.
link checked 17 Sept 2026Christophe Andrieu, Nando de Freitas, Arnaud Doucet & Michael I. Jordan
Surveys rejection sampling, importance sampling, and the main MCMC algorithms as one family of solutions to the same problem, drawing samples from a distribution too complex to sample from directly.
link checked 17 Sept 2026Roger D. Peng
Argues that a published computational result should ship with the code and data needed to regenerate it, since peer review alone cannot catch most analysis errors.
link checked 17 Sept 2026John C. Nash & Ravi Varadhan
Compares R's competing general-purpose optimisers on a common set of test problems and shows none of them dominates, which is why optimx wraps rather than replaces them.
link checked 17 Sept 2026Bradley Efron & Trevor Hastie
Traces the arc from classical inference through the bootstrap, generalised linear models, survival analysis and machine learning as one continuous story about what computation made possible.
link checked 17 Sept 2026Tim P. Morris, Ian R. White & Michael J. Crowther
Sets out a checklist — aims, data-generating mechanisms, estimands, methods, performance measures — for a simulation study to be reported so another researcher can trust and reproduce it.
link checked 17 Sept 2026Arthur P. Dempster, Nan M. Laird & Donald B. Rubin
Formalises the EM algorithm as a general way to maximise a likelihood when data are incomplete, alternating between guessing the missing part and refitting the model.
link checked 17 Sept 2026Bradley Efron & Robert Tibshirani
Extends the bootstrap from standard errors to bias, prediction error and confidence intervals, and works through complicated data, censored, regression, time series, where a closed-form formula would not exist at all.
link checked 17 Sept 2026