- General Overview
- The Central Thesis
- Statistics as problem-solving: data becomes knowledge through a disciplined cycle, not a bag of techniques.
- PPDAC cycle: Problem, Plan, Data, Analysis, Conclusion and communication repeat as questions evolve.
- Uncertainty is central: estimates, variability, and chance must be quantified, not hidden.
- Data literacy: everyone needs to analyze data and critique statistical claims in modern life.
- Real-world cases: Shipman, Bristol, statins, Titanic, and polls illustrate each stage.
- Data, Populations, Measurement
- Constructed data: definitions and measurement choices shape counts like trees, unemployment, and GDP.
- Inductive chain: raw data → sample truth → study population → target population, each link uncertain.
- Reliability and validity: repeatable measures must also measure what is intended without systematic bias.
- Sampling: random selection supports internal validity; non-response and bias threaten representativeness.
- External validity: study population must match the target population for generalization.
- Categorical data: counts and percentages need careful framing to avoid distorted risk perceptions.
- Summarizing and Visualizing
- Distributions: strip-charts, box plots, and histograms reveal shape, skew, and outliers.
- Averages: mean, median, and mode answer different questions; the mean is easily dragged by extremes.
- Spread: range, IQR, and standard deviation describe variability; robust measures resist outliers.
- Correlation: Pearson and Spearman describe relationships but do not establish causation.
- Visualization: good graphs show reliable patterns honestly, while log scales tame skew and aid comparison.
- Communication: know the audience, resist inserting opinion, and let context shape the message.
- Causation and Trials
- Correlation is not causation: associations may reflect ascertainment bias, confounding, or reverse causation.
- Counterfactual cause: X causes Y if intervening on X would change Y, probabilistically and not deterministically.
- Randomized trial: random allocation balances known and unknown confounders; placebo, blinding, and intention-to-treat protect fairness.
- Heart Protection Study: statins reduced vascular events; non-adherence diluted but mechanisms and meta-analysis confirmed benefit.
- Observational alternatives: cohort and case-control designs plus adjustment help when randomization is impossible.
- Bradford Hill: guidelines like dose-response, reversibility, and mechanism judge causation from observational evidence.
- Regression Models
- Least squares: fit a line minimizing squared residuals to predict an outcome from explanatory variables.
- Regression to the mean: extreme values tend toward average, so before-and-after comparisons can mislead.
- Multiple regression: adjusts coefficients for other variables, helping control confounders but not lurking factors.
- Logistic regression: models binary outcomes with curves bounded between 0 and 1.
- Model choice: simple statistics, simulations, machine learning, and causal models serve different purposes.
- Box's warning: all models are wrong; some are useful, so match model detail to the question.
- Algorithms and Prediction
- Algorithms: learn associations to classify or predict, not to explain mechanisms.
- Big data: large n and p enable clustering, feature engineering, and deep learning.
- Titanic example: training/test splits, baselines, and classification trees show practical prediction.
- Performance: ROC/AUC and Brier scores assess discrimination and calibration, not just accuracy.
- Over-fitting: complex models fit noise; cross-validation and regularization trade bias against variance.
- Simplicity and ethics: black boxes resist explanation and audit; prefer comprehensible, accountable algorithms.
- Uncertainty, Probability, Inference
- Sampling variability: statistics vary from sample to sample; standard errors and bootstrap intervals quantify that uncertainty.
- Probability language: rules for conditional probability, expected frequencies, and random variables clarify chance.
- Central Limit Theorem: averages become normal in large samples, enabling confidence intervals and tests.
- Confidence intervals: estimate ± margin of error; intervals express uncertainty better than headlines alone.
- Hypothesis testing: P-values measure compatibility with a null, but thresholds, multiple testing, and selective reporting mislead.
- Significance caution: non-significance is not zero effect; practical importance and replication matter.
- Bayesian Learning and Better Practice
- Bayes' theorem: prior odds × likelihood ratio = posterior odds, updating beliefs as evidence arrives.
- Bayesian applications: MRP polls, forensic likelihood ratios, and Bayes factors reinterpret evidence and uncertainty.
- Reproducibility crisis: questionable research practices, P-hacking, HARKing, and publication bias inflate false discoveries.
- Fraud and error: computational mistakes, wrong units, selective reporting, and fabrication all corrupt conclusions.
- Better practice: pre-registration, replication, transparency, data ethics, and the ten questions for claims improve statistics.
- Exit poll lesson: careful problem, plan, data, analysis, and communication can yield powerful, credible predictions.
- The Central Thesis
- Deep Dive
- Introduction
- Why We Need Statistics
- Harold Shipman case: forensic data analysis exposed a family doctor who murdered at least 215 patients.
- Visual exploration: scatter-plots and bar-charts revealed victim age, gender, and the 1992 gap in murders.
- Time-of-day pattern: Shipman's patients died overwhelmingly in early afternoon, pointing to after-lunch home visits.
- Statistics illuminates reality: turning tragedies into data enables questions, patterns, and better judgements.
- Turning the World Into Data
- Data requires definitions: "tree" is officially a plant with a woody stem meeting a diameter-at-breast-height threshold.
- Counting trees: biome sampling plus satellite imaging estimated 3.04 trillion trees, half the historical total.
- Constructed measures: unemployment definitions changed thirty-one times; GDP revisions include illegal drugs and prostitution.
- Data limitations: measurements are imperfect proxies and vary across people, places, and times.
- Statistical Science and Data Literacy
- Historical roots: modern statistics began in the 1650s when Pascal and Fermat formalized probability.
- Changing discipline: data science and big data demand skills beyond traditional statistical tools, including programming and subject knowledge.
- Reproducibility crisis: many published discoveries fail replication; inappropriate statistical methods share the blame.
- Data-dredging risk: analyzing many patterns and reporting only interesting findings inflates false discoveries.
- Data literacy: the ability to analyze data and critique others' statistical conclusions is essential in modern life.
- The PPDAC Cycle and This Book
- PPDAC framework: Problem, Plan, Data, Analysis, Conclusion and communication form a repeating problem-solving cycle.
- Problem and plan: statistical inquiry starts with questions and requires careful design, often neglected in practice.
- Data and analysis: good data needs cleaning; analysis may be as simple as a useful visualization.
- Conclusions: acknowledge limitations and communicate clearly, then let new questions restart the cycle.
- Book's approach: real-world problems introduce concepts, requiring no mathematics and little formal software.
- Why We Need Statistics
- CHAPTER 1 Getting Things in Proportion: Categorical Data and Percentages
- Bristol and the Limits of Counting
- Bristol inquiry: statistical comparison found around 30 excess child heart surgery deaths between 1991 and 1995.
- Definitions matter: "child," "open surgery," and "death within 30 days" all needed explicit choices before counting.
- Conflicting sources: national records and the surgical registry disagreed on deaths and operations; no source was authoritative.
- Bristol's legacy: public reporting of hospital survival rates began, but presentation can still shape how data are read.
- Presenting Counts and Proportions
- Binary data: yes/no outcomes are summarized by the count and percentage of events.
- Framing: "5% mortality" sounds worse than "95% survival," even though the risk is identical.
- Two manipulative tricks: switching from positive to negative framing, then converting percentages into counts of people.
- Row and column order: ordering hospitals by mortality rates implies league-table comparisons, ignoring case mix and chance.
- Axis choice: starting a bar chart at 95% exaggerates differences; starting at 0% makes them invisible.
- Pie charts: distort area comparisons, especially in 3D; bar charts based on length are clearer.
- Categorical Variables
- Variable: any measurement that takes different values; binary variables are the simplest yes/no case.
- Categorical types: categories may be unordered, ordered, or grouped numbers such as BMI thresholds.
- Visual comparison: height and length are easier to judge than areas, making bar charts preferable to pie charts.
- Relative versus Absolute Risk
- Relative risk: an 18% increase in bowel cancer from daily bacon sounds alarming.
- Absolute risk: that means one extra case, from 6 to 7 out of 100 lifetime bacon-eaters.
- Expected frequencies: presenting outcomes for 100 or 1,000 people improves understanding and appropriate concern.
- "1 in X" formats: many people misjudge which risk is larger, such as 1 in 10 versus 1 in 1,000.
- The Odds Ratio Trap
- Odds ratio: the ratio of the chance of an event happening to the chance of it not happening.
- Statins example: an odds ratio of 1.18 was misreported as a 20% increase, when absolute risk only rose from 85% to 87%.
- Communication rule: use absolute risks for general audiences; reserve odds ratios for scientific contexts.
- Bristol and the Limits of Counting
- CHAPTER 2 Summarizing and Communicating Numbers. Lots of Numbers
- Summaries, Spread, and Crowd Wisdom (CHAPTER 2 Summarizing and Communicating Numbers. Lots of Numbers · I)
- Wisdom of Crowds
- Galton’s ox experiment: 787 guesses, median 1,207 lb; actual dressed weight 1,198 lb.
- Jelly-bean replication: 915 guesses, median 1,775; true count 1,616, beating 90% of individuals.
- Crowd wisdom: robust summary emerges despite bizarre guesses, errors, and outliers.
- Seeing the Distribution
- Data distribution: pattern of sample values; shown by strip-chart, box plot, or histogram.
- Skewness: jelly-bean guesses are highly skewed, with a long right-hand tail and round-number clusters.
- Graph choices: strip-charts show individual points; box plots give rapid summaries; histograms reveal shape.
- Log scale: re-plots skewed data into a symmetric pattern, avoiding arbitrary exclusion of outliers.
- Variable types: counts are integers; continuous variables can be measured to arbitrary precision.
- What Does “Average” Mean?
- Mean: sum of values divided by cases; easily dragged by extremes.
- Median: middle ordered value; robust and Galton’s democratic choice.
- Mode: most common value; in jelly beans, 10,000 reflects round-number preference.
- Mean can mislead: for skewed data like income or partners, it diverges from typical experience.
- Average ambiguity: media blur mean and median; house-price “averages” are usually medians.
- Measuring Spread
- Spread matters: one summary is not enough; shoe sizes and plane seats need variability.
- Range: natural but highly sensitive to extreme values like the 31,337 guess.
- Inter-quartile range: distance between 25th and 75th percentiles; contains the central half.
- Standard deviation: common measure, but appropriate only for symmetric data; outliers distort it.
- Robust measures: median and IQR resist odd observations; ideal for messy real-world data.
- Comparing Groups: Sexual Partners
- Natsal survey: large, careful UK sexual behaviour study created to model AIDS spread.
- Summary comparison: men’s median is 3 partners higher; mean 6 higher; relative gap about 60%.
- Mathematical check: in a closed population, mean opposite-sex partners must be equal for men and women.
- Reporting bias: men may overplay and women underplay; raw data shows heaping on round numbers.
- Wisdom of Crowds
- Seeing Data Through Honest Visualization (CHAPTER 2 Summarizing and Communicating Numbers. Lots of Numbers · II)
- Describing Relationships Between Variables
- Visualization: no substitute for looking at data properly; summaries can hide the pattern.
- Volume effect: busier hospitals claim better survival; scatter-plots put that claim to the test.
- Outliers: Bristol’s poor survival distorts the pattern; removing it strengthens the association.
- Pearson correlation: runs −1 to 1; measures how close points fall to a straight line.
- Spearman rank correlation: uses ranks and captures steady increasing curves Pearson misses.
- Correlation limits: single points shift coefficients; association is not causation; zero need not mean none.
- Describing Trends
- Line-graphs: show world population tripling; standard scale hides continental contrasts.
- Log scale: separates continents and reveals Africa’s steeper growth.
- Relative increase: country-level summary exposes clusters and outliers by continent.
- Good visualization: reliable information, noticeable patterns, honest attractive design, exploration.
- Interactivity: lets viewers trace personal trends, like name popularity over decades.
- Communication
- First rule: shut up and listen; know your audience and its limitations.
- Second rule: know what you want to achieve; encourage open informed debate.
- Objectivity: statisticians declared facts-only but kept inserting opinions; resist the temptation.
- Numbers don’t speak: context, language and graphic design shape how messages are received.
- Storytelling honesty: pre-empt gut reactions, knowing people will judge and compare anyway.
- Storytelling with Statistics
- Dataviz vs infoviz: exploratory plots for insight; infographics for telling a story.
- Infographics: combine selected data to guide viewers, like the contraception gap.
- Dynamic graphics: Hans Rosling’s bubbles show countries moving toward health and prosperity.
- Computing: visualization is easier and more flexible; summaries hide as well as illuminate.
- Describing Relationships Between Variables
- Summaries, Spread, and Crowd Wisdom (CHAPTER 2 Summarizing and Communicating Numbers. Lots of Numbers · I)
- CHAPTER 3 Why Are We Looking at Data Anyway? Populations and Measurement
- From Sample to Population via Induction (CHAPTER 3 Why Are We Looking at Data Anyway? Populations and Measurement · I)
- Four Steps of Inductive Inference
- Inductive inference: generalizing from particular observations to wider truths is uncertain, unlike deduction.
- Four-stage chain: raw data → sample truth → study population → target population.
- Weakest links: measurement error, non-representative response, and mismatched target population all threaten conclusions.
- Statistical science: smooths each stage and states honestly what can and cannot be learned.
- Measurement: Reliability and Validity
- Reliability: low variability across occasions; repeatable, precise answers.
- Validity: measures what is intended, without systematic bias.
- Social acceptability bias: men may overstate, women understate sexual partners, distorting raw data.
- Question framing: logically identical wording can flip opinion, e.g. “giving 16-year-olds the vote” vs “reducing the voting age”.
- Priming: earlier questions can inflate answers, as in the BBC loneliness survey.
- Random Sampling and Internal Validity
- Internal validity: sample accurately reflects the study population; depends on random sampling.
- Stirring the soup: Gallup’s analogy — a representative spoonful, not the whole pot, is enough.
- Vietnam draft lottery: inadequate mixing produced biased birthdays; random sampling requires genuine randomization.
- Non-response bias: low response rates can make even large samples unrepresentative, as in 2015 UK polls.
- External Validity and Complete Data
- External validity: study population must match the target population we want to describe.
- Off-label extrapolation: trials on adult men may not apply to women and children; assumptions required.
- Routinely collected data: when all data are available, sample-population gap disappears, but measurement issues remain.
- Crime statistics: survey reports and administrative records each carry their own biases.
- Four Steps of Inductive Inference
- Populations, Biases, and Bell Curves (CHAPTER 3 Why Are We Looking at Data Anyway? Populations and Measurement · II)
- Trusting Data Sources
- Inductive path: data → study sample → study population → target population, with risks at each stage.
- Crime Survey: sample survey used as a designated national statistic for monitoring long-term trends.
- Police records: administrative census, but misses unreported and unrecorded crimes; recording changes doubled sexual offences.
- Divergent trends: survey estimated crime fell 9%; police recorded 13% more offences.
- Statisticians’ preference: survey more trusted; police data lost national-statistic status in 2014.
- Systematic biases: allocation bias, volunteer bias, and many other threats to reliable claims.
- The Bell-Shaped Curve
- Population distribution: the pattern made by the entire group of interest, not just the sample.
- Normal distribution: arises when many small influences combine, as in birth weight and height.
- Parameters: mean and standard deviation summarise populations just as statistics summarise samples.
- Empirical rule: about 95% of a normal population lies within ±2 SD; 99.8% within ±3 SD.
- Z-score: expresses a data point’s distance from the mean in standard deviation units.
- Friend’s baby: 2,910 g, 1.2 SD below the mean, on the 11th percentile.
- Percentiles and Probabilities
- Percentiles: median, quartiles, and inter-quartile range measure location and spread of populations.
- Dual meaning: a shaded area is both a population proportion and the probability for a random draw.
- Low birth weight: under 2,500 g; normal curve predicted 1.7%, observed 1.3%.
- Group differences: low-birth-weight rates are 1.3% in white full-term births, 8% overall, 13% in black women.
- Inference goal: usually we know data, not parameters; move from samples to populations.
- Three Kinds of Population
- Literal population: an identifiable group from which we can sample, like voters or YouTube viewers.
- Virtual population: all possible repeat measurements, such as successive blood-pressure readings.
- Metaphorical population: an imaginary space of possible outcomes, as in complete heart-surgery data.
- Value of the idea: treating observed data as a random pick makes sampling mathematics applicable.
- Mind-stretching: whether chance is real or divine, the mathematics remains the same.
- Trusting Data Sources
- From Sample to Population via Induction (CHAPTER 3 Why Are We Looking at Data Anyway? Populations and Measurement · I)
- CHAPTER 4 What Causes What?
- From correlation to controlled trials (CHAPTER 4 What Causes What? · I)
- Association is not explanation
- Linked registries: Scandinavian identity numbers connect health, education, and tax records.
- Brain-tumour finding: over 4 million Swedes showed higher socioeconomic position linked to slightly more diagnoses.
- Press distortion: "socioeconomic position" became "university education" via a press release and headline.
- Ascertainment bias: educated, affluent people are more likely to be diagnosed and have tumours registered.
- Mantra: correlation does not imply causation, statisticians have warned since 1900.
- Apophenia: humans spin causal stories from coincidences, like mozzarella and doctorates.
- What counts as causation?
- Counterfactual: cause means the outcome would have differed if the earlier event had not happened.
- Probabilistic cause: X causes Y when repeated interventions forcing X make Y tend to happen more often.
- Not deterministic: smoking causes lung cancer, yet most smokers escape and some non-smokers develop it.
- Individual uncertainty: a statin lowers LDL directly, but long-term personal benefit cannot be known.
- Two consequences: inferring cause requires intervening in experiments and repeating them to amass evidence.
- Designing a fair trial
- RCT: a fair test comparing treated and control groups under identical care apart from the intervention.
- Randomization: random allocation balances known and unknown factors that could affect outcomes.
- Placebo controls: sugar pills provide a baseline so benefits are not credited to expectation alone.
- Intention to treat: analyse participants as allocated, measuring the effect of being prescribed treatment.
- Blinding: participants, clinic staff, and outcome assessors should not know group allocation.
- Complete follow-up: full tracking avoids dropout bias; the Heart Protection Study reached 99.6%.
- The statin trial in action
- Heart Protection Study: 20,536 high-risk patients were randomised to simvastatin or a dummy tablet.
- Observed benefit: fewer major vascular events in the statin group, attributable to prescribed treatment.
- Non-adherence: 18% stopped statins and 32% crossed over, diluting the apparent effect.
- Corrected estimate: the true effect of taking statins is about 50% larger than the allocated comparison.
- Absolute benefit: 31 heart attacks avoided per 1,000; about 30 need five years of treatment to prevent one.
- Beyond a single study
- Replication: robust conclusions require multiple trials, not one isolated result.
- Systematic review: pool every relevant study, ignoring none.
- Meta-analysis: 27 statin trials showed 21% fewer major vascular events per mmol/L LDL reduction.
- Mechanism assumption: linking reduced LDL to benefit estimates the real treatment effect despite non-compliance.
- Chance ruled out: huge, combined trials make random variation an unlikely explanation.
- Roots of randomisation
- First modern RCT: the 1948 streptomycin trial for tuberculosis randomly allocated patients.
- Ethical fairness: random allocation was acceptable when the drug was too scarce for everyone.
- Dramatic decisions: even mastectomy versus lumpectomy has essentially been settled by random allocation.
- Blocking: adapted from agricultural experiments to balance treatments across soil, skill, or time.
- Hernia trial: blocked patient sequences balanced open and keyhole surgery as skills improved.
- Association is not explanation
- From Randomization to Causal Reasoning (CHAPTER 4 What Causes What? · II)
- Randomized Trials as the Gold Standard
- Randomized trials: gold standard for testing treatments, now applied to education and policing
- Study supporter trial: students with nominated text-message supporters had 27% higher pass rates
- Prayer trial: STEP's 1,800 bypass patients showed no benefit; aware group suffered more complications
- A/B testing: web experiments with huge samples reliably detect small, profitable effects
- When Randomization Is Impossible
- Observational data: arises without experiments; requires good design plus healthy scepticism
- Old men's ears: PPDAC cycle frames the problem and guides study-design choices
- Cohort studies: prospective or retrospective designs trace whether ears grow with age
- Case-control study: compare survivors to matched deceased to test whether big ears predict longevity
- Confounders and Simpson's Paradox
- Confounder: common cause, like weather driving both ice-cream sales and drownings
- Adjustment: stratifying by confounder levels reveals the relationship within each level
- Simpson's paradox: Cambridge 1996 admitted more men overall, yet more women in every subject
- Waitrose claim: "store adds £36,000" is a correlation driven by wealthy locations, not causation
- Lurking Factors and Reverse Causation
- Reverse causation: non-drinkers die more because illness stops drinking, not abstinence
- Lurking factors: unmeasured common causes that trip up naïve conclusions from data
- Vaccination and autism: coincidental timing at similar ages, not a causal link
- Left-handed seniors: scarce because past children were forced to switch, not because they die young
- Pope longevity: candidates are selected from those who already survived to old age
- Bradford Hill and Causal Guidelines
- Bradford Hill criteria: seven guidelines for judging causation from observational evidence
- Direct evidence: large effects, cause preceding effect, dose responsiveness and reversibility
- Mechanistic evidence: plausible causal chain, as aspirin's acidity explains mouth ulcers
- Parallel evidence: fits known science, replication, and similar studies
- Forensic epidemiology: courts infer individual causation when relative risk exceeds two
- Mendelian randomization: genes randomly assigned at conception act as a natural experiment
- Randomized Trials as the Gold Standard
- From correlation to controlled trials (CHAPTER 4 What Causes What? · I)
- CHAPTER 5 Modelling Relationships Using Regression
- Regression, Models, and Misattribution (CHAPTER 5 Modelling Relationships Using Regression · I)
- From Galton to Regression to the Mean
- Galton's question: predict an adult offspring's height from parents' heights using family data.
- Least squares line: choose the line minimizing the sum of squared residuals; developed by Legendre and Gauss.
- Regression to the mean: tall fathers tend to have shorter sons; short fathers tend to have taller sons.
- Origin of the term: Galton called the pattern "regression to mediocrity"; later any line-fitting became regression.
- Anatomy of a Regression Model
- Dependent variable: quantity to predict or explain, plotted on the y-axis; also called response variable.
- Independent variable: quantity used for prediction or explanation, plotted on the x-axis.
- Regression coefficient: the gradient indicates expected change in the dependent variable for a one-unit difference in the independent variable.
- Gradient and correlation: when standard deviations are equal, the gradient equals the Pearson correlation coefficient.
- Causal interpretation is stronger: requires experimentation or formal causal modelling, not just observational data.
- Statistical Models as Signal Plus Noise
- Model definition: a simplified mathematical representation of some aspect of the world.
- Two components: deterministic formula plus residual error.
- Residual error: not a mistake, but the inevitable mismatch between model and observation.
- Core equation: observation = deterministic model + residual error; classic signal-and-noise idea.
- Regression to the Mean in Real Life
- Before-and-after comparisons are treacherous: we over-attribute changes to intervention.
- Speed cameras: placed after accidents; later declines may simply be reversion to the mean.
- Sacked managers: successors get credit for a return to normal after a run of bad luck.
- Fund managers and Sports Illustrated cover: post-peak declines reflect luck, not a curse or lost skill.
- League tables: good teams tend to decline next year; bad teams improve — beware crediting new methods.
- From Galton to Regression to the Mean
- Regression, Adjustment, and Model Limits (CHAPTER 5 Modelling Relationships Using Regression · II)
- Regression to the Mean
- League-table shifts: observed correlation −0.60 vs −0.71 expected by chance; ranking changes mostly noise.
- Clinical-trial improvements: much apparent placebo benefit may be regression-to-the-mean, since patients enrol while symptomatic.
- Speed cameras: randomised studies estimate about two-thirds of apparent benefit is regression-to-the-mean, not a genuine camera effect.
- Multiple Regression and Adjustment
- Multiple linear regression: predicts a response from several explanatory variables via least squares.
- Coefficient adjustment: mother’s coefficient drops from 0.33 to 0.30, father’s from 0.45 to 0.41, once the other parent is included.
- Confounder control: Swedish brain-tumour study adjusted for age, region, income and more, yet lurking factors remain possible.
- Randomised trials: allocation balances confounders; regression is still used to check for residual imbalances.
- Response Variables and Logistic Regression
- Non-continuous responses: proportions, counts and survival times each need their own form of regression.
- Logistic regression: fits a curve that can never predict survival below 0% or above 100%.
- Child heart surgery: each extra 100 operations predicts about 10% lower relative mortality; correlation does not prove causation.
- Policy controversy: the 2001 finding shaped unresolved disputes over how many UK hospitals should perform the surgery.
- Models: Maps, Not Territory
- Four modelling cultures: simple statistics, physical simulations, machine-learning black boxes, and causal econometrics.
- Box’s aphorism: “All models are wrong, some are useful” warns against believing models too literally.
- Map analogy: a simple map works for motorways; a detailed map is needed for countryside walking.
- Financial crisis: over-trusted mortgage models assumed moderate correlations; failures rose together and risks were vastly underestimated.
- Regression to the Mean
- Regression, Models, and Misattribution (CHAPTER 5 Modelling Relationships Using Regression · I)
- CHAPTER 6 Algorithms, Analytics and Prediction
- From Science to Predictive Algorithms (CHAPTER 6 Algorithms, Analytics and Prediction · I)
- From Science to Algorithms
- From science to technology: algorithms aim at practical decisions, not at understanding how the world works.
- Classification vs prediction: both map current observations to a conclusion, whether present or future.
- Narrow AI: machine learning from vast examples powers speech, translation, vision, and game champions.
- General AI: absent; these systems are judged only on their limited task performance.
- Historical roots: statistical algorithms have aided decisions since Halley’s annuities in the 1690s.
- Big Data and Data Reduction
- Big data: scale grows in both cases (n) and features (p); genomics was small n, large p.
- Large n, large p: billions of posts, likes, and dislikes feed personalised ad algorithms.
- Clustering: unsupervised learning finds groups without being told they exist, then labels them for targeting.
- Feature engineering: reduces many raw measures to useful predictors, or forms composite measures.
- Deep learning: complex models may process raw data without initial feature reduction.
- Classification and Prediction Methods
- Method diversity: statisticians preferred regressions; computer scientists favoured rules or neural networks.
- Menu-driven software: makes performance, not modelling philosophy, the basis for choosing techniques.
- Kaggle contests: training sets build algorithms; held-out test sets decide winners.
- Data don’t speak: skills and care are still needed to avoid naïve-algorithm pitfalls.
- The Titanic Challenge
- PPDAC problem: though the Titanic won’t recur, survival prediction is classification, not future prediction.
- Francis Somerton: third-class passenger from Ilfracombe; algorithm can assess his slim survival chances.
- Data: 1,309 passengers; predictors include name, sex, age, class, fare, family, and port.
- Training/test split: 897 cases build the algorithm, 412 assess it; test set stays untouched.
- Pre-processing: missing fares and messy titles need judgement before modelling begins.
- Baselines: “nobody survived” scores 61%; “women survive, men don’t” scores 78%.
- Classification Trees and Performance
- Classification tree: a series of yes/no questions ends with a majority-outcome prediction.
- Somerton’s branch: as “Mr”, he is in a 58% group with only 16% survival.
- Tree predictions: women and children in first/second class survive, barring rare titles.
- Accuracy: tree scores 82% on training data and 81% on the held-out test set.
- Error matrix: sensitivity detects true survivors; specificity detects true non-survivors.
- Imperfect discrimination: every branch mixes survivors and non-survivors; confidence matters.
- From Science to Algorithms
- Probability, Calibration, and Over-fitting (CHAPTER 6 Algorithms, Analytics and Prediction · II)
- ROC Curves and Discrimination
- Probabilistic predictions: assign a survival probability rather than a blunt majority-vote category, e.g. “Mr” gets 16%.
- ROC curve: sensitivity vs false-positive rate across all thresholds; developed for WWII radar analysis.
- AUC summary: a random algorithm gives a diagonal curve with area 0.5; perfect discrimination gives 1.
- Titanic result: test-set AUC is 0.82 — 82% chance the algorithm ranks a true survivor above a true non-survivor.
- Probability Forecasts and Calibration
- Ensemble forecasts: weather models run from slightly adjusted initial conditions; rain in 5 of 50 futures → 10% probability.
- Calibration: probabilities should mean what they say, so around 70% of “70% rain” days should actually be wet.
- Calibration plots: predicted probability vs observed proportion; the diagonal marks perfectly reliable probabilities.
- 95% bars: vertical bars show expected sampling variation; diagonal within them signals a well-calibrated algorithm.
- Brier Score and Skill
- Brier score: mean squared error treating rain as 1 and no rain as 0; lower is better.
- Reference forecast: skill-less climate-based probabilities provide a baseline, e.g. a Brier score of 0.28 vs 0.11 for the forecast.
- Skill score: proportional reduction over the reference; 61% better than relying on climate alone.
- Titanic comparison: naive 39% survival rule scores 0.232; simple tree’s 0.139 is a 40% improvement.
- Over-fitting and the Bias/Variance Trade-off
- Over-fitting: adding tree branches raises training accuracy but worsens test-set Brier from 0.139 to 0.150.
- Fitting noise: too-complex algorithms chase idiosyncrasies in the training data rather than signal.
- Munroe’s presidents: plausible electoral rules such as “Catholics can’t win” are later broken by Kennedy.
- Bias/variance trade-off: matching ever more characteristics lowers bias but shrinks sample size, raising variance.
- Regularization: an alternative protection that permits complexity while pulling variable effects toward zero.
- Cross-validation and Model Tuning
- Independent test set: provides a final check but does not itself improve the algorithm.
- Tenfold cross-validation: remove 10%, train on 90%, repeat ten times, and average predictive performance.
- Complexity parameter: grow a deliberately over-deep tree, then prune using cross-validated performance.
- Full-data fit: the chosen complexity is applied to the complete training set; all models in the chapter are tuned this way.
- Regression Models
- Logistic regression: appropriate for a yes/no survival outcome, as with the child heart-surgery data.
- Fitted model: Table 6.3 presents logistic regression results, offering a compact formula alternative to classification trees.
- ROC Curves and Discrimination
- Algorithms, Prediction, and Accountability (CHAPTER 6 Algorithms, Analytics and Prediction · III)
- Logistic Regression Details
- Boosting: iterative procedure upweights misclassified training cases; iterations tuned by tenfold cross-validation.
- Coefficient scoring: passenger features add or subtract from survival score, converted to probability.
- Interactions: combined features can offset extreme individual coefficients, giving a more nuanced score.
- LASSO: estimates coefficients while selecting predictors, shrinking irrelevant ones to zero.
- More Complex Techniques
- Random forests: many trees bagged; final classification by majority vote.
- Support vector machines: seek linear feature combinations that best split outcomes.
- Neural networks: layered weighted nodes, essentially stacked logistic regressions; deep models have many layers.
- K-nearest-neighbour: assigns by majority outcome among close training cases.
- Performance and Simplicity
- Accuracy is not enough: the simple rule “all women survive” beats or nearly beats complex algorithms.
- No clear winner: random forest best ROC; simple tree best Brier; margins may be chance variation.
- Black boxes: complex winners resist implementation, explanation, and bias audit.
- Simplicity trade-off: once performance is good enough, prefer comprehensible algorithms.
- Challenges of Algorithms
- Lack of robustness: algorithms infer associations, not mechanisms; Google flu trends failed after search changes.
- Statistical variability: limited data make rankings unreliable; teacher value-added scores swing wildly year to year.
- Implicit bias: algorithms use proxies like postcodes; vision models learned snow, race, and other unintended cues.
- Lack of transparency: proprietary algorithms such as COMPAS hide weightings; reverse-engineering can reveal influences.
- Accountability: fairness demands transparent, controlled algorithms with rights of appeal.
- Predict as a Beneficial Algorithm
- Problem: after breast cancer surgery, predict survival benefits of adjuvant therapy.
- Design: use 5,700 registry cases but constrain treatment effects by clinical-trial estimates.
- Predict 2.1: open-source algorithm gives five- and ten-year survival proportions by treatment.
- Use: supports multidisciplinary meetings and shared care; caveat: estimates are ballpark guides.
- Artificial Intelligence and Causality
- Narrow AI: statistical learning powers achievements in vision, speech, and games; decision support becomes AI.
- Causal gap: algorithms answer “what next?”, not interventions or counterfactuals; Pearl urges causal models.
- More data, more care: big data reduces sample-size worry but worsens bias and generalization challenges.
- Humility: essential when building algorithms.
- Logistic Regression Details
- From Science to Predictive Algorithms (CHAPTER 6 Algorithms, Analytics and Prediction · I)
- CHAPTER 7 How Sure Can We Be About What Is Going On? Estimates and Intervals
- Real-World Estimates Hide Uncertainty
- Unemployment headlines: UK fall of 3,000 and US rise of 108,000 are sample estimates, not full counts.
- Hidden margins of error: published totals carried uncertainty intervals of ±77,000 and ±300,000.
- Media confusion: journalists treat survey estimates as exact population facts, blurring sample mean and population mean.
- Statistical duty: an estimate must be accompanied by a realistic assessment of its possible error.
- Sample Statistics Versus Population Parameters
- Statistic and parameter: sample summaries (m, s) approximate population values (μ, σ).
- Sample size effect: larger random samples produce statistics closer to the population truths.
- Natsal-3 example: mean lifetime partners for men 35–44 stabilizes toward 15 as sample size grows from 10 to 796.
- Bootstrapping: Resampling from the Sample
- Core idea: repeatedly draw new samples of the same size from the original data, replacing each point.
- No population assumptions: bootstrapping learns variability without assuming a mathematical shape for the population.
- Sampling distribution: histogram of 1,000 bootstrap means reveals how much the estimate would vary.
- Central Limit Theorem glimpse: bootstrap distributions turn symmetric and normal, regardless of original data skew.
- Quantifying Uncertainty with Intervals
- 95% uncertainty interval: the range containing 95% of bootstrap estimates, also called a margin of error.
- Larger samples, narrower intervals: bootstrap spreads shrink as sample size increases, so intervals tighten.
- Regression lines too: bootstrapping Galton’s 433 mother–daughter pairs gives a 95% interval for the gradient of 0.22 to 0.44.
- Limits of Bootstrapping
- Impractical for massive surveys: resampling 100,000 unemployment records is clumsy and unnecessary.
- Traditional theory next: probability theory provides formulae for uncertainty intervals, introduced in Chapter 9.
- Real-World Estimates Hide Uncertainty
- CHAPTER 8 Probability–the Language of Uncertainty and Variability
- The Language of Uncertainty and Variability (CHAPTER 8 Probability–the Language of Uncertainty and Variability · I)
- The Gambler's Puzzle and the Birth of Probability
- Chevalier de Méré compared winning on a six in four throws versus double-six in 24 throws.
- Simulation showed Game 1 wins about 52% of the time, Game 2 about 49%.
- Pascal and Fermat created formal probability theory from this gambling puzzle.
- Probability theory now underpins physics, insurance, pensions, trading, and forecasting.
- Expected Frequencies and Probability Trees
- Expected frequency asks what you would expect over many trials, making probability reasoning intuitive.
- Probability trees label each branch split with its fraction; multiply along branches for a sequence.
- Coin example gives two heads a probability of ½ × ½ = ¼.
- Conditional assumptions always apply: coin fairness, proper flipping, no external shocks.
- The Formal Rules of Probability
- Probability scale runs from 0 for impossible to 1 for certain events.
- Complement rule: probability of an event is 1 minus the probability it does not happen.
- OR rule: add probabilities of mutually exclusive events.
- AND rule: multiply probabilities of independent events in a sequence.
- Conditional Probability and the Prosecutor's Fallacy
- Conditional probability: with 1% prevalence and 90% accuracy, a positive mammogram means only an 8% chance of cancer.
- Expected frequency tree shows 108 positive mammograms per 1,000 women, only 9 truly have cancer.
- Prosecutor's fallacy confuses probability of evidence given innocence with probability of innocence given evidence.
- Reversal flaw is as invalid as inferring "if Catholic then Pope" from "if Pope then Catholic".
- The Meaning of Probability
- Measurement problem: there is no probability-ometer; probability is virtual, never directly measured.
- Philosophical disagreement: experts accept the math but dispute what probability numbers mean.
- Practical relevance: interpretations shape how probability is applied in statistical inference.
- The Gambler's Puzzle and the Birth of Probability
- Probability Interpretations and Acting as Random (CHAPTER 8 Probability–the Language of Uncertainty and Variability · II)
- Five Interpretations of Probability
- Classical probability: ratio of favourable outcomes to equally likely total outcomes; definition is circular because “equally likely” is assumed.
- Enumerative probability: counts chances from a physical set, such as drawing 3 white socks from 7; extends classical ideas to random choices from populations.
- Long-run frequency: proportion of successes in an infinite sequence of identical experiments; impractical for unique events like tomorrow’s weather.
- Propensity or chance: supposed objective tendency of a situation to produce an event; attractive but offers no basis for humans to estimate a metaphysical “true chance.”
- Subjective probability: personal judgement on a specific occasion, expressed through reasonable betting odds; the author’s preferred view and foundation of Bayesian inference.
- Why Probability Works Even Without Random Devices
- Three situations: randomness may generate data, select pre-existing data, or be absent entirely—yet we often act as if it were present.
- Random variable: maps possible outcomes to quantities, giving a probability distribution; observation collapses many potential futures into one actual value.
- Pseudo-random numbers: logically deterministic yet complex enough to be indistinguishable from true randomness, supporting the “as if” stance.
- Metaphorical population: includes all possible eventualities that might have occurred but did not; we deliberately treat observed data as drawn from it.
- Poisson Distribution and Homicide Clusters
- Homicide example: 1,545 incidents across 1,095 days average 1.41 per day; no observed day reached seven incidents.
- Poisson distribution: models events with a huge number of opportunities but tiny individual chances; depends only on its mean.
- Validation: Poisson predicted 12.4 days with exactly five homicides; the actual count was 13, showing an almost disturbingly close fit.
- Cluster probability: seven or more homicides in a day has probability 0.07%, or about once every four years—unlikely but not impossible.
- Key insight: individually unpredictable tragedies aggregate into data that behave as if produced by a known random mechanism.
- Regularity Within Unpredictability
- Quetelet: Belgian astronomer and sociologist who saw normal distributions in natural phenomena and imagined “l’homme moyen.”
- Social physics: unstable individual lives combine into stable national statistics, just as random molecules produce predictable gas properties.
- Natural variability: “chance” is a practical label for unavoidable unpredictability in birth weights, surgeries, exams, and homicides alike.
- Boundary: probability now covers both pure randomness and natural variability, setting the stage for formal statistical inference.
- Five Interpretations of Probability
- The Language of Uncertainty and Variability (CHAPTER 8 Probability–the Language of Uncertainty and Variability · I)
- CHAPTER 9 Putting Probability and Statistics Together
- Sampling Variation and Statistical Inference (CHAPTER 9 Putting Probability and Statistics Together · I)
- Statistics as Random Variables
- Fundamental step: sample statistics are random variables from their own distributions, not just single data points.
- Formulae vs bootstrap: theory gives insight and convenience; simulation is clumsy for large or complex problems.
- Known to unknown: first map how known populations produce sample variation; then reverse to infer populations.
- Assumptions matter: algebra can deceive; theory is valid only if its conditions hold.
- Binomial Sampling Distributions
- Binomial distribution: governs probabilities of observed proportions from random samples of a known population.
- Expectation: sampling distributions center on the population proportion (0.2) and tighten as sample size grows.
- Tail areas: probabilities such as “at least 30% left-handers” come from adding the distribution’s upper bars.
- Standard error: the standard deviation of a statistic, measuring sampling variability and shrinking with n.
- Funnel Plots and Chance Variation
- Bowel cancer headline: an apparent threefold UK variation in death rates proved to be expected chance variability.
- Funnel plot: rates plotted against population size, with control limits from an assumed binomial risk.
- Control limits: 95% and 99.8% bounds; only Glasgow City clearly falls outside the expected range.
- Small districts: extreme rates arise from few cases—Rossendale's rate was based on 7 deaths.
- Data journalism: open data still requires statistical principles to avoid misleading patterns.
- Follow-up: new data raised fresh questions, restarting the problem-solving cycle.
- Central Limit Theorem
- Law of Large Numbers: sample proportions converge to the true probability through accumulating Bernoulli trials.
- Gambler's fallacy: the coin has no memory; past imbalances are overwhelmed, not compensated, by independent flips.
- Central Limit Theorem: averages from virtually any population distribution become normal when samples are large.
- Normal approximation: mean matches original mean; standard error follows from population standard deviation.
- Galton's marvel: the normal law is a “supreme law of Unreason,” order latent in chaos.
- From Sample to Unknown Population
- Reversing inference: move from a single sample to the unknown population, not from known population to samples.
- Inductive inference: the path from sample to population, introduced in Chapter 3, underlies estimating accuracy.
- Coin illustration: probability stays 50:50 after a hidden flip; even another person's view doesn't change your information.
- Information dependence: probability statements reflect what the observer knows, not just an inherent state.
- Statistics as Random Variables
- Turning Probability into Confidence (CHAPTER 9 Putting Probability and Statistics Together · II)
- Aleatory vs Epistemic Uncertainty
- Aleatory uncertainty: chance before observing a random event like a coin flip or lottery.
- Epistemic uncertainty: personal ignorance about a fixed but unknown outcome, as with a covered coin or scratch card.
- Parameter: fixed but unknown population quantity; statistics estimate it via random sampling.
- Probability theory: tells what to expect in future; statistical inference flips it to extract past learnings.
- Confidence Intervals: Basic Principle
- Sampling distributions: probability theory predicts intervals where an observed statistic should lie for each parameter value.
- 95% prediction interval: range for a statistic before data collection, with 95% probability under a hypothesized parameter.
- 95% confidence interval: after observing statistic, range of parameter values for which the statistic lies in the prediction interval.
- Interpretation trap: a confidence interval is a procedure result; 95% of such intervals contain the true value, not probability for this interval.
- Origin: formalized by Jerzy Neyman and Egon Pearson at University College London in the 1930s.
- Calculating Confidence Intervals
- Bootstrap vs exact: bootstrapped 95% intervals from Chapter 7 nearly match exact intervals from standard software.
- Central Limit Theorem: makes large-sample estimates near-normal, so normal-based intervals are acceptable.
- Convention: 95% interval is usually estimate ± two standard errors; 80% or 99% intervals are sometimes used.
- Institutional choices differ: US Bureau of Labor Statistics uses 90% intervals; UK ONS uses 95%—always state which.
- Margins of Error in Surveys
- Rule of thumb: margin of error for a percentage is at most ±100 divided by the square root of the sample size.
- Example: 1,000 respondents give about ±3%; 40% in sample means the true percentage lies roughly 37–43%.
- Assumptions matter: valid margin requires random sampling, full response, and truthful opinions—rarely all achieved.
- Real polls exceed margins: 2017 UK election polls based on ~1,000 respondents varied far more than ±3%.
- Skeptical heuristic: doubling any quoted poll margin of error is a personal heuristic for systematic polling errors.
- Metrology splits error: Type A statistical error shrinks with more observations; Type B systematic error needs external judgment.
- When All Data Are Observed
- Complete counts: a fully accurate homicide count has no margin of error; uncertainty concerns the underlying true rate.
- Poisson model: yearly homicide count acts as one Poisson observation with mean m; its standard error is √m.
- Example: 497 homicides give a 95% confidence interval of roughly 453–541 for the true annual rate.
- Overlap criterion: overlapping intervals don’t prove no change; better to use a confidence interval for the difference.
- Change inference: a rise of 60 homicides has 95% confidence interval −4 to +124; includes zero, so no confident real change.
- Two interval meanings: unemployment intervals express epistemic uncertainty about actual count; homicide intervals express uncertainty about underlying risk.
- Aleatory vs Epistemic Uncertainty
- Sampling Variation and Statistical Inference (CHAPTER 9 Putting Probability and Statistics Together · I)
- CHAPTER 10 Answering Questions and Claiming Discoveries
- Hypothesis Testing and P-Values (CHAPTER 10 Answering Questions and Claiming Discoveries · I)
- Arbuthnot’s Test of Divine Providence
- Sex ratio: London baptisms 1629–1710, 82 years, more boys each year, ratio ~107.
- Null reasoning: 82 consecutive male excesses akin 82 heads in a row, odds 1/2^82.
- Claim: first recorded significance test; used data as evidence for divine providence.
- Modern view: natural sex ratio ~105; supernatural conclusion not justified.
- What Is a Hypothesis?
- Definition: provisional working assumption about a model’s deterministic or stochastic component, not truth.
- Null hypothesis: simplified model assumed until evidence against it; notation H₀.
- Relentlessly negative: denies progress and change; can be disproved, never proved.
- Legal analogy: guilty vs not proven; failing to reject is not accepting truth.
- Why Formal Testing?
- Apophenia: humans see patterns where none exist; false discoveries undermine science.
- Protection: hypothesis testing guards against mistaking random noise for discovery.
- Specific questions: from homicides to Higgs boson—all need a null hypothesis to test.
- P-Values and Permutation Tests
- Arm-crossing example: 54 students; 7% difference in right-arm proportions between genders.
- Permutation test: randomly relabel arm-crossing 1,000 times; null distribution centred at zero.
- P-value: probability of observing a result at least as extreme, if null hypothesis were true.
- One-tailed vs two-tailed: 0.45 for female excess; 0.89 for either direction.
- First two-sided P-value: Arbuthnot’s 1/2^81 for either sex exceeding in all years.
- Statistical Significance and Fisher
- Significance threshold: small P-value means statistically significant; Fisher popularized P < 0.05 and 0.01.
- Tea-tasting test: Muriel Bristol guessed all eight cups; hypergeometric probability 1/70.
- Steps of NHST: set H₀, choose test statistic, derive null distribution, compute P-value, declare significance.
- Conditionality: P-values depend on null hypothesis and all model assumptions, e.g. no bias, independent observations.
- Using Probability Theory
- Chi-squared test: compares observed vs expected counts under the no-association null.
- Arm-crossing example: expected female left-armers 5.7; chi-squared 0.02, P = 0.90.
- Convenience: theoretical approximations match exact permutation tests when available.
- Critique: formula-centric courses gave statistics a reputation for table-lookup.
- Arbuthnot’s Test of Divine Providence
- Significance, Power, and False Discoveries (CHAPTER 10 Answering Questions and Claiming Discoveries · II)
- Tests and Confidence Intervals Are One
- Goodness-of-fit: homicide counts fit a Poisson distribution with P=0.96; no evidence against the null.
- Unemployment example: quarterly change −3,000 had 95% CI −80,000 to +74,000; containing zero meant not significant.
- Equivalence: a two-sided P<0.05 corresponds exactly to a 95% confidence interval excluding the null value.
- Confidence interval as accepted nulls: the interval is the set of null hypotheses not rejected at P<0.05.
- Non-significance ≠ null true: an interval containing zero only says the data cannot rule out zero.
- Statistical Significance in Practice
- Statin trial: HPS confidence intervals exclude zero; P for 27% heart-attack reduction is about 1 in 3 million.
- Summary choice: HPS emphasizes proportional reduction because it stays fairly constant across subgroups.
- t-statistic: estimate divided by standard error; Galton height coefficients give vanishingly small P-values.
- Prediction competitions: winning tree's Brier margin over neural network was not significant (t = −0.54, P = 0.59).
- Significance Hunting and Multiple Testing
- Significance hunting: running many tests and reporting the most extreme invites false discoveries.
- False-positive inflation: 5% per true-null test; ten useless trials give about 40% chance of at least one P<0.05.
- Salmon cautionary tale: dead salmon fMRI showed 16 significant brain sites out of 8,064, pure chance.
- Bonferroni correction: demand P < 0.05/n; genetics uses a 1 in 20 million threshold.
- False discovery rate: softer methods control the proportion of claimed discoveries that are false.
- Independent replication: FDA requires two trials at P<0.05, reducing approval error to 1 in 400.
- Five Sigma and the Higgs Boson
- Five sigma: Higgs hump lay five standard errors from null; P below 1 in 3.5 million.
- Physics sigmas: reported σ values are the same as t-statistics.
- Look-elsewhere effect: CERN corrected for searching all energy levels—multiple testing under another name.
- Replication intent: CERN waited for extreme P and repeatable evidence before declaring discovery.
- New baseline: the established boson becomes the null hypothesis for future, deeper theories.
- Designing Studies: Neyman–Pearson
- Two traditions: Fisher's P-value measures evidence; Neyman-Pearson added alternatives and accept/reject rules; practice mixes both.
- Type I and II errors: rejecting a true null vs missing a real effect; like convicting the innocent vs acquitting the guilty.
- Size and power: α is the chosen significance threshold; power is the chance to detect a true effect.
- Trade-off: easing the significance threshold raises power but increases Type I errors.
- Sample-size planning: HPS set α=1% and 90% power to detect 25% mortality reduction, requiring 20,000 patients.
- Underpowered fields: psychology/neuroscience samples as low as 20 per condition can miss true effects.
- Tests and Confidence Intervals Are One
- Shipman, P-Values, and False Discoveries (CHAPTER 10 Answering Questions and Claiming Discoveries · III)
- Shipman: detecting excess mortality
- Excess mortality: observed minus expected death certificates, controlling for local conditions like flu and temperature.
- Shipman's toll: by 1998, statistical analysis predicted 174 women and 49 men aged 65+ over expected.
- Statistical accuracy: the estimate matched confirmed victims almost exactly, without using individual case data.
- Naive monitoring: a yearly P-value test would have flagged Shipman as significant as early as 1979.
- Multiple-testing trap: with ~25,000 GPs, one in 20 innocent doctors would look “significant” by chance.
- Sequential testing and the SPRT
- Bonferroni correction: demanding P < 0.05/25,000 would have alerted on Shipman in 1984.
- Repeated testing flaw: the Law of the Iterated Logarithm means continued tests eventually reject a true null hypothesis.
- SPRT: developed for wartime quality control, it monitors accumulating evidence against simple alert thresholds.
- Error control: thresholds bound overall Type I and Type II error rates while evidence accumulates.
- Shipman application: the SPRT crossed the million-to-one threshold in 1985 for women, and in 1984 combined.
- Lives saved: routine monitoring could have led to prosecution around 1984, saving roughly 175 lives.
- Outlier caveat: statistics detect unusual outcomes but not reasons; a flagged GP was caring for elderly patients in retirement homes.
- The replication problem and P-value critiques
- P-value in practice: Fisher imagined one test on one dataset; research now publishes countless P-values.
- False discovery arithmetic: with 1,000 tests, 80% power, and 10% true effects, 36% of discoveries are false positives.
- Ioannidis's claim: most published research findings are false, due to weak studies and publication bias.
- Institutional backlash: a psychology journal banned null-hypothesis significance testing; the ASA issued six principles.
- Six ASA principles for P-values
- Compatibility: P-values indicate how incompatible data are with a specified statistical model.
- Not posterior probability: P-values do not measure the probability the hypothesis is true; five-sigma reports often commit the prosecutor's fallacy.
- No threshold alone: decisions should not rest only on whether a P-value passes a threshold; a non-significant result is not evidence of zero effect.
- Transparency: full reporting is required, including how many tests were done and avoiding selective reporting.
- Effect size: statistical significance does not measure practical importance; a 19% relative risk increase meant 5 vs 6 tumours per 3,000 men.
- Evidence strength: a P-value near 0.05 is weak evidence; lowering the threshold to 0.005 would cut false discoveries from 36% to 5%.
- Shipman: detecting excess mortality
- Hypothesis Testing and P-Values (CHAPTER 10 Answering Questions and Claiming Discoveries · I)
- CHAPTER 11 Learning from Experience the Bayesian Way
- Probability, Evidence, and Belief Updating (CHAPTER 11 Learning from Experience the Bayesian Way · I)
- Epistemic Uncertainty and Subjective Probability
- Aleatory uncertainty: randomness in future events; epistemic uncertainty: fixed facts unknown to us.
- Bayes’ first contribution: probability expresses ignorance, not just random chance.
- Bayesian probabilities are subjective: they depend on our knowledge and must change with new evidence.
- Data never speaks for itself: background knowledge and judgement enter through a formal mathematical route.
- Contested legacy: many statisticians reject subjective judgement as unscientific.
- Bayes’ Theorem and Reversing the Tree
- Bayes’ theorem: a probability rule for continuously revising beliefs from experience.
- Expected frequency trees: make Bayesian updating intuitive by counting what happens in repeated trials.
- Three-coin puzzle: seeing heads makes it 2/3 likely you picked the two-headed coin, not 1/2.
- Doping paradox: a 95% accurate test with 1 in 50 prevalence means positive results are only 28% true.
- Reversing the tree: restructures outcomes by evidence first, then truth — the essence of inverse probability.
- Inverse fallacy: P(evidence | hypothesis) is not P(hypothesis | evidence); this confuses P-values and prosecutors’ arguments.
- Odds and Likelihood Ratios
- Odds: probability of an event divided by probability of it not happening.
- Likelihood ratio: compares evidence under two hypotheses: P(evidence | A) / P(evidence | B).
- Bayes’ theorem in odds form: prior odds × likelihood ratio = posterior odds.
- Doping example: prior odds 1/49 times likelihood ratio 19 gives posterior probability 28%.
- Repeated updating: posterior odds become prior odds; independent likelihood ratios multiply.
- Likelihood Ratios and Forensic Science
- Richard III prior: cautious assumptions gave only 1 in 400 chance the first skeleton was the king.
- Individual findings: dating, sex, scoliosis, mutilation, and DNA each gave modest likelihood ratios.
- Combined evidence: independent likelihood ratios multiply to 6.7 million, yielding posterior odds around 17,000 to 1.
- DNA evidence: likelihood ratios often reach millions or billions via the random match probability.
- UK courts: Bayes’ theorem is essentially prohibited; individual likelihood ratios are allowed but combining is left to the jury.
- Archbishop’s royal flush: 72,000 likelihood ratio is outweighed by prior odds of a million to one, leaving 7% suspicion.
- Bayesian Statistical Inference
- From hypotheses to parameters: Bayesian inference handles unknown quantities that can take many values.
- Bayes’ original question: given past successes, what probability should we assign to the next occurrence?
- Scientific updating: Bayes’ theorem is the coherent way to change our minds from evidence.
- Third approach: a Bayesian alternative to Fisher and Neyman–Pearson inference, only prominent in recent decades.
- Epistemic Uncertainty and Subjective Probability
- Bayesian Learning from Prior to Practice (CHAPTER 11 Learning from Experience the Bayesian Way · II)
- Bayes’ Billiard Table: Prior, Likelihood, Posterior
- Thumbtack opener: Bayes’ 16/22 estimate differs from frequentist 15/20 by incorporating prior knowledge.
- Billiard metaphor: hidden line from a random white ball; red-ball counts update beliefs about its position.
- Bayesian estimate: (left + 1)/(total + 2), e.g. 2/5 becomes 3/7, never 0 or 1.
- Shrinkage: estimates are pulled toward the prior centre, here 1/2, making sparse-data answers more sensible.
- Core mechanics: prior + likelihood → posterior distribution; prior encodes physical knowledge or judgement.
- No true prior: use sensitivity checks across plausible prior distributions.
- Hierarchical Models: MRP and Election Polling
- Online panels are non-random, so local estimates are unreliable; MRP pools strength across cells.
- Cells define homogeneous voter groups by demographics and location, with background counts from census data.
- Bayesian assumption: area coefficients drawn from a common prior distribution, smoothing local estimates toward each other.
- 2016 US election: MRP correctly called 50 of 51 states plus DC using only 9,485 interviews.
- 2017 UK election: YouGov predicted a hung parliament and Conservatives at 42%, matching the result.
- Caveat: systematic misrepresentation by respondents cannot be fixed by statistical modelling.
- Bayesian Brains and Self-Driving Cars
- Bayesian Brain: perception starts from prior expectations and updates only with unexpected features.
- Self-driving cars maintain probabilistic mental maps, continually updated by sensors.
- Bayesian smoothing also models disease spread over space and time.
- Bayes Factors for Scientific Hypotheses
- Bayes’ theorem: posterior odds = likelihood ratio × prior odds; Bayes factors provide the likelihood-ratio component.
- Prior dependence: unknown parameters require averaging over priors, making priors central to the result.
- Kass–Raftery scale: Bayes factors above 150 are very strong; P = 0.05 corresponds to 2.4–3.4, weak evidence.
- Legal contrast: courts require likelihood ratios of 10,000 for very strong evidence, reflecting “beyond reasonable doubt.”
- Symmetric support: Bayes factors can actively favour the null, unlike conventional significance tests.
- Higgs boson: P = 1/3.5 million converts to Bayes factor ~80,000; prior odds 1 gives posterior odds 80,000 to 1.
- The Ideological Battle Between Bayesians and Frequentists
- Frequentist vs Bayesian: long-run sampling properties versus posterior distributions from priors and data.
- Bayesian 95% uncertainty interval states the probability the true value lies inside; frequentist confidence interval does not.
- Historical clash: Bowley called Neyman’s confidence intervals a “confidence trick”; Fisher then fought Neyman and Pearson.
- Modern compromise: Neyman–Pearson design with Fisherian P-value analysis; Cornfield called it solid but lacking a firm logical foundation.
- Practical priority: statistical failures come from bad design, biased data, and poor scientific practice, not philosophy.
- Ecumenical truce: methods are chosen by context, not ideology.
- Bayes’ Billiard Table: Prior, Likelihood, Posterior
- Probability, Evidence, and Belief Updating (CHAPTER 11 Learning from Experience the Bayesian Way · I)
- CHAPTER 12 How Things Go Wrong
- Statistical Failure and Reproducibility Crisis (CHAPTER 12 How Things Go Wrong · I)
- An Opening Failure: The ESP Study
- Bem's precognition design: positions were randomized after each choice, so the null success rate was 50%.
- Bem's reported result: 53% correct choices with erotic images (P=0.01); eight of nine experiments showed significant precognition.
- The opening challenge: a prominent paper's striking claims frame the chapter's inquiry into statistical fallibility.
- The Reproducibility Crisis
- Ioannidis's thesis: most published research findings are false, a claim that has spread across medicine and social science.
- Reproducibility Project: only 36% of 100 replicated psychology studies were significant versus 97% of originals.
- Gelman's warning: the difference between significant and non-significant is not itself statistically significant.
- Replication effects: typically in the same direction but about half the original magnitude; only 23% of replications differed significantly.
- Regression to the null: early exaggerated estimates shrink toward the null, driven by pressure to publish discoveries.
- Breakdowns in Problem, Planning, and Data
- Problem stage: some questions cannot be answered from observed data, as with falling UK teenage pregnancy rates.
- Planning flaws: convenience samples, leading questions, and unfair comparisons corrupt the evidence.
- Design flaws: low power, missing confounders, and lack of blinding weaken studies from the start.
- Data stage: missing responses, dropouts, slow recruitment, and coding failures betray poor piloting.
- Fisher's warning: consulting a statistician after an experiment is a post-mortem; he can say what it died of.
- Analysis and Conclusions Go Wrong
- Computational errors: a spreadsheet mistake left five countries out of Reinhart-Rogoff's influential austerity analysis.
- Model errors: AXA Rosenberg's mis-coded statistical model understated risk by a factor of 10,000; SEC fine total $242 million.
- Wrong unit of analysis: cluster-randomized trials cannot be analyzed as if individuals had been randomized.
- Testing interaction: baseline changes within groups do not prove a difference; formally test whether groups differ.
- Non-significant is not zero: confidence intervals often show no real difference despite one P-value being significant.
- Selective reporting: airing only significant results inflates false positives; Harkonen's subset claim led to a wire-fraud conviction.
- Fraud and Its Detection
- Fabrication rarity: about 2% of scientists admit falsifying data, and detected cases are surely an underestimate.
- Chocolate hoax: Bohannon's fake institute, tiny samples, and selective reporting passed peer review and made headlines.
- Simonsohn's method: impossible similarities in reported standard deviations exposed fabricated randomized data.
- Cyril Burt: IQ twin correlations stayed almost identical as sample size grew, leaving suspected fraud unresolved.
- An Opening Failure: The ESP Study
- Questionable Research and Communication Practices (CHAPTER 12 How Things Go Wrong · II)
- Questionable Research Practices
- Researcher degrees of freedom: minor tweaks in design, stopping, exclusions, and analyses inflate statistical significance.
- Exploratory vs confirmatory: flexible exploration is fine; confirmatory studies need a pre-specified, preferably public, protocol.
- P-hacking: multiple testing plus subtle data-dependent choices to force P < 0.05.
- HARKing: inventing Hypotheses After the Results are Known blurs the line between exploration and confirmation.
- A Deliberate Demonstration
- When I'm Sixty-Four study: Simonsohn manufactured a Beatles-age link through flexible analyses after 34 subjects.
- Method: kept enrolling participants, tweaking comparisons and adjustments until a significant association emerged.
- Selective reporting: only the significant analysis was reported before the tweaks were revealed at the end.
- Purpose: a classic demonstration of how QRPs can create false findings.
- How Common Are These Practices?
- Survey of 2,155 psychologists: 94% acknowledged at least one questionable research practice.
- Common admissions: 67% failed to report all responses; 58% collected more data after seeing results.
- HARKing admitted by 35%; outright data falsification was rare at 2%.
- Authors saw practices as defensible: but they undermine any claim to be confirming a hypothesis.
- The Pipeline to the Public
- Communication chain: data → authorities → press offices → journalists → headlines → public; distortion occurs at every stage.
- File-drawer effect: unpublished null or unflattering studies leave the published literature positively biased.
- Press office exaggeration: 40% of UK university press releases contained exaggerated advice; 33% had causal overclaims.
- Framing trick: a 10% protective gene became the negatively-framed headline “nine in ten carry a risk gene.”
- Questionable Media Practices
- Journalists depend on press releases: headline writers and editors control framing, so blame is often misplaced.
- Spice-up tactics: eleven practices include ignoring uncertainty, implying causation, and using vivid graphics.
- Relative risk manipulation: TV box-set headline quoted a relative risk of 2.5; absolute rate was 13 per 158,000 person-years.
- Novelty drives coverage: media favor unusual, exaggerated claims over solid statistical evidence.
- Bem’s Precognition Case
- Bem encouraged replication and supplied materials, but the journal refused to publish failed replications.
- His methods exploited degrees of freedom: designs were changed, and favorable groups were highlighted.
- Gelman’s critique: Bem’s P-values assumed analyses would be unchanged if data differed—false across his nine studies.
- Catalyst for soul-searching: the 2011 paper triggered the reproducibility debate in psychology and science.
- Questionable Research Practices
- Statistical Failure and Reproducibility Crisis (CHAPTER 12 How Things Go Wrong · I)
- CHAPTER 13 How We Can Do Statistics Better
- Shared Responsibility for Better Statistics
- Three groups must act: producers, communicators, and audiences all improve how statistics shape society.
- Reproducibility manifesto: pre-registration, replication, transparent reporting, and diversified peer review make science more reliable.
- Exploratory vs confirmatory distinction: flexibility in analysis is fine if the sequence of choices is clearly reported.
- Open Science Framework: promotes data-sharing and pre-registration to curb adapting hypotheses to data.
- Improving What Is Produced
- Pre-specification tension: fixed analyses can prove inappropriate once unexpected data arrive.
- Ovarian cancer screening trial: prespecified primary analysis was non-significant, while reasonable post hoc analyses showed mortality benefit.
- Design limitation: researchers failed to anticipate screening’s late effect, constraining the primary outcome.
- Media misreading: non-significant results were wrongly reported as proof that screening did not work.
- Improving Communication
- Narrative risk: standard story arcs tempt oversimplification; stories should reflect evidence’s strengths, weaknesses, and uncertainties.
- Nuance can engage: Christie Aschwanden’s breast-screening story respected evidence while showing personal values drive decisions.
- Data journalism potential: statistician–journalist collaborations enrich storytelling, but need reporting guidelines and training.
- Researchable questions: how to communicate uncertainty without losing trust, tailored to different audiences.
- Calling Out Poor Practice
- Referees’ role: demand robust, replicated results but tolerate imperfection to encourage honest reporting.
- Suspicious signals: implausibly large effects in small samples and selective reporting of many comparisons raise red flags.
- Attractive-parents study: statistically significant but biologically implausible; external knowledge exposed it.
- P-curve test: significant P-values uniformly spread 0–0.05 hint at bias; clusters near 0.05 suggest massaging.
- Choice-overload effect: P-curve analysis found publication bias and no good evidence for the claimed effect.
- Ten Questions for Statistical Claims
- Assess the numbers: check study rigor, statistical uncertainty, and appropriateness of summaries.
- Assess the source: consider bias, conflicts of interest, spin, and exaggerated headlines.
- Assess the interpretation: fit with prior knowledge, causal explanations, relevance, and practical importance.
- Most important question: “What am I not being told?” exposes cherry-picking and missing context.
- No simple checklist: judging trustworthiness demands experience, scepticism, and intelligent transparency.
- Data Ethics and Trustworthiness
- Trustworthiness standards: honesty, competence, and reliability, demonstrated through transparent practice.
- Intelligent transparency: claims should be accessible, intelligible, assessable, and useable.
- Growing discipline: data ethics covers algorithmic fairness, privacy, consent, and honest science; training must include it.
- Exemplary Statistical Science: 2017 UK Exit Poll
- Problem: predict seat totals within minutes of polls closing, before official counts are known.
- Plan: exit polls at 144 carefully selected polling stations, with voters asked how they voted now and last time.
- Analysis: estimated swings at sampled stations, then used demographics and regression to predict each constituency—MRP.
- Results: predicted seat counts missed by at most four seats; in 2015 they forecast the Liberal Democrat collapse.
- Lesson: meticulous attention across the whole problem-solving cycle produced powerful, credible predictions.
- Shared Responsibility for Better Statistics
- CHAPTER 14 In Conclusion
- Frame the Question
- Ten rules: senior statisticians summarize self-evident non-technical issues rarely taught in courses
- Question first: statistical methods should let data answer scientific questions, not chosen techniques
- Plan ahead: pre-specify confirmatory analyses to avoid researcher degrees of freedom
- Data quality: everything rests on the data—worry about it
- Embrace Uncertainty
- Signal vs noise: variability is inevitable; use probability models to separate the two
- Assess variability: provide margins of error, remembering they are usually bigger than claimed
- Check assumptions: test assumptions and state clearly when checking was impossible
- Think, Don't Just Compute
- Beyond computation: statistical analysis is more than computations; know why for each step
- Keep it simple: communicate plainly; avoid complex modelling unless truly necessary
- Verify and Share
- Replicate: encourage replication whenever possible
- Make reproducible: others should be able to access your data and code
- Why It Matters
- Statistics in society: the field constantly evolves with increasing quantity and depth of data
- Personal enrichment: engaging with statistics can enrich individual lives
- Frame the Question
- Discover More
- Promotional Front Matter
- Discovery page: a marketing call-to-action offering sneak peeks, recommendations, and author news — no substantive book content.
- Promotional Front Matter
- Introduction
- Core Conclusion and Practical Takeaways
- Core Statistical Mindset
- Question first: let the scientific problem determine data collection and analysis, not a preferred technique.
- PPDAC cycle: Problem, Plan, Data, Analysis, Conclusion; repeat as new questions emerge.
- Data are constructed: definitions, proxies, and measurement choices shape what any statistic means.
- Signal versus noise: variability is inevitable; probability models separate real patterns from chance.
- Cause requires counterfactuals: correlation is not causation; fair trials randomize and control confounders.
- Practical Analysis Habits
- Look before summarizing: visualize distributions and relationships; summaries can hide crucial patterns.
- Match summary to shape: use median and IQR for skewed data, mean and SD for symmetric data.
- Quantify uncertainty: report intervals or margins of error alongside every estimate.
- Prefer absolute risks: communicate absolute frequencies to general audiences; reserve odds ratios for technical contexts.
- Check models as maps: models are useful simplifications, not truth; validate assumptions and avoid over-trust.
- Avoiding Statistical Pitfalls
- Beware multiple testing: many comparisons inflate false positives; correct, pre-register, or replicate.
- Avoid p-hacking and HARKing: do not tweak analyses to force significance or invent hypotheses after results.
- Adjust for confounders: common causes and Simpson's paradox can reverse apparent relationships.
- Expect regression to the mean: extreme results normalize; before-and-after comparisons can overstate effects.
- Watch selection bias: non-response, convenience samples, and missing data undermine representativeness.
- Communicating and Questioning Claims
- Know the audience: listen first, use plain language, and frame numbers around decisions.
- Use honest visuals: bar charts beat pie charts; axes and framing can exaggerate or conceal differences.
- State limitations clearly: report design, uncertainty, and what cannot be concluded.
- Tell accurate stories: narrative should reflect evidence strength, not oversimplify into false certainty.
- Ask what is missing: the key critical question is "What am I not being told?"
- Better Science and Bayesian Updates
- Replicate findings: robust conclusions need independent repetition, not one significant result.
- Pre-specify confirmatory work: register protocols and separate exploration from confirmation.
- Share data and code: reproducible research lets others check and extend results.
- Update from evidence: Bayesian inference combines prior knowledge and likelihood into posterior belief.
- Choose methods by context: frequentist and Bayesian tools answer different inferential questions.
- Core Statistical Mindset
opening map…