- General Overview
- Central Thesis: Thinking Under Uncertainty
- Statistical thinking: simplify a complex world while tracking uncertainty about conclusions.
- Not mere procedures: methods are tools for describing, deciding, and predicting from data.
- Uncertainty irreducible: evidence stays tentative; statistics rarely proves, it estimates.
- Learning by doing: real, messy data and computation teach statistical reasoning.
- Reproducibility crisis: misuse of testing and flexible analyses motivate critique and reform.
- Working With Data: Types and Measurement
- Qualitative vs quantitative: labels describe qualities; numbers measure quantities and are not labels.
- Measurement levels: nominal, ordinal, interval, ratio scales constrain which statistics are legitimate.
- Reliability vs validity: consistency of measurement differs from measuring the intended construct.
- Measurement error: better instruments or averaging reduce error, but no measure is perfect.
- Summarizing and Visualizing
- Aggregation: summarizing is deliberately discarding data to reveal structure and support generalization.
- Distributions: frequency tables, histograms, normal curves, skew, and long tails characterize data.
- Idealized models: the normal distribution and central limit theorem explain recurring shapes.
- Visualization: show raw data, maximize data/ink, avoid chartjunk, and respect perceptual limits.
- Axis honesty: lie factor and baseline choices can distort or clarify the same data.
- Models: Fit, Error, and Simplicity
- data = model + error: statistics splits observations into expected pattern and residual deviation.
- Mean as model: central tendency is a fitted model; mean minimizes squared error, median resists outliers.
- Variance and df: sample variability corrects for estimated parameters using n−1.
- Z-scores: standardizing places variables on a common scale for comparison.
- Overfitting: complex models capture noise and generalize worse; simplicity trades fit for prediction.
- Probability, Sampling, and Simulation
- Probability axioms: experiments, sample spaces, and events ground all statistical inference.
- Conditional probability: Bayes’ theorem reverses known conditionals and exposes base-rate neglect.
- Sampling: representative samples estimate parameters; sampling error and standard error quantify uncertainty.
- Central limit theorem: sample means become normal as n grows, even from skewed data.
- Resampling: Monte Carlo and the bootstrap replace analytic formulas with repeated computation.
- Hypothesis Testing and Effect Size
- NHST: tests ask how likely data are under a null, not how likely the hypothesis is.
- Errors and power: Type I, Type II, alpha, beta, and power govern decisions.
- p-value limits: significance is not practical importance, replication, or proof of the alternative.
- Multiple testing: familywise error inflates false positives; Bonferroni and Bayes factors address it.
- Effect size: Cohen’s d, correlation, and odds ratios quantify magnitude beyond significance.
- Bayesian and Multivariate Approaches
- Bayesian inference: priors and likelihoods yield posteriors, credible intervals, and Bayes factors.
- Evidence for null: Bayes factors can support the null, unlike frequentist tests.
- Categorical modeling: chi-squared tests independence; residuals and odds ratios locate effects.
- Continuous relationships: correlation, regression, causation, and Simpson’s paradox require care.
- Multivariate methods: clustering, PCA, and factor analysis discover structure without a target.
- General linear model: regression, t tests, ANOVA, and extensions share one framework.
- Practical Modeling and Reproducible Research
- Workflow: specify, prepare, visualize, choose, fit, diagnose, then test and quantify effects.
- Assumptions and diagnostics: normality, independence, outliers, confounders, and colliders shape validity.
- Mixed-effects models: random effects handle clustered and repeated observations.
- Reproducibility crisis: fraud, QRPs, p-hacking, HARKing, and winner’s curse distort literature.
- Reform tools: preregistration, replication, open data, and scripted analyses restore trust.
- Ethical duty: analyses embed values; answer nature’s questions truthfully, not just seek significance.
- Central Thesis: Thinking Under Uncertainty
- Deep Dive
- Preface��������������
- Statistical Thinking as the Core Subject
- Beyond tool training: most introductory books teach procedures; this one teaches a way of thinking about data.
- Statistical thinking defined: a systematic approach to describing the world and making decisions under uncertainty.
- Uncertainty is inherent: real-world conclusions always carry irreducible doubt, not just measurement error.
- Descriptive and predictive goals: the same framework serves both summarizing the world and forecasting it.
- Computational Power and Simulation
- New methods, newly feasible: analyses that took years in the 1950s now run in seconds on a laptop.
- Simulation as inquiry: cheap computation lets researchers ask questions that were previously unanswerable.
- Curricular shift implied: statistical methods should be taught in light of what modern computers enable.
- A Response to the Reproducibility Crisis
- The crisis context: since 2010, many scientific fields have faced failures to replicate published findings.
- Hypothesis testing implicated: misuse and abuse of significance testing is a root cause of replication failure.
- Statistical education connects directly: how testing is taught shapes how researchers use it.
- A critical stance: current methods are examined for problems, and alternatives are proposed.
- Capstone critique: the book's final chapter details the abuses of hypothesis testing.
- Learning by Doing with Real Data
- Learning requires doing: statistics is only truly learned by performing analyses, not reading about them.
- Code your own analyses: open-source languages replaced point-and-click software, deepening understanding.
- R and Python companions: two free online resources accompany the text, plus figure-recreation examples.
- The golden age of data: governments, scientists, and companies now share vast open datasets freely.
- Real data over idealized examples: genuine datasets prepare students for actual practice.
- Data need wrangling: real data rarely arrive analysis-ready and must be reshaped first.
- Assumptions often break: real distributions—such as Facebook friend counts—can have long tails.
- Acknowledged limitation: datasets are heavily US-centric, relying on NHANES and FiveThirtyEight.
- Statistical Thinking as the Core Subject
- 1. Introduction
- Statistical Thinking Versus Intuition
- Statistical thinking: describing a complex world in simple terms that capture its essential structure — plus our uncertainty about it.
- Interdisciplinary roots: grounded chiefly in mathematics and statistics, but also in computer science and psychology.
- Heuristics: intuition relies on mental shortcuts that often answer data questions wrongly.
- Availability heuristic: we judge how common something is by how easily an example comes to mind.
- Crime paradox: Americans believed violent crime was rising while data showed steady decline — coverage, not reality, drove the perception.
- Dealing with Statistics Anxiety
- Anxiety is normal: in a recent class, two-thirds of students rated themselves nervous about statistics; a quarter strongly agreed.
- Arousal can help: emotional arousal focuses attention and can improve performance on many tasks.
- Reframe it: knowing others feel the same can turn nervousness into a tool that cements the material.
- What Statistics Can Do
- Describe: condense a dataset too complex to grasp into a few numbers that reveal its structure.
- Decide: judge whether a pattern exceeds what chance would produce — often mistaken for statistics' whole purpose.
- Predict: generalize beyond the sample to new situations; a study limited to its own subjects is useless.
- PURE study: 135,335 people in 18 countries over a median 7.4 years, compressed by quintiles into 10 numbers.
- Relative risk: quintile death rates compared with the lowest-intake group; above one means higher risk.
- Surprise: higher saturated fat tracked with lower mortality, higher carbohydrate with higher mortality.
- Learning from Data
- Hypotheses first: every analysis begins from prior ideas about what might be the case.
- Prior knowledge: expectations vary in strength — weak for a first restaurant visit, strong after ten.
- Updating: statistics describes how new data should revise belief; psychology's learning theories echo this.
- Machine learning: a field at the statistics–computer science interface, building algorithms that learn from experience.
- Two cultures: Breiman's phrase for how differently statisticians and machine learners approach the same problems; this book blends them.
- Aggregation, Uncertainty, and Sampling
- Aggregation: "the science of throwing away data" — condensing 100,000 numbers into ten, revolutionary when first proposed.
- Caution: summaries can go too far and paint a misleading picture of the data behind them.
- Uncertainty: causation is probabilistic — lifelong heavy smokers face roughly 15% lung cancer risk, yet many never develop it.
- No proof: statistics yields evidence that is always tentative, never the certainty of a mathematical proof.
- Sampling: a small sample can characterize an entire population if obtained in the right way.
- Diminishing returns: accuracy grows with the square root of sample size — quadrupling a sample only doubles precision.
- Causality and Statistics
- Correlation is a hint: Tufte's amendment to the maxim — association alone never establishes causation.
- Confounds: the PURE result fits equally well a third factor, such as wealth, raising both fat intake and longevity.
- Randomized controlled trial: randomly assign participants to treatment or control, then compare outcomes.
- Why randomization works: it removes systematic group differences; residual imbalances still occur and are hard to adjust away.
- The verdict: RCTs consistently found no appreciable effect of reducing saturated fat on death rates.
- Statistical Thinking Versus Intuition
- 2. Working with Data
- Qualitative versus Quantitative Data
- Qualitative variables: describe a quality, not a numeric quantity; numbers assigned are labels, not values.
- Averaging labels is meaningless: computing a mean of food preferences or zip codes makes no sense.
- Quantitative data: numerical measurements; counting qualitative responses yields a quantitative summary.
- Types of Numbers and Measurements
- Binary and logical: zero/one or FALSE/TRUE representing presence, absence, or truth.
- Integers: whole numbers without fractional parts, used for counting and Likert-scale responses.
- Real numbers: carry decimals, measurable to arbitrary precision, like weight in micrograms.
- Discrete measurements: finite set of distinct values with no middle ground, like number of friends.
- Continuous measurements: any real value in a range, though tools cap the precision actually recorded.
- Measuring Unobservable Constructs
- Construct: an unobservable theoretical concept, such as knowledge or extroversion, we try to measure.
- Measurement error is inevitable: reduce it with better instruments or by averaging many measurements.
- Gold standard: the most accurate reference measure, often costlier and swapped for cheaper, error-prone alternatives.
- Reliability: Consistency of Measurement
- Reliability: how consistently a measurement reproduces itself across repetitions.
- Test-retest reliability: the same measurement performed twice should yield similar results.
- Interrater reliability: subjective ratings should not depend on which rater performed them.
- Reliability caps relationships: a measure's correlation with anything cannot exceed its correlation with itself.
- Validity: Measuring What We Intend
- Validity: whether the measurement captures the construct it claims to capture; reliability alone is insufficient.
- Face validity: does the measure make sense on its surface, like a blood-pressure cuff versus tongue color?
- Construct validity: convergent measures of the same construct should correlate; divergent measures of different constructs should not.
- Predictive validity: valid measures should forecast related real-world outcomes, like sensation seeking predicting risk taking.
- Scales of Measurement
- Four properties: identity, magnitude, equal intervals, and absolute zero define how values relate.
- Four scales: nominal (labels only), ordinal (ordered), interval (equal intervals, no true zero), ratio (all four properties).
- Scale constrains operations: nominal allows only equality; ordinal adds ordering; interval adds arithmetic; ratio adds multiplication and division.
- Scale constrains statistics: mode works anywhere, median needs ordinal, mean needs interval — though researchers often ignore this.
- Qualitative versus Quantitative Data
- 3. Summarizing Data
- Why Summarize Data
- Summarizing discards information: deliberately throwing data away is how statistics builds understanding.
- Generalization: summaries let us make statements that extend beyond specific observations.
- Borges' Funes: "To think is to forget a difference, to generalize, to abstract."
- Categorization: we recognize an ostrich, robin, and chicken as "birds" despite surface differences.
- Prediction: generalization yields useful, if imperfect, expectations about new individuals.
- Tables and Frequency Distributions
- Frequency distribution: a table counting how often each possible value occurs.
- Absolute frequency: raw counts per value, but hard for judging imbalances.
- Relative frequency: each count divided by the total; multiplied by 100 it becomes a percentage.
- Missing data matters: why values are missing shapes whether analysis is safe.
- Biased missingness: if missingness relates to the true value, dropping NAs distorts results.
- Ancient tool: frequency tables have summarized data for thousands of years.
- Cumulative Distributions and Histograms
- Histogram: plots the number of cases per value; the density version plots relative frequencies.
- Cumulative frequency: counts observations at or below a value, summed up to that point.
- Monotonic increase: a cumulative curve can only rise or stay flat, never fall.
- Binning: continuous values are grouped into ranges; bins = range ÷ bin width.
- Bin width is a judgment call: no fixed rule; trial and error, or methods like Freedman–Diaconis.
- Check provenance: NHANES caps age at 80 to protect identity, creating a false spike.
- Idealized Distributions: The Normal
- Recurring shapes: every dataset differs, yet patterns repeat, permitting idealized models.
- Normal (Gaussian) distribution: symmetric with a rounded central peak, as in adult height and pulse.
- Two parameters: the mean locates the center; the standard deviation sets the width.
- Two numbers suffice: a normal curve characterizes a distribution surprisingly well.
- Deep reason: the central limit theorem explains why many real variables take this form.
- Skewness and Long Tails
- Skewness: asymmetry where one tail is denser than the other.
- Right skew: common for counts and elapsed times bounded at zero, like SFO security wait times.
- Left skew is rarer, arising when values are capped above.
- Long-tailed distributions: extreme right tails far beyond typical members — Facebook friends average 24, maximum 1043.
- Matthew effect: "the rich get richer" — compounding advantage generates long tails.
- Tools can fail: many methods assume normality; Taleb's The Black Swan links this to the 2008 crisis.
- Why Summarize Data
- 4. Data Visualization
- Why Visualization Matters
- Challenger disaster: engineers' handwritten slides failed to persuade NASA; a proper plot of O-ring damage against temperature might have postponed the launch.
- Presentation is evidence: the same data can persuade or fail depending on how it is displayed.
- Anatomy of a Plot
- Axes: x-axis horizontal, y-axis vertical; values plotted may be continuous or categorical.
- Bar graph: shows group means but hides spread and overlap between distributions.
- Beeswarm plot: overlays individual points on bars, revealing distribution shape and overlap.
- Violin plot: displays each condition's smoothed distribution.
- Box plot: median, interquartile range, and individually marked outliers.
- Preference: techniques revealing the full distribution beat means-only summaries.
- Show the Data and Make It Stand Out
- Show raw data: plotting only a fitted line hides what the underlying data actually look like.
- Too-perfect patterns: real data are messy; suspiciously tidy points signal fabrication.
- Outliers: a single extreme point can manufacture an apparent relationship.
- Golden rule: always inspect the raw data before trusting any summary.
- Data/Ink Ratio and Chartjunk
- Data/ink ratio: ink devoted to data divided by total ink; higher is better.
- Chartjunk: decorations thematically related to content but unrelated to data—avoid at all costs.
- Common offenders: three-dimensional bars, background textures, and clip art distort or distract.
- Spreadsheet defaults: popular chart tools invite chartjunk; resist their templates.
- Distorting the Data: Axis Choice and Lie Factor
- Axis scaling: the same crime data can appear flat or plummeting depending on the y-axis range.
- Huff vs. Tufte: Huff insists on always including zero; Tufte prefers a baseline that shows the data without wasted space.
- When zero is absurd: body temperature never reaches zero, so including it wastes space and hides real change.
- Lie factor: physical difference in the graphic relative to the difference in the data; near 1 is honest.
- Bar charts: should nearly always include zero, since truncated bars exaggerate area differences.
- Practical rule: for line and scatter plots, use the full space unless zero genuinely matters.
- Accommodating Human Limitations
- Color blindness: never rely on color alone; vary brightness and texture too.
- Pie charts: humans poorly judge volume and angle; perspective distorts; legends force memory lookups.
- Bars on a common scale: length comparisons are the most accurate perceptual judgment available.
- Correct for confounders: adjust for inflation, population size, or season—unadjusted trends mislead.
- Trade-off: the clearer bar chart beats the flashier pie chart in accuracy and readability.
- Why Visualization Matters
- 5. Fitting Models to Data
- Models, Error, and Generalization (5. Fitting Models to Data · I)
- Models as Simplified Descriptions
- Model: simplified fiction capturing data structure; “All models are wrong but some are useful.”
- Data = model + error: splits observed values into expected pattern and residual deviation.
- Prediction: model estimates data as ŷ; error = observed y − predicted ŷ.
- Parameters: tunable values (β) estimated from data; hats denote estimates.
- Summarizing with a Single Number
- Mode: most common value; simple but poor height predictor, RMSE ~39 cm.
- Mean: arithmetic average; guarantees average error zero, RMSE ~27 cm.
- Zero average error: positive and negative errors cancel, so not a good fit criterion.
- Squared error: counts positive/negative errors as bad; basis for MSE and RMSE.
- Measuring Model Error
- Sum of squared errors: total squared residuals; depends on sample size.
- MSE: mean squared error; in squared units, hard to interpret directly.
- RMSE: square root of MSE; returns error to original units, such as centimeters.
- Model comparison: lower error generally better, but not the sole criterion.
- Adding Variables and Intercepts
- Age model: height rises with age; strong relationship in NHANES.
- No-intercept trap: ŷ = β·age forces zero height at age zero, worsening RMSE.
- Intercept: β0 lets model match overall magnitude even if zero is implausible.
- Age + constant: dramatically reduces RMSE to 8.36 cm.
- Gender: adds slight improvement, mainly after puberty; diminishing returns.
- Error Sources and Overfitting
- Model misspecification: wrong form or missing variables increases error.
- Noise/measurement error: random variation persists even when model is correct.
- Overfitting: complex model fits current data better but generalizes worse.
- Generalization: aim for low error on new datasets, not just training data.
- Bias-variance tension: fit versus generalizability often conflict; simpler model can win.
- Models as Simplified Descriptions
- Overfitting, Central Tendency, and Standardized Scores (5. Fitting Models to Data · II)
- Overfitting and the Cost of Complexity
- Overfitting: a complex model captures random noise, so it fits new data from the same source worse
- Simplicity principle: make models as simple as possible, but not simpler
- Fit trade-off: fit must be good, but "too good" is a warning sign, not a virtue
- The Mean as a Fitted Model
- Mean as model: computing an average is fitting a model to data, not just arithmetic
- Sample vs population: X̄ and μ use identical math but different symbols, marking statistic versus parameter
- SSE minimization: the mean is the estimator that minimizes the sum of squared errors
- Outlier sensitivity: adding one extreme income makes the mean represent nobody in the room
- Robust Alternatives: Median and Mode
- Median: the middle sorted value; it minimizes absolute errors, so outliers barely move it
- Central tendency trade-off: the mean varies less across samples; the median resists extreme values
- Mode: the most common value — the only sensible summary for categorical data like phone models
- Variability: Variance and Degrees of Freedom
- Variance and standard deviation: squared errors per observation; the square root restores original units
- Sample variance: dividing by n−1 instead of N corrects the bias in estimating population variance
- Degrees of freedom: estimating the mean leaves one data value no longer free to vary
- Bias consequence: underestimating variance makes statistical decisions overconfident
- Simulation: resampling a real dataset confirms n−1 estimates track the true population value
- Z-Scores: Expressing Position in a Distribution
- Z-score: how far a data point sits from the mean, measured in standard deviations
- Crime example: raw counts make California look dangerous; per-capita rates place it near average
- Standard normal: about 16% of values fall beyond Z = ±1, only about 2.3% beyond ±2
- Standardized scores: rescaling Z-scores to a new mean and SD yields IQ-style scores
- Comparing Distributions on a Common Scale
- Common scale: Z-scoring different variables makes otherwise incomparable measures directly comparable
- Difference scores: standardized variables let you compute relative rates between variables across units
- Small-sample noise: extreme difference scores in small states reflect variability, not real causes
- Overfitting and the Cost of Complexity
- Error, Overfitting, and Standardized Scores (5. Fitting Models to Data · III)
- Model, Data, and Error
- Three parts: statistics always combines data, a model, and the error between them.
- Error versus mode: errors from the mode do not average to zero, so it is a weaker estimate.
- Two error sources: prediction error when a model meets data, plus variability inherent in the data.
- Why the Mean Minimizes Error
- Least squares: the mean is the estimate minimizing the sum of squared errors.
- Guaranteed ranking: squared error from the mean is always less than or equal to that from the median.
- Zero-sum proof: deviations from the mean sum to zero because they equal the total minus n times the mean.
- Outlier sensitivity: the mean shifts with extreme values; the median does not.
- Overfitting and Generalization
- Overfitting: a model fitted too closely to one dataset captures noise, not signal.
- Trade-off: the best-fitting model on one sample is generally not best on a new sample.
- Diagnosis: fit worsens on fresh data — relevance tied most directly to overfitting.
- Degrees of Freedom and Variability
- Hat notation: a hat marks a value estimated from data rather than a true parameter.
- Population versus sample: sample standard deviation divides by n−1, correcting for the estimated mean.
- Degrees of freedom: estimating parameters from data consumes independent information.
- Standardized Scores
- Z-score: (value − mean) divided by standard deviation, expressing position in units of variability.
- Comparability: Z-scores place different variables—violence and property crime—on one shared scale.
- Figure 5.15: symbol size encodes population, exposing how crime rates and city size covary.
- Model, Data, and Error
- Models, Error, and Generalization (5. Fitting Models to Data · I)
- 6. Probability
- Probability Foundations and Rules (6. Probability · I)
- Foundations: Experiments, Outcomes, Axioms
- Experiment: any activity that produces or observes an outcome.
- Sample space: the set of all possible outcomes for an experiment.
- Event: a subset of the sample space; an elementary event has exactly one outcome.
- Kolmogorov axioms: probabilities are nonnegative, sum to one, and none exceeds one.
- Three Routes to Assigning Probability
- Personal belief: subjective knowledge or opinion when an experiment cannot be run.
- Empirical frequency: repeat an experiment and divide outcome count by total trials.
- Rain example: 73 rainy days in San Francisco 2017 gives P=73/365≈0.2.
- Classical probability: known mechanism with equally likely outcomes gives each outcome 1/n.
- Law of Large Numbers vs Small Samples
- Law of large numbers: empirical probability converges on true probability as sample size grows.
- Coin simulation: estimates of heads approach 0.5, but small samples can be far off.
- Law of small numbers: people wrongly expect small samples to behave like large ones.
- Alabama 2017 election: early precinct returns swung wildly before settling, showing small-sample volatility.
- Rules of Probability and De Méré's Dice
- Rule of subtraction: P(not A)=1−P(A), because A and not-A exhaust the sample space.
- Intersection of independents: P(A∩B)=P(A)×P(B) only when A and B are independent.
- Addition rule: P(A∪B)=P(A)+P(B)−P(A∩B); De Méré erred by ignoring the intersection.
- Pascal's complement trick: compute "no success" then subtract from one for "at least one."
- Dice outcomes: at least one six in four rolls ≈0.517; at least one double-six in 24 rolls ≈0.491.
- Probability Distributions
- Probability distribution: describes probabilities of all possible outcomes in an experiment.
- Binomial distribution: gives P(k successes in n Bernoulli trials) with success probability p.
- Binomial coefficient: "n-choose-k" counts ways to choose k items from n.
- Steph Curry example: with p=0.91, hitting only 2 of 4 free throws has probability ≈0.040.
- Cumulative distribution: gives probability of a value as extreme or more extreme than a target.
- Conditional Probability
- Conditional probability: P(A|B) is the probability of A given that B has occurred.
- Simple vs conditional: simple probability ignores prior information; conditional uses it.
- 2016 election example: P(Republican)=0.44 and P(Trump voter)=0.46 are simple probabilities.
- Joint probability needed: computing P(Trump|Republican) requires the probability of both events.
- Foundations: Experiments, Outcomes, Axioms
- Conditioning, Independence, and Belief Updating (6. Probability · II)
- Conditional Probability
- Definition: P(A|B) = P(A∩B)/P(B) — probability both hold given B is true
- Graphical view: conditioning on B narrows analysis to a subset of the data
- From data: combine simple and joint probabilities; P(diabetes|inactive) = 0.141
- Computation trick: logical true/false acts as 1/0, so an event's probability is the mean
- Joint tables: row proportions reveal conditionals; simple proportions show marginals
- Statistical Independence
- Definition: knowing one variable tells you nothing about another; P(A|B) = P(A)
- Common usage differs: everyday "independence" often means mutually exclusive, not statistical independence
- Jefferson example: Jeffersonians cannot be Californians, so the two are not statistically independent
- Testing: compare simple vs. conditional probability; NHANES shows 0.16 vs. 0.13 for bad mental health
- Caveat: differences may stem from sampling variability; formal tests appear in Chapter 12
- Reversing Conditional Probability: Bayes' Theorem
- Purpose: turn known P(A|B) into wanted P(B|A) — central to medical screening
- Sensitivity: P(positive test|disease); specificity: P(negative test|no disease)
- Bayes' theorem: P(B|A) = P(A|B)·P(B) / P(A), derived from the conditional-probability rule
- PSA example: with 80% sensitivity and 70% specificity, P(cancer|positive) is only 0.14
- Base rate neglect: people systematically ignore overall prevalence, producing surprising results
- Prosecutor's fallacy: confusing P(evidence|innocence) with P(innocence|evidence)
- Learning from Data
- Bayesian updating: Bayes' rule revises prior beliefs according to the evidence
- Prior: initial P(B) before data; likelihood: P(A|B) of data under the hypothesis
- Marginal likelihood: P(A), the data's overall probability across all hypotheses
- Posterior: P(B|A), the updated belief after seeing the data
- Update rule: multiply the prior by how much more likely the data are given B than overall
- Odds and Odds Ratios
- Odds: P(A)/P(¬A), the relative likelihood of an event versus its absence
- PSA odds: given a positive test, odds of cancer are 0.16; dice odds of a six are 0.2
- Odds ratio: posterior odds ÷ prior odds = 2.62, an effect size quantifying the shift
- Screening warning: for uncommon conditions, most positives are false positives, driving needless follow-up
- What Do Probabilities Mean?
- Frequentist: probabilities as long-run frequencies; ill-suited to one-time events
- Bayesian: probabilities as degree of belief, answering questions with no relevant frequencies
- Gamble framing: a probability equals the odds at which you would accept the bet
- Relevance: these interpretations ground the two approaches to hypothesis testing in later chapters
- Conditional Probability
- Probability Foundations and Rules (6. Probability · I)
- 7. Sampling
- Population, Sample, and Why We Sample
- Population vs. sample: infer a feature of an entire population from a small, carefully selected subset
- Parameter vs. statistic: the population parameter is unknown; the sample statistic estimates it
- Efficiency motive: measuring every individual wastes effort when a small sample suffices
- Nate Silver's polls: ~21,000 sampled voters predicted the behavior of over 125 million voters
- Representative Sampling
- Representative sample: every member of the population has an equal chance of being selected
- Bias: a nonrepresentative selection makes the statistic systematically differ from the parameter
- With replacement: sampled members return to the pool; without replacement: they become ineligible
- Sampling without replacement is standard; bootstrapping (chapter 8) uses sampling with replacement
- Sampling Error and Sampling Distribution
- Sampling error: the near-inevitable gap between a sample statistic and the population parameter
- Sampling distribution: the distribution of a statistic across many repeated samples
- NHANES example: population mean height 168.35 cm; samples of 50 vary slightly around it
- 5000 sample means of size 50 averaged 168.38—almost exactly the true population mean
- Unbiased statistic: its expected value equals the parameter, yet any single estimate still differs
- Standard Error of the Mean
- SEM = estimated standard deviation ÷ √n, the standard deviation of the sampling distribution of the mean
- Two determinants: population variability and sample size; only sample size is under our control
- Diminishing returns: quality improves with √n, so doubling the sample improves it only about 1.41×
- Population SEM (1.436) closely matched the observed spread of 5000 sample means (1.435)
- Estimating SEM from a small sample (under ~30) requires caution
- The Central Limit Theorem
- CLT: as sample size grows, the sampling distribution of the mean becomes normal, even from non-normal data
- The heavily skewed AlcoholYear variable yielded a near-normal sampling distribution at n = 50
- Why it matters: it licenses normal-based techniques and explains why normal data are so common
- Combining many independent factors into one number tends to produce a normal distribution, as with height
- Normal distribution: fixed bell shape, varying only in mean (location) and standard deviation (width)
- Population, Sample, and Why We Sample
- 8. Resampling and Simulation
- The Monte Carlo Method
- Origin: Ulam and Metropolis devised it during the Manhattan Project to trace neutron collisions.
- Casino analogy: estimate an outcome's probability by playing the game hundreds of times.
- Four steps: define a domain, draw random numbers, compute, then combine across many repetitions.
- Purpose: replace complicated analytic mathematics with straightforward repeated computation.
- Randomness in Statistics
- Random defined: a process is random when it is unpredictable, not necessarily undetermined.
- Deterministic in principle: coin flips obey physics; too many factors make them practically unpredictable.
- Pareidolia: humans perceive familiar patterns, such as faces in clouds, where none exist.
- Gambler's fallacy: we wrongly expect random processes to self-correct, feeling "due for a win."
- Generating Random Numbers
- Pseudorandom: software produces numbers that seem random, but the sequence eventually repeats.
- Ample cycle: R's generator repeats only after 2^19937 − 1 numbers, far beyond any practical need.
- Quantile function: inverts the CDF to map uniform draws onto any target distribution.
- Random seed: fixing it regenerates the identical sequence, making simulations reproducible.
- Simulation in Practice: An Example
- The question: how long should a quiz period be so students finish 99% of the time?
- The method: simulate 5000 quizzes, recording the distribution of maximum finishing times.
- The answer: 8.74 minutes, the 99th percentile of that simulated distribution.
- Assumptions govern: a wrong distributional model makes the simulation's answer useless.
- The Bootstrap
- Core idea: resample the data themselves to estimate a statistic's sampling distribution.
- With replacement: each resample draws with replacement, so data points often repeat.
- Bootstrap principle: treat your sample as the population and pull yourself up by your own bootstraps.
- Efron's method: recovers the standard error of the mean without assuming normality.
- Best use: statistics lacking a standard-error formula or violating normal-distribution assumptions.
- Limitation: presumes the sample mirrors the population, which breaks down for small samples.
- The Monte Carlo Method
- 9. Hypothesis Testing
- Null Hypothesis Statistical Testing (9. Hypothesis Testing · I)
- NHST as a Decision Framework
- NHST: backbone of psychological research, used widely but deeply flawed and widely misunderstood.
- Three statistical goals: describe, decide, predict—hypothesis testing serves the decide goal.
- Flipped question: NHST asks how likely data are under null, not how likely hypothesis is given data.
- Null assumption: operate as if null is true until data are sufficiently unlikely to reject it.
- Rejection logic: unlikely data under null license rejecting null in favor of alternative.
- Retention: insufficient evidence means retain/fail to reject null, not prove it true.
- Specifying Hypotheses Before Data
- Hypothesis first: state prediction before seeing data; never choose test after looking.
- Null hypothesis (H0): baseline equality/≤/≥, representing no effect or no difference.
- Alternative hypothesis (HA): inequality (≠, >, <), representing effect of interest.
- Directional test: predicts direction; more powerful only with strong prior reason.
- Nondirectional test: predicts any difference; generally preferred as more conservative.
- Preregistration: formalizes writing hypotheses before data to avoid bias.
- Steps and Test Statistic
- Six-step process: hypothesize, specify H0/HA, collect data, fit model, compute probability, assess significance.
- Blood pressure example: test whether physical activity relates to systolic blood pressure using NHANES.
- Test statistic: quantifies effect size relative to sampling variability; must have null distribution.
- t statistic: compares two means when population SD unknown and samples small.
- Welch t test: adjusts degrees of freedom for unequal variances and sample sizes, penalizing imbalance.
- t distribution: resembles normal but has wider tails when degrees of freedom are small.
- Computing p-values
- p-value question: probability of observed statistic or one more extreme, assuming null true.
- Continuous caveat: exact value probability is zero, so count at least as extreme outcomes.
- Coin-flip example: 70+ heads in 100 flips has p≈0.00004 under fair coin.
- Simulation: repeatedly generate null data to approximate p-value, e.g., 0.00002 by simulation.
- t-test p-value: t=1.89, df=246.41 gives one-tailed p=0.0299.
- Two-tailed p-value: double the one-tailed p when direction is not predicted; p=0.0597.
- Randomization and Resampling
- Randomization: generate null-like data by randomly rearranging observed data to make null true.
- Squat example: compare football players vs cross-country runners; H0: μFB ≤ μXC, HA: μFB > μXC.
- Null simulation logic: ask how extreme observed data are relative to simulated null data.
- Relation to bootstrapping: both use your own data, but randomization imposes null rather than resampling for uncertainty.
- t.test(): standard R function for t-test; in example shows large group difference.
- NHST as a Decision Framework
- Randomization, Thresholds, and Decision Errors (9. Hypothesis Testing · II)
- Randomization Tests: Shuffling the Null
- Core idea: if the null hypothesis is true, group labels carry no information and can be shuffled freely
- Procedure: randomly relabel group membership, recompute the statistic, repeat thousands of times
- Null distribution: shuffled t-values center on zero and roughly trace the theoretical t distribution
- p-value: the proportion of shuffled statistics at least as extreme as the observed one
- Two examples agree: squat data gave 0.0041 by randomization versus 0.0002 by t-test; blood pressure gave 0.028 versus 0.030
- Advantages: no normality assumption, and p-values become possible for statistics lacking a theoretical distribution
- Exchangeability: The One Assumption
- Exchangeability: all observations are distributed the same way, so labels can be interchanged without changing the distribution
- Where it breaks: related observations, such as siblings clustered within families, are not exchangeable
- Safeguard: data obtained by random sampling generally satisfies exchangeability across observations
- Step 6: Where to Draw the Significance Line
- Step 6: decide whether the p-value is small enough to reject the null hypothesis
- No correct answer: the threshold demands subjective judgment about how much evidence is required
- Fisher's .05: offered as a convenient conventional line, never intended as a fixed rule
- Ritual origin: pre-computer p-value tables all listed .05, turning a convention into a habit
- Modern proposal: lower the default to .005, because significance at .05 is relatively weak evidence
- Neyman–Pearson: Testing as Decision Making
- Reframing: tests are not evidence about a hypothesis but rules that govern long-run behavior
- Two realities, two decisions: H0 true or false, crossed with reject or retain
- Correct outcomes: rejecting a false H0 (a hit) and retaining a true H0 (a correct rejection)
- Type I error (false alarm): rejecting H0 when it is true; long-run probability is α, conventionally .05
- Type II error (miss): retaining H0 when it is false; long-run probability is β, conventionally .2
- What a Significant Result Does Not Mean
- Not P(H0): the p-value is the probability of the data given the null, never the probability of the null
- Not the error probability: it is not P(H0 | data), the chance that your decision is wrong
- Not a replication forecast: it says nothing about the likelihood of obtaining the same result again
- Statistical ≠ practical significance: a genuine effect can still be too small to matter
- Sample size effect: with about 65,000 subjects, a half-point blood-pressure drop is "significant" over 90% of the time
- Multiple Testing in Modern Science
- Modern scale: genome-wide and brain-imaging studies test millions of hypotheses at once
- Expected false positives: at α = .05, a million null tests yield roughly 50,000 seemingly significant results
- Familywise error: the per-test alpha does not control the error rate across the whole family of tests
- Bonferroni correction: divide α by the number of tests — 0.000005 for one million tests
- Result: the probability of making any error in the study is brought back to about .05
- Randomization Tests: Shuffling the Null
- Errors, Significance, and Multiplicity (9. Hypothesis Testing · III)
- Error Types and Their Probabilities
- Type I error: rejecting the null hypothesis when it is actually true
- Type II error: failing to reject the null hypothesis when it is actually false
- Alpha: the probability of committing a Type I error
- Beta: the probability of committing a Type II error
- Power: the probability of not making a Type II error, equal to 1 − beta
- Statistical vs Practical Significance
- Two separate questions: a result can be statistically significant yet practically trivial
- Broccoli and income: a p-value of .01 with a 4-cent average gap is statistical, not practical, significance
- Large samples: with 200 million observations, even negligible differences produce tiny p-values
- Effect size matters: significance testing alone cannot tell you whether a finding is worth acting on
- The t Distribution
- Wider tails: the t distribution has fatter tails than the normal distribution
- Sample size: as n increases, the t and normal distributions become more alike
- Convergence: t tails narrow toward normal tails as sample size grows — they do not widen
- Fixed normal: only the t distribution changes shape with sample size
- One-Tailed vs Two-Tailed Tests
- Two-tailed p-value: larger than the one-tailed p-value for the same observed statistic
- Reason: two-tailed tests count extreme values in both directions, doubling the tail area
- Choice up front: the test form should follow the hypothesis, not the desired p-value
- Reading a p-Value Correctly
- Correct interpretation: a .01 p-value means a 1% chance of a result at least as extreme, given a true null
- Not the null's probability: it is not the chance that the null hypothesis is true
- Not replicability: it is not the chance a repeat experiment fails to find significance
- Not proof: the alternative hypothesis is not proved true
- Not practical meaning: significance level says nothing about whether an effect is large enough to matter
- Familywise Error and Bonferroni
- Familywise error: testing many hypotheses inflates the chance of at least one false positive
- Control's purpose: keeps the overall Type I error rate at the intended level
- Bonferroni correction: divide alpha by the number of tests, or multiply each p-value by that number
- Worked example: with four tests, the threshold becomes .0125, so only p = .01 survives
- Error Types and Their Probabilities
- Null Hypothesis Statistical Testing (9. Hypothesis Testing · I)
- 10. Quantifying Effects and Designing Studies
- Confidence Intervals and Estimate Uncertainty (10. Quantifying Effects and Designing Studies · I)
- Point Estimates and Their Limits
- Point estimate: a single sample value offered as the best guess at an unknown population parameter
- Standard error: describes how much uncertainty surrounds that estimate
- Two components: standard error depends on population standard deviation and sample size
- Only sample size is controllable — larger samples shrink uncertainty, up to measuring the entire population
- Margin of error: political polls express this as an estimate ± a few percentage points
- What a Confidence Interval Really Means
- Neyman's caution: the parameter is an unknown constant; no probability statement about its value may be made
- Long-run interpretation: a 95% interval procedure captures the true parameter 95% of the time
- Not this interval: any specific interval either contains the fixed parameter or it does not
- Simulation check: across 100 NHANES samples, 95 intervals captured the known true mean
- Computing Intervals: Normal and t
- General form: confidence interval = point estimate ± critical value × standard error
- Normal case: when the population standard deviation is known, the critical value is ±1.96 for 95%
- t distribution: appropriate when the standard deviation is estimated, giving wider intervals
- Extra uncertainty: the broader t distribution accounts for estimating parameters from small samples
- Critical value: taken from the percentiles of the relevant sampling distribution
- Sample Size and Interval Width
- Narrowing bounds: bigger samples produce progressively tighter intervals around the estimate
- Diminishing returns: width shrinks only in proportion to the square root of sample size
- Full population: measuring everyone eliminates uncertainty entirely
- Bootstrap Confidence Intervals
- When assumptions fail: used when normality or the sampling distribution cannot be assumed
- Procedure: repeatedly resample the data with replacement and recompute the statistic
- Surrogate distribution: the resampled statistics stand in for the sampling distribution
- Percentile method: the NHANES bootstrap gave roughly (78, 84), close to the t-based interval
- Intervals Versus Hypothesis Tests
- Direct link: if the interval excludes the null value, the associated test is statistically significant
- Single mean: testing against zero reduces to checking whether zero lies inside the interval
- Clear case, no difference: if each mean falls inside the other's interval, there is no significant difference
- Clear case, difference: non-overlapping intervals imply a significant difference, conservatively
- Ambiguous case: overlapping intervals that miss the other mean have no general answer
- Avoid the eyeball test: never judge significance by visually comparing overlapping intervals
- Point Estimates and Their Limits
- Effect Sizes, Power, and Study Design (10. Quantifying Effects and Designing Studies · II)
- Effect Size as Standardized Magnitude
- Effect size: a standardized measure comparing an effect to a reference quantity such as data variability
- Signal-to-noise ratio: the same concept as named in science and engineering
- Magnitude over significance: ask not whether a treatment affects people but how much
- Many measures: the appropriate effect size depends on the nature of the data
- Cohen's d
- Cohen's d: difference between two means divided by their pooled standard deviation
- Sample-size independence: unlike the t statistic, d stays constant as the sample grows
- Interpretation scale: negligible below 0.2, small to 0.5, medium to 0.8, large above
- Height example: the gender difference in adult height, d = 2.05, is enormous yet still overlaps
- Overlap caveat: even a huge effect leaves individuals from each group resembling the other
- Rarity in science: such obvious effects need no research; very large reported effects often signal questionable practices
- Other Effect Size Measures
- Correlation coefficient r: strength of the linear relationship between two continuous variables, ranging −1 to 1
- Anchors: 1 is perfect positive, 0 no relationship, −1 perfect negative
- Odds: the relative likelihood of an event versus its absence, P(A)/P(¬A)
- Odds ratio: the ratio of two odds, the natural effect size for binary variables
- Smoking and lung cancer: odds ratio 23.22, smokers' cancer odds roughly 23 times never-smokers'
- Case-control caution: such designs reveal relationships but not prevalence in the general population
- Statistical Power
- Power: the likelihood of finding a positive result given that it exists, equal to 1 − β
- Type II error: β, the false negative rate whose complement is power
- Type I priority: false positives are feared most, so α is set low, usually 0.05
- Alternative hypothesis: you must specify the effect size you hope to detect, or β is uninterpretable
- Three levers: sample size, effect size, and Type I error rate jointly determine power
- Power Analysis and Study Design
- Power analysis: a tool for computing the needed sample size before a study begins
- Worked example: detecting d = 0.5 with 80% power at α = 0.05 requires 64 per group
- Large effects need few: an effect of d = 2 needs only about five subjects per group
- Futility: an underpowered study will likely find nothing even when a true effect exists
- False findings: positive results from underpowered studies are more likely to be false
- Effect Size as Standardized Magnitude
- Confidence Intervals, Power, and Effect Sizes (10. Quantifying Effects and Designing Studies · III)
- Reading Confidence Intervals Correctly
- Frequentist meaning: a 95% interval captures the true parameter in 95% of repeated samples
- Not a probability claim: no 95% chance the true mean sits inside one observed interval
- Not a forecast: the interval does not bound where a future sample mean will land
- Overlap as a test: non-overlapping 95% intervals for two means imply a difference at p < .05
- Common failure: even professional scientists routinely misread intervals
- Width and Spread Measures
- Sample size: the confidence interval narrows as the sample size grows
- Standard error: interval width scales with the standard error of the estimate
- Effect Sizes by Data Type
- Mean difference: Cohen's d divides by the standard deviation
- t statistic: divides by the standard error of the mean
- Continuous relationship: a correlation coefficient measures linear association between two continuous measures
- Binary difference: an odds ratio expresses the difference between two binary variables
- Statistical Power and Its Drivers
- Definition: the likelihood of detecting a true effect when it genuinely exists
- Misconception: power is not the likelihood of getting the right answer overall
- Sample size: power increases as the sample size increases
- Effect size: power increases as the true effect size increases
- Type I error rate: power increases as the Type I error rate increases
- Effect Sizes in Power Analysis
- Plausibility rule: specify an effect size that is scientifically interesting and supported by prior research
- Obvious effects: no study is needed to detect gaps as large as 16-year-olds versus 6-year-olds
- Winner's curse: published effect sizes are likely inflated relative to true effects
- Practical consequence: shifting the target from d = 0.5 to d = 0.2 requires a larger sample
- Reading Confidence Intervals Correctly
- Confidence Intervals and Estimate Uncertainty (10. Quantifying Effects and Designing Studies · I)
- 11. Bayesian Statistics
- Bayesian Inference and Generative Models (11. Bayesian Statistics · I)
- Generative Models and Latent Processes
- Generative model: a latent, unseen process produces observed data, usually with randomness.
- Inference goal: use observed data to learn the latent variable that generated them.
- Coin example: knowing P_heads = 0.5 lets you generate expected samples; usually you infer P_heads instead.
- Friend example: a silent greeting invites inference about unseen causes like distraction or anger.
- Sampling link: sample statistics estimate population parameters that exist in the latent world.
- Bayes’ Theorem as Inverse Inference
- Bayes’ theorem: inverts conditional probability to infer hypotheses from data.
- Prior P(H): degree of belief about a hypothesis before seeing data.
- Likelihood P(D|H): how likely the data are under a given hypothesis.
- Marginal likelihood P(D): overall data likelihood across all hypotheses weighted by priors.
- Posterior P(H|D): updated belief about the hypothesis after seeing data.
- Frequentist contrast: frequentists treat hypotheses as fixed and data as random; Bayesians assign probabilities to both.
- Bayesian Estimation: Airport Screening
- Screening question: find P(explosive | positive test), not P(positive test | explosive).
- Prior: FAA 2017 data implied ~1 explosive per 971 million passengers; post-9/11 staff used 1 per million.
- Data: three explosive screening tests all return positive.
- Likelihood: test sensitivity 0.99; three positives modeled as Bernoulli/binomial trials.
- Marginal likelihood: weighted average of positive-test likelihood with and without explosive, using specificity 0.99.
- Posterior: three positives yield just under 50% probability of explosive; rare events produce false positives; posterior can become next prior.
- Estimating Posterior Distributions
- Drug example: estimate proportion of patients for whom a new pain drug is effective.
- Uniform prior: no prior information means uniform distribution over 99 discrete values from .01 to .99; avoids calculus.
- Data and likelihood: 100 patients, 64 respond; binomial likelihood is higher under P_respond = 0.7 than 0.5 or 0.3.
- Bayesian principle: upweight parameter values in proportion to data likelihood, balanced against prior knowledge.
- Marginal likelihood: sum likelihoods across all hypotheses so posterior values remain true probabilities.
- Posterior distribution: computed across all possible values of P_respond.
- Maximum a Posteriori Estimation
- MAP estimate: the parameter value with the highest posterior probability.
- Point estimate: summarizes the posterior distribution for decisions about P_respond.
- Generative Models and Latent Processes
- Bayesian Estimation, Priors, and Evidence (11. Bayesian Statistics · II)
- Posterior Estimation and MAP
- MAP estimate: the peak of the posterior distribution, representing the most probable parameter value.
- Uniform prior leaves the data in charge: the MAP (0.64) is simply the sample proportion of responders.
- Posterior blends prior and likelihood, weighting each source by its relative strength.
- Credible Intervals
- Credible interval: an interval with a stated probability (e.g., 95%) that the true parameter falls inside it.
- Direct interpretation: unlike confidence intervals, it answers "where is the parameter?" plainly.
- Computation by sampling: draw from the posterior, then take quantiles — works without a closed form.
- Inference payoff: in the drug example, the interval confirms high confidence that the response rate exceeds zero.
- Effects of Different Priors
- Prior strength pulls the posterior: strong, concentrated priors shift estimates toward their center.
- Sequential updating: the posterior from one analysis becomes the prior for the next.
- Absolute priors dominate: setting prior probability to zero for some values forces zero posterior density there.
- Choosing a Prior
- Uninformative priors: attempt to influence the resulting posterior as little as possible, as with a uniform prior.
- Weakly informative (default) priors: nudge results only slightly, e.g., a binomial based on one head in two flips.
- Empirical priors: built from prior studies or the scientific literature.
- Practical default: prefer uninformative or weakly informative priors to minimize concerns about bias.
- Bayes Factors for Hypothesis Testing
- Bayes factor: ratio of how well two hypotheses predict the data, quantifying relative evidence.
- Senator example: a BF of 3325 strongly favors Smith's 40% claim over Jones's 60%.
- Statistical hypotheses need marginal likelihood: an integrated average over parameters weighted by priors.
- Trial example: a significant t-test yielded BF of 3.4 — statistically significant but weak evidence.
- One-Sided Tests and Interpreting Bayes Factors
- Directional testing: dividing the positive by the negative Bayes factor assesses one-sided evidence (~29 here).
- Kass–Raftery scale: 1–3 barely worth mentioning, 3–20 positive, 20–150 strong, >150 very strong.
- Beyond significance: Bayes factors quantify strength of evidence, not merely whether an effect exists.
- Posterior Estimation and MAP
- Weighing Evidence, Priors, and Sampling (11. Bayesian Statistics · III)
- Evidence Beyond Significance
- Significance ≠ strong evidence: a statistically significant result can rest on weak evidence for the alternative versus a point null
- Directional hypotheses: evidence for a one-sided hypothesis can be far stronger than for the point null
- Bayes factor's payoff: it quantifies how much evidence the data actually provide, not just whether an effect is detectable
- Evidence for the Null
- Frequentist blind spot: standard null hypothesis testing assumes the null is true, so it can never support the null
- Bayes factor advantage: comparing two hypotheses directly permits evidence in favor of the null
- Nonsignificant results: the Bayes factor reveals whether they mean a truly negative result or merely insufficient evidence
- Priors, Learning, and Inference
- Bayesian learning: prior beliefs are updated by data into a posterior, making inference a form of learning
- Dangerous priors: a sufficiently strong prior can overwhelm the data entirely
- Subjective priors: a legitimate source of bias, weighed against the frequentist ideal of objectivity
- Central contrast: frequentists estimate the probability of data under the null; Bayesians estimate the probability of a hypothesis given the data
- Appendix: Rejection Sampling
- Purpose: generate samples from a posterior distribution that cannot be sampled directly
- Algorithm: draw candidate values of x and y uniformly, and accept only when y < f(x)
- Result: the histogram of accepted samples approximates the true posterior density
- Credible interval: the 2.5th and 97.5th percentiles bound the 95% credible interval — 0.54 to 0.73 in the pain drug example
- Evidence Beyond Significance
- Bayesian Inference and Generative Models (11. Bayesian Statistics · I)
- 12. Modeling Categorical Relationships
- Categorical Data and Counts
- Categorical variables: relationships among qualitatively measured variables, modeled as counts per category or combination
- Null expectation: expected counts derived from hypothesized proportions, then compared to what was observed
- Candy example: 30 chocolates, 33 licorices, 37 gumballs against a stated equal split, 33 each
- The Chi-Squared Test
- Chi-squared statistic: sum over categories of (observed − expected)² ÷ expected
- Chi-squared distribution: the sum of squares of standard normal variables; shape depends on degrees of freedom
- Degrees of freedom: k − 1 for one variable, since estimating the mean costs one
- Candy result: χ² = 0.74, p = 0.691 — counts unsurprising, so retain the null of equal proportions
- Simulation check: sums of squares of random normals closely match the theoretical chi-squared distribution
- Contingency Tables and Independence
- Contingency table: counts or proportions for every combination of values of two categorical variables
- Statistical independence: P(X ∩ Y) = P(X)·P(Y), so expected cell counts are products of marginal probabilities
- Two-way degrees of freedom: (nRows − 1)·(nColumns − 1); a 2×2 table leaves only one value free
- Police search data: χ² = 828, p ≈ 3.79×10⁻¹⁸² — race and searches are clearly not independent
- Interpreting a Significant Result
- Standardized residuals: (observed − expected) ÷ √expected, interpretable as Z-scores per cell
- Direction of effect: searches of Black drivers far exceed expectation (+26.6); White drivers fall below (−10.4)
- Odds ratio: odds of being searched are 2.59 times higher for Black than White drivers
- Bayes factor: ratio of data likelihoods under alternative vs. null; 1.8×10¹⁴² is overwhelming evidence
- Complementary tools: residuals, odds ratios, and Bayes factors supply what the chi-squared test alone cannot
- Larger Tables and Simpson's Paradox
- Beyond 2×2: chi-squared and Bayes factor analyses extend to tables with many categories per variable
- NHANES example: depression level vs. sleep trouble, χ² = 191 with df = 2, Bayes factor 1.8×10³⁵
- Simpson's paradox: a pattern in combined data may not appear in any subset of that data
- Baseball case: Justice out-hit Jeter each year 1995–1997, yet Jeter's combined average was higher
- Lurking variable: unequal at-bats across years drove the reversal; always check for such confounders
- Categorical Data and Counts
- 13. Modeling Continuous Relationships
- Quantifying Association: Covariance and Correlation
- Covariance: measures whether two variables deviate from their means in the same or opposite directions.
- Correlation coefficient: scales covariance by both standard deviations, yielding a unitless measure.
- Bounded interpretation: r ranges from -1 to 1; 0 means no linear relationship.
- Toy example: covariance 17.05 becomes r = 0.89 after standardization.
- Testing Correlation and Its Uncertainty
- Null hypothesis: test H0: r = 0 using t = r sqrt(N-2)/sqrt(1-r^2), df = N-2.
- Hate crime example: r = 0.42, t = 3, p = .002; reject zero correlation.
- Confidence interval: wide 95% CI [0.16, 0.63] shows imprecision.
- Randomization test: shuffling one variable builds null distribution; p ≈ .007.
- Assumption: t-test assumes both variables normally distributed.
- Outliers and Robust Correlation
- Outlier sensitivity: a single outlier can flip correlation from -1 to 0.83.
- Spearman correlation: correlation on ranks reduces outlier influence.
- Hate crime reanalysis: Spearman rho = 0.033, p = .8, undermining original claim.
- Lesson: always inspect scatterplots and check robustness before trusting r.
- Correlation Versus Causation
- Causation as manipulation: X causes Y if experimentally manipulating X changes Y.
- Koch's postulates: causal organism present in disease, elimination cures, infection causes.
- Marshall example: self-infection with H. pylori demonstrated bacterial cause of ulcers.
- Confounding: association may reflect a third variable influencing both.
- Correlation's limit: it says something is probably causing something, not what causes what.
- Causal Graphs and Mediation
- Causal graph: circles are variables, arrows are causal relationships.
- Latent variable: unmeasured knowledge mediates study time's effects on grades and finishing time.
- Mediation: holding mediator constant should remove the effect of cause.
- Spurious causal reading: grades and finishing time correlate negatively, but faster finishing does not cause better grades.
- Caution: inferring causation without experiments requires strong assumptions.
- Appendix: Inequality and Bayesian Correlation
- Gini index: quantifies income inequality via Lorenz curve / relative mean absolute difference.
- Lorenz curve examples: perfect equality G = 0; high inequality G ≈ 0.89.
- Bayesian correlation: posterior probability r > 0 plus prior regularization.
- Hate crime Bayesian result: rho ≈ 0.38, BF > 20, still suggests positive correlation.
- But: Bayesian analysis is not robust to the outlier.
- Quantifying Association: Covariance and Correlation
- 14. The General Linear Model
- Regression, Prediction, and Model Interpretation (14. The General Linear Model · I)
- The General Linear Model Framework
- data = model + error: statistics' core equation; seek the model minimizing error under a simplicity constraint
- General linear model: expresses the dependent variable as a linear combination of independent variables
- Dependent vs. independent variables: Y is the outcome to explain; X variables do the explaining
- Beta weights (β): each predictor is multiplied by a weight setting its relative contribution
- Three fundamental activities: describe a relationship, decide whether it is significant, predict new values
- Scope: nearly every statistical model is a GLM or an extension of one
- Simple Linear Regression
- Linear regression: the one-predictor GLM, ŷ = x·β̂x + β̂0 — simply the equation of a line
- Slope (βx): expected change in y for a one-unit change in x
- Intercept (β0): predicted y when x = 0; models overall magnitude even if x never reaches zero
- Residuals: whatever the fitted model leaves unexplained; called residuals since "error" is ambiguous
- Correlation vs. regression: slope equals r times the ratio of y and x standard deviations; identical for Z-scores
- Advantage over correlation: only regression can incorporate multiple independent variables
- Regression to the Mean and Study Design
- Regression to the mean: Galton's discovery that extreme values drift closer to the average on re-measurement
- Selection trap: recruiting only the poorest readers guarantees apparent improvement with no intervention
- Simulation: scores "rose" from 88 to 101 under pure chance, from two independent random draws
- Control group: necessary to separate real intervention effects from regression to the mean
- Random assignment: prevents systematic differences between treatment and control groups
- Inference and Goodness of Fit
- Residual variance: how much variability in Y the model fails to explain
- MSerror: squared residuals summed, then divided by N minus the number of estimated parameters
- Standard errors: SE of the model is √MSerror, rescaled per parameter by the spread of x
- t statistic: parameter estimate divided by its standard error, tested against β = 0
- R² (coefficient of determination): SSmodel/SStotal, the fraction of variance the model accounts for
- Caution: a statistically significant model with small R² may explain very little
- Building More Complex Models
- Multiple predictors: ŷ = β̂1·studyTime + β̂2·priorClass + β̂0
- Dummy coding: binary 1/0 variable whose coefficient is the difference in means between groups
- Shared slope assumption: the model assumes the study-time effect is identical in both groups
- Interaction: when the effect of one variable depends on the value of another
- Example: caffeine's effect on public speaking looks null until an interaction is modeled
- Values Embedded in Statistical Tools
- Complicated legacy: Galton, Fisher, and Pearson all promoted eugenics and its grave harms
- Methods endure: the founding tools remain essential and cannot be discarded for their creators' views
- Analysis is value-laden: data selected and questions asked embed the analyst's preconceptions
- Ethical duty: consider the impact of analysis on populations least able to defend themselves
- The General Linear Model Framework
- Interactions, Model Fit, and Prediction Limits (14. The General Linear Model · II)
- Interactions Change the Model
- Interaction: allowing a predictor's slope to differ across groups, not merely shifting the intercept
- Anxiety example: caffeine plus anxiety alone finds nothing; adding their interaction reveals opposite slopes
- Main effects: hard to interpret when a significant interaction is present, since the effect depends on another variable
- Baseline category: the alphabetically first group (anxious) becomes the reference the other slope is compared against
- Interaction coefficient: measures the difference in slope between groups, not a slope itself
- Comparing and Choosing Models
- Model comparison: asking which of two models fits better, not merely whether one fits
- Analysis of variance: the F test on the drop in residual sum of squares favors the interaction model
- Nested models: one model is a simplified version of the other, making comparison straightforward
- Nonnested models: comparison becomes much more complicated when neither model contains the other
- Specifying and Checking Assumptions
- Garbage in, garbage out: a model must be properly specified and paired with appropriate data
- Misspecification: omitting an intercept or relevant variables makes the model misrepresent the data
- Normality of residuals: the general linear model assumes the differences between predictions and data are normally distributed
- Q-Q plot: plots data quantiles against normal quantiles; divergence from the line signals nonnormality
- "Linear" Means Linear in Parameters
- The name clarified: "linear" describes parameters entering linearly, not straight-line responses
- Curves allowed: the same machinery fits nonlinear shapes because they stay linear in their parameters
- Generalized linear models: extensions that handle binary and other non-continuous outcomes
- Prediction Versus Fitting
- "Predict" misused: regression labels fitted values predictions, implying they forecast unseen data
- Overfitting: a 32-parameter model on 48 children fits the sample yet errs by 25 kg on new children
- Shuffling test: even with the true relationship destroyed, a complex model fits the noise and looks accurate
- Skepticism: prediction-accuracy claims deserve doubt unless generated by appropriate methods
- Cross-Validation
- Cross-validation: repeatedly fit the model leaving out a subset, then test it on that withheld subset
- Sixfold example: split data into six subsets and fit six times, each tested on the left-out group
- Honest accuracy: cross-validation's R² (0.35) tracks new data (0.33), not the inflated original fit (0.95)
- Data leakage: test information must never enter training; consult an expert before using it in practice
- The Matrix View of Fitting
- Matrix form: the general linear model is written Y = Xβ + E, with capitals marking matrices
- Design matrix: X holds a column of ones for the intercept plus each predictor's values
- Dimensions: an 8×2 design matrix times a 2×1 β yields the 8×1 fitted values
- Estimation: solving for β requires multiplying by the matrix inverse of X, not ordinary division
- Interactions Change the Model
- Regression, Prediction, and Model Interpretation (14. The General Linear Model · I)
- 15. Comparing Means
- Testing a Single Mean
- Sign test: asks whether the proportion of values above a hypothesized mean differs from chance's 0.5.
- Binomial test: turns that proportion into a p-value using the binomial distribution.
- t statistic: deviation of the sample mean from the hypothesized value, scaled by the standard error of the mean.
- Large p-value is not evidence for the null: the null was assumed true to begin with, so it cannot be confirmed.
- Bayes factor: quantifies evidence for the null itself; here it showed exceedingly strong support for no effect.
- Comparing Two Independent Means
- Two-sample t statistic: difference in group means divided by the standard error of that difference.
- Welch degrees of freedom: used when group sizes and variances differ, rather than assuming they are equal.
- One-tailed test: justified when the direction is specified in advance, as with marijuana use and TV watching.
- Illustrative result: regular marijuana users watched about 0.8 more hours of TV daily, a significant difference.
- The t Test as a Linear Model
- Reframing: the t test is simply the general linear model with one binary predictor.
- Dummy variable: coded 1 for the group of interest and 0 for the baseline group.
- Intercept: the predicted mean for the group coded zero.
- Slope: the difference in means between the two groups.
- Two- vs one-tailed: the linear model's p-value is exactly twice the one-tailed t test p-value.
- Effect Sizes and Evidence for Mean Differences
- Cohen's d: expresses the mean difference in standard deviation units; 0.55 counts as medium.
- Computing d: the model slope divided by the residual standard deviation.
- R²: just 0.05 of the variance explained — statistically significant yet practically small.
- Bayes factor: a value of 61 indicates quite strong evidence against the null.
- Comparing Paired Observations
- Within-subjects design: the same person is measured repeatedly, so observations are not independent.
- Independent t test misleads: it ignores that paired data points come from the same individual.
- Difference scores: the right representation — analyze within-person change, not raw measurements.
- Paired t test: a one-sample t test on whether the mean within-person difference is zero.
- Sign test trade-off: discarding the magnitude of differences means real effects can be missed.
- Evidence gap: a significant paired result still yielded a Bayes factor of only 3, weak evidence.
- Comparing More Than Two Means
- Omnibus hypothesis: tests whether any group mean differs, not which ones or by how much.
- Variance partitioning: total sum of squares splits into model and error components, each divided by its df.
- F distribution: describes how ratios of sums of squares behave under the null; it has two degrees of freedom.
- F statistic: asks whether the model beats a model containing only an intercept.
- Dummy coding extended: k groups need k−1 dummy variables, with the omitted group forming the baseline.
- Testing a Single Mean
- 16. Multivariate Statistics
- Finding Structure Without Supervision (16. Multivariate Statistics · I)
- Two Families of Multivariate Analysis
- Multivariate analysis: treats multiple random variables as equals, not as dependent and independent roles
- Structure discovery: asks which variables or observations relate, using distance as the measure
- Dimensionality reduction: compresses many variables into fewer while retaining maximum information
- Unsupervised learning: clustering and dimensionality reduction lack a target value to predict
- No single right answer: methods rest on differing assumptions and goals, unlike supervised learning's agreed "best" model
- The Self-Control Dataset
- Nine dimensions: 522 participants measured on four SSRT variables and five UPPS-P impulsivity variables
- Response inhibition: SSRT estimates how long an individual takes to stop an action
- Impulsivity: UPPS-P survey assesses five facets of acting without regard to consequences
- Purpose: a small, freely available dataset makes the methods concrete before scaling up
- Visualizing Multivariate Structure
- Hard limit: human vision cannot directly grasp data beyond three dimensions
- Scatterplot of matrices: histograms on the diagonal, pairwise scatterplots with regression lines below, correlations above
- Practical use: effective for roughly ten variables or fewer
- Heatmap: color encodes correlation values, scaling to much larger matrices
- fMRI example: correlated activity across 316 brain regions reveals diagonal blocks matching major brain networks
- Clustering by Distance
- Clustering: identifies groups of observations or variables whose values are most similar
- Euclidean distance: the straight-line length between points, generalizing the Pythagorean theorem to any dimensions
- Unlike correlation: Euclidean distance is sensitive to each variable's overall mean and variability
- Standard practice: Z-score or scale variables before computing distances
- K-Means Clustering
- Procedure: seed K centroids, assign each point to its nearest, recompute centers, iterate to stability
- No correct K: the number of clusters depends on goals and trade-offs, not ground truth
- Instability: random starting points can yield different solutions across runs
- Diagnostic: rerun with many random seeds and discount unstable results
- Confusion matrix: tabulating clusters against known labels, as with continents, shows only partial matches
- Self-control result: K=2 stably separates SSRT from impulsivity; higher K is inconsistent
- Two Families of Multivariate Analysis
- Clustering, Components, and Latent Factors (16. Multivariate Statistics · II)
- Hierarchical Clustering
- Agglomerative clustering: every point begins as its own cluster; the two closest clusters merge repeatedly
- Average linkage: distance between clusters equals the average distance between all their point pairs
- Dendrogram: a treelike map showing how variables relate across every distance scale at once
- Cutting the tree: cut height sets cluster count — 25 yields two, 20 yields three, 19 yields four
- Cross-method agreement: the hierarchical solution matched the majority of K-means runs
- Substantive reading: SSRT and UPPS variables resemble each other within sets, differ sharply between sets
- Clustering Is a Judgment Call
- No single right number: methods rely on different assumptions and can return different clusters
- Test robustness: display the data clustered at several levels and confirm conclusions do not shift
- Random starts matter: in repeated K-means runs, shaded variables are cluster-mates; agreement signals stability
- Principal Component Analysis
- Purpose: replace many correlated variables with fewer components carrying most of the variance
- Variance as signal plus noise: components chase the strongest shared signal among variables
- Sequential components: each explains the maximum remaining variance while staying uncorrelated with earlier ones
- Versus regression: PCA minimizes perpendicular distance to the line; regression minimizes vertical distance
- Loadings: each variable's weight on a component; sign is arbitrary, so read magnitudes
- Putting PCA to Work
- Summary score: when the first component explains enough variance, it becomes a one-number summary
- Separate reductions: SSRT's first component captured about 60% of variance, UPPS's about 55%
- Testing two constructs: correlating the two summary scores gave r ≈ −0.01 with a narrow confidence interval
- Scree plot: variance per component guides how many to retain; the first two dominated the full dataset
- Component meaning: the first component tracked impulsivity, the second tracked response inhibition
- Factor Analysis
- Latent variables: observed measures arise from unobservable factors plus measurement error on each variable
- EFA versus PCA: factor analysis permits correlated factors and models measurement error; standard PCA forbids both
- Covariance structure models: parameters are chosen to match model-implied covariances to observed covariances
- Path diagram: ellipses for latent factors, rectangles for observed variables, arrows for substantial loadings
- Fit check: RMSEA quantifies covariance mismatch; values below 0.08 are often considered adequate
- Choosing the Number of Factors
- Complexity always wins: more factors fit better, so fit metrics must penalize the parameter count
- SABIC: balances fit, number of parameters, and sample size; lower is better, absolute values are uninterpretable
- Recovering structure: both simulated and real self-control data peaked at three factors
- Real-data payoff: the working memory and fluid reasoning factors correlated even more strongly than simulated
- Hierarchical Clustering
- Finding Structure in Multivariate Data (16. Multivariate Statistics · III)
- Visualizing Multivariate Data
- Scatterplot matrix: pairwise plots of every variable pair expose relationships at a glance
- Heatmap: color-codes a correlation matrix, superior when many variables make scatterplots unreadable
- Trade-off: scatterplots show individual points and shape; heatmaps show the overall pattern compactly
- Iterative Analysis Methods
- Iterative method: refines an estimate through repeated passes rather than one closed-form computation
- Compare runs: repeat the analysis and compare results, since outcomes can shift with starting conditions
- Where they appear: clustering and factor analysis both proceed iteratively
- K-Means Clustering
- Algorithm: assign each observation to the nearest of K centers, then recompute centers from members
- Convergence: alternate assignment and center-update steps until assignments stop changing
- Dependence: the solution depends on the chosen K and on initial center placement
- Choosing the Number of Clusters
- No unique answer: clustering has no single correct solution; it is as much art as science
- Variance criterion: favor the solution that accounts for the most variance in the data
- Similarity criterion: favor the number that maximizes similarity within clusters
- Visual inspection: pick the most interpretable solution — legitimate, but a judgment call
- Principal Component Analysis
- Ordered variance: the first component accounts for more variance than the second, and so on
- Uncorrelated components: components are constructed to be uncorrelated with one another
- Arbitrary sign: the sign of component values is meaningless — flipping them all changes nothing
- Not regression: a principal component is not identical to a regression line
- Exploratory Factor Analysis
- Contrast with PCA: EFA explains shared variance via latent factors rather than summarizing total variance
- Choosing factors: the number of factors must be determined by the analyst, like the cluster-number problem
- Separate modeling: working memory and fluid reasoning are closely related, yet modeling them separately has utility
- Visualizing Multivariate Data
- Finding Structure Without Supervision (16. Multivariate Statistics · I)
- 17. Practical Statistical Modeling
- Modeling Workflow, Diagnostics, and Bias (17. Practical Statistical Modeling · I)
- The Modeling Workflow
- Seven steps: specify the question, get data, prepare and visualize, choose a model, fit, diagnose, then test hypotheses.
- Iterative, not linear: violated assumptions send you back to transform data or respecify, then refit.
- Diagnostics decide trust: assumption checks determine whether p-values and effect sizes are meaningful at all.
- Data Wrangling and Reproducibility
- Wrangling: renaming, merging, and reshaping raw data is unavoidable, and best done in R or Python.
- Spreadsheets hide errors: manual edits are difficult to reproduce or audit, and can silently corrupt conclusions.
- Reinhart–Rogoff: one spreadsheet error in a debt-growth analysis helped justify global austerity programs.
- Choosing a Model from the Dependent Variable
- Residual normality: linear regression p-values assume roughly normal residuals, most critical in small samples.
- Central limit theorem: with hundreds or thousands of observations, sampling distributions stay normal despite nonnormal residuals.
- Q-Q plots: inspect visually for deviation; obvious departures in small samples must be addressed.
- Two remedies: transform the data (logarithm, square root) or adopt a model built for that distribution.
- Logistic regression: an extension of linear regression designed for binary outcomes, such as collapsed skewed data.
- Misspecification: nonnormal residuals often reveal a missing important variable rather than the wrong model family.
- Clustered Data and IID Violations
- IID assumption: residuals must be independent and have equal variance for unbiased estimates and valid inference.
- Hidden structure: shared teachers or repeated measurements make observations alike within clusters, breaking independence.
- Consequence: ignoring clusters biases effect estimates and flaws inference, as with the paired t test.
- Mixed-effects models: the standard remedy across fields for nested, repeated, or otherwise clustered data.
- Outliers and Influential Observations
- Tukey's criterion: an outlier falls more than 1.5 interquartile ranges beyond the first or third quartile.
- Cook's D: quantifies how much the regression parameters would shift if each point were omitted.
- Leverage: extreme x values can swing the fit, like weight far from a seesaw's fulcrum; central points cannot.
- Three origins: measurement error, a different data-generating process, or the genuine tail of a skewed distribution.
- Fix before deleting: log transforms tame extremes; removal needs scientific justification set before analysis.
- Instructive blunder: a published age–worldview finding required correction after ages of 5 and 32,757 appeared.
- Independent Variables, Confounders, Colliders
- Include plausible causes: omitting a variable that influences the outcome leaves unexplained structure in the residuals.
- Confounder: a common cause of both independent and dependent variables, like beach crowds linking ice cream and shark attacks.
- Causal caution: without confounders in the model, regression coefficients cannot be read as causal effects.
- Collider bias: conditioning on a common effect of two variables, such as NBA status, manufactures a spurious height–speed relationship.
- The Modeling Workflow
- From Question to Logistic Model (17. Practical Statistical Modeling · II)
- Selection Bias as Collider Bias
- Selection bias: a form of collider bias arising when sample inclusion depends on a modeled variable
- Collider rule: never condition on a plausible common effect of the dependent and independent variables
- Helmet paradox: analyzing only hospitalized riders made helmets appear to worsen injuries
- Sampling vigilance: interpretation requires knowing how the sample was collected and what biases shaped it
- Step 1–2: Question and Data
- Question of interest: is self-reported impulsivity related to the likelihood of ever being arrested?
- Impulsivity: the tendency to decide with little forethought, estimated via self-report questionnaires
- BIS11: subscales for motor impulsivity, nonplanning, and attentional impulsivity
- UPPS-P: urgency, lack of premeditation and perseverance, sensation seeking
- Additional instruments: Dickman inventory (functional/dysfunctional impulsivity) and the Sensation-Seeking Scale
- Outcome measure: one self-report item on lifetime arrests or charges
- Step 3: Prepare and Visualize
- Skewed outcome: 78.5% never arrested; most arrests occurred once — normality assumptions fail
- Clustered heatmap: correlations reordered by clustering revealed two groups of impulsivity variables
- Cluster one: classic impulsivity — nonplanning and urgency
- Cluster two: sensation seeking, thrill seeking, and boredom susceptibility
- Composite scores: average each cluster, then standardize to Z-scores for easier interpretation
- Scatterplot matrix: composites moderately correlated, slightly skewed, roughly normal
- Step 4: Confounds and Model Choice
- Binarize the outcome: collapse arrest counts to yes/no, since linear models fit skewed counts poorly
- Independence caveat: random sampling helps, but Mechanical Turk users may differ psychologically
- Confound age: older individuals had more opportunity to be arrested
- Confound sex: 27.4% of males versus 15.6% of females had been arrested
- Preliminary t tests: arrested participants were more impulsive (d = 0.33) and more sensation seeking (d = 0.5)
- Step 5: Fitting a Logistic Regression
- Binary outcome problem: linear regression predicts values outside the 0–1 range for extreme predictors
- Generalized linear model: the linear predictor links to the outcome through a function beyond the linear case
- Logit link: models the log odds of the outcome, so the equation is "linear in log odds"
- Predictions: fitted values now fall within the observed range of the binary data
- Adjustment: including age and sex tests whether impulsivity matters after removing their effects
- Trade-offs and Next Steps
- Dichotomizing costs information: collapsing arrests to yes/no discards the count structure
- Hurdle models: more complex generalized linear models capture both arrest likelihood and arrest counts
- Consultation advised: ask a statistical consultant whether such a model suits your data
- Diagnostics prerequisite: check model fit and residuals before trusting logistic regression estimates
- Selection Bias as Collider Bias
- Logistic Regression Diagnostics and Evidence (17. Practical Statistical Modeling · III)
- Diagnosing the Logistic Model
- Overdispersion: residual variance beyond expectation shrinks standard errors and p-values, inflating the Type I error rate
- Dispersion test: the fitted model shows no overdispersion (dispersion = 1, p = 0.3)
- Linearity of log odds: predictors must relate linearly to unobserved log odds, checked via local regression
- Cook's D: all values far below one with no visual outliers, so hypothesis testing could safely proceed
- Handling Nonlinearity in a Linear Model
- Inverted-U age effect: age's log odds curve peaks in midlife, violating the linearity assumption
- Squared predictor: centering age and adding Age² captures the curve while keeping the model linear
- Still linear: Age² enters by being multiplied by one regression parameter, so the model form is unchanged
- Testing Hypotheses and Quantifying Effect Size
- All predictors significant: impulsivity, sensation seeking, age, age², and sex each reached p < .05
- Dummy coding: SexFemale captures females' offset from the male baseline; its negative sign means lower odds
- Intercept: the mean log odds for a nonexistent male at zero on every predictor
- Robust effects: impulsivity and sensation seeking survive adjustment, so confounding does not explain them away
- Odds ratios: exponentiate coefficients to express effect size; compare variables only on the same scale
- Confidence intervals: each odds ratio is presented with its range, showing values consistent with the data
- Quantifying Evidence with Bayes Factors
- Bayes factor: measures evidence for the alternative hypothesis, not merely whether the null is rejected
- BIC approximation: BF ≈ e^((BIC₁ − BIC₀)/2), valid when the compared models have equal prior probabilities
- Computed result: BIC 532 (full) versus 542 (baseline) gave BF₀₁ = 0.004, inverted to a Bayes factor of 242
- Interpretation: a Bayes factor of 242 constitutes strong evidence that impulsivity relates to arrest
- Example: Mask Wearing and Face Touching
- Research question: does wearing a mask make people touch their face more, a COVID-era transmission concern
- Preregistration: design and analysis plan were registered on the Open Science Framework before data collection
- Replication: a follow-up study repeated the design, adding trustworthiness to the findings
- Open data: shared raw data let other researchers reproduce results and run additional analyses
- Chi-squared tests: neither study showed a significant mask-wearing effect (p = 0.9 and p = 0.3)
- Clustered Data and Mixed-Effects Models
- Clustered data: people filmed in the same segment resemble one another, violating independently distributed residuals
- Fixed effect: mask wearing — its two levels are the comparison of real interest
- Random effect: video segment — sampled from a population, included only to absorb between-segment variability
- Random intercept and slope: each segment gets its own baseline face-touching level, and its own mask-wearing effect
- Final result: mask wearing stayed non-significant, observation duration strongly predicted touching, and the odds-ratio interval still spanned nearly half to double
- Diagnosing the Logistic Model
- Building and Diagnosing Statistical Models (17. Practical Statistical Modeling · IV)
- The Analysis Workflow
- Recurring template: specify question, gather data, prepare, choose model, fit, diagnose, then test and size effects
- Bayes factor: quantifies evidence for the null itself — the mask study gave 24.2 favoring no effect
- Consult a statistician: engage one before starting; free university consulting prevents major headaches
- Asthma: Question and Data
- Hypothesis: PM2.5 air pollution relates to asthma prevalence across US regions
- Data: CDC 500 Cities Project — 26,448 census tracts drawn from 494 cities
- Confounds addressed: median income, median age, and population included alongside pollution
- Weak raw signal: asthma prevalence correlates only 0.08 with PM2.5
- Asthma: Diagnostics and Model Choice
- Continuous outcome: linear regression is the natural starting model for asthma prevalence
- Violated independence: census tracts cluster within cities, requiring a mixed-effects model
- Misspecification clue: residual nonnormality signals unmodeled structure, not mere noise
- Random effects cure: city-level random intercepts and slopes recenter residuals around zero
- Large-sample caution: with ~26,000 observations, almost any effect turns significant
- Asthma: Effect Size
- Per-unit effect: each μg/m³ PM2.5 associates with 0.22–0.50% higher asthma prevalence
- Practical translation: roughly 3–4 additional asthma cases per 1000 people
- Comparative scale: one μg/m³ PM2.5 ≈ a $3500 drop in median income
- Variance explained: PM2.5 accounts for nearly 5%, income about 3.5%
- Plants: Skew, Interaction, and Clustering
- Skewed outcome: raw plant abundance is long-tailed and right-skewed, calling for log transformation
- Question posed: does the nitrogen-fertilizer effect on growth differ between tilled and untilled soil?
- Double clustering: variation across plots and 132 species each demands random effects
- Diagnostics drive revision: severe nonnormality and species differences push from linear to mixed model
- Plants: Interaction and Post Hoc Testing
- Significant interaction: fertilization helps tilled plots more; main effects alone mislead
- Post hoc contrasts: Tukey adjustment controls Type I error across pairwise comparisons
- Standardized effects: untilled soil ~1/3 SD higher abundance; fertilization adds ~1/3 SD
- Modest variance: interaction adds ~1.1%, full model ~2.1% over a year-only baseline
- The Analysis Workflow
- Modeling Workflow, Diagnostics, and Bias (17. Practical Statistical Modeling · I)
- 18. Doing Reproducible Research
- Reproducibility Crisis and Questionable Research Practices (18. Doing Reproducible Research · I)
- The Ideal Scientific Workflow
- The standard model: hypothesize, collect data, test the null hypothesis, draw conclusions.
- Wansink's Elmo study: branded apples raised selection over cookies from 20.7% to 33.8% (P = .02).
- The assumption: preplanned comparisons and honest reporting yield reliable, reusable results.
- Why it matters: medicine, aviation, and daily trust in science rest on this self-correcting ideal.
- The Wansink Case: Fraud in Practice
- "Wizardry" needed: collaborators admitted early data showed no strong effect and required tweaking.
- The p-value push: an original p = .06 was massaged below .05 before publication.
- Fabrication uncovered: the study claimed 8- to 11-year-olds but actually used preschoolers.
- Consequences: 18+ papers retracted and Wansink resigned from Cornell in 2018.
- Lesson: outright fraud is rare, but its exposure reveals weaknesses in how results are validated.
- The Reproducibility Crisis
- Reproducibility Project: of 100 psychology studies re-run, only 37% replicated versus 97% originally significant.
- Not just psychology: failures span cancer biology, chemistry, economics, and the social sciences.
- Ioannidis's prediction: null hypothesis testing in modern science necessarily produces many false findings.
- Base rate effect: when hypotheses are unlikely, most positive outcomes are false positives.
- Positive Predictive Value
- PPV defined: the proportion of significant results that are actually true.
- Formula: PPV = [p(hTrue)·power] ÷ [p(hTrue)·power + (1−p(hTrue))·α].
- High-prior case: when p(hTrue) = 0.8, PPV = 0.98 — but such fields are rarely interesting.
- Low-prior case: when p(hTrue) = 0.1, PPV = 0.31, so most positives are false.
- Key insight: interesting fields test risky hypotheses, so their PPV is inherently low.
- Cookbook review: 80% of foods were linked to cancer risk, yet evidence was weak and pooled results null.
- The Winner's Curse
- Definition: significant results systematically overestimate the true effect size.
- Origin: economics auctions where the winning bid necessarily exceeds the item's value.
- Simulation: with true d = 0.2 and low power, published effect estimates are highly inflated.
- Mechanism: only extreme samples cross the significance threshold, skewing published effects high.
- Remedy: only with high power and large effects do estimates approach the true value.
- Questionable Research Practices
- The Compleat Academic: an APA career guide advising aspiring researchers.
- Bem's chapter: offers writing suggestions now recognized as deeply problematic.
- QRPs named: selective reporting and similar tactics undermine the reliability of findings.
- The framing: these are choices made under career pressure, distinct from outright fraud.
- The Ideal Scientific Workflow
- Questionable Practices and Reproducibility Reform (18. Doing Reproducible Research · II)
- The Winner's Curse
- Winner's curse: with low power, only inflated effect sizes reach significance, so published estimates overshoot the truth
- Simulation: as statistical power rises, the estimated effect size converges on the actual one (figure 18.2A)
- Selective reporting: significant results cluster at extreme estimates, nonsignificant ones near zero (18.2B)
- Practical consequence: across many studies of one effect, the true size is smaller than the largest reported
- HARKing, p-Hacking, and the ESP Affair
- HARKing: hypothesizing after results are known, reframing a post hoc conclusion as an a priori prediction (Kerr 1998)
- Goalpost problem: rewriting theory to fit data makes incorrect ideas nearly impossible to disconfirm
- p-hacking: running many analyses until one turns significant, then reporting only that one
- Common tactics: stopping at p < .05, cherry-picking variables or conditions, excluding participants, transforming data
- False-positive inflation: Simmons, Nelson, and Simonsohn (2011) showed these practices sharply raise the real false positive rate
- ESP case: Bem's nine experiments (d = 0.22) showed every practice Yarkoni later catalogued, from one-tailed tests to near-.05 p-values
- Preregistration
- Preregistration: submit a detailed study design and analysis plan to a repository before analyzing data
- Effect: precommitment makes p-hacking and hidden analytic flexibility far harder to pass off as planned
- Natural experiment: the NHLBI required clinical trials to register at ClinicalTrials.gov beginning in 2000
- Result: Kaplan and Irvin (2015) found positive clinical trial outcomes dropped markedly after registration began
- Rules for Reproducible Reporting
- Stopping rule: decide and report the data-collection termination rule before data collection begins
- Sample size: collect at least 20 observations per cell, or justify the cost of collecting fewer
- Transparency: list all variables collected and report all conditions, including failed manipulations
- Sensitivity checks: report results with eliminated observations included, and without any covariate
- Replication and Its Limits
- Replication: others should be able to run the same study and obtain the same result
- Self-replication first: confirm your own finding in a new, adequately powered sample, often larger than the original
- Failure isn't falsity: at 80% power, a true effect still yields a nonsignificant result roughly one time in five
- Weight of evidence: require multiple replications, ideally higher-powered than the original, before believing a finding
- p-value limits: p says nothing about whether a finding is true or will replicate; that needs its prior probability
- Test case: Galak and colleagues' seven-study attempt failed to replicate Bem's ESP results
- Computational Reproducibility and Doing Better Science
- Computational reproducibility: share data and analysis code so others can rerun your exact analyses
- Scripted analyses: prefer R scripts over point-and-click software, and free open-source tools others can actually run
- Sharing infrastructure: version control on GitHub, Zenodo for larger datasets, OpenNeuro for neuroimaging data
- Incentives: journals such as Psychological Science award badges for shared materials, data, code, and preregistration
- Scientist's responsibility: the goal is not a significant result but answering nature's questions truthfully
- The Winner's Curse
- Reproducibility Crisis and Questionable Research Practices (18. Doing Reproducible Research · I)
- Preface��������������
- Core Conclusion and Practical Takeaways
- Core Ideas: How to Think About Data
- Statistics defined: describing a complex world in simple terms, plus honest uncertainty about it.
- Data = model + error: every analysis splits observations into expected pattern and residual deviation.
- Aggregation is deliberate forgetting: throwing data away is how understanding and generalization are built.
- Evidence is tentative, never proof: statistical conclusions carry irreducible doubt, unlike mathematical proof.
- Generalize, don't memorize: the point is predicting beyond the sample, not fitting the sample itself.
- Base rates rule: when hypotheses are unlikely, most "significant" findings are false positives.
- Practical Habits for Working with Data
- Plot raw data first: inspect distributions and individual points before trusting any summary or fit.
- Show the full distribution: prefer beeswarm, violin, or box plots over means-only bar charts.
- Check assumptions and outliers: use Q-Q plots, Cook's D, and leverage before trusting p-values.
- Wrangle reproducibly: script analyses in R or Python; spreadsheets hide unreproducible errors.
- Simulate when formulas fail: bootstrap and randomization tests replace shaky normality assumptions.
- Cross-validate predictions: honest accuracy comes only from data the model never saw during fitting.
- Interpreting Evidence Correctly
- p-value ≠ probability of null: it is the data's probability given the null, never the reverse.
- Significance ≠ practical importance: with large samples, trivial effects turn statistically significant.
- Report effect sizes: Cohen's d, odds ratios, and correlations convey magnitude, not mere existence.
- Intervals answer better questions: confidence intervals bound estimates; credible intervals state parameter probability.
- Bayes factors weigh evidence: they can support the null where frequentist tests cannot.
- Respect familywise error: with many tests, correct alpha or expect a flood of false positives.
- Designing Studies Right
- Specify hypotheses before data: preregistration blocks HARKing and hidden p-hacking.
- Run power analysis first: choose sample size from the effect you actually hope to detect.
- Control confounders and colliders: randomize or model common causes; never condition on common effects.
- Beware regression to the mean: extreme groups drift toward average with no intervention at all.
- Favor simplicity: overfit models capture noise and generalize worse; simpler often wins honestly.
- Replicate before believing: one significant result proves little; require independent, well-powered repeats.
- Mindset Shifts
- Embrace uncertainty: irreducible doubt is inherent to real-world conclusions, not a personal failing.
- Statistics is learned by doing: only running your own analyses on real, messy data builds skill.
- Reframe statistics anxiety: nervousness is normal and can sharpen focus instead of blocking learning.
- Correlation is a hint, not proof: association never establishes causation without an experiment.
- Analysis is value-laden: the questions asked and data selected embed the analyst's preconceptions.
- Core Ideas: How to Think About Data
opening map…