- General Overview
- The Central Thesis: Error Splits in Two
- Judgment as measurement: the human mind is an instrument that assigns value on a scale
- Bias: systematic, directional deviation — the average error of a set of judgments
- Noise: unwanted variability, scatter that should not exist and cannot be explained
- The equation: MSE = bias² + noise² — the two contribute to error independently
- The imbalance: bias stars in research and public debate; noise is the unacknowledged bit player
- Bounded disagreement: judgment sits between computation, where disagreement is banned, and taste
- The Evidence: A Noisy World
- Sentencing: identical cases drew probation to ten years — a genuine judgment lottery
- Insurance audit: executives expected 10% divergence; actual median gaps reached 55%
- Medicine: doctors disagree on X-rays, angiograms, biopsies, and basic diagnoses
- Forecasting: experts disagree with each other and, on repeated occasions, with themselves
- Everywhere: asylum, bail, custody, hiring, patents, and forensics all show it
- Measurable without truth: scatter is visible from behind the target, without a bull's-eye
- The Anatomy of Noise
- System noise: unwanted variability among interchangeable professionals judging identical cases
- Level noise: some judges are consistently harsher, others consistently more lenient
- Pattern noise: idiosyncratic reactions to particular cases — generally the largest component
- Occasion noise: the same person judging the same case differently at different moments
- The decomposition: noise squared splits into level, stable pattern, and occasion components
- Two lotteries: which professional you draw, and which moment that professional is in
- Why Noise Stays Invisible
- Naive realism: we assume others view the world much as we do, so disagreement seems rare
- Causal craving: causal stories explain single events; noise is statistical, visible only in ensembles
- Valley of the normal: successful outcomes are self-explaining and never prompt questions
- Confidence signals: a coherent story feels like accuracy but guarantees nothing
- Objective ignorance: much is genuinely unknowable, and we systematically underestimate how much
- Excessive coherence: early impressions contaminate later assessments through the halo effect
- Humans Versus Models
- Meehl's verdict: simple mechanical rules generally beat clinical judgment, study after study
- Noiseless core: rules win partly because they introduce no random variability at all
- Simplicity suffices: equal weights rival optimized regression; frugal rules match elaborate tools
- The model of the judge: experts act as if using simple formulas — and the model beats them
- Mindless consistency: randomly weighted models beat experts; consistency, not insight, drives the gain
- Resistance: algorithm aversion, a perfection double standard, and fear of technological replacement
- Better Judges and Debiasing
- Better judges: ability, cognitive style, and active open-mindedness reduce both bias and noise
- True versus respect experts: skill can be verified against outcomes; esteem-expertise rests on peer respect
- Debiasing ex ante: reshape the environment or train the judge before the judgment occurs
- Debiasing ex post: correct a known, predictable directional error after the fact
- The limit: bias-specific fixes fail when several biases work at once or differ across judges
- Decision observer: a live checklist-watcher can flag biases as a decision unfolds
- Decision Hygiene and Structure
- Decision hygiene: prevent unnamed errors, as handwashing prevents unnamed germs
- Structure: break complex judgments into independent, separately assessed components
- Independence: elicit judgments separately — cascades and deliberation amplify noise
- Aggregation: averaging independent judgments divides noise by the square root of their number
- Outside view: treat the case as one of a reference class and regress toward its mean
- Relative judgments: pairwise comparisons beat absolute scales and reliably reduce noise
- Costs, Limits, and a Less Noisy World
- Optimal noise: the right level is not zero — reduction can cost more than it saves
- Dignity: people want a human to weigh their particular circumstances, not a rigid rule
- Rules or standards: rules cut noise but risk rigidity; standards preserve latitude and invite noise
- Contested objections: gaming, deterrence, moral evolution, and demoralization are real, often overstated
- The remedy: replace bad noise reduction with better noise reduction, not with resignation
- The invitation: audits, algorithms, decomposition, and hygiene should become standard practice
- The Central Thesis: Error Splits in Two
- Deep Dive
- Introduction: Two Kinds of Error
- The Shooting Arcade: Bias, Noise, and Error
- The shooting arcade: four teams, one rifle, five shots each — a metaphor for human judgment
- Bias: systematic deviation from the target; shots cluster tightly but land consistently off-center
- Noise: random scatter around the target, with no pattern and no obvious explanation
- Team D: real-world error often combines bias and noise at once
- Predictability: consistent bias lets you forecast the next shot; variability destroys that power
- Measuring Noise Without Knowing the Truth
- Back of the target: scatter remains fully visible and measurable without seeing the bull's-eye
- Core property: noise can be measured without knowing the right answer or who is right
- Why it matters: judgments with unknowable answers — diagnoses, film forecasts — can still be studied
- The imbalance: bias is the star of research and public debate; noise is the unacknowledged bit player
- Noise Everywhere: The Evidence
- Medicine: doctors disagree about the same patient, even when reading X-rays
- Child custody: heavy-handed case managers send more children to foster care, with worse later outcomes
- Forecasts: professionals disagree with each other and themselves — identical task estimates varied by 71%
- Asylum and bail: outcomes hinge on the assigned judge — one admitted 5% of applicants, another 88%
- Personnel, forensics, patents: ratings reveal the assessor, fingerprint experts contradict themselves, patent grants track the examiner
- The Book's Remedy: Audits and Decision Hygiene
- Noise audit: the instrument for measuring disagreement among professionals judging identical cases
- Occasion noise: the same person or group judges the same case differently on different occasions
- Rules over humans: their advantage lies chiefly in noiselessness, not superior insight
- Decision hygiene: techniques for reducing noise across medicine, business, education, and government
- Mediating assessments protocol: a general approach to evaluating options with less noise and more reliability
- Optimal noise: the right level is not zero — eliminating it can be infeasible, costly, or harmful to morale
- The Shooting Arcade: Bias, Noise, and Error
- Part I: Finding Noise
- 1. Crime and Noisy Punishment
- The Scandal of Discretionary Sentencing
- Judicial discretion: long celebrated as humane, letting sentences fit the defendant's unique character and circumstances.
- The outrage: similar offenders, convicted of the same crime, receive probation, two years, or ten years.
- The principle: a sentence should not depend on which judge the case happens to be assigned to.
- System noise: unwanted variability in judgments that should ideally be identical.
- Key theme: wherever there is judgment, there is noise — and more of it than you think.
- Frankel's Crusade
- Marvin Frankel, 1973: a federal judge outraged that a bank robber's 0-to-25-year range reflected the judge, not the case.
- Anecdotes over statistics: a $58.40 counterfeit check drew 15 years; a $35.20 one drew 30 days.
- "Arbitrary cruelties": unchecked sweeping judicial power was intolerable in a government of laws, not of men.
- His remedy: detailed checklists, objective grading, computers, and a sentencing commission.
- Twin targets: inexplicable variation (noise) and racial and socioeconomic disparity (bias).
- The Experiments Confirm the Noise
- 1974 study, chaired by Frankel: 50 judges given identical reports — "absence of consensus was the norm."
- Astounding ranges: a heroin dealer got one to ten years, a bank robber five to eighteen, an extortionist twenty years plus a $65,000 fine down to three years with none.
- Prison or not: in 16 of 20 cases, judges disagreed on whether to incarcerate at all.
- 1981 study: 208 federal judges, 16 cases; only 3 yielded unanimous prison terms, one fraud case reaching life.
- Understatement: controlled experiments likely understate real-world noise, since real judges see far more information.
- When Irrelevant Factors Decide
- Hunger: parole is granted more readily at the start of the day and after food breaks.
- Football: after a local team's weekend loss, judges are harsher — and Black defendants bear the brunt.
- Birthdays: French judges are more lenient when the defendant's birthday falls on the hearing day.
- Temperature: immigration judges grant asylum significantly less often on hot days.
- Reform: The Sentencing Reform Act of 1984
- Kennedy's persistence: shocked by the disparity evidence, he pressed for reform from 1975 and won in 1984.
- The act: restricted the "unfettered discretion" of judges and parole authorities to cut unjustified disparity.
- Sentencing Commission: created to issue mandatory guidelines anchored in average sentences from 10,000 real cases.
- Two inputs: crime severity (43 offense levels) and criminal history, yielding narrow ranges requiring appellate justification to depart from.
- Breyer's admission: no defensible rank ordering of crimes could be agreed on; guidelines simply followed history.
- Backlash and the Return of Noise
- Objection: guidelines were too mechanical — "the need is not for blindness, but for insight, for equity."
- Effect: studies found guidelines cut the variation attributable to the identity of the sentencing judge.
- 2005: the Supreme Court struck the guidelines down on unrelated technical grounds, making them advisory.
- Judges' verdict: 75% preferred the advisory regime; only 3% preferred mandatory.
- Yang's data: across nearly 400,000 defendants, interjudge disparity doubled after 2005 and racial disparities widened.
- The Scandal of Discretionary Sentencing
- 2. A Noisy System
- The Insurance Company's Noise Audit
- Origin: executives doubted noise was a serious problem, so they settled it with a simple experiment
- Noise audit: the same case is judged by many professionals, making their variability visible
- Goldilocks value: for any risk a right price exists; deviations above or below it both cost money
- The shock: results shattered executives' assumption that qualified colleagues would judge similarly
- The Judgment Lottery
- Random assignment: chance picks which underwriter or adjuster handles a given case
- High stakes: early estimates set implicit negotiation goals and bind the firm to financial reserves
- Unjustified lottery: unlike lotteries allocating goods or bads, this one allocates nothing—only uncertainty
- Quantified cost: one senior executive estimated the annual cost of underwriting noise in the hundreds of millions
- What Noise Audits Reveal
- Expected divergence: most executives, and 828 surveyed CEOs, guessed a 10% difference or less
- Actual divergence: median differences were 55% in underwriting, 43% for adjusters, 41% for asset managers
- Medians understate: in half of case pairs, the gap between two judgments was even larger
- Scatter not bull's-eye: you needn't know the correct answer to see that the variability is a problem
- Customer expectation: people expect organizations to deliver consistent judgments, not system noise
- Unwanted Variability vs Wanted Diversity
- Taste and markets: differences in preferences, innovation, and forecasting are welcome and productive
- Variation enables selection: in markets as in nature, the best judgment cannot triumph without variation
- System noise defined: unwanted variability within an organization—a problem of systems, not markets
- Errors don't cancel: noisy systems judge different cases; overpricing and underpricing both cost
- Pervasive lottery: it decides which doctor, judge, patent examiner, or service rep you encounter
- The Illusion of Agreement
- Naive realism: we assume others view the world much as we do, rarely imagining alternatives
- Confidence from fluency: professionals grow confident by agreeing with their past selves, not their colleagues
- Vague shared rules: common language eliminates some options but never specifies a common answer
- Discomfort of disagreement: organizations prefer harmony; masking ratings caused "so many disagreements" they reverted
- Postmortems mislead: easy consensus about egregious errors reinforces the illusion on acceptable judgments
- The verdict: wherever there is judgment, there is noise—and more of it than you think
- The Insurance Company's Noise Audit
- 3. Singular Decisions
- Singular Versus Recurrent Decisions
- Recurrent decisions: interchangeable professionals judge similar cases repeatedly, so unwanted variability can be measured.
- Singular decisions: made once, with no prepackaged response and genuinely unique features.
- Obama's 2014 Ebola response: refusing to close borders and sending 3,000 people to West Africa was a one-time call.
- A continuum, not a category: a fourth house purchase starts to feel recurrent; going to war never does.
- Two Separate Traditions of Study
- Recurrent decisions: social scientists take a statistical view, studying many similar cases to find patterns and measure accuracy.
- Singular decisions: historians and management gurus take a causal view, working backward from what happened.
- The split left high-stakes singular judgment largely unexamined for noise.
- Counterfactual Proof of Hidden Noise
- No second run: singular problems are never exactly repeated, so direct comparison of judgments is impossible.
- Counterfactual test: imagine an equally competent decider, with the same goals and facts, reaching a different conclusion.
- Irrelevant influences: mood, weather, or how facts were framed could have changed the outcome.
- COVID-19 evidence: nations struck at the same time in similar ways responded very differently, revealing noise.
- Invisible is not absent: had only one country been hit, no variability would appear—yet the noise would remain.
- Controlling Noise in Singular Decisions
- Reframe the decision: a singular decision is a recurrent decision that happens only once.
- Resist the one-of-a-kind instinct: treating it as wholly unique invites the noise you would otherwise filter out.
- Reject the exemption claim: the rules of probabilistic thinking are not irrelevant to one-time choices under uncertainty.
- Same remedies work: practices that reduce bias and noise in repeated judgments apply to your single shot.
- Speaking of Singular Decisions
- "The way you approach this unusual opportunity exposes you to noise."
- "Remember: a singular decision is a recurrent decision that is made only once."
- "The personal experiences that made you who you are are not truly relevant to this decision."
- Singular Versus Recurrent Decisions
- 1. Crime and Noisy Punishment
- Part II: Your Mind Is a Measuring Instrument
- Noisy Judgment and Bounded Disagreement (4. Matters of Judgment · I)
- Judgment as Measurement
- Human mind as instrument: judgment assigns value on a scale, so the mind measures with error.
- Goal is accuracy: judgment aims at truth, not persuasion, impressiveness, or taking a stand.
- Judgment as conclusion: a word or phrase summarizing a mental process; not synonymous with thinking.
- Judge as technical term: anyone making professional judgments, not only courtroom judges.
- Error is inevitable: measurement and judgment never achieve perfection; error splits into bias and noise.
- Bias, Noise, and the Stopwatch
- Stopwatch exercise: trying to produce five ten-second laps reveals unavoidable variability.
- Noise defined: uncontrolled variability around a target; bias is the mean’s distance from ten seconds.
- Universal variability: physiological and psychological noise is ordinary, within and between people.
- Standard deviation: the common statistical measure for quantifying noise in judgments.
- Predictive judgments: forecasts and diagnoses aim at a true value, so differing judgments cannot both be right.
- Matters of Judgment
- Professional judgment: assumes a competent judge aiming to get it right, yet never certain the judgment is right.
- Bounded disagreement: judgment lies between fact/computation and taste/opinion; reasonable people may disagree, but only within limits.
- Judgment where agreement expected: the word applies mainly when people believe they should converge.
- Unacceptable extremes: absurd judgments—one-dollar fine or life sentence for routine fraud—produce agreement.
- Acceptable amount: how much disagreement is tolerable is itself a judgment call, depending on difficulty.
- Matters of taste: unresolved differences in music or food are fully acceptable, unlike professional judgment.
- The Gambardi Exercise
- Task: estimate probability a CEO candidate keeps his job two years, using conflicting evidence.
- Selective attention: you weight some cues and ignore others without full awareness; recall varies across readers.
- Informal integration: cues are combined into a coherent impression quickly, without a formal plan, creating noise.
- Impression to number: mapping a holistic impression onto 0–100 adds further unexplained variability.
- Wide scatter: 115 MBA students gave probabilities from 10 to 95—large between-judge noise.
- Noise, Reliability, and Predictive Aims
- Stopwatch noise: variability across repeated attempts by one person shows within-person unreliability.
- Gambardi noise: variability across different people judging one case shows between-person unreliability.
- Two reliabilities: measurement language distinguishes within-person from between-person reliability.
- Predictive judgments: forecasts and diagnoses aim at a true value, so differing judgments cannot both be right.
- Verifiable outcomes: for temperature, games, or elections, disagreement can eventually be settled by results.
- Judgment as Measurement
- Verifiability, Coherence, and Evaluating Judgment (4. Matters of Judgment · II)
- Nonverifiable Judgments
- The Gambardi problem: because he is fictitious, no outcome can ever reveal who judged correctly
- Unconfirmable probabilities: any judgment between 0 and 100% escapes confirmation or disconfirmation alike
- Outcome is not ex ante probability: a 90% forecast that fails may still have been sound — 10% events occur 10% of the time
- Two routes to nonverifiability: invented subjects and probabilistic answers each block later verification
- Professionals live here: underwriters will never know whether a policy was over- or underpriced
- Conditional and distant forecasts: "if we go to war, we will be crushed" and century-scale climate estimates will likely stay untested
- The Internal Signal of Completion
- Verifiability does not change the experience: people tackle plausible hypothetical problems much as real ones
- Not error minimization: with no outcome in view, the judge strives only to land on an answer worth trusting
- Signal of coherence: the answer feels right when it fits comfortably with the messy evidence
- 0 and 100 are unavailable: absolute confidence would contradict ambiguous, conflicting facts
- Research payoff: this is why made-up problems in the lab behave like real ones
- Evaluative Judgments
- Not predictions: sentencing, grading, and wine judging match a response to the severity or quality of a case
- Trade-off decisions: hiring, strategy, and policy choices weigh pros and cons of multiple options
- Predictive inputs remain: every such decision still rests on forecasts of performance, markets, or epidemic spread
- Bounded disagreement expected: informed colleagues sharing goals should mostly agree, not wildly diverge
- Not mere taste, boundary fuzzy: values shape these judgments, yet judges and graders cannot always say which type they are making
- Two Ways to Evaluate Judgment
- Outcome scoring: for verifiable judgments, error is the measured gap between judgment and what happened
- Process scoring: the alternative, and the only one that works when no true value exists
- Ensemble test: a forecaster calling 100 candidates at 70% should see roughly 70 elected
- Normative test: does the process conform to logic and probability theory?
- Dart-throwing chimpanzee: in a single quarter, luck can beat a skilled forecaster with the best tools
- Scholars' advice: judge the process, not the single-case outcome — though real life rewards the opposite
- What's Wrong with Noise
- Predictive noise signals error: two disagreeing doctors or forecasters cannot both be right
- Evaluative noise offends fairness: interchangeable judges assigned quasi-randomly must not disagree wildly
- Frankel's "arbitrary cruelties": disagreement that turns sentencing into a lottery is indefensible
- Credibility damage: judgments should reflect the system's values, not those of an individual judge
- Everyday stakes: unequal refunds, promotions, grades, and disability rulings erode trust
- Bounded tolerance: "it's a matter of judgment" excuses some disagreement, but judgments too far out are simply wrong
- Undesirable but Measurable
- No true value needed: measuring noise requires only multiple judgments of the same problem
- Shooting range: the bull's-eye is invisible, but the scatter of the shots is plain to see
- Noise audit: ask all forecasters for next quarter's sales; the spread of their answers is the noise
- Bias versus noise: separating the two is essential — and once measured, noise can often be reduced
- Nonverifiable Judgments
- 5. Measuring Error
- Bias and Noise as Equal Contributors
- Core claim: when accuracy is the goal, bias and noise play identical roles in overall error
- Consistent bias: a scale adding constant weight, or a perpetually optimistic manager, produces costly errors
- Noise costs too: errors scattered in both directions do not cancel out — they add up
- Priority: measuring and reducing noise deserves the same standing as measuring and reducing bias
- The GoodSell Noise Audit
- Setup: many forecasters independently estimate market share for one region
- Result: a bell curve peaking at 44%, standard deviation 10 points — one number captures system noise
- Key insight: noise is visible without knowing the true value, just as shot spread is visible from behind the target
- Later revealed: the outcome was 34%; average error (bias) was 10%, equal to noise in this case
- Boss's mistake: waiting to learn the bias before acting on noise is a false choice
- Mean Squared Error
- Gauss's rule (1795): overall error is the average of squared individual errors
- Intuition anchor: the mean is the best estimate — MSE is the only error measure that the mean minimizes
- Squaring's effect: large errors dominate; a 3 mm shift in an estimate can double MSE
- Legacy: least squares solved the rediscovery of Ceres and remains the standard wherever accuracy matters
- The Error Equations
- Single measurement: error = bias + noisy error, where noisy errors average zero
- Overall: MSE = bias² + noise²
- Pythagorean picture: MSE, bias², and noise² as the three squares on a right triangle
- Interchangeability: reducing bias or noise by the same amount lowers overall error equally, independent of the other
- The Counterintuitive Case for Noise Reduction
- Halving bias: the distribution shifts toward the truth — an obvious improvement
- Halving noise: forecasts concentrate and 98% now overshoot; it looks worse but is not
- Illusion decoded: bias is average error, the peak-to-truth distance — not the imbalance of positive and negative errors
- MSE vs intuition: people prize perfect hits but barely distinguish two large errors — the mirror image of accuracy
- Reduce both: cutting noise makes bias impossible to miss, putting it next on the agenda
- Surprise: matching noise's contribution required 84% of forecasts erring the same way — noise often rivals bias
- Where the Equation Applies
- Predictive judgments: forecasts and estimates aim at a true value — least bias, least noise
- Not evaluative: no true value, and costs are asymmetrical — overestimating an elevator's load can be catastrophic
- Symmetric costs absent: being one minute or five minutes late for a train costs the same
- Two steps: neutral prediction first, then values set the safety margin or acceptable risk
- Wide stakes: pandemic response and military offensives rest on value-neutral predictions — keep facts and values separate
- Bias and Noise as Equal Contributors
- 6. The Analysis of Noise
- The Sentencing Noise Audit
- Noise audit design: 208 federal judges sentenced 16 hypothetical robbery or fraud cases, yielding 3,328 judgments.
- Mean sentence benchmark: the average of 208 sentences per case is treated as the "just" sentence.
- Bias vs. noise: average bias remains a distinct source of error and unfairness, not this chapter's focus.
- System noise: variability within a case's column is undesirable; a perfect world would make each column identical.
- Sentencing lottery: mean sentence was 7.0 years, standard deviation 3.4 years, so random judges differed by 3.8 years.
- Understated reality: artificial cases and fuller courtroom information likely make actual sentencing noise even larger.
- Level Noise: Judges' Average Severity
- Level noise: variability in judges' average sentencing severity; its standard deviation was 2.4 years.
- Level errors: deviations from average severity may be errors even if the average itself is biased.
- Personality trait: average severity behaves like a stable disposition, unrelated to the case or defendant.
- Predictors: rehabilitation goals predicted leniency; Southern location and conservative ideology predicted severity.
- Beyond courts: level noise appears wherever evaluators differ in generosity, optimism, or aggressiveness.
- Pattern Noise: Judge × Case Interaction
- Pattern noise: variability in how judges respond to particular cases, beyond their average severity.
- Judge × case interaction: the statistical term for judges being inconsistently harsh or lenient across cases.
- Additive model fails: case mean plus judge severity does not predict most actual sentences.
- Idiosyncratic reactions: one judge may be lenient toward white-collar criminals but harsh toward recidivists.
- Different rankings: judges would not rank cases similarly by deserved prison time.
- Pervasive: pattern noise appears in hospitalization, hiring, legal case selection, and television production decisions.
- Decomposing System Noise
- Equation: System Noise² = Level Noise² + Pattern Noise².
- Approximate balance: in the sentencing study, level and pattern noise contributed about equally.
- Stable patterns: measured pattern noise reflects stable attitudes but includes some occasion noise.
- Occasion noise: within-person variability from mood, recent events, or weather; treated as random error here.
- General method: the same decomposition applies to any noise audit in business, medicine, or government.
- The Sentencing Noise Audit
- The Variability Within Our Own Judgment (7. Occasion Noise · I)
- The Free Throw Lottery
- Free throw as lottery: the same practiced player still produces variable outcomes
- Range of skill: NBA shooters average ~75%; the best near 90%, the worst near 50%
- Within-player noise: variability appears inside a single shooter, not only between shooters
- Unknowable causes: fatigue, pressure, and crowd effects are invoked but rarely verified
- The Second Lottery
- First lottery: system noise picks which professional happens to judge the case
- Second lottery: occasion noise picks the moment, the mood, the mental residue
- Abstract alternatives: we see the judgment made, never the cloud it was drawn from
- Same facts, different verdicts: physicians, wine judges, fingerprint examiners, software estimators all disagree with themselves
- Why Occasion Noise Hides
- Justification illusion: a considered opinion arrives wrapped in reasons that appear to explain it
- Consistency pull: recognizing a repeated case, experts simply reproduce their earlier answer
- Test-retest reliability: studies repeating judgments in one session mostly show self-agreement
- Blinding the expert: blind tastings and unrecognized delayed repeats bypass recognition and expose noise
- Big-data probe: significant effects of time of day or temperature on decisions betray occasion noise
- The Crowd Within
- One answer is a sample: your estimate is one point in a cloud of plausible answers
- Galton's ox: 787 villagers averaged 1,200 pounds against a true weight of 1,198
- Averaging cancels noise: combining independent judgments yields less noise, not less bias
- Vul and Pashler: averaging two guesses from one person beats either guess alone
- Distance pays: three weeks between guesses raises the gain to a third of a second opinion
- Dialectical Bootstrapping
- Argue against yourself: assume the first estimate is wrong, find why, then revise
- Active divergence: forcing a new perspective samples a more distant second self
- Bigger gain: consecutive dialectical estimates capture roughly half the value of a second opinion
- Practical sequence: seek independent opinions; failing that, build an inner crowd and average
- Sampling insight: responses are drawn from an internal distribution, not deterministically selected
- Occasion noise is universal: all our judgments fluctuate, all the time
- Sources of Occasion Noise: Mood
- Control requires mechanism: only by understanding what produces noise can we hope to reduce it
- Mood is the visible source: our judgments track how we feel, and so do others'
- Induction techniques: recalling happy or sad memories, or watching funny versus sad film clips
- Forgas's program: decades of research mapping mood's effect on judgment and decision
- The Free Throw Lottery
- Mood, Fatigue, and Moment-to-Moment Noise (7. Occasion Noise · II)
- Mood Colors What We Think
- Mood congruence: Good mood retrieves happy memories, approves of people, gives generously; bad mood reverses this
- Perception shift: The same smile reads friendly or awkward depending on the observer's mood
- Two effects of mood: It changes what you notice and retrieve—and also how you think
- The Mixed Blessings of Mood
- Good mood in negotiation: Cooperative and reciprocated, yielding better outcomes; anger mid-negotiation can work too
- Good mood breeds credulity: First impressions pass unchallenged and stereotypes bite harder
- Bullshit receptivity: Induced good mood makes people agree with vacuous pseudo-profound statements
- Bad mood as safeguard: Negative mood helps eyewitnesses disregard misleading information
- Moral judgment: Positive mood tripled willingness to push the man off the footbridge
- Stress, Fatigue, Weather
- End-of-day drift: Physicians prescribe more opioids and antibiotics, fewer flu shots, late in the day
- Quick fix under pressure: Time pressure pushes physicians toward fast solutions despite serious downsides
- Weather as mediator: Weather barely decides directly; it shifts mood, which shifts judgment
- Clouds make nerds look good: Admissions officers weight academics more on cloudy days, nonacademics on sunny ones
- Order and Sequence Effects
- Implicit frame: Each case is judged against the decisions that immediately preceded it
- Balance seeking: After a streak, judges lean opposite beyond what the case justifies
- Asylum data: Judges are 19% less likely to grant asylum when the prior two cases were approved
- Gambler's fallacy: We underestimate how often streaks occur by chance alone
- How Large Is Occasion Noise?
- Smaller than individual differences: In every measurable case, occasion noise contributed less than differences between people
- Asylum comparison: A 19% streak effect pales beside 88% versus 5% between judges in one Miami courthouse
- Self-consistency: Fingerprint examiners and physicians disagree with themselves less than with others
- Reassuring asymmetry: You resemble yesterday's you more than you resemble someone else today
- Inner Causes of Occasion Noise
- Kahana's memory study: Seventy-nine subjects, twenty-three sessions; all predictors explained only 11% of variation
- Best predictor is internal: Performance on one list predicts the next, ebbing and flowing with no external cause
- Endogenous neural efficiency: Moment-to-moment variability is intrinsic to how the brain functions
- The free-throw analogy: Neurons never fire identically, just as muscles never repeat a gesture exactly
- Control what you can: Occasion noise cannot be eliminated, but undue influences can be reduced—especially in groups
- Mood Colors What We Think
- Social Influence Amplifies Group Noise (8. How Groups Amplify Noise · I)
- Groups Add a Layer of Noise
- Group noise: similar groups reach very different verdicts; one group's judgment is merely one in a cloud of possibilities
- Irrelevant factors: who speaks first or last, who seems confident, seating, gestures—all shift outcomes
- Aggregation's limit: averaging independent judgments reduces noise, but group dynamics can create it instead
- Everyday stakes: hiring, promotion, admissions, regulations, and strategy all inherit this variability
- Wise crowds or herds: crowds can be accurate, or can follow tyrants, bubbles, and shared illusions
- The Music Downloads Experiment
- Design: thousands rated 72 unknown songs; eight parallel "worlds" could additionally see peers' download counts
- Finding: rankings across the eight groups were wildly disparate—social influence generated a great deal of noise
- Floor and ceiling: the worst songs never topped and the best never sank, but otherwise almost anything could happen
- Self-reinforcing popularity: early downloads snowballed, so outcomes hinged on who saw what first
- Manipulation: inverting the rankings made most unpopular songs popular and most popular songs flops
- One exception: the control group's single most popular song still rose—merit constrains but does not decide
- Social Influence Beyond Music
- UK referenda: an early burst of support is self-reinforcing, while a quiet first day essentially dooms a proposal
- Partisan tipping: the same view won or lost Democratic support depending on whether Democrats or Republicans endorsed it first
- Early movers: "chance variation in a small number of early movers" can tip large populations toward unrelated clusters of views
- Comment votes: an artificial first up vote made the next viewer 32% more likely to approve
- Persistence: five months later, that single fake vote still raised a comment's mean rating by 25%
- Group analogue: members' early agreement, neutrality, or dissent works like an up vote on plans, products, and verdicts
- Social Influence Undermines Crowd Wisdom
- Independence requirement: the wisdom of crowds presupposes members judging without seeing others' answers
- Learning versus herding: sharing what people know can help, but copying opinions produces conformity
- Estimation studies: crowds estimating crimes, populations, and borders were wise only when they registered views independently
- Exposure to averages: learning a twelve-person group's estimate made the crowd's answers worse
- The irony: social influence reduces group diversity without diminishing collective error
- Informational Cascades
- Cascade defined: pervasive dynamics that let similar groups move in multiple directions, so small changes yield different outcomes
- Hiring scenario: Arthur names Thomas first; Barbara, unsure, may simply trust him and agree
- Lost signals: deference to predecessors swallows private information the group might have needed
- Clouds of possibilities: history runs only once, but many group decisions had other plausible endings
- Noise link: because arbitrary early moves harden into group truth, similar groups diverge
- Groups Add a Layer of Noise
- Cascades, Social Pressure, and Group Polarization (8. How Groups Amplify Noise · II)
- Informational Cascades
- Respectful listening: Charles defers to Arthur and Barbara not from cowardice but because he assumes they hold evidence he lacks.
- Cascade defined: Once David follows his predecessors without better information of his own, he is in an informational cascade.
- Private information lost: Late speakers' unique insights never surface, so early views go unchallenged and unrevised.
- Order decides outcome: Had Barbara spoken first for Sam, or Arthur preferred Julie, the group would have converged there instead.
- Not irrational: Following a growing majority is often sensible—the larger the crowd, the better the bet.
- The Two Failures of Cascades
- Illusion of collective wisdom: People underestimate how much apparent agreement merely reflects members copying earlier speakers.
- Terrible directions: A cascade can carry a group unanimously toward a judgment that is simply wrong.
- Noise across groups: Identical groups cascade to different answers purely because different speakers went first.
- Social Pressure Cascades
- Silence for belonging: People suppress dissent so as not to appear uncongenial, truculent, obtuse, or not a team player.
- Bandwagon effect: An early or powerful endorsement imposes strong pressure on later speakers to fall in line.
- Exaggerated conviction: Members read others' compliance as genuine preference, adding their own voice to the illusion.
- Unanimity without conviction: Groups can unanimously support a judgment that few participants privately hold.
- Divergent groups: Similar groups end up in different places because of who spoke first or who stayed silent.
- Group Polarization
- Polarization defined: Discussion moves people to a more extreme point in line with their original inclination.
- Two mechanisms: Dominant views generate more arguments on their side, and reputational caution pushes members toward them.
- Shift, not averaging: Groups become more unified, more confident, and more extreme than their median member.
- Beyond juries: Professional teams making judgments polarize too, not only criminal juries.
- Deliberating Versus Statistical Juries
- Statistical jury: The median of six independent judgments, aggregated mechanically, which substantially reduces noise.
- Averaging reduces noise: Combining independent judgments cancels errors; for that purpose, more judgments are always better.
- Deliberation increases noise: Real six-person juries that discussed the case were far noisier than matched statistical juries.
- Accentuation: Lenient medians produced more lenient verdicts; severe medians produced more severe ones.
- Severity drift: 27% of juries awarded as much as, or more than, their most severe member.
- Amplification chain: Level, pattern, and occasion noise in individuals grow louder once group dynamics work on them.
- Implications for Organizations
- First speakers hold outsized power: Outcomes hinge on the judgment of whoever happens to speak first.
- Manage deliberation: Leaders should reduce noise in individual judgments and design group processes not to amplify it.
- Warning signs: Early popularity deciding a new release, ideas spreading because others like them, teams uniting confidently around a chosen course.
- Informational Cascades
- Noisy Judgment and Bounded Disagreement (4. Matters of Judgment · I)
- Part III: Noise in Predictive Judgments
- Simple Models Outpredict Human Judgment (9. Judgments and Models · I)
- Measuring Predictive Accuracy
- Percent concordant (PC): the probability that the higher-rated case also performs better
- Reference points: perfect prediction gives PC = 100%, useless prediction only 50%
- Correlation coefficient (r): the standard measure, running from 0 to 1 for positive relations
- Shared determinants: r reads as the percentage of causal factors two variables have in common
- Height and foot size: r ≈ .60, PC = 71%; the two measures move in lockstep
- Clinical Judgment Versus a Formula
- Clinical judgment: consider the information, compute roughly, consult intuition — that is judgment
- Monica and Nathalie: two candidates rated on five dimensions; predict their performance two years on
- Study result: doctoral-level psychologists achieved r = .15 (PC = 55%), barely above chance
- Mechanical alternative: multiple regression on the same ratings reached r = .32 (PC = 60%)
- Least squares principle: the optimal weights minimize mean squared error of the prediction
- How Simple Models Work
- Fixed rule: identical weights apply to every case, with no case-by-case tailoring
- Linearity: a one-unit gain in a predictor always produces the same effect
- Weighting logic: predictors strongly correlated with the target get large weights, useless ones zero
- Negative weights: unpaid traffic tickets would predict against managerial success
- The workhorse: mechanical prediction spans simple rules to AI, but linear models dominate
- Meehl's Verdict and the Illusion of Validity
- 1954: Clinical Versus Statistical Prediction reviewed twenty studies pitting judgment against rules
- Consistent result: simple mechanical rules generally beat human judgment
- The sting: professionals are weakest precisely where they claim strength — integrating information
- Illusion of validity: confidence in evaluating cases is mistaken for confidence in predicting outcomes
- Cases versus predictions: judging who looks stronger is easy; knowing who will succeed is not
- Measuring Predictive Accuracy
- Simple Models Beat Noisy Judges (9. Judgments and Models · II)
- Meehl's Verdict and the Clinicians' Shock
- Meehl's finding: trivial formulas, consistently applied, outdo clinical judgment — clinicians met it with shock and contempt
- Experience over evidence: judgment feels valid subjectively, so we trust our experience over a scholar's claim
- Meehl misunderstood: a practicing psychoanalyst with Freud on his wall, not a hard-nosed numbers guy
- "Massive and consistent": Meehl's own words for the evidence favoring mechanical aggregation of inputs
- The 2000 review: across 136 studies, mechanical prediction won 63, tied 65, lost only 8
- Understated advantage: mechanical prediction is faster and cheaper, and humans often had private information models lacked
- The Model of the Judge
- Eugene, Oregon: Paul Hoffman's institute gathered researchers who made it a world center for judgment study
- Goldberg's twist: model the judge, not reality — same predictors, same regression, different target variable
- As-if theory: experts act as if using simple formulas, like billiard players solving mechanics equations
- How close: across 237 studies, the model of a judge correlated .80 with that judge's own judgments
- The ersatz beats the original: models of judges out-predict the professionals they were built from
- Ninety-eight for ninety-eight: every participant's model predicted graduate GPA better than the participant did
- Subtlety Versus Noise
- Two eliminations: the model of you deletes your subtle rules and your pattern noise
- Models only subtract: a model of your judgments cannot add information, only simplify what is there
- Valid subtlety lost: complex combinations, like skill times motivation, do beat weighted averages when truly valid
- Noise always hurts: removing noise from judgment always improves predictive accuracy
- The arithmetic: a .50 correlation with half the variance noisy would rise to .71 if noise-free
- The verdict: you believe you are subtler than the caricature, but you are mostly noisier
- Why Complex Rules Fail
- Rarely true: many complex rules people invent are not generally valid
- Unobserved conditions: even valid subtle rules apply in conditions too rare to be reliably checked
- Original candidates: exceptionally original hires are rare, originality scores unreliable, and flukes abound
- Measurement error: errors at both ends of a scale attenuate validity; true subtlety drowns
- Illusion of validity: complexity and richness feel like insight without improving predictions
- Mindless Consistency Beats Experts
- Yu and Kuncel: 847 executive candidates, seven dimensions, clinical overall scores — unimpressive results
- Ten thousand random models: random weights, applied consistently, tested against expert predictions
- The result: 77% of random models beat experts in one sample, 100% in the other two
- Blunt conclusion: it proved almost impossible to build a linear model that did worse than the experts
- Mindless consistency: mechanical consistency itself, not better weights, drives the improvement
- The caveat: this does not mean any model beats any human, but it shows noise's massive effect
- Meehl's Verdict and the Clinicians' Shock
- The Robust Beauty of Simple Rules (10. Noiseless Rules · I)
- The Algorithmic Advantage
- Broad definition: any process or set of rules followed in calculation counts as an algorithm, not just machine learning.
- Noise-free core: all mechanical approaches outperform human judgment partly because rules introduce no random variability.
- The spectrum: the journey runs from multiple regression toward extreme simplicity, then back toward greater sophistication.
- Dawes's Improper Linear Models
- Equal weights: instead of optimizing each predictor's weight, give every predictor the same weight.
- Heretical claim: Dawes found equal-weight models about as accurate as proper regression, and far superior to clinical judgment.
- Disbelief: editors refused the paper; the result seemed contrary to statistical intuition.
- Why Equal Weights Survive
- Overfitting: multiple regression minimizes error in the original sample, adjusting itself to every random fluke.
- Cross-validated correlation: true accuracy is measured on new data, so it almost always falls below in-sample performance.
- Small samples: flukes loom larger, so the advantage of "optimal" weighting disappears in most social-science research.
- Robust beauty: equal weights resist accidents of sampling; "we do not need models more precise than our measurements."
- Frugal Models and Simple Rules
- Correlated predictors: combining two correlated predictors is barely more predictive than the best one alone.
- Bail example: age plus missed court dates produced a risk score needing no computer, even no calculator.
- Matched complexity: the frugal rule equaled larger regression models and beat virtually all human bail judges.
- Recidivism: two inputs matched an existing tool that used 137 variables.
- Transparency payoff: simple rules are easy to apply at little cost in accuracy.
- Broken Legs and Machine Learning
- Promise of AI: huge data sets let machines spot relationship patterns no human could detect.
- Broken-leg exception: decisive private information the model cannot use justifies overriding its recommendation.
- Restraint rule: disagreeing without such information reflects an invalid personal pattern—do not override.
- Machine advantage: machine learning discovers far more broken legs than humans can think of.
- The Algorithmic Advantage
- When Algorithms Outpredict Human Judges (10. Noiseless Rules · II)
- Pattern Finding Without Understanding
- AI as pattern finding: machine learning involves no magic and no understanding, only detection of regularities in data
- Broken-leg patterns: like hospital visits predicting a missed movie night, decisive signals are learnable from vast data
- Less supervision: predicting rare events well reduces the need for human oversight
- The Bail Study: Better Predictions
- Scale: 758,027 bail decisions, using offense, rap sheet, and prior failures to appear
- Score, not verdict: the model outputs flight risk; the acceptable-risk threshold remains an evaluative judgment
- Striking gains: at equal detention rates crime falls up to 24%; at equal crime, detentions fall up to 42%
- Beyond linear: machine learning outperforms linear models by finding signal in combinations of variables
- Noise in Judges' Bail Decisions
- Level noise: the most lenient quintile released 83% of defendants; the least lenient only 61%
- Pattern noise: judges disagree about which defendants are high flight risks, not merely about severity
- Variance split: cases explained 67% of variance, system noise 33%—and 79% of that noise was pattern noise
- Accuracy Without Sacrificing Fairness
- Fairness risk: race-correlated predictors or biased training data can make an algorithm discriminate
- Bail result: at an equal crime rate, the algorithm jailed 41% fewer people of color than judges did
- Cowgill's résumés: algorithm-selected engineers were 14% more likely to get offers, 18% more likely to accept, and more diverse
- Proper weighting: humans favored résumés matching the "typical" profile; the algorithm weighted each predictor correctly
- Caveat: algorithms trained on biased outcomes, such as past promotions, replicate human bias
- Why We Resist Algorithms
- Persistent resistance: medicine, hiring, Hollywood, publishing, and sports still trust gut judgment
- Meehl's survey: seventeen rebutted objections, driven partly by fear of technological unemployment and dislike of computers
- Algorithm aversion: people often prefer algorithmic advice—until they see it make a mistake
- Perfection double standard: we forgive human errors but discard machines that err, trusting inferior judgment instead
- What Human Judges Can Borrow
- Emulate, not compete: we cannot match AI's information use, but can copy its simplicity and noiselessness
- Noise reduction: adopting methods that curtail system noise directly improves predictive judgment
- Guiding question: ask whether there is a broken leg, or whether you simply dislike the prediction
- Pattern Finding Without Understanding
- 11. Objective Ignorance
- The Internal Signal of Judgment Completion
- Knowing without knowing why: intuition arrives with an aura of conviction but no articulable reasons
- Signal as reward: reaching closure on a judgment produces a pleasing sense of coherence — the pieces fit
- Feeling mistaken for belief: "the evidence feels right" masquerades as rational confidence in validity
- Confidence ≠ accuracy: many confident predictions turn out wrong
- The deeper limit: the largest source of error is the ceiling on how good predictions could be
- Objective Ignorance Defined
- Two sources: intractable uncertainty (what cannot be known) and imperfect information (what could be known but isn't)
- Not bias or noise: ignorance is an objective property of the predictive task itself
- Terminology: "ignorance" distinguishes world-uncertainty from noise, the variability across judgments that should be identical
- Domain-dependent: doctors and lawyers often predict well, but predictability is generally lower than assumed
- Systematically underestimated: overconfidence is among the best-documented cognitive biases; limited information breeds excessive certainty
- The calibration gap: executives guess 75–85% concordance; personnel judges actually achieve a .28 correlation (PC = 59%)
- Overconfident Pundits
- Tetlock's study: nearly three hundred experts over two decades; their forecasts barely beat chance
- Dart-throwing chimpanzee: expert accuracy roughly equaled random choice among three outcomes
- Storytellers, not seers: experts spun compelling narratives but could not predict the future
- Confident and wrong: pundits with clear theories were the most confident and the least accurate
- Compounding ignorance: unforeseeable events have unforeseeable consequences, so ignorance grows with the time horizon
- Superforecasters: short-term forecasting is difficult but not impossible; some consistently outperform intelligence professionals
- Poor Judges and Barely Better Models
- Modest gap: models consistently beat people, but not by much; no case shows humans failing while models excel with the same information
- Hit rates: in the median study, clinicians were right 68% of the time, formulas 73%
- Correlations: clinicians achieved .32, formulas .56; even mechanical validity stays strikingly limited
- Bail algorithm: cuts crime by up to 24% — impressive, but perfect prediction would do far more
- Heart attack AI: 2,400 variables, 4.4 million visits; superior to physicians, yet top-decile risk was only 30%
- Conclusion: physician error is bounded at least as much by objective ignorance as by flawed judgment
- The Denial of Ignorance
- Denying the unpredictable: believing in the predictability of unpredictable events amounts to a denial of ignorance
- Explains Meehl's puzzle: decision makers keep trusting their gut because the internal signal rewards them emotionally
- Uncertainty invites intuition: leaders turn to gut feeling most when situations feel highly uncertain
- Dismissing modest gains: "if models aren't perfect, why bother?" — yet .44 versus .28 has real value
- The real price: intuition supplies an emotional certainty that imperfect algorithms cannot match
- Implication: algorithms will never be perfect everywhere, so human judgment won't be replaced — it must be improved
- The Internal Signal of Judgment Completion
- 12. The Valley of the Normal
- The Fragile Families Prediction Challenge
- Common task method: 160 teams competed to predict six life outcomes from a vast family database
- Single events stay unpredictable: best correlations of .17–.24 (PC = 55–58%)
- Aggregates fare better: child's GPA at .44, material hardship at .48
- Residual objective ignorance: rich data cannot buy granular prediction of individual lives
- Typical of social science: a review of 25,000 studies found effects averaging r = .21
- "Significant" misleads: it means unlikely by chance, not strong, large, or important
- Understanding and Prediction
- To understand is to describe a causal chain: claims of understanding are claims about causes
- Causation implies correlation: where causality exists, prediction must be possible
- Correlation gauges understanding: predictive accuracy measures how much causality we grasp
- Shoe size and math: correlation supports prediction without implying causation
- Stark admonition: researchers must reconcile "understanding life trajectories" with consistently inaccurate predictions
- Causal Thinking
- Causal thinking: builds stories in which specific people and events affect one another
- Inevitability illusion: an outcome feels like the logical end of a foreordained chain
- Alternate narratives: every fork could have gone otherwise, yet each ending feels equally explicable
- Cheap and natural: even low correlations become satisfying causal explanations
- Statistical thinking is costly: ensembles and base rates demand System 2 attention and specialized training
- The Valley of the Normal
- Valley of the normal: most events are neither actively expected nor especially surprising
- Understanding runs backward: an unanticipated event triggers a memory search for a candidate cause
- Search stops at the first good narrative: the opposite outcome would have explained itself just as well
- Self-explanatory events: the occurrence itself tells you its cause, and the cause fits perfectly
- Genuine surprise is rare: it occurs only when routine hindsight fails to explain
- Illusion of predictability: understanding the past inflates confidence in predicting the future
- Test question: we think we understand what is going on here — but could we have predicted it?
- Inside and Outside
- Inside view: the causal, case-specific mode; effortless, automatic, and our default
- Outside view: the statistical mode; each case is one instance of a broader class
- Noise goes unseen: being a statistical notion, it never surfaces in causal narratives
- Resolved uncertainty is erased: memories of past doubt fade once the outcome is known
- The Fragile Families Prediction Challenge
- Simple Models Outpredict Human Judgment (9. Judgments and Models · I)
- Part IV: How Noise Happens
- Heuristics, Substitution, and Predictable Bias (13. Heuristics, Biases, and Noise · I)
- Fast Thinking and the Heuristics and Biases Program
- Heuristics: System 1 answers difficult questions with simplifying operations of fast, intuitive thinking
- Usually useful and adequate, but they sometimes produce biases — systematic, predictable errors
- The program studied what people share, not how they differ: error-producing processes are widely shared
- Shared biases create statistical bias; differing biases create system noise — both create error
- This book extends half a century of research reviewed in Thinking, Fast and Slow
- Diagnosing Bias When Truth Is Unknown
- Bias is measured against a true value — e.g., the planning fallacy underestimates completion times
- Without a true value: find an irrelevant factor that has an effect, such as sleek paper swaying investors
- Or find a relevant factor that has none — scope insensitivity for two- versus three-year job survival
- Reserve "bias" for specific, identifiable errors and the mechanisms that produce them
- Blaming every failure on unspecified "bias" means only that "mistakes were made"
- Substitution: Similarity for Probability
- The Bill puzzle: people give identical answers to probability and resemblance questions
- Substitution: a heuristic answers the hard question by answering an easier one instead
- Similarity is easier to judge than probability, so similarity is what people actually assess
- Added detail can raise representativeness while lowering probability — the conjunction fallacy
- Venn-diagram logic constrains probability, not similarity, so substitution yields predictable errors
- Base-Rate Neglect and the Outside View
- The Gambardi case: matching a description to a success image neglects base rates
- Outside view: treat the case as one member of a class and think statistically about it
- CEO turnover near 15% annually gives a new CEO roughly 72% survival at two years
- Base rates are only a starting point, but ignoring them discards key information
- The authors themselves missed this bias in their own case — base-rate thinking is not automatic
- Availability and the Reach of Substitution
- Availability heuristic: ease of recalling instances substitutes for judging frequency
- Risk perception spikes after publicized airplane crashes and hurricanes; recent events get extra weight
- Substitution is general-purpose, extending far beyond probability and frequency
- Climate belief becomes do I trust the people; surgeon competence becomes does she sound authoritative
- Project schedule becomes is it on schedule now; nuclear necessity becomes do I recoil at "nuclear"
- Fast Thinking and the Heuristics and Biases Program
- Biases That Create Error and Noise (13. Heuristics, Biases, and Noise · II)
- Substitution and Misweighting
- Substitution: answering an easier question than the one asked, then treating the answer as if it fits.
- Misweighting: substituting one question for another gives evidence the wrong weights, inevitably producing error.
- Mood for life satisfaction: a full answer requires more than current mood, yet mood is heavily over-weighted.
- Similarity for probability: judging by similarity neglects base rates that are properly irrelevant to similarity judgments.
- Irrelevant aesthetics: document polish should carry little weight in valuing a company; any effect is error.
- Conclusion Bias: Prejudgment
- Prejudgment: judgment often begins with an inclination toward a particular conclusion rather than a search.
- Lucas and Kasdan: "I don't like that and I don't believe that" shows not liking preceding not believing.
- System 1 first: intuition suggests a conclusion; System 2 is then mobilized to argue for it.
- Confirmation and desirability bias: evidence is collected and interpreted selectively to fit what we believe or wish.
- Rationalization: people invent plausible reasons and then believe those reasons caused their belief.
- Test it: imagine the supporting arguments suddenly proven invalid — a genuine judgment would change.
- Affect Heuristic and Anchoring
- Affect heuristic: people decide what they think by consulting their feelings, liking most things about favored politicians.
- Brand and ratings: companies cultivate positive affect; high teaching marks also lift ratings of identical readings.
- Anchoring effect: an arbitrary number, even your Social Security digits, shifts a quantitative judgment.
- Robust and involuntary: recipients of an anchor automatically seek reasons to make the number plausible.
- Negotiating edge: going first with an anchor is an advantage, drawn from the bazaar to the boardroom.
- Excessive Coherence
- Excessive coherence: we form coherent impressions quickly and revise them too slowly.
- Order matters: Intelligent, Persistent, Cunning, Unprincipled reads far better than the same words reversed.
- Halo effect: an early positive impression colors evaluation of everything that follows.
- Calorie labels: labels placed left of the item matter more where people read left to right — reversed for Hebrew readers.
- First reactions dominate: whichever comes first, the food or the calorie count, shapes the choice.
- Inadmissible evidence: jurors cannot truly unhear testimony; instruction to disregard is not realistic.
- When Biases Become Bias and Noise
- Shared bias, no noise: if every respondent makes the same substitution error, the result is bias without variability.
- Substitution noise: "Do I trust those who say climate change is real?" varies by social circle and politics.
- Occasion noise: answering life satisfaction from mood makes the same person's answer vary by the hour.
- Prejudgment disparity: one asylum judge admits 5%, another 88% — differing biases create massive system noise.
- Coherence two ways: identical sequences give shared bias; haphazardly ordered data gives random distortion.
- Debiasing: psychological biases are universal, so reducing them improves judgment — elaborated in part 5.
- Substitution and Misweighting
- 14. The Matching Operation
- Matching: The Versatile Core of Judgment
- Matching defined: finding a value on a judgment scale that fits a subjective impression
- Effortless conversion: the sky's darkness becomes a probability of rain without deliberation
- Cross-dimension mapping: enthusiasm for a singer can be matched to a building's height
- Context sets the range: "large" or "a lot of money" means nothing outside a frame of reference
- Unlimited targets: given a thin description, we can match it to nearly any question asked
- Matching and Coherence
- Bill's professions: probability questions are silently answered as similarity judgments
- Conflicting cues: incompatible details block a coherent impression and make matching harder
- Noise locus: complex judgments with conflicting evidence are exactly where noise concentrates
- Two scale types: qualitative scales are unordered; quantitative scales always answer "which is more"
- The Bias of Matching Predictions
- Julie's GPA: reading precocity at age four instantly suggests 3.7–3.8
- Two-step mechanism: evaluate the intensity of the evidence, then match a prediction to that intensity
- Substitution: an easy evaluation question answers a hard prediction question
- Nonregressive error: matching assumes perfect correlation, ignoring later life events
- Outside view: anchor on the class average, move toward the evidence only as far as it warrants
- Costly illusion: hit directors, star hires, and top salespeople are expected to repeat themselves
- Optimism, Pessimism, and System 2
- Extreme in both directions: matching predictions overshoot when hopeful and when bleak
- Asymmetry: people resist unfavorable matching predictions until more information arrives
- Veto power: System 1 proposes the match, but System 2 must endorse it before it becomes belief
- Noise in Matching: Limits of Absolute Judgment
- Magical number seven: we can assign only about seven distinct labels on an intensity scale
- Errors past the limit: length, loudness, and brightness judgments begin reversing beyond seven
- Label versus compare: judging cases side by side is far more sensitive than rating them alone
- Hierarchical refinement: finer categories help only if defined in advance with clear boundaries
- Ranking remedy: sort into five categories, rank within each, then subdivide further
- Accountability: comparative ratings like "top 20%" work only when misuse of the scale is detectable
- Matching: The Versatile Core of Judgment
- 15. Scales
- The Punitive Damages Study
- Joan Glover case: a jury scenario built as a noise audit, testing how people punish corporate wrongdoing
- Punitive damages: supplemental awards meant to deter and warn, notorious for their unpredictability
- Three scales, one question: outrage, punitive intent, and dollar awards measure the same severity in different units
- Design: 899 participants each rated a single scale across ten cases, never comparing scenarios side by side
- The Outrage Hypothesis
- Substitution: people answer the easy "how angry am I?" instead of the hard question of just punishment
- Near-perfect correlation: mean outrage and punitive intent ratings correlated 0.98, confirming outrage drives intent
- Harm's asymmetric role: greater harm raises punishment but leaves outrage almost unchanged
- Automatic weighting: raters unknowingly took a stand on retributive justice, weighting evidence by task alone
- Ambiguity Breeds Noise
- Variance decomposition: judgment variance = variance of just punishments + level noise + pattern noise
- Noise share: punitive intent 51% noise, outrage 71%, dollars 94% — identical content, wildly different reliability
- Bounded scales: "extremely severe punishment" is clearer than "absolutely outrageous" because law caps it
- Core lesson: ambiguous scales are inherently noisy; the choice of scale shapes how much noise enters judgment
- Dollars and Anchors
- Ratio scales: dollars have a meaningful zero and no upper bound, unlike verbal rating scales
- Modulus: a single intermediate anchor, such as $1.5 million, ties down the entire dollar scale
- Coherent arbitrariness: one arbitrary anchor fixes absolute values while relative judgments stay consistent
- Anchorless scaling: without guidance, people arbitrarily anchor on their first answer and stay consistent thereafter
- Real juries: see only one case, receive no anchor, and the arbitrary number becomes the verdict
- Ranking Reveals Agreement
- Rank transformation: converting dollar awards to ranks eliminates juror-level scale errors
- Noise drop: dollar-award noise falls from 94% to 49% once awards are ranked
- Hidden agreement: rankings show jurors largely agree on relative punishment; absolute dollars are meaningless
- The Unfortunate Conclusion
- Psychological nonsense: the law assumes jurors can move directly from offense to the correct dollar punishment
- Institutional fix: justice systems should acknowledge human limits and supply reference anchors
- Lesson one: choosing a scale can sharply change the amount of noise in judgments
- Lesson two: replacing absolute judgments with relative ones, where feasible, reliably reduces noise
- The Punitive Damages Study
- 16. Patterns
- Ambiguity Makes Hard Problems
- Conflicting cues: hard judgment problems are defined by ambiguity — multiple cues pointing in different directions.
- Easy problems: one cue, or several in agreement, lets matching and System 1 deliver a fast, shared answer.
- Coherent stories: when evidence paints a unified picture, judgment is easy and most people agree.
- Noise from ambiguity: if there is more than one way to see a case, people will differ; complex problems are noisier.
- Julie 2.0: incoherent life stories are realistic, yet they force judges to select different evidence as their narrative core.
- Confidence Is Not Accuracy
- Two conditions: confidence requires a comprehensively coherent story and no attractive alternative interpretations.
- Cheap coherence: you can also feel confident by ignoring or explaining away the details that do not fit.
- Real expertise: the true expert can explain why rival stories are wrong, not merely why her own is right.
- Illusion of agreement: suppressing alternatives leads people to assume other observers must reach the same conclusion.
- Overconfidence: subjective confidence guarantees nothing about accuracy, and most of us are more confident than we should be.
- Decomposing Pattern Noise
- Pattern error: an individual's deviation on a case not explained by the separate effects of case and judge.
- Example: the normally lenient judge who is unusually severe with traffic offenders, or unusually lenient toward young women.
- Equation: pattern error = stable pattern error + transient (occasion) error.
- Variance decomposition: pattern noise squared = stable pattern noise squared + occasion noise squared, graphed as a right triangle.
- Sources of Stable Pattern Noise
- Model of the judge: each recruiter weights the same ratings differently, producing different candidate rankings.
- Idiosyncratic associations: personal memories make Julie evoke different people in different judges — stable but unpredictable.
- Hidden accuracy: a forecaster with genuine insight deviates from the average, and without outcome data that looks like pattern noise.
- Differential skill: specialists attending to different dimensions of a case (skills, injuries, motivation) generate pattern noise.
- Skills as asset: when teams judge together, diversity of skills complements rather than merely adds noise.
- Two Lotteries of Judgment
- First lottery: which professional you draw brings a kaleidoscope of values, beliefs, memories, and associations — your unique baggage.
- Second lottery: the moment, mood, and recent circumstances of the judgment produce occasion noise.
- No sharp line: stable and transient pattern noise differ only in whether the trigger is permanent or fleeting.
- Newspaper example: a fresh article on campus drug use makes Julie's pot smoking loom larger — transient, not durable.
- The Personality Analogy
- Weak prediction: broad traits correlate only about .30 with specific behaviors; situations matter strongly.
- Non-additive joining: behavior is a function of personality and situation, not a mechanical sum.
- Signature patterns: people differ in which situations trigger which responses, and those signatures are stable.
- Level vs. pattern: equal average aggressiveness can coexist with entirely different profiles of reaction.
- Cases as situations: ranking judges by quality varies by case because sensitivity to features differs.
- Professional caveat: personality uniqueness is a cause for celebration, but in professional judgment variation is error.
- Ambiguity Makes Hard Problems
- 17. The Sources of Noise
- The Nested Breakdown of Error
- Three successive splits: error into bias and system noise; system noise into level and pattern noise; pattern into stable pattern and occasion noise.
- MSE decomposition: total error squared equals the squares of bias plus three distinct noise components.
- Core conclusion: noise is often a larger component of error than bias, and far more worthy of study.
- Pattern Noise Outweighs Level Noise
- Insurance underwriters: differences in average premiums set accounted for only 20% of system noise; 80% was pattern noise.
- Federal judges: level noise was slightly less than half of total system noise; pattern noise was larger.
- Punitive damages: pattern noise held roughly constant across three scales — 63%, 62%, and 61%.
- Hidden magnitude: level noise is the only component organizations can monitor without audits, so shocking findings understate the problem by at least a factor of two.
- Exception: asylum judge variability is almost certainly driven more by level noise, though pattern noise is likely large there too.
- Stable Pattern Noise Dominates Occasion Noise
- Abandoned default: statistics treated residual pattern error as random occasion noise, constraining thinking for years.
- Sentencing implausibility: if all pattern noise were transient, the same judge would vary by about 2.8 years on identical cases — grotesque.
- Consistency without consensus: Todorov's face-rating studies found stable pattern noise largest, level noise second, occasion noise smallest.
- Bail judge simulation: 173 modeled judges deciding 24 million cases showed stable pattern noise nearly four times level noise (26% vs. 7%).
- Triangulation required: few studies capture all error components at once, so conclusions remain tentative but consistent.
- Noise Is Interaction, Not Just Level
- Level noise is tractable: abnormally tough graders or cautious officers can be equalized, as with mandated grade distributions.
- What level noise misses: noise is mostly how different judges, teachers, and social workers deal with particular cases, students, and families.
- Noise as uniqueness: it is largely a by-product of our individual judgment personality, not across-the-board severity.
- Incomplete fix: reducing level noise alone leaves most of the system noise problem unsolved.
- Why Noise Stays Invisible
- Valley of the normal: successful outcomes are unsurprising, self-explaining, and seldom questioned.
- Explaining the abnormal: failures invite causal stories — fundamental attribution error, hindsight, motives, incompetence.
- Vacuous bias claims: an explanation that applies only to one event, like after-the-fact overconfidence, provides satisfying illusion without prediction.
- Rosenzweig's business critique: empty bias explanations proliferate because we crave causal stories.
- Causal Thinking Versus Statistical Thinking
- Noise is statistical: it becomes visible only when thinking statistically about an ensemble of similar judgments.
- Causally nowhere: no single case reveals noise; statistically it is everywhere — the scatter of hits on the target.
- Figure and ground: bias is the compelling figure; noise is the background we ignore.
- Effort required: causes come naturally, statistical thinking must be learned and remains difficult.
- The Nested Breakdown of Error
- Heuristics, Substitution, and Predictable Bias (13. Heuristics, Biases, and Noise · I)
- Part V: Improving Judgments
- What Makes a Judge Good (18. Better Judges for Better Judgments · I)
- The Three Ingredients of Good Judgment
- Judges differ: some people reliably out-perform others in any task requiring judgment
- Three levers: good judgment depends on what you know, how well you think, and how you think
- Training, intelligence, cognitive style: together they make judgments both less noisy and less biased
- Better crowds: wisdom-of-crowds aggregation improves when the crowd is made of more able people
- Open-mindedness: good judges are experienced and smart, but also actively learn from new information
- True Experts: Skill You Can Verify
- Skill shows: skilled radiologists diagnose pneumonia better; superforecasters out-predict their peers
- Convergent mastery: specialist lawyers make similar, and good, predictions about routine legal disputes
- Double benefit: highly skilled judges are less noisy and also show less statistical bias
- Track records: outcome data lets you choose whoever has been right most often in the past
- Practical caveat: verification is often hard; no one should subject their family doctor to a proficiency exam
- Respect-Experts: Credentialed by Esteem
- Unverifiable judgments: underwriting, sentencing, wine tasting, and essay grading have no uncontested true value
- Respect-experts: professionals whose credibility rests entirely on the respect of their peers
- Not a criticism: many professors, scholars, and consultants legitimately work in this regime
- Productive disagreement: rival analysts, or Kantian, Benthamite, and Aristotelian philosophers, are all respected—yet not all are right
- Doctrine, Experience, and the Seeds of Noise
- Shared norms: professional doctrine teaches experts which inputs count and how to justify judgments
- Doctrine is not a recipe: it leaves room for interpretation, so identically trained experts drift apart
- Noise is inevitable: doctrine yields judgments, not computations; claims adjusters agreed on the checklist yet diverged widely
- Experience required: credibility takes years—there are no young prodigies in underwriting
- Confidence, Coherence, and Intelligence
- Confidence heuristic: groups give more weight to confident people, even when confidence is groundless
- Coherent stories: experience enables pattern recognition, reasoning by analogy, and fast hypothesis confirmation
- The trap: training, experience, and confidence command trust without guaranteeing the quality of judgment
- Intelligence: general intelligence correlates with good performance in virtually all domains
- The Three Ingredients of Good Judgment
- Selecting Judges by Ability and Open-Mindedness (18. Better Judges for Better Judgments · II)
- GMA: The Strongest Single Predictor
- GMA's predictive power: predicts occupational level and performance within an occupation better than any other trait or disposition
- Crystallized vs. fluid intelligence: solving from stored knowledge versus solving wholly novel problems
- Complexity matters: test-performance correlations reach .50 for complex jobs—very strong by social-science standards
- Innateness misconception: tests measure developed abilities, shaped by both heritable traits and environment
- Beyond GMA: conscientiousness, grit, practical intelligence, and creativity also matter, but predict less
- The Threshold Myth
- No ceiling effect: within the top 1% of ability, the top quartile is two to three times likelier to earn a doctorate, publish, or patent
- Hyper-elite: Fortune 500 CEOs and billionaires come from the most able—and ability still tracks pay and net worth
- Dropout illusion: famous billionaire dropouts hide the forest; 88% of billionaires earned college degrees
- Practical upshot: if you must choose judges, choosing the highest mental ability makes a lot of sense
- Why Ability Alone Misleads
- Guesswork problem: you cannot test everyone, so you must infer who is high-GMA
- Persuasion confound: high GMA also improves convincing others, so the able become respect-experts without reality feedback
- Astrologer warning: medieval astrologers were plausibly among the highest-GMA people of their time
- Trusting eloquence: sensible but insufficient—articulate confidence may backfire as a selection strategy
- Cognitive Style: Reflective Versus Impulsive
- Cognitive reflection test: bat-and-ball items measure overriding the first, wrong answer that comes to mind
- Real-world correlates: low CRT scores predict belief in astrology, ESP, and falling for blatant fake news
- System 1 vs. System 2: CRT indexes the propensity for slow reflective thought over fast impulse
- Need for cognition: high scorers enjoy mental effort and resist biases—but self-reports are transparently gameable
- Skill-based alternatives: Adult Decision Making Competence and Halpern assessments predict fewer adverse life events
- Actively Open-Minded Thinking
- Surprising null: need for cognition did not predict who forecasts better or seeks more information
- The lone predictor: Baron's actively open-minded thinking scale was the only style measure predicting forecasting performance
- Core behavior: actively search for information contradicting your preexisting hypotheses, including dissent
- Humility: judgment as work in progress; being convinced by an opposing argument signals good character
- Teachable: evidence suggests open-mindedness can be learned; it marks the very best forecasters
- Choosing Better Judges
- Domain matters: treat judgments verifiable against true values differently from those of respect-experts
- Chess vs. pundit: a timid grandmaster deserves more trust than an articulate political analyst
- Individual differences: some judges are better than equally qualified peers—less biased and less noisy
- The profile to seek: people who methodically seek disconfirming information and eagerly change their minds
- Counter-stereotype: openness, not bone-deep certainty, reduces error; decisiveness belongs at the end of a process
- GMA: The Strongest Single Predictor
- 19. Debiasing and Decision Hygiene
- Two Ways to Debiase a Judgment
- Measurement analogy: fix a biased scale by adjusting the dial or subtracting the known error from each reading
- Ex ante debiasing: intervene before the judgment, by reshaping the environment or training the judge
- Ex post debiasing: correct after the fact—buffer a three-month estimate by planning for four
- HM Treasury's Green Book: applies historical optimism-bias percentages to project costs and timelines
- Known direction of error: correction works when bias is a clear, predictable statistical tilt
- Nudges and Boosting
- Nudges: modify choice architecture so the right decision becomes the easy default
- Automatic enrollment: overrides inertia and optimism bias, sharply raising pension participation
- Save More Tomorrow: earmarks future wage increases for savings before they are ever felt
- Salience and friction: make hidden fees visible; cut administrative burdens on access to care
- Boosting: train statistical literacy and general reasoning capacity
- Transfer problem: forecasters calibrated on rain stay overconfident on general-knowledge questions
- The Limitation of Bias-Specific Debiasing
- One-bias assumption: most interventions target a single bias assumed to be present
- Wrong target: a team burned by a past project may err conservatively, the opposite of expectation
- Offsetting forces: loss aversion can cancel overconfidence; status quo bias can mute it
- Unpredictable net effect: with several biases at work, the direction of error is unknowable
- Where debiasing fails: not when biases vary among judges—precisely the situation that produces noise
- Real-Time Detection: The Decision Observer
- Bias blind spot: people see biases in others far more easily than in themselves
- Decision observer: watches a live decision with a checklist, flagging biases as they emerge
- Prerequisite: leaders must initiate and back the role; a self-appointed observer wins no friends
- Three models: supervisor, internal "bias buster," or outside facilitator—each with real tradeoffs
- Checklist power: agencies ignored a fifty-page rulebook until a 1.5-page checklist enforced it
- Customization: checklists target the most frequent and consequential biases, not every known one
- Decision Hygiene: Preventing Invisible Error
- Bias vs. noise: bias is visible and directional; noise is unpredictable error we cannot see or explain
- Hygiene analogy: handwashing prevents unnamed germs; decision hygiene prevents unspecified errors
- Invisible victory: noise-reduction procedures prevent many errors, but you never learn which ones
- Thankless by design: like handwashing, compliance is hard to enforce because the benefit is unseen
- Worth the battle: despite its invisibility, the damage noise causes justifies the discipline
- What follows: forensic science, forecasting, medicine, and HR supply the concrete hygiene strategies
- Two Ways to Debiase a Judgment
- Fingerprint Judgment, Bias, and Noise (20. Sequencing Information in Forensic Science · I)
- The Mayfield Case
- Madrid, March 2004: bombs on commuter trains killed 192 people and injured over 2,000.
- False match: a latent print on a plastic bag went out via Interpol; days later the FBI identified it as Brandon Mayfield's.
- Plausible suspect: a Muslim convert, former army officer, and lawyer who had represented Taliban-linked clients; already on a watch list.
- Unraveling: Spain had already called the print a negative match; surveillance and arrest produced no evidence and no charges.
- Aftermath: freed after two weeks, he received $2 million and an inquiry blaming human error, not method or technology.
- Fingerprint Examination as Judgment
- Latent prints: smudged, partial, overlapping impressions from a scene, unlike clean prints taken under controlled conditions.
- Not mechanical: matching a latent print to an exemplar requires expert human judgment, not an automated scan.
- ACE-V: analysis, comparison, evaluation, then verification—yielding identification, exclusion, or inconclusive.
- Asymmetric verification: only identification decisions trigger review by a second examiner.
- Origins: Faulds, Galton, Bertillon, and Vucetich built the systems; Vucetich made the first crime-scene comparison in 1892.
- The Myth of Infallibility
- Untouchable evidence: until 2002, fingerprint evidence had never been successfully challenged in a US courtroom.
- FBI's own words: fingerprints were advertised as "an infallible means of personal identification."
- Errors excused: the rare failures were blamed on incompetence or deliberate fraud, never on the method.
- Ground-truth gap: who actually left the print is usually unknown, so a suspect who disputes it makes the evidence look stronger.
- Noise still measurable: not knowing the truth is no obstacle to measuring how often examiners disagree.
- Occasion Noise
- Dror's audits: a cognitive neuroscientist tested a field that assumed it had no noise problem.
- Definition: occasion noise is the same expert judging the same evidence differently on a later occasion.
- Ideal test bed: prints are hard to memorize, and the studies ran inside examiners' routine casework, unnoticed.
- Stakes of self-inconsistency: if experts are not reliable with themselves, the basis of their professionalism is in question.
- Forensic Confirmation Bias
- Biasing context: examiners were told the suspect had an alibi, or that firearms evidence excluded him—or that he had confessed.
- Judgments moved: four of five experts reversed an earlier identification in one study; four of twenty-four decisions changed in another.
- Perception changed too: biased examiners literally observed fewer minutiae when a target exemplar accompanied the latent print.
- Reinforcing conditions: shifts are likelier on difficult decisions, with strong context, and toward inconclusive verdicts.
- Beyond fingerprints: documented in blood pattern and arson analysis, skeletal remains, forensic pathology, and complex DNA mixtures.
- Independence and Decision Hygiene
- Built-in exposure: transmittal letters and direct contact with police, prosecutors, and other examiners feed biasing information.
- Broken safeguard: verifying examiners know the first conclusion was an identification, so ACE-V aggregation is not truly independent.
- Mayfield cascade: three FBI experts concurred, the first apparently swayed by the automated database's powerful correlation.
- Remedy: control the flow of information reaching the examiner—sequencing what is known, and when, is the decision hygiene.
- The Mayfield Case
- Bias Cascades and Sequenced Information (20. Sequencing Information in Forensic Science · II)
- Bias Cascades Across Forensic Disciplines
- Mayfield case: one erroneous identification tainted every subsequent examination
- Authority amplifies error: a respected supervisor's conclusion made disagreement nearly impossible
- Even the defense expert concurred, confirming the FBI identification
- Cross-discipline contagion: firearms and odontology experts biased by knowing other results
- Bias cascade: one expert's error becomes the next expert's biasing information
- Occasion Noise in Fingerprint Judgment
- Self-inconsistency: examiners sometimes reverse judgments on prints they have seen before
- FBI 2012 replication: 72 examiners re-evaluated 25 print pairs about seven months later
- One decision in ten altered, mostly to or from inconclusive; none false identifications
- Troubling implication: convictions may rest on identifications inconclusive on another day
- Inconsistency persists even when context is held as constant as possible
- How Much Error? The Available Evidence
- Innocence Project review: misapplied forensic science contributed to 45% of 350 exonerations
- Wrong question: jurors need error rates, not exoneration statistics
- PCAST 2016 report: the most authoritative review of forensic science in criminal courts
- FBI 2011 study: 169 examiners, ~100 pairs each, false-positive rate about one in six hundred
- Low but surprising: far higher than jurors believe given claims of fingerprint infallibility
- Likely underestimate: test conditions lacked bias and examiners knew they were tested
- Florida study found much higher false positives; the literature demands more research
- Caution as an Asymmetry of Error
- Mated pairs: fewer than one-third were accurately judged as identifications
- Asymmetric costs: erroneous identification is the deadly sin; other errors are cheaper
- Exclusion equals inconclusive: a missing print does not exonerate; a present print convicts
- Biased direction: easier to push examiners toward "inconclusive" than toward "identification"
- The Bias Blind Spot in Forensics
- Fingerprint Society chair: biased examiners "should seek employment in Disneyland"
- Case information defended as personal satisfaction that leaves judgment unaltered
- FBI internal review denied that knowing prior examiners' results influences conclusions
- Survey of 400 scientists: 71% see cognitive bias as a field-wide concern, only 26% in themselves
- Noise is invisible, even to people whose job is seeing the invisible
- Sequencing Information as Decision Hygiene
- Core strategy: sequence information to prevent premature intuitions
- More information is not better when it biases judgment
- Linear sequential unmasking: reveal only what examiners need, when they need it
- Document each step: analyze the latent print before viewing exemplars; record every change of mind
- Independent verification: the second examiner must not know the first's judgment
- Broad applicability: shield judgments from obvious occasion-noise triggers, including accurate information
- Bias Cascades Across Forensic Disciplines
- Selection and Aggregation Beat Noisy Forecasts (21. Selection and Aggregation in Forecasting · I)
- Forecasting's Twin Enemies: Bias and Noise
- High stakes: forecasts of unemployment, elections, climate, and earnings drive major private and public decisions
- Bias vs. noise: analysts sharply distinguish systematic error (bias) from inconsistency (noise) in forecasts
- Official optimism: agencies routinely project unrealistically high growth and unrealistically low deficits
- Motive is irrelevant: whether bias comes from politics or cognition, the resulting error is the same
- Overconfidence and Forecast Noise
- Overconfidence: forecasters set confidence intervals far narrower than reality warrants
- CFO failure: an 80% confidence interval for S&P returns contained the actual outcome only 36% of the time
- Occasion noise: forecasters disagree with their own earlier judgments
- Between-person noise: specialists disagree with each other — law professors on Supreme Court rulings, economists on growth
- Amplitude of disagreement: benefit projections for air pollution regulation ranged from $3 billion to $9 billion
- Averaging: The Engine of Crowd Wisdom
- Two levers: select better judges, and aggregate multiple independent estimates
- Averaging law: averaging divides noise by the square root of the number of judgments
- Scale of gain: averaging 100 judgments cuts noise 90%; 400 cuts it 95%
- Bias untouched: averaging reduces only noise, so its benefit depends on the bias/noise mix
- Independence matters: the wisdom of crowds works best when judgments are independent and lack shared bias
- Evidence: unweighted consensus forecasts beat most and sometimes all individual forecasts
- Structured Aggregation: Select-Crowd, Markets, Delphi
- Select-crowd: averaging the five best recent performers can equal straight averaging, and reassures expertise-respecting decision makers
- Prediction markets: incentivized betting aggregates views well; 70% prices correspond to roughly 70% occurrence
- Delphi method: anonymous multi-round estimates with reasons converge, combining aggregation with social learning
- Mini-Delphi: estimate-talk-estimate runs the same logic inside a single meeting
- Combining methods: averaging forecasts from different methods cut errors by 12.5% across thirty comparisons
- The Good Judgment Project
- Origin: Tetlock, Mellers, and Moore recruited tens of thousands of ordinary volunteers in 2011
- Scope: hundreds of geopolitical questions, posing the same forecasting problems as mundane business judgments
- Many forecasts: performance averaged across numerous predictions, so luck cannot explain success
- Probabilistic form: probability estimates, not binary calls, suit a world of objective ignorance
- Continuous updating: forecasters revise as news arrives; each update is scored as a new forecast
- Calibration, Resolution, and Superforecasters
- Calibration: a forecaster is well calibrated when events rated 60% likely happen 60% of the time
- Calibration's trap: always predicting 60% rain is well calibrated but practically useless
- Resolution: the willingness to take bold, differentiated stands that actually discriminate among outcomes
- Brier score: rewards calibration and resolution together; based on mean squared error logic, lower is better
- Superforecasters: most volunteers did poorly, but about 2% predicted far better than chance
- Forecasting's Twin Enemies: Bias and Noise
- Superforecaster Skill, Training, and Aggregation (21. Selection and Aggregation in Forecasting · II)
- What Sets Superforecasters Apart
- Better than intelligence analysts: trained professionals with classified data still underperform superforecasters.
- Intelligence helps, but modestly: superforecasters score higher on cognitive tests, yet many high scorers never qualify.
- Analytical style over math talent: their edge is thinking probabilistically, not raw numerical skill.
- Disaggregation: they split big geopolitical questions into subsidiary questions—what would make the answer yes or no?
- Outside view first: they anchor on base rates (how often do border disputes escalate?) before weighing current specifics.
- Active open-mindedness: they seek disconfirming evidence and update on new information without overreacting.
- Perpetual Beta
- Tetlock's strongest predictor: perpetual beta—commitment to belief updating and self-improvement—outranks all else.
- Doing, not being: excellence comes from research, self-criticism, synthesizing others' views, and relentless updating.
- The cycle: try, fail, analyze, adjust, try again—an endlessly improvable process, never a finished one.
- Three Interventions Tested
- Training: taught probabilistic reasoning, biases, averaging diverse predictions, and reference classes.
- Teaming (aggregation): forecasters debated one another's predictions, forcing engagement with opposing arguments.
- Selection: top 2% by accuracy were designated superforecasters and worked together the next year.
- Ranked effect: training helped, teaming helped more, selection helped most—confirming aggregation and selection.
- The BIN Model: Where the Gains Come From
- Three error sources: forecasts differ by information skill, statistical bias toward change or stability, and noise.
- Surprising verdict: every intervention improved accuracy mainly by suppressing random error—noise, not bias.
- Training paradox: designed to fight psychological biases, it worked by reducing noise—since those biases hit people differently, they produce noise, not statistical bias.
- Teaming's bonus: it also improved information extraction, because several brains find more signals than one.
- Selection's lesson: superforecasters owe success more to discipline in tamping down measurement error than to incisive news reading.
- Combining Selection and Aggregation
- Build hybrid teams: select judges who are both highly valid and complementary to one another.
- Multiple regression analogy: after the best predictor, add the test or judge that adds the most non-redundant predictive power.
- Diversity beats redundancy: a moderately valid but different judge may outrank a similar, stronger one.
- Four-witness image: independent witnesses seeing a crime from four angles give far better pooled information.
- Paradox of noise: a diverse, disagreeing team's average beats a unanimous team's average.
- The Independence Caveat
- Aggregation requires independence: pooled judgments reduce noise only if they are truly independent.
- Deliberation can backfire: group discussion often adds more bias than the noise it removes.
- Welcome disagreement: organizations must let members form judgments separately before combining them.
- Cheapest hygiene: eliciting and aggregating independent, diverse judgments is often the easiest and broadest strategy.
- What Sets Superforecasters Apart
- The Magnitude of Medical Diagnostic Noise (22. Guidelines in Medicine · I)
- Judgment Versus Calculation in Diagnosis
- Paul's white coat syndrome: years of futile medication dissolved once a new doctor questioned the diagnosis
- Diagnosis is judgment: deciding whether a patient is ill is inferential, and therefore noisy
- Routine calls are quiet: dislocated shoulders, tendon degeneration, and breast core biopsies show little noise
- Replacing judgment with tests: strep antigen, HbA1c, and COVID testing collapse noise by making diagnosis mechanical
- Second opinions: sometimes mandatory; divergence between two doctors is noise, though not proof of who errs
- The surprise is magnitude: not that medical judgment is noisy, but how large the noise turns out to be
- Measuring Agreement: The Kappa Statistic
- Kappa coefficient: medical interrater reliability is measured by agreement above chance
- Scale of noise: 1 is perfect agreement; 0 is monkeys throwing darts at a diagnosis list
- Typical findings: reliability across many diagnostic domains is only "slight," "poor," or "fair"
- Noise persists in technical fields: drug interactions, kidney staging, breast pathology, and spinal MRI all show fair agreement
- Diagnosis as lottery: whether a serious disease is found can depend on which doctor happens to be seen
- Documented Noise Across Specialties
- Coronary angiograms: physicians disagreed 31% of the time on whether a vessel was over 70% blocked
- Endometriosis: 108 surgeons shown three laparoscopy videos disagreed sharply on lesion number and location
- Tuberculosis: seventy-five years of chest X-ray studies still yield only moderate or fair agreement
- Melanoma: pathologists reached only moderate agreement; one study found 64% accuracy, another missed 36%
- Mammography: false-negative rates ranged from 0% to over 50%, false positives from under 1% to 64%
- Skill, Not Only Noise
- Radiology's Achilles' heel: radiologists themselves name diagnostic variation their central vulnerability
- Documentation bias: radiology and pathology may look noisier only because scans and slides can be re-reviewed
- Skill explains much: 44% of diagnostic variation traced to skill, favoring training over uniform decision rules
- Noise and bias differ: training and selection address error, but noise reduction is a separate problem
- Occasion noise: radiologists contradict their own earlier readings, on angiograms 63% to 92% of the time
- Fatigue, Time, and the Doctor's Clock
- Afternoon effect: cancer screening orders ran 63.7% at 8 a.m., falling to 47.8% at 5 p.m.
- Same patient, different order: later appointment times made guideline-recommended screening less likely
- The mechanism: clinics run behind, and complex patients consume the standard twenty-minute slot
- Stress and fatigue: known occasion-noise triggers, now visible in something as routine as ordering a test
- Guidelines as Decision Hygiene
- Diagnostic guidelines: one structured strategy the profession uses to shrink unwanted variability
- Treatment is noisy too: the Dartmouth Atlas documented glaring geographic variation in care, beyond diagnosis
- A gold mine: medicine's prescriptive noise-reduction literature offers ideas transferable to other fields
- Unsolved despite awareness: corrective steps have not eliminated variability in reading angiograms or TB films
- Judgment Versus Calculation in Diagnosis
- Guidelines, Noise, and Medical Judgment (22. Guidelines in Medicine · II)
- Where Noise Lives in Medicine
- Noise in diagnosis: the same patient, same criteria, different clinicians—different conclusions
- Fatigue effects: doctors late in a shift skip preventive discussions and wash hands less often
- Mechanical diagnoses: some conditions allow no room for judgment at all
- Straightforward diagnoses: any clinician with training reliably reaches the same conclusion
- Specialized cases: lung cancer specialists agree closely; noise exists but is minimal
- Open-ended criteria: in psychiatry, judgment room is vast and noise is substantial
- Three Routes to Less Noise
- Training: builds skill, and skill genuinely reduces variability
- Aggregation: second opinions and pooled expert judgments improve accuracy
- Deep learning: algorithms detect breast cancer lymph node metastases better than the best pathologist
- Other successes: AI matches radiologists on mammograms and flags diabetic eye disease
- Trade-off: algorithms promise to cut both bias and noise, saving lives and money
- How Guidelines Reduce Noise: Apgar and Beyond
- Apgar score: Virginia Apgar's 1952 guideline replaced clinical judgment about newborn distress
- Five cues: appearance, pulse, grimace, activity, respiration, each scored 0, 1, or 2
- Judgment decomposed: only heart rate is numeric; simple components keep disagreement low
- Three mechanisms: focus on relevant predictors, simplify cue judgments, aggregate weights mechanically
- Centor score: four signs score strep throat risk, cutting unnecessary testing and treatment
- BI-RADS: raised interrater agreement on mammograms; pathology saw similar guideline gains
- Psychiatry: The Extreme Case
- Persistent disagreement: psychiatrists using identical criteria frequently disagree about the same patient
- Early studies: agreement between two psychiatrists ran only 50%, 54%, and 57%
- Pattern noise: some psychiatrists lean toward diagnosing depression, others toward anxiety
- Physician inconstancy: differing schools, training, experience, and interview styles drive divergence
- Nomenclature inadequacy: large, vague diagnostic categories force interpretation of open-ended criteria
- Why DSM Revisions Fell Short
- DSM-III 1980: introduced explicit, detailed criteria; it spurred research and reduced some noise
- Still noisy: even after the 2000 revision of DSM-IV, noise levels remained high
- DSM-5 2013: objective scaled criteria were expected to help; psychiatrists still disagree
- Depression agreement: field trials found highly trained specialists agreed only 4–15% of the time
- Worse in places: mixed anxiety-depressive disorder proved "so unreliable as to appear useless"
- Root cause: criteria stay vague and subjective, lacking objective measures like blood tests
- The Case for More Guidelines
- Proposals: clarify criteria, define symptom references, add structured interviews to open conversation
- Screening tool: one 24-question interview guide improves reliability for anxiety, depression, eating disorders
- Verdict: guidelines have succeeded broadly in medicine at cutting bias and noise—the profession needs more
- Where Noise Lives in Medicine
- Performance Ratings Are Mostly Noise (23. Defining the Scale in Performance Ratings · I)
- The Rating-Scale Exercise
- Simple exercise: rate three acquaintances on kindness, intelligence, and diligence, then compare with another knowledgeable rater.
- Level noise: raters differ in scale use—one treats 5 as extraordinary, another as merely unusually good.
- Definition noise: raters may define the traits themselves differently, altering judgments of the same person.
- High stakes: if promotion or bonus depends on ratings, policy and scaling differences can produce even more noise.
- Performance Reviews as Judgment Tasks
- Ubiquitous reviews: almost all large organizations evaluate performance regularly, though nearly everyone hates the process.
- Bias and noise: reviews are widely known to suffer both, but their extreme noisiness is less recognized.
- Objective ideal: evaluation would avoid judgment if measurable outputs captured performance.
- Knowledge work: CFOs, researchers, and doctors balance multiple, sometimes contradictory objectives.
- Context required: sales and coding metrics must be interpreted against customer difficulty and project differences.
- Judgment remains: many roles cannot be evaluated entirely by objective performance metrics.
- Signal Lost in System Noise
- Weak relationship: true performance variance explains only 20–30% of rating variance; 70–80% is system noise.
- Level noise: some raters are consistently lenient or harsh toward everyone they rate.
- Pattern noise: a rater’s stable, idiosyncratic reaction to a specific person adds unique noise.
- Occasion noise: transient moods or events—a dented car or surprise bonus—shift a rating.
- Strategic motives: raters may inflate to avoid feedback, favor a promotion, or help a transfer.
- Developmental reviews: even 360-degree feedback used only for development remains highly noisy.
- Reforms and 360-Degree Feedback
- Aggregation: most organizations combine multiple ratings to reduce system noise.
- Broader view: 360-degree feedback was meant to capture peers, subordinates, and customers, not just a boss.
- Project-based fit: its rise matched fluid organizations; evidence suggests it predicts some objective performance.
- Overengineered surveys: computerization and proliferating objectives created absurdly complex questionnaires.
- Halo effect: supposedly separate dimensions bleed together, with early answers pulling later ones.
- Time burden: middle managers complete dozens of questionnaires, harming the quality of information supplied.
- Inflation and the Case for Rankings
- Creeping inflation: one company rated 98% of managers as fully meeting expectations, undermining ratings’ value.
- Forced ranking: standardization forces a predetermined distribution and curbs inflation.
- G.E. example: Jack Welch championed forced ranking for candor; many firms adopted then abandoned it.
- Morale costs: side effects on teamwork and morale made forced ranking unpopular.
- Relative advantage: rankings are less noisy than absolute ratings.
- Punitive-damages parallel: relative judgments showed less noise, a pattern that applies to performance ratings too.
- The Rating-Scale Exercise
- Forced Ranking, Absolute Scales, Rating Noise (23. Defining the Scale in Performance Ratings · II)
- Absolute Versus Relative Scales
- Panel A: absolute scales demand a matching operation — the score closest to your impression
- Panel B: relative scales rank each person against a group on one specified dimension
- Percentile scale: the supervisor states an employee's rank within a defined population
- Two Advantages of Ranking
- Structuring: rating one dimension at a time splits a complex judgment into manageable parts
- Halo effect: separate dimensions stop one impression from keeping all scores in a narrow band
- Condition: structuring works only if each dimension is ranked separately, not as aggregate quality
- Pattern noise: comparing two team members makes a rater less inconsistent than grading each alone
- Level noise: lenient and tough raters differ in averages but not in rankings
- Forced Ranking and Its Two Flaws
- Forced ranking: a mandated distribution — no more than 20% top, at least 15% bottom
- Stated aim: equalizing every rater's mean and distribution is the main objective
- Flaw one: forcing a relative scale onto absolute performance is illogical
- Absolute expectations: nearly everyone can meet expectations if defined ex ante and absolutely
- Absurd mandate: requiring a set percentage to fail is cruel and absurd, not merely harsh
- Legitimate case: relative ratings fit fixed-slot contests, like colonels competing for general
- The Distribution Assumption Fails
- Distribution assumption: forced curves presume ratings mirror the true spread of performance
- Small samples: ten random people need not reproduce the population — only a 30% chance
- Non-random teams: some units are stacked with stars, others with subpar performers
- Undifferentiated reality: forcing distinctions among five equals increases rather than reduces error
- Fatal flaw: the defect is the "forced," not the "ranking" — a wrong scale adds noise mechanically
- The Disappointing Record of Evaluation
- Escalating cost: Deloitte spent 2 million hours a year evaluating 65,000 people
- Dreaded ritual: givers and receivers alike hate reviews; ~90% say the process fails
- Demotivating: development-linked feedback helps, but ratings as practiced demotivate as often
- Radical option: some firms, largely tech, are dropping evaluations or going numberless
- Reducing Noise: Anchors, Training, Resistance
- Bottom line: noise pervades performance ratings, and no simple technological fix exists
- Common frame: better rating formats plus rater training raise consistency in scale use
- Behaviorally anchored scales: each point describes specific behaviors, yet noise persists
- Frame-of-reference training: raters practice vignettes, then compare with experts' "true" ratings
- Case scales: new ratings compare to anchor cases, and comparisons are less noisy than numbers
- Why rare: complex, customized, time-consuming — and raters resist losing control over outcomes
- Absolute Versus Relative Scales
- Why Unstructured Interviews Fail (24. Structure in Hiring · I)
- The Interview as a Judgment Task
- Universal ritual: nearly everyone is hired through some form of interview, rarely questioned as a method
- Deep-seated belief: organizations trust intuitive judgment over formal evidence when choosing people
- A perfect test case: no complex judgment task has drawn more field research than personnel selection
- Stakes stated early: the first Journal of Applied Psychology (1917) called hiring the "supreme problem"
- Unstructured Interviews Predict Poorly
- Unstructured vs structured: the familiar free-form interview is distinguished from its structured counterpart
- Weak correlation: interview ratings correlate only .20–.33 with later job performance
- Coin-flip odds: knowing who interviewed better gives only a 56–61% chance of predicting who performs better
- Measurement caveat: success is judged by supervisor ratings or tenure—questionable, but the only available standard
- Secondary purposes: interviews also sell the company and build rapport, yet selection remains their main job
- Two Sources of Error
- Objective ignorance: performance depends on unforeseeable events, limiting any selection technique's predictive validity
- Shared biases: similarity preferences (gender, race, education) and physical appearance skew all interviewers alike
- Bias versus noise: shared biases create common error; interviewers disagreeing with each other is a separate problem
- Bias training: many companies now train recruiters, addressing shared error but not variability
- Noise in Interviewing
- Inter-interviewer disagreement: correlations between two interviewers' ratings of the same candidate run only .37–.44
- Panel interviews help, incompletely: correlation rises to .74—still disagreeing about one candidate in four
- Pattern noise: each interviewer reacts idiosyncratically to the same interviewee behavior
- Occasion noise: the first two or three minutes of rapport-building shape the final recommendation
- Superficial signals: extraversion, verbal skill, even handshake quality predict hiring recommendations
- The Psychology of Interviewers
- Self-confirming questions: interviewers steer toward evidence that fits their initial impression rather than testing it
- Positive-impression shortcut: interviewers who like a candidate ask fewer questions and shift into selling the company
- Coherence hunger: in an experiment, no interviewer noticed candidates answering yes-or-no questions at random
- Interpretation follows attitude: the same departure after "strategic disagreement with the CEO" read as integrity or as immaturity
- Facts aren't neutral: how evidence is read depends on the evaluator's prior attitude toward the candidate
- The Interview as a Judgment Task
- Structuring Hiring to Cut Noise (24. Structure in Hiring · II)
- How Traditional Interviews Mislead
- Vivid impressions: interviewers feel confident about conclusions drawn from a chat that predicts almost nothing
- Overweighting the interview: impressions crowd out more predictive data such as test scores
- Case in point: a teaching demo's bad performance outweighed a résumé of teaching awards
- The verdict: limitations of informal interviews cast serious doubt on drawing meaningful conclusions
- Aggregating Inputs: Judgment or Formula
- Two routes: inputs can be combined clinically (by judgment) or mechanically (by formula)
- Established superiority: mechanical aggregation wins generally and for work performance prediction
- Default failure: most HR professionals favor clinical aggregation, adding yet more noise
- Precondition: aggregation only works if the judgments being combined are independent
- Google's Redesign
- The audit: recruiting interviews showed "zero relationship… a complete random mess"
- Fewer interviews: twenty-five cut to four, since extra interviews added almost no validity
- Independence enforced: interviewers rate candidates separately before any discussion
- Four assessments: general cognitive ability, leadership, googleyness, role-related knowledge
- Decomposition: Defining What You Seek
- Decomposition: break the decision into mediating assessments, focusing judges on important cues
- Filtering function: a road map of needed data that excludes irrelevant information
- Bloated descriptions: vague consensus wish lists offer no way to calibrate or trade off traits
- Invest in problem definition: agree on a specific job description before meeting any candidate
- Independence and Structured Interviews
- Interdependence problem: when assessments influence one another, each becomes very noisy
- Structured behavioral interviews: predefined questions, recorded answers, unified rubric, BARS
- Predictive gain: structured interviews correlate .44–.57 with performance; PC rises to 65–69%
- Other inputs: work sample tests are top predictors; backdoor references add unsolicited evidence
- Unpopular truth: both interviewers and candidates prefer unstructured formats, yet structure wins
- Delaying Holistic Judgment
- Prescription: do not exclude intuition—delay it until evidence is collected and analyzed
- Hiring committee: final decision stays a judgment, anchored on the average of four interview scores
- Precedent: Kahneman's 1956 Israeli army process formalized dimensions, scored each in turn, then judged
- Persistence of illusion: recruiters and candidates alike underestimate the noise in hiring judgments
- How Traditional Interviews Mislead
- Structuring Judgment to Reduce Noise (25. The Mediating Assessments Protocol · I)
- A Protocol for Noise Reduction
- Purpose: a decision method designed with noise mitigation as its primary objective
- Foundation: it bundles the decision hygiene strategies introduced in earlier chapters
- Scope: applicable whenever a plan or option must be weighed on multiple dimensions
- Analogy: evaluating a strategic option is treated as evaluating a job candidate
- The First Meeting: Structuring the Debate
- Structured interviews: structure produces far more accurate evaluations than free-form discussion
- Core analogy: options are like candidates, so hiring discipline transfers to strategic choices
- Agenda first: agree in advance on the assessments, then discuss them one by one
- Postponed closure: keeping the verdict open stops debate from confirming first impressions
- Intermediate goals: reaching a conclusion on one dimension should not color an unrelated one
- Objection answered: the protocol reshapes the agenda rather than delaying the decision
- Defining the Mediating Assessments
- Comprehensive: every relevant fact should influence at least one assessment
- Independent: ideally each fact influences only one, minimizing redundancy
- The hard part: keeping the list short, comprehensive, and nonoverlapping
- Familiar surface: the list resembles a normal report's contents, but each item is a separate chapter and discussion
- The rating question: how strongly does the evidence on this dimension alone argue for or against the deal?
- The Deal Team: Objectivity and the Outside View
- Mission: provide an objective, independent rating of each assessment, not of the deal as a whole
- Outside view: anchor estimates in base rates drawn from comparable cases
- Reference class: define which deals count as comparable before estimating approval odds
- Comparative judgment: rank against peers ("second quintile") rather than calling performance "good"
- Relative over absolute: comparative evaluations are more accurate than absolute ones
- Independence and Honest Reporting
- Halo effect: a general impression distorts ratings on specific dimensions
- Witness rule: witnesses do not compare notes before testifying, and neither should analysts
- Structural separation: assign different analysts to different assessments and forbid cross-talk
- Sequencing: when one analyst holds two assessments, make them dissimilar and complete one first
- External input: an outside HR expert rates management quality to break the link with recent results
- No concealment: include contradicting evidence rather than smoothing it into the rating
- A Protocol for Noise Reduction
- Independent Assessments, Delayed Intuition (25. The Mediating Assessments Protocol · II)
- Keeping Assessments Independent
- Truth over advocacy: Analysts are told to represent the truth, not to sell a recommendation.
- Honest confidence: Tell the board when you are really in the dark; they know you lack perfect information.
- Instant escalation: Any potential deal breaker must be reported immediately, not buried in the report.
- Desirable incoherence: Independent assessments produce uneven ratings — reality is less coherent than board presentations make it seem.
- Reviewing Each Assessment Separately
- Distinct agenda items: Each assessment is considered on its own before any holistic view is formed.
- Artificial but valuable: Insiders had to withhold their overall views; some quietly changed their minds during the meeting.
- Discrepancies as fuel: Uneven ratings raise questions and trigger the debate that makes the decision better.
- Estimate–Talk–Estimate
- Facts first: The deal team briefly summarizes key facts the board has already read in detail.
- Anonymous temperature check: A phone vote projects the distribution of independent ratings without identifying raters.
- Blunting cascades: Getting independent opinions before discussion reduces social influence and information cascades.
- Focus on divides: Joan spends most time where views oppose, urging facts, arguments, nuance, and humility.
- Reasonableness frame: "We are all reasonable people and we disagree" — so this is a legitimate disagreement.
- Second estimate: Re-voting after discussion usually yields more convergence than the first round.
- Delaying Intuition, Not Banning It
- Profile of the deal: Averaged ratings on the whiteboard frame the final deliberation.
- Why formulas are rejected: People resist schemes that tie their hands, and game formulas to reach desired conclusions.
- Hidden broken legs: Unanticipated deal breakers or clinchers could be missed by mechanical averaging.
- Anchored intuition: With fact-based assessments known to all, the final judgment is safely anchored rather than free-floating.
- The Protocol in Recurring Decisions
- Define once: The list of mediating assessments is set once and applied to every investment.
- Experience as reference class: Hundreds of prior evaluations give a shared basis for comparative judgment.
- Case scale: Anchor cases — "as good as ABC, not quite DEF" — make relative judgments, which beat absolute ratings.
- Upkeep required: Anchor cases must be known to all participants and periodically updated.
- What the Protocol Changes
- Six steps: Structure assessments, use the outside view, keep them independent, review separately, estimate–talk–estimate, delay intuition without banning it.
- Content versus process: Content is specific and fun; process is generic and unglamorous — hence the resistance.
- Not bureaucracy: Decision hygiene promotes challenge and debate, unlike stifling consensus.
- Invisible noise: Leaders are typically unaware of noise in their biggest decisions, so they take no measures against it.
- Handwashing analogy: Decision hygiene will not prevent all mistakes, but it addresses an invisible, pervasive, damaging problem.
- Keeping Assessments Independent
- What Makes a Judge Good (18. Better Judges for Better Judgments · I)
- Part VI: Optimal Noise
- The Case Against Cutting Noise (26. The Costs of Noise Reduction · I)
- Objections to Noise Reduction
- Noise is unwanted variability: if unwanted, why not eliminate it — unless eliminating it costs more than it saves.
- Seven objections: expense, new errors, lost dignity, frozen values, gaming, deterrence, and demoralization.
- Not a blanket rejection: each objection applies to specific strategies, not to noise reduction as a goal.
- Selective agreement: one may reject rigid guidelines yet welcome aggregating independent judgments.
- The Cost Objection
- Too expensive: noise reduction demands time, effort, and money that may exceed its benefits.
- Teacher's essays: double-grading, checklists, and re-reading reduce noise but consume scarce hours.
- Stakes decide: elaborate hygiene suits senior theses affecting admissions, not low-stakes ninth-grade work.
- Feasibility first: a remedy must be practical before it can be judged worthwhile.
- Costly test: an invasive, expensive diagnostic may not be justified when variability is mild.
- Legitimate but overstated: the expense argument is often an excuse, sometimes shortsightedly self-serving.
- Weighing Benefits Against Costs
- Disciplined analysis: compare accuracy gained, its importance, and the resources required.
- Noise audits: they reveal when noise produces outrageous unfairness, very high costs, or both.
- Then resist the excuse: when audits show real harm, expense is no reason to do nothing.
- Not always wrong: sometimes tolerating noise genuinely is the rational choice.
- Less Noise, More Mistakes
- Blunt instruments: crude noise-reduction tools can generate unacceptably high levels of error.
- False positives: banning all posts containing vulgar words cuts noise but removes posts that should stay.
- Bias, not just noise: uniform rules can convert random error into directional error.
- Cures worse than the disease: many well-motivated reforms curbing discretion backfire badly.
- Hirschman's Three Objections
- Perversity: the reform aggravates the very problem it was meant to solve.
- Futility: the reform changes nothing at all.
- Jeopardy: the reform endangers other important values.
- Rhetoric or substance: perversity and jeopardy are the most powerful — and often deployed to derail reforms that would do real good.
- Objections to Noise Reduction
- The Price of Taming Noise (26. The Costs of Noise Reduction · II)
- The Objection from Discretion
- The judges' counterclaim: cutting discretion may produce more mistakes, not fewer
- Havel's warning: human situations are not a puzzle awaiting a universal solution
- Varied cases: good judgment addresses particulars, which may mean tolerating some noise
- The trade-off: noise is not the only cost; rigidity and error count too
- Crude Rules That Silence Error-Prone Judgment
- Airline chess program: a noiseless rule—always check the king—played consistently and badly
- Consistency is not accuracy: identical output every time can still be wrong every time
- Three strikes: life sentences remove sentencing noise but ignore nonviolence, circumstances, rehabilitation
- Woodson v. North Carolina: mandatory death sentences struck down because they were rules
- Individualized justice: offenders must not be a "faceless, undifferentiated mass"
- Rigidity Outside the Courtroom
- Wide reach: teachers, doctors, employers, underwriters, coaches risk rigid-rule mistakes
- Narrow scorecards: simple evaluation rules may ignore important aspects of performance
- The asymmetry: a noise-free score missing key variables can be worse than noisy judgment
- The Objection Is Weaker Than It Seems
- No false choice: one error-prone strategy does not justify accepting high noise
- Better rules exist: admissions formulas weighing scores, grades, and background beat crude cutoffs
- Professional guidelines: complex diagnostic rules cut noise without intolerable cost
- Decision hygiene: aggregate judgments or use structured protocols when rules fail
- The real remedy: replace bad noise reduction with better noise reduction
- Noiseless but Biased Algorithms
- Dual failure: algorithms can eliminate noise and embed unacceptable discrimination
- COMPAS: ProPublica found the recidivism tool biased against racial minorities
- Overt bias: algorithms can use race, gender, or pregnancy directly
- Proxy predictors: correlated variables like height, weight, or neighborhood smuggle bias in
- Biased training data: predictive policing perpetuates the overpolicing it learns from
- Transparency and Better Algorithms
- Auditability: algorithms can be tested for inadmissible inputs and disparate impact
- Human opacity: unconscious discrimination by judges is far harder to detect
- Combined criteria: accuracy, noise reduction, nondiscrimination, and fairness judged together
- Evidence: bail and résumé algorithms can be both more accurate and less discriminatory
- Central conclusion: costs of noise reduction are often an excuse—design better strategies, don't give up
- The Objection from Discretion
- 27. Dignity
- The Dignity Objection
- Individualized treatment: people want a real human being to weigh their particular circumstances, not a firm rule.
- Due process ideal: face-to-face discretion signals respect, even when it necessarily produces noise.
- Portia's plea: mercy is unbound by rules and therefore noisy; the appeal echoes in countless organizations.
- Not all noise reduction offends: aggregation and guidelines still allow hearings; only rigid rules strip discretion.
- Changing Values
- Irrebuttable presumption: LaFleur condemned a five-month pregnancy leave rule for lacking individualized determination, not for its length.
- Frozen norms: rigid rules can lock in values that ought to evolve, as an outdated and sexist expense policy showed.
- Defense of flexibility: noisy systems let judges and firms shift as moral values change.
- Rebuttal: shared scales and aggregation accommodate change, rules can be revised regularly—and hiring or diagnostic noise rarely reflects evolving morality.
- Gaming the System
- Clear edges: precise rules let the clever evade them through conduct technically exempt but equally harmful.
- Tax code tradeoff: some vagueness discourages opportunistic behavior but purchases that gain with noise.
- Wrongdoing gap: unspecified prohibitions breed noise; specific lists tolerate everything not listed.
- The real question: weigh how much evasion against how much noise—little evasion and much noise favors reducing noise.
- Deterrence and Risk Aversion
- Expected value: a 50% chance of a $5,000 fine equals a certain $2,500 fine, in the abstract.
- Risk attitude decides: risk-averse wrongdoers are more deterred by the lottery, risk-seeking ones less.
- Better route: raise the penalty and eliminate the noise, deterring more while removing unfairness.
- Creativity and Morale
- Discretion as dignity: judges rebelled against sentencing guidelines, feeling diminished, even humiliated, by lost judgment.
- Cogs in a machine: rule-bound work feels mechanical and squelches independent decision-making.
- Demoralization is a cost: disengaged, uninspired employees perform worse.
- Howard's principles: general principles like "be reasonable" free expertise but invite noise in interpretation and enforcement.
- Design choice: structuring complex judgments cuts noise while keeping room for fresh ideas; challenge rules by process, never by case-by-case breach.
- Speaking of Dignity
- Objections summarized: dignity, moral evolution, deterrence, evasion, and creativity are each invoked to defend noise.
- Verdict: dignity, moral evolution, and human creativity can all be honored without tolerating noise's unfairness and cost.
- The Dignity Objection
- 28. Rules or Standards?
- The Central Distinction
- Rules: designed to eliminate discretion; those who apply them answer questions of fact
- Standards: designed to grant discretion; vague terms require judges to give them content
- Noise consequence: rules should sharply cut noise; standards invite it
- Bias and noise: biased rules still cut noise if followed; standards permit both
- Delegation: those who devise standards export decision-making authority to others
- Why Standards Persist
- Divisions: divided groups can agree on standards without agreeing on specifics
- Ignorance: when no one knows the right rule, standards lean on trusted experts
- Compromise: lawmakers may accept noise as the price of enacting any law at all
- Constitutions: broad standards protect rights that diverse peoples can endorse
- Hidden cost: agreement on words buys silence on meaning, producing variability
- Controlling Subordinates
- Principal-agent choice: any leader must choose specificity or latitude for subordinates
- Trust and discretion: the amount of discretion granted tracks trust in the agent
- Facebook's mix: public Community Standards read as standards; 12,000-word Implementation Standards as rules
- Scale forces rules: thousands of reviewers made blunt, mechanical rules necessary
- Judgment's place: pilots and doctors may do better with standards despite the noise
- A Framework for Choosing
- Costs of decisions: standards are burdensome to apply; rules are cheap once in place
- Costs of errors: neither form is universally safer; weigh number and magnitude of mistakes
- Producing rules: writing an accurate rule in advance can be prohibitively costly
- Repeated decisions: high volume makes ad hoc judgment's burdens intolerable
- Smart organizations: reduce noise with rules, then invest ahead to make rules accurate
- The Return of the Repressed
- Bureaucratic justice: Mashaw's term for the disability matrix's noise-eliminating fairness
- Underground discretion: judges, employees, and agencies quietly ignore rigid rules
- Jury nullification: juries refuse senselessly harsh law, hiding discretion from view
- Invisible noise: when discretion goes underground, variability returns unpoliced
- Monitor and revise: noise in outcomes may signal rules aren't working as intended
- Should the Law Outlaw Noise?
- Weber's Kadi justice: ad hoc judgment that "knows no rational rules of decision"
- Kadi persists: long after Weber, informal case-by-case judgment remains pervasive
- Noise as injustice: sentencing, asylum, licenses, child custody can be frightfully noisy
- Bias seen, noise unseen: organizations treat bias as villain; they should treat noise the same
- Legal systems should act: the law should do far more to combat noise and its unfairness
- The Central Distinction
- The Case Against Cutting Noise (26. The Costs of Noise Reduction · I)
- Review and Conclusion: Taking Noise Seriously
- Noise as Unwanted Variability in Judgment (Review and Conclusion: Taking Noise Seriously · I)
- Judgment as a Form of Measurement
- Judgment defined: a measurement whose instrument is a human mind, far narrower than thinking.
- Assigns a score: the score need not be numeric—"probably benign" is a judgment too.
- Informal integration: judgment blends diverse cues into an overall assessment without exact rules.
- Professional judges: coaches, cardiologists, lawyers, underwriters—their quality affects everyone.
- Invisible bull's-eye: judges act as if a true value exists even when none can be known.
- Bounded disagreement: judgment calls sit between computation, where disagreement is banned, and taste, where it is expected.
- Bias and Noise as Distinct Errors
- Bias: when most errors lean the same direction—the average error of a set.
- Noise: the unwanted divergence remaining after bias is removed; variability that should not exist.
- System noise: variability among interchangeable professionals deciding identical cases.
- Independent and additive: as measured by MSE, bias and noise each contribute separately to error.
- Zero scatter is best: reducing noise always improves accuracy, even when bias persists.
- Which dominates: noise outweighs bias whenever bias is smaller than one standard deviation of error.
- Measuring Bias and Noise
- MSE standard: used for two centuries; penalizes large errors and ignores asymmetric costs.
- Forecast accuracy: MSE suits predictive judgments, where objective accuracy is the goal.
- Noise audit: several professionals independently judge the same cases; scatter is visible without truth.
- Bias unknowable: for unverifiable judgments, the average of many judges is a convenient assumption, not a fact.
- Audit yields: quantified system noise, and sometimes revealed gaps in skill or training.
- Why Noise Is a Problem
- Variance welcomed: diverse opinions drive ideas, markets, and innovation—sometimes disagreement is the point.
- But not in judgment: if two doctors diagnose differently, at least one is wrong.
- Errors don't cancel: scattered shots are not a bull's-eye; over- and underpricing both cost the insurer.
- Injustice: similarly situated people treated differently, and systems perceived as inconsistent lose credibility.
- Scale of surprise: system noise and its damage both far exceed common expectations.
- Level Noise, Pattern Noise, Occasion Noise
- Level noise: some judges are generally harsher, others more lenient; ambiguous scales fuel it.
- Pattern noise: judges differ in which cases they treat harshly, producing different rankings.
- Stable component: idiosyncratic responses to cases—often reflecting unacknowledged personal principles.
- Largest source: stable pattern noise is generally the single biggest contributor to system noise.
- Occasion noise: transient shifts—same image, different diagnosis; harsher before lunch, lenient after a win.
- Psychology Behind Judgment Error
- Objective ignorance: some facts are unknowable; most people underestimate this limit.
- Confidence signal: subjective certainty is a self-generated reward for a coherent story, not a measure of accuracy.
- Models win: simple rules and linear models beat human judges mainly because they are noise-free.
- Biases create noise: unshared or unevenly applied biases add scatter, not just systematic error.
- When agreement holds: single-dimension and same-direction cues produce consensus.
- When it breaks: large individual differences emerge when multiple conflicting cues must be weighted.
- Judgment as a Form of Measurement
- Seeing Noise, Reducing Judgment Error (Review and Conclusion: Taking Noise Seriously · II)
- Why Noise Stays Obscure
- Inconsistent cues: when evidence resists a coherent story, judges weight different cues — and pattern noise follows.
- Explanatory charisma: bias supplies a causal story for bad decisions; noise supplies none, so bias monopolizes blame.
- Causal craving: the mind craves causes and lacks statistical intuition, so noise never surfaces in hindsight.
- No feedback: professionals judge alone, never learn whether colleagues would agree, and treat rare disagreement as isolated.
- Organizational embarrassment: institutions suppress evidence of expert divergence because noise is an embarrassment.
- Decision Hygiene: Prevention Against an Unidentified Enemy
- Good judges, not identical judges: skill, intelligence, and active open-mindedness help, but human variety makes some noise inevitable.
- Debiasing's third way: beyond correcting and taming biases, detect them in real time with a designated decision observer.
- Decision hygiene: prevention against an unidentified enemy, like handwashing — it reduces errors without knowing what they are.
- Noise audit first: begin with an audit to win organizational commitment and quantify the separate types of noise.
- Unglamorous but vital: there is no glory in preventing an unidentified harm, but there are results.
- Principles of Accuracy: Rules, Outside View, Structure
- Accuracy, not expression: judgment is no place for individuality, which belongs to goals, options, and creativity.
- Algorithms eliminate noise: the only approach that removes it completely, though unlikely to replace the final human decision.
- Guidelines constrain discretion: decision guidelines promote homogeneity across diagnoses and rulings, cutting noise.
- Outside view: treat the case as one of a reference class; judges sharing a reference class agree more.
- Regressive prediction: anchor forecasts in similar cases, moderate them, and hold predictive humility.
- Principles of Discipline: Delay, Structure, Aggregate, Compare
- Excessive coherence: impressions of one aspect contaminate others; unfitting information gets distorted or ignored.
- Structure judgments: break complex decisions into independent tasks — structured interviews, Apgar scores, mediating assessments protocol.
- Protect independence: assign assessments to separate teams and minimize communication among them.
- Delay intuition: intuition is acceptable if informed, disciplined, and delayed; sequence information so bias cannot enter early.
- Aggregate independent judgments: collect opinions before discussion; averaging reduces system noise but not bias.
- Relative judgments: pairwise comparisons beat absolute scales; case scales anchor on instances everyone knows.
- Noise in Singular Decisions
- Recurrent in disguise: one-off decisions still carry noise; treat them as recurrent judgments made only once.
- Shooters analogy: a team's scatter is invisible when you watch only the first shooter — so it is with lone decisions.
- Hygiene applies here too: the same principles improve major singular decisions, not just repeated ones.
- The Costs and Limits of Noise Reduction
- Optimal noise may exceed zero: when reduction costs outweigh benefits, some noise should be tolerated.
- Legitimate downsides: algorithms err stupidly, facelessly, or on bad data; hygiene can bureaucratize and demoralize.
- Judge each strategy separately: an objection to aggregating judgments may not apply to guidelines.
- The unmeasured excuse: without a noise audit, citing the difficulty of reduction is an excuse not to measure.
- Take noise seriously: random error damages no less than bias; better decisions require confronting it.
- Why Noise Stays Obscure
- Noise as Unwanted Variability in Judgment (Review and Conclusion: Taking Noise Seriously · I)
- Epilogue: A Less Noisy World
- Organizations Redesigned for Noise
- Noise as a central flaw: hospitals, courts, agencies, and firms would treat unwanted variability as a first-order problem
- Noise audits: routine, perhaps annual, across organizations that make consequential judgments
- Algorithms: used to replace or supplement human judgment far more widely than today
- Decomposition: complex judgments broken down into simpler, separately assessed components
- Decision hygiene: a known discipline, with its prescriptions followed as standard practice
- Independent judgments: elicited separately and then aggregated rather than merged in discussion
- Structured Decision Making
- Meetings: reshaped around structure, with discussion ordered rather than free-form
- Outside view: systematically integrated into the decision process instead of invoked occasionally
- Constructive conflict: overt disagreements become both more frequent and more productively resolved
- Process design: the architecture of judgment, not the individual judge, drives better outcomes
- The Payoff and the Invitation
- Money: a less noisy world would save a great deal of it
- Public safety and health: both improved by more consistent decisions
- Fairness: increased, because similar cases receive similar treatment
- Avoidable errors: many prevented through better judgment processes
- The book's aim: to draw attention to this opportunity and invite readers to seize it
- Organizations Redesigned for Noise
- Appendix A: How to Conduct a Noise Audit – Appendix C: Correcting Predictions
- Appendix A: How to Conduct a Noise Audit
- Purpose: measures noise, and also exposes bias, blind spots, and training gaps
- Cast of characters: project team, executive clients, interchangeable judges, subject experts, senior project manager
- Task selection: written cases with numerical judgments; simplified cases whose noise convinces insiders real work is noisier
- Commitment before results: executives approve the design and document their expectations in advance
- Quiet administration: frame as a decision-making study, guarantee anonymity, test all judges at once
- Corrective follow-through: quantify level and pattern noise, apply decision hygiene, report to leadership
- Appendix B: A Checklist for a Decision Observer
- Design your own: the generic checklist is inspiration, not a form to be used as it stands
- Approach to judgment: substitution, inside view, and correlated or missing perspectives
- Prejudgments: conflicts of interest, prior commitment, suppressed dissent, escalating commitment
- Premature closure: were alternatives fully explored and uncomfortable evidence actively sought?
- Information processing: availability, anecdote over data, anchoring, nonregressive extrapolation
- The decision: planning fallacy, confidence intervals, loss aversion, present bias
- Appendix C: Correcting Predictions
- Matching predictions: forecasting as if our information were perfectly predictive of the outcome
- Regression to the mean: extremes become less extreme because past and future performance correlate imperfectly
- Steps 1–2: write the intuitive guess, then find the class mean—the outside view of the case
- Steps 3–4: estimate diagnostic value as a correlation (above .50 is rare, .20 common), then adjust from the mean in that proportion
- The result: corrected predictions are never as extreme as intuition, always closer to the mean
- Why it wins: outliers are rare by definition, and predicting they persist is the more frequent error
- Appendix A: How to Conduct a Noise Audit
- Acknowledgments – Discover More
- Core Team and Production
- Linnea Gandhi: chief of staff who gave substantive guidance, kept the authors organized, and lifted morale
- Dan Lovallo: coauthored one of the articles that seeded this book
- John Brockman: agent who was enthusiastic, hopeful, sharp, and wise at every stage
- Tracy Behar: principal editor and guide; made the book better in large and small ways
- Arabella Pike and Ian Straus: added superb editorial suggestions
- Expert Readers and Advisers
- Draft readers: dozens of scholars commented on chapters, some on the entire manuscript
- Julian Parris: offered invaluable help on many statistical issues
- Machine-learning chapters: enabled by Mullainathan, Kleinberg, Ludwig, Stoddard, and Chang
- Consistency of judgment: indebted to Todorov's Princeton team and to Highhouse and Broadfoot
- Domain experts: Bock, Cowgill, Dana, Krueger, Mauboussin and others shared their expertise
- Responsibility: any misunderstandings or errors remain the authors' own
- Research Support and Collaboration
- Research army: Bhardwaj, Fisher, Goyal, Grabel, Heinrich, Johnson, Mehta and others contributed over the years
- Payoff: their excellent work left the book with less bias and less noise
- Bespoke analyses: research teams ran special analyses specifically for the authors
- Remote tools: Dropbox and Zoom sustained a three-author, two-continent team through 2020
- Discover More
- Reader promotion: sneak peeks, book recommendations, and news about favorite authors
- Call to action: a link inviting readers to learn more
- Core Team and Production
- Introduction: Two Kinds of Error
- Core Conclusion and Practical Takeaways
- The Central Diagnosis
- Noise as invisible flaw: unwanted variability in judgment rivals bias as a source of error
- MSE equation: total error = bias² + noise² — each contributes independently and equally
- Noise usually dominates: it exceeds bias whenever bias falls below one standard deviation
- Errors don't cancel: judgments scattered in both directions cost as much as a consistent tilt
- Injustice: similar people treated differently by interchangeable judges erodes credibility
- The blind spot: bias monopolizes blame; noise is the unacknowledged bit player
- Measure Before You Fix
- Noise audit: many professionals judge identical cases; their scatter reveals system noise
- Truth not needed: noise can be measured without ever knowing the right answer
- Level noise: some judges are consistently harsher or more lenient than others
- Pattern noise: judges disagree about which specific cases deserve harshness — the largest source
- Occasion noise: the same judge varies with mood, fatigue, time of day, and sequence
- Decomposition: split system noise into level, pattern, and occasion components
- Decision Hygiene in Practice
- Structure judgments: break complex decisions into independent, separately rated dimensions
- Independence first: elicit judgments separately before any discussion to prevent cascades
- Aggregate mechanically: averaging independent judgments divides noise by the square root of their number
- Delay intuition: collect and analyze evidence first; let holistic judgment come last
- Simple rules win: linear models beat experts mainly by being noise-free, not smarter
- Relative over absolute: pairwise comparisons and case scales are less noisy than absolute ratings
- Mindset Shifts
- Confidence ≠ accuracy: the feeling of coherence is a self-generated reward, not a measure of truth
- Objective ignorance: many outcomes are simply unpredictable; overconfidence is well documented
- Judgment is measurement: like a stopwatch, the human mind carries irreducible error
- Singular decisions have noise: treat a one-time decision as a recurrent one made only once
- Take the outside view: anchor on base rates and reference classes, then adjust toward the evidence
- Better judges: seek high ability plus actively open-minded thinking, not bone-deep certainty
- The Limits and the Call
- Optimal noise isn't zero: eliminating it can be infeasible, costly, or demoralizing
- Weigh each strategy: an objection to one remedy does not apply to others
- The excuse trap: without an audit, citing difficulty is an excuse not to measure
- Algorithms err too: they can embed bias, so design better strategies rather than give up
- Take noise seriously: confront random error as vigorously as we confront bias
- The Central Diagnosis
opening map…