- General Overview
- The Optimistic Skeptic
- Central thesis: foresight is real but modest—practiced, measured, and bounded by genuine complexity.
- Optimistic skeptic: accept hard limits of chaos while recognizing much of daily life is predictable.
- Clocks and clouds: how predictable something is depends on what, how far ahead, and under what conditions.
- Measurement gap: most forecasting is never scored, so it can never improve.
- Measured doubt: the cure is finely calibrated uncertainty, reduced by evidence but never zero.
- Illusions of Knowledge
- God complex: experts trust intuition alone, as medicine did for centuries.
- System 1: fast and automatic, jumps to conclusions without flagging its own errors.
- WYSIATI: intuition treats whatever evidence is available as complete and reliable.
- Bait and switch: hard questions get silently replaced by easier, unasked proxies.
- Confirmation bias: we gather supporting evidence and become motivated skeptics against the rest.
- Falsification: "What would convince me I am wrong?" tests attachment to a belief.
- Keeping Score
- Ambiguity kills scorekeeping: forecasts need precise odds, unambiguous terms, and deadlines.
- Wrong-side-of-maybe: judging a 70% forecast only by whether it happened is incoherent.
- Calibration and resolution: good scoring rewards matching frequencies and bold correct calls.
- Brier score: distance between forecast and outcome; lower is better, 0.5 is guessing.
- Foxes beat hedgehogs: eclectic, self-updating thinkers outperformed Big Idea ideologues.
- Who the Superforecasters Are
- Ordinary people: uncredentialed volunteers beat professional intelligence analysts.
- Threshold intelligence: above-average ability helps, but genius territory is not required.
- Talent is not the secret: their edge is what they do—research, self-criticism, synthesis.
- Learnable skill: a one-hour tutorial improved accuracy about 10% across a whole year.
- Defying regression: superforecasters widened their lead in years two and three, unlike lucky streaks.
- How They Think
- Fermi-ize: break unanswerable questions into small, inspectable subquestions.
- Outside view first: anchor on base rates before exploring case specifics.
- Inside view as investigation: then test each pathway with evidence for and against.
- Synthesis: fuse both views like binocular vision; neither alone suffices.
- Granularity: single-percentage-point precision predicts accuracy; rounding even to 5% hurts.
- Active open-mindedness: beliefs are hypotheses to test, and hard thinking is enjoyable.
- Perpetual Beta
- Living judgments: forecasts are revised as information arrives, not buy-and-forget bets.
- Underreaction: giving news too little weight stalls good forecasts; identity blocks revision.
- Overreaction: treating stale or noisy evidence as meaningful destroys accuracy just as badly.
- Small increments: Minto updated daily, averaging 3.5% moves—Bayesian without equations.
- Postmortems: study misses and wins alike; luck can masquerade as skill.
- Grit and growth mindset: perpetual beta predicts improvement about three times better than intelligence.
- Teams and Leadership
- Superteams: grouping elite forecasters gained them 50% accuracy in later years.
- Psychological safety: explicit invitations to disagree beat polite silence.
- Constructive confrontation: precision questioning lowers heat and exposes weak evidence.
- Mission command: tell people what to do, not how, and they improvise brilliantly.
- Superleader: decide deliberately, then act with resoluteness—Moltke resolves the leader's dilemma.
- Intellectual humility: confidence before the task's complexity, not self-doubt.
- Limits and What's Next
- Kahneman's caution: even the best are one System 2 slip from a blown forecast.
- Black swans: history jumps as well as crawls; radical indeterminacy is real.
- Fat-tailed risk: extreme events far exceed bell-curve odds, warranting humility.
- Human-scale foresight: puny but not negligible—modest probabilities still guide decisions.
- Big questions: decompose them into many small, scorable, relevant forecasts.
- The stakes: better forecasting separates prosperity from bankruptcy and sound policy from waste.
- The Optimistic Skeptic
- Deep Dive
- 1. An Optimistic Skeptic
- Prediction's Limits and Possibilities (1. An Optimistic Skeptic · I)
- Untested Pundits, Verified Superforecasters
- Bill Flack: Nebraska retiree with ~300 recorded, independently scored geopolitical forecasts and an excellent record.
- Superforecaster label: reliable track records, not credentials, earned by the top ~2% of volunteer forecasters.
- Tom Friedman: celebrity pundit whose influential forecasts have never been rigorously tested for accuracy.
- Media norms: pundits rise on storytelling conviction, not on reconciling past forecasts with events.
- Hidden risk: leaders and investors make critical decisions on forecasts of unknown quality.
- Cultivable skill: forecasting is not fixed talent; diligent practice can improve it.
- The Dart-Throwing Chimpanzee Study
- Expert-forecast study: two-decade project (1984–2004) assessed thousands of political and economic predictions.
- Average expert: barely beat random guessing on many questions, fueling the dart-throwing chimpanzee joke.
- Horizon effect: accuracy beat chance at one year, then faded to chimp level three to five years out.
- Media distortion: retellings mutated findings into "all expert forecasts are useless"—nonsense.
- Balanced stance: debunkers rightly attack untested pundits, but wrongly dismiss all forecasting as futile.
- The Skeptic's Case for Limits
- Bouazizi trigger: one Tunisian vendor's self-immolation set off cascading protests, toppling dictators.
- Butterfly effect: tiny changes in initial conditions can mushroom in complex nonlinear systems (Lorenz).
- Laplace's demon: classical ideal that total knowledge yields total predictability—undermined by chaos theory.
- Lorenzian cloud: exact shapes of clouds remain unpredictable despite complete knowledge of cloud formation.
- Hard limits: no one can foresee distant futures in a world where small events ripple globally.
- Narrative hindsight: connecting Bouazizi to the Arab Spring is storytelling, not foresight.
- The Optimist's Everyday Predictability
- Mundane forecasts: traffic, restaurant hours, red lights, and meetings are anticipated reliably every day.
- Electricity grids: producers anticipate routine demand surges and adjust output accordingly.
- Actuarial science: insurers profitably forecast disability and death across large populations, not individuals.
- Algorithms: Amazon and Google personalize predictions from past behavior and millions of other users.
- Astronomical precision: sunrise and sunset times are known to the minute decades ahead.
- Optimistic skeptic: accept hard limits yet recognize much of daily life is eminently predictable.
- Untested Pundits, Verified Superforecasters
- Measuring Foresight in a Cloudy World (1. An Optimistic Skeptic · II)
- Clocks and Clouds
- False dichotomy: reality is neither fully clocklike nor cloudlike; predictability and unpredictability coexist.
- Context matters: how predictable something is depends on what, how far ahead, and under what circumstances.
- Time and chaos: the farther out the forecast, the more opportunity for nonlinearity and butterfly effects.
- No certainties: even sunrise forecasts assume no asteroid or black-hole catastrophe intervenes.
- The Weather Model
- Forecast, measure, revise: meteorologists routinely score forecasts and refine models, yielding slow improvement.
- Nonlinear limits: beyond a week, weather forecasts approach dart-throwing chimp; computing gains yield shrinking payoffs.
- Knowing limits: recognizing current predictability boundaries is itself a forecasting success.
- The Measurement Gap
- Most forecasting is unmeasured: outside tech-driven fields, accuracy is seldom checked with rigor.
- Demand-side failure: governments, business, and public rarely demand evidence of forecast accuracy.
- No revision, no progress: without measurement, forecasting cannot improve; clear goals and measures drive progress.
- Mixed motives: forecasts can entertain, persuade, impress, or comfort—not just aim for accuracy.
- The Good Judgment Project
- GJP origin: launched with Barbara Mellers in 2011; thousands of lay volunteers forecast global affairs.
- IARPA tournament: five teams forecast the same questions over four years, creating a level playing field.
- Stunning results: GJP beat control by 60–78% and outperformed competitors and professional intelligence analysts.
- Massive data: nearly 500 questions and over one million individual judgments about the future.
- Superforecasters
- Foresight is real: measurable skill exists for judging high-stakes events months to a year out.
- Not who they are: intelligence, numeracy, and news consumption matter only up to a threshold.
- What they do: open-minded, careful, curious, self-critical thinking—plus commitment to self-improvement.
- Learnable habits: a one-hour tutorial improved accuracy about 10% for the whole tournament year.
- Small edges compound: sustained modest gains separate consistent winners from those going broke.
- Machines and Judgment
- Algorithms beat experts: since Meehl, 200+ studies show statistical methods usually outperform subjective judgment.
- Rare validated algorithms: the finding rarely applies because real-world forecasting algorithms are seldom available.
- Machine limits: Watson can retrieve historical facts but not judge Putin’s intentions or Russian political dynamics.
- Human-computer synthesis: freestyle-chess-style teams may beat both pure humans and pure machines.
- Guru model obsolete: experts paired with computers can overcome human cognitive limitations and biases.
- Clocks and Clouds
- Prediction's Limits and Possibilities (1. An Optimistic Skeptic · I)
- 2. Illusions of Knowledge
- Knowledge, Intuition, and Certainty (2. Illusions of Knowledge · I)
- The Doctor Who Was Certain
- Archie Cochrane: renowned physician told he had terminal cancer in 1956, then learned he never had it.
- Specialist's error: no one doubted the diagnosis or waited for the pathologist's report before surgery.
- Cochrane's blind spot: despite championing evidence, he accepted one man's judgment and prepared to die.
- Double riddle: why trust intuition over evidence, and why did a skeptic rush to judgment?
- Medicine's Long Stagnation
- Historical treatments: from ancient Egypt to Washington's doctors, most were useless or worse; the sick often fared better alone.
- Galen: authority for over a millennium, never ran experiments, and read every outcome as confirmation.
- Charismatic quacks: Thomsonianism, orificial surgery, and orthodox physicians all "blind men arguing over the colors of the rainbow."
- No doubt, no progress: absence of skepticism kept medicine unscientific for centuries.
- The Road to Testing
- James Lind: 1747 scurvy trial hit on citrus cure, yet he failed to grasp its significance.
- Austin Bradford Hill: randomized trials let differences among patients balance out, isolating treatment effects.
- Randomized controlled trials: seemed obvious after WWII but were slow and resisted by the establishment.
- Cargo Cult Science
- Cargo cult science: Richard Feynman's term for imitating science's form without its substance.
- Missing ingredient: medicine had the outward trappings of science but lacked doubt.
- Doubt's value: uncertainty prompts new directions and new things to test, driving progress.
- The God Complex
- God complex: Cochrane's label for physicians' belief that their judgment alone revealed truth.
- Cardiac care unit trial: Cochrane proposed randomizing patients; cardiologists called it unethical.
- His trick: told colleagues home care did slightly better, then revealed he had reversed the results.
- Uncontrolled policy: "short, sharp, shock" made its effectiveness unknowable.
- Thinking About Thinking
- System 1: fast, automatic, and always running; the source of instant answers.
- System 2: deliberate scrutiny of System 1's answers; often never gets involved.
- Introspection limits: conscious thought captures only a tiny fraction of the processes behind decisions.
- Bat and ball problem: instant "ten cents" feels right but fails sober reflection — the ball costs 5 cents.
- The Doctor Who Was Certain
- Intuition's Illusions and Doubt (2. Illusions of Knowledge · II)
- System 1's Shortcuts
- Cognitive Reflection Test: most people, even smart ones, blurt "ten cents" without reflecting.
- System 1: designed to jump to conclusions from little evidence; if it feels true, it is.
- WYSIATI (What You See Is All There Is): intuition treats available evidence as reliable and sufficient without quality check.
- Availability heuristic: easily recalled shadow-lion examples shape judgments, not evidence quality.
- Tip-of-your-nose perspective: automatic, subjective view is compelling but often wrong.
- Coherent Stories and Confirmation Bias
- Explanatory urge: the brain demands order and confabulates stories that impose coherence.
- Split-brain patient: with no idea why he pointed at a shovel, he invented "chicken shed" logic.
- Market journalism: "Dow rose on news…" invents a cause even when the news came after the rise.
- Oslo 2011: Islamist narrative seemed to fit perfectly; the killer was a lone anti-Muslim Norwegian.
- Confirmation bias: we gather supporting evidence and become motivated skeptics against the rest.
- Bait and Switch
- Attribute substitution (bait and switch): hard questions are replaced by easy proxies without noticing.
- Availability heuristic: "Can I recall a lion attack?" replaces real risk—a classic bait-and-switch.
- Conscious substitution: laypeople may deliberately use climatologists' consensus as a proxy for their own ignorance.
- No immunity: even skeptical physician Cochrane skipped the pathologist's wait when the story felt true.
- Blinking and Thinking
- Blink vs think: false dichotomy; blend intuition with analysis as situations evolve.
- Fire commander: inexplicable unease before floor collapse was pattern recognition, not ESP.
- Valid-cue condition: trust intuition only where patterns can be learned from valid cues; otherwise luck or magic.
- Carlsen's double-checking: trusts his ten-second intuition, then verifies because intuition can fail.
- Peggy Noonan: predicted Romney from "vibrations" at rallies—intuition's pull in election forecasting.
- Doubt as the Cure
- Scientific doubt: "What would convince me I am wrong?" is the test of attachment to belief.
- Falsification: strongest evidence comes from experiments that try and fail to disprove a hypothesis.
- Kahneman's warning: high confidence mainly signals a coherent story in the mind, not truth.
- Forecasting today: confident authority and little experimentation, like nineteenth-century medicine.
- Measured doubt: the cure is finely measured uncertainty, reduced by better evidence but never zero.
- System 1's Shortcuts
- Knowledge, Intuition, and Certainty (2. Illusions of Knowledge · I)
- 3. Keeping Score
- The Elusive Truth of Forecasts (3. Keeping Score · I)
- Scorekeeping demands precision
- Ambiguity kills scorekeeping: forecasts must be judged, so meaning must be unambiguous.
- Missing timelines: forecasts without dates become untestable as tacit frames fade.
- Vague terms: "significant" or "serious possibility" can mean opposite things to different readers.
- Probability problem: even "very likely" can't be scored without rerunning history many times.
- Ballmer's iPhone forecast
- Ballmer's infamous call: "No chance the iPhone gets any significant market share" in 2007.
- Full context matters: he meant global mobile phone market, not US smartphone share.
- Outcome: iPhone took ~6% of global mobile phone sales, above his "2% or 3%" but not absurd.
- Tone vs. words: brash dismissal, but words too ambiguous to declare spectacularly wrong.
- The Bernanke letter
- Open letter to Bernanke: quantitative easing "risks currency debasement and inflation."
- No deadline: letter gives no time frame, so critics could always say "just wait."
- Risk language: calling it a "risk" allows non-occurrence without being wrong.
- Experts and hindsight
- Nuclear-war panel: 1984 experts expected neo-Stalinist continuity, not Gorbachev.
- Retrospective rationalization: both liberals and conservatives claimed it was predictable afterward.
- Kahneman's bet: knowledge helps forecasting, but benefits taper off fast.
- Sherman Kent's probability lexicon
- Vague estimates: NIE "serious possibility" meant odds from 80-20 to 20-80 among Kent's team.
- Bay of Pigs: Joint Chiefs' "fair chance" was intended as 3-1 against, but Kennedy read it positively.
- Kent's fix: map words like "probable" to numerical probability ranges.
- Not adopted: the system reduced confusion but intelligence community never embraced it.
- Scorekeeping demands precision
- Numbers, Calibration, and Foxes (3. Keeping Score · II)
- Why Numbers Beat Elastic Language
- Numerical probabilities: people resist them as awkward, unaesthetic, or sounding like a bookie.
- Sherman Kent's retort: "I'd rather be a bookie than a goddamn poet" champions precise odds.
- Numbers are estimates: they express subjective judgment, not objective fact, just like words.
- Metacognition: forcing vague terms into numbers compels forecasters to clarify their thinking.
- Rubbery language: words like "may" let forecasters avoid accountability and exploit ambiguity.
- Historical shift: intelligence agencies adopted numeric odds only after the Iraq WMD debacle.
- The Wrong Side of Maybe
- Wrong-side-of-maybe fallacy: people judge a 70% forecast by whether the event happened, not by the stated odds.
- Single forecasts are unjudgeable: a 70% rain forecast can be right even if it doesn't rain.
- Sophisticated errors: even prediction markets and media misread a 75% chance as a certainty.
- Perverse incentives: vague language lets forecasters stretch or shrink their words after the outcome.
- Clear terms and timelines: forecasts must be precise and numerous before scoring can begin.
- Calibration and Resolution
- Calibration: over many forecasts, events should happen with the same frequency as predicted probabilities.
- Calibration chart: underconfident curves sit above the diagonal; overconfident curves sit below it.
- Perfect calibration alone can be cowardly: a forecaster who only says 40%–60% never strays or informs.
- Resolution: skill at assigning high probabilities to what happens and low probabilities to what doesn't.
- Good scoring combines both: bold correct calls earn more, while overconfident misses are punished.
- Brier Scores and Benchmarks
- Brier score: measures distance between forecast and outcome; lower is better; 0 is perfect, 0.5 guessing, 2.0 worst.
- Betting analogy: an 80% forecast is 4-to-1 odds; claiming 95% puts 19-to-1 at stake.
- Benchmarks matter: a 0.2 Brier score in variable Springfield may beat 0.2 in predictable Phoenix.
- No-change baseline: predictions of "same as last time" are a tough benchmark; Silver's 50-state call barely beat it.
- Level playing field: comparing forecasters fairly requires identical questions, timelines, and difficulty.
- The EPJ Verdict: Foxes Beat Hedgehogs
- EPJ study: 284 experts made 28,000 predictions; results published in Expert Political Judgment.
- Average expert: roughly as accurate as a dart-throwing chimpanzee; many lost to simple algorithms.
- Hedgehogs: Big Idea ideologues were overconfident, dismissive of contrary evidence, and reluctant to update.
- Foxes: pragmatic experts used many tools, expressed probabilities, admitted errors, and changed their minds.
- How to think mattered more than credentials, access, ideology, or optimism.
- Berlin's fox: "The fox knows many things, but the hedgehog knows one big thing" captures the split.
- Why Numbers Beat Elastic Language
- Foxes Win by Aggregating Perspectives (3. Keeping Score · III)
- Hedgehogs and Their Big Idea
- Hedgehogs know one Big Idea; foxes beat them on calibration and resolution.
- Green-tinted glasses: the Big Idea distorts reality, like Oz’s Emerald City.
- Kudlow: supply-side faith made him insist the “Bush boom” was real and deny recession.
- More information only raised hedgehog confidence, not accuracy.
- Expertise backfired: forecasters got worse on their own specialties.
- Fame penalty: the more famous the expert, the less accurate — media reward confident hedgehogs.
- The Wisdom of Crowds
- Galton’s ox: crowd average missed dressed weight by one pound — earliest wisdom-of-crowds demo.
- Why it works: dispersed scraps of valid information pile up; scattered errors cancel out.
- Quality matters: aggregating ignoramuses yields nothing; diverse expertise yields most.
- Poll of polls: aggregating polls and models, as PollyVote does, beats any single source.
- Dragonfly Eyes: Aggregate Perspectives
- Foxes aggregate: they seek many sources and perspectives, then synthesize.
- Guess-the-number: rational Nash answer 0 lost; winning guess was 13.
- Logic plus psycho-logic: Thaler’s game needed Vulcan Spock, human McCoy, and Captain Kirk synthesis.
- Dragonfly metaphor: thousands of lenses turn unique views into one superb vision.
- Inside-one-skull aggregation: foxes do in their heads what crowds do collectively.
- Perspective-Taking Is Hard
- Tip-of-your-nose: default view feels objective, so alternatives are ignored.
- Poker blind spot: even expert students assume a raise means a strong hand.
- Duke’s exercise: they wouldn’t raise with strong hands, yet misread opponents’ raises.
- Effort matters: escaping one’s perspective is a struggle foxes attempt anyway.
- Fox–Hedgehog Spectrum
- Spectrum, not dichotomy: foxes and hedgehogs are poles; hybrids and shifting styles exist.
- Context matters: people can be calculating at work, impulsive at the mall.
- Models: “All models are wrong, but some are useful” — George Box.
- Core finding: modest real foresight exists; thinking style is the critical ingredient.
- Hedgehogs and Their Big Idea
- The Elusive Truth of Forecasts (3. Keeping Score · I)
- 4. Superforecasters
- From Hubris to Evidence-Based Forecasting (4. Superforecasters · I)
- Iraq WMD: A Colossal Intelligence Failure
- National Intelligence Estimate 2002-16HC: declared Saddam had WMDs with certainty, yet none existed.
- Hubris, not politicization: official probes and Robert Jervis found sincere but unquestioned conclusions.
- No red teams: postmortems showed nobody even considered the "no WMD" possibility.
- Stakes: certainty helped pass the invasion resolution; modest odds might have stopped it.
- The Reasonable-Wrong Trap
- Reasonable vs correct: Jervis's Why Intelligence Fails judged the WMD call reasonable on available evidence, though wrong.
- Bait and switch: people replace "was it reasonable?" with "was it correct?".
- Poker lesson: good bets can lose, bad bets can win; evaluate decisions, not outcomes.
- Correctable errors: better analysis would not have changed the conclusion, only lowered confidence.
- Lowered confidence matters: 60–70% odds might have failed the "beyond reasonable doubt" bar.
- Accountability Demands Accuracy Metrics
- No accuracy tracking: the $50 billion intelligence community never systematically assesses forecast accuracy.
- Process accountability: analysts are judged on checklists, not on whether their calls proved right.
- Cochrane principle: don't trust intuitively appealing methods; test them, as medicine learned.
- Untested training: the CIA's bias-awareness manual may help, but its effectiveness is unknown.
- Meaningful accountability: requires systematic accuracy tracking, not just post-failure congressional hearings.
- IARPA Tournament: An Experiment
- IARPA response: after the WMD shock, funded a forecasting tournament to test analytic methods.
- Goldilocks questions: neither too easy for attentive readers nor impossible for anyone.
- Competitive design: teams had to beat a wisdom-of-the-crowd control, by margins growing to 50%.
- Randomized trials: teams ran Archie Cochrane-style experiments to see what truly works.
- Good Judgment Project: thousands of volunteers recruited to forecast global political events.
- Ordinary People Beat Professional Forecasters
- Winning method: average crowd forecasts, upweight top performers, and extremize probabilities.
- Extremizing rationale: simulate informed confidence by pushing combined odds closer to 0 or 100.
- Upset victory: a few hundred volunteers plus simple algorithms beat professional intelligence analysts.
- Doug Lorch: a retired programmer with no expertise exemplifies the unassuming ordinary superforecaster.
- Iraq WMD: A Colossal Intelligence Failure
- Superforecasters, Skill, and Luck (4. Superforecasters · II)
- Doug Lorch, the Amateur Oracle
- Profile: retired with no forecasting credentials; daily hour at the dining table reading news and updating judgments
- Brier scale: measures gap between forecasts and reality; 0 is perfect, 0.5 is random guessing
- Volume: ~1,000 separate forecasts in year 1 because every update counted as a new judgment
- Accuracy: year 1 Brier 0.22, year 2 0.14 — best among 2,800 volunteers
- Benchmarks: beat prediction market, extremizing algorithm, and control crowd by over 60%
- Threat: outperformed salaried professionals for a $250 gift certificate
- The Superforecaster Cohort
- 58 others: first class of superforecasters; group beat regulars by 60% over four years
- Foresight edge: superforecasters at 300 days were more accurate than regulars at 100 days
- Vision analogy: 60% improvement equals going from 20/100 to 20/40 — life-changing
- Professional comparison: reported ~30% better than intelligence analysts reading intercepts
- Institutional threat: agencies avoid tournaments because testing could humble insiders
- The Luck Challenge
- Randomness risk: with 2,800 people and 104 coin flips, extreme streaks arise by chance
- Illusion of prediction: Langer's Yale students overrated skill after an early lucky streak
- Pundit fallacy: a called crash proves nothing when many other forecasts went wrong
- Streak fallacy: six good years on Wall Street is plausible if thousands of investors tried
- Business books: success tales rarely prove causes or acknowledge luck
- Regression to the Mean
- Skill-luck mix: Mauboussin's The Success Equation shows performance blends both, not either/or
- Core concept: extreme results tend to move toward average on retest
- Father-son case: with 0.5 correlation, best guess for a six-foot father's son is five-ten
- Everyday trap: back pain improves after homeopathy due to regression, not treatment
- Diagnostic rule: slow regression signals skill; rapid regression signals chance
- Defying Statistical Gravity
- Opposite pattern: superforecasters increased their lead in years 2 and 3
- Anointment effect: labeling and teaming superforecasters erased expected regression
- Fresh cohorts: later superforecasters also sustained or improved performance
- Individual churn: ~30% of top performers fall each year; 70% persist
- Odds: consistency under skill correlation 0.65 ~1 in 3; under pure luck <1 in 100 million
- Conclusions: not infallible, but mostly skilled — why remains open
- Doug Lorch, the Amateur Oracle
- From Hubris to Evidence-Based Forecasting (4. Superforecasters · I)
- 5. Supersmart?
- Intelligence Is No Substitute for Method (5. Supersmart? · I)
- A Superforecaster's Profile
- Sandy Sillman: atmospheric scientist with multiple sclerosis, joined GJP as a meaningful "transition project" after disability leave.
- Credentialed mind: PhD in applied physics from Harvard, plus fluency in five languages and voracious reading.
- Tournament champion: tied for first place with a Brier score of 0.19, beating roughly 2,800 forecasters.
- Central question: does Sandy's remarkable mind explain his remarkable forecasting results?
- Intelligence and Knowledge Tests
- Sample bias: GJP volunteers were self-selected, not representative of the general population.
- Two intelligences measured: fluid intelligence via pattern puzzles, crystallized intelligence via factual knowledge questions.
- Results: regular forecasters beat 70% of the population; superforecasters beat 80%.
- Threshold effect: above average helps, but genius territory is not required—consistent with Kahneman's hunch about attentive NYT readers.
- McNamara warning: brilliant "best and brightest" made gravely flawed forecasts by never critically analyzing assumptions.
- Bottom line: it is not raw crunching power that counts, but how you use it.
- Fermi-ize: Decompose the Question
- Fermi problems: break unanswerable questions into smaller, inspectable subquestions instead of guessing from a black box.
- Separate knowable and unknowable: decomposition brings guesses into the light where they can be examined.
- Dare to be wrong: Fermi-izing requires overcoming the fear of looking dumb; crude guesses can land surprisingly close.
- Chicago piano tuners: estimated via population, piano ownership, tuning frequency, and work hours—63 tuners, near the yellow pages' ~83 listings.
- Superforecaster habit: Sandy Sillman says Fermi estimation became "part of my natural way of thinking."
- Arafat Poisoning: Question Substitution
- The real question: would French or Swiss inquiries find elevated polonium in Arafat's remains?
- Tip-of-the-nose trap: gut hunches answer "Did Israel do it?" not the question actually asked.
- System 1 bait and switch: the hard question is replaced by an easier, unasked one.
- Fermi-ize the puzzle: Bill Flack, with no Middle East expertise, asked what would make the answer yes or no.
- Apolitical first step: polonium decays quickly, so exhumed remains may not show elevated levels.
- A Superforecaster's Profile
- Base Rates, Hypotheses, Synthesis, Open-Mindedness (5. Supersmart? · II)
- Fermi-Style Decomposition
- Fermi-ize first: break Arafat question into possible pathways to “yes”—Israeli poisoning, Palestinian enemies, postmortem contamination.
- Each pathway raises odds: more ways contamination could happen means higher probability of a positive polonium test.
- Avoid the tiger trap: decomposition exposed the bait-and-switch assumption that Israel must be the only explanation.
- Road map for research: hypothesis list guided subsequent investigation instead of aimless immersion.
- Outside View First
- Start with base rates: ask how common this class of event is before examining case specifics.
- Renzetti pet test: anchor at 62% of American households, then adjust with family details.
- Outside view beats storytelling: concrete inside details tempt narratives; bare base rates are often ignored.
- Noonan’s error: Bush’s approval rebound was normal post-presidency rise, not a meaningful Democratic warning.
- Meaningful anchor matters: anchoring underadjustment makes gut numbers dangerous; a base rate is a better start.
- Fermi can create outside view: if no data, bound the probability—Arafat exhumation case set 20%-80%, midpoint 50%.
- Inside View as Investigation
- Finally, dig inside: after outside view is set, explore case-specific politics and history.
- Targeted hypotheses, not amble: investigate each pathway with evidence pro and con.
- Operationalize conditions: “Israel poisoned Arafat” requires polonium access, motive, and capability.
- Methodical detective work: real investigation is slow and demanding, unlike TV detective clarity.
- Merge, Synthesize, Seek Perspectives
- Dialectical synthesis: fuse outside and inside views like binocular vision; neither alone is enough.
- Rogg’s terrorism forecast: base rate 1.2/year raised to 1.8 for ISIS and security, yielding 34% probability.
- Crowd within: assume your judgment is wrong, re-estimate, and combine; nearly as good as a second person.
- Self-distancing: Bill writes and critiques his own judgment; Soros steps back to judge his own thinking.
- Reword the question: ask “Will South Africa deny the Dalai Lama?” to counter confirmation bias.
- Dragonfly eye: juggle many “on the one hand” perspectives; mentally demanding but superforecasters persist.
- Active Open-Mindedness
- Need for cognition: superforecasters enjoy hard mental work like puzzles.
- Openness to experience: curiosity about Ghana makes unfamiliar questions inviting.
- Behavior beats raw intelligence: self-critical thinking matters more than cognitive horsepower.
- Doug Lorch’s diversity feed: program curates ideologically varied sources to force contrary perspectives.
- Beliefs are hypotheses: to be tested, not guarded treasures—no bumper sticker needed.
- Fermi-Style Decomposition
- Intelligence Is No Substitute for Method (5. Supersmart? · I)
- 6. Superquants?
- Numbers, Judgment, and Maybe (6. Superquants? · I)
- Numeracy Isn't the Secret
- Big Data magic: Clarke's law makes data science look like wizardry from outside.
- Superquant suspicion: superforecasters' numeracy invites a Wall Street quant analogy that fails.
- Levine's proof: math professor and superforecaster who deliberately did it without math.
- Judgment, not models: an occasional Monte Carlo, but forecasting is mostly careful thought and nuanced judgment.
- No moat: numeracy helps, but no castle wall separates ordinary people from superforecasting.
- The Wisdom in Disagreement
- Fictional Panetta craved consensus and admired Maya's certainty—both instincts are wrong.
- Crowd effect: independent, varied estimates are gift-wrapped wisdom; average or weight them.
- Groupthink warning: uniform agreement means minds aren't working independently.
- Real Panetta welcomed estimates from 30% to 90% and demanded honest beliefs over pleasing answers.
- Right but unreasonable: Maya's 100% outran the evidence; correct outcomes don't validate bad probabilities.
- Obama's Fifty-Fifty
- Situation Room split: CIA estimates ranged from 30% to 95%; Obama answered, "This is fifty-fifty."
- Literal 50%: if meant literally, he discarded the room's collective judgment without a basis.
- "Maybe" as reason: as shorthand for uncertainty, fifty-fifty can be a sane threshold for action.
- Ignorance prior: letting disagreement breed distrust wastes evidence already on the table.
- Three settings: Tversky's joke—gonna happen, not gonna happen, maybe—captures human judgment.
- Probability for the Stone Age
- Late arrival: probability theory is a recent invention; ancestors made decisions without it.
- Coarse dials: System 1 says lion or no lion; it cannot feel 60% versus 80%.
- Fast = survival: three quick settings beat fine-grained analysis when a lion may be in the grass.
- Certainty premium: people pay far more to cut risk from 5% to 0 than from 10% to 5.
- Worry-free zones: constant alertness cost too much, so small chances were ignored.
- The Modern Price of Maybe-Phobia
- Confidence bias: people equate confidence with competence though the confidence-accuracy link is weak.
- Hedgehog appeal: media prefer certain forecasters over accurate ones, whatever their records.
- Rain confusion: "70% chance of rain" only makes sense over repeated days; "it will rain" is the natural reading.
- Leonhardt's trap: a 74% forecast can be right and still "wrong"; the 26% branch also happens.
- Rubin's briefing: policy makers heard 80% as certainty, frustrating officials who meant probability.
- Numeracy Isn't the Secret
- Probability, Granularity, and Fate (6. Superquants? · II)
- Uncertainty Is the Baseline
- Certainty is illusory: scientists accept uncertainty because all knowledge is tentative.
- Nothing is 100%: Leon Panetta's remark captures the scientific view of reality.
- Epistemic vs aleatory: some uncertainty is knowable, some unknowable in principle.
- Cloud-like questions: aleatory uncertainty makes life surprising no matter how carefully we plan.
- Superforecasters hedge: for irreducible uncertainty, they stay inside 35–65% and move cautiously.
- Two/three-setting dials fail: yes/no express certainty, leaving only subdivided "maybe."
- Probabilistic Thinking in Action
- Rubin's axiom: rejecting certainty made every judgment a probability, as in In an Uncertain World.
- Precision demanded: a Treasury aide learned to say 60%, not "absolutely," and argue 59 vs 60.
- Binary intuition persists: people grasp 60/40 but translate 80% into "it will happen."
- Fish vs birds: probabilistic and binary thinking rest on different assumptions about reality.
- Intelligence community lags: NIC's five- or seven-degree scale still falls short of superforecaster precision.
- Granularity Is a Superpower
- Superforecasters use fine scales: a third of forecasts use single percentage points.
- Labatte's casual precision: 70/30 corrected to 65/35 shows trained granularity.
- Granularity predicts accuracy: Mellers found finer forecasters beat ten-point users, so precision isn't bafflegab.
- Rounding hurts experts: superforecasters lose accuracy even rounding to nearest 5%.
- Fifty-fifty signals maybe: forecasters who lean on 50% as "maybe" are less accurate.
- Munger's warning: ignoring probability math makes life an ass-kicking contest.
- Chance and Fate Are at War
- "Why me?" is Earthling thinking: Vonnegut's aliens reject the question; probabilistic thinkers ask "why not me?"
- Meaning-making is human: fate beliefs comfort and build resilience after trauma.
- Counterfactual thinking breeds fate: imagining alternate paths turns choices into "meant to be."
- Love story fallacy: tiny odds plus happening becomes "100% meant"—incoherent logic.
- Probabilistic "how" beats metaphysical "why": Shiller sees his existence as indeterminate history, not fate.
- Superforecasters Reject Fate
- Lowest fate scores: superforecasters firmly reject "everything happens for a reason."
- Fate belief undermines accuracy: higher fate scores correlate with worse Brier scores.
- Accuracy vs wellbeing trade-off: meaning is good for resilience but bad for foresight.
- Outside view applies to identity: even life-defining events are quasi-random draws.
- Uncertainty Is the Baseline
- Numbers, Judgment, and Maybe (6. Superquants? · I)
- 7. Supernewsjunkies?
- The Discipline of Updating (7. Supernewsjunkies? · I)
- Forecasts Are Living Judgments
- Initial process: unpack the question, separate known from unknown, take outside/inside views, synthesize, then state precise probabilities.
- No lottery tickets: forecasts are live judgments to revise as information changes, not buy-and-forget draws.
- Updating advantage: superforecasters revise far more often; better-informed forecasts are usually more accurate.
- Not the whole story: even with no updating, superforecasters' initial forecasts were 50% more accurate than regular forecasters'.
- What Smart Updating Looks Like
- Bill Flack, polonium: Swiss delay hinted at extra tests to rule out lead decay, so he raised 60% to 65% — right before others.
- Early and right: Flack's Brier score beat the tournament prediction market by five times on a question that shocked experts.
- Demanding skill: good updating uses the same tough mental work as the initial forecast, and sometimes more.
- Subtle signals: the edge comes from spotting diagnostic details most people miss, not from reacting to what everyone knows.
- Underreaction
- Definition: underreaction gives new evidence too little weight and can destroy a good forecast.
- Distraction case: Joshua Frankel stayed at 82% after Obama's Syria strike made intervention near-certain because he was too swamped.
- Substituted question: Flack asked "If I were PM, would I visit Yasukuni?" instead of "Will Abe?" so he dismissed the aide's signal.
- Belief perseverance: people rationalize endlessly to avoid admitting facts that upset settled beliefs, as with Japanese internment.
- DeWitt's paradox: "no sabotage yet" became proof that sabotage would come — stubborn underreaction at its extreme.
- Overreaction
- Definition: overreaction treats new evidence as more meaningful than it is and adjusts too radically.
- Doug Lorch, Arctic ice: spun from 55% to 95% on a month-old consensus; ice loss slowed, sinking his forecast.
- Timeliness matters: a one-month-old report is weak evidence for a 28-day forecast in a fast-changing system.
- Balance needed: both under- and overreaction distort accuracy; updating requires weighing how much information really changes odds.
- Why Beliefs Resist Change
- Jenga model: beliefs stack like blocks; peripheral ones are easy to toss, identity-laden core blocks resist removal.
- Public commitment: the stronger and more public the commitment to a belief, the harder it is to revise.
- Identity over evidence: Beugoms admits days of denial on military forecasts because West Point and his dissertation make expertise central.
- Protective cognition: Kahan shows judgments about risks follow identity, not evidence — psycho-logic trumps logic.
- Apocryphal Keynes: the famous "facts change" quote has no source; the author, lacking identity investment, easily confessed his error.
- Forecasts Are Living Judgments
- The Art of Belief Updating (7. Supernewsjunkies? · II)
- Why Beliefs Resist Change
- Identity-laden beliefs: Gun-control partisans often refuse to change even when handed conclusive disconfirming evidence.
- Foundational commitments: A belief at the base of the mental tower cannot be removed without collapsing everything above it.
- Earl Warren case: Admitting he had unjustly imprisoned 112,000 people would have shattered his self-image as a civil libertarian.
- Superforecasters' edge: Less ego in any single forecast makes it easier to adjust and avoid underreaction.
- Overreaction: The Other Danger
- Dilution effect: Irrelevant facts weaken stereotypical judgments, showing people overreact to noise they should ignore.
- Market churn: Frequent trading destroys returns; buy-and-hold investors beat hyperactive traders.
- Scylla and Charybdis: Forecasters must steer between underreaction and overreaction; good updating is the middle passage.
- Captain Minto's Updating Style
- Constant revision: Tim Minto re-examined forecasts daily, often making dozens of updates per question.
- Small increments: His average update was just 3.5%, with no dramatic swings—many small steps.
- Sizing an update: Break possible reactions into tiny ranges and settle on a modest, evidence-weighted shift.
- Balancing old and new: Frequent small updates capture the value of prior knowledge plus fresh information.
- Bayesian Wisdom Without Equations
- Bayes' theorem: New belief depends on prior odds multiplied by the diagnostic value of new evidence.
- Hagel case: Base rate gave 96% confirmation odds; a poor hearing moved the forecast to 83%, not 50-50.
- Core insight: Superforecasters internalize Bayesian updating without computing formulas.
- Minto's approach: He is a "Bayesian who does not use Bayes' theorem."
- When Small Updates Are Wrong
- No skeleton key: Many small updates usually work, but sometimes evidence demands a sharp break.
- Doug Lorch's reversal: A discredited report justified dropping from 95% to 15% in one decisive move.
- Orwell's sixth rule: Principles apply, but "break any of these rules sooner than saying anything outright barbarous."
- Why Beliefs Resist Change
- The Discipline of Updating (7. Supernewsjunkies? · I)
- 8. Perpetual Beta
- Growth Mindset, Failure, and Feedback (8. Perpetual Beta · I)
- Growth Mindset and Motivation
- Mary Simpson: missing the 2008 crisis sparked frustration and resolve; she joined the Good Judgment Project and became a superforecaster
- Growth mindset: abilities are largely products of effort, so failure becomes opportunity to improve
- Fixed mindset: abilities are fixed; "I'm bad at math" becomes a self-fulfilling prophecy
- Brain scans: fixed mindsets care only about right/wrong; growth mindsets prioritize information that stretches knowledge
- Keynes: Failure as Fuel
- Consistently inconsistent: Keynes changed his mind ungrudgingly and took pride in admitting mistakes
- Bouncing back: near ruin in 1920 and the 1929 crash became chances to rethink and retry
- Value investing: after 1929, he sought real underlying value; Graham named it, Buffett built on it
- Relentless cycle: try, fail, analyze, adjust, try again — ceaseless self-correction
- Learning by Doing
- Learning by doing: the baby flops backward, absorbs the lesson, and sits steadier next time
- Tacit knowledge: Polanyi's precise bicycle physics cannot teach riding; only bruising practice can
- Forecasting requires practice: reading books is no substitute for making real predictions
- Informed practice: the training booklet boosts accuracy ~10%; practice and reading reinforce each other
- Feedback That Works
- Clear, timely feedback: necessary to learn from failure — you must know when you've failed
- Police lie detection: delayed, messy outcomes make experienced officers overconfident, not more accurate
- Calibration: confidence should match accuracy; overconfidence is common but not an immutable law
- Good calibrators: meteorologists and bridge players improve because results come quickly and clearly
- Barriers to Feedback
- Vague language: elastic words like "probably" cannot be judged; Forer's astrology profiles seemed personal
- Time lag: long-horizon forecasts let ordinary forgetfulness creep in before feedback arrives
- Hindsight bias: knowing what happened distorts your memory of how likely it seemed
- Growth Mindset and Motivation
- Feedback, Grit, and Perpetual Beta (8. Perpetual Beta · II)
- Hindsight Bias Distorts Memories
- Hindsight bias: knowing an outcome skews memory of what you once thought likely.
- Fischhoff experiments: people recalled pre-event estimates slanted toward the actual outcome.
- Soviet collapse study: experts recalled probabilities, on average, 31 points higher than original.
- “I knew it all along”: the effect can be large—20% remembered as 70%.
- Precise Feedback Is Essential
- Ambiguous language + flawed memory: blocks clear feedback, so forecasters cannot learn from results.
- Free throws in the dark: without visible outcomes, practice builds confidence, not skill.
- Brier scores: precise, unambiguous feedback lets forecasters see and own misses.
- Calibration does not transfer: bridge expertise won’t make you a better political forecaster; train in the target domain.
- Analyze and Adjust
- Postmortems: superforecasters question assumptions and study what went wrong after every close.
- Question expertise: Devyn’s polonium lesson—challenge assumptions and seek outside experts.
- Document reasoning: Jean-Pierre left detailed comments so he could reconstruct and critique his thinking.
- Own your luck: Devyn credited good outcomes partly to luck, avoiding the trap of equating outcome with decision quality.
- Counterfactual asymmetry: experts embraced “almost right” stories but dismissed “almost wrong” ones; superforecasters resist this.
- Grit and the Growth Mindset
- Grit: passionate perseverance toward long-term goals despite frustration and failure.
- Elizabeth Sloane: fought brain cancer and volunteered for GJP to “re-grow her synapses.”
- Anne Kilkenny: a housewife with no geopolitical background who pursued rigorous research and self-critique.
- Deeper than praise: Kilkenny saw partisans’ “I knew it” as feelings masquerading as knowledge.
- Perpetual beta: forecasting is a never-ending try-fail-analyze-adjust cycle, not a final product.
- The Superforecaster Profile
- Philosophic outlook: cautious, humble, nondeterministic—nothing is certain, reality is complex.
- Thinking style: actively open-minded, reflective, numerate, and intellectually curious.
- Methods: pragmatic, analytical, dragonfly-eyed, probabilistic, thoughtful updaters, good intuitive psychologists.
- Work ethic: growth mindset plus grit; perpetual beta is the strongest predictor of improvement.
- Perpetual beta ~3x intelligence: superforecasting is about 75% perspiration, 25% inspiration.
- Missing element: other people—decisions are social; what happens when superforecasters work in groups?
- Hindsight Bias Distorts Memories
- Growth Mindset, Failure, and Feedback (8. Perpetual Beta · I)
- 9. Superteams
- Group Judgment From Fiasco to Forecasting (9. Superteams · I)
- Groupthink's Double Edge
- Groupthink: cohesive groups unconsciously maintain shared illusions that block critical reality testing.
- Bay of Pigs: unanimous assent let a secret invasion plan survive front-page exposure and obvious flaws.
- Cuban missile crisis: same advisers, redesigned process, produced a negotiated peace under extreme pressure.
- Process matters: group wisdom or madness depends on culture and method, not just membership.
- Fixing Decision Culture
- Post-fiasco fixes: Kennedy mandated skepticism, generalist questioning, and relentless devil's advocacy.
- Intellectual watchdogs: Sorensen and Robert Kennedy probed every bone of contention, rudely if needed.
- Withheld preferences: JFK kept his preferred airstrike option private so alternatives could get real debate.
- Fresh voices: new advisers and presidential absences kept hierarchy from stifling honest exchange.
- To Team or Not to Team?
- Wisdom of crowds: independent judgments make errors cancel; group discussion can destroy independence.
- Group risks: cognitive loafing and groupthink can reinforce into self-righteous complacency.
- Group benefits: shared information, multiple perspectives, and aggregation improve accuracy.
- Precision Questioning
- Constructive confrontation: Andy Grove's phrase for disagreeing without being disagreeable.
- Precision questioning: dissect vague claims with targeted queries, lowering heat and revealing evidence.
- Counterfactual test: eighty miles of swamp would have exposed the Bay of Pigs escape plan as absurd.
- Superteam Experiment
- Year 1 result: random teams beat solo forecasters by 23% accuracy on average.
- Superteams: top forecasters were formed into teams to test whether elite groups can exceed individuals.
- Online teams: distance helps manage disputes but makes it easier to disregard unseen teammates.
- Hubris risk: acclaim can undermine the habits that produced success, known as CEO disease.
- Groupthink's Double Edge
- From Strangers to Superteams (9. Superteams · II)
- Forging a Team from Strangers
- Quiet start: Elaine withheld opinions despite good forecasts, lacking credentials and confidence.
- Fear of offense: teammates danced around disagreements, avoiding taboo questions like Arafat-polonium.
- Psychological safety: explicit invitations for push-back and thanks for criticism reduced the dancing.
- Norms emerged: experience taught strangers that politeness could hinder critical examination of views.
- Leading from Behind
- No leaders or norms: superteams started unstructured, unlike typical teams with hierarchy fixes.
- Leading by example: Marty Rosenthal modeled detailed explanations and invited comments to shape behavior.
- Initiative without takeover: he organized workload calls, careful not to seem to be seizing control.
- Face-to-face bonds: conferences and Marty's barbecue deepened commitment and willingness to share.
- Superteam Performance
- Information advantage: teams cover far more ground than any individual; Paul Theron's Honduras find paid off.
- Accuracy jump: forecasters placed on superteams became 50% more accurate in years 2 and 3.
- Against prediction markets: superteams beat markets by 15–30%, while ordinary teams beat crowd by ~10%.
- Market caveat: thin prediction-market liquidity may explain part of superteams' edge; markets still strong.
- Collective Open-Mindedness
- Emergent property: team AOM depends on communication patterns, not just individual members' AOM.
- Not sum of parts: open-minded people who don't care can underperform; opinionated truth-seekers can overperform.
- Shared purpose: best teams said "our" not "my," resembling Edmondson's psychologically safe surgical teams.
- Avoiding extremes: superteams escaped both groupthink and flame wars by respectful challenge and admitting ignorance.
- Givers, Diversity, and Limits
- Givers win: pro-social contributors like Marty Rosenthal improve group behavior; none are chumps.
- Giving spreads: Doug Lorch and Tim Minto shared tools and analyses, boosting both their teams and own scores.
- No simple recipe: replicating superteams in organizations risks division, disruption, and uncertain results.
- Obama's advisers: averaging diverse guesses gave roughly 70%; sharing scraps could lift extremized estimate to 80–85%.
- Decision tools: extremized forecasts are cheap enough to deserve a place on presidential desks.
- Forging a Team from Strangers
- Group Judgment From Fiasco to Forecasting (9. Superteams · I)
- 10. The Leader’s Dilemma
- Forecasting Leaders and Mission Command (10. The Leader’s Dilemma · I)
- The Apparent Contradiction
- Leadership demands confidence, decisiveness, vision: nothing can be accomplished without belief it can.
- Superforecasters see uncertainty: nothing certain, deliberation slow, self-critical, ready to change.
- Forecasting rests on humility: Churchill and Jobs are not called humble; mistakes are inevitable.
- Superteams were anarchic: they flourished without hierarchy, but real organizations need structure and leaders.
- Moltke’s Legacy
- Everything is uncertain: no plan survives contact; two cases never alike, improvisation essential.
- Education for judgment: critical thinking and open debate replaced memorization in Prussian war academies.
- Disobey when warranted: Seydlitz refused three direct orders, then attacked at the decisive moment.
- Deliberation and action: decide as circumstances allow; once decided, act with resoluteness and assurance.
- Unwavering yet adaptive: hold the decision until changing circumstances clearly demand a new one.
- Mission Command
- Auftragstaktik: push decisions down; tell subordinates the goal, not how to achieve it.
- Why, not how: the captain knows the intent, so he can improvise around what he finds.
- Independent thinking: every soldier must make calculated, daring use of the situation.
- Short, simple orders: “I don’t care how you do it” suited the German command system.
- Lessons for Leaders
- Eben Emael: lost leaders and gliders could not stop sergeants from improvising and winning.
- Hitler’s failure: holding Normandy reserves for his personal order paralyzed response during invasion.
- Ike’s warning ignored: U.S. Army dismissed his tank ideas as wrong, dangerous, and court-martial worthy.
- Superleader equals superforecaster plus commander: Moltke’s model resolves the apparent contradiction.
- The Apparent Contradiction
- Mission Command and Intellectual Humility (10. The Leader’s Dilemma · II)
- Auftragstaktik and Mission Command
- Auftragstaktik: tell people what to do, not how—they will surprise you with ingenuity.
- Eisenhower's uncertainty: knew nothing is certain; wrote a failure-responsibility note before D-Day.
- Calm confidence: Eisenhower projected steady optimism with troops despite private moodiness.
- Open debate: Eisenhower welcomed well-founded criticism and conceded mistakes, as with Ridgway.
- From Wehrmacht to NATO: Israel used initiative ("plans are platform for change"); US adopted mission command in 1982.
- Petraeus: Developing Flexible Leaders
- Thunder run: bold armored raid into Baghdad succeeded because ground commanders made key calls.
- Petraeus in Mosul: improvised "secure and serve" counterinsurgency on his own initiative, without permission.
- Surprise training: replace scripted live-fire drills with safe, surprising exercises to build flexible leaders.
- Graduate education: encountering different assumptions at universities trains mental flexibility, not just knowledge.
- Doer-thinker divide: false dichotomy—leaders must figure the right move, then execute it boldly.
- Mission Command in Business
- 3M: tell people what to accomplish, not how—near-perfect mission command from R&D chief.
- Amazon: "disagree and commit" principle requires respectful challenge, then wholehearted commitment.
- Walmart: leadership academy built on mission command to develop store managers faster.
- Military vs business: ex-officers find corporations more command-and-control than militaries.
- Intellectual Humility
- Self-confidence doesn't preclude humility: Churchill, Jobs, Petraeus, Eisenhower all had big egos.
- Annie Duke's distinction: be humble before the game's complexity, not before your opponents.
- Intellectual humility: recognition that reality is complex and judgment is fallible—not self-doubt.
- Lincoln's formula: fierce conviction plus "as God gives us to see the right"—humble limits.
- The Wehrmacht Question
- Dissonance is necessary: acknowledge the Wehrmacht was both evil and organizationally effective.
- No moral-competence link: assuming evil equals inept leads to underestimating adversaries.
- First-rate intelligence: hold two opposed ideas and still function, per Fitzgerald.
- Superforecasters struggle too: wishful hope skewed forecasts on Aleppo and North Korea.
- Why use Wehrmacht?: precisely because it makes us squirm and tests forecasters' perspective-taking.
- Auftragstaktik and Mission Command
- Forecasting Leaders and Mission Command (10. The Leader’s Dilemma · I)
- 11. Are They Really So Super?
- Cognitive Illusions and Black Swans (11. Are They Really So Super? · I)
- Kahneman's Challenge
- Kahneman's question: Are superforecasters different kinds of people, or people doing different things? — a bit of both.
- Their edge is what they do, not what they are: research, self-criticism, synthesizing perspectives, granular updating.
- Fragility: System 2 monitoring is exhausting; the feeling of knowing is seductive, so even the best slip into System 1.
- WYSIATI's Trap
- WYSIATI: What You See Is All There Is — the mother of cognitive illusions from the tip-of-your-nose perspective.
- General Flynn's error: daily bad news felt unprecedented, yet conflict and battle-death data show a long decline.
- Unstoppable illusions: Müller-Lyer illusion proves knowing it's an illusion cannot switch it off; only monitor and check with a ruler.
- Scope Insensitivity
- Kahneman's classic finding: willingness to pay the same $10 for cleaning a few lakes or all 250,000 Ontario lakes.
- Bait and switch: people answer "how bad it feels," not the asked valuation; all bird counts yield about $80.
- Forecast scope: Assad's fall probability depends on the time frame; regular forecasters gave 40% for 3 months, 41% for 6.
- Superforecasters: 15% for 3 months and 24% for 6 — not perfect, but good enough to surprise Kahneman.
- Inoculation through practice: deliberate practice can move System 2 corrections into System 1, like a golfer internalizing swing mechanics.
- Thought experiments: mentally varying time frames and targets stress-tests mental models; superforecasters already did this before training.
- Superforecasting Is Hard Work
- Not bulletproof: Kahneman and Tetlock agree — superforecasters are always one System 2 slip from a blown forecast.
- Expectation to stumble: those who do it well appreciate fragility, draw lessons, and keep forecasting.
- Optimism vs Kahneman: training and recruiting can partly inoculate people against cognitive illusions.
- Tools lighten load: software like Doug Lorch's new-source selection program corrects System 1's like-minded bias.
- Taleb's Black Swan Critique
- Black swan claim: history jumps; impactful rare events determine history, so forecasting is a fool's errand.
- False dichotomy: "follow my formula" vs "forecasting is bunk"; both flawed.
- Stringent black swans are rare: 9/11 was anticipated by 1994 and 1998 scenarios; WWI preceded by years of fretting.
- Milder definition: highly improbable consequential event; but data would take decades, centuries, or millennia to accumulate.
- Kahneman's Challenge
- Black Swans, Fat Tails, and Humility (11. Are They Really So Super? · II)
- Limits of Tournament Evidence
- First IARPA tournament: no test of superforecasters spotting gray or black swans.
- History crawls as well as jumps: slow incremental change like life expectancy and economic growth matters profoundly.
- Black swan investing: Khosla spreads many bets; most fail, but rare start-ups yield Google-size fortunes.
- Poker approach: sharper probability estimates beat competitors more often but amass modest fortunes slowly.
- Neither superior: black swan hunting and accurate forecasting are different paths, not competing.
- Black Swans Are Forecastable in Parts
- Black swans are event-plus-consequences: the Bastille means the storming plus the French Revolution it triggered.
- 9/11 sequence: attacks plus Afghanistan invasion made it a black swan; each step was arguably foreseeable.
- Conditional forecasts: superforecasters can handle questions like Taliban compliance and bin Laden’s flight before the invasion.
- Anticipating consequences: forecasting the immediate aftermath can foresee what later becomes a black swan.
- Long-Term Prediction and Planning for Surprise
- Agreement with critics: no evidence forecasts beyond ten years beat obvious truisms; butterfly dynamics limit predictability.
- Ritual forecasting: the Pentagon’s Quadrennial Defense Review demands twenty-year predictions; each decade surprised Wells.
- Plan for surprise: Danzig urges adaptability and resilience; Eisenhower: “Plans are useless, but planning is indispensable.”
- Antifragility: Taleb wants critical systems strengthened by shocks, but resilience is costly.
- Probability judgments unavoidable: building codes and the two-war doctrine hinge on likelihoods and priorities.
- Explicit probabilities: saying what we know and don’t know beats hidden guesses in long-term planning.
- The Mind’s Craving for Certainty
- Kahneman’s insight: minds crave certainty and impose it when they don’t find it.
- Hindsight bias: brushing off surprises makes the past look predictable and the future more predictable than it is.
- Gorbachev surprise: experts quickly rationalized it as explicable after failing to predict it.
- Japan/China caution: confident 1980s Japan dominance forecasts warn against certainty about China’s rise.
- Fat-Tailed History
- Taleb’s core claim: history’s probabilities are distributed like wealth, not height; extreme events far exceed bell-curve odds.
- Wealth example: billionaire frequency is one in 700,000, not one in trillions under normal assumptions.
- War casualties are fat-tailed: a 1914 official using normal deaths would dismiss a 70-million-casualty war as virtually impossible.
- Policy impact: realistic fat-tail risk would have made pre-1914 leaders try harder to avert catastrophe.
- Radical contingency: Hitler, Stalin, and Mao had an 87.5% chance at least one was female; tiny chance reshapes history.
- Counterfactual humility: history is one path among vast possible worlds; butterfly effects can upend best-laid plans.
- Humility and Human-Scale Foresight
- Agreement with Taleb and Kahneman: radical indeterminacy is real; humility is warranted.
- Foresight is puny but not negligible: accurate forecasts on some matters are possible with considerable effort.
- Human scale matters: modest probabilities still guide decisions even amid vast uncertainty.
- Limits of Tournament Evidence
- Cognitive Illusions and Black Swans (11. Are They Really So Super? · I)
- 12. What’s Next?
- From Kto-Kogo to Scorekeeping (12. What’s Next? · I)
- The Drezner Moment
- Scotland referendum: superforecasters beat British betting markets, calling no's 55.3% win
- Drezner's confession: pundit was unsure, surprised by the margin, called it a teachable moment
- Pledge: clear predictions with confidence intervals — "I want to keep score"
- Why Scorekeeping Matters
- Feedback loop: only unambiguous, scorable forecasts yield the clear feedback that improves foresight
- High stakes: forecasting separates prosperity from bankruptcy, sound policy from waste, peace from war
- Iraq 2003: false certainty about WMD helped enable a disastrous invasion
- Fuzzy thinking: vague expectations can never be proven wrong, so mental models never update
- Core cycle: forecast, measure, revise — the surest path to seeing better
- The Kto-Kogo Status Quo
- Lenin's law: politics is "who, whom" — a contest to wield power; accuracy is expendable
- Nate Silver: praised for Obama calls, reviled for Senate calls — same forecaster, same track record
- Santander analyst: fired for an accurate forecast that threatened a rising political power
- Schell and Morris: forecasts that fail as predictions can succeed by mobilizing their tribe
- Change: When Evidence Won
- Codman's End Result System: track ailments, treatments, and outcomes; let evidence judge doctors and hospitals
- Medical resistance: hospitals hated scorekeeping; Codman lost his Mass General and Harvard posts
- Vindication: the core insight triumphed, becoming evidence-based medicine
- Movement spreads: evidence-based policy, Gates Foundation rigor, and Moneyball analytics followed
- IT catalyst: cheap counting and testing accelerate the shift from authority to analysis
- The Intelligence Community's Blame-Game
- Whipsaw: blamed for missing 9/11, then for overhyping Iraq WMD; each rebuke flips the IC to extremes
- Wrong-side-of-maybe: numerical forecasts leave analysts defenseless when the unlikely happens
- Sherman Kent's lesson: fuzzy phrasing survived because numbers invite unfair blame
- Hopeful sign: the IC funded IARPA's tournament and leaders eye scorekeeping
- Attentive public: readers, not pundits, decide whether change or status quo wins
- The Drezner Moment
- Forecasting, Questions, and Collective Wisdom (12. What’s Next? · II)
- The Humanist Objection
- Numbers as tools: quantification is useful but never sacred; quality ranges from wretched to superb.
- Wieseltier's warning: “Where wisdom once was, quantification will now be” names the risk of metric-driven analysis.
- Brier scoring: false alarms equal misses in basic scoring, but weights can be set in advance for real-world stakes.
- Perfection unattainable: progressive improvement is possible; scoring systems remain works in progress.
- Credit-score analogy: imperfect scoring still beats whimsical judgment and podium performance.
- The Big-Question Dilemma
- Saffo's challenge: the question “How does this all turn out?” is most important but too vague to score.
- The dilemma: we must choose between big unscorable questions and small scorable ones—unless we decompose.
- Bayesian question clustering: answer big questions through many small, relevant, scorable questions.
- Pointillist approach: each tiny forecast dot adds little; accumulated dots reveal a larger picture.
- Tournament questions: IARPA's were screened for difficulty and relevance, not trivialities.
- Superforecasters and Superquestioners
- Good judgment: forecasting is only one element; moral judgment and asking good questions are equally essential.
- Smack-the-forehead test: a good question earns “If only I had thought of that before” once hindsight arrives.
- Friedman's Iraq question: correctly identified sectarianism as a key driver, despite his wrong forecast about invasion.
- Friedman's oil column: vague warning of surprises is better read as a question, not a forecast.
- Different mindsets: superquestioners tend toward hedgehog confidence; superforecasters toward foxy uncertainty.
- Tom-Bill symbiosis: the opening frame of the book is a false dichotomy; both questioning and forecasting are needed.
- Why Public Debate Fails
- 2010 signatories' rigidity: Bloomberg follow-up found every respondent still insisted they were right.
- Unfalsifiable forecasts: vague “risk” language with no date lets forecasters dodge disconfirmation.
- Keynesians vs Austerians: after 2008–09 predictions, no camp changed its mind when evidence came in.
- Brute-force debate: Krugman and Ferguson traded failure-catalogs; nobody learned beyond defending positions.
- Adversarial Collaboration
- Kahneman-Klein precedent: rival theorists collaborated under scientific ground rules to reconcile disagreements.
- Precise forecast questions: specify amount, benchmark, and time frame to reduce ambiguity to a minimum.
- Split decisions are progress: if both sides are partly right, reality is more mixed than either thought.
- Good faith required: adversarial collaboration needs participants who want truth more than victory.
- Keep score publicly: clear tests force rationalizers to pay reputational price; audiences learn with them.
- The Humanist Objection
- From Kto-Kogo to Scorekeeping (12. What’s Next? · I)
- Epilogue
- Bill Flack's Example
- Bill Flack: superforecaster who models precision, humility, and self-improvement
- Precision: asks whether "native Cornhusker" fits; exactness fuels forecasting skill
- Success: outstanding Brier score on real-world questions; pundits lack comparable records
- Uncertainty: luck plays a role; stats demand regression to mean in unpredictable phases
- Intellectual Humility
- Self-awareness: knows he cannot lecture on all world regions; defers to genuine experts
- Davos test: seats belong to deep specialists, not those with strong track records
- No cockiness: outperforming pundits doesn't make him dismiss their expertise
- Humility vs confidence: knowing gaps prevents overconfident extreme forecasts
- Using Pundits
- Two-way use: pundits are useful despite worse forecasting records
- Bad pundits: predictions without arguments or based on anecdotes add little
- Good pundits: argue cases like lawyers, enabling adversarial consideration
- Weighted synthesis: superforecasters weigh pundit arguments and background into own forecast
- Baseball metaphor: strategic thinkers can pitch questions and be judged like batters
- Perpetual Beta
- Try, fail, analyze, adjust: cycle drives continuous forecasting improvement
- Brier feedback: scores reveal over- or underconfidence; forecasters modify behavior
- Grit: persistent cycles build resilience and improvement
- Slumps inevitable: irreducible uncertainty means even skilled forecasters will regress
- Bill Flack: perpetual beta — open to adjustment, not fixed identity
- Bill Flack's Example
- Appendix: Ten Commandments for Aspiring Superforecasters
- Choose the Right Battles
- Triage: invest effort in Goldilocks-zone questions, not clocklike or cloud-like ones.
- Break problems down: Fermi-ize intractable questions into knowable parts and rough guesstimates.
- Balance inside and outside views: ask how often things of this sort happen in situations of this sort.
- Expose assumptions: better to discover errors quickly than to hide them behind vague verbiage.
- Update Beliefs Skillfully
- Incremental updating: move probabilities in small, precise steps as evidence trickles in.
- Jump when warranted: respond fast to diagnostic signals, not just noisy news flow.
- Hunt lead indicators: look for nonobvious signs of what would have to happen before X.
- Granular doubt: use numeric probability dials instead of vague words like "maybe" or "likely."
- Apply rigor everywhere: treat national-security odds with the same discipline as sports betting.
- Synthesize and Balance
- Clashing forces: list in advance the signs that would nudge you toward the opposing view.
- Dove-hawk synthesis: merge competing arguments into nuanced judgment, not cookie-cutter positions.
- Balance under- and overconfidence: avoid both rushing to judgment and dawdling near "maybe."
- Tamp both error types: manage misses and false alarms, not just the most recent mistake.
- Learn Honestly
- Own your failures: conduct unflinching postmortems without excuses or hindsight bias.
- Postmortem successes too: you may have lucked out through offsetting errors.
- Beware overlearning: a minor technical slip with big consequences need not invalidate your worldview.
- Practice With Others
- Team skills: cultivate perspective taking, precision questioning, and constructive confrontation.
- Hold the dove: manage groups tightly enough to focus, loosely enough to preserve initiative.
- Error-balancing bicycle: balancing opposing errors is learned by doing with clear feedback.
- Deliberate practice: deep effortful forecasting, not casual news-reading and probability-tossing.
- Guidelines, not rules: no two cases are exactly alike; stay mindful even while following commandments.
- Choose the Right Battles
- Acknowledgments
- A Profoundly Collaborative Project
- First person, collective work: the book's singular voice conceals a large, deeply collaborative research enterprise
- Complementary expertise: statisticians, programmers, political scientists, and administrators each kept the project alive
- Invisible scaffolding: unglamorous operational and administrative work repeatedly prevented collapse
- Institutional Courage Behind the Science
- IARPA's gamble: a David-versus-Goliath tournament funded a small academic team against a bureaucratic giant
- Radical openness: the competition was wholly unclassified, with zero constraints on publishing results
- A rare condition: no other intelligence agency was known to permit such transparency
- Forecasting Skills Are Teachable
- The great discovery: real-world forecasting ability can be taught, not merely measured
- Proof in practice: junior researchers demonstrated that the skill is genuinely learnable
- Meaning and Motivation
- Rooted in grief: the project began amid personal tragedy, filling empty lives with a measure of meaning
- Higher purpose: if its message is heeded, a crazy world might become a bit saner
- Unexpected results: ordinary smart people pushing themselves to the limit surprised everyone, including the researchers
- The Discipline of Simplicity
- Professorial "complexify": the instinct to complicate fundamentally simple points had to be fought for two years
- Foxes will parse: fine-grained detail is left to those who want the endnotes
- Editors won: clarity reached readers because coauthor and editor prevailed
- A Profoundly Collaborative Project
- 1. An Optimistic Skeptic
- Core Conclusion and Practical Takeaways
- The Central Finding
- Forecasting is a trainable skill: measurable foresight exists months to a year out, and deliberate practice improves it
- Superforecasters are made, not born: intelligence, numeracy, and news consumption help only up to a modest threshold
- Process beats pedigree: how you think — open-minded, self-critical, probabilistic — outweighs credentials and fame
- Foxes beat hedgehogs: many tools, admitted uncertainty, and willingness to change minds outperform one Big Idea
- Mindset Shifts
- Doubt is the cure: ask "What would convince me I am wrong?"; measured uncertainty beats confident certainty
- Beliefs are hypotheses: hold them to be tested, not treasures to be guarded
- Perpetual beta: treat every judgment as a work in progress, revised as evidence arrives
- Reject fate: outcomes are quasi-random draws, not "meant to be"; fate beliefs degrade accuracy
- Intellectual humility: be humble before the complexity of the game, not before your opponents
- Daily Practices
- Triage questions: invest in Goldilocks problems, neither clocklike certainties nor cloudlike chaos
- Fermi-ize: decompose intractable questions into small, inspectable subquestions and rough guesstimates
- Outside view first: anchor on base rates before letting case-specific details tell a story
- Then go inside: investigate each pathway with evidence for and against, then synthesize
- Speak in numbers: say 65%, not "likely"; granular probabilities force clearer thinking
- Hedge the irreducible: when uncertainty is unknowable, stay inside roughly 35–65% and move cautiously
- The Learning Loop
- Try, fail, analyze, adjust: relentless self-correction, not any single judgment, produces improvement
- Demand precise feedback: score forecasts with Brier scores so misses become unambiguous
- Postmortem everything: examine failures without excuses, and successes for hidden luck
- Update in small steps: most evidence justifies modest shifts; jump only on truly diagnostic signals
- Steer between errors: avoid underreaction and overreaction — the middle passage is good updating
- Working With Others
- Aggregate perspectives: synthesize many viewpoints like a dragonfly's compound eye
- Crowd within: assume your judgment is wrong, re-estimate, and combine the two
- Teams beat individuals: forecasters placed in superteams grew roughly 50% more accurate
- Build psychological safety: invite push-back, thank critics, question precisely without heat
- Beware both extremes: groupthink and flame wars both destroy accuracy; respectful challenge preserves it
- Leading, Limits, and the Public Good
- Mission command: tell people what to achieve, not how; independent initiative beats rigid orders
- Confidence in action, humility in judgment: decide firmly, then revise when circumstances clearly change
- Plan for surprise: plans are useless, planning indispensable — build resilience and adaptability
- Black swans are partly forecastable: anticipate the event and its aftermath, piece by piece
- Respect the horizon: beyond about ten years, forecasts add little beyond truisms
- Keep score publicly: clear, dated forecasts discipline pundits and let audiences learn alongside
- The Central Finding
opening map…