Within Calibration

When a Wrong Forecast Was Still Reasonable

Missed forecasts are most useful when the review separates bad evidence, bad framing, slow updates, and reasonable uncertainty.

34 sources 3 graphics

On this page

  • Why wrong does not always mean badly judged
  • Five failure types to look for in reviews
  • How to avoid punishing good uncertainty
Preview for When a Wrong Forecast Was Still Reasonable

Introduction

Reviewing a failed prediction is one of the fastest ways to improve judgement—provided the review asks whether the reasoning was sound at the time rather than whether the outcome happened to be favourable. Hindsight bias, often described as the “I knew it all along” effect, makes past events appear more predictable than they really were. Once the outcome is known, people unconsciously rewrite their memory of what they expected, making good decisions look obvious and reasonable uncertainty look like incompetence.[The Decision Lab+2Wikipedia]thedecisionlab.comThe Decision Lab Hindsight BiasAlso called the “knew-it-all-along” effect…Read more…

Miss Reviews illustration 1 Within a prediction journal, the purpose of reviewing missed forecasts is therefore not to punish every incorrect prediction. It is to identify whether the forecast failed because the evidence was poor, the question was framed badly, important information arrived later, or because an unlikely event simply occurred. That distinction is essential for improving confidence calibration rather than merely becoming more cautious.

Why wrong does not always mean badly judged

A prediction should be evaluated against the information that existed when it was made, not against information revealed afterwards. This principle is sometimes called evaluating the process rather than the outcome.

Consider two forecasts:

  • A prediction given a 70% probability fails.
  • A prediction given a 99% probability fails.

The first may represent perfectly reasonable judgement. If many independent 70% predictions are made, roughly three out of ten should fail. The second deserves much closer examination because the stated confidence implied that very little uncertainty remained.

This distinction is closely related to the difference between accuracy and calibration. Good calibration accepts that uncertain events will sometimes go the “wrong” way. Poor calibration usually appears as confidence that was too high, too low, or insufficiently responsive to new evidence. Research from forecasting tournaments such as the Good Judgment Project consistently emphasises continual feedback and recalibration rather than judging forecasters solely by isolated successes or failures.[Wharton Faculty Platform+2Good Judgment]faculty.wharton.upenn.edu2015 superforecastersIn this article, we describe the best-performing strategy of the winning research program: the GJP. To preempt confusion, we…

The opposite mistake is outcome bias: judging a decision purely because the result was good or bad. A poor decision can produce a lucky success, while an excellent decision can end with an unfavourable outcome simply because uncertainty was genuine. Daniel Kahneman has repeatedly argued that outcome bias systematically rewards luck and punishes sound reasoning.[Air Force]af.milnobel laureate gives important decision making tipsAir ForceNobel Laureate gives important decision-making tips23 Apr 2013 — This leads to another bias he called the "outcome bias." This m…

Five failure types to look for in reviews

Instead of asking only “Was I wrong?”, a prediction review becomes much more useful when it classifies why the forecast missed.

1. Bad evidence

Sometimes the available information genuinely pointed in the wrong direction because sources were unreliable, incomplete or misunderstood.

Questions to ask include:

  • Which evidence turned out to be misleading?
  • Did I rely on a single source instead of independent confirmation?
  • Were stronger base rates available but ignored?

The lesson is usually to improve information gathering rather than to reduce confidence across every future prediction.

2. Bad framing

Many forecasting errors begin before any evidence is collected because the question itself was poorly specified.

Examples include:

  • predicting something too vaguely to resolve clearly;
  • combining several independent events into one prediction;
  • failing to define success before the outcome.

Prediction journals work best when every forecast has explicit resolution criteria written in advance, preventing later reinterpretation.

3. Slow updating

Some forecasts start reasonably but remain unchanged after meaningful new information appears.

Review questions include:

  • When did important evidence become available?
  • Did I notice it promptly?
  • If I noticed it, why did my probability remain unchanged?

Research on high-performing forecasters consistently finds that they revise probabilities incrementally as evidence accumulates rather than treating forecasts as positions to defend.[Good Judgment+2Good Judgment]goodjudgment.comGood JudgmentGood Judgment: See the future sooner with SuperforecastingReports that Superforecasters were 30% more accurate than intellig…

4. Overconfidence or underconfidence

The prediction itself may have pointed in roughly the correct direction while the assigned confidence was poorly matched to reality.

Examples include:

  • assigning 95% confidence where 70% was more appropriate;
  • repeatedly using probabilities clustered around 50% because of excessive caution;
  • treating complex events as virtually certain.

Modern calibration research suggests that confidence errors vary across tasks rather than always taking the form of overconfidence. On difficult questions people often become too certain, while on easier questions they may even become underconfident.[Wiley Online Library]onlinelibrary.wiley.comWiley Online LibraryThe effect of calibration training on…Sep 7, 2024 — Results indicate that the training shifted analyst bias toward…

5. Reasonable uncertainty

Some forecasts simply lose because low-probability events happen.

If an event assigned a 20% probability occurs, the prediction is not automatically poor. Across many similar predictions, one in five such events should happen eventually.

The review should therefore ask:

  • Was the event genuinely foreseeable?
  • Was its probability underestimated?
  • Or did a legitimate tail event simply occur?

Treating every surprise as evidence of incompetence eventually encourages unrealistic certainty or excessive timidity—both forms of poor calibration.

Miss Reviews illustration 2

How hindsight bias distorts reviews

[Hindsight bias]WikipediaHindsight bias does more than create embarrassment. It changes memory itself.

Researchers distinguish several related effects:

  • Memory distortion: remembering that you assigned a probability closer to the true outcome than you actually did.
  • Inevitability: believing the outcome “had to happen.”
  • Foreseeability: believing you personally could easily have predicted it.[Wikipedia]WikipediaHindsight biasHindsight bias

These distortions interfere directly with learning because they erase the gap between what was genuinely known beforehand and what became obvious afterwards.

This is one reason written prediction journals outperform retrospective reflection alone. A dated record fixes the original evidence, confidence level and assumptions before memory begins reconstructing them.

Conducting a review without rewriting history

A useful review follows the timeline rather than the outcome.

A practical sequence is:

  1. Read the original prediction exactly as written.
  2. Ignore the eventual outcome initially.
  3. Reconstruct what information was genuinely available at that time.
  4. Identify which assumptions proved correct and which failed.
  5. Ask whether a different probability—not merely a different answer—would have been more appropriate.
  6. Record one concrete change to future forecasting practice.

Notice that the final step concerns improving the forecasting process rather than defending or condemning the original judgement.

Miss Reviews illustration 3

Questions that produce better learning

Certain review questions encourage calibration, while others mainly encourage self-criticism.

Helpful questions include:

  • Which assumption contributed most to the error?
  • Which evidence deserved greater weight?
  • Which evidence did I overweight?
  • What information would genuinely have changed my mind at the time?
  • Did I confuse confidence with conviction?
  • If I faced the identical evidence tomorrow, would I assign a different probability?

Less useful questions include:

  • “How could I have missed something so obvious?”
  • “Why didn’t I know the answer?”
  • “Why was I completely wrong?”

The first set investigates mechanisms; the second merely reacts to outcomes.

How to avoid punishing good uncertainty

One danger in reviewing misses is becoming afraid to make confident predictions at all. Another is becoming so cautious that every forecast drifts toward 50%.

Neither improves calibration.

Instead, keep several principles in mind:

  • Judge forecasts in groups rather than individually. One surprising event says little about overall calibration.
  • Separate confidence errors from directional errors. A correct direction with excessive certainty is different from a wrong direction held cautiously.
  • Reward honest probability estimates, even when they fail.
  • Treat large surprises as opportunities to search for missing evidence rather than automatic proof of poor reasoning.
  • Update forecasting habits only after recurring patterns appear across many reviews.

Forecasting research repeatedly shows that improvement comes from continuous feedback, careful decomposition of mistakes and willingness to revise beliefs—not from trying to eliminate uncertainty altogether. The most effective forecasters treat their beliefs as hypotheses to test rather than identities to defend.[Good Judgment+2Wharton Faculty Platform]goodjudgment.comGood JudgmentBeliefs as Hypotheses: The Superforecaster's MindsetThe best calibrated forecasters treat their beliefs not as sacrosanct tr…

A practical example

Imagine you predicted an 80% chance that a product launch would occur on schedule.

The launch is delayed.

A hindsight-driven review concludes:

“The delays were obvious. I should have seen this coming.”

A calibration-focused review instead asks:

  • Were supplier risks already visible when the prediction was made?
  • Did new information emerge afterwards that could not reasonably have been anticipated?
  • Was 80% too high given the historical rate of delays?
  • Did I ignore warning signs because I trusted optimistic internal reports?
  • Should similar projects receive lower confidence until independent evidence appears?

The second review extracts reusable lessons. The first mainly creates the illusion that the future had been obvious all along.

The real objective of reviewing misses

The purpose of reviewing failed predictions is not to maximise the number of forecasts that turn out correct. It is to make future confidence levels better matched to reality.

That requires preserving uncertainty instead of erasing it with hindsight. A wrong forecast can still represent excellent judgement if it reflected the evidence available at the time and assigned uncertainty appropriately. Conversely, a correct forecast reached through weak reasoning or unjustified certainty may deserve more scrutiny than a thoughtful miss.

Prediction journals become valuable when they preserve that distinction. They transform mistakes from personal verdicts into diagnostic information, helping confidence gradually become better calibrated instead of merely more cautious.

Amazon book picks

Further Reading

Books and field guides related to When a Wrong Forecast Was Still Reasonable. Use these as the next step if you want deeper reading beyond the article.

BookCover for Superforecasting

Superforecasting

By Philip Tetlock, Dan Gardner

Explains probabilistic forecasting, calibration, feedback and evaluating predictions beyond simple outcomes.

eBay marketplace picks

Marketplace Samples

Example marketplace items related to this page. Use the search link to explore similar finds on eBay.

UsingUSA

Endnotes

1. Source: Wikipedia
Title: Hindsight bias
Link:https://en.wikipedia.org/wiki/Hindsight_bias

2. Source: Wikipedia
Link:https://en.wikipedia.org/wiki/The_Good_Judgment_Project

Source snippet

The Good Judgment ProjectThe Good Judgment Project (GJP) is an organization dedicated to "harnessing the wisdom of the crowd to foreca...

3. Source: onlinelibrary.wiley.com
Link:https://onlinelibrary.wiley.com/doi/10.1002/acp.4236?af=R

Source snippet

Wiley Online LibraryThe effect of calibration training on...Sep 7, 2024 — Results indicate that the training shifted analyst bias toward...

4. Source: thedecisionlab.com
Title: The Decision Lab Hindsight Bias
Link:https://thedecisionlab.com/biases/hindsight-bias

Source snippet

Also called the “knew-it-all-along” effect...Read more...

5. Source: faculty.wharton.upenn.edu
Title: 2015 superforecasters
Link:https://faculty.wharton.upenn.edu/wp-content/uploads/2015/07/2015—superforecasters.pdf

Source snippet

In this article, we describe the best-performing strategy of the winning research program: the GJP. To preempt confusion, we...

6. Source: goodjudgment.com
Link:https://goodjudgment.com/

Source snippet

Good JudgmentGood Judgment: See the future sooner with SuperforecastingReports that Superforecasters were 30% more accurate than intellig...

7. Source: af.mil
Title: nobel laureate gives important decision making tips
Link:https://www.af.mil/News/Article-Display/Article/109322/nobel-laureate-gives-important-decision-making-tips/

Source snippet

Air ForceNobel Laureate gives important decision-making tips23 Apr 2013 — This leads to another bias he called the "outcome bias." This m...

8. Source: goodjudgment.com
Link:https://goodjudgment.com/superforecasters-toolbox-beliefs/

Source snippet

Good JudgmentBeliefs as Hypotheses: The Superforecaster's MindsetThe best calibrated forecasters treat their beliefs not as sacrosanct tr...

9. Source: goodjudgment.com
Link:https://goodjudgment.com/philip-tetlocks-10-commandments-of-superforecasting/

Source snippet

Ten Commandments for Aspiring SuperforecastersIn Superforecasting: The Art and Science of Prediction, Good Judgment co-founder Philip Tet...

10. Source: jstor.org
Link:https://www.jstor.org/stable/44280792

Source snippet

Hindsight Biasby NJ Roese · 2012 · Cited by 885 — A meta-analysis of research on hindsight bias. Basic and. Applied Social Psychology, 26...

11. Source: thedecisionlab.com
Link:https://thedecisionlab.com/thinkers/political-science/philip-tetlock

Additional References

12. Source: effectivealtruism.org
Link:https://www.effectivealtruism.org/articles/fireside-chat-with-philip-tetlock

Source snippet

Fireside Chat with Philip TetlockPhilip Tetlock is an expert on forecasting. He's spent decades studying how people make predictions — fr...

13. Source: youtube.com
Link:https://www.youtube.com/watch?v=qjkywetAdxQ

Source snippet

The Hindsight Bias - I Knew It All Along Phenomenon - Psychology in 5 Minutes is highly relevant because it explains how keeping a decisi...

14. Source: journals.sagepub.com
Title: The Consequences of the Hindsight Bias in Medical Decision Making. Show
Link:https://journals.sagepub.com/doi/10.1177/0272989X8800800406

Source snippet

Sage JournalsHindsight Bias: An Impediment to Accurate Probability...by NV Dawson · 1988 · Cited by 133 — Hindsight Bias: An Impediment...

15. Source: esotericlibrary.weebly.com
Title: philip e. tetlock superforecasting the art and science of prediction
Link:https://esotericlibrary.weebly.com/uploads/5/0/7/7/5077636/philip_e.tetlock-_superforecasting_the_art_and_science_of_prediction.pdf

Source snippet

weebly.comSuperforecasting“Philip Tetlock is renowned for demonstrating that most experts are no better than 'dart- throwing monkeys' at...

16. Source: aiimpacts.org
Link:https://aiimpacts.org/evidence-on-good-forecasting-practices-from-the-good-judgment-project/

Source snippet

He calls it “Bayesian Question Clustering.” (Superforecasting 263) The idea is to take the...Read more...

17. Source: pmc.ncbi.nlm.nih.gov
Title: Omission bias is the preference for harm
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC8763848/

Source snippet

Impact of Cognitive Biases on Professionals' Decision...by V Berthet · 2022 · Cited by 287 — Hindsight bias is a propensity to perceive...

18. Source: efcog.org
Link:https://efcog.org/wp-content/uploads/Wgs/Project%20Delivery%20Working%20Group/_Project%20Management%20Subgroup/Risk%20Management%20Task%20Team/Documents/Improve%20Planning%20and%20Forecasting%20and%20Reduce%20Risk%20Using%20Behavioral%20Science%20-%20R0.pdf

Source snippet

was found that many other studies provided evidence of the ability to plan and...

19. Source: medium.com
Link:https://medium.com/%40westofthesun/a-peek-into-the-future-6398ca5c5569

Source snippet

and the insights he's gained from following the predictions...Read more...

20. Source: online.ucpress.edu
Title: Overprecision in the Survey of Professional
Link:https://online.ucpress.edu/collabra/article/10/1/92953/200113/Overprecision-in-the-Survey-of-Professional

Source snippet

University of California PressOverprecision in the Survey of Professional ForecastersFeb 28, 2024 — We find forecasts are overly precise...

21. Source: researchgate.net
Title: Expert Political Judgment: How Good is It?
Link:https://www.researchgate.net/publication/286362106_Expert_Political_Judgment_How_Good_is_It_How_can_We_Know

Source snippet

How can We...Here, Philip E. Tetlock explores what constitutes good judgment in predicting future events, and looks at why experts are o...

Topic Tree

Follow this branch

Parent topic

Calibration Is Your Confidence Matched to Evidence?

Related pages 5