Within Probabilities

Are Your Confidence Levels Honest?

Calibration shows whether confidence levels match what actually happens over many predictions.

25 sources 3 graphics

On this page

  • What calibration means in forecasting
  • How repeated scoring reveals overconfidence
  • How to improve estimates with feedback
Preview for Are Your Confidence Levels Honest?

Introduction

Forecast calibration asks a simple but powerful question: when you say something has a 70% chance of happening, does it actually happen about 70% of the time across many similar predictions? This is different from merely being right or wrong on individual forecasts. A forecaster can occasionally make accurate predictions while consistently overstating or understating their confidence. Calibration therefore provides an evidence-based way to judge whether confidence claims deserve trust.

Calibration illustration 1 Within analytical thinking, calibration is valuable because it separates persuasive certainty from demonstrated reliability. Instead of judging confidence by tone, credentials or intuition, it evaluates whether stated probabilities match reality over repeated observations. This makes calibration a cornerstone of probabilistic reasoning in fields ranging from weather forecasting and medicine to intelligence analysis and everyday decision-making.[Good Judgment]goodjudgment.comGood JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin…

What calibration means in forecasting

Calibration measures the relationship between predicted probabilities and observed outcomes.

Imagine a forecaster makes one hundred predictions, each assigned a probability of 80%. If around eighty of those events actually occur, the forecaster is well calibrated at the 80% confidence level. If only sixty occur, they have been overconfident. If ninety-five occur, they have been underconfident.

This differs from asking whether a forecaster is simply “accurate”. Accuracy concerns whether predictions are correct overall. Calibration concerns whether confidence levels honestly represent uncertainty.

A useful way to visualise this is through a calibration curve (sometimes called a reliability diagram). Predictions are grouped into probability ranges—for example 10%, 30%, 50%, 70% and 90%—and compared with the proportion of events that actually happened. Perfect calibration produces points close to a diagonal line where predicted and observed frequencies match.[scikit-learn]scikit-learn.org1.16. Probability calibrationThe calibration module allows you to better calibrate the probabilities of a given model, or to…

Good calibration has two important characteristics:

  • Reliable confidence. A stated probability has the meaning it claims.
  • Consistency over time. Calibration is assessed across many forecasts, not isolated successes or failures.

A single correct 95% prediction proves nothing about calibration. Only repeated measurement can reveal whether confidence estimates are honest.

Why repeated scoring reveals overconfidence

Human judgement is often more confident than justified by the evidence. Without systematic feedback, this tendency is difficult to notice because memorable successes overshadow forgotten mistakes.

Forecasting research addresses this by requiring participants to assign explicit probabilities before outcomes are known, then scoring those predictions after events resolve. Repeated scoring exposes patterns that intuition alone misses.

For example:

  • Someone who frequently assigns 95% confidence but is correct only 75% of the time is substantially overconfident.
  • Someone who rarely exceeds 60% confidence despite being correct almost every time is underconfident.
  • A well-calibrated forecaster can still be wrong on individual predictions because unlikely events sometimes occur.

The emphasis on repeated measurement also guards against hindsight bias. Once an outcome becomes known, people often feel they “knew it all along”, making their original uncertainty seem smaller than it really was. Recording probabilities beforehand preserves an objective record for later evaluation.[Good Judgment]goodjudgment.comGood JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin…

Calibration is not the same as forecasting skill

Calibration is necessary but not sufficient for excellent forecasting.

Consider two forecasters:

  • Forecaster A predicts every event with 50% confidence.
  • Forecaster B confidently distinguishes likely from unlikely events while remaining well calibrated.

Forecaster A may appear reasonably calibrated because about half of all events occur, but their forecasts provide little useful information. They never express meaningful differences between situations.

Forecaster B is both calibrated and informative, assigning high probabilities when evidence strongly supports an outcome and low probabilities when it does not.

Forecast researchers therefore distinguish between:

  • Calibration (reliability): Do stated probabilities match observed frequencies?
  • Resolution or sharpness: Do forecasts meaningfully separate likely from unlikely events?

An ideal forecaster achieves both. Honest probabilities alone are valuable, but the greatest practical value comes from combining reliable confidence with informative discrimination.[Wikipedia+2Emergent Mind]WikipediaBrier scoreBrier score

Calibration illustration 2

What scoring rules actually measure

Forecast tournaments commonly use proper scoring rules, mathematical measures that reward honest probability estimates rather than exaggerated confidence.

The best-known example is the Brier score, which measures the squared difference between predicted probabilities and actual outcomes. Lower scores indicate better probabilistic forecasts.

However, an important subtlety is that a low Brier score alone does not prove good calibration. Recent methodological work shows that the score combines several properties—including calibration and discrimination—and can sometimes appear favourable even when systematic probability bias exists. Analysts therefore often examine calibration curves alongside numerical scores instead of relying on a single metric.[PMC+2Wikipedia]pmc.ncbi.nlm.nih.govWe can have perfect predictions, but low or high Brier score due to the underlying…Read more…

This distinction matters because different forecasting systems may have similar overall scores while displaying very different confidence behaviour.

Evidence that calibration can improve

Research suggests calibration is a learnable skill rather than a fixed personal trait.

The Good Judgment Project, originally conducted as part of a large forecasting programme, found that forecasting performance improved through structured training, repeated practice, teamwork, aggregation of independent judgements and continual feedback. Participants repeatedly compared their stated probabilities with actual outcomes, allowing them to adjust their confidence over time.[Good Judgment]goodjudgment.comGood JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin…

More recent experimental work has also found that interactive calibration training can reduce overconfidence. In one study involving repeated sports forecasting over an eleven-week period, participants receiving structured calibration training improved the quality of their probabilistic judgements compared with their starting performance.[Wiley Online Library]onlinelibrary.wiley.comWiley Online LibraryCalibration training for improving probabilistic judgments…by R Gruetzemacher · 2024 · Cited by 4 — We describe an…

The common feature across these studies is not simply making more forecasts, but receiving systematic feedback about whether confidence levels matched reality.

How to improve your own estimates with feedback

Improving calibration requires creating a feedback loop between predictions and outcomes.

A practical approach is:

  1. Assign explicit probabilities rather than vague phrases such as “probably” or “unlikely”.
  2. Record predictions before outcomes are known.
  3. Review groups of forecasts, not individual successes or failures.
  4. Compare predicted probabilities with observed frequencies.
  5. Adjust future confidence levels when persistent overconfidence or underconfidence appears.

Many people also benefit from using a standard probability scale—for example 10%, 30%, 50%, 70% and 90%—instead of making arbitrary confidence statements. This encourages consistent interpretation and makes later evaluation easier.

Another useful habit is to separate evidence from confidence. New evidence may justify changing a probability, but confidence should increase only when the evidence genuinely becomes stronger, not simply because time has passed or a preferred outcome feels more plausible.

Calibration illustration 3

Common mistakes when judging confidence

Several misconceptions make calibration harder than it appears.

Confusing confidence with expertise. Experts may possess valuable knowledge while still expressing probabilities that are systematically too high or too low. Calibration measures performance rather than reputation.

Evaluating isolated predictions. One spectacular success or failure says little about calibration. The question is how confidence performs across many cases.

Ignoring sample size. Reliable calibration requires enough predictions within each confidence range. Ten predictions at 90% confidence reveal much less than several hundred.

Treating certainty as persuasive. Strongly stated opinions often appear convincing even when historical calibration is poor. Measuring outcomes over time provides a more dependable guide than rhetorical confidence.

Why calibration strengthens analytical thinking

Analytical thinking improves when confidence becomes something that can be tested rather than merely asserted.

Calibration transforms probability estimates into claims that reality can verify. Instead of asking whether someone sounds certain, it asks whether their confidence levels consistently correspond to what actually happens. That shift encourages intellectual humility without abandoning quantitative judgement.

Over time, calibrated forecasters become better not because they eliminate uncertainty, but because they learn to represent it more honestly. Their confidence becomes evidence-based, making decisions easier to compare, improve and trust.

Amazon book picks

Further Reading

Books and field guides related to Are Your Confidence Levels Honest?. Use these as the next step if you want deeper reading beyond the article.

eBay marketplace picks

Marketplace Samples

Example marketplace items related to this page. Use the search link to explore similar finds on eBay.

UsingUSA

Endnotes

1. Source: Wikipedia
Title: Brier score
Link:https://en.wikipedia.org/wiki/Brier_score

2. Source: scikit-learn.org
Link:https://scikit-learn.org/stable/modules/calibration.html

Source snippet

1.16. Probability calibrationThe calibration module allows you to better calibrate the probabilities of a given model, or to...

3. Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC12818272/

Source snippet

We can have perfect predictions, but low or high Brier score due to the underlying...Read more...

4. Source: Wikipedia
Title: The Good Judgment Project
Link:https://en.wikipedia.org/wiki/The_Good_Judgment_Project

5. Source: onlinelibrary.wiley.com
Link:https://onlinelibrary.wiley.com/doi/abs/10.1002/ffo2.177

Source snippet

Wiley Online LibraryCalibration training for improving probabilistic judgments...by R Gruetzemacher · 2024 · Cited by 4 — We describe an...

6. Source: Wikipedia
Link:https://en.wikipedia.org/wiki/Forecasting

Source snippet

ForecastingForecasting is the process of making predictions based on past and present data. These forecasts can later be compared with...

7. Source: forum.effectivealtruism.org
Title: In one
Link:https://forum.effectivealtruism.org/posts/pnpnqA4hijnr59p7d/efforts-to-improve-the-accuracy-of-our-judgments-and

Source snippet

to Improve the Accuracy of Our Judgments and...Oct 25, 2016 — Evidence indicates that the calibration of judgment can be substantially e...

8. Source: goodjudgment.com
Link:https://goodjudgment.com/about/the-science-of-superforecasting/

Source snippet

Good JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin...

9. Source: emergentmind.com
Title: brier score term
Link:https://www.emergentmind.com/topics/brier-score-term

Source snippet

Brier Score: Calibration, Resolution, and Uncertainty25 Jul 2025 — The Brier score term evaluates probabilistic forecasts by decomposing...

10. Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC10189590/

Source snippet

improves forecasting - PMC - NIHby DN Ferreiro · 2023 · Cited by 4 — We test this by analysing 5 years of data from the Good Judgement Pr...

11. Source: mdpi.com
Title: Forecasting | An Open Access Journal from MDPIForecasting
Link:https://www.mdpi.com/journal/forecasting

Source snippet

Forecasting is an international, peer-reviewed, open access journal on all aspects of forecasting published bimonthly online by MDPI. Ope...

12. Source: aws.amazon.com
Link:https://aws.amazon.com/what-is/forecast/?tag=searcht-20

Source snippet

ing Models ExplainedA forecast is a prediction made by studying historical data and past patterns. Businesses use software tools and syst...

Additional References

13. Source: accaglobal.com
Link:https://www.accaglobal.com/gb/en/student/exam-support-resources/professional-exams-study-resources/p5/technical-articles/forecasting-data.html

Source snippet

Forecasting with dataThis article looks at some important techniques which could be used in forecasting. You should already be familiar w...

14. Source: hbr.org
Link:https://hbr.org/1971/07/how-to-choose-the-right-forecasting-technique

Source snippet

How to Choose the Right Forecasting TechniqueMany forecasting techniques have been developed in recent years. Each has its special use, a...

15. Source: researchgate.net
Link:https://www.researchgate.net/figure/Mean-standardized-[Brier-scores

Source snippet

hich reality is coded as 1 for the event and 0 otherwise), ranging from 0 (...

16. Source: alphaus.cloud
Title: forecasting understanding its core meaning
Link:https://www.alphaus.cloud/en/blog/forecasting-understanding-its-core-meaning

Source snippet

Forecasting: Understanding Its Core Meaning28 Mar 2025 — Definition and Importance: Forecasting is the process of making informed predict...

17. Source: investopedia.com
Title: It can involve projections for specific business metrics.Read more
Link:https://www.investopedia.com/articles/financial-theory/11/basics-business-forcasting.asp

Source snippet

Business Forecasting: Key Methods and Models for SuccessBusiness forecasting is the process of making informed predictions about future b...

18. Source: commoncog.com
Title: how do you evaluate your own predictions
Link:https://commoncog.com/how-do-you-evaluate-your-own-predictions/

Source snippet

?Dec 17, 2019 — This post provides a comprehensive summary of the technique that Tetlock and Gardner presents in Superforecasting...

19. Source: indeed.com
Title: What Is Forecasting?
Link:https://www.indeed.com/career-advice/career-development/what-is-forecasting

Source snippet

Definition, Methods and Examples15 Dec 2025 — Forecasting is a method of making informed predictions by using historical data as the main...

20. Source: dataopsschool.com
Title: What is Brier Score?
Link:https://dataopsschool.com/blog/brier-score/

Source snippet

Meaning, Architecture, Examples, Use...17 Feb 2026 — The Brier Score measures the accuracy of probabilistic predictions. Reliability dia...

21. Source: youtube.com
Link:https://www.youtube.com/watch?v=oc0LWUvjlL4

Source snippet

Sensitivity and Specificity Explained Clearly (Biostatistics)...

22. Source: ibm.com
Link:https://www.ibm.com/think/topics/forecasting

Source snippet

n previous and current data.Read more...

Topic Tree

Follow this branch

Parent topic

Probabilities How Do You Think Through Uncertainty?

Related pages 5