Within Probabilities
Are Your Confidence Levels Honest?
Calibration shows whether confidence levels match what actually happens over many predictions.
On this page
- What calibration means in forecasting
- How repeated scoring reveals overconfidence
- How to improve estimates with feedback
Page outline Jump by section
Introduction
Forecast calibration asks a simple but powerful question: when you say something has a 70% chance of happening, does it actually happen about 70% of the time across many similar predictions? This is different from merely being right or wrong on individual forecasts. A forecaster can occasionally make accurate predictions while consistently overstating or understating their confidence. Calibration therefore provides an evidence-based way to judge whether confidence claims deserve trust.
Within analytical thinking, calibration is valuable because it separates persuasive certainty from demonstrated reliability. Instead of judging confidence by tone, credentials or intuition, it evaluates whether stated probabilities match reality over repeated observations. This makes calibration a cornerstone of probabilistic reasoning in fields ranging from weather forecasting and medicine to intelligence analysis and everyday decision-making.[Good Judgment]goodjudgment.comGood JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin…
What calibration means in forecasting
Calibration measures the relationship between predicted probabilities and observed outcomes.
Imagine a forecaster makes one hundred predictions, each assigned a probability of 80%. If around eighty of those events actually occur, the forecaster is well calibrated at the 80% confidence level. If only sixty occur, they have been overconfident. If ninety-five occur, they have been underconfident.
This differs from asking whether a forecaster is simply “accurate”. Accuracy concerns whether predictions are correct overall. Calibration concerns whether confidence levels honestly represent uncertainty.
A useful way to visualise this is through a calibration curve (sometimes called a reliability diagram). Predictions are grouped into probability ranges—for example 10%, 30%, 50%, 70% and 90%—and compared with the proportion of events that actually happened. Perfect calibration produces points close to a diagonal line where predicted and observed frequencies match.[scikit-learn]scikit-learn.org1.16. Probability calibrationThe calibration module allows you to better calibrate the probabilities of a given model, or to…
Good calibration has two important characteristics:
- Reliable confidence. A stated probability has the meaning it claims.
- Consistency over time. Calibration is assessed across many forecasts, not isolated successes or failures.
A single correct 95% prediction proves nothing about calibration. Only repeated measurement can reveal whether confidence estimates are honest.
Why repeated scoring reveals overconfidence
Human judgement is often more confident than justified by the evidence. Without systematic feedback, this tendency is difficult to notice because memorable successes overshadow forgotten mistakes.
Forecasting research addresses this by requiring participants to assign explicit probabilities before outcomes are known, then scoring those predictions after events resolve. Repeated scoring exposes patterns that intuition alone misses.
For example:
- Someone who frequently assigns 95% confidence but is correct only 75% of the time is substantially overconfident.
- Someone who rarely exceeds 60% confidence despite being correct almost every time is underconfident.
- A well-calibrated forecaster can still be wrong on individual predictions because unlikely events sometimes occur.
The emphasis on repeated measurement also guards against hindsight bias. Once an outcome becomes known, people often feel they “knew it all along”, making their original uncertainty seem smaller than it really was. Recording probabilities beforehand preserves an objective record for later evaluation.[Good Judgment]goodjudgment.comGood JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin…
Calibration is not the same as forecasting skill
Calibration is necessary but not sufficient for excellent forecasting.
Consider two forecasters:
- Forecaster A predicts every event with 50% confidence.
- Forecaster B confidently distinguishes likely from unlikely events while remaining well calibrated.
Forecaster A may appear reasonably calibrated because about half of all events occur, but their forecasts provide little useful information. They never express meaningful differences between situations.
Forecaster B is both calibrated and informative, assigning high probabilities when evidence strongly supports an outcome and low probabilities when it does not.
Forecast researchers therefore distinguish between:
- Calibration (reliability): Do stated probabilities match observed frequencies?
- Resolution or sharpness: Do forecasts meaningfully separate likely from unlikely events?
An ideal forecaster achieves both. Honest probabilities alone are valuable, but the greatest practical value comes from combining reliable confidence with informative discrimination.[Wikipedia+2Emergent Mind]WikipediaBrier scoreBrier score
What scoring rules actually measure
Forecast tournaments commonly use proper scoring rules, mathematical measures that reward honest probability estimates rather than exaggerated confidence.
The best-known example is the Brier score, which measures the squared difference between predicted probabilities and actual outcomes. Lower scores indicate better probabilistic forecasts.
However, an important subtlety is that a low Brier score alone does not prove good calibration. Recent methodological work shows that the score combines several properties—including calibration and discrimination—and can sometimes appear favourable even when systematic probability bias exists. Analysts therefore often examine calibration curves alongside numerical scores instead of relying on a single metric.[PMC+2Wikipedia]pmc.ncbi.nlm.nih.govWe can have perfect predictions, but low or high Brier score due to the underlying…Read more…
This distinction matters because different forecasting systems may have similar overall scores while displaying very different confidence behaviour.
Evidence that calibration can improve
Research suggests calibration is a learnable skill rather than a fixed personal trait.
The Good Judgment Project, originally conducted as part of a large forecasting programme, found that forecasting performance improved through structured training, repeated practice, teamwork, aggregation of independent judgements and continual feedback. Participants repeatedly compared their stated probabilities with actual outcomes, allowing them to adjust their confidence over time.[Good Judgment]goodjudgment.comGood JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin…
More recent experimental work has also found that interactive calibration training can reduce overconfidence. In one study involving repeated sports forecasting over an eleven-week period, participants receiving structured calibration training improved the quality of their probabilistic judgements compared with their starting performance.[Wiley Online Library]onlinelibrary.wiley.comWiley Online LibraryCalibration training for improving probabilistic judgments…by R Gruetzemacher · 2024 · Cited by 4 — We describe an…
The common feature across these studies is not simply making more forecasts, but receiving systematic feedback about whether confidence levels matched reality.
How to improve your own estimates with feedback
Improving calibration requires creating a feedback loop between predictions and outcomes.
A practical approach is:
- Assign explicit probabilities rather than vague phrases such as “probably” or “unlikely”.
- Record predictions before outcomes are known.
- Review groups of forecasts, not individual successes or failures.
- Compare predicted probabilities with observed frequencies.
- Adjust future confidence levels when persistent overconfidence or underconfidence appears.
Many people also benefit from using a standard probability scale—for example 10%, 30%, 50%, 70% and 90%—instead of making arbitrary confidence statements. This encourages consistent interpretation and makes later evaluation easier.
Another useful habit is to separate evidence from confidence. New evidence may justify changing a probability, but confidence should increase only when the evidence genuinely becomes stronger, not simply because time has passed or a preferred outcome feels more plausible.
Common mistakes when judging confidence
Several misconceptions make calibration harder than it appears.
Confusing confidence with expertise. Experts may possess valuable knowledge while still expressing probabilities that are systematically too high or too low. Calibration measures performance rather than reputation.
Evaluating isolated predictions. One spectacular success or failure says little about calibration. The question is how confidence performs across many cases.
Ignoring sample size. Reliable calibration requires enough predictions within each confidence range. Ten predictions at 90% confidence reveal much less than several hundred.
Treating certainty as persuasive. Strongly stated opinions often appear convincing even when historical calibration is poor. Measuring outcomes over time provides a more dependable guide than rhetorical confidence.
Why calibration strengthens analytical thinking
Analytical thinking improves when confidence becomes something that can be tested rather than merely asserted.
Calibration transforms probability estimates into claims that reality can verify. Instead of asking whether someone sounds certain, it asks whether their confidence levels consistently correspond to what actually happens. That shift encourages intellectual humility without abandoning quantitative judgement.
Over time, calibrated forecasters become better not because they eliminate uncertainty, but because they learn to represent it more honestly. Their confidence becomes evidence-based, making decisions easier to compare, improve and trust.
Amazon book picks
Further Reading
Books and field guides related to Are Your Confidence Levels Honest?. Use these as the next step if you want deeper reading beyond the article.
Thinking in Bets
Frames decisions as probabilistic bets and teaches readers to express uncertainty more honestly.
The Signal and the Noise
Explores prediction, uncertainty, probability, and why some forecasters outperform others.
How to Decide
Includes tools for decision quality, uncertainty, feedback, and avoiding overconfidence.
Superforecasting: The Art and Science of Prediction
Directly covers forecasting accuracy, probabilistic thinking, calibration, feedback, and the Good Judgment Project.
Endnotes
1.
Source: Wikipedia
Title: Brier score
Link:https://en.wikipedia.org/wiki/Brier_score
2.
Source: scikit-learn.org
Link:https://scikit-learn.org/stable/modules/calibration.html
Source snippet
1.16. Probability calibrationThe calibration module allows you to better calibrate the probabilities of a given model, or to...
3.
Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC12818272/
Source snippet
We can have perfect predictions, but low or high Brier score due to the underlying...Read more...
4.
Source: Wikipedia
Title: The Good Judgment Project
Link:https://en.wikipedia.org/wiki/The_Good_Judgment_Project
5.
Source: onlinelibrary.wiley.com
Link:https://onlinelibrary.wiley.com/doi/abs/10.1002/ffo2.177
Source snippet
Wiley Online LibraryCalibration training for improving probabilistic judgments...by R Gruetzemacher · 2024 · Cited by 4 — We describe an...
6.
Source: Wikipedia
Link:https://en.wikipedia.org/wiki/Forecasting
Source snippet
ForecastingForecasting is the process of making predictions based on past and present data. These forecasts can later be compared with...
7.
Source: forum.effectivealtruism.org
Title: In one
Link:https://forum.effectivealtruism.org/posts/pnpnqA4hijnr59p7d/efforts-to-improve-the-accuracy-of-our-judgments-and
Source snippet
to Improve the Accuracy of Our Judgments and...Oct 25, 2016 — Evidence indicates that the calibration of judgment can be substantially e...
8.
Source: goodjudgment.com
Link:https://goodjudgment.com/about/the-science-of-superforecasting/
Source snippet
Good JudgmentThe Science Of SuperforecastingGood Judgment research discovered four keys to accurate forecasting: talent-spotting, trainin...
9.
Source: emergentmind.com
Title: brier score term
Link:https://www.emergentmind.com/topics/brier-score-term
Source snippet
Brier Score: Calibration, Resolution, and Uncertainty25 Jul 2025 — The Brier score term evaluates probabilistic forecasts by decomposing...
10.
Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC10189590/
Source snippet
improves forecasting - PMC - NIHby DN Ferreiro · 2023 · Cited by 4 — We test this by analysing 5 years of data from the Good Judgement Pr...
11.
Source: mdpi.com
Title: Forecasting | An Open Access Journal from MDPIForecasting
Link:https://www.mdpi.com/journal/forecasting
Source snippet
Forecasting is an international, peer-reviewed, open access journal on all aspects of forecasting published bimonthly online by MDPI. Ope...
12.
Source: aws.amazon.com
Link:https://aws.amazon.com/what-is/forecast/?tag=searcht-20
Source snippet
ing Models ExplainedA forecast is a prediction made by studying historical data and past patterns. Businesses use software tools and syst...
Additional References
13.
Source: accaglobal.com
Link:https://www.accaglobal.com/gb/en/student/exam-support-resources/professional-exams-study-resources/p5/technical-articles/forecasting-data.html
Source snippet
Forecasting with dataThis article looks at some important techniques which could be used in forecasting. You should already be familiar w...
14.
Source: hbr.org
Link:https://hbr.org/1971/07/how-to-choose-the-right-forecasting-technique
Source snippet
How to Choose the Right Forecasting TechniqueMany forecasting techniques have been developed in recent years. Each has its special use, a...
15.
Source: researchgate.net
Link:https://www.researchgate.net/figure/Mean-standardized-[Brier-scores
Source snippet
hich reality is coded as 1 for the event and 0 otherwise), ranging from 0 (...
16.
Source: alphaus.cloud
Title: forecasting understanding its core meaning
Link:https://www.alphaus.cloud/en/blog/forecasting-understanding-its-core-meaning
Source snippet
Forecasting: Understanding Its Core Meaning28 Mar 2025 — Definition and Importance: Forecasting is the process of making informed predict...
17.
Source: investopedia.com
Title: It can involve projections for specific business metrics.Read more
Link:https://www.investopedia.com/articles/financial-theory/11/basics-business-forcasting.asp
Source snippet
Business Forecasting: Key Methods and Models for SuccessBusiness forecasting is the process of making informed predictions about future b...
18.
Source: commoncog.com
Title: how do you evaluate your own predictions
Link:https://commoncog.com/how-do-you-evaluate-your-own-predictions/
Source snippet
?Dec 17, 2019 — This post provides a comprehensive summary of the technique that Tetlock and Gardner presents in Superforecasting...
19.
Source: indeed.com
Title: What Is Forecasting?
Link:https://www.indeed.com/career-advice/career-development/what-is-forecasting
Source snippet
Definition, Methods and Examples15 Dec 2025 — Forecasting is a method of making informed predictions by using historical data as the main...
20.
Source: dataopsschool.com
Title: What is Brier Score?
Link:https://dataopsschool.com/blog/brier-score/
Source snippet
Meaning, Architecture, Examples, Use...17 Feb 2026 — The Brier Score measures the accuracy of probabilistic predictions. Reliability dia...
21.
Source: youtube.com
Link:https://www.youtube.com/watch?v=oc0LWUvjlL4
Source snippet
Sensitivity and Specificity Explained Clearly (Biostatistics)...
22.
Source: ibm.com
Link:https://www.ibm.com/think/topics/forecasting
Source snippet
n previous and current data.Read more...
Topic Tree



