Within Deliberate Practice
How sure should you really be?
Calibration practice teaches people to attach numbers to uncertainty, check outcomes, and notice when confidence outruns accuracy.
On this page
- Why confidence feels clearer than accuracy
- Using probabilities, intervals and outcome checks
- Keeping a calibration log without overcomplicating it
Page outline Jump by section
Introduction
Calibration is the habit of matching your confidence to the strength of the evidence rather than to how convincing an idea feels. Someone who is well calibrated is not necessarily more knowledgeable than everyone else; they are simply better at saying “I am about 70% sure” when they are right about 70% of the time, and reserving near-certainty for genuinely exceptional cases. This makes calibration one of the most practical reasoning skills to train because it converts vague confidence into something measurable and improvable. Research from psychology, forecasting and decision science consistently shows that structured feedback on probability judgements can reduce overconfidence and improve the quality of decisions, even with relatively short training programmes.[Wiley Online Library]onlinelibrary.wiley.comWiley Online LibraryCalibration training for improving probabilistic judgments…by R Gruetzemacher · 2024 · Cited by 4 — We describe an…
Within deliberate practice, calibration is valuable because it creates a feedback loop. Instead of asking only whether you were correct, you also ask whether your level of confidence deserved to be that high. Over time, this helps separate genuine expertise from the illusion of certainty.
Why confidence feels clearer than accuracy
Human intuition naturally produces feelings of certainty, but those feelings are not direct measurements of truth. Familiarity, fluency, strong narratives and recent experiences can all make an answer feel more convincing than the available evidence justifies. As a result, people often display overprecision—being too certain that a particular answer is correct—even when their overall knowledge is reasonable.[Wikipedia]WikipediaOverconfidence effectOverconfidence effect
A useful distinction is between accuracy and calibration.
- Accuracy asks whether the answer was correct.
- Calibration asks whether the stated confidence matched the long-run frequency of being correct.
Imagine answering 100 questions while claiming 90% confidence each time. Perfect calibration means approximately 90 of those answers should be correct. If only 70 are correct, confidence has outrun reality. Conversely, consistently being correct 90% of the time while claiming only 60% confidence indicates underconfidence rather than good judgement.[Wikipedia]WikipediaCalibrated probability assessmentCalibrated probability assessment
This distinction matters because many important decisions involve uncertainty that cannot be eliminated. Better judgement often comes not from becoming certain, but from expressing uncertainty honestly enough that later feedback can improve future decisions.
Using probabilities, intervals and outcome checks
Calibration drills work because they force uncertainty into the open. Instead of recording only conclusions, they require explicit numerical estimates that can later be checked.
Assign probabilities instead of certainty
A simple exercise is to replace statements like “I’m sure” or “I think probably” with numerical probabilities.
For example:
- “I am 60% confident this meeting will be cancelled.”
- “I am 80% confident this diagnosis is correct.”
- “There is a 30% chance this project finishes before Friday.”
The exact number is less important than committing to one before the outcome is known. Once the result arrives, compare prediction with reality.
Over dozens or hundreds of predictions, patterns emerge. Some people discover that every 90% judgement behaves like an actual 70% judgement. Others find that they avoid expressing high confidence even when evidence supports it. Both patterns become visible only when confidence is quantified.[Wikipedia]WikipediaCalibrated probability assessmentCalibrated probability assessment
Practise with confidence intervals
Not every judgement is yes-or-no. Many involve estimating quantities.
Instead of predicting a single value, estimate a range that you believe has a specified chance of containing the truth.
Examples include:
- “I am 90% confident the population lies between X and Y.”
- “I am 90% confident this task will take between three and six hours.”
- “I am 90% confident the answer falls within this numerical range.”
Most people make these intervals far too narrow. Studies repeatedly find that intervals intended to capture the correct answer 90% of the time often succeed only about half the time, revealing substantial overconfidence. A practical correction is to widen intervals until observed success approaches the intended confidence level.[Wikipedia]WikipediaOverconfidence effectOverconfidence effect
Separate judgement from outcome
One incorrect prediction does not necessarily indicate poor reasoning, just as one lucky success does not prove good judgement.
Suppose you assign a 20% chance that a supplier misses a deadline. If the supplier does miss it, your prediction was not “wrong”; a one-in-five event should occur occasionally. Likewise, assigning 99% confidence to an outcome that fails reveals a serious calibration problem even if similar predictions often succeed.
Calibration therefore evaluates many predictions together rather than judging isolated successes or failures.[Wikipedia]WikipediaBrier scoreBrier score
Keeping a calibration log without overcomplicating it
An effective calibration log can remain remarkably simple. The goal is consistency rather than elaborate record-keeping.
For each prediction, record:
ItemExampleDate25 JunePrediction”Client approves proposal this week.”Confidence70%Reason”Strong prior relationship; minor revisions outstanding.”OutcomeApproved after four daysReflectionConfidence slightly high because I ignored scheduling delays.
After accumulating around 30–50 predictions, review them by confidence level.
For example:
Confidence statedActual success rate60%58%70%73%80%66%90%74%
The objective is not perfection but convergence. If your 90% predictions repeatedly succeed only about three-quarters of the time, reserve 90% confidence for much stronger evidence.
Forecasting practitioners often recommend reviewing predictions in batches rather than emotionally reacting to individual mistakes. This encourages learning from statistical patterns rather than memorable anecdotes.[Commoncog]commoncog.comhow do you evaluate your own predictionsHow Do You Evaluate Your Own Predictions?Dec 17, 2019 — This post provides a comprehensive summary of the technique that Tetlock…
Calibration drills that target overconfidence
Different drills expose different forms of misplaced certainty.
The “Could I be wrong?” drill
Before finalising an answer, ask:
- What evidence would change my mind?
- What probability would I assign to the strongest alternative?
- What assumption am I treating as certain without checking?
The aim is not endless doubt but preventing unwarranted certainty from becoming invisible.
The outside-view adjustment
After making an initial estimate, compare it with similar past situations.
Instead of asking only, “How will this project unfold?” also ask, “How long have comparable projects actually taken?”
External reference points frequently reveal that initial confidence relied too heavily on the unique story rather than historical evidence.
Forced probability bins
Restrict yourself to confidence categories such as:
- 55%
- 65%
- 75%
- 85%
- 95%
This discourages meaningless precision such as claiming 83.7% confidence while still requiring genuine distinctions between weak, moderate and strong evidence.
Delayed answer drill
When possible, write down a prediction before gathering additional opinions or seeing the eventual result.
This prevents hindsight from rewriting your memory of how certain you really were. Hindsight bias often makes people remember having been “obviously right” even when their original judgement was much less confident.
Measuring improvement without chasing perfect scores
Calibration improves through repeated prediction, feedback and adjustment rather than through occasional dramatic insights.
One common evaluation method is the Brier score, which measures how closely probability estimates match actual outcomes. Unlike simple right-or-wrong scoring, it rewards honest probabilities and penalises unwarranted certainty. Researchers frequently use it when evaluating forecasters because it captures both correctness and the quality of confidence estimates.[Wikipedia]WikipediaBrier scoreBrier score
Importantly, calibration is only one dimension of good judgement. A perfectly calibrated forecaster who always predicts 50% confidence provides little useful information. Good judgement combines calibration with discrimination—assigning higher probabilities when events are genuinely more likely and lower probabilities when they are not. Strong forecasting balances both rather than maximising one at the expense of the other.[PMC]pmc.ncbi.nlm.nih.govOn misconceptions about the Brier score in binary prediction…by L Hoessly · 2026 · Cited by 11 — Calibration refers to how well pre…
Recent experimental work also suggests that calibration itself can be trained. In one forecasting study, participants using an interactive calibration programme showed modest but measurable reductions in overconfidence and improvements in both calibration and overall probabilistic accuracy after less than half an hour of structured training, demonstrating that calibration is a learnable reasoning skill rather than a fixed personal trait.[Wiley Online Library]onlinelibrary.wiley.comWiley Online LibraryCalibration training for improving probabilistic judgments…by R Gruetzemacher · 2024 · Cited by 4 — We describe an…
Common mistakes that weaken calibration practice
Several habits make calibration exercises less effective than they appear.
- Recording only memorable predictions. Successful calibration requires keeping ordinary predictions as well as dramatic ones.
- Changing confidence after the outcome. Confidence must be recorded before feedback arrives.
- Using only 0% or 100%. Absolute certainty is rarely justified outside logical or mathematical facts.
- Confusing confidence with commitment. You can act decisively while acknowledging uncertainty.
- Stopping after a handful of predictions. Reliable calibration emerges from repeated measurement across many judgements, not isolated examples.[Wikipedia]WikipediaCalibrated probability assessmentCalibrated probability assessment
Over time, these drills cultivate a subtle but important shift in reasoning. Instead of asking only, “Am I right?”, the calibrated thinker repeatedly asks, “Was I as sure as the evidence warranted?” That question produces more realistic confidence, more informative predictions and a clearer understanding of where judgement genuinely deserves trust.
Amazon book picks
Further Reading
Books and field guides related to How sure should you really be?. Use these as the next step if you want deeper reading beyond the article.
Superforecasting
Directly explains probabilistic thinking, calibration and improving confidence judgments.
Thinking, Fast and Slow
Provides foundational understanding of overconfidence, judgment biases and decision-making.
The Scout Mindset
Strong fit for learning calibrated confidence, intellectual humility and evidence-based reasoning.
How to Measure Anything
Shows practical methods for expressing uncertainty and improving estimates.
Endnotes
1.
Source: onlinelibrary.wiley.com
Link:https://onlinelibrary.wiley.com/doi/abs/10.1002/ffo2.177
Source snippet
Wiley Online LibraryCalibration training for improving probabilistic judgments...by R Gruetzemacher · 2024 · Cited by 4 — We describe an...
2.
Source: Wikipedia
Title: Overconfidence effect
Link:https://en.wikipedia.org/wiki/Overconfidence_effect
3.
Source: Wikipedia
Title: Calibrated probability assessment
Link:https://en.wikipedia.org/wiki/Calibrated_probability_assessment
4.
Source: commoncog.com
Title: how do you evaluate your own predictions
Link:https://commoncog.com/how-do-you-evaluate-your-own-predictions/
Source snippet
How Do You Evaluate Your Own Predictions?Dec 17, 2019 — This post provides a comprehensive summary of the technique that Tetlock...
5.
Source: Wikipedia
Title: Brier score
Link:https://en.wikipedia.org/wiki/Brier_score
6.
Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC12818272/
Source snippet
On misconceptions about the Brier score in binary prediction...by L Hoessly · 2026 · Cited by 11 — Calibration refers to how well pre...
Additional References
7.
Source: gwern.net
Link:https://gwern.net/doc/statistics/prediction/2023-atanasov.pdf
Source snippet
Talent Spotting in Crowd PredictionCalibration and discrimination examine different facets of accuracy and can be obtained using Brier sc...
8.
Source: reddit.com
Link:https://www.reddit.com/r/deeplearning/comments/klmx70/understanding_brier_score_and_probability/
Source snippet
Understanding Brier Score and Probability CalibrationCalibrating your model is a crucial step to increase its prediction performance, esp...
9.
Source: cameronrwolfe.substack.com
Title: confidence calibration for deep networks why and how e2cd4fe4a086
Link:https://cameronrwolfe.substack.com/p/confidence-calibration-for-deep-networks-why-and-how-e2cd4fe4a086
Source snippet
Calibration for Deep Networks: Why and How?Confidence calibration is defined as the ability of some model to provide an accurate probabil...
10.
Source: researchgate.net
Link:https://www.researchgate.net/publication/376940026_Calibration_training_for_improving_probabilistic_judgments_using_an_interactive_app
Source snippet
ive app and a novel training process for improving calibration and reducing...
11.
Source: coefficientgiving.org
Title: calibration scores resulting from the training in their study.Read more
Link:https://coefficientgiving.org/research/efforts-to-improve-the-accuracy-of-our-judgments-and-forecasts/
Source snippet
Efforts to Improve the Accuracy of Our Judgments and...Oct 25, 2016 — The improvement in mean [Brier scores]({{ 'brier-scores/' | relative_url }}) from probability-training and...
12.
Source: medium.com
Link:https://medium.com/%40iamban/why-probabilities-lie-the-need-for-probability-calibration-in-ml-fad3c9801be4
Source snippet
how the part that's just calibration. But don't...
13.
Source: harry-cheslaw.medium.com
Title: super forecasting dd146e441c1c
Link:https://harry-cheslaw.medium.com/super-forecasting-dd146e441c1c
Source snippet
medium.comSuper-Forecasting. By Philip Tetlock and Dan GardnerCalibration-If you predict something with a 70% accuracy then it will happe...
14.
Source: arxiv.org
Link:https://arxiv.org/html/2305.03780v3
Source snippet
Boldness-Recalibration for Binary Event PredictionsJan 31, 2024 — The purpose of this paper is to develop boldness-recalibration that ena...
15.
Source: lesswrong.com
Link:https://www.lesswrong.com/posts/hJPwgwjWRSJ9DK4XN/in-forecasting-how-do-accuracy-calibration-and-reliability
Source snippet
In forecasting, how do accuracy, calibration and reliability...Sep 11, 2022 — In footnote 10, I believe they decompose Brier score into...
16.
Source: youtube.com
Title: Charting the Thicket: Using Argument Mapping to Explore Controversial Topics
Link:https://www.youtube.com/watch?v=f833pHMlJjk
Source snippet
Critical Thinking - 1.6 Argument Mapping...
Topic Tree

