Within Sharper Thinking
Is Your Confidence Matched to Evidence?
Recording predictions with confidence levels turns judgment into a feedback loop instead of a private feeling.
On this page
- Why calibration matters
- How prediction journals work
- Learning from misses
Page outline Jump by section
Introduction
Confidence calibration is the practice of checking whether your certainty matches reality. If you say ten things are 70% likely, roughly seven should happen over time. A prediction journal makes that test possible: before an outcome is known, you write down the claim, the probability you assign to it, the evidence behind it, and the date when it will be checked. That small habit turns judgement from a private feeling into a feedback loop.
This matters because memory is a poor scorekeeper. People often remember the broad direction of their past views, but not the exact confidence they felt, the alternatives they dismissed, or the conditions under which they said they would change their mind. Calibration does not require you to become cautious about everything. The aim is sharper confidence: low when the evidence is thin, higher when the evidence is strong, and revisable when new information arrives. Forecasting research, especially the Good Judgment Project, suggests that probabilistic prediction, feedback, practice and clear scoring can improve judgement in real-world questions, not just in classroom puzzles.[Good Judgment]goodjudgment.comGood Judgment The impact of training and practice on judgmental accuracyGood JudgmentThe impact of training and practice on judgmental accuracy…September 30, 2016 — by W Chang · 2016 · Cited by 153 — Althou…
Why calibration matters
A person can be knowledgeable and still miscalibrated. The issue is not simply whether a conclusion is right or wrong; it is whether the stated confidence was appropriate. A 55% judgement that turns out wrong may have been reasonable. A 95% judgement that turns out wrong deserves much closer review, because the error was not only in the prediction but in the strength of belief attached to it.
Calibration is easiest to understand through weather forecasts. If a forecaster says there is a 70% chance of rain on many similar days, rain should occur on about 70% of those days. The same idea applies to personal and professional judgement: “I am 80% sure this project will ship by Friday”, “There is a 30% chance this client renews”, or “I am 65% confident this news report will be corrected within a week”. The Brier score, originally developed for probability forecasts, measures the gap between predicted probabilities and actual outcomes; lower scores mean better probabilistic accuracy.[Wikipedia]WikipediaBrier scoreBrier score
The practical value is that calibration separates three things people often blur together:
- Accuracy: did the thing happen?
- Confidence: how strongly did you expect it?
- Learning signal: what should change in your future judgement?
Without calibration, a person can protect almost any self-image. A failed 80% prediction becomes “I only said it was likely, not certain”. A successful 55% prediction becomes “I knew it”. Written probabilities make that escape harder. They also make improvement less moralistic. Instead of asking “Am I good at judgement?”, you ask “Where am I overconfident, underconfident, vague, or slow to update?”
Research on overconfidence is more nuanced than the slogan “people are always overconfident”. Reviews of probability calibration show that miscalibration varies by task difficulty, knowledge and evidence quality. People may be overconfident on hard questions yet underconfident on easier ones, which means the useful intervention is not blanket humility but better matching between evidence and certainty.[Warrington College of Business]bear.warrington.ufl.eduOpen source on ufl.edu.
How prediction journals work
A prediction journal is a simple record of claims before reality has had a chance to make them look obvious. It can be a spreadsheet, notebook, database, forecasting platform or shared team log. The format matters less than the discipline: the prediction must be specific enough to resolve, probabilistic enough to reveal confidence, and reviewed often enough to create feedback.
A useful entry normally includes five parts:
- A clear question. “Will the supplier deliver the parts by 5 pm on 30 June 2026?” is better than “Will the supplier be reliable?”
- A probability. Use numbers, not vague words such as “probably” or “unlikely”. A 60% forecast and an 85% forecast should not be treated as the same belief.
- The evidence and assumptions. Record the base rate, current signals, reasons for doubt, and any key dependency.
- A resolution rule. State what will count as yes, no, partial, or unresolved.
- A review date. Decide when the prediction will be scored or revisited.
The Good Judgment Project is the strongest modern demonstration of this style of judgement training at scale. In a multi-year geopolitical forecasting tournament, forecasters made probabilistic predictions on real-world questions, updated them as new information arrived, and were scored using measures such as the Brier score. Research from the project found that brief probability training improved forecasting accuracy, with one study reporting Brier-score improvements of 6% to 11% over a control condition after training lasting less than an hour.[Good Judgment]goodjudgment.comGood Judgment The impact of training and practice on judgmental accuracyGood JudgmentThe impact of training and practice on judgmental accuracy…September 30, 2016 — by W Chang · 2016 · Cited by 153 — Althou…
The point is not that everyone needs a formal tournament. The transferable lesson is that judgement improves when forecasts are explicit, repeated, scored, and compared against outcomes. Good Judgment describes its evidence-based process as combining talent spotting, training, teamwork and aggregation; for an individual or small team, the most accessible parts are training, practice, and careful review.[Good Judgment]goodjudgment.comOpen source on goodjudgment.com.
A prediction journal also changes the emotional texture of thinking. When a belief stays private, being wrong can feel like a threat. When predictions are logged routinely, wrong forecasts become data. A journal makes it normal to say, “I gave this 75%, it failed, and here is what I missed.” That is a better learning environment than one in which only confident-sounding conclusions are rewarded.
What to predict
The best journal questions sit in the middle zone: uncertain enough to teach you something, but concrete enough to resolve. Very obvious predictions teach little. Extremely vague predictions cannot be scored. The useful targets are decisions and recurring judgement clusters where better calibration would change behaviour.
Good candidates include:
- Project delivery: whether a task will finish by a date, exceed budget, or need rework.
- Hiring and management: whether a candidate will accept, whether a team member will hit a milestone, or whether a meeting will produce a decision.
- Research and analysis: whether a source will be confirmed, whether a claim will survive checking, or whether an initial explanation will remain the best one.
- Personal planning: whether a habit will be maintained, whether a trip will stay within budget, or whether a deadline estimate is realistic.
- External events: public questions in politics, economics, technology, sport or policy, where outcome criteria are available.
For improving analytical skill, repeated small predictions are usually more valuable than rare dramatic ones. A person who makes two major forecasts a year gets little feedback. A person who makes ten small forecasts a week can begin to see patterns: perhaps they are too optimistic about delivery dates, too deferential to confident colleagues, too slow to update after contrary evidence, or too inclined to put everything in the safe 55% to 65% range.
Forecasting platforms such as Metaculus and Good Judgment Open show how this can work in public settings: questions are stated in advance, probabilities can be updated, and track records can be compared over time. Their broader value for a private prediction journal is cultural as much as technical: they normalise the idea that good judgement is not a single impressive call, but a long-run record of probabilistic claims.[metaculus.com]metaculus.comcomparing forecasting track records for ai benchmarking and beyondcomparing forecasting track records for ai benchmarking and beyond
Learning from misses
The review is where a prediction journal becomes a thinking intervention rather than a diary. It is tempting to look only at wrong predictions, but calibration requires a wider view. You need to ask whether your 60% predictions happen about 60% of the time, whether your 80% predictions happen about 80% of the time, and whether your very high confidence forecasts are rare and justified.
A useful review separates different kinds of failure:
The evidence was weak. You may have relied on a vivid anecdote, a single source, or a recent example instead of a base rate.
The question was badly framed. If the resolution rule was unclear, the forecast cannot teach much. Ambiguous predictions are often a sign that the original thinking was also ambiguous.
The probability was too extreme. A 90% forecast means the outcome should fail only about once in ten comparable cases. Many people use high numbers to express emphasis rather than measured probability.
The update was too slow. Forecasting skill is not only about the first estimate. In Good Judgment-style forecasting, updating in response to new evidence is part of the discipline; strong forecasters often revise as facts change rather than defending the first number.[Commoncog]commoncog.comhow do you evaluate your own predictionshow do you evaluate your own predictions
The miss was reasonable. Some low-probability events happen. A good review does not punish every wrong forecast; it asks whether the probability was fair at the time.
This last point is important. Calibration is not hindsight perfection. If you predicted a 20% chance of an event and it happened, that does not automatically mean your forecast was bad. Low-probability events should occur sometimes. The question is whether, across many similar calls, your 20% bucket behaves like a 20% bucket.
Scoring without overcomplicating it
A full scoring system is useful, but it should not become a barrier to practice. For most people, the first step is simply to group forecasts by confidence band and compare them with outcomes.
For example:
Confidence bandNumber of predictionsNumber that happenedWhat to check50–60%2012Close to expected; look for vague hedging70–80%2010Possible overconfidence90–100%106Serious overconfidence unless sample is unusual
The Brier score adds a stricter numerical penalty: confident wrong predictions hurt more than cautious wrong predictions. That is exactly why it is useful. Saying “95%” should carry more accountability than saying “60%”. At the same time, researchers note that Brier scores can be decomposed into different properties, including calibration and resolution. Resolution matters because a person who always says 50% may be well protected from embarrassment but is not adding much decision value.[Cambridge University Press & Assessment]cambridge.orgUniversity Press & Assessment Weighted Brier score decompositions for topicallyUniversity Press & Assessment Weighted Brier score decompositions for topically
For everyday use, the scoring rule can be modest:
- Review monthly, not after every single outcome.
- Track confidence bands before calculating complex scores.
- Keep original predictions visible after updates.
- Note whether misses came from bad information, bad framing, poor base rates, or emotional commitment.
- Watch for both overconfidence and underconfidence.
This approach keeps the focus on better judgement rather than score-chasing. A prediction journal should make decisions clearer, not turn thinking into a game where people avoid useful forecasts because they fear damaging their record.
Using calibration in teams
Prediction journals are especially powerful in teams because they reduce the social distortions around confidence. In many organisations, the most fluent or senior person can sound “right” before evidence has been tested. A forecasting habit asks everyone to put a number on the claim and record the reason. That makes disagreement more productive: two people can both favour the same outcome but differ sharply between 55% and 85%, which reveals different assumptions.
For a team, the intervention can be simple:
- Identify a recurring decision cluster, such as project deadlines, sales renewals, policy risks or product launches.
- Require a small number of written probabilistic forecasts before major decisions.
- Record assumptions and resolution rules.
- Review outcomes in batches.
- Adjust decision rules, not just individual opinions.
The batch review is crucial. If a product team repeatedly gives 80% confidence to delivery dates that succeed only half the time, the lesson is not merely “be less confident”. The team may need to change planning buffers, dependency checks, escalation triggers or how it treats optimistic estimates. Calibration turns judgement errors into process evidence.
Structured forecasting also helps distinguish confidence from authority. A junior analyst with a strong track record on a specific class of questions may deserve more weight than a senior person who speaks firmly but has not been scored. Good Judgment’s work on forecasting tournaments reflects this broader idea: track records, training and aggregation can reveal signal that status alone may hide.[Good Judgment]goodjudgment.comOpen source on goodjudgment.com.
Common traps
The first trap is writing predictions that cannot lose. “The launch may face challenges” is not a prediction journal entry; it is a fog machine. A better version is: “There is a 70% chance the launch is delayed by more than five working days, using the currently announced date as the baseline.”
The second trap is treating calibration as pessimism. Good calibration can make someone more confident when the evidence supports it. The goal is not to lower every probability, but to stop using confidence as a mood, identity signal or negotiation tactic.
The third trap is reviewing only spectacular errors. Most calibration problems are mundane and repeated: deadline optimism, exaggerated certainty from small samples, failure to update after new evidence, or treating a preferred outcome as more likely than it is.
The fourth trap is ignoring sample size. Ten predictions are enough to start a habit, not enough to diagnose your whole mind. Calibration curves and confidence bands become more informative as the number of forecasts grows. Older calibration research has long emphasised that confidence and accuracy need to be compared across many judgements rather than inferred from isolated examples.[California State University Long Beach]home.csulb.eduOpen source on csulb.edu.
The fifth trap is using scoring in a punitive way. If a manager uses prediction records mainly to shame people, forecasts will become timid, political or vague. The healthier norm is accountability without humiliation: strong claims should be tested, but honest uncertainty should be protected.
A simple starting routine
A practical routine can be small enough to use immediately. At the end of each working day, write one to three predictions about decisions already on your mind. Use percentages in increments such as 5% or 10%. Add a one-sentence reason and a review date. Once a week, update any forecasts where new evidence has appeared. Once a month, review resolved predictions by confidence band.
The most valuable prompt is often: “What would I expect to see if I were wrong?” That question pushes the journal beyond betting and into analysis. It asks you to name the evidence that would weaken your current view before the world has delivered it.
Over time, the journal should reveal a personal calibration map. You may find that you are well calibrated about technical estimates but overconfident about people’s availability. You may be accurate about short-term deadlines but poor at three-month planning. You may be too cautious in public forecasts and too bold in private assumptions. These patterns are exactly the point. They show where better thinking has to be implemented, not merely admired.
The real payoff
Confidence calibration is valuable because it changes the unit of improvement. Instead of trying to become “a better thinker” in the abstract, you build a record of specific judgements, confidence levels, outcomes and lessons. Prediction journals make thinking visible enough to correct.
The deeper benefit is intellectual honesty under uncertainty. Calibrated thinkers can still be bold, but their boldness is earned. They know the difference between a strong signal and a strong feeling. They can say “I am 60% confident” without sounding weak, and “I was 90% confident and wrong” without pretending the mistake never happened. For improving analytical skill, that shift is hard to beat: it makes confidence answerable to evidence.
Amazon book picks
Further Reading
Books and field guides related to Is Your Confidence Matched to Evidence?. Use these as the next step if you want deeper reading beyond the article.
Thinking, Fast and Slow
Explains cognitive biases and overconfidence underlying poor calibration.
The Signal and the Noise
Focuses on probabilistic prediction, uncertainty and learning from forecasts.
How to Measure Anything
Shows practical methods for quantifying uncertainty and evidence.
The Scout Mindset
Encourages updating beliefs in line with evidence and better judgment.
Endnotes
1.
Source: Wikipedia
Title: Brier score
Link:https://en.wikipedia.org/wiki/Brier_score
2.
Source: metaculus.com
Title: comparing forecasting track records for ai benchmarking and beyond
Link:https://www.metaculus.com/notebooks/28552/comparing-forecasting-track-records-for-ai-benchmarking-and-beyond/
3.
Source: commoncog.com
Title: how do you evaluate your own predictions
Link:https://commoncog.com/how-do-you-evaluate-your-own-predictions/
4.
Source: cambridge.org
Title: University Press & Assessment Weighted Brier score decompositions for topically
Link:https://www.cambridge.org/core/services/aop-cambridge-core/content/view/8172E04F2DBC601DA5D953D4685CA346/S1930297500007099a.pdf/weighted_brier_score_decompositions_for_topically_heterogenous_forecasting_tournaments.pdf
5.
Source: metaculus.com
Title: exploring metaculuss ai track record
Link:https://www.metaculus.com/notebooks/16708/exploring-metaculuss-ai-track-record/
6.
Source: metaculus.com
Title: why i reject the comparison of metaculus to prediction markets
Link:https://www.metaculus.com/notebooks/17599/why-i-reject-the-comparison-of-metaculus-to-prediction-markets/
7.
Source: cambridge.org
Link:https://www.cambridge.org/core/journals/judgment-and-decision-making/article/weighted-brier-score-decompositions-for-topically-heterogenous-forecasting-tournaments/8172E04F2DBC601DA5D953D4685CA346
8.
Source: dictionary.cambridge.org
Link:https://dictionary.cambridge.org/us/dictionary/english/confidence
9.
Source: cambridge.org
Link:https://www.cambridge.org/core/books/judgment-under-uncertainty/calibration-of-probabilities-the-state-of-the-art-to-1980/9F0C9EC2997AEEB6DDDB304C2F935A16
10.
Source: Wikipedia
Link:https://en.wikipedia.org/wiki/Confidence
11.
Source: Wikipedia
Title: The Good Judgment Project
Link:https://en.wikipedia.org/wiki/The_Good_Judgment_Project
12.
Source: goodjudgment.com
Title: Good Judgment The impact of training and practice on judgmental accuracy
Link:https://goodjudgment.com/wp-content/uploads/2018/12/jdm16511.pdf
Source snippet
Good JudgmentThe impact of training and practice on judgmental accuracy...September 30, 2016 — by W Chang · 2016 · Cited by 153 — Althou...
Published: September 30, 2016
13.
Source: bear.warrington.ufl.edu
Link:https://bear.warrington.ufl.edu/brenner/mar7588/Papers/koehlerbrennergriffin2002.pdf
14.
Source: goodjudgment.com
Link:https://goodjudgment.com/about/the-science-of-superforecasting/
15.
Source: goodjudgment.com
Link:https://goodjudgment.com/
16.
Source: home.csulb.edu
Link:https://home.csulb.edu/~cwallis/382/certainty/chapter19.html
17.
Source: goodjudgment.com
Link:https://goodjudgment.com/about/
18.
Source: goodjudgment.com
Title: Goldstein et al GJP vs ICPM
Link:https://goodjudgment.com/wp-content/uploads/2020/11/Goldstein-et-al-GJP-vs-ICPM.pdf
19.
Source: goodjudgment.com
Link:https://goodjudgment.com/resources/case-studies/
20.
Source: psychologytoday.com
Link:https://www.psychologytoday.com/us/basics/confidence
21.
Source: jclinepi.com
Link:https://www.jclinepi.com/article/S0895-4356%2809%2900363-1/pdf
22.
Source: home.csulb.edu
Title: Training to Improve Calibration
Link:https://home.csulb.edu/~cwallis/382/certainty/overconfidence/Training%20to%20Improve%20Calibration.pdf
23.
Source: pure.mpg.de
Link:https://pure.mpg.de/rest/items/item_2220390_1/component/file_2220389/content
24.
Source: goodjudgment.substack.com
Link:https://goodjudgment.substack.com/p/a-primer-on-good-judgment-inc-and
25.
Source: stat.berkeley.edu
Link:https://www.stat.berkeley.edu/~ryantibs/statlearn-s23/lectures/calibration.pdf
26.
Source: youtube.com
Link:https://www.youtube.com/watch?v=pedNak4S9IE
27.
Source: alexandria.unisg.ch
Link:https://www.alexandria.unisg.ch/server/api/core/bitstreams/50fe560a-bf6b-47dd-827c-6432801ea15a/content
Additional References
28.
Source: youtube.com
Title: LLM confidence calibration. Confidence Gap in high stakes decision making
Link:https://www.youtube.com/watch?v=7trfF0BV3xo
Source snippet
The Brier Score That Proves PropsBot.AI Beats Vegas — AI Sports Betting Tutorial #6...
29.
Source: youtube.com
Title: ‘Superforecasting’: The people that [predict]({{ ‘predict/’ | relative_url }}) the future – BBC REEL
Link:https://www.youtube.com/watch?v=SAzTP2A634g
Source snippet
Model Calibration - Brier Score Explained...
30.
Source: stanford.edu
Link:https://stanford.edu/~knutson/jdm/mellers15.pdf
Source snippet
Stanford UniversityIdentifying and Cultivating Superforecasters as a Method of...by B Mellers · 2015 · Cited by 332 — Accurate probabili...
31.
Source: youtube.com
Title: Model Calibration
Link:https://www.youtube.com/watch?v=BiaebXlgfNQ
Source snippet
LLM confidence calibration. Confidence Gap in high stakes decision making...
32.
Source: researchgate.net
Link:https://www.researchgate.net/publication/320911494_Confidence_Calibration_in_a_Multiyear_Geopolitical_Forecasting_Competition
33.
Source: stata.com
Link:https://www.stata.com/manuals15/rbrier.pdf
34.
Source: merriam-webster.com
Link:https://www.merriam-webster.com/dictionary/calibration
35.
Source: merriam-webster.com
Link:https://www.merriam-webster.com/dictionary/confidence
36.
Source: openreview.net
Link:https://openreview.net/forum?id=6DDaTwTvdE
37.
Source: aiimpacts.org
Link:https://aiimpacts.org/evidence-on-good-forecasting-practices-from-the-good-judgment-project/
Topic Tree
Follow this branch
Parent topic
Sharper ThinkingRelated pages 29
- Brier Scores A Simple Scorecard for Forecasting Skill
- Good Judgment What Superforecasting Teaches Everyday Thinkers
- Miss Reviews When a Wrong Forecast Was Still Reasonable
- Probability Words Why Probably Is Not Precise Enough
- Scorable Questions Can Your Forecast Actually Be Scored?
- +1 more in sidebar



