Calibration, not confidence: the only score that has ever mattered
Two weather forecasters, same city, same decade. Forecaster A says "70% chance of rain" nearly every day she feels unsure, and over ten years, on the days she says 70%, it rains about seven days in ten. Forecaster B never says a number he doesn't believe with his whole chest — "it's DEFINITELY going to rain," "I am one hundred percent sure this front stalls" — and he is memorably, gloriously confident, roughly half the time.
You are planning an outdoor wedding. Which forecaster do you actually want?
Everyone picks A, immediately, and then goes right back to running their engineering org on Forecaster B. Because in a standup, in a sprint review, in a hallway conversation about whether the retry logic is the real problem, the person who says "I am ONE HUNDRED PERCENT SURE" gets treated like the person who knows, and the person who says "probably, maybe seventy percent, I'd want to check the fault rate before Thursday" gets treated like the person who's still figuring it out. We have the wedding-planning instinct exactly backwards at work, every single day, and it is costing us more than a rained-out reception.
The forecaster's oldest problem, solved in three pages
In 1950, a meteorologist named Glenn W. Brier published a short paper in Monthly Weather Review with a genuinely unglamorous title — "Verification of Forecasts Expressed in Terms of Probability" — that quietly solved a problem the entire forecasting profession had been dodging for decades: how do you score a probability?
You can't just check whether it rained. A forecaster who says "70% chance of rain" and gets a dry day hasn't necessarily been wrong — 30% isn't 0%. What Brier built instead is now called the Brier score: take the forecast probability, subtract the actual outcome (1 if it happened, 0 if it didn't), square the difference, and average that over every forecast the person ever made. It rewards forecasters who say 90% when they're nearly certain and punishes forecasters who say 90% when they're actually just enthusiastic, because the squared-error penalty for a confident miss is brutal, and the only way to keep your long-run score low is to say what you actually believe, as precisely as you actually believe it. Statisticians call this a proper scoring rule: the honest answer is always the best-scoring answer. No amount of vocal commitment beats it.
Run the arithmetic once, because it's genuinely this simple. Forecaster A says 70% and it rains: her squared error is (1 − 0.7)² = 0.09. She says 70% and it doesn't rain: (0 − 0.7)² = 0.49. Average enough of those over a career and a well-calibrated forecaster lands comfortably below 0.25, the score pure chance would produce. Forecaster B, saying 99% and being wrong about half the time, racks up (0 − 0.99)² ≈ 0.98 on every miss — one confident wrong call costs him roughly as much as ten of Forecaster A's honest, hedged misses combined. The math doesn't punish uncertainty. It punishes uncertainty dressed up as certainty, and it does so in proportion to how loudly the certainty was dressed.
The genius of the Brier score isn't the arithmetic. It's what it makes irrelevant. It doesn't care how the forecast was delivered, how senior the forecaster is, or how many people nodded along in the room. It cares about exactly one thing: across enough forecasts, did the stated confidence match the observed frequency. That property has a name — calibration — and it turns out to be almost entirely unrelated to how confident someone sounds while forecasting.
The tournament that proved it at scale
For decades this stayed a meteorologist's tool. Then, starting in 2011, the intelligence community's research arm ran an actual controlled experiment on human judgment at a scale nobody had attempted: the Good Judgment Project, led by Philip Tetlock and Barbara Mellers, recruited thousands of ordinary volunteers to forecast real geopolitical events — will this ceasefire hold, will that election happen on schedule — for years, scored every single forecast with a Brier score, and pitted the results against career intelligence analysts with classified access.
The volunteers won. Not the loudest volunteers, and not the ones with the most impressive résumés walking in — the ones whose Brier scores, across hundreds of forecasts, were consistently low. Mellers, Tetlock and their coauthors published the mechanism behind the winners in a 2014 Psychological Science paper: the best forecasters weren't smarter in some general sense. They were better at three specific habits — training themselves out of known cognitive biases, working in small teams that could challenge a bad take before it calcified, and relentlessly tracking their own prior forecasts against what actually happened, so the next forecast started from an updated position instead of a repeated mistake. Tetlock's later name for the population that consistently produced these scores — superforecasters — has nothing to do with charisma. It's a Brier-score percentile.
Confidence is a performance. Calibration is a track record. One of them you can fake in a single sentence. The other takes hundreds of checked claims to build, and one honest look at the ledger to verify.
What calibration is not
It would be easy to walk away from this thinking the fix is simple: hedge everything, say "maybe" a lot, and let the math sort out the rest. It doesn't work that way, and the Brier score is specifically built to catch it. A forecaster who says "50% chance of rain" every single day, forever, regardless of the actual weather, will eventually post a mediocre but stable score — better than the confidently wrong Forecaster B, considerably worse than anyone actually paying attention to the sky. Hedging isn't calibration. It's a quieter way of refusing to say anything checkable, and a proper scoring rule punishes it almost as reliably as it punishes overconfidence, just less dramatically. The forecasters who actually won the Good Judgment Project's tournament weren't the hedgers. They moved their probabilities — 15% one week, 40% the next, 8% the week after that — as fast as the evidence did, which is precisely what a person who's genuinely updating on the world looks like from the outside. Calibration rewards being specific and right over time, not vague and technically unfalsifiable.
This distinction matters the first time a team encounters a claim ledger, because the instinctive defensive move — and every group eventually tries it — is to retreat into hedges precisely calibrated to be impossible to score. "It might be the retry logic, hard to say, could be a few things." That sentence sounds humble. It is actually the standup equivalent of Forecaster A quietly becoming a third forecaster who never says anything specific enough to be wrong, and therefore never says anything specific enough to be useful either. An unfalsifiable hedge has to be treated the same way a claim with no nameable subject is treated — as unresolvable, not as a clever dodge — because the alternative is training an entire team to speak exclusively in fog, which produces a spotless, meaningless calibration score for everyone.
Now put a standup through the same filter
A standup is a room full of uncalibrated Forecaster Bs, and it's nobody's fault, because nobody ever built the scoreboard. "I'm sure the timeout is in the retry logic." "That refactor is definitely low risk." "We'll have this wrapped by Thursday, no problem." Every one of these is a probabilistic forecast wearing the costume of a fact, delivered with a confidence level chosen almost entirely for its social effect — the flat declarative sentence that ends a conversation reads as more competent than the hedged one that invites scrutiny — and then the meeting moves on, because nothing in the room's design asks anyone to write the number down and check it later.
This is precisely the gap this arc has been circling for thirty parts: every one of those sentences is a claim with a truth value that reality will eventually supply, whether or not anyone bothers to look. A commitment resolves against the commit log. A causal claim about "why" resolves against the telemetry. And once enough of a person's claims have resolved — not one, not two, but a real number, always shown with its denominator — you get something a standup has never had before: an actual calibration score, built on the same premise Brier proved in 1950 and the Good Judgment Project proved at scale six decades later. Being right, tracked over time, is a completely different quantity from sounding right in the room, and the two have never been more confusable than they are in a meeting run entirely on memory.
What this changes, once you can see it
The uncomfortable part isn't that confident people are usually wrong. Plenty of confident people are excellent forecasters, and plenty of hedgers are just as unreliable as the loud ones — hedging is not calibration, it's a different kind of performance, and it can be just as uncorrelated with the truth. The uncomfortable part is that you currently cannot tell the difference, in either direction, because nobody in the room is keeping the one number that would show you. You are running Forecaster A and Forecaster B side by side, forever, with no scoreboard, and defaulting — because it's the only signal available — to whoever says it loudest.
Cadence's calibration score is exactly this arithmetic, run on the sentences a standup already produces: the share of a person's resolved claims that came true, weighted by how much of production the claim was actually about, never by how it was delivered. Brier didn't invent honesty. He invented a way to make honesty the only strategy that pays off over a long enough run. That's the whole trick, and it's exactly the trick a standup has been missing since the format was invented.
References
- Brier, G.W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1).
- Mellers, B., Ungar, L., Baron, J. et al. & Tetlock, P.E. (2014). Psychological Strategies for Winning a Geopolitical Forecasting Tournament. Psychological Science, 25(5).
- The Good Judgment Project — Philip Tetlock & Barbara Mellers.