CodeNSM
The Standup · Part 22

Two of two is not one hundred percent

2026-06-26· 7 min read· by Think North

Here is a sentence that will show up, unedited, in a sprint retro slide somewhere near you: "Priya's calibration this sprint: 100%."

It sounds like an award. It is actually a confession, and the confession is: Priya made exactly two checkable claims this sprint, and both of them happened to come true. Two out of two. Flip a coin twice and get two heads and you have not discovered a two-headed coin — you have discovered that flipping a coin twice is not enough flips to discover anything.

The percentage is doing something to your brain

Round the fraction two-out-of-two into a percentage and something happens that wouldn't happen if you'd just written "2/2." A percentage looks like it belongs on the same axis as every other percentage you've ever seen — a 100% uptime SLA, a 100% test pass rate, a 100% approval rating — and your brain, quite reasonably, files it in the same drawer: flawless, reliable, done. But a 100% uptime SLA over a year is backed by roughly thirty-one million seconds of observation. A "100%" calibration score from a standup can be backed by two sentences. Feeding both numbers through the same word invites you to trust them the same amount, and that invitation is the entire mechanism of the mistake.

n=2 n=10 n=40 how much a "100%" claim actually tells you, by sample size

What a mathematician would actually tell you

This isn't a new problem, and it has a startlingly old fix. In his Essai philosophique sur les probabilités, Pierre-Simon Laplace worked out what to do when you've observed a small number of trials and want an honest estimate of the underlying rate — the same question, dressed differently, as "the sun has risen every day of recorded history; what's the probability it rises tomorrow?" His answer, now called the rule of succession, says: don't report k successes out of n trials as the raw rate k/n. Report (k+1)/(n+2) instead — effectively padding the count with one imagined success and one imagined failure, to keep a short early streak from being mistaken for certainty. Run Priya's two-for-two through Laplace's rule and "100%" becomes 3/4, or 75%. Still good! Just no longer indistinguishable from perfection.

The historian of probability Sandy Zabell traced the rule's genealogy — from Bayes and Price through Laplace to the twentieth-century philosophers of induction who kept rediscovering how badly it's needed — in a 1989 paper in Erkenntnis, and the throughline across two and a half centuries of that literature is the same one Priya's slide violates: a short streak is real evidence, but it is weak evidence, and the correct way to report weak evidence is not to round it up into a number that looks exactly like strong evidence.

Why the mistake feels so natural

There's a second reason "2 of 2" reads as "100%" instead of "barely any data," and it isn't just the percentage sign — it's a well-documented failure mode called the base-rate fallacy. Maya Bar-Hillel's classic 1980 paper in Acta Psychologica showed, across a range of judgment tasks, that people reliably underweight general statistical information — how rare or common something is on average — in favor of specific, vivid, individuating information about the case right in front of them. Two specific, concrete, nameable instances — this claim, about this function, came true; that claim, about that risk, also came true — feel like overwhelming evidence precisely because they're vivid and personal. The boring background fact that two data points barely constrain anything statistically doesn't compete, emotionally, with two stories you can actually remember telling.

A calibration score without its denominator isn't a lie. It's a number wearing a costume borrowed from a much larger number, and nobody in the room checked underneath.

Small samples don't just mislead — they mislead in a specific direction

There's a third piece of research worth adding here, because it explains something the base-rate fallacy alone doesn't quite cover: why small samples don't just produce noisy conclusions, they produce overconfident ones. Amos Tversky and Daniel Kahneman named this in a 1971 Psychological Bulletin paper as the belief in the law of small numbers — a tendency, which they found even among trained research scientists, to treat a small sample as though it obeys the same statistical regularities as a large one, when in fact small samples are exactly where those regularities are weakest. Their subjects, mathematically sophisticated psychologists among them, consistently overestimated how much a short run of results could tell you and underestimated how much a short run of results could be pure noise. Two out of two isn't just "not enough data" in some vague sense — it's data drawn from precisely the range where intuition about what data means is at its least reliable, for everyone, including people whose job is statistics.

Put this together with Bar-Hillel's base-rate finding and you get the full mechanism, not just half of it: a short streak feels vivid and personal (base-rate neglect), and separately, the streak itself gets misread as more statistically meaningful than it is (belief in the law of small numbers). Neither error alone would be enough to turn "2 of 2" into a boardroom-ready percentage. Together, they make it feel not just acceptable but obvious — which is exactly why it keeps happening, quietly, on slide after slide, without anyone involved feeling like they've done anything wrong.

The rule this has to become

This is exactly why Cadence's design refuses to ever print a calibration percentage without the count it came from. Not "100%." Not even "75% (Laplace-adjusted)" — that's a magic trick with a fancier hat. The honest sentence is "2 of 2 resolved claims came true," reported plainly, every time, no matter how tempting the rounder number sitting one keystroke away looks. A person with two resolved claims has a calibration of two claims. A person with forty has a calibration of forty claims, and forty is a number you can actually start to trust the way you trust an SLA — because by forty, Laplace's correction and the raw percentage have converged close enough that the distinction stops mattering. That crossover point is exactly what a bare percentage hides and a stated denominator reveals.

The practical test is almost insultingly simple, and you can run it on any dashboard in your company right now, engineering or otherwise: does the number ever appear without the count of things it was computed from? If the answer is no, you have a real metric. If the answer is yes, you have a percentage doing a magic trick, and the trick works precisely because "100%" and "2 of 2" describe the identical mathematical fact while producing completely different feelings in the person reading them.

None of this means small samples are worthless — Priya's two correct calls are real information, and dismissing them entirely would be its own mistake, the mirror image of over-trusting them. The fix isn't suspicion of small numbers. It's honesty about what they are: an early, genuinely encouraging first read that needs more resolved claims before anyone, including Priya, should trust it the way they'd trust a number built on forty.

It's worth noticing where this leaves you, practically, at the start of a new project or with someone new to the team: everyone's calibration starts at zero denominator, and stays uninformative for a while, and that's not a flaw to engineer around — it's the correct, honest state of not-yet-knowing. The temptation, especially with a new hire eager to prove themselves, is to let the first few resolved claims stand in for a verdict on their judgment. Two of two from someone new is exactly as uninformative as two of two from someone who's been on the team for years; the only thing that fixes it, for either person, is more resolved claims, patiently accumulated, with the denominator printed next to every single one of them until it's actually large enough to mean something.

That patience is the whole ask, and it runs against a very human instinct to want a verdict now, this sprint, about who's reliable. Laplace didn't have that luxury either — his rule of succession exists precisely because he was trying to reason honestly about the sun rising with a sample size of "every day anyone has ever recorded," which sounds enormous until you remember it's still just one long, unbroken streak, mathematically closer to Priya's two claims than most people are comfortable admitting.

References

  1. Zabell, S.L. (1989). The Rule of Succession. Erkenntnis, 31.
  2. Bar-Hillel, M. (1980). The base-rate fallacy in probability judgments. Acta Psychologica, 44(3).
  3. Tversky, A. & Kahneman, D. (1971). Belief in the law of small numbers. Psychological Bulletin, 76(2).

See your own codebase as an office.

One pip install and every function reports for duty — archetype, live state, debt tier, and a single Code-Health North-Star. Free plan, no card.

Read next