CodeNSM
The Standup · Part 26

Unresolvable is a finding: why 'we don't know yet' has to stay on the scoreboard

2026-07-02· 7 min read· by Think North

Somebody says, in passing, "that refactor's low risk, nothing calls it anymore." Nobody checks. The sprint ends. Nothing visibly breaks. What happens to that claim in the team's collective memory?

Usually one of two things, and both of them are wrong. Either it quietly gets counted as correct — nothing broke, so the claim must have been right — or, if something unrelated goes sideways two weeks later, it gets dragged in retroactively as evidence the claim was reckless. Neither of those is what actually happened. What actually happened is that nobody checked, and "nobody checked" is not the same finding as either "true" or "false." It's its own finding, and the moment you silently convert it into one of the other two, you've manufactured a result nobody earned.

A distinction medicine had to learn the hard way

Clinical research ran headfirst into this exact confusion for decades, and the correction, when it finally got written down plainly, became one of the most quoted lines in medical statistics. Douglas Altman and J. Martin Bland's 1995 BMJ "Statistics Notes" entry — titled, with maximum bluntness, "Absence of evidence is not evidence of absence" — pointed out that trials failing to find a statistically significant difference were routinely reported as "negative," a word that quietly implies the treatments were shown to be equivalent. Usually, they argued, nothing of the sort had been shown. The trial simply lacked the statistical power to detect a real difference if one existed, which is a completely different, much less informative outcome than "we looked carefully and found nothing there." Reporting the second as though it were the first has misled treatment decisions for years at a time.

The mechanism behind their argument is worth spelling out, because it maps onto a standup almost exactly. A trial's ability to detect a real effect — its statistical power — depends on how much data it collected and how big the effect actually is. A small trial can easily fail to find a real, important difference not because the difference isn't there, but because the trial never had the sensitivity to see it. Altman and Bland's point was that "no significant difference" gets reported in the passive voice, as a property of the world, when it's really a property of the study's ability to look. A standup claim has an equivalent limitation: a horizon of one sprint, on a function nobody happened to call much that week, is a low-power test almost by construction — passing quietly through the horizon proves considerably less than the room tends to assume it does.

The failure mode isn't unique to medicine, and it's not new even there. Robert Rosenthal named the publication-side version of it in a 1979 Psychological Bulletin paper as the file-drawer problem: studies that find "nothing" — no significant effect either way — are far less likely to get written up and published than studies that find something, which means the visible research record is systematically missing its inconclusive results, not because they didn't happen, but because inconclusive is a much less compelling thing to publish than a verdict. The scoreboard everyone can see ends up quietly purged of exactly the entries that say "we don't actually know."

The standup version of the file drawer

A standup has its own file drawer, and it's the same drawer every unrecorded, unchecked, half-remembered claim from the last sprint falls into. A claim with no nameable subject — "we should really look into that at some point" — was never checkable to begin with. A claim with a real subject but a horizon that quietly passed with nobody following up is checkable in principle and unchecked in practice. Both get discarded from the team's working memory the same way a null result gets discarded from a researcher's file drawer: silently, and without anyone deciding to do it on purpose.

This is exactly why unresolvable has to be treated as a third, permanent, first-class outcome — never collapsed into did_not, never quietly dropped from the count. A claim that named nothing checkable, or whose horizon passed with no evidence either way, is reported as exactly what it is: something the team said and never actually found out about. In CodeNSM's own early usage, that bucket is routinely large — often a meaningful share of everything said out loud in a given standup — not because teams are being careless, but because most of what gets said in a meeting was never built to be checkable in the first place. That number, reported honestly, is not a scold. It's one of the most useful sentences a team can be handed about its own meeting.

Silently converting "we never checked" into "it was wrong" doesn't make the record more honest. It makes it more confident, which is a worse kind of dishonest, because it looks like rigor.

Why the temptation to fold it away is so strong

Unresolvable claims are unsatisfying in a specific way that pulls people toward mislabeling them. A ledger with three clean outcomes — right, wrong, unknown — is honest but visually messy; a ledger with only two — right or wrong — resolves cleanly into a single percentage that's much easier to put on a slide. That aesthetic pressure is exactly the mechanism Rosenthal described in a different field: the tidy, resolved result crowds out the messy, unresolved one, not through anyone's dishonesty, but through everyone's mild preference for a cleaner story.

Resisting it is unglamorous and non-negotiable. If a punished "did_not" and an ignored "unresolvable" produce the same social consequence — a black mark against whoever spoke — the entire population of claims made in a room will drift, quietly and rationally, toward vagueness, because vague claims are the only ones that can never be caught being wrong. That is the exact opposite of what a scoring system is supposed to produce, and it's why the unresolvable count isn't a footnote to the real numbers. It's one of the real numbers.

What an honest ledger actually looks like

Picture a sprint with eleven claims made across three standups. Four resolve came_true. Two resolve did_not. The remaining five — nearly half — resolve unresolvable: two named no checkable subject at all, two had horizons that passed with no corresponding commit or telemetry movement either way, and one referenced a department rather than a function, which nothing in the repository or the runtime can currently check. A team seeing that ledger for the first time almost always reaches for the same instinct: round the five away, report "4 of 6," and call it a 67% sprint. Resist that instinct and the actual finding is more useful than any single percentage could be — nearly half of what got said out loud in this sprint's standups was never going to be checkable no matter what happened next, which is a fact about how the team talks, not about how the team performs, and it's fixable in a completely different way than a low calibration score would be. You don't coach someone out of an unresolvable claim. You coach them toward naming a subject next time.

That reframing matters more than it looks, because it changes what a bad number is asking a team to do about it. A low calibration score is a signal to check judgment. A high unresolvable share is a signal to check vocabulary — whether the standup format itself gives people an easy way to say something checkable, or whether the habit of the room defaults to generalities because generalities have always been safe. Two different diagnoses, two different fixes, and conflating them, the way rounding the number away conflates them, means treating a communication-format problem as if it were a competence problem. That's not just statistically sloppy. It's aimed at the wrong target entirely.

It's worth ending where Altman and Bland ended, because their closing point translates almost word for word: an inconclusive result reported honestly, with its limitations stated plainly, is more valuable to the next decision-maker than a confident conclusion the evidence never actually supported. Their audience was clinicians deciding on treatments. This arc's audience is a team deciding what to build next. The stakes are different. The statistical honesty required is exactly the same.

So the next time a claim from last sprint quietly drops out of the conversation because nobody ever circled back to it, resist the urge to let it drop silently. Ask which bucket it actually belongs in. Most of the time the honest answer will be the least satisfying one on offer — we don't know, and we still don't — and that answer, said out loud and counted, is worth more to the next sprint than either of the two tidier lies it could have been replaced with.

References

  1. Altman, D.G. & Bland, J.M. (1995). Statistics notes: Absence of evidence is not evidence of absence. BMJ, 311(7003).
  2. Rosenthal, R. (1979). The file drawer problem and tolerance for null results. Psychological Bulletin, 86(3).

See your own codebase as an office.

One pip install and every function reports for duty — archetype, live state, debt tier, and a single Code-Health North-Star. Free plan, no card.

Read next