The day velocity became a target (and stopped being a measure)
Somewhere in your sprint history, probably without a specific dramatic moment attached to it, your team crossed a line. Before the line: story points were a private scratchpad an engineer used to reason about how much work a ticket really was, relative to other tickets they'd done. After the line: story points became the number leadership watches trend upward on a chart, the number that shows up in the quarterly review, the number a team's morale quietly starts tracking. Nothing about the arithmetic changed. Everything about what the number is for did.
What happens next is not a story about your team being unusually cynical. It's one of the best-documented regularities in the social sciences, and getting the credit for it right matters, because this exact idea gets misattributed constantly, usually to whichever name is easiest to remember.
Three people, three different discoveries, one conflated legend
Start with the actual originator. Charles Goodhart was a Bank of England economist, and in a 1975 paper — "Problems of Monetary Management: The U.K. Experience," delivered at a Sydney monetary-economics conference and later reprinted in the collection Monetary Theory and Practice — he was writing specifically about UK monetary policy, arguing that once regulators started targeting a particular statistical relationship between money-supply measures and the economy for control purposes, that relationship reliably broke down. This is a precise, technical, econometric observation about central banking. Goodhart never wrote the punchy one-liner people now attribute to him.
That one-liner belongs to someone else entirely. Marilyn Strathern, an anthropologist studying audit culture in British academia, restated Goodhart's underlying insight in a completely different context — the UK's university research-assessment exercises — in a 1997 European Review paper, "'Improving ratings': audit in the British university system," and it's her sentence, not Goodhart's, that everyone actually quotes: "When a measure becomes a target, it ceases to be a good measure." Strathern gave Goodhart's economics a portable, general-purpose phrasing, applied it to institutions far from monetary policy, and — through one of history's small ironies — the sentence keeps getting handed back to Goodhart, who never wrote it, almost as often as Strathern goes uncredited for having written it.
And there's a third, independent strand that deserves its own name rather than folding into the other two. Donald T. Campbell, a social scientist studying evaluation research, arrived at a closely related but separately derived conclusion in a 1979 Evaluation and Program Planning paper, "Assessing the Impact of Planned Social Change," studying things like standardized testing and crime statistics: "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Where Goodhart's point is essentially statistical — a regularity breaks down once you lean on it for control — Campbell's is essentially institutional: people, under pressure, actively game the indicator, whether that's teachers narrowing instruction to the test or police departments reclassifying crimes to improve the numbers. Both things are true. They are not the same mechanism, and treating "Goodhart's Law" as a single catch-all for both, in Strathern's clothing, erases two genuinely different research traditions that happened to converge on the same warning from different directions.
Story points, corrupted on schedule
Watch what happens to an estimate once it becomes a target, and both mechanisms show up right on cue. Goodhart's version: the statistical relationship that made story points useful in the first place — that a five-point ticket really does, on average, take about as long as another team's five-point ticket — quietly stops holding, because the number is no longer being generated for the private reasoning purpose it was designed for. Campbell's version, layered on top: points start inflating (why size something a three when a five looks better on the burndown), tickets get sliced to pad the count, and "toil" work that soaks up real capacity gets quietly folded into estimates in ways that keep the chart's upward trend intact regardless of what actually shipped. Neither of these requires anyone to be acting in bad faith. It requires only that the number became a target, and people — entirely rationally, exactly as both Goodhart and Campbell separately predicted — started optimizing the number instead of the thing it used to represent.
The part that makes this so hard to catch from inside a team is that nobody involved has to be acting in bad faith at any single step. The engineer who sizes a ticket a five instead of a three isn't lying — the ticket genuinely does feel bigger once you know the burndown is being screenshotted for a stakeholder update. The manager who quietly stops logging "toil" as its own category isn't hiding anything — it's just tidier not to have a line item that makes the trend look worse. Each individual decision is locally reasonable, sometimes even generous. The corruption isn't a person. It's the aggregate, and it's precisely the kind of thing Campbell was describing when he wrote about indicators being distorted by the social processes they're meant to monitor rather than by any one villain inside those processes.
Goodhart described the statistics breaking. Campbell described the people bending. Strathern gave the whole phenomenon a sentence short enough to fit on a slide — which is, not coincidentally, exactly the kind of measure that becomes a target next.
What doesn't automatically escape this
It would be convenient to claim that a claim ledger — the calibration and alignment numbers this arc has been describing — is somehow immune to Goodhart, Campbell and Strathern simply because the resolving evidence lives downstream, in commits and telemetry, further from easy manipulation than a self-reported story point. That's a real structural advantage, and it's worth stating plainly: it is much harder to fabricate a commit history than to round an estimate upward. But "harder to game" is not "impossible to game," and the honest position is that the day people start speaking for the scoreboard instead of just being checked against it afterward — hedging every claim into unfalsifiability, making only claims about code nobody's watching, gaming which claims even get logged — is the day this exact argument needs to be turned back on the thing making it. Every measurement invites this. Pretending otherwise is how a measure becomes a target in the first place.
What that gaming would actually look like, concretely, is worth naming rather than leaving abstract: claims phrased just vaguely enough to dodge a clean subject, so they land in the unresolvable pile instead of at risk of a did_not — which is exactly why Part 26 of this arc insists that bucket be measured and watched, not just protected. Or claims made only about code far from any real traffic, where being wrong costs nothing on a reach-weighted score. A team that starts noticing these patterns isn't discovering a flaw unique to this measurement. It's discovering that Goodhart, Campbell and Strathern were right about something more general than economics or auditing — they were right about measurement itself, and the honest response is to keep watching the watcher, indefinitely, rather than to declare any single system finally immune.
The test that actually tells you where you are
Ask, plainly, in your next retro: does anyone privately optimize this number, separately from optimizing the thing it's supposed to represent? If the honest answer is yes for velocity, you're already past Strathern's line, and the chart going up and to the right has stopped telling you what you think it's telling you.
None of this is an argument against measuring anything. It's an argument against measuring one thing and rewarding it, then acting surprised when the thing you wanted — real capacity, real throughput, real progress — quietly stops correlating with the number you built to track it. Goodhart, Campbell and Strathern spent three separate careers, in three separate fields, proving the same uncomfortable law from three different directions. The least you can do, having read all three, is stop crediting the wrong one for the sentence you actually quote.
And the smallest, most concrete step available to any team reading this today costs nothing: the next time someone quotes velocity in a planning meeting, ask what it was originally for. Not accusingly — genuinely. If the honest answer has drifted from "helping this team reason privately about its own capacity" to "showing leadership the trend is up," the number has already changed jobs, whether or not anyone updated the label on the chart.
References
- Goodhart, C.A.E. (1975/1984). Problems of Monetary Management: The U.K. Experience. In Monetary Theory and Practice.
- Campbell, D.T. (1979). Assessing the Impact of Planned Social Change. Evaluation and Program Planning, 2(1).
- Strathern, M. (1997). 'Improving ratings': audit in the British university system. European Review, 5(3).