CodeNSM
The Standup · Part 16

Talking time is not signal

2026-06-19· 8 min read· by Think North

Every team has an unspoken leaderboard, and it is measured in airtime. The person who talks the most in standup reads, to the room, as engaged, on top of things, driving the work. The person who says four sentences and sits back reads as either coasting or — in the more generous version — mysteriously wise, saving their words for when they matter. Neither read is measuring anything real. Both are measuring volume, and this post exists to make the case, carefully, for why volume should never enter a scoring system — and to correct a popular misuse of the one famous study people reach for to make that case.

What the Science paper actually found

In 2010, Anita Woolley, Christopher Chabris, Alex Pentland, Nada Hashmi, and Thomas Malone published a study in Science that gets cited constantly and read carefully rather less often. They gathered small groups, had them work through a diverse battery of collaborative tasks — brainstorming, moral reasoning, visual puzzles, negotiation — and looked for something like the "g factor" of individual IQ testing, but for the group as a whole: a single statistical factor, which they called collective intelligence (c), that predicted a group's performance across the whole range of unrelated tasks.

The design choice worth appreciating here is the same one that made individual-IQ research convincing in the first place: a single group, tested on a single task, tells you almost nothing about whether that group is generally good at working together, versus just lucky on that particular puzzle. Woolley and her co-authors ran each group through several genuinely different kinds of tasks specifically so they could ask whether performance on one predicted performance on the others — and it did, which is what justified calling it a factor at all, rather than a coincidence dressed up as one.

They found one. And here is the part worth being precise about, because it is the part that gets flattened into a slogan: c was only weakly correlated with the average or maximum individual intelligence of the group's members. Stacking a group with high-IQ individuals did not reliably buy you a high-performing group. What did correlate with c, more strongly, was the average social sensitivity of the group's members — essentially, how well people could read each other's mental and emotional states — along with the evenness of conversational turn-taking, and the proportion of women in the group (a finding the authors linked back to social sensitivity, which tends to score higher on average among women in the standard measures they used). A later, much larger replication and extension by Christoph Riedl, Young Ji Kim, Pranav Gupta, Thomas Malone, and Anita Woolley — pooling data across 22 studies and over five thousand people — held up the core result: the collaboration process predicted group performance more reliably than the individual skill of whoever happened to be in the room.

What it does not say, and why the distinction matters here

Now the correction, because it's the whole reason this post exists. The Woolley study is about predicting a group's performance on a shared task, using group-level statistical patterns measured across many different groups. It says nothing — literally nothing, the study wasn't designed to and doesn't claim to — about whether any specific individual who talks less in any specific meeting is smarter, more insightful, or more likely to be right about a specific technical claim. "Equal turn-taking correlates with group collective intelligence" is not the same statement as "the quiet person in your standup is secretly correct" or "loud people are bad at their jobs," and treating it as a personality-typing tool for judging your coworkers is a popular misreading the original authors never endorsed and the data never supported.

What the study does license, carefully stated, is this: at the group level, a pattern where a few people dominate the conversation tends to depress collective performance relative to a pattern where contribution is more evenly distributed — not because the quiet members were right and the loud ones were wrong, but because dominance crowds out information that other members were holding, regardless of whose information turns out to matter more. That's a claim about process, not about any one person's individual correctness. It is not a personality test, and it is not evidence for or against any specific claim any specific person makes.

The Woolley study is not a scientific excuse to trust your quietest engineer more. It's evidence that a standup where three people talk and six people don't is probably leaving information on the table — which is a completely different, and more useful, thing to know.

The trap CodeNSM's design refuses to walk into

Here's why this distinction has teeth for anyone trying to score a standup honestly. It would be extremely tempting to build a "collective intelligence score" straight out of the popular misreading — track talk time per person, flag the loud ones, praise the quiet ones, call it science. It would also be exactly the mistake CADENCE's founding rule exists to block: understanding is never scored from what someone said, only from whether what they said turned out to be true. Talk time, turn-taking, tone of voice, confidence — none of it is allowed anywhere near the number that says whether a person's claims held up, because none of it is what the Woolley research actually measured about individuals, and because using it that way would rebuild, with better branding, the exact talk-equals-competence bias this whole series keeps circling back to.

What the research is genuinely useful for is a completely different, structural question: not "who should I trust," but "is this meeting's format itself suppressing information?" A standup where two people say ninety percent of the words is a standup where — per the actual mechanism Woolley's data points to — other people's claims, corrections, and blockers are statistically more likely to go unsaid. That's worth noticing and fixing at the level of the meeting's shape. It is not evidence about any individual's Grip, and treating it as such is the misreading this post was written to head off.

Why the misreading is so tempting in the first place

It's worth asking why this particular misreading spreads so easily, because the answer says something about standups specifically. Talk time is the single easiest thing to observe in a meeting — you don't need telemetry, a call graph, or a commit log, you just need to have been in the room and have a sense of who filled the silence. Everything else this series has argued actually matters — whether a claim named a subject, whether it later came true, whether a blocker's latency was climbing — requires waiting, checking, and record-keeping. Talk time requires none of that. It's available instantly, for free, at the end of every single meeting, which makes it an almost irresistible stand-in for the harder, slower, actually-informative signal, the same way a resting heart rate is an irresistible stand-in for a full cardiac workup: cheap, available, and not remotely the same measurement.

The Woolley research is seductive to misuse for exactly this reason — it's a genuinely rigorous, Science-published finding that happens to involve the word "talking," and it's much easier to remember "quiet groups are smarter" than to remember the actual, narrower, group-level claim about turn-taking equality and social sensitivity. Popular science journalism did some of this flattening on its own, long before any given engineering manager got hold of it. The correction matters here specifically because a standup-scoring product is exactly the kind of place where the flattened version, if it snuck in unchecked, would do real damage — quietly re-introducing the talk-equals-signal bias this whole series exists to push back against, with a research citation attached that makes it look load-bearing.

A short thought experiment to keep the two straight

Imagine your quietest engineer says exactly one sentence in a standup, all sprint: "the retry logic is causing the faults, not the queue." Loud or quiet, confident or mumbled — none of that changes whether the fault rate on the retry function actually dropped after it was fixed. Now imagine your most talkative engineer says fifteen confident sentences, and eleven of them named no checkable subject at all. Fifteen sentences of airtime and one falsifiable claim in there is not a better contribution than one sentence with a checkable subject. It's eleven unresolvable claims and one that will actually get scored — the same claim ledger either way, regardless of who said more words to get there.

Score the claim. Not the minutes it took to say it.

References

  1. Woolley, A.W., Chabris, C.F., Pentland, A., Hashmi, N. & Malone, T.W. (2010). Evidence for a Collective Intelligence Factor in the Performance of Human Groups. Science, 330(6004), 686–688.
  2. Riedl, C., Kim, Y.J., Gupta, P., Malone, T.W. & Woolley, A.W. (2021). Quantifying Collective Intelligence in Human Groups. PNAS, 118(21), e2005737118.

See your own codebase as an office.

One pip install and every function reports for duty — archetype, live state, debt tier, and a single Code-Health North-Star. Free plan, no card.

Read next