Argument
Individual Productivity Metrics Break Teams. The Evidence Is Not Ambiguous.
The argument against per-developer measurement, the 2023 debate that made it public, and the peer-reviewed research on what monitoring does to teams.
Every framework in this space says the same thing: measure teams and systems, not people. DORA is a delivery-system diagnostic. SPACE goes out of its way to warn against judging developers by story points or lines of code. Core 4 calculates diffs per engineer as an organisational average and states explicitly that it isn't an individual measure.
Organisations do it anyway. It's worth being precise about why that fails, because "it's bad for morale" is both true and too weak to survive a conversation with a CFO.
Why does the pressure exist at all?
Because engineering is the only large cost centre that claims to be unmeasurable.
This is the uncomfortable part of the argument, and Kent Beck and Gergely Orosz named it directly in 2023: CEOs and CFOs are frustrated by engineering leaders who respond to measurement questions by saying software is too nuanced to measure. Sales and recruiting both report individual and team performance in numbers nobody disputes. Engineering, spending comparable or larger sums, says the question is malformed.
That vacuum is what gets filled. When McKinsey published its developer productivity framework in August 2023 — claiming an approach already deployed at nearly twenty companies — it was answering a real demand, not inventing one.
Any argument against individual metrics that doesn't offer an alternative answer to the CFO's question loses. It should lose.
What was actually wrong with the McKinsey framework?
Beck, who created extreme programming and has spent forty years on this problem, and Orosz, a former Uber engineering manager, published a two-part response that became the reference document for the field's position.
Their central objection is structural rather than emotional. Software work runs through a chain: effort, then output, then outcome, then impact. The framework measured the first two — code volume, activity levels, deployment frequency — and skipped the second two. It captured half the lifecycle and called the result productivity.
Dave Farley, co-author of Continuous Delivery, was blunter: setting aside its use of DORA metrics, he compared the rest of the framework to astrology.
The mechanism of harm both identified is specific. Productivity scores do not stay in a dashboard. They migrate into performance reviews — particularly during cutbacks — and once they do, the numbers stop describing the work. Beck's observation is that this distorts the measurement itself, not merely the culture around it. And Beck and Orosz noted the motive that makes this predictable: a common reason executives want individual productivity data is to decide who to let go.
What does the research say?
Peer-reviewed work published in Communications Psychology in 2024 by Schlund and Zitek, with roughly 1,200 participants, tested algorithmic monitoring directly. Monitoring quadrupled complaints and reduced idea generation.
The finding that matters most is the exception. Those effects largely disappeared when the monitoring was framed as developmental rather than evaluative. Identical data, different stated purpose, different outcome.
That result is useful in both directions. It undermines the position that all measurement is surveillance — it isn't, and framing changes the result materially. It also undermines the position that intent alone makes individual metrics safe, because framing only holds while the data stays out of evaluation. The moment it enters a performance review, the developmental framing is no longer true, and people know it.
The failure modes, concretely
Optimisation against the metric. Developers will optimise for what's measured. Diff counts rise by splitting pull requests. Story points inflate. Cycle time improves by avoiding hard work. None of this requires bad faith; it's the rational response to being ranked.
The invisible work disappears. Mentoring, unblocking colleagues, preventing an incident that then doesn't happen, deleting code, simplifying a system so future work is cheaper. Frequently the highest-leverage contributions on a team, and structurally invisible to every metric in this category.
Context is not comparable. A developer on a legacy service with a two-week release train will always look slower than one on a greenfield service. The gap measures the systems, not the people.
Collaboration becomes irrational. If your number depends on your output, helping someone else costs you. Individual metrics impose a tax on exactly the behaviour that makes teams effective.
What to offer instead
The alternative that works is not "you can't measure this." It's a different unit of analysis.
Measure at team and system level, and say so. DORA and delivery metrics answer the throughput and stability question honestly at the level they were designed for.
Add outcome and impact. This is where the McKinsey critique cuts both ways: effort and output are insufficient, so bring the other half. What shipped, what changed for users, what the business got. This is the answer to the CFO question, and it's an answer engineering leaders too often decline to give.
Add developer experience data. Surveys covering friction, focus time and review burden explain why the delivery numbers look the way they do.
Draw an explicit line, in writing. State that individual metrics do not enter performance evaluation, and mean it. The research says framing does real work — but only while it remains accurate.
Answer the underlying question directly. When someone asks for a per-developer breakdown, the productive response is to find out what decision they're trying to make. It's usually a resourcing, performance or investment question that has a better answer.
Frequently asked
Are individual metrics ever legitimate? In a developmental context a developer controls themselves, yes — the research supports it. As an input to evaluation or comparison, no.
What do I tell an executive demanding per-developer numbers? That the data exists and is misleading, that the framework authors say so, and that the question they're actually asking has a better answer. Then give it.
Doesn't refusing to measure make engineering look evasive? Yes, which is why refusal is the wrong move. Measure at the right level and report on outcomes, rather than declining the conversation.
Isn't Core 4's diffs per engineer an individual metric? It's calculated per engineer and reported as an organisational average. The framework's authors state directly that it is not tracked at the individual level. Using it that way is a misuse, not a feature.
Get new analysis by email
Independent work on engineering measurement. No vendor sponsorship, no affiliate placement, no weekly cadence padded with links.