Goodhart’s Law: How to Measure Your Software Team Without Gaming the Numbers

Charles Goodhart, an economist at the Bank of England, articulated in the 1970s what practitioners in many fields already knew intuitively: when a measure becomes a target, it ceases to be a good measure. The act of making something a target changes the behavior around it in ways that undermine what the measure was supposed to track.

Software engineering is possibly the richest environment Goodhart’s Law has ever encountered. We have more ways to measure developer work than almost any other discipline. Story points, lines of code, pull request count, deployment frequency, test coverage, velocity, lead time, bug counts, code review turnaround. Every one of these can be gamed once it becomes a performance target. Most teams have at least one story about when one of them was.

The Classic Examples

Story point velocity is the one most Agile teams encounter first. Velocity is a useful planning tool: if a team has averaged 30 points per sprint for six sprints, they’ll probably complete around 30 points next sprint. The moment velocity becomes a performance measure, things change. Estimates inflate. Stories get split to create more points from the same work. Teams compare velocities between teams, which is meaningless since points aren’t calibrated across teams. The number goes up. The actual delivery doesn’t.

Code coverage is another common one. 80 percent coverage sounds rigorous. Teams trying to hit 80 percent coverage find ways to hit 80 percent coverage. They write tests that call code without asserting anything meaningful. They mock everything so the tests can’t possibly fail. The coverage number reaches 80 percent. The test suite doesn’t actually catch more bugs.

Pull request count is a simple one. If someone decides PRs are a proxy for productivity, developers write more, smaller, less thoughtful PRs. The number goes up. The quality of the work doesn’t improve.

Each of these is Goodhart’s Law in action. The measure tracked something real before it became a target. Once it became a target, it stopped tracking what it was supposed to.

What LEAN Offers: Flow Metrics

LEAN’s answer to Goodhart’s Law is to measure the system, not the individuals, and to use metrics that describe flow through the system rather than output from individuals.

Cycle time is the time from when work starts to when it’s delivered. Lead time is the time from when work is requested to when it’s delivered. Throughput is the number of items completed per unit time.

These are harder to game than individual metrics. A developer can write more PRs, but they can’t individually reduce the team’s cycle time without actually improving the flow of work. The metrics describe something systemic, and improving them requires systemic improvements.

The WIP level is another flow metric. When WIP is high, cycle time tends to be high. Reducing WIP by setting limits improves cycle time, which improves lead time. The connection between the lever and the outcome is direct and honest.

What DORA Measures

The DevOps Research and Assessment group spent years studying what engineering metrics correlate with high-performing software organizations. They identified four: deployment frequency, lead time for changes, change failure rate, and mean time to recovery.

These metrics are a direct response to Goodhart’s Law. They can’t easily be gamed in isolation, because they’re in tension. You can increase deployment frequency by shipping smaller things, but if those smaller things have a high failure rate, change failure rate goes up and you’ve traded one metric for another. Improving all four simultaneously requires genuine system improvement, not measurement manipulation.

The DORA metrics are also outcome-oriented rather than output-oriented. They describe whether the software delivery system is working, not how many things people produced.

OKRs vs KPIs

The OKR framework distinguishes between objectives (qualitative outcomes you want to achieve) and key results (measurable signals that you’re achieving them). The distinction matters because key results are lagging indicators, not the goal itself.

This is more resistant to Goodhart’s Law than traditional KPI-based management. If the objective is “users find our onboarding experience smooth,” then the key result might be “onboarding completion rate above 70 percent.” The team knows the number serves the objective. Gaming the number without improving the experience would be visible in user research, in support tickets, in churn data.

The key is that the team understands what they’re actually trying to achieve and uses the metric as evidence, not as the thing itself.

Building a Measurement Culture

Teams that use metrics well have a few things in common. They measure multiple things simultaneously so that gaming one metric degrades another. They treat metric changes as signals to investigate rather than verdicts to act on. They distinguish between metrics used for team planning and metrics used for external reporting.

They also have explicit conversations about what each metric is supposed to tell them and what it doesn’t. Velocity tells us about sprint capacity planning. It doesn’t tell us about team health, technical quality, or user satisfaction. Knowing what a metric can’t tell you is as important as knowing what it can.

Want to build a metrics approach that drives real improvement? Let’s talk.

Goodhart’s Law: How to Measure Your Software Team Without Gaming the Numbers

Leave a Reply

Your email address will not be published. Required fields are marked *