All videos

Video · 10 min

Why measuring developer productivity keeps failing

By pressing "Play on YouTube", you agree to load this video from YouTube, a Google service. Google then receives data such as your IP address and may store data in your browser.
Details in our Cookie Policy and Privacy Policy.

Also available with audio in Arabic, French, German, Hebrew, Italian, Japanese, Korean, and Portuguese: choose one in the player's settings.

Watch on YouTube (Opens in a new tab)More from Oren Keinan on YouTube (Opens in a new tab)

Measuring developer productivity keeps failing when you count one activity. Apple's managers stopped asking one engineer for his weekly lines of code in 1982 after he wrote minus two thousand. Ron Jeffries, who says he may have invented story points, regrets how they're misused. Kent Beck and Gergely Orosz answered McKinsey's 2023 framework: nearly every measure it added counted effort or output. And Amazon shut down its token leaderboard in 2026 after employees cheated to climb it. Andy Grove wrote the fix in 1983: measure output, not activity, and pair every measure with a measure of its counter-effect.

In this video

  1. 0:00Amazon's token leaderboard
  2. 0:50Lines of code, 1982
  3. 1:33Story points, tickets and tokens
  4. 3:32McKinsey 2023 and the response
  5. 6:07Goodhart's law and the fix from 1983
  6. 8:02Ranking done right
  7. 9:25What to take to your next meeting

Transcript

Show the full transcript

0:00 Amazon's token leaderboard

In May 2026, Amazon shut down an internal leaderboard that tracked how many AI tokens its employees used. Some employees later told a news site called 404 Media that they'd cheated to climb it. If you lead engineers, you're probably already measuring them. Or you'll be soon. What you choose to count decides what happens next. Either your engineers start cheating the numbers, or they start doing better work.

By the end of this video, you'll know what to count so they do better work. And you'll know why the number should start a conversation, never become a verdict. Amazon isn't the first company to count the wrong thing. One famous case is 44 years old, and it starts with a weekly form at Apple. Later in this video, I'll also tell you what happened inside Facebook when one simple survey score started to matter.

0:50 Lines of code, 1982

-2000 lines of code. That's what Bill Atkinson wrote on a weekly form at Apple in 1982. Some managers on the Lisa team had asked each engineer to report the lines of code they wrote every week. The Lisa was the computer Apple was building at the time, and Atkinson was the main designer of its user interface. He'd just rewritten part of his own code.

The new version ran almost 6 times faster, and it was about 2,000 lines shorter. So when the form asked how many lines he'd written, he wrote -2000. After a couple more weeks, the managers stopped asking him to fill it in. To me, that one form shows the whole problem. A week that made the software faster and smaller made the number go negative.

1:33 Story points, tickets and tokens

Lines of code weren't the last single activity that teams counted. Many teams today estimate their work in story points. Ron Jeffries was the first coach of Extreme Programming, a team style of building software. He has written about where story points came from. A story point was simply a renamed ideal day, the time a pair of programmers would need without interruptions.

Points had one job, which was helping a team decide how much work to take on. In 2019, Jeffries wrote this. "I may have invented story points, and if I did, I'm sorry now." He wrote that comparing teams on their estimates or on the points they finish is harmful. Tickets have the same problem. In 2026, a paper from Microsoft's Engineering Thrive team described the pattern.

Engineering Thrive is Microsoft's own system for measuring the conditions its developers work in. Companies track activities like lines of code, pull requests and tasks. Then they treat those activities as if they were results. Notice that pull requests are on that list. A pull request someone opened is activity. A merged pull request whose code is still standing months later is much closer to a result.

I call counting the activity measuring the paperwork instead of the work. And to be clear, your ticket tracker isn't the problem. The mistake is measuring people through it. The newest single count is AI tokens, the small pieces of text that an AI tool takes in and gives back. That's what Amazon's leaderboard counted, according to the news site Business Insider.

It reported that the leaderboard encouraged some staff to do tasks that didn't necessarily solve problems, just so they could climb the ranks. It also reported that Amazon does track how many tokens it uses, but to measure its costs. Amazon was right to treat tokens as a cost. A token count is a fuel bill, not an achievement. And like the counts before it, it stopped measuring real work the moment people were ranked by it.

3:32 McKinsey 2023 and the response

In August 2023 the consulting firm McKinsey published an article called "Yes, you can measure software developer productivity." I'll give their case in its strongest form first, because it deserves that. McKinsey pointed out that other parts of a company get measured all the time, like sales. But software development gets measured far less. More and more companies are turning into software companies.

So their leaders need to know they're using their most valuable people well. And to be fair, McKinsey itself warned against simple measures like lines of code. McKinsey said nearly 20 companies were already using its approach. It kept some measures from well-known research and added its own. One was called contribution analysis. It looked at how much each person added to the team's backlog, the list of planned work.

It started from data in tools like Jira. About 2 weeks later, 2 respected engineers answered. Gergely Orosz writes The Pragmatic Engineer, a newsletter for software engineers that says it has more than a million readers. Kent Beck created Extreme Programming and test-driven development. Together they split software work into 4 stages, from effort and output to outcome and impact.

Effort is the planning and the coding. And output is what that produces, like the feature and its code. Then the feature reaches the customers. The outcome is how they behave differently because of the feature. And the impact is the value that comes back to the company, like revenue. And they wrote that nearly every measure McKinsey added counted effort or output.

Their central point was simple. Measuring developers changes how they work, because people start to cheat the system. And here's the Facebook story I promised. Kent Beck described what happened to a developer survey there. The survey was meant to show how things were going. Then the scores became goals, and managers started bargaining with their engineers for better scores.

In Beck's telling, one manager offered this deal. "Give me a 5, and I'll make sure you get an 'exceeds expectations'." That's a strong rating in a performance review. In the end the company knew even less about how things were going, as Beck and Orosz put it. And things were getting worse, because people were now cheating the survey.

Here's how I see that debate. McKinsey got the question right. Leaders do need to know where engineering time and money go. Actually, no. Let me put that more carefully. McKinsey asked a fair question and answered it with the wrong measures. You can't learn where the work went by counting how busy people look.

6:07 Goodhart's law and the fix from 1983

So can developer productivity be measured at all? Kent Beck says no. In the second part of their response to McKinsey, he wrote his own answer. "Measure developer productivity? Not possible." There's a law behind that answer, called Goodhart's law. An adviser at the Bank of England named Charles Goodhart described it in 1975. The best-known short version comes from a Cambridge University professor named Marilyn Strathern, in 1997.

"When a measure becomes a target, it ceases to be a good measure." I think Kent Beck and Goodhart are both right. But only about one kind of measure. Count one activity and reward it, and people will raise that number without doing better work. But measuring people changes how they work, and that's not always bad. It's the reason to measure at all.

What matters is what you count. One year after the Apple form, Andy Grove published a book called "High Output Management." That was 1983, and Grove was the president of Intel at the time. He wrote that a good measure covers the output of the work, not simply the activity involved. His example was a salesman. You measure him by the orders he gets, not by the calls he makes.

And because measures direct what people do, he wrote that you should pair them. You measure the effect, and you measure the counter-effect next to it. He even gave a software example. Track when each part of a program will be ready, and also what it can do. Watching both keeps you from building a perfect program that's never ready.

It also keeps you from rushing out one that isn't good enough. DORA stands for DevOps Research and Assessment, and it's Google's research program on how teams deliver software. It works the same way today. It puts how fast a team ships next to how often its changes fail. DORA publishes a guide to those measures, last updated in January 2026.

And that guide says speed and stability aren't a trade-off, because for most teams they rise and fall together.

8:02 Ranking done right

A ranking of engineers needs 4 things if you want it to make people better, not make them cheat. First, count the work that shipped and how long it lasted. Not how busy people looked. Second, pair every speed measure with a quality measure. Third, take points away when the work breaks. And more when it breaks sooner. And fourth, treat the number as the start of a conversation.

Never as a verdict. Measuring people still changes how they work. With a ranking built like this, it changes how they work in the direction you want. The way to climb is to do better work: write clean code that ships and doesn't need fixing soon after. Own the code you approve, not just the code you write. And don't keep your teammates waiting for a review.

[On screen: illustrative example, not real data]

I run a company called Aidealy, and we built the Aidealy Productivity Rank around these 4 rules. Among other things, it counts each engineer's merged pull requests. A change that leaves the code cleaner earns a little more, and one that leaves it messier earns a little less. It takes points away when that code has to be fixed after the merge, and more when it breaks sooner.

It gives credit for reviewing other people's work, and it counts how quickly someone reviews. The lines an engineer writes earn no points, and neither do tokens. We count tokens. We just don't rank by them. And the rank is where a conversation starts, never a verdict.

9:25 What to take to your next meeting

Now, there's something no rank can tell you. It shows how the work went, but it can't tell you whether the work was worth doing. That's a judgment, and it stays with you. So here's what to take to your next leadership meeting. Count one activity, and sooner or later people will cheat it. Rank on shipped work that lasted, pair speed with quality.

Take points away when the work breaks, and let the number start a conversation. The next video to watch is called Developer productivity metrics that get people fired. It goes through the measures that get cheated most often, one by one. Subscribe if you want it. The sources I used are all linked below, so read them yourself before you take my word for any of this.

See what your R&D is really doing.

Tell us what you're trying to figure out, and we'll get you set up to answer it on your own R&D: your engineers, your AI agents, and the code they ship together.

Talk to us