Video · 10 min
Developer productivity metrics that get people fired
By pressing "Play on YouTube", you agree to load this video from YouTube, a Google service. Google then receives data such as your IP address and may store data in your browser.
Details in our Cookie Policy and Privacy Policy.
Also available with audio in Arabic, French, German, Hebrew, Italian, Japanese, Korean, and Portuguese: choose one in the player's settings.
Watch on YouTube (Opens in a new tab)More from Oren Keinan on YouTube (Opens in a new tab)
Developer productivity metrics that count a single activity get cheated, and when they decide reviews or layoffs, the wrong people get hurt. Never read one alone: put next to each count the measure that goes down when the work is bad, and before you act on a number about a person, ask them.
In this video
- 0:00When a count decides who stays
- 0:54A number shows where to look
- 1:10Lines of code: pair them with what survived
- 1:54Commits: don't count them at all
- 2:46Story points and tickets: keep them in the team
- 3:50Pull request targets: pair them with fixes and waits
- 5:09Token counts and AI targets: adoption is not impact
- 7:35Tokens: judge a cost by what it bought
- 8:13Why the pair matters
- 9:01A ranking built on pairs
- 10:01What to look at instead
Transcript
Show the full transcript
0:00 When a count decides who stays
In July 2026, a few current and former Amazon employees spoke to the business news channel CNBC. They said some teams there now count AI use in performance reviews. If you lead engineers, you're probably measuring them already. And the numbers you pick may soon decide who stays. By the end of this video, you'll know which measures get cheated most often.
And you'll know what to put next to each one, so the cheating shows. I know the title sounds dramatic. Every case in this video is real, dated and linked below. At Amazon, it's performance reviews. At Coinbase, it was people's jobs. About a year earlier, Coinbase's chief executive fired engineers who hadn't set up their new AI coding tools within a week.
He told the story himself on a podcast. Some of those engineers had a good reason. The ones who didn't got fired.
0:54 A number shows where to look
Now, I think you should measure your engineers. But a number can show you where to look. It can't tell you why. So before you act on a number about a person, ask them. The real question is what you count, and what you let the number decide once you have it.
1:10 Lines of code: pair them with what survived
The oldest measure on the list is lines of code. The idea was that more lines meant more work done. It never did. And AI has made it worse. According to DX's report from mid-2026, AI now writes more than half of all the code in the companies it studies. DX is a company that measures how software teams work.
Faros AI is a company that analyzes engineering data. In its 2026 report, it said this about counting how much work comes out. "Throughput measures what was shipped, not what survived." So here's what I'd put next to a line count: how much of that code is still there a few months later. That one number tells you if the lines were work, or just typing.
1:54 Commits: don't count them at all
[On screen: illustrative example, not real data]
There's a free script on GitHub with more than 4,000 stars that fills a whole year of your commit history. Its own description says it does it instantly. To be fair, its authors say it's meant for learning and shouldn't be used to make anyone's work look bigger. But it shows how little a commit count proves. A commit count tells you how busy someone's history looks.
It can't tell you whether anything shipped. Take an engineer who spent a whole week on one hard fix. That's a single commit. In 2021, a group of researchers published a well-known framework for measuring developer productivity called SPACE. They wrote that counts like commits and pull requests "should never be used in isolation either to reward or to penalize developers."
So my advice for commits is short. Don't count them at all. Count the changes that got merged.
2:46 Story points and tickets: keep them in the team
Many teams estimate their work in story points. A story point is a guess about how big a task is, made before anyone writes the code. So when you count finished points, you're counting guesses. Closed tickets have a different weakness. Split one task into 5 tickets, and the count goes up 5 times. In 2023, Gergely Orosz and Kent Beck wrote about the leaders who use numbers like these.
Orosz writes The Pragmatic Engineer, a newsletter for software engineers. Beck created Extreme Programming, a team style of building software. They described a chief technology officer who wants to decide which engineers to fire. One way they listed was to pick a few easy-to-measure numbers, like pull requests and closed tickets. And they wrote this. "Decisions are often made based on metrics known to be incorrect."
So keep story points and ticket counts inside the team, for planning the next few weeks. To be clear, you can keep your ticket tracker. Just don't measure people through it.
3:50 Pull request targets: pair them with fixes and waits
Pull requests are next, and they need more care than anything else on this list. Uber built a dashboard that counted diffs per engineer. A diff is a code change sent for review, much like a pull request. Gergely Orosz asked engineers there how the dashboard changed things. People started creating far more diffs. That raised the load on the systems that build and test every change, and the cost of running them.
Some managers and directors then set targets for the number of diffs per month. People met them easily. They just made each diff smaller. And small changes are a good thing. DORA stands for DevOps Research and Assessment, and it's Google's research program on how teams deliver software. It recommends small changes. Its guide on working in small batches was last updated in December 2025.
It says to "avoid the temptation to generate massive pull requests." Small changes are easier to review and to test. And they're safer to merge. So the trouble at Uber was the target on the count. Break a big change into small ones that each pass review and stay in production, and that's good engineering. What matters is what you put next to the count.
How many of those pull requests needed a fix soon after they were merged. And how long each one waited for someone to review it.
5:09 Token counts and AI targets: adoption is not impact
The newest measures count AI itself. Every token an AI tool uses costs money. Some leaders now treat token spending as a sign of good work. Jensen Huang is the chief executive of Nvidia. In March 2026, he said this on a podcast. "If that $500,000 engineer did not consume at least $250,000 worth of tokens, I am going to be deeply alarmed."
His argument makes sense. An engineer with good AI tools can get more done, and the company is paying for those tools anyway. DORA makes a similar case for counting use. In June 2026, it wrote that a token leaderboard can get hesitant developers to try AI. And a team can't get results from AI if nobody uses it. I agree with that part.
Now look at what happens when use becomes the target. In April 2026, the newsletter The Pragmatic Engineer reported on Salesforce. It said the company expects its engineers to spend at least $170 a month on AI tokens. Anyone who spends less gets flagged. That same month, the tech news site The Information reported on a leaderboard at Meta. The leaderboard ranked more than 85,000 employees by the tokens they used.
In 30 days, they used about 60 trillion. Some left AI agents running for hours just to raise their numbers. Meta took the leaderboard down. Then in July 2026, a group of 26 former Meta employees sued the company. The news agency Reuters reported on their lawsuit. It says Meta relied on factors like productivity and AI token use when it chose whom to lay off.
And it says that hurt people who'd missed work for medical reasons. Meta denies the claims. A Meta spokesperson told Reuters this. "Workforce management and organizational decisions were and are made by people, not AI." According to DX, here's what all that use delivered. It sampled data from 400 companies, from late 2024 to early 2026. AI use went up 65% on average.
The number of pull requests they merged went up by less than 8%. Amazon shut down its own token leaderboard in May 2026. Dave Treadwell is one of its senior vice presidents, and he told his staff this. "Please don't use AI just for the sake of using AI."
7:35 Tokens: judge a cost by what it bought
So here's where I disagree with Jensen Huang. I want engineers to use AI, and to use a lot of it. Well, a lot of it where it helps. But a token bill tells you what the work cost, and only that. What the work delivered is a different number. DORA's own article from June 2026 also says what to put next to a token count.
It suggests looking at token use by team, by application or by use case. And it suggests the cost of each accepted change, meaning what you spend on AI for each change that got merged. That's the pair. Tokens are a cost, and you judge a cost by what it bought.
8:13 Why the pair matters
Here's why the pair matters so much right now. Faros AI followed about 22,000 developers for 2 years. As their AI use grew, the pull requests each developer got merged went up about 16%. Bugs per developer went up 54%. Count only the first number, and you'd count that as a success. And Faros wrote a warning for anyone planning layoffs based on that first number.
"The engineers being considered for cuts are in many cases the ones absorbing the quality gap AI is creating." In plain words, those are the people who review and fix what the AI wrote. That's what a pair is for. It shows you the cost next to the output, and the people who carry it. And to be honest, a pair makes cheating show.
It doesn't make it impossible.
9:01 A ranking built on pairs
[On screen: illustrative example, not real data]
I run a company called Aidealy. We built the Aidealy Productivity Rank on pairs like these. The lines an engineer writes earn no points. Neither do commits, closed tickets or tokens. We count tokens. We just don't rank by them. Among other things, a merged pull request earns credit, and more when it increases code quality or adds tests. Points come off when its code has to be fixed, and more when it breaks sooner.
Reviewing a teammate's work earns credit, and keeping them waiting costs points. Working across more languages and repositories earns more too. So does working with several AI chats at once, one measure among many. When someone's rank drops, it shows you where to look. You can ask Aidealy why it moved, and see the work behind the answer. Then you ask the person.
So would a rank like this get people fired? That's a choice leaders make. We built Aidealy so you can investigate the data behind every number, and understand where your team needs help.
10:01 What to look at instead
Here's what to look at instead, all on one screen, so you can take it to your next leadership meeting. The last column is a cost. Track it, but don't rank anyone by it. If a number can go up while the work gets worse, don't use it alone. The sources I used are all linked below, so read them yourself before you take my word for any of this.
If you want to keep going, pick one of the videos on the screen now, or subscribe for the next one.
Sources
- Annie Palmer, CNBC, "Burnout, frustration and heartbreak: Amazon layoffs take their toll in saturated job market", 2026 (Opens in a new tab)
- Julie Bort, TechCrunch, "Coinbase CEO explains why he fired engineers who didn't try AI immediately", 2025 (Opens in a new tab)
- DX, "DX releases Q2 2026 State of AI Impact in Engineering report", 2026 (Opens in a new tab)
- Faros AI, "The AI Engineering Report 2026: The AI Acceleration Whiplash", 2026 (Opens in a new tab)
- github-activity-generator, GitHub (Opens in a new tab)
- Forsgren, Storey, Maddila, Zimmermann, Houck and Butler, "The SPACE of Developer Productivity", Communications of the ACM, 2021 (Opens in a new tab)
- Gergely Orosz and Kent Beck, The Pragmatic Engineer, "Measuring developer productivity? A response to McKinsey", 2023 (Opens in a new tab)
- DORA (DevOps Research and Assessment), "Working in small batches", 2025 (Opens in a new tab)
- Lee Chong Ming, Business Insider, "Jensen Huang says he would be 'deeply alarmed' if his $500,000 engineer did not consume at least $250,000 of tokens", 2026 (Opens in a new tab)
- Evan Conaway, DORA, "Finding balance in the era of tokenmaxxing", 2026 (Opens in a new tab)
- Gergely Orosz, The Pragmatic Engineer, "The Pulse: 'Tokenmaxxing' as a weird new trend", 2026 (Opens in a new tab)
- Matthias Bastian, The Decoder, "Meta employees compete for token consumption on an internal AI leaderboard" (reporting The Information), 2026 (Opens in a new tab)
- Reuters, "Meta used AI to target workers with medical conditions for layoffs, lawsuit claims", 2026 (as published by Deccan Herald) (Opens in a new tab)
- Justin Reock, DX, newsletter post on how much AI use and merged pull requests grew across 400 companies, 2026 (Opens in a new tab)
- Brent D. Griffiths, Business Insider, "Amazon says it shut down a token leaderboard", 2026 (Opens in a new tab)
See what your R&D is really doing.
Tell us what you're trying to figure out, and we'll get you set up to answer it on your own R&D: your engineers, your AI agents, and the code they ship together.
Talk to us