Video · 10 min
DORA metrics when half your code is AI-written
By pressing "Play on YouTube", you agree to load this video from YouTube, a Google service. Google then receives data such as your IP address and may store data in your browser.
Details in our Cookie Policy and Privacy Policy.
Also available with audio in Arabic, French, German, Hebrew, Italian, Japanese, Korean, and Portuguese: choose one in the player's settings.
Watch on YouTube (Opens in a new tab)More from Oren Keinan on YouTube (Opens in a new tab)
DORA (originally DevOps Research and Assessment) metrics still work when half your code is AI-written, but only in pairs: a speed number is safe to read only next to a stability number. This video reads DORA's 2024 and 2025 findings on AI adoption and shows what to change in your next quarterly review.
In this video
- 0:00Faster and less stable, in the same report
- 0:56What the DORA measures were built to tell you
- 3:01Why 2024 and 2025 disagree
- 4:47Read speed only next to stability
- 7:33Three changes for your next quarterly review
- 8:53What these numbers cannot tell you
- 9:36What to watch next
Transcript
Show the full transcript
0:00 Faster and less stable, in the same report
According to DORA's 2025 survey, nine out of ten people who build software now use AI at work. By the way, DORA is Google's research program on how companies build and ship software. And in that same survey, software delivery got faster. And less stable. I know how that sounds. If you're hoping for a sales pitch built around the DORA metrics you already use, this isn't that video.
Here's why this matters now. DX is a company that measures engineering teams. According to its data from more than five hundred companies, AI was writing more than half of all code by the middle of 2026. By the end of this video, you'll know which DORA metric you can still read on its own. And which one misleads you when it stands alone.
And there's one finding from DORA's 2025 research that I haven't seen anyone quote. It's about teams that adopt AI without paying attention to their users. I'll come back to it in a few minutes. So let's talk about DORA properly.
0:56 What the DORA measures were built to tell you
The name stands for DevOps Research and Assessment. It's a research team inside Google, and every year it surveys thousands of people who build software. For about ten years, DORA has been known for four numbers that describe how well software gets delivered. Meaning how fast and how safely changes reach the users. People call them the four key metrics, or simply the DORA metrics.
And they were always meant to be read together. The first DORA metric is called change lead time. It measures how long a change takes from a developer's finished code to running in production, which means live in front of real users. The second metric is deployment frequency. A deployment is when you put a new version of the software live, so this is simply how often you ship.
Once a month, or ten times a day. The third one is called failed deployment recovery time. That's how fast you recover when a deployment goes wrong. Those three metrics are about throughput, meaning how much work gets through to production. Then there's the fourth metric, change fail rate. It counts how often a deployment needs an urgent fix or a rollback, and a rollback means putting the old version back.
And there's a fifth one now. DORA publishes an official guide to these metrics, and the latest version from January 2026 lists a fifth metric called deployment rework rate. It counts the deployments nobody planned, because something broke in production and forced them. Those last two metrics are about instability. For most teams, throughput and stability moved together. Teams that shipped more often were also more stable, because the habits that let you ship often are the same habits that catch problems early.
DORA's guide says it plainly: speed and stability aren't tradeoffs, meaning you don't have to give up one to get the other. That connection is what made the DORA metrics useful as a diagnosis, a way to spot what's wrong. If your speed went up but your stability didn't, you knew something was wrong. That connection is what AI broke.
3:01 Why 2024 and 2025 disagree
Let's go back to DORA's 2024 report. That year, DORA found something it called contrary to its expectations. More AI adoption went together with worse delivery. For every 25 percent rise in AI adoption, throughput dropped by about one and a half percent. And stability dropped by more. Those are estimates from survey data, and DORA says so. DORA's explanation was that people were producing a lot more code in the same time.
So each change was probably getting bigger. And bigger changes have always been slower and more likely to break. Back in 2024, Google's DORA research estimated that every 25 percent rise in AI adoption went with about seven percent lower delivery stability. Then the 2025 report came out, and half of that picture reversed. AI adoption now went with higher delivery output, meaning more work getting through.
But the other half didn't reverse. AI adoption still went with lower stability, meaning more failures. Two years, one research program. Same warning. Faster, and less stable. So your dashboard shows deployment frequency going up, and you call that good news. You're reading half the story. The half you're not reading is the stability half, and that's the half that produces the incidents.
DORA's own explanation is in the 2025 report. AI speeds up development, and that extra speed exposes weaknesses later in the delivery process. More changes flowing through turn into more instability, unless you've got strong safeguards in place. DORA calls those safeguards control systems and names three of them: automated testing, version control and fast feedback. Fast feedback means finding out quickly when something breaks.
So here's the change in how you read the DORA metrics from now on.
4:47 Read speed only next to stability
Each metric now needs a partner. A throughput number only makes sense next to a stability number. And a stability number only makes sense next to the safeguards behind it, like your automated tests. A metric that used to tell you what was wrong now only tells you where to look. The partner matters because of the size of each change.
LinearB is a company that tracks software delivery, and its 2026 benchmarks cover about eight million pull requests from nearly five thousand teams. Three out of four pull requests written with AI help stay under about four hundred lines. Pull requests written without AI stay under about a hundred and sixty. And when the whole pull request was written by an AI agent, meaning an AI working without a developer.
It waited almost eighteen hours before anyone started reviewing it. Pull requests written without AI waited about three and a half hours. So the changes got bigger, the waits got longer. And the reviewers stayed the same. That's the mechanism DORA guessed at in 2024, now measured. Faros AI is another company that measures engineering work, and it checked the same thing a different way.
Not with a survey. Faros pulled the numbers out of the engineering tools themselves. Over two years, across twenty-two thousand developers. As companies moved from low to high AI use, bugs per developer went up by more than half. And pull requests merged with no review at all went up by nearly a third. Then Faros says something DORA doesn't.
Companies with good DORA metrics got worse just like everyone else. Faros puts it bluntly. Perception lags reality. In plain words, a survey tells you how people feel. And feelings catch up late. The measured numbers don't wait. Now, here's the finding I promised you earlier. Alongside its 2025 report, DORA published a companion document it calls the AI Capabilities Model.
It names seven things a company needs to be good at, from a clear position on AI to a focus on its users. And it says that without a focus on the users, adopting AI can hurt a team's performance. To me, that says adopting AI is a decision for the whole organisation. And the fastest way to get that decision wrong?
Read one number on its own. This is where I disagree with the way DORA describes all this. DORA's 2025 report calls AI an amplifier, meaning it makes good teams better and struggling teams worse. That's the strongest version of the idea, and honestly it's mostly right. But remember what Faros measured: the good teams got worse too. So I think the amplifier idea is true, and also dangerous.
It lets a strong team believe it's safe. My rule? Look at your stability numbers even when your team is good. Especially then.
7:33 Three changes for your next quarterly review
So, three things you can change before your next quarterly review. First, never report deployment frequency on its own. Put change fail rate right next to it, every single time. So speed and stability are always read together. Second, split your delivery numbers by whether AI helped write the change. Do that per service or per repository, meaning for each piece of software.
Never per person. That part is DORA's own rule. The third change comes before you buy more AI tools. Check whether your automated tests and rollback process can absorb much more change. If they can't, more AI will make your numbers worse.
[On screen: illustrative example, not real customer data]
That second change is the hard one, because it needs two numbers your dashboard probably doesn't have. Which merged changes were written with AI help? And how many of those had to be corrected soon after? That gap is the one thing I'll say about Aidealy, the company I run. Delivery tools measure how fast and how often you ship, and Aidealy shows that too.
But its main focus is different. It measures the workforce, meaning the people and the AI agents. On the quality and lasting impact of what they ship. That includes the bugs that surface later. So a question like which repositories had the most AI-assisted changes corrected within thirty days this quarter has an answer.
8:53 What these numbers cannot tell you
Here's what none of this tells you. DORA's 2025 numbers are survey answers from nearly five thousand people, and the 2024 figures are estimates. Both show that two things moved together, AI adoption and delivery stability. They don't tell you what will happen in your organisation. Faros measures real tool data instead of asking people, but one dataset isn't the whole industry.
I do not know which of the two describes your team. And neither do they. That's exactly why reading the metrics in pairs matters. Your own stability number next to your own speed number is the only evidence in this video that's about you. And even that only tells you how the work went. Whether the work was worth doing is a different question.
That judgment stays with you.
9:36 What to watch next
So yes, the DORA metrics still work. For ten years you could read them as one package, and any single number on its own still told you something. Now they only work as two pairs. And a speed number on its own is the most misleading thing on your dashboard. The next video to watch is called Why measuring developer productivity keeps failing.
It's about why counting how busy people are fails as a way to measure productivity. Subscribe if you want it. The reports I mentioned are linked below, so read them yourself before you take my word for any of this.
Sources
- DORA, Accelerate State of DevOps Report 2024, announcement (Opens in a new tab)
- DORA, State of AI-assisted Software Development 2025, announcement (Opens in a new tab)
- DORA, software delivery performance metrics guide (Opens in a new tab)
- DORA AI Capabilities Model (Opens in a new tab)
- LinearB, 2026 Software Engineering Benchmarks Report (Opens in a new tab)
- Faros AI, The AI Engineering Report 2026 (Opens in a new tab)
- DX, State of AI Impact in Engineering, second quarter (Q2) 2026 (Opens in a new tab)
See what your R&D is really doing.
Tell us what you're trying to figure out, and we'll get you set up to answer it on your own R&D: your engineers, your AI agents, and the code they ship together.
Talk to us