Skip to content
Talk to an engineer

Delivery

The typing was never the expensive part

AI adoption now correlates with higher delivery throughput and with lower delivery stability at the same time. Both are true, and together they decide what is worth paying for.

6 min readComputing America

In short

  • In DORA’s 2025 data, AI adoption correlates positively with delivery throughput and negatively with delivery stability: the writing got faster and the controls downstream of it did not.
  • Veracode’s 2025 tests across more than 100 models found 45% of working samples introduced an OWASP Top 10 vulnerability, and that rate stayed flat while the same models got measurably better at functional correctness.
  • GitClear’s commit-corpus analysis finds the share of changes that touch code last modified more than twelve months ago fell from 1.7% to 0.46% between 2023 and 2026, which is what a maintenance problem looks like before anyone calls it one.
  • The decision in front of a buyer is no longer whether to use these tools, because almost everyone does. It is whether tests that can fail a build, review by someone who did not write the code, and a way back exist before change volume rises.
  • The most-quoted statistic in this argument, METR’s finding that experienced developers were 19% slower with AI, has been qualified by its own authors and by METR’s own follow-up, and should not be used as evidence by anyone, including us.

Everybody has seen the demo by now, and it was not a trick. Somebody described a screen and a working version of it existed before the meeting ended. That is real, it is not going away, and any argument that starts by disputing it has already lost the room.

The useful question is what happens next. Two years of evidence on this does not divide into people for whom the tools work and people for whom they do not; it divides into what happens before the code exists and what happens after. In Stack Overflow’s 2025 survey (opens in a new tab), 84% of developers use or plan to use these tools and the most-cited frustration, at 66%, is output that is almost right but not quite. Almost right is not a complaint about a demo. It is a description of where the work moved.

Two findings that have to be held at once

DORA’s 2025 report on AI-assisted software development (opens in a new tab), from nearly 5,000 technology professionals, reports 90% using AI at work and more than 80% believing it has made them more productive. Unlike the previous year, it finds a positive relationship between AI adoption and both delivery throughput and product performance. It also finds, again, a negative relationship between AI adoption and delivery stability.

Those two are not in tension. More change, arriving faster, moves through the same review, the same test suite and the same release process, and the thing that gives is the rate at which changes fail. The generation step got cheaper and nothing downstream of it did.

“AI doesn’t fix a team; it amplifies what’s already there.”
DORA, State of AI-Assisted Software Development 2025 (opens in a new tab)

That is the most useful sentence published on the subject, because it is not a verdict on the tools at all. It says they resolve to whatever was already true about how you ship: teams with real controls convert the speed into throughput, teams without them convert it into incidents, and both groups report feeling faster. If you are deciding whether to buy engineering help, that sentence is the decision, restated.

The number we are not going to use

One widely quoted study on developer productivity would be convenient for a firm like ours, and we are not going to use it. METR ran a randomized controlled trial (opens in a new tab) with sixteen experienced open-source developers across 246 real tasks in repositories they knew well, and found them 19% slower with AI access than without, while those same developers predicted they would be 24% faster. It is an RCT, it is the most-cited figure in this debate, and we are not going to use it.

Three reasons, and they are the authors’ own. The paper states that it is not evidence AI fails to speed up most developers, that it does not generalize past experienced developers on code they already know, and that it is a historical result. METR then published a correction to the experiment design (opens in a new tab) in February 2026, describing selection effects that understated the benefit: developers who declined to take part because they did not want to work without AI, and participants who withheld the tasks they expected AI to help with most. METR’s own position now is that developers are likely more sped up than that study measured.

We raise it because somebody else will, and because a number that needs a footnote to survive contact is a talking point rather than evidence. Our own pricing already concedes the speed: we do not bill by the hour, for the plain reason that tooling made implementation materially faster and an hourly rate would charge you less the better we got at the work.

Capability improved. Security did not come with it.

Veracode tested 80 coding tasks across more than 100 models (opens in a new tab) in Java, JavaScript, Python and C#. Of the samples that worked, 45% introduced an OWASP Top 10 vulnerability; by language, Java failed 72% of the time, and cross-site scripting was left undefended in 86% of the cases where it was relevant.

The percentage is not the finding. The derivative is: across models of increasing size and sophistication, functional correctness improved and security performance stayed flat. Veracode sells application security testing, which is a reason to read the shape of that result rather than bank the exact figure. The shape is the part that matters, and it is what a demo cannot show you. A demo tests whether the thing works. Nothing in a working demo is evidence about the failure modes that carry contractual or regulatory weight.

The part that arrives in year two

GitClear’s analysis of 623 million code changes (opens in a new tab) between 2023 and 2026 carries one measure worth understanding, because it is the only one that speaks to the second year: the share of changes that remove or update code last touched more than twelve months ago. It fell from 1.7% to 0.46%. New code is increasingly being written beside the old rather than through it.

Two caveats, stated rather than buried. This is correlation across a corpus of commits, not a controlled trial, and GitClear sells developer analytics. Their headline refactoring figure is also quoted against two different baselines across their own materials, which is why the number above is the one we cite: it is stated consistently everywhere it appears. Read it as a signal with a direction, not as a measurement of your codebase.

What the signal describes is familiar to anyone who has inherited a system. Everything works. Nothing is documented. The same logic exists in four places with small differences, and nobody can say which one is authoritative or why any of it does what it does. That now includes the person who commissioned it, because they did not write it either.

What to check before the volume rises

None of this is an argument for generating less code. It is an argument that three specific things have to exist before change volume goes up, and that their absence is invisible right up until it is expensive.

  1. 1.A test suite that can fail a build. Not a test suite, but one that blocks. The distinction sounds pedantic and is the whole control: a suite that reports is a suite somebody overrides at 5pm on a Thursday.
  2. 2.Review by somebody who did not write it, which now means somebody who did not prompt it. The author’s blind spot is inherited by whoever accepted the output, and accepting output is faster than writing it, which is exactly why more of it arrives.
  3. 3.A way back. A release you can undo without a meeting. Stability is not the absence of failures; it is how quickly a failure stops being everyone’s problem.

There is a fourth that is less obvious. Generated code encodes decisions nobody typed, so the reasoning that used to live in somebody’s head now lives nowhere at all. Keep it in whatever form suits you: a decision record, a comment on the pull request, three lines in a file. The question to be able to answer in a year is not what the code does, which is readable, but what it was supposed to do.

What it changes about buying

For a genuinely small internal tool, build it yourself: one screen, a handful of people in one building, nothing regulated, nobody outside depending on the answer. That was true before these tools existed and it is more true now, and a firm that will not say so is protecting its own scope.

The boundary is not size and it is not sophistication. It is the first moment somebody outside the team that built it depends on the answer, or the first time being wrong costs money. Past that line the thing being bought is not code, which is now the cheap part. It is the review, the tests, the evidence and the accountability that the change volume has already outrun.

If you already built something this way and it is starting to matter, the first thing we hand back is an assessment rather than code: what is worth keeping, what the remaining work actually is, and what finishing costs against what replacing costs, with three options priced and doing nothing among them. It is priced to be worth doing on its own, which is the only arrangement under which the answer “keep what you have” is credible.

Sources

  1. 1.State of AI-Assisted Software Development 2025 (opens in a new tab), DORA (Google Cloud),
  2. 2.2025 GenAI Code Security Report (opens in a new tab), Veracode,
  3. 3.The Maintainability Gap: 2026 AI code quality research (opens in a new tab), GitClear
  4. 4.Measuring the impact of early-2025 AI on experienced open-source developer productivity (opens in a new tab), METR,
  5. 5.We are changing our developer productivity experiment design (opens in a new tab), METR,
  6. 6.AI (opens in a new tab), Stack Overflow Developer Survey 2025

Next step

Send us the thing you already built.

Tell us what it does and who depends on it now. We will name which of the three downstream controls it is missing (tests that can fail a build, review by someone who did not write it, or a way back) and say whether that is worth fixing before it grows.

Reply
A person replies, not a sequence: within one business day, from someone who would be on the engagement.