AI coding measurement has become surprisingly sophisticated in a very short time. We can measure active AI users. We can see which tools developers use. We can count prompts, accepted suggestions, tokens, AI-assisted commits, and increasingly even identify which parts of a codebase were created with AI assistance.
At the other end of the software delivery system, we already have mature engineering metrics:
Yet there is a large gap between these two worlds.
We know how much AI is being used.
We know how software is being delivered.
What we often don't know is what happened in between.
That missing middle is where I think most of the interesting AI productivity questions now live.
This was reasonable. When companies started rolling out GitHub Copilot, Cursor, Claude Code, and similar tools, engineering leaders first needed answers to basic questions:
These are important operational metrics.
But they answer a rollout question: Are developers using AI?
They don't answer the much harder question: What changed in the engineering system because developers used AI?
The distinction becomes more important as adoption increases. Once most developers in an organization use an AI coding tool, increasing adoption from 70% to 80% tells an engineering leader relatively little about whether the investment is working. At that point, usage stops being the interesting variable. The outcomes become interesting.
Software doesn't become valuable when code is generated. It still has to survive:
Coding → Review → Testing → Merge → Deployment → Production → Maintenance
This sounds obvious, but many AI productivity dashboards effectively stop at the first step.
Imagine an AI assistant reduces coding time by 30%.
Great.
Now imagine that the resulting pull requests are larger, review takes 25% longer, and rework increases.
Did productivity improve?
Maybe.
Maybe not.
It depends on where the time went.
This is the first measurement principle I think engineering organizations need in the AI era:
A local productivity gain is not necessarily a system productivity gain.
Software engineering is a pipeline. Accelerating one stage can expose or create a bottleneck somewhere else.
Consider a simplified engineering workflow:
Work Item
↓
Coding
↓
Pull Request
↓
Review
↓
Testing
↓
Deployment
↓
Production
Suppose AI dramatically increases the rate at which developers produce pull requests.
The arrival rate into code review increases. But reviewer capacity hasn't changed. You can end up with something like this:
| Metric | Change |
|---|---|
| Coding Time | -25% |
| PR Throughput | +30% |
| PR Pickup Time | +18% |
| Review Time | +27% |
Looking only at coding activity makes AI look highly successful.
Looking at the entire system tells a more complicated story.
The bottleneck moved.
This is not necessarily a failure of AI. In fact, it may mean the AI tool is doing exactly what it should.
The organization simply hasn't adapted the rest of its engineering system to the new throughput.
I actually think AI code attribution is becoming an important engineering signal. But not for the reason people sometimes assume.
If 60% of a team's changes are AI-assisted, that does not mean the team is 60% more productive. It doesn't even mean those developers saved 60% of their coding time. What it gives us is something much more useful:
An analytical dimension.
Now we can ask:
How do highly AI-assisted changes behave compared with less AI-assisted changes?
For example:
| Dimension | Question |
|---|---|
| Coding | Are AI-assisted changes completed faster? |
| PR Size | Are AI-assisted PRs larger? |
| Review | Do they require more review time? |
| Rework | How much code changes again shortly after merge? |
| Quality | Do quality issues change? |
| Delivery | Does lead time improve? |
| Stability | What happens to failed changes? |
AI contribution becomes useful when it is joined with engineering outcomes.
By itself, it is just another activity metric.
One of the mistakes we made with traditional developer productivity metrics was assuming visible activity corresponded closely to useful work: commits, lines of code, tickets closed, pull requests created.
AI makes that assumption even more dangerous because producing artifacts is becoming dramatically cheaper.
The question therefore changes from:
How much did developers produce?
to:
Where did engineering effort move?
This is closely related to a broader shift in software engineering measurement: activity metrics become far more useful when they are interpreted alongside flow, quality, delivery, and developer experience. I explored that broader measurement model in a separate guide to software engineering metrics, including why isolated activity counts can create misleading conclusions.
An AI assistant might reduce:
At the same time, it might increase:
If 40 minutes disappear from implementation but 25 minutes appear in verification, the productivity gain is not 40 minutes. And if that verification work falls on another developer, looking only at the original developer's metrics will completely miss it.
Instead of a single AI productivity metric, I find it more useful to think of measurement as a funnel.
Who can use AI?
Examples:
Who actually uses it?
Examples:
Where does AI participate in engineering work?
Examples:
What happens to development?
Examples:
What happens after that code is created?
Examples:
Does the organization ship differently?
Examples:
Did something economically meaningful change?
Examples:
The further down this funnel you go, the closer you get to actual organizational impact.
The downside is that attribution becomes harder.
That is unavoidable.
AI Engineering Funnel
Executives understandably like summary numbers. But I would be cautious about creating something like:
AI Productivity Score: 83
unless everyone understands exactly what went into it.
AI affects multiple dimensions that can move in opposite directions. For example:
| Metric | Change |
|---|---|
| Coding Time | -21% |
| PR Throughput | +17% |
| Review Time | +14% |
| Rework | +9% |
| Deployment Frequency | +8% |
| Change Fail Rate | +3% |
| Developer Satisfaction | +16% |
Is AI working?
That is a much more interesting engineering discussion than whether a score changed from 76 to 83.
It also exposes something important:
Productivity is multidimensional.
A productivity improvement may appear as:
Trying to compress all of that into one number can destroy the information engineering leaders actually need.
This is another reason the debate around AI productivity often becomes confusing.
Some controlled experiments have shown large improvements in task completion speed when developers use AI coding assistants.
Other real-world studies involving experienced developers working in mature repositories have found smaller gains, no gains, or even temporary slowdowns.
Those results sound contradictory.
They aren't necessarily.
They measure different environments.
A bounded programming task is different from changing a mature production system containing:
This makes universal statements like:
"AI makes developers 30% faster"
almost meaningless without context.
The better question is:
Which developers, performing which tasks, in which codebases, using which AI tools, measured at which part of the delivery system?
Suppose an organization sees:
It is tempting to connect those numbers. But many things may have changed simultaneously.
A better analysis starts segmenting.
AI may be extremely effective in one codebase and much less useful in another.
Feature development, maintenance, testing, refactoring, and incident fixes are different activities.
Different engineering practices can dramatically change AI outcomes.
Compare lower and higher AI-assisted work.
Otherwise, AI-heavy work may simply be larger or smaller.
Compare stable periods rather than only the week immediately before and after rollout.
The goal isn't to manufacture causality.
The goal is to eliminate obviously misleading comparisons.
Engineering analytics sometimes falls into another extreme.
If we cannot prove causality perfectly, some people conclude we shouldn't measure the relationship at all.
I disagree.
Observational engineering data can still reveal useful patterns.
Suppose high-AI pull requests repeatedly show:
That doesn't prove AI caused the pattern. But it gives an engineering leader a very useful hypothesis:
Maybe our AI-enabled development workflow needs smaller pull requests or stronger automated validation.
You can now change the system and observe what happens.
Measurement becomes an improvement loop rather than a performance judgment.
That's much more useful.
AI telemetry makes extremely granular measurement possible. That doesn't mean every possible metric should become a management metric.
I would be especially cautious about metrics such as:
Once people realize a metric affects how they are evaluated, the metric stops behaving like neutral telemetry.
It becomes a target. And targets get optimized.
Instead, start with teams and workflows.
Ask:
Is this team's engineering system improving?
Then use deeper data diagnostically when the team needs to understand why.
The objective should be improving the engineering system, not building a more sophisticated surveillance system.
This is perhaps the simplest practical rule.
Whenever AI appears to improve one metric, place a balancing metric beside it.
| If you measure... | Also measure... |
|---|---|
| Coding Time | Rework |
| PR Throughput | Review Time |
| PR Size | Review Load |
| Deployment Frequency | Change Fail Rate |
| AI Contribution | Quality |
| Cycle Time | Developer Experience |
| AI Cost | Delivery Outcome |
Why?
Because optimization in engineering frequently transfers cost.
A team can increase deployment frequency by making smaller deployments.
Good.
Or by bypassing necessary controls.
Not good.
The number alone can't tell you which happened.
Not 40 AI metrics. I'd start with something much smaller.
Active AI Developers
Are people actually using the tools?
AI-Assisted Development Rate
Where is AI participating in actual engineering work?
PR Cycle Time
Is work moving through development faster?
Review Time
Did the bottleneck move downstream?
Rework Rate
Are we creating additional follow-up work?
Change Lead Time
Is local acceleration reaching production?
Change Fail Rate / Deployment Rework
Are faster changes remaining reliable?
Developer Perception
Do developers actually feel that the workflow improved?
AI Cost
What are we paying to create those changes?
That's already enough to have a substantially better conversation.
A few years ago, the interesting question was:
Can AI write useful production code?
Then it became:
Will developers adopt AI coding assistants?
Both questions are becoming less interesting.
The next question is harder:
What happens to a software engineering system when AI becomes a normal participant in development?
An AI tool's usage dashboard can't answer that question. And it cannot be answered by traditional engineering metrics alone.
You need both.
AI attribution without engineering outcomes tells you what AI did. Engineering outcomes without AI context tell you what changed. Connecting the two is where we begin to understand impact. And that missing middle may turn out to be the most important part of measuring software engineering in the AI era.