A team I worked with got access to AI coding tools. Good engineers, real delivery pressure. Within a few weeks, they were writing code noticeably faster.
Then they looked at their sprint results.
Nothing had changed.
The same number of features were reaching users. The same delays at testing. The same wait for approvals. AI made the coding faster. Everything around it stayed exactly the same.
This is the story of that gap – between AI freeing up time and your business actually getting something from it. It is a gap most organisations are not measuring. And it is where most AI transformation programmes quietly fail.
The freed time goes somewhere — just not where you think
When AI helps a developer write a feature in two hours instead of four, those two hours are real. They have been freed. That is not nothing.
But where do those two hours actually go?
In most large organisations, they get absorbed. The existing queue of work fills them up. Technical debt that was never prioritised gets a little attention. The next feature on the backlog gets started slightly earlier.
The system does not produce more value just because one part of it got faster.
Think of a pipe with three blockages in it. If you widen one section, more water does not come out the other end. The water just hits the next blockage harder and faster.
AI widened one section of the pipe. Whether more water comes out depends on where the other blockages are – and whether anyone has fixed them.
The wrong scoreboard
When organisations try to measure whether AI is working, most look at things like: how fast are we shipping? How often are we deploying? How quickly are bugs being caught?
These are useful numbers. But they are not the right question.
They tell you how fast the engine is running. They do not tell you where the vehicle is going.
The right question is: did customers do something different because of what we built?
Did they use the product more? Did complaints go down? Did more of them stay? Did the thing we shipped actually solve the problem we thought it would?
That is the measurement that matters. And it requires completely different instruments than the delivery dashboard most teams are looking at.
Here is the mistake most AI programmes make: they measure delivery speed and call it transformation. It looks like evidence. It fits on a slide. But it answers the wrong question.
AI is touching the fast part. Your constraint is probably somewhere else.
In any system, the slowest part determines the speed of everything.
In large banks, insurance companies, and telecoms — organisations that have been around for decades — the slowest part is almost never the code.
It is the three weeks it takes to get requirements signed off. It is the approval committee that has to clear every production deployment. It is the regulatory review that adds four months to any meaningful change. It is the decision that travels up three levels of management before it comes back down.
AI makes the code layer faster. It does not touch any of those things.
So if your organisation deploys AI coding tools and sees engineers writing code faster — but cycle times and business results staying flat — you have learned something important. The constraint was never the code.
That is actually valuable learning. The mistake is not discovering it. The mistake is not asking the question before you spend the budget.
The old systems problem
Most of the evidence that AI boosts developer productivity comes from companies starting fresh. Clean codebases. Modern tools. Small teams. Clear requirements.
That is not most large enterprises.
A bank running 40-year-old systems, with regulatory constraints on every significant change, with layers of technical debt built up over decades — that is a completely different environment.
AI productivity gains there are not a given. They might be real. They might be much smaller than the research suggests. The only way to know is to test it honestly, in your actual environment, with your actual constraints.
An honest experiment looks like this: one team, one defined constraint, a clear prediction about what AI will move, a measurement date agreed in advance, and named people responsible for the result.
Not a programme. Not a rollout. An experiment that can tell you something true.
What actually works
The organisations that close the gap — between AI freeing capacity and the business actually getting something from it — are not the ones with the highest AI adoption rates.
They are the ones that built the right sequence.
Find the actual constraint first. Not where AI can help. Where work actually slows down or stops. That is the starting point.
Form a specific hypothesis. “If we apply AI here, this specific thing will move, and we will see evidence of it by this date.” Vague predictions produce vague results.
Run a small experiment. Small enough to be honest. Large enough to be real.
Measure whether customer behaviour changed. Not whether delivery got faster. Whether the customer did something different as a result.
Decide deliberately what to do with the freed capacity. This is the step most programmes skip entirely. Freed capacity does not automatically become customer value. It gets absorbed by existing demand unless someone deliberately redirects it — and that redirection is an organisational decision, not a technical one.
The question to ask in your next review
If you are running an AI programme, or sitting in a review of one, there is one question worth asking out loud:
What did customers do differently because of this investment?
If the honest answer is “we don’t have that measurement yet” — the programme is measuring its own activity. It knows how hard the engine is working. It does not know where the vehicle is going.
Capacity freed is the starting point. What you do with it is the work that actually matters.
The author is an Enterprise Architect and Agile Transformation Leader writing a book on AI transformation in banks and large enterprises.

Leave a Reply