Talk to us ↗

Explaining AI ROI to Your Board Without Getting Caught Out

The standard slide — 90% adoption, 40% self-reported productivity, annualised to millions — will not survive a numerate board member. Here is an answer that does.

At some point in the next two quarters, somebody on your board will ask what the AI spend is returning. It is a fair question and most engineering leaders answer it badly — not through evasion, but because the honest answer is more complicated than the format allows.

Here is how to give an answer that survives scrutiny, and why the easy version will not.

The answer that fails

The standard slide says something like: 90% of engineers use AI daily; they report a 40% productivity improvement; annualised, that is worth $X million.

Every element of that is defensible in isolation. Together they will not hold, and a numerate board member will find the seam.

Adoption is an input, not a result. DORA's research puts industry adoption at around 90% of technology professionals. Reporting your own 90% says you are normal, not that you are winning.

Self-reported productivity has been measured against reality, once, carefully, and it did not survive. METR's randomised controlled trial had sixteen experienced developers complete 246 real tasks on repositories they knew well. With AI they were 19% slower — and estimated they had been 20% faster. A 39-point gap, persisting after they had finished the work.

That result does not prove AI is useless. It proves that asking people how much faster they feel produces a number uncorrelated with what happened. If your ROI rests on such a survey, it rests on the same mechanism.

Annualising a perception compounds the error. Multiplying an unreliable percentage by headcount and salary produces a large, precise, meaningless figure — and precision is exactly what invites the question you cannot answer.

What a defensible answer looks like

Three parts, in this order.

One: what we can actually measure.

Name the things you have instrumented and give the numbers, including flat ones. Cycle time on a repeated class of work. Rework share — the proportion of effort spent after a first version was shown. Cost per completed outcome, not cost per call. Time to first meaningful contribution for new hires.

A leader who says "cycle time is flat, rework fell four points, cost per merged change fell eleven percent" is more credible than one claiming forty percent, because the numbers are uneven. Real results are uneven. Uniform improvement across every metric is the signature of a number that came from a survey.

Two: what we are deliberately not claiming.

Say plainly that self-reported productivity is not evidence, and why. This costs you nothing — the board will discover it eventually — and buys a great deal, because you become the person who told them first.

Then name the cost that is not on the invoice. DORA calls it the verification tax: time saved in creation re-spent on auditing. With 96% of developers reporting they do not fully trust AI-generated code to be functionally correct, somebody is checking, and that checking does not appear in any budget line. If you have measured it, give the number. If you have not, say when you will.

Three: what we are actually buying.

This is where most decks are weakest, because the honest answer is not "speed".

The bottleneck in software has never been typing. Research consistently finds developers spend 58% to 70% of their time understanding existing code rather than writing it. Making writing cheaper addresses the smaller half and increases the volume of material to be understood — which is why DORA observes copy-paste rates rising from 8.3% to 12.3% between 2021 and 2024 while refactoring collapsed from about 24% to under 10%.

So the investment case is not "we type faster". It is: we are trying to reduce what it costs this organisation to understand its own systems, because that is where 60-plus percent of engineering time goes and where the expensive failures originate.

Deloitte attributes 64% of defects to the requirements and design phases, and the IBM System Science Institute's figures put a defect fixed in maintenance at roughly 100× its cost at design. That is a board-legible argument, and it is true.

Three questions to be ready for

"Why is our delivery rate flat if everyone is faster?" Because production got cheaper and comprehension did not, and comprehension was the constraint. This is the moment to introduce the verification tax rather than to defend the survey.

"What would make you stop spending this?" Have a real answer. "If cost per completed outcome does not fall by X over two quarters, we reduce the licences." A leader with a stopping condition is trusted on the continuing condition.

"Are our competitors ahead?" On adoption, nobody is ahead — it is at 90%. The differentiator is not whether a company uses these tools; it is whether it has instrumented the effect. Most have not, which is an opportunity rather than a threat, and it is worth saying so.

The trap worth naming

There is pressure to produce a big number, because other departments do. Marketing reports attributed pipeline; sales reports closed revenue; engineering feels obliged to match.

Resist it. An inflated AI ROI figure has a specific failure mode: it is checked in twelve months, against a delivery rate that did not move, by people who now discount everything else you have told them. The cost is not the correction — it is the credibility, which you need for the next investment case.

A smaller number you can defend is worth more than a large one you cannot. That is true generally, and it is especially true here, because this particular claim will be audited by reality on a schedule you do not control.

Put a number on your own workflow

Bring us the recurring workflow where AI still needs expensive people to supervise, review and correct. We baseline what it costs and put the target in writing before we build. Or connect a repository and see it on your own code first — 1,000 credits free, no card.

Build My Business Case →   Start free →

Frequently asked questions

How should I report AI ROI to a board?

In three parts: what you have actually measured, including the flat numbers; what you are deliberately not claiming, and why self-reported productivity is not evidence; and what the investment is really buying — a reduction in what it costs the organisation to understand its own systems, which is where most engineering time and most expensive failures sit.

Why can't I use developer surveys as evidence of AI productivity?

Because the one controlled trial available found the opposite of what participants believed. Sixteen experienced developers were 19% slower with AI while estimating they were 20% faster, and the gap persisted after they had completed the tasks. Any figure derived from asking people how much faster they feel inherits that mechanism.

What metrics are defensible in an AI ROI discussion?

Cycle time on a repeated class of work, rework share, cost per completed outcome rather than per call, and time to first meaningful contribution for new hires. Uneven results across these is a sign of real measurement; uniform improvement everywhere is the signature of a survey.

What is the strongest argument for continuing AI investment?

That comprehension, not typing, is the constraint — developers spend 58% to 70% of their time understanding existing code, and 64% of defects originate in requirements and design rather than coding. Investment aimed at reducing the cost of understanding addresses the larger half of the work and the more expensive class of failure.

Should I report adoption rates to the board?

Only as context. Industry adoption is around 90%, so your own high number says you are normal rather than ahead. What distinguishes companies now is whether they have instrumented the effect, and most have not.

Vladimir Miroshnichenko
Vladimir Miroshnichenko
Founder, GitMir

Founder of GitMir, the intelligence layer that gives AI real context about a company's software. I write about AI agents, context engineering, workflow economics and keeping AI-generated work under control.

LinkedIn →

← More articles