Case Study: Putting a Defensible Number on AI Adoption
An AI programme everyone liked nearly lost its budget because nobody could prove it worked. Two baselined workflows later, it survived the cut — and earned an expansion.

"The team really likes it" is the sentence that kills AI budgets1. Not because it is false — because it is unpriceable, and it always loses to a line item that isn't.1A composite case, anonymised from several Fellow engagements. Numbers are representative of the pattern, not one client's audited results.
The situation
A 600-person services organisation had run AI licences and training for a year. Usage was genuinely healthy: daily activity, enthusiastic champions, anecdotes everywhere. Then the CFO asked the reasonable question — what did we get for it? — and the programme discovered it had no answer. No baseline existed for any workflow the tools had changed. The renewal went into review with sentiment as its only defence.
The review meeting was not hostile, which somehow made it worse. The CFO was not sceptical about AI — she used it herself — she was sceptical about a benefit nobody had written down anywhere she could check. The programme lead's first instinct was to reconstruct the past: pull calendar data, ask senior consultants how long a proposal used to take before the tools arrived. Two days of that produced recalled figures ranging from three hours to eleven for the same document type. Retrospective recall is advocacy, not evidence, and it does not survive a follow-up question.
Their readiness profile showed the signature: solid Applied use, near-zero on impact measurement◦.By the numbers0 of 14AI-touched workflows with any before-measurement, at the point the budget question arrived
The anecdotes also disagreed with one another. Two teams reported that drafting time had halved. A third had quietly stopped using the tools months earlier and nobody had noticed. When nothing is counted, the good news and the bad news are equally invisible — which is why an unmeasured programme cannot defend itself even when it is working.
Most AI programmes cannot answer this question
This is not a local failure of one services firm. MIT's Project NANDA, in its 2025 report on the state of AI in business, estimated that around 95% of enterprise generative-AI pilots deliver zero measurable profit-and-loss return1◦. The nuance matters more than the headline, and most people repeating the number drop it: "zero return" means no impact that shows up in the accounts, not that the tools were useless. The same report found individual productivity gains were common. What was missing was the path from a person saving forty minutes to a line the finance team can see. NANDA's own diagnosis of the failure driver was a learning gap in workflow integration◦ — not model quality. The tools were good enough; the workflows around them had not changed.SourceMIT Project NANDA, "The GenAI Divide: State of AI in Business 2025" (via Fortune)By the numbers~95%MIT Project NANDA's 2025 estimate of enterprise generative-AI pilots with no measurable P&L returnDefinitionLearning gap MIT Project NANDA's term for the gap between a capable tool and a workflow that has actually been rebuilt around it — the failure driver it identified, rather than model quality
McKinsey's March 2025 State of AI survey of 1,491 organisations found the same shape from the finance side: more than 80% reported no tangible enterprise-level EBIT impact from generative AI, and fewer than one in five tracked well-defined KPIs for their generative-AI solutions2◦. Read those two numbers together and the second largely explains the first. An organisation that has not defined a KPI has not built the instrument that would detect an impact, so the absence of evidence is guaranteed in advance.SourceMcKinsey, "The State of AI", March 2025By the numbers<1 in 5organisations tracking well-defined KPIs for their generative-AI solutions, McKinsey State of AI, March 2025
BCG's October 2024 study of 1,000 senior executives across 59 markets put it in blunter terms: 74% of companies had yet to show tangible value from AI, and only 4% were consistently generating significant value3. But BCG also recorded the flip side, and it is the part worth quoting to a CFO — the companies it classed as AI leaders achieved 1.5 times the revenue growth and 1.6 times the shareholder returns of their peers over three years. The gap between 4% and 74% is not mostly a technology gap.SourceBCG, "Where's the Value in AI?", October 2024
Two more data points frame the timing risk. In July 2024 Gartner predicted that at least 30% of generative-AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value — that is a forward-looking prediction rather than a measured outcome, and should be cited as one4. And in Deloitte's Wave 4 survey of 2,773 director-to-C-suite leaders, published in January 2025, more than two-thirds expected that 30% or fewer of their current generative-AI experiments would be fully scaled within three to six months5. The window in which an unproven programme is safe is roughly one budget cycle long.SourceGartner press release, 29 July 2024SourceDeloitte, "State of Generative AI in the Enterprise", Wave 4, January 2025
The counter-example is instructive because it is so ordinary. Lumen Technologies, working with Microsoft Copilot, reported that its sellers were saving about four hours per week — customer research that had taken roughly four hours came down to about fifteen minutes — which Lumen's own arithmetic put at approximately $50 million annually6. Treat that figure with care: it is a company estimate published through the vendor's marketing channel, not an audited result. Its value here is not the dollar number at all. It is that somebody had written down how long the task took before, in a specific unit, so that an after-number meant something. That is the entire difference between Lumen's story and the 95%.SourceLumen Technologies on Microsoft Copilot — a company estimate published through the vendor
Why value goes unmeasured
Nobody plans to skip measurement; it is skipped by default. In the excitement phase, measuring feels like bureaucracy slowing the fun down. Later, the before-state is gone — nobody recorded how long a report took in the era before the tools, so the improvement became unprovable the moment it succeeded◦.DefinitionBaseline The before-measurement — time, quality, adoption — taken prior to a change, without which no after-number means anything
The cost is not hypothetical. Unmeasured programmes lose budget reviews to measured ones regardless of relative merit, and the people who quietly gained hours a week lose them back when the licences go.
There is a second reason, and it is the one the external research keeps pointing at: the thing that is easy to feel and the thing that shows up in the accounts are not the same object. Forty minutes saved on a draft is real and immediately invisible, because saved minutes only become value if they are redeployed into work the business counts — more proposals sent, faster response to clients, a role not backfilled. Programmes that never make that second step explicit are honestly reporting a benefit their own finance function has no way to see.
The third reason is ownership. A measure that belongs to everyone is produced by no one, and the first month it is inconvenient it simply stops being produced. In every version of this pattern we have seen, the measurement survived exactly as long as one identifiable person's name was attached to it◦.RelatedWhat changes when adoption gets one named owner
What we did
We did not build a measurement programme. We baselined exactly two workflows — proposal drafting and client-meeting summaries — chosen for volume, repeatability, and how visibly they mattered to the business.
The selection took one ninety-minute workshop with fourteen candidate workflows on a wall. Two survived three questions: does it happen at least weekly, does it produce a comparable unit of output, and would a director notice if it got better? Most of the fourteen failed the second question, which is the one that quietly disqualifies most measurement attempts — you cannot compare "a piece of analysis" to another piece of analysis, but you can compare one first draft of one proposal type to another.
For each: two weeks of honest before-numbers (time per unit, revision counts, who used AI at which step), one named owner of the number, and a one-page monthly delta report designed in the CFO's own format. Nothing else was measured on purpose — two credible numbers beat ten estimates◦.Figure
One person owning one honest number is a measurement system; a committee with a dashboard is usually not
The resistance was immediate and legitimate. Senior consultants heard "log your time per draft" as surveillance, and one said so in the kickoff. The design changed in that conversation rather than after it: logging was recorded at workflow level rather than per person, results were reported as team medians, the window had an announced end date two weeks out, and no individual's figures went to a line manager. Had we argued the objection down instead, the logs would have been quietly padded and we would never have known which numbers were real.
It nearly went wrong anyway. In the first week several entries came back suspiciously round — two hours exactly, repeatedly, from different people. They were logging what a draft ought to take rather than what it took. The tell was the variance: real professional work is messy, and a baseline that is too tidy is almost always wrong. We rebriefed, discarded the first three days, and extended the window by three. A contaminated baseline is worse than no baseline, because it produces a confident number that collapses under one question.
The whole thing cost roughly three person-weeks spread over a quarter: the ninety-minute selection workshop, about half a day a week from the named owner during the baseline fortnight, then around two hours a month to produce the one-page report. The largest line was not analysis. It was the twenty-odd short conversations needed to get people to log honestly, and the one conversation with the CFO's analyst that got us the template used for capital requests. That last detail mattered more than expected: the same number in the finance team's own layout is read as a claim to be checked, while in a slide deck it is read as marketing.
What changed
One quarter later, proposal first-draft time was down by roughly half against its baseline, with revision rounds flat — quality had held◦. The renewal conversation took ten minutes: the programme survived the cut that year, and the same evidence format earned an expansion into two more workflows the next.By the numbers≈50%reduction in proposal first-draft time against a two-week baseline, quality flat
Ten minutes, because there was almost nothing to argue about. The CFO asked two questions — what is the unit, and who counted it — and both had answers with names attached. The claim being made was also deliberately modest: one workflow, one quarter, one measured delta, with the revision-count figure included precisely because it was the number most likely to have moved in the wrong direction.
The deeper change was cultural: with a delta report on the table, adoption stopped being a belief and became a managed quantity — which also exposed one workflow where AI genuinely was not helping, and it was retired without drama. Proof cuts both ways, and that is the point.
Specifically, the meeting-summary workflow held up well for internal reviews and badly for the committee meetings whose minutes had to be signed off: the draft saved perhaps twenty minutes and cost thirty in correction. Nobody had wanted to say so out loud while the programme's survival depended on good news. Once the reporting format could carry a negative result without threatening the budget, saying so became easy — and the credibility that bought is the reason the following year's expansion was approved on the strength of the format as much as the result.
The part that is harder than it sounds
Baselining is easy to describe and awkward to do, and three things are harder than this write-up makes them sound.
First, by the time anyone wants a baseline, the tools are already in use — so you are measuring a contaminated before-state, not a clean one. The honest description of our two-week baseline is "current practice including whatever AI use already existed", which understates the total gain and is the conservative direction to be wrong in. If we were starting again we would say that out loud in the first report rather than in the third.
Second, two weeks is too short for anything seasonal. Proposal volume and complexity move together, and a quarter-on-quarter delta partly reflects what came through the door. We now log volume and a rough complexity band alongside time, so the obvious challenge — "were these just easier proposals?" — can be answered rather than absorbed.
Third, and least comfortable: a delta is not causation. We can say first-draft time halved after the change. We cannot rule out that a revised template, a new hire, or simply a year's accumulated practice did part of the work. The number's job is to survive scrutiny in a budget meeting, not to prove a mechanism — and a programme that claims more than that will eventually meet someone who takes the claim apart. Overclaiming is how measurement loses the credibility it was built to create.
What to steal
- This month: pick two workflows — high-volume, repeatable, visible — and baseline them before your next change: time, quality proxy, adoption.
- Name one owner per number. A committee owns nothing.
- Report the delta monthly, one page, in the budget-holder's format — before they ask.
- Let the number kill weak use cases. A programme that can retire failures is one that gets believed about successes.
If your readiness report flagged unproven value, your programme is currently one budget review away from resetting to zero — the two-workflow baseline is the cheapest insurance that exists◦.RelatedWhy measurement is what keeps a rollout alive
Sources
- MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025" (via Fortune)
- McKinsey, "The State of AI", March 2025
- BCG, "Where's the Value in AI?", October 2024
- Gartner press release, 29 July 2024
- Deloitte, "State of Generative AI in the Enterprise", Wave 4, January 2025
- Lumen Technologies on Microsoft Copilot — a company estimate published through the vendor
Keep reading

Case Study: Turning Individual AI Wins Into Shared Capability
A capable, well-trained team where nothing compounded: every AI win stayed private. Five harvested workflows and one co-built assistant later, the median user caught up with the best.

Case Study: Mapping Shadow AI Without Driving It Underground
A conglomerate's official AI adoption was 12%. The real number was closer to 60% — running through personal accounts. The amnesty that surfaced it changed the whole roadmap.