← All insights
8 July 2026
Case StudyAI Adoption

Case Study: Cutting Report First-Draft Time by 60%

How a 400-person professional services firm turned a scattered set of AI experiments into a measured, repeatable workflow.

A professional services firm of about 400 people had the usual picture: a handful of power users doing impressive things with AI, and everyone else watching.1 Report writing — a core, time-heavy deliverable — was still done the old way by most of the team.1A composite case, anonymised and simplified from several Fellow engagements. The numbers are representative of what we measure in this pattern, not one client's audited results.

The partner who called us framed it as a training problem. Twelve months and two vendor workshops later, the firm had a licence for everyone and a usage curve that looked like a single spike in week one.

The situation

Three things were true at once, and they are worth separating because they usually get collapsed into "adoption is slow".

People understood the tools. Nobody needed convincing that a language model could draft a section of a report. The scepticism was narrower and better founded than that: they did not believe an AI-drafted section could pass the firm's review without costing more time in editing than it saved in drafting.

The power users had genuinely good workflows, and none of them were written down. Each of the four or five people doing impressive work had built a personal set of prompts and habits in their own notes app. Nothing was shared, so nothing compounded.RelatedTurning individual wins into shared capability

And report quality was the firm's product. This is the constraint that makes professional services different from most AI-adoption stories: a bad deliverable is not an inconvenience, it is the thing the client is paying not to receive.

This is not a local problem

Report drafting is one of the few AI use cases with strong experimental evidence behind it, which is a large part of why we chose it.

Noy and Zhang's 2023 study in Science gave 453 professionals mid-level writing tasks and found time to completion fell about 40% while output quality rose about 18%1. It also found something more useful for a rollout: the gap between stronger and weaker writers narrowed.2SourceNoy & Zhang, Science, 2023 (via MIT News)By the numbers−40% / +18%writing time and quality change across 453 professionals — Noy & Zhang, Science, 20232These were lab-style, incentivised tasks rather than real client deliverables, so read the direction as robust and the magnitude as generous.

Dell'Acqua and colleagues, in the 2023 jagged-frontier study run with 758 BCG consultants, found tasks completed 25.1% faster and over 40% higher in quality inside AI's capability — and that below-average performers gained 43% against 17% for above-average ones2. Brynjolfsson, Li and Raymond's NBER study of more than 5,000 support agents found the same shape: +14% overall, +34% for novices, and close to nothing for the most experienced3.SourceDell'Acqua et al., "Navigating the Jagged Technological Frontier", 2023SourceBrynjolfsson, Li & Raymond, "Generative AI at Work", NBER WP 31161, 2023

The consistent finding across all three is that this kind of work compresses most for the people currently slowest at it. That is an argument for aiming a drafting workflow at the middle of the firm, not at the partners.FigureThe gains concentrate where the drafting is slowest — the middle of the firm, not the top.The gains concentrate where the drafting is slowest — the middle of the firm, not the top.

The counter-evidence matters just as much, and it comes from exactly this industry. In October 2025, Deloitte Australia refunded part of a A$440,000 contract with the Department of Employment and Workplace Relations after a delivered report was found to contain AI-hallucinated citations, including a fabricated quote attributed to a federal court judgment4. In Mata v. Avianca in 2023, two New York lawyers and their firm were sanctioned $5,000 for a filing built on ChatGPT-invented case citations5.SourceDeloitte Australia's partial refund on an AI-error report, 2025 (via CFO Dive)SourceMata v. Avianca, Inc., 2023

Both are the same failure: AI-drafted text entering a client deliverable without a review step that was actually designed to catch invention. The firm's sceptics were right about the risk. They were wrong that the answer was to avoid the tool.

Why more training was the wrong move

The instinct after a flat adoption curve is to run better training. We have watched this fail often enough to name why.

Generic training teaches the tool, and the barrier here was not the tool. Microsoft's 2024 Work Trend Index found only 39% of AI users had received any company training6 — but BCG's 2025 study, across 10,600 employees, found that even among the trained, only about a third felt properly prepared7. More hours of the same content does not convert people whose actual objection is "this will not survive review".SourceMicrosoft, 2024 Work Trend IndexSourceBCG, "AI at Work 2025"

The literacy assessment made this concrete. Foundations scored high; workflow integration scored low; judgment scored in the middle with a wide spread. In other words: they knew what the tools were, some of them knew when to distrust output, and almost nobody had woven either into a deliverable. A curriculum addressing foundations would have taught the firm what it already knew.RelatedHow that assessment works

Report drafting was the obvious first target — high volume, high time cost, consistent structure, and a quality bar that already existed in written form.

What we did, in order

Six weeks, roughly five person-weeks of Fellow time and about the same again from the firm.

Week 1: baseline, honestly. We timed first drafts across three practice groups before touching anything. This was the least popular part of the engagement — timing your own work feels like surveillance, and two of the three groups pushed back hard. What resolved it was reporting by group rather than by person, and saying so before collecting anything. Without that baseline we would have had no defensible number at the end, and the programme would have been re-litigated at the next budget round.DefinitionBaseline The before-measurement you compare results against. Without it, every later claim about improvement is an assertion.

Weeks 2–3: co-built the prompt library. Not written for the teams — written with them, in sessions where a partner brought a real report and we built the prompts against it live. The library was tuned to the firm's actual report formats and house tone. The unglamorous detail that mattered most: it lived where people already worked, not in a new tool that required a new login.DefinitionPrompt library A shared, maintained set of prompts tuned to the firm's formats and tone, owned by a named person.

Week 3: the review checklist, which nearly did not ship. This was the contested piece. The drafting side of the workflow was popular immediately; a mandatory check on AI-drafted sections read to several people as a statement that their judgment was not trusted. It survived because we scoped it to the two failure modes that actually appear — invented citations and confidently wrong specifics — rather than as a general quality gate, and because a partner rather than a consultant introduced it.

In hindsight this was the most important component in the engagement, and we nearly traded it away to keep the room happy.

Weeks 4–5: named owners. One person per practice group, accountable for maintaining the library and coaching peers, with the role written into their objectives and a named backup. One of the three original owners had moved off the role within four months; the backup is why that did not show up as a decline.FigureThe workflow lived with the practice groups, not the consultants.The workflow lived with the practice groups, not the consultants.RelatedWhat changed when adoption got a named owner

Week 6: instrumented and handed over. A monthly read on real usage and draft times, owned by the firm's operations lead — deliberately someone whose job was not AI.

What changed in one quarter

Average first-draft time for standard reports fell by roughly 60%.By the numbers≈60%drop in average first-draft time over one quarter — not total time to a client-ready report

Two qualifications belong immediately next to that number. It is first-draft time, not total cycle time: review and partner sign-off did not compress nearly as much, and total time to a client-ready report fell by closer to a quarter. And it is a composite figure across engagements showing this pattern, not one firm's audited result.

Adoption held, which matters more than the percentage.3 Three months after the engagement ended, usage had not decayed. The reason is structural rather than cultural: the workflow belonged to the practice groups, the library had a maintainer, and the monthly read had an owner who was still there.3"Held" is the whole point — the workflow belonged to the teams, so it survived our leaving. Most of what we are asked to fix is a decayed win from a previous programme.

The review checklist caught two invented citations in the first quarter. Neither reached a client. That is the number the partners actually talk about.

The lesson repeats across every engagement: adoption follows a specific, owned workflow far more reliably than it follows a good tool.

What we would do differently

We would build the review step first. We built drafting first because it was the easy win and it bought us credibility. But for a firm whose product is the deliverable, the check is the load-bearing part, and treating it as a phase-two concession nearly cost us it entirely. Start where the risk is and earn the drafting gains second.

We under-invested in the jagged frontier. The library made people faster on report sections that resemble other report sections. It gave them nothing for the genuinely novel engagement — and that is exactly where Dell'Acqua's finding bites: on tasks outside AI's capability, consultants using it were 19 percentage points less likely to be correct. A library of good prompts quietly teaches people that the tool is reliable. We should have paired it with explicit guidance on when to close it.By the numbers−19pphow much less likely AI-assisted consultants were to be correct on a task outside AI's capability — Dell'Acqua et al., 2023

Sixty per cent was the wrong headline internally. It travelled well and it made the programme feel finished. A quarter later, the firm's attention had moved on while review time — now the actual constraint — went unexamined. If we ran it again, the reported metric would be total time to a client-ready report from the start, even though the number would have looked far less impressive.

The prompt library needed a review date, not just an owner. Models change; a prompt tuned to one model's habits ages into a workaround for a problem that no longer exists. Ownership keeps a library alive. Only a scheduled review keeps it correct.

What to steal

  • Baseline before you build, report by group, and say so before you collect. No baseline, no defensible number.
  • Pick a deliverable with high volume, consistent structure, and a quality bar that already exists in writing.
  • Build the review step for the two failure modes that actually occur — invented sources and confident specifics — not as a general quality gate. Narrow checks get used.
  • Have a partner introduce the check, not a consultant.
  • Name an owner per group, write it into their objectives, and name a backup. You will need the backup.
  • Report the metric that reflects the real constraint, even when a flattering one is available.
  • Give the library a review date. An owned artefact still goes stale.RelatedWhat a rollout system has to contain to survive

Sources

  1. Noy & Zhang, Science, 2023 (via MIT News)
  2. Dell'Acqua et al., "Navigating the Jagged Technological Frontier", 2023
  3. Brynjolfsson, Li & Raymond, "Generative AI at Work", NBER WP 31161, 2023
  4. Deloitte Australia's partial refund on an AI-error report, 2025 (via CFO Dive)
  5. Mata v. Avianca, Inc., 2023
  6. Microsoft, 2024 Work Trend Index
  7. BCG, "AI at Work 2025"

Keep reading

Want this in your organisation?

Talk to our team