Case Study: Turning Individual AI Wins Into Shared Capability
A capable, well-trained team where nothing compounded: every AI win stayed private. Five harvested workflows and one co-built assistant later, the median user caught up with the best.

There is a readiness profile that looks like success and behaves like a plateau: strong foundations, capable people, real daily usage — and nothing compounding1. The organisation is paying enterprise prices for individual productivity.1A composite case, anonymised from several Fellow engagements. Numbers are representative of the pattern, not one client's audited results.
The situation
A professional-services team of about 300 had done the hard part: licences everywhere, training completed, genuine enthusiasm. Yet output gains were confined to a handful of power users. The assessment showed the signature gap — high Foundations and Applied use, low Workflow integration — and one number told the whole story: the gap between the best user's productivity and the median user's was wide, and widening◦.By the numbers5×difference in AI-assisted output between the team's best user and its median, at engagement start
Every win was private. Prompts lived in personal chat histories. Techniques spread by lunch conversation, or not at all. When a power user went on leave, their throughput went with them.
The working hypothesis when we arrived was that the answer was more training — a second, deeper wave of the same course, already budgeted. It was the wrong diagnosis drawn from the right instinct. The people who had gained least were not the people who knew least; several of them had scored well on the assessment and could describe good practice fluently in the abstract. What they lacked was not knowledge of AI. It was access to what the colleague two desks away had already worked out.
Why the gap widens on its own
This is not a local pathology, and the research is unusually specific both about why the gap opens and about the one mechanism that closes it.
The finding this whole case rests on comes from Brynjolfsson, Li and Raymond's 2023 NBER working paper, number 31161, which followed more than 5,000 customer-support agents given a generative-AI assistant1. On average, agents resolved 14% more issues per hour. The average concealed the real result: novice and lower-skilled agents improved by about a third, while experienced high performers gained little or nothing◦. The authors' explanation is the sentence this article exists to act on: the assistant appeared to disseminate the best practices of more able workers. It had been shaped on the organisation's own successful conversations, so what a novice received was, in effect, the tacit expertise of the firm's best people, arriving at the moment of use.SourceBrynjolfsson, Li & Raymond, "Generative AI at Work", NBER WP 31161, 2023By the numbers+34%the productivity gain for novice agents in Brynjolfsson, Li and Raymond's 2023 NBER study, against 14% on average and near-zero for the most experienced
That mechanism deserves to be stated carefully. Brynjolfsson, Li and Raymond present it as suggestive evidence, not a proven causal claim, and we treat it the same way. But it names precisely what a harvest does by hand. In their setting a model carried the best practice invisibly; in ours we wrote it down, gave it a name, and gave it an owner. The technology is different. The mechanism is the same one, and it is the only mechanism in this literature that raises a whole distribution rather than a few individuals.
The levelling effect replicates. In the 2023 Harvard, BCG, Wharton and MIT study led by Dell'Acqua, 758 BCG consultants using AI completed 12.2% more tasks, 25.1% faster, at more than 40% higher quality — and consultants who had started below average gained 43% against 17% for those already above it2. Noy and Zhang, publishing in Science in 2023, ran 453 professionals through mid-level writing tasks and found time down 40%, quality up 18%, and inequality between workers decreased; their caveat, which we keep, is that these were lab-style incentivised tasks rather than live client work3.SourceDell'Acqua et al., "Navigating the Jagged Technological Frontier", 2023SourceNoy & Zhang, Science, 2023 (via MIT News)
So why did our client's gap widen instead of narrow? Because in every one of those studies the best practice was handed to people. Nobody had to go and find it. Left to distribute itself, capability does the opposite: Microsoft's 2024 Work Trend Index found power users saving more than 30 minutes a day where sceptics saved 10 or fewer, while only 39% of AI users had received any training at all from their employer4◦. BCG's AI at Work 2025 survey found frontline regular use at 51% against more than 75% among leaders and managers, with only about a third of employees saying they had been properly trained5. By the 2026 edition frontline regular use had reached 74%6 — the access gap closed inside a year while the training gap did not, which is exactly how an organisation ends up with far more people using AI and no more shared practice than it had before5. And PwC's 2025 Global AI Jobs Barometer put the wage premium on AI skills at 56%, double the previous year's 25% — an industry-level correlational analysis rather than a measurement of individuals, but an unambiguous signal about which way the incentives now point7. The gap is not a temporary artefact of an early rollout. It pays.SourceMicrosoft, 2024 Work Trend IndexBy the numbers39%the share of AI users who had received any company training, in Microsoft's 2024 Work Trend IndexSourceBCG, "AI at Work 2025"SourceBCG, "AI at Work 2026: Strategy Matters More Than Tools", June 2026 (via BCG press release)SourceBCG, "AI at Work 2025"SourcePwC, 2025 Global AI Jobs Barometer
One further finding from the BCG consultant study belongs here rather than quietly in a footnote. On a task the researchers had deliberately designed to sit outside AI's capability, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. The same tool that lifted the bottom of the distribution on suitable work walked people confidently past the right answer on unsuitable work. Hold that; we come back to it at the end, because it is the sharpest limitation on everything we did.
Why capability doesn't compound on its own
Individual skill decays and departs; only shared assets accumulate◦. Without a place for practices to live, every person re-solves the same problems: the team had, we counted, eleven separately-invented prompt variants for the same client-letter task — none of them the best one.DefinitionCompounding When this quarter's improvements become the floor for next quarter — requiring assets that persist beyond the individuals who created them
We found the eleven by asking, in a single message, for anyone who drafted client letters to paste what they currently used into one document. It took an afternoon. The revealing part was not that eleven versions existed. It was that not one person in the team, including its best drafter, had ever seen another person's.
There had also been an attempt at a fix before we arrived: a "share your prompts" channel, opened with enthusiasm, busy for a fortnight, silent from week six. It failed the way those channels usually fail. Posting was voluntary, nothing was curated, nothing was maintained, and using anything in it required believing that a stranger's undocumented prompt would survive contact with your own task. A channel is a place to put things. It is not an asset.
Training alone cannot fix this either. Training raises individuals; compounding requires infrastructure — however light. The design constraint is that the infrastructure has to stay small enough that maintaining it is nobody's second job.
What we did
The whole intervention cost about eighteen person-days of concentrated effort spread across a quarter: six running and writing up the harvest, five co-building the assistant, four coaching the new owners through the first month, and the remainder in fortnightly reviews. On top of that sat the two workshop days the room itself gave up. The number is worth saying plainly, because the usual objection to harvesting is that it sounds like a programme. It is not a programme. It cost less time than the training the team had already completed.
- Harvested the top five practices. Two workshop days: power users demonstrated their real workflows on live work, the team voted on the five with widest application, and each was written up as a shared, named asset — prompt, checklist, worked example, and a section on when not to use it.
- Gave each asset an owner. Not the original inventor by default — the person whose role most depended on the workflow. Owners maintain, coach, and retire their asset◦.RelatedWhy a named owner is what makes adoption stick
- Co-built one assistant for the heaviest recurring workflow — client-letter drafting — with the people who own that work, not for them. Building it together is what made it theirs; assistants delivered as gifts get admired and abandoned◦.Figure
Co-building is the transfer mechanism: the team that builds the assistant is the team that keeps it - Changed the definition of done. Any AI win worth mentioning in a team meeting now ends with the same question: where does it live now? Sharing became a norm with a place to put things, instead of a value statement.
Inside the two-day harvest
Day one was demonstrations. Eight power users, forty minutes each, screens shared, working a live task rather than describing one. The single rule was that nothing could be prepared — no cleaned-up prompt, no rehearsed example. We wanted the real session, including the false starts, because the false starts are where the judgement is visible.
Day two was selection and writing. Twenty-two people, deliberately weighted towards consumers rather than inventors, scored every demonstrated practice on two axes: how often would you personally use this, and how confident are you that you could use it tomorrow without help◦. The second axis did most of the work. Two of the most impressive practices scored badly on it and were cut. The five that survived were not the cleverest; they were the most transferable.Figure
Day two scores each practice on transferability rather than cleverness — the second axis is what does the work
Each surviving practice was then written up in the room, by a pair: the power user who invented it, and the person who had scored it lowest on confidence.
The sceptic held the pen. If the sceptic could not write the practice down clearly enough to use it unaided, it was not finished.
That pairing did more for the quality of the assets than any template we could have handed out.
What the power users could not say
The first hour of day one nearly ended the exercise. Asked to explain what they were doing, the best users produced generalities. "I just give it a lot of context." "I iterate." One said, with complete honesty, that he did not think he did anything special. This is the normal condition of expertise rather than evasion◦.DefinitionTacit knowledge Skill a person performs reliably but cannot readily put into words — the usual state of genuine expertise
What worked was to stop asking for explanations and start watching, then interrogate the moments of judgement. The most productive question by a distance was: why did you reject that draft? Every rejection is a criterion, and the criterion is the asset. In one forty-minute session, a partner who claimed to do nothing special rejected six drafts for six different reasons — the tone was too eager; the second paragraph conceded a point she had not yet conceded; a figure appeared without its source; the closing asked for two things where one would land. None of that was in her prompt. All of it went into the checklist.
Two other questions earned their place. What do you always paste in before you start? — which surfaced context files and precedent bundles nobody else knew existed. And what would you never use this for? — which produced the "when not to use it" section of every asset, and in hindsight prevented more harm than the prompts prevented rework.
The resistance, and how we handled it
Three objections came up, in roughly this order, with rather different amounts of honesty in them.
- "This is just how I work." The most common and the least hostile: a genuine report of tacit skill, not a refusal. We handled it by not arguing. Nobody was asked to write anything on day one — only to work while being watched.
- The quiet fear of losing an edge. Rarely said aloud, and entirely reasonable. In a firm that ranks people, being several times more productive than your peers is worth something real, and we were asking people to give it away. Reassurance would not have touched this, so we changed the structure instead: contributors were named on the assets they seeded, participation in the harvest went into performance conversations, and the two heaviest contributors were given the first two owner roles. The edge was not confiscated; it was converted into a different kind of standing.
- Seniority discomfort. Several senior staff disliked being taught by an analyst a decade their junior, which is exactly who two of the best users were. We changed the frame rather than the facts: the asset carries the practice, not the person. Nobody was taught by anybody — they were handed a checklist with a name on it. That sounds cosmetic. It was not. The same content, delivered as a document rather than a lesson, met almost no resistance at all.
What nearly failed
The assistant. We built the first version in a week, with two people who were good at building and were not the people who write the letters. It drafted better than the average human first pass, and almost nobody used it. We rebuilt it over three working sessions with four actual letter-writers arguing about wording, refusing two of our defaults, and insisting on a step we thought was redundant. By our own measure the second version drafted slightly worse than the first. It got used, because they had argued about it.
The fortnightly reviews nearly went the same way. They were skipped twice in month two, and the pause was enough for two owners to quietly stop updating anything. Restarting took a direct conversation rather than a calendar invitation. Half an hour a fortnight sounds too small to be load-bearing. It is the whole load.
What changed
One quarter later the best-to-median gap had halved — not by slowing the best users but by raising the middle◦. As in the research, the distribution matters more than the average: the people who moved furthest were the ones who had moved least before.By the numbers5× → ≈2×best-to-median output gap after one quarter of shared assets and one co-built assistant
New joiners reached working proficiency in days instead of months, because the assets were the onboarding. A new starter no longer has to reconstruct the firm's judgement from scratch; they inherit six months of it in an afternoon, complete with the reasons behind it.
The original power users, freed from being the team's informal help desk, went back to inventing — now with somewhere for the inventions to land. Two of the five harvested assets have since been replaced by better versions, and both replacements were proposed by people who had not demonstrated anything on day one. That is what compounding looks like in practice: the floor moved, and the new floor started producing its own improvements.
Not all five worked. One — a research-synthesis practice — has been used by essentially nobody outside its owner. It was voted in for being impressive rather than transferable, which is exactly the failure mode day two was designed to catch and only partly caught. We left it in this account because a harvest that produces five winners is not a harvest; it is a brochure.
The part that is harder than it sounds
Two things about this work are consistently underestimated, and one of them is a genuine risk rather than an inconvenience.
Harvested assets decay without a named maintainer. An asset is not a document; it is a standing commitment to keep a document true. Models change, clients change, the regulation the checklist encodes changes, and a prompt that was excellent in March is quietly misleading by September. Ours drifted within a quarter. The fix is unglamorous: a review date written on the asset itself, a standing fifteen minutes a fortnight, and an explicit retirement path — anything nobody has used in two months should be deleted rather than tidied. Deleting is the part organisations find hardest, and the part that keeps the collection trustworthy.
A shared prompt spreads a blind spot exactly as efficiently as it spreads a best practice. This is the Dell'Acqua finding coming back: on work sitting outside the model's competence, AI users were 19 percentage points less likely to be right, and confident while being wrong. Harvesting scales whatever it captures. If the practice we wrote up embeds a bad assumption — a compliance step skipped, a client nuance flattened — we have now distributed that error to 300 people with an internal endorsement attached to it, which makes it harder to question than the eleven private prompts ever were. So an asset needs a review date, not just an owner2. And its "when not to use it" section has to be treated as the most important part of the document rather than the appendix.2A review date is the only thing that reliably separates a shared asset from a shared assumption.
The measure itself is also cruder than it looks. Best-to-median output gap is a useful management number and a poor scientific one; it moves when the median improves and it also moves when the best users take on harder work. We report it because the team acted on it, not because it would survive a referee.
What to steal
- This month: run a two-day harvest — power users demo live work, the team votes on transferability rather than cleverness, and the top five practices become named shared assets.
- Ask "why did you reject that draft?" until the criteria come out. Experts cannot explain what they do; they can always explain what they refused.
- Let the sceptic hold the pen. If the least confident person cannot write the practice down usably, it is not finished.
- Assign each asset an owner whose role depends on it. Inventorship is not ownership.
- Give every asset a review date and a retirement rule, not just an owner — and treat its "when not to use it" section as the important half.
- This quarter: co-build one assistant for your heaviest recurring workflow, with its owners in the room and permission to overrule you.
- Add "where does it live now?" to every AI win discussion. A win without an address is a story, not an asset.
Capable people with nothing compounding is momentum quietly expiring — the harvest is how you bank it before it does◦.RelatedA full worked example: one owned workflow, measured end to end
Sources
- Brynjolfsson, Li & Raymond, "Generative AI at Work", NBER WP 31161, 2023
- Dell'Acqua et al., "Navigating the Jagged Technological Frontier", 2023
- Noy & Zhang, Science, 2023 (via MIT News)
- Microsoft, 2024 Work Trend Index
- BCG, "AI at Work 2025"
- BCG, "AI at Work 2026: Strategy Matters More Than Tools", June 2026 (via BCG press release)
- PwC, 2025 Global AI Jobs Barometer
Keep reading

Case Study: Putting a Defensible Number on AI Adoption
An AI programme everyone liked nearly lost its budget because nobody could prove it worked. Two baselined workflows later, it survived the cut — and earned an expansion.

Case Study: Mapping Shadow AI Without Driving It Underground
A conglomerate's official AI adoption was 12%. The real number was closer to 60% — running through personal accounts. The amnesty that surfaced it changed the whole roadmap.