Case Study: The Audit That Turned “Not Sure” Into a Plan
A division head answered “not sure” to seven of twelve readiness indicators. Four weeks of asking — not surveying — turned the blind spots into the year's adoption roadmap.

The most honest readiness result we see is also the most uncomfortable one: "not sure", seven times out of twelve1. It is uncomfortable because of what it actually means: decisions about AI in your area are being made every day — just not within your sight.1A composite case, anonymised from several Fellow engagements. Numbers are representative of the pattern, not one client's audited results.
The situation
A division head at a large Thai enterprise took our readiness assessment and answered "not sure" on seven of twelve indicators — tool visibility, data policy, verification, measurement, ownership among them. Her first reaction was to apologise for the result. Our reading was the opposite: she was the only leader in the cohort whose answers could be fully trusted, because everyone else had guessed◦.By the numbers7 of 12indicators answered "not sure" — the score wasn't low, it was invisible
Her position was not unusual for the level. She ran a division of roughly four hundred people across three sites, two of which she visited monthly. Everything she knew about AI in her area was what had been escalated to her: a procurement request for an enterprise licence, one complaint about an inaccurate summary, and a slide in a quarterly pack asserting that "AI adoption is in progress". None of it constituted a picture.
The five indicators she did answer confidently were the ones with an artefact behind them — a signed policy, a budget line, a vendor contract. The seven she could not answer were the ones that exist only as behaviour: what people actually do on a Tuesday. That split is the diagnostic worth keeping. Leaders can see what has been documented. They cannot see what is being done, unless somebody has made a point of looking.
An unscored assessment is not a failed assessment. It is a map of exactly where your line of sight ends — and hers ended, as it does for most senior leaders, precisely where the real work happens.
Leaders systematically underestimate this
Seven blank answers look like an individual gap. The research says it is structural — and that the leaders who feel confident are usually the ones who are furthest off.
- McKinsey's Superagency in the Workplace study, published in January 2025, asked both sides the same question1. C-suite respondents estimated that 4% of employees use generative AI for at least 30% of their daily work. The employees themselves reported 13% — roughly three times the leadership estimate◦.SourceMcKinsey, "Superagency in the workplace", 2025By the numbers4% vs 13%what the C-suite estimates about heavy gen-AI use versus what employees report, in McKinsey's January 2025 Superagency study
- The same study found that only 1% of leaders describe their company as mature in AI deployment — a rare piece of executive candour, and one that sits oddly beside the confidence with which adoption is usually reported upward.
- Microsoft's 2024 Work Trend Index found that 78% of AI users bring their own AI tools to work and that 52% are reluctant to admit using AI for their most important tasks2. The activity a leader is least likely to hear about is the activity on the work that matters most.SourceMicrosoft, 2024 Work Trend Index
- Slack's Fall 2024 Workforce Index found that 48% of desk workers would be uncomfortable telling their manager they had used AI, while workers who are comfortable disclosing are 67% more likely to use AI at all3◦. The silence is not neutral: it distorts the leader's picture and suppresses the practice at the same time.SourceSlack (Salesforce), Fall 2024 Workforce IndexBy the numbers48%of desk workers would be uncomfortable telling their manager they used AI, in Slack's Fall 2024 Workforce Index — and those who are comfortable are 67% more likely to use it
Gartner's 2025 survey of 302 cybersecurity leaders puts a number on the resulting exposure: 69% either suspect or have evidence that employees are using prohibited public generative-AI tools, and Gartner expects more than 40% of organisations to face a shadow-AI security or compliance incident by 20304. Its prescribed response is not stronger prohibition but regular audits2 — which is, in effect, an endorsement of the four weeks described below.SourceGartner on shadow AI, 2025 (via ITPro)2The Gartner figures here were confirmed through attributed secondary coverage rather than the primary release, so treat them as directional rather than exact.
The gap widens as the technology gets more autonomous. Deloitte's State of AI in the Enterprise finds only one in five companies has a mature governance model for autonomous AI agents — systems that act rather than answer5. A leader who cannot describe what today's assistants are doing has no foundation on which to govern something that takes actions unsupervised.SourceDeloitte, State of AI in the Enterprise, 2026 edition
The cost of not looking is not hypothetical. In Moffatt v. Air Canada (2024 BCCRT 149), a Canadian tribunal held the airline liable for a bereavement-fare policy its website chatbot had invented, and rejected the argument that the chatbot was a separate legal entity answerable for its own statements6. The sum was small — CA$812.02 — and the precedent is not: an organisation owns what its AI says, whether or not anyone is checking. Verification was one of the seven indicators the division head could not answer.SourceMoffatt v. Air Canada, 2024 BCCRT 149 (via McCarthy Tétrault)
Why guessing is worse than not knowing
A leader who guesses "we're probably fine" converts a visibility problem into a false confidence problem◦. Budgets get set, policies get written, and risk gets accepted on a picture that is part data, part folklore. "Not sure", said out loud, is the first governance act — it makes the gap workable.DefinitionBlind-spot risk Usage, spend and exposure running in the unobserved parts of a scope — unmanageable not because it is hidden deliberately, but because nobody is looking
Guessing also compounds. The guess gets written into a board slide, the slide gets quoted in a budget submission, and by the third repetition the number has acquired a provenance it never had. We have watched an invented adoption percentage survive two planning cycles because nobody could remember where it came from and everyone assumed somebody else had measured it.
Underneath all of it is a measurement vacuum. McKinsey's March 2025 survey of 1,491 organisations found that fewer than one in five track well-defined KPIs for their generative-AI solutions7◦. Where there is no instrument, the confident answer and the honest answer cost exactly the same to give — which is why so many leaders give the confident one.SourceMcKinsey, "The State of AI", March 2025By the numbers<1 in 5organisations track well-defined KPIs for their gen-AI work, in McKinsey's March 2025 survey of 1,491 organisations
The four-week audit
We designed the lightest audit that would close her specific blind spots — asking, not surveying:
- Week 1 — access and tools. With IT: which AI tools have licences, who activated them, what browser tools appear in network logs. Desk research, no interviews.
- Weeks 2–3 — fifteen conversations. Thirty minutes each with the people doing the work, selected across levels, with one script: show me where AI touches your week. No judgement, no forms◦.Figure
Fifteen conversations produced what two years of dashboards had not: an accurate picture - Week 4 — the one-page map. Tools in real use (thirteen, against four officially known), the five workflows that mattered, where data was actually flowing, and — for every practice found — a name: who runs this, who should own it.
The desk week was cheaper than anyone expected and more revealing than anyone wanted. About six hours of one IT analyst's time produced a list of tools nobody in the division had ever seen written down. Two of them were being paid for on individual expense claims, coded as "software subscription", which is how they had stayed invisible to a procurement review for eleven months.
The conversations were the expensive half — fifteen half-hour sessions plus scheduling and write-up came to roughly two person-weeks — and they nearly failed. The first three produced almost nothing. People described the approved tools, in the register of a performance review. What changed the fourth conversation was putting the notebook away and asking a different question: not "do you use AI?" but "show me the last thing you did that took less time than it used to". The screen came out. After that we opened every session by sharing something small and unglamorous from our own week, and the sessions ran over rather than short.
There was one moment when the audit almost became the thing it was designed not to be. Midway through week two, a senior manager asked for the interviewee list and what each person had said. Handing it over would have ended the audit — the fifth conversation would have been the last useful one. The division head refused, and put her reasoning in writing to her leadership team: the audit reports patterns and workflows, never names attached to behaviour, and the only names on the final map are owners who agreed to be there. That note did more for the quality of the remaining eleven conversations than any assurance we could have offered.
The one-page constraint on week four was deliberate, and it hurt. The working draft ran to nine pages. Everything below the top five workflows was true, and none of it would have changed a decision — a thirty-page report would have been read by nobody and actioned by nobody. Cutting to a single page forced a ranking, and the ranking is what made the document a roadmap rather than an inventory.
Each of her original "not sure" answers became a line on that page with a finding and an owner attached.
What changed
The map became the division's adoption roadmap for the year: two shadow workflows promoted into supported ones, one genuine data exposure closed within a fortnight of being seen, and an adoption owner named — the first structural fix, because permanent visibility needs a person, not an annual audit. On re-assessment a quarter later, eleven of twelve indicators were answerable, and the score that emerged was lower than her peers' guesses had been — and, unlike theirs, true◦.By the numbers13 vs 4AI tools actually in use versus officially known — found in four weeks of asking
The data exposure was the uncomfortable finding. A team had been pasting customer records into a consumer AI account to reformat them — a habit roughly eight months old, invented by somebody trying to be helpful, and never flagged because no rule obviously covered it. Closing it took a fortnight and cost nobody their job: a supported tool for the same task, a data rule written in a single paragraph, and a conversation rather than a disciplinary process. Cyberhaven Labs' 2026 AI Adoption and Risk Report, which measures real enterprise data flows rather than asking people, found that 39.7% of data movements into AI tools involve sensitive data and that the average employee enters sensitive data into an AI tool every three days8◦. This team was ordinary, not exceptional.SourceCyberhaven Labs, 2026 AI Adoption & Risk ReportBy the numbersevery 3 dayshow often the average employee enters sensitive data into an AI tool, in Cyberhaven Labs' 2026 measurement of enterprise data flows
The owner appointment was the fix with the longest half-life, and the one that took the most arguing. It was resisted as "another hat on an already full head" until the role was scoped honestly: two hours a week, a standing list, and the authority to say which workflows get supported next. Four weeks buys you an accurate picture with a shelf life of about two quarters. A named owner with two hours a week keeps it accurate indefinitely — which is the whole difference between running an audit and having a capability.
What we would do differently
Three things, in order of what they cost us.
The audit measured presence, not value. We came out of week four able to name thirteen tools, five workflows and a data exposure — and unable to say whether any of it had saved a single hour. Nobody had baselined anything before the AI arrived, and by the time we thought to ask, the "before" was a year old and remembered rather than recorded. A quarter later the division could describe its AI use with real precision and still could not defend its AI budget. Run again, week one would include picking two workflows and capturing a crude baseline — how long, how many, how often — because a rough number recorded beats a precise number reconstructed from memory◦.RelatedWhy the measurement has to start before the tool does
Fifteen conversations across a division of four hundred is a sample, and a self-selecting one. The people who agreed to a thirty-minute session about their working habits were, on the whole, people comfortable talking about their working habits. We are reasonably confident the map caught every workflow that more than a handful of people were running. We are not at all confident it caught the quiet, individual, slightly embarrassed uses — which, on the evidence above, is exactly where a disproportionate share of the risk sits.
And four weeks produces a photograph, not a feed. The map was accurate in month one, roughly accurate in month four, and would have been actively misleading by month twelve had the owner not kept it current. The most common failure we see after an audit like this is not a bad map. It is a good map that nobody updates, still being quoted with confidence eighteen months later — which is precisely the guessing problem the audit was run to fix.
What to steal
- This week: list your own "not sure" answers — from a readiness report or honest reflection. Each one is a question someone in your organisation can answer.
- Before you start asking, pick two workflows and write down a crude baseline — how long the task takes now, how often it happens. You will not get another chance at the "before".
- Weeks 1–2: do the desk half — licences, activations, network-visible tools, and the expense lines coded as software. It is faster than it sounds.
- Weeks 2–4: hold fifteen show-me conversations across levels. Conversations, not surveys — forms collect what people think you want to hear. Promise, in writing, that no names are attached to behaviour, and keep that promise when someone senior asks.
- Then: write the one-page map with a name against every finding, and appoint the owner who keeps it current. One page, ranked — the ranking is the roadmap.
If your readiness result was mostly unknowns, do not retake the assessment hoping for a better guess — run the audit, then measure what you can finally see◦.RelatedWhat the asking usually surfaces: the shadow-AI map
Sources
- McKinsey, "Superagency in the workplace", 2025
- Microsoft, 2024 Work Trend Index
- Slack (Salesforce), Fall 2024 Workforce Index
- Gartner on shadow AI, 2025 (via ITPro)
- Deloitte, State of AI in the Enterprise, 2026 edition
- Moffatt v. Air Canada, 2024 BCCRT 149 (via McCarthy Tétrault)
- McKinsey, "The State of AI", March 2025
- Cyberhaven Labs, 2026 AI Adoption & Risk Report
Keep reading

Case Study: Turning Individual AI Wins Into Shared Capability
A capable, well-trained team where nothing compounded: every AI win stayed private. Five harvested workflows and one co-built assistant later, the median user caught up with the best.

Case Study: What Changed When AI Adoption Got a Named Owner
A retailer's AI momentum collapsed the month its champion resigned. The rebuild took one owner, three workflows, one measure each — and a monthly review leadership actually attends.