Case Study: Putting a Defensible Number on AI Adoption
An AI programme everyone liked nearly lost its budget because nobody could prove it worked. Two baselined workflows later, it survived the cut — and earned an expansion.

"The team really likes it" is the sentence that kills AI budgets1. Not because it is false — because it is unpriceable, and it always loses to a line item that isn't.1A composite case, anonymised from several Fellow engagements. Numbers are representative of the pattern, not one client's audited results.
The situation
A 600-person services organisation had run AI licences and training for a year. Usage was genuinely healthy: daily activity, enthusiastic champions, anecdotes everywhere. Then the CFO asked the reasonable question — what did we get for it? — and the programme discovered it had no answer. No baseline existed for any workflow the tools had changed. The renewal went into review with sentiment as its only defence.
Their readiness profile showed the signature: solid Applied use, near-zero on impact measurement◦.By the numbers0 of 14AI-touched workflows with any before-measurement, at the point the budget question arrived
Why value goes unmeasured
Nobody plans to skip measurement; it is skipped by default. In the excitement phase, measuring feels like bureaucracy slowing the fun down. Later, the before-state is gone — nobody recorded how long a report took in the era before the tools, so the improvement became unprovable the moment it succeeded◦.DefinitionBaseline The before-measurement — time, quality, adoption — taken prior to a change, without which no after-number means anything
The cost is not hypothetical. Unmeasured programmes lose budget reviews to measured ones regardless of relative merit, and the people who quietly gained hours a week lose them back when the licences go.
What we did
We did not build a measurement programme. We baselined exactly two workflows — proposal drafting and client-meeting summaries — chosen for volume, repeatability, and how visibly they mattered to the business.
For each: two weeks of honest before-numbers (time per unit, revision counts, who used AI at which step), one named owner of the number, and a one-page monthly delta report designed in the CFO's own format. Nothing else was measured on purpose — two credible numbers beat ten estimates◦.Figure
One person owning one honest number is a measurement system; a committee with a dashboard is usually not
What changed
One quarter later, proposal first-draft time was down by roughly half against its baseline, with revision rounds flat — quality had held◦. The renewal conversation took ten minutes: the programme survived the cut that year, and the same evidence format earned an expansion into two more workflows the next.By the numbers≈50%reduction in proposal first-draft time against a two-week baseline, quality flat
The deeper change was cultural: with a delta report on the table, adoption stopped being a belief and became a managed quantity — which also exposed one workflow where AI genuinely was not helping, and it was retired without drama. Proof cuts both ways, and that is the point.
What to steal
- This month: pick two workflows — high-volume, repeatable, visible — and baseline them before your next change: time, quality proxy, adoption.
- Name one owner per number. A committee owns nothing.
- Report the delta monthly, one page, in the budget-holder's format — before they ask.
- Let the number kill weak use cases. A programme that can retire failures is one that gets believed about successes.
If your readiness report flagged unproven value, your programme is currently one budget review away from resetting to zero — the two-workflow baseline is the cheapest insurance that exists◦.RelatedWhy measurement is what keeps a rollout alive
“ทีมชอบมากเลยครับ” คือประโยคที่ฆ่างบประมาณ AI มานักต่อนักแล้ว1 ไม่ใช่เพราะมันไม่จริง แต่เพราะมันตีเป็นตัวเงินไม่ได้ และมันแพ้ให้กับรายการที่มีตัวเลขเสมอ1กรณีศึกษานี้เป็นการประกอบขึ้นจากหลายงานจริงของ Fellow ปรับข้อมูลให้ไม่ระบุตัวตน ตัวเลขสะท้อนรูปแบบที่พบซ้ำ ไม่ใช่ผลตรวจสอบของลูกค้ารายเดียว
สถานการณ์
องค์กรบริการขนาด 600 คนใช้ลิขสิทธิ์และจัดอบรม AI มาหนึ่งปี การใช้งานดีจริง: มีการใช้ทุกวัน มีแชมเปียนกระตือรือร้น มีเรื่องเล่าความสำเร็จเต็มไปหมด จนกระทั่ง CFO ถามคำถามที่สมเหตุสมผลที่สุด: เราได้อะไรกลับมา — และโครงการก็พบว่าตัวเองไม่มีคำตอบ ไม่มี baseline ของเวิร์กโฟลว์ใดเลย การต่ออายุลิขสิทธิ์เข้าสู่การทบทวนโดยมีเพียงความรู้สึกเป็นเกราะป้องกัน
โปรไฟล์ความพร้อมของพวกเขาแสดงลายเซ็นชัดเจน: Applied use แข็งแรง แต่การวัดผลกระทบเกือบเป็นศูนย์◦ตัวเลข0 จาก 14เวิร์กโฟลว์ที่ AI เข้าไปแตะต้องแล้วมีการวัดก่อนหลัง ณ วันที่คำถามงบประมาณมาถึง
ทำไมคุณค่าจึงไม่ถูกวัด
ไม่มีใครวางแผนจะข้ามการวัดผล มันถูกข้ามโดยค่าเริ่มต้น ในช่วงตื่นเต้น การวัดผลดูเหมือนระบบราชการที่มาถ่วงความสนุก พอถึงภายหลัง สภาพ“ก่อนใช้”ก็หายไปแล้ว — ไม่มีใครบันทึกไว้ว่ารายงานหนึ่งฉบับเคยใช้เวลาเท่าไหร่ ความสำเร็จจึงพิสูจน์ไม่ได้ทันทีที่มันเกิดขึ้นจริง◦นิยามBaseline การวัดสภาพก่อนเปลี่ยนแปลง — เวลา คุณภาพ การใช้งาน — ซึ่งหากไม่มี ตัวเลขหลังเปลี่ยนก็ไร้ความหมาย
ต้นทุนนี้ไม่ใช่เรื่องสมมติ โครงการที่ไม่วัดผลจะแพ้การทบทวนงบให้โครงการที่วัดผล โดยไม่เกี่ยวกับคุณค่าจริง และคนที่เคยได้เวลาคืนมาหลายชั่วโมงต่อสัปดาห์ จะเสียมันกลับไปเมื่อลิขสิทธิ์ถูกตัด
สิ่งที่เราทำ
เราไม่ได้สร้างระบบวัดผลขนาดใหญ่ เราทำ baseline แค่สองเวิร์กโฟลว์ — การร่าง proposal และการสรุปประชุมลูกค้า — เลือกจากปริมาณงาน ความทำซ้ำได้ และความสำคัญต่อธุรกิจ
แต่ละเวิร์กโฟลว์: เก็บตัวเลขก่อนใช้อย่างตรงไปตรงมาสองสัปดาห์ (เวลาต่อชิ้น จำนวนรอบแก้ ใครใช้ AI ขั้นตอนไหน) มีเจ้าของตัวเลขหนึ่งคน และรายงาน delta หน้าเดียวทุกเดือนในฟอร์แมตของ CFO เอง นอกนั้นตั้งใจไม่วัด — ตัวเลขน่าเชื่อถือสองตัว ชนะการประมาณสิบตัว◦ภาพประกอบ
คนเดียวที่เป็นเจ้าของตัวเลขที่ซื่อสัตย์หนึ่งตัว คือระบบวัดผลที่ใช้ได้จริง
สิ่งที่เปลี่ยนไป
หนึ่งไตรมาสต่อมา เวลาร่าง proposal ฉบับแรกลดลงประมาณครึ่งหนึ่งจาก baseline โดยจำนวนรอบแก้ไม่เพิ่มขึ้น — คุณภาพคงเดิม◦ การประชุมต่ออายุใช้เวลาสิบนาที โครงการรอดจากการตัดงบปีนั้น และฟอร์แมตหลักฐานเดียวกันนี้ทำให้ได้งบขยายไปอีกสองเวิร์กโฟลว์ในปีถัดมาตัวเลข≈50%เวลาร่าง proposal ฉบับแรกที่ลดลง เทียบกับ baseline สองสัปดาห์ โดยคุณภาพไม่ตก
การเปลี่ยนที่ลึกกว่าคือวัฒนธรรม: เมื่อมีรายงาน delta อยู่บนโต๊ะ การนำ AI มาใช้เลิกเป็นความเชื่อ และกลายเป็นปริมาณที่บริหารได้ — ซึ่งเผยให้เห็นด้วยว่ามีหนึ่งเวิร์กโฟลว์ที่ AI ไม่ช่วยจริง และถูกปลดระวางออกโดยไม่มีดราม่า หลักฐานตัดได้สองทาง และนั่นคือประเด็น
สิ่งที่นำไปใช้ได้เลย
- เดือนนี้: เลือกสองเวิร์กโฟลว์ — ปริมาณสูง ทำซ้ำได้ มองเห็นชัด — แล้ววัด baseline ก่อนการเปลี่ยนแปลงครั้งถัดไป: เวลา ตัวแทนคุณภาพ การใช้งาน
- ตั้งเจ้าของหนึ่งคนต่อหนึ่งตัวเลข คณะกรรมการไม่เคยเป็นเจ้าของอะไร
- รายงาน delta ทุกเดือน หน้าเดียว ในฟอร์แมตของผู้ถืองบ — ก่อนที่เขาจะถาม
- ให้ตัวเลขฆ่า use case ที่อ่อนได้ โครงการที่กล้าปลดสิ่งที่ล้มเหลว คือโครงการที่คนเชื่อเรื่องความสำเร็จ
ถ้ารายงานความพร้อมของคุณติดธงเรื่องการพิสูจน์คุณค่าไม่ได้ โครงการของคุณอยู่ห่างจากศูนย์แค่การทบทวนงบหนึ่งรอบ — baseline สองเวิร์กโฟลว์คือประกันที่ถูกที่สุดที่มี◦บทความที่เกี่ยวข้องทำไมการวัดผลคือสิ่งที่ทำให้การขยายผลอยู่รอด
Keep reading

Case Study: Turning Individual AI Wins Into Shared Capability
A capable, well-trained team where nothing compounded: every AI win stayed private. Five harvested workflows and one co-built assistant later, the median user caught up with the best.

Case Study: Mapping Shadow AI Without Driving It Underground
A conglomerate's official AI adoption was 12%. The real number was closer to 60% — running through personal accounts. The amnesty that surfaced it changed the whole roadmap.