Skip to module content
Module 10 ยท ~8 min

Measuring What AI Is Worth

The ROI question destroys more AI programs than any technical failure.

Reading progress
0/7 ยท 0%

The big idea

๐Ÿ’กKey idea
Measure like the research says value actually arrives: system โ†’ workflow โ†’ business, leading before lagging, baseline before build, learning asset counted, verdict time-boxed. An engagement instrumented this way can prove itself at every tier โ€” which is also, not coincidentally, the evidence pack for your next case study.
Quick check
1 question ยท instant feedback
0/1
  1. 'Value is real, early, diffuse, and compounding; P&L attribution is late, precise, and linear.' The consequence:

Deep dive

1/1 open

The 'missing ROI' problem isn't really a results problem โ€” it's a measurement design problem. Once you understand how value from AI actually arrives, the mystery mostly dissolves.

MIT research found that P&L-level attribution fails for roughly 95% of AI deployments โ€” not because the value isn't there, but because it's diffuse and arrives gradually across many processes rather than in one measurable line item. Duke's CFO survey confirms that finance chiefs expect AI to deliver productivity gains, faster decisions, and satisfaction improvements โ€” but they don't expect to see measurable cost savings or headcount reductions in the near term.

The agentic-enterprise research adds a structural explanation: agents get better over time (appreciating through accumulated learning) while also degrading if not maintained (depreciating through model drift). Conventional financial models, which treat value as either a one-time asset purchase or a recurring employee cost, aren't designed to capture that compounding dynamic โ€” so they systematically undervalue the return.

And shadow-AI findings show real value accruing in channels no official metric is watching at all. The pattern is consistent: value is real, arrives early, spreads across many people and workflows, and compounds over time. P&L attribution is late, precise, and linear. Programs get killed in the gap between those two realities โ€” which is exactly what the three-tier measurement framework is designed to bridge.

Quick check
1 question ยท instant feedback
0/1
  1. The three metric tiers to report separately:

How to run it

  1. 1 ยท Baseline before build โ€” always
    Cycle time, error rate, unit cost, volume, satisfaction for the target workflow, captured during discovery. Without a baseline there is no honest ROI claim.
  2. 2 ยท Three metric tiers, reported separately
    System metrics (eval pass rates, cost per task, latency) prove it works. Workflow metrics (cycle time, throughput, error rate, rework vs baseline) prove it changed the process. Business metrics (cost per unit, revenue effect, capacity released) prove it mattered โ€” reported with honest attribution ranges.
  3. 3 ยท Leading indicators first
    Adoption rate by process owners, eval-verified quality, human-edit rates trending down. Morgan Stanley's 98% daily adoption and 20%โ†’80% document access are exactly this tier โ€” OpenAI's flagship case leads with them, not P&L.
  4. 4 ยท Count the learning asset
    Every reviewer correction captured, every golden-set case added, every workflow instrumented is proprietary data a competitor cannot buy. Report as an asset line: '214 validated edge cases; error taxonomy covering 96% of observed failures.'
  5. 5 ยท Time-box the verdict
    Pre-agree the review point ('workflow metrics at 90 days decide scale-up') so the program is judged by design rather than by nervousness.
Quick check
1 question ยท instant feedback
0/1
  1. The 'learning asset' shows up on a report as:

In the field

๐Ÿ”ฌWorked example
Second tab in every ROI model: the measurement plan โ€” baseline fields captured in discovery, the three metric tiers with owners and cadence, and the 90-day verdict criteria. Selling the measurement plan alongside the ROI model converts your weakest moment (defending projections) into your strongest (proposing how you'll be held accountable). Almost nobody selling AI services does this; every CFO notices.
๐ŸšซWhen not to reach for it
Resist inflated hours-saved math (everyone's '2 hours/week' ร— loaded cost ร— 52) โ€” CFOs discount it on sight; your sensitivity-analysis discipline (best/base/worst) is the antidote. And never launder system metrics into business claims โ€” '40% higher quality on evals' is not '40% more revenue.' The three-tier separation exists precisely so each claim carries only the weight its evidence supports.
Quick check
1 question ยท instant feedback
0/1
  1. Selling the measurement plan alongside the ROI model does what?

Pitfalls & takeaways

Failure modes

  • No baseline captured during discovery โ€” no honest ROI claim possible later.
  • Inflated hours-saved math without sensitivity analysis.
  • Laundering system metrics into business claims.
  • Skipping leading indicators; waiting for lagging financials that never arrive in time.
  • Not counting the learning asset โ€” the part that actually compounds.
  • Open-ended verdicts; program judged by whichever quarter someone gets nervous.

Durable takeaways

  • Baseline captured in discovery is mandatory, not optional.
  • System / workflow / business โ€” three tiers, reported separately.
  • Leading indicators lead financials โ€” report both.
  • Learning asset is an asset line, not a footnote.
  • Verdict criteria pre-agreed, before politics arrive.

Do the work

๐Ÿ“ฆArtifact to produce
ROI model with a measurement-plan second tab: baselines, three tiers, cadence, 90-day verdict.

Sources

  • ยท MIT NANDA + Duke Fuqua/Federal Reserve CFO Survey + Volume 1's evals