Skip to module content
Module 20 ยท ~13 min

Personal Metrics & Review Ritual

Measure your own AI ROI before you sell it, scale it, or preach it.

Reading progress
0/5 ยท 0%

The big idea

๐Ÿ’กKey idea
Four honest numbers โ€” hours saved, quality delta, cost, and failure count โ€” turn a vague sense that AI is helping into a defensible ledger, provided the hours-saved column also counts setup, debugging, and review time. A weekly 15-minute check-in and a monthly keep/kill/add audit turn that ledger into an ongoing habit rather than a one-time exercise.
Quick check
1 question ยท instant feedback
0/1
  1. The four numbers are:

Deep dive

8/8 open

Measuring your own AI use builds credibility with yourself before it builds credibility with anyone else. It's easy to feel like AI is saving time without being able to say by how much, or in which specific workflows, or at what cost in setup and review.

That vague confidence doesn't survive contact with a skeptical client or boss asking for numbers โ€” and it shouldn't have to, once you've built the habit of tracking honestly.

Four figures capture the essential picture: hours saved, quality delta, cost, and failure count. Hours saved is the headline number but means little alone; quality delta tracks whether output got better or worse, measured through evals or real outcomes like response rates and error rates; cost is straightforward subscription and usage spend; failure count tracks how many times something genuinely went wrong.

Together these four numbers turn "AI has been great" into a specific, checkable claim.

The credibility of the whole ledger rests on one discipline: include setup, debugging, and review time as a minus column against gross hours saved. A workflow that saves five hours a week but costs two hours of debugging and review nets three โ€” a very different number from five, and the honest one.

Call this the skeptic's ledger: assume someone will challenge every number and make sure it survives the challenge.

Speed without quality tracking is a trap โ€” a workflow can feel fast while quietly getting worse. Quality deltas come from two sources: structured evals of the kind covered in earlier modules, and real-world outcomes like response rates, error rates, and client feedback.

Whichever source you use, measure it consistently enough over time that a trend, not just a single data point, becomes visible.

A fifteen-minute weekly ritual keeps the ledger alive: score the week against the four numbers, note any failures specifically, and pick one improvement to make before next week. Fifteen minutes is short enough to actually happen every week, which matters more than any more elaborate system that gets skipped.

The note-one-failure habit is doing real work here โ€” it's what feeds the monthly audit with concrete detail instead of vague impressions.

Once a month, review the whole stack with three verbs: keep, kill, add. A daily journaling automation that scores zero uses in three weeks despite "loving the idea" gets killed without guilt, freeing attention for a workflow that's actually compounding.

The audit's job is protecting your attention, not your ego โ€” every tool kept out of sunk-cost attachment is attention taken from something that would actually pay off.

A few months of honest ledger data becomes something more valuable than the numbers themselves: a before/after case study only you can write, because only you have the honest minus column behind it. That story is what actually persuades a client or a boss, far more than a headline number without context.

Writing it down periodically also forces the same honesty the ledger requires โ€” vague enthusiasm doesn't survive being put into a specific before/after narrative.

Some signals matter precisely because they don't show up as a positive number: skills atrophying, verification habits slipping, or thinking starting to take on an AI-shaped uniformity. These are worth watching for deliberately, since nothing in the four-number ledger surfaces them automatically.

The pitfall to avoid across all of this is vanity accounting โ€” counting gross time saved while ignoring setup, review, and rework, then quoting an inflated number to someone who later measures honestly. Your credibility as an operator rests entirely on that minus column.

Quick check
1 question ยท instant feedback
0/1
  1. Honest hours accounting includes:

Pitfalls & takeaways

Failure modes

  • Counting gross time saved while ignoring setup, debugging, and review time
  • Never running a monthly audit, so dead tools linger out of habit or attachment
  • Skipping quality measurement and tracking only speed
  • Quoting inflated time-savings numbers to clients who later measure honestly
  • Ignoring anti-metrics like slipping verification habits or atrophying skills

Durable takeaways

  • The four numbers โ€” hours saved, quality delta, cost, failures โ€” turn a vague sense of AI value into a defensible ledger
  • Honest hours accounting includes setup, debugging, and review time, not just gross time saved
  • A weekly 15-minute check-in and a monthly keep/kill/add audit turn measurement into a sustained habit
Quick check
1 question ยท instant feedback
0/1
  1. The monthly audit's verbs are:

Do the work

๐Ÿ‹๏ธProve you learned it

Start the ledger today with a simple sheet tracking the four numbers, updated in a weekly 15-minute calendar block. After four weeks, run your first monthly audit โ€” keep, kill, or add across your stack โ€” and write a 200-word honest before/after account of your single best workflow.

0 chars
Quick check
1 question ยท instant feedback
0/1
  1. An anti-metric worth watching is:

Sources

  • ยท Ethan Mollick (oneusefulthing.org)
  • ยท Lenny's Newsletter AI issues (lennysnewsletter.com โ€” measuring AI impact)
  • ยท Every.to "How I use AI" series