How to Measure AI Automation ROI Without Fooling Yourself
Count baseline minutes, volume, rework, review, model cost, and platform cost for one loop. Payback is a date. Hoped hours are not a business case.

You can save ten hours on a slide and still spend twelve in review. Return on investment (ROI) for an automation is not a vibe. It is baseline minutes, volume, rework, review, model cost, and platform cost, for one loop, in a dated window.
Pick the loop with which business processes to automate with AI. Write the standard operating procedure (SOP) first (SOPs before automation). Print the AI automation ROI worksheet and fill it before you buy another seat. Evaluation before production is a different job (evaluate an AI workflow before production). The process scorecard decides rules versus model. This page decides whether the loop is worth running at all.
What counts, and what is a costume
| Input | Counts | Costume |
|---|---|---|
| Volume | Jobs in a dated window | Yearly hope divided by 12 |
| Human minutes | Timed or sampled | “It feels like an hour” |
| Rework | Jobs that come back, minutes to fix | Zero because the demo was clean |
| Review rate | Share of jobs a human still opens | Ignored because the model “usually” is fine |
| Run cost | Model invoice plus platform for that window | A blog’s average cents per 1,000 tokens |
The arithmetic, named
flowchart LR
B["Baseline minutes"] --> N["Net minutes"]
A["After minutes including review"] --> N
R["Rework delta"] --> N
N --> V["Value of minutes"]
C["Model plus platform"] --> P["Payback"]
V --> P
H["Build hours"] --> P
Baseline minutes = volume × minutes each, plus rework minutes. After minutes = volume × (review minutes × review rate), plus the new rework. Net minutes saved can be negative. Value = net minutes × loaded cost per minute. Monthly net = that value minus model invoice minus the share of platform you would not pay without this loop. Payback months = build hours × loaded hourly cost, divided by monthly net, if monthly net is positive. Otherwise NEVER.
Loaded minute cost is salary plus overhead, divided by minutes actually worked, not by 40 hours of calendar. A guessed number is allowed if you write GUESSED. A missing number is not ROI.
Baseline before you touch a model
Time the work as it exists. Ten samples beat a round number from a kickoff. Include the minutes spent waiting on a spreadsheet, not only the minutes typing. If nobody can time it, you do not have a process. You have a folk tale. Do not automate a folk tale and then claim savings against a number you invented.
- One owner: the person who currently does the job. Not the vendor.
- One window: last 30 days if volume is weekly. Last 90 if it is rare. Label the dates.
- Rework is a cost: a wrong tag that a human later fixes is not “almost automated.”
- Exceptions stay in the count: the 8% that still go to a human are part of after-minutes, not a footnote.
After: include the human who still sits there
A human-in-the-loop (HITL) review is part of the new minutes, not a footnote. If every send is reviewed, your new human minutes are read + edit + click. That can still win. It can also be slower than the old paste. A 100% review rate is a design choice. Price it.
Model cost is the invoice for that window. Platform cost is the seat or runtime you would not pay without this loop. Do not amortize a whole n8n cloud plan onto one toy workflow unless that plan exists only for this loop. Build hours are once. If you skip them, a two-week build that saves 20 minutes a month looks free. It is a hobby.
A 30-day measurement sprint
- Days 1–3: Name the loop. Fill volume and sample ten baseline jobs with a timer.
- Days 4–7: Write GUESSED on loaded cost if needed. Count last-window rework from the inbox, not from memory.
- Days 8–21: Run the candidate loop next to the old path, or on a slice. Log review minutes and new rework. Do not turn off the old path yet.
- Days 22–30: Fill after-cells from invoices and the log. Write a payback month or NEVER. Kill, cut review, or keep. Evaluation of quality is the next article, not this worksheet.
Hypothetical, labeled as such
Hypothetical. A studio tags 80 inbound emails a month. Baseline: 8 minutes each, 10% come back, 15 minutes to fix. After a classifier: 3 minutes of review on 100% of jobs, rework 12%, model plus platform $40 for the month. Loaded minute cost $0.80 (guessed, labeled). Baseline minutes are 80 × 8 plus 8 × 15 = 760. After minutes are 80 × 3 plus about 10 × 15 = 390. Net is 370 minutes, about $296 of guessed time, minus $40 run cost. If the build took 12 hours at $48 loaded, payback is a couple of months if the volume holds. If review climbs to 8 minutes because the model is messy, net collapses. The worksheet exists so you see that before you hire a prompt.
What that example does not prove: a payback for your firm, or that classifiers always win. It proves the cells can be filled without a case-study percentage.
Failure modes
- Counting hoped hours: the slide used 20 minutes. The timer said 6.
- Ignoring review: the model drafts. A senior rewrites. You paid twice.
- Quality as ROI: fewer rude emails can be a goal. Call it quality. Do not convert it to dollars you did not measure.
- Build hours forgotten: a two-week build that saves 20 minutes a month is a hobby.
- Averaging loops: the winner hides the loser. One worksheet each.
Observed, this studio, 2026: we refuse a seat until the loop has a named owner, a dated window, and a payback cell that is allowed to say NEVER. That is a decision rule, not a savings claim. What it does not prove: that every workflow we ship is cheaper than a human. Some we kill.
Fill the worksheet. If you want help scoring a loop without a fantasy number, that is an automation conversation. Contact.
Frequently asked questions
What is ROI in this article?
Return on investment (ROI) here means: did this loop return more value than it cost to build and run, in a dated window you can show. It is not a vendor slide. If you cannot name the minutes and the invoices, you do not have ROI. You have a story.
Do I need a dollar figure for human time?
You need a loaded minute cost you can defend, even if it is labeled Guessed. Salary plus overhead, divided by minutes actually worked. If you refuse to put a number on time, stop calling it ROI. Call it a quality experiment.
What if review ate all the saved minutes?
Then payback is NEVER for that design. Cut the review surface, stop the send, or kill the loop. A model that drafts and a human that rewrites the whole thing is a new job, not automation.
Can I average three processes into one payback?
No. One loop, one worksheet. Averages hide the loop that is losing money.