What belongs in an AI ROI model?
Build an AI ROI model that separates capacity from cash, counts adoption and risk costs, and gives the board auditable investment ranges.

Table of Contents
An AI investment deserves the same financial discipline as a new sales channel or a factory tool. A board should see the cash committed, the operational change that creates a return, the delay before that return appears, and the evidence behind every assumption. If the case depends on multiplying survey hours by salary, it is not ready.
I have watched teams turn a promising pilot into a fictional payback period by treating every saved minute as spend removed. The software may save time and still produce no cash return because nobody changes staffing, throughput, delivery dates, or error rates. A useful AI ROI model makes that conversion explicit. It also gives management a way to stop, narrow, or expand the investment as real data replaces estimates.
The method below works for coding copilots, support agents, internal research assistants, sales tools, and workflow agents. The details differ, but the finance does not: establish the baseline, model incremental cash flow, discount uncertain gains, include the cost of organizational change, and assign an owner to each benefit.
Separate technical activity from financial value
A board funds an operating outcome, not tokens, prompts, generated code, or hours that employees say they saved. Those technical measures help diagnose the system, but none belongs in the return line until it changes a result the company pays for.
The field routinely blurs three different ideas. Time saved is a task estimate: a developer says a test took 30 minutes instead of 50. Capacity created is usable time returned to a constrained team. Cash value appears only when the company converts that capacity into lower spend, more gross profit, faster receipt of revenue, or avoided loss. Calling all three "savings" produces a large number that finance will cut apart in minutes.
Suppose 20 engineers save four hours per week. At a loaded cost of $90 per hour for 48 working weeks, the apparent annual value is $345,600. That number is valid only if the company can point to what changes. If the same team ships the same roadmap on the same dates, payroll and revenue remain unchanged. The company has activity improvement, not an economic return.
There are four defensible conversion paths:
- The company removes contractor spend, open roles, overtime, or planned hires.
- A constrained team releases paid work sooner or handles more paid volume with the same headcount.
- Better quality reduces refunds, cloud waste, support labor, incident cost, or contractual penalties.
- Faster delivery moves a dated revenue event forward, and the model counts contribution margin during the time gained.
Name the path for every benefit. If a benefit has no conversion path, keep it as an operating metric outside the financial total. That does not make it useless. It keeps the investment memo honest.
Build the baseline before the pilot changes it
A credible baseline uses observable work from a fixed period, segmented by workflow and team. Do not ask people to remember how long work used to take after they have tried the tool. Memory bends toward whatever story the sponsor wants to tell.
Choose a period that reflects normal work. Eight to twelve weeks often covers enough variation for product engineering or support, but a seasonal business may need the matching period from the prior year. Record volume, cycle time, paid labor, quality failures, rework, and business output. Use medians for cycle time when a few stuck items distort the average. Keep the raw count beside every rate so a jump from one failure to two does not masquerade as a stable trend.
Segment the baseline where the mechanism changes. A coding copilot can help routine test creation while slowing an unfamiliar migration through review churn. A support agent may resolve password resets but escalate billing disputes. One blended productivity percentage hides both facts and encourages deployment into work where the tool performs badly.
Your baseline table should have one row per use case and these columns:
- Monthly work volume and the unit being counted.
- Median handling time or cycle time, including review and rework.
- Loaded internal labor cost and external spend.
- Defect, escalation, refund, or incident rate.
- Revenue or contribution margin tied to completed work, if any.
Freeze definitions before the pilot. If "resolved ticket" excludes reopened tickets in the pilot but included them in the baseline, the comparison is broken. Put the data owner and source beside each metric: payroll system, issue tracker, support platform, finance ledger, or a sampled time study. Boards accept estimates when management labels them and shows how actuals will replace them. They reject precision that nobody can trace.
Count the full cost of ownership
The license is usually the easiest cost and rarely the whole cost. Model one-time implementation spending, recurring operation, usage that grows with volume, human review, security work, and the management time required to change the workflow.
For a copilot, recurring cost may include seats, model usage, code scanning, and extra review of generated changes. For an agent, add orchestration, test environments, observability, retries, exception handling, and the employee time spent approving sensitive actions. If an agent touches customers or production data, include evaluation maintenance and incident response. A cheap model call can sit inside an expensive operating process.
Do not bury internal labor because the employees already work for you. Separate committed cash from allocated labor, then show both. Cash tells the board how much budget is at risk. Allocated labor shows what other work the implementation displaces. A security engineer spending six weeks on access controls has a real opportunity cost even if payroll does not change.
Use a cost schedule by month, not one annual lump:
Total cost[m] = license[m] + usage[m] + infrastructure[m]
+ implementation_labor[m] + review_labor[m]
+ training[m] + security_and_compliance[m]
+ vendor_and_exit_cost[m]
Include an exit allowance for data export, workflow replacement, and contract overlap. It is not a prediction that the vendor will fail. It prices the option to leave. I also add a contingency line to costs rather than quietly lowering every benefit. That makes the source of caution visible and prevents the same risk from being deducted twice.
Training is not a launch-day event. Teams need examples, office hours, review rules, and time to remove workflows that do not work. Manager time often spikes after launch because queues and responsibilities change. Count it. A model that assumes free organizational change will miss cash outlay and the delivery dip that appears before gains.
Convert capacity into cash with explicit gates
Every projected dollar should pass through a realization gate that management can observe and control. The gate is the action that turns operational improvement into cash flow. Without it, the benefit stays in a separate capacity ledger.
For labor reduction, the gate might be a contractor contract ending, a vacant role closed, or a planned hire removed from the approved plan. For throughput, it might be a contracted backlog, a sales pipeline limited by delivery capacity, or usage with known contribution margin. For faster launch, it is a dated release tied to a forecast that finance already recognizes. For quality, it is a measured reduction in paid rework, credits, infrastructure cost, or expected incident loss.
Apply a realization factor to avoid pretending that every available hour converts. The factor combines workflow coverage, adoption, task success, and the share of released capacity that management can actually redeploy. Keep those inputs separate in the working sheet even if the board slide shows their product.
Capacity hours[m] = eligible volume[m] * baseline hours per unit
* measured time reduction
Realized hours[m] = capacity hours[m] * adoption[m]
* task success[m] * redeployment[m]
Labor benefit[m] = realized hours[m] * loaded hourly cost
Loaded labor value is appropriate when released hours replace paid work. It is not automatically appropriate for spare employee time. If management plans to remove two contractor roles, cap the cash benefit at the contractor spend actually removed. If management plans to avoid three hires, start the benefit in the months those hires appeared in the approved plan, not on the pilot start date.
Revenue needs stricter treatment. Extra output has value only when demand exists and the company can deliver, sell, and collect. Use contribution margin, not bookings or gross revenue. If an agent lets a team launch one month earlier, count the margin expected in that month, adjusted for launch probability and collection timing. Do not count the same outcome again as labor capacity.
Assign a named executive owner to every realization gate. The tool owner can report adoption; the finance owner confirms spend removed; the revenue owner confirms the dated commercial event. When nobody owns conversion, the benefit belongs in the upside case, not the base case.
Use measured evidence instead of enthusiasm
A pilot should estimate causal improvement in a defined workflow, not collect applause from early adopters. Volunteers tend to have stronger motivation, cleaner use cases, and more patience than the people who receive a company-wide rollout. Their satisfaction can guide product choices, but it is weak evidence for a budget.
Run a phased comparison when you can. Match users or work queues by role, skill, work type, and baseline performance. Give one group the tool first and keep the comparison group on the existing process for the same period. If random assignment is impractical, stagger the rollout and compare changes within each group. Remove the first learning week from steady-state measurement, but show that productivity dip in implementation cost.
Measure the whole workflow. Coding time can fall while pull request review, test repair, or production defects rise. Support handling time can fall while reopen rates and escalations rise. Research can finish sooner while verification takes longer. The output boundary should end where the business accepts the work, not where the model stops generating.
Use a compact evidence register for each assumption:
Assumption: 18% lower accepted-ticket handling time
Source: 1,240 eligible tickets, six-week staggered pilot
Owner: VP Support
Confidence: medium
Recheck: after 90 days at more than 70% active adoption
Kill threshold: no reduction after reopen and escalation time
That artifact does two useful things. It lets finance trace a number back to evidence, and it tells the operating team what would disprove the case. I prefer confidence labels tied to evidence rules: high for ledger data or repeated controlled measurement, medium for a representative pilot, and low for vendor claims, interviews, or analogies. The labels should change the modeled range, not decorate the slide.
Do not average away failure. Report median improvement, the spread across users or tasks, and the share of eligible work that the tool can handle. A 30 percent gain on 20 percent of work is a 6 percent gross opportunity before adoption, quality, and realization. That calculation often saves a board from approving a rollout whose headline came from a narrow demo.
Put uncertainty into scenarios and cash flow
Boards can accept uncertainty when the model shows where it lives. Give them a downside, base, and upside case with explicit assumptions. Do not make the downside case a cosmetic five percent haircut across every line. Change the variables most likely to fail: eligible workflow share, adoption, task success, implementation delay, review burden, and realization.
A monthly model matters because AI projects often pay costs before adoption and benefits ramp slowly. Annual totals can hide a nine-month cash trough behind an attractive year-two return. Show monthly net cash flow for at least the committed contract period and long enough to see steady operation.
These formulas are enough for a board model:
Net cash flow[m] = realized benefit[m] - total cost[m]
Cumulative cash[m] = cumulative cash[m-1] + net cash flow[m]
NPV = sum(net cash flow[m] / (1 + annual discount rate)^(m/12))
ROI = (total realized benefit - total cost) / total cost
Payback month = first month cumulative cash is greater than or equal to zero
Return both ROI and net present value. ROI is easy to read but insensitive to scale and timing. Net present value captures timing and gives finance a cash amount to compare with other investments. Payback shows exposure duration. None replaces the others. Ask finance for the company discount rate instead of inventing one for the AI project.
Correlate assumptions when they share a cause. Weak adoption can lower benefit while leaving fixed license and implementation cost unchanged. Poor model performance can reduce task success and increase review labor together. A simplistic sheet that changes each cell independently can produce combinations that could never happen in operation. Three coherent scenarios usually communicate this better than a theatrical Monte Carlo chart built on guesses.
For the downside case, include the cost of stopping: contract commitments, decommissioning, duplicated workflow, and retained implementation labor. Exclude sunk pilot spending from a forward approval decision, but show it separately so nobody confuses total program spend with the incremental decision now before the board.
A worked model exposes the weak assumptions
Consider a company evaluating an AI support agent for a team handling 12,000 eligible tickets each month. Baseline handling and after-work time averages 12 minutes per ticket. Loaded labor cost is $48 per hour. A six-week pilot measures a 25 percent reduction in accepted handling time after reopen and escalation work.
The base case assumes 70 percent active adoption, 85 percent successful task completion, and 60 percent redeployment of released capacity. The gross monthly capacity is 12,000 times 0.2 hours times 25 percent, or 600 hours. After the three gates, realized capacity is about 214 hours. At $48 per hour, that is $10,282 of monthly labor value.
Now force the cash question. Management has a contractor renewal worth $7,500 per month and can end it in month five if service levels hold. It has no approved reduction for employee roles. The base case therefore counts no labor cash benefit in months one through four and $7,500 per month from month five. The remaining calculated capacity stays in the operating ledger. This is more conservative than multiplying all 214 hours by loaded cost, and it matches an action the COO can approve.
The company also expects fewer quality reviews, worth $2,000 per month in contractor spend, but the pilot sample is small. Put $1,000 in the base case from month four, $2,000 in upside, and zero in downside. Do not call an unmeasured improvement certain because it sounds plausible.
Implementation costs $42,000 across the first three months. Licenses, model usage, infrastructure, and evaluation cost $8,500 per month after launch. Human exception review starts at $4,000 per month and falls to $2,500 by month six as the team removes weak use cases. Training and security work add $18,000 in the first two months. Under these assumptions, steady base benefits of $8,500 merely equal steady operating costs. The project never repays the initial investment.
That result is useful. The team has four honest choices: negotiate lower recurring cost, find another measured cash conversion, increase eligible volume without harming quality, or stop. It should not rescue the proposal by pricing unallocated employee minutes as cash. In the upside case, the full $2,000 quality saving and an additional $5,000 monthly avoided hire beginning in month nine may create a return. Finance should require proof that the hire sits in the approved plan.
The worked case also reveals a portfolio issue. A support agent can be operationally successful and financially neutral by itself, yet provide infrastructure that lowers the cost of a second use case. Model shared platform cost once, allocate it consistently, and show the incremental economics of each added workflow. Never charge the first project all shared cost while letting later projects look free.
Give the board a decision, not a sales deck
A good approval memo makes the decision reversible where possible. Ask for staged capital tied to evidence, with authority to expand only when the operating and financial gates pass. That protects the downside without forcing the team to pretend it already knows the upside.
The board pack needs one page of economics and a short appendix. On the main page, show committed cash, base NPV, downside exposure, payback month, the two largest assumptions, and the next funding gate. In the appendix, include baseline definitions, the monthly model, evidence register, security and compliance costs, scenario assumptions, and benefit owners. Keep model activity metrics there unless one directly controls a payment or gate.
A practical approval sequence looks like this:
- Fund a bounded pilot with a named workflow, baseline, data owner, and maximum cash exposure.
- Approve limited production when quality thresholds pass and the company identifies a specific realization action.
- Release broader funding after finance verifies the first cash conversion and operating data holds at higher adoption.
- Reforecast quarterly, removing benefits that missed their gates and updating costs from invoices and labor records.
Write kill criteria before enthusiasm and sunk cost cloud the decision. Examples include quality below the existing process, review labor above the planned ceiling, adoption below the viable threshold after training, a blocked security requirement, or the loss of the contractor or hiring action that created the cash case. A kill criterion should name who decides and what happens to contracts, data, and workflow afterward.
Do not present vendor benchmarks as company forecasts. They can help set the range before a pilot, but the approval should depend on your volume, workflow, cost, and controls. The NIST AI Risk Management Framework asks organizations to govern, map, measure, and manage AI risk. For an investment memo, that means risk work belongs in the operating plan and cost model, not in a final slide labeled "considerations."
Govern the benefits after approval
The investment case becomes a management instrument after approval. Replace assumptions with actuals on a fixed schedule and reconcile benefits to the financial ledger. If the spreadsheet disappears once funding arrives, it was persuasion, not a model.
Track leading and lagging measures separately. Adoption, eligible task share, model success, review time, and quality predict whether value may appear. Contractor invoices, approved headcount, contribution margin, credits, and infrastructure bills show whether cash did appear. The operating owner reports the first group; finance verifies the second.
Avoid cumulative victory metrics that never reverse. If an agent saves an estimated 10,000 hours while headcount, delivery, and quality do not change, the counter has become theater. Reset estimates when workflow volume changes, subtract added review and rework, and retire benefits when the planned conversion action disappears.
For a portfolio, maintain a shared register with use case, owner, current stage, committed cash, forecast NPV, actual net cash, evidence grade, next gate, and stop condition. This prevents five teams from claiming the same avoided hire or charging shared infrastructure inconsistently. It also lets management move money toward workflows that convert and away from those that only produce impressive demos.
A Team & AI Audit at oleg.is uses this operating view to identify where a smaller AI-augmented team can produce measurable savings before a company commits to a larger transformation. The stated audit is fixed at $5,000 over five business days, with a guarantee of at least $50,000 a year in identified savings or it is free.
Benefit governance also needs a rule for attribution. If faster delivery and lower rework come from the same change, define which metric receives the benefit before the review. Count the direct cash effect once. A release that arrives earlier may add contribution margin, while fewer defects protect part of that margin. Adding both full estimates can double count the same customer outcome. Finance should keep an attribution note beside the value and test whether removing one assumption changes the other.
Compare forecasts with actuals by cohort and use case, not only for the program total. A strong support workflow can conceal a weak sales workflow when their numbers roll into one portfolio return. Cohort reporting also shows whether later users match the early experts. If performance drops as adoption expands, lower the forecast before the next contract commitment. The purpose of a reforecast is to make a better decision, not to defend the number approved six months earlier.
Set variance rules in advance. A ten percent cost overrun may require an owner explanation, while a missed quality threshold may freeze expansion immediately even if the financial total still looks positive. Different variances carry different operating risk. Record the response, the owner, and the date when management will decide whether to correct, narrow, or stop the workflow. This turns governance into a routine operating process instead of a quarterly argument over whose spreadsheet is right.
Track displaced work as well. When staff use released capacity for another approved priority, name that priority, record its output, and avoid pricing it twice. If the capacity absorbs previously unfunded maintenance, report the operational gain without inventing cash. If it brings forward a contracted delivery, use the contribution margin and dated receipt. The distinction can feel severe, but it tells management which benefits improve the income statement and which improve the company without changing near-term cash.
Keep the board model alive for at least two planning cycles. At each review, ask one blunt question for every benefit: what changed in the ledger because we deployed this system? If the owner cannot answer, move that amount back to the capacity ledger until the company completes the conversion. That discipline does more for AI returns than another optimistic productivity survey.
Frequently Asked Questions
What is a reasonable ROI target for an AI project?
There is no universal target. Compare the project's NPV, payback period, downside exposure, and management burden with the company's other uses of capital, using the discount rate finance already applies.
Should saved employee hours count as AI savings?
Count them as capacity first. Move them into financial benefit only when the company removes spend, avoids an approved hire, increases contribution margin, or reduces a measured loss.
How long should an AI ROI pilot run?
Run it long enough to cover normal workflow variation and the learning period. Six to twelve weeks often works for steady operational queues, while seasonal or low-volume work needs a longer comparison.
How do you calculate ROI for an AI copilot?
Subtract the full cost of licenses, usage, implementation, review, training, security, and operation from realized financial benefits, then divide by total cost. Also report NPV and payback because ROI alone hides timing and scale.
What costs do companies miss in AI business cases?
Human review, workflow redesign, evaluation maintenance, security work, retry costs, management time, and exit work are frequent omissions. The early productivity dip during training also belongs in the monthly model.
Can productivity survey results support a board investment?
They can explain user experience, but they should not carry the financial case. Pair surveys with workflow data, accepted output, quality, paid cost, and a documented action that converts capacity into cash.
How should AI risk appear in an ROI model?
Price the controls and expected operating losses that the company can estimate, then use gates and scenarios for uncertainty that cannot be priced cleanly. Do not hide security, compliance, or quality work outside the investment total.
When does avoided hiring count as a benefit?
Count it when the role exists in an approved hiring plan and management removes or delays it because the deployed workflow handles the demand. Begin the benefit in the month the hire would have started.
What should make a company stop an AI rollout?
Stop or narrow it when quality misses the existing process, review costs break the ceiling, adoption stays below the viable level, a required control fails, or the cash conversion disappears. Set those thresholds and decision owners before launch.
How often should an AI investment model be updated?
Update operating measures monthly during rollout and reconcile financial benefits at least quarterly. Replace estimates with invoices, payroll decisions, margin data, and measured quality costs as they become available.


