Enterprise AI ROI: measuring what is not on the invoice
Baseline, metrics per use case, efficiency gain versus accumulated asset — and how to avoid endless pilots that never reach production.
*Tenth article in our series on enterprise AI, legacy integration and information governance.*
Executive summary
AI ROI fails on two fronts: nobody measured how the process performed before, and the calculation ignores the accumulated asset. Without a baseline, every gain is anecdote. Without counting the reusable knowledge produced, the company underestimates precisely the part that compounds. The method below applies the rigor of an industrial automation business case to knowledge work.
Start with the baseline
Before any pilot, measure for four to six weeks in the chosen process:
- average end-to-end cycle time;
- people involved per unit of work;
- rework rate and errors found after delivery;
- monthly volume and seasonality;
- satisfaction of whoever receives the output.
If the company cannot measure today, that is already the project's first deliverable — and it is worth doing on its own.
Metrics by use-case type
| Use-case type | Primary metric | Quality metric |
|---|---|---|
| Document generation (proposals, reports) | Time to first approvable draft | % approved without significant rewrite |
| Service and support | Resolution time | Ticket reopen rate |
| Document analysis (contracts) | Documents analyzed per hour | Correct findings vs. omissions |
| Internal knowledge search | Time to the correct answer | Answers with a valid cited source |
| Process automation | Cases completed without intervention | Exception and reversal rate |
Pick one primary and one quality metric per case. Two well-measured metrics beat a dashboard with thirty.
The two kinds of return
Efficiency (linear). Hours saved, cycles shortened, capacity released. It shows up fast and is easy to defend, provided the freed hour is reallocated — savings that become neither capacity nor cost reduction are just comfort.
Asset (compounding). Validated templates, a curated archive, evaluations and a decision trail. It never appears on the invoice, and it is what lowers the cost of the next use case. An honest proxy: how many new cases reused archive material instead of starting from scratch.
Total cost, unvarnished
- 1.Licenses and inference (the visible part).
- 2.Legacy integration and data modeling.
- 3.Curation and archive maintenance.
- 4.Governance: review, audit and impact assessment.
- 5.Change management and training.
A practical reserve rule: 15% to 20% of implementation value per year for ongoing evolution, the same discipline used in industrial automation.
How to escape the endless pilot
Pilots drag on when there is no written exit criterion. Define before starting:
- the minimum value of the primary metric that justifies production;
- the cost ceiling per unit of work;
- the governance requirements that must be in place (trail, permissions, policy);
- the decision date — and the willingness to shut the case down if criteria are missed.
Killing a weak use case quickly is a positive result: it frees budget for the next one.
What to do on Monday
- Pick a process and start measuring its baseline this month, even with no AI involved.
- Define two metrics and a written exit criterion for every active pilot.
- Include items 2 through 5 of total cost; a proposal with only item 1 is incomplete.
- Publish the result 90 days after go-live, positive or negative — transparency is what enables the second project.
Conclusion
AI ROI is not an exercise in optimism but in traceability: measured baseline, attributable gains, counted assets and honest cost. With that, the conversation stops being about technology and becomes about capital allocation — the language of the people who approve. The next step is sequencing adoption waves, the subject of A 12-month roadmap.
Further reading
Engineering track:
- Instrumenting cost and quality in AI production — the technical deep dive on this topic.
