How do you verify the value of an AI automation before paying for it?
Write the number down before the build, measure the same number after a full cycle on your own data, and pay only on cash removed, never on hours saved.
You verify the value of an AI automation by writing down one number before anything is built, then measuring the same number after the automation has run for a full cycle, on your own data, with your own people counting. The number is cost removed or cash recovered in a named process. If the vendor sets the baseline or runs the measurement, it is a demo, not a verification.
Most AI value is claimed, and most of it never reaches the P&L
MIT NANDA's 2025 report on generative AI in business, as reported by Fortune (2025), found that about 95 percent of enterprise pilots delivered little or no measurable impact on the P&L. BCG (2024) surveyed 1,000 executives and found 74 percent had yet to show tangible value from AI. IBM's 2025 survey of 2,000 CEOs (IBM Institute for Business Value, 2025) put the share of initiatives that delivered the expected return at 25 percent.
Those projects were verified backwards: the tool went live, then someone looked for a benefit.
Write one page before anything is built
Name the process, the metric, the baseline value and where it came from, the period, and the person who will count. One page, agreed by the team that does the work and the team that builds. If that page cannot be written, the opportunity is not ready, and no tooling changes that.
PwC (2025) puts about 20 percent of an initiative's value in the technology and 80 percent in redesigning the work around it. The one page is where that 80 percent gets designed.
Measure in units your finance team already tracks
A vendor case study is measured in the vendor's units: percentage faster, hours saved per user, satisfaction scores. None of those appear in your accounts. For finance work, count invoices per person per week and days to close. For procurement, count contract value renegotiated and renewals caught before the date. The unit should be boring. Boring units survive an audit.
| Method | Who sets the baseline | What it proves | When to trust it |
|---|---|---|---|
| Vendor case study | The vendor | The tool worked somewhere once | Never, as evidence for your company |
| Pilot dashboard | The vendor's tool | Activity inside the tool | Only alongside your own count |
| Before-and-after on your data | Your team | Change in your metric over time | When the period covers a full cycle |
| Side-by-side run | Your team | Difference between old and new on the same work | The strongest option, and the most work |
Run old and new side by side for one full cycle
Before-and-after comparisons invite a simple objection: something else changed. Close it by running the old process and the new one on the same work for one cycle. Half the invoices through the agent, half through the team, same week, same approvers.
A cycle means a real cycle. For accounts payable that is a month-end. For collections it is a quarter. A two-week pilot in a quiet month proves very little.
Hours become value only when they become cash
Our view, and here we disagree with most of the industry, which reports value in hours saved: an agent that returns 400 hours a month to a team of ten has created value only if those hours turn into something. A role not backfilled, a contractor released, more volume handled with the same people. If the ten people are still there doing the same work, the hours are capacity until management decides what to do with it.
Track both. Pay only on the second. Anyone who invoices you on the first has moved the risk back to you.
Sign-off belongs to the people who did the work
The person who confirms the number is the one who used to do the task, or their direct manager. Neither the project sponsor nor the vendor. They know whether they now check the agent's output for two hours a day instead. That checking time is the most common hidden cost. If the agent saved eight hours and created three of review, the number is five, and the sign-off should say five.
A pay-from-value contract should state, in plain words: the metric, the baseline, the period, who counts, and what happens if the number is zero. It should say that verification comes before invoicing. Anything left to be agreed later postpones the argument without removing your risk.
Questions people ask next
What if the process has no baseline data?
Build the baseline first. Two to four weeks of counting by hand, on a shared sheet, is usually enough. An automation without a baseline cannot be verified, so do not start it.
Should we count time saved as money saved?
Only once the time has turned into a decision: a role not filled, a contractor ended, a cost line removed. Until then it is capacity, and it should be reported as capacity.
Who should own verification inside the company?
Finance, with the process owner counting. Finance already owns the numbers that will be argued over later, and it has no stake in the project looking good.
Can a proof of concept count as verification?
No. A proof of concept shows the automation can work. Verification shows that it did work, on your volume, for a full cycle, with the hidden review time counted.
Lightbloom AI writes up every problem with a number before anything is built: Transcript builds the world model of the company and values each problem in the company's own hours and rates. How we work is at /service-offering.
References
- Fortune, MIT report: 95% of generative AI pilots at companies are failing, reporting on MIT NANDA, The GenAI Divide: State of AI in Business 2025 (2025), https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
- BCG, AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value (2024), https://www.bcg.com/press/24october2024-ai-adoption-in-2024-74-of-companies-struggle-to-achieve-and-scale-value
- IBM Institute for Business Value, IBM Study: CEOs Double Down on AI While Navigating Enterprise Hurdles (2025), https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles
- PwC, 2026 AI Business Predictions (2025), https://www.pwc.com/us/en/tech-effect/ai-analytics/ai-predictions.html
Keep reading.
All notes- How do you capture tribal knowledge before someone leaves? Tribal knowledge is how the company actually runs, held by the people who run it. Why handovers fail, what works, and why it is the first layer of a world model. Read →
- Why AI agents need a world model of your company AI agents stall inside companies because nobody has described how the work runs. What a company world model gives an agent, and why to build it before you deploy. Read →
- What is a company world model? A company world model is a living picture of how a company really works: every process as it runs, what hurts most, and who everything depends on. Why AI agents need it. Read →