Measure AI agent ROI at the workflow level. Compare the financial benefit with the full investment, then count only work that somebody accepts.
Do not start with model price or time saved. Both can hide the cost of review, repair, failed runs, and work that never reaches the finish line.
#The formula
Use one period and one workflow.
ROI = (financial benefit - new workflow cost) / new workflow cost
Financial benefit can include avoided cost and added profit. The new workflow cost includes the platform, usage, setup, review, repair, and exceptions.
If you only measure cost reduction, use the current workflow cost as the financial benefit. Label the result as a cost-saving ROI estimate.
Use the same scope on both sides. Do not compare one month of current work with one week of agent work.
#A transparent example
Assume the current workflow needs 40 hours each month at $40 per hour. Its current monthly cost is $1,600.
Assume the new workflow costs $300 for the platform and usage. It also needs ten human hours for review and repair, which costs $400. The full new cost is $700.
The avoided cost is $1,600. The net benefit is $900. The cost-saving ROI estimate is $900 divided by $700, or 129%.
These numbers only show the calculation. They are not Alfera prices, customer results, or market benchmarks. Replace every input with your own measured data.
#Choose one result to count
Pick a unit that a manager already understands. Examples include:
- One reconciled payment batch.
- One support request resolved.
- One CRM record updated after a call.
- One weekly report delivered.
- One late invoice followed up.
Then define "accepted." A completed run is not enough. The result must meet the same standard as the current process.
#Measure the current workflow first
Record the baseline before the agent changes the work. Use real work for at least one normal cycle.
Track these inputs:
| Input | What to record |
|---|---|
| Volume | Results completed in the period |
| Active time | Minutes people spend doing the work |
| Waiting time | Time from request to accepted result |
| Review | Minutes another person spends checking |
| Rework | Minutes spent fixing errors |
| Exceptions | Cases that need senior judgment |
| Other cost | Contractors and workflow-specific software |
Ask the person doing the job. A process document often omits the small checks that make the result usable.
#Measure the agent workflow with the same table
Keep the unit and acceptance standard unchanged. Add four agent-specific results:
- Finished without repair.
- Finished after a person intervened.
- Wrong but detected.
- Wrong and not detected.
Also count blocked work. A blocked action can show that a permission or approval limit worked. It becomes a cost when nobody clears it or the explanation is not useful.
The AI agent evaluation guide explains how to test the failure paths.
#Use accepted results as the denominator
Cost per run can reward retries and unusable output. Cost per accepted result measures the thing the company wanted.
Cost per accepted result = total new workflow cost / accepted results
Run count still matters for diagnosis. It should not be the main business measure.
#Include every new cost
The agent cost has more than one line.
- The platform subscription.
- Model or usage charges.
- Integration and setup time.
- Human review.
- Human repair.
- Exception handling.
- Security and procurement work.
- Workflow maintenance.
Alfera lists its subscription and included usage on the pricing page. It charges one workspace fee plus the work that runs. Read why we charge for work, not seats for the reason behind that structure.
Use the same complete method for every vendor. A low model price can be a small part of the total.
#Separate cash savings from released time
One saved hour does not always put money back in the bank. Report three types of value separately.
Cash savings
Count spending that stops. This can include avoided contractor fees, removed software, lower error costs, or a hire the company no longer needs.
Added capacity
Count extra accepted results with the same team. Use the value of those results only when the company can explain it.
Released time
Report the hours people can use elsewhere. Do not turn every hour into salary savings unless payroll or hiring changes.
This separation makes the business case harder to inflate and easier to trust.
#Measure risk as an outcome
An average can hide one expensive mistake. Track failure severity as well as failure count.
Record whether a wrong action was detected before it reached another system. Also record whether the system respected the permission and approval rules.
Alfera checks permissions outside the model. The related engineering article explains why the location of that check matters.
#Run a two-week test
Use one real job. Avoid a task invented for the demo.
- Record the current process for one normal cycle.
- Write the job, exception, and approval limit.
- Connect only the sources and tools it needs.
- Run the agent for two weeks.
- Review accepted, repaired, wrong, and blocked results.
- Calculate the cost per accepted result and the ROI.
Use a month when the workflow runs weekly or changes at month end. The first month guide gives a wider rollout plan.
#Use a decision rule before the test
Set the threshold before you see the result. This prevents the team from moving the goal after a promising demo.
For example:
- The agent must lower the cost per accepted result.
- It must not increase undetected errors.
- It must complete a set share without human repair.
- It must show every blocked and failed case.
Choose the values from your workflow. Do not copy a benchmark from another company with different work and risk.
#Where ROI does not answer the decision
ROI does not settle legal, security, or customer-trust requirements. A positive number cannot approve a use case that breaks one of those rules.
Some work also has a value that the formula misses. Faster customer replies or less repetitive work can matter without a clean cash value. Report that value separately. Do not force it into the financial result.