Economics
The cost that matters is the cost per completed case, not the price of a token.
Teams budget for a model subscription and are surprised by retries, long context, tool calls, and the people still reviewing the work.

How I see it
AI agent cost
Agent cost is a stack. Inference. Tools. Infrastructure. Monitoring. Human review. Failure handling. The only number that should go into the business case is what it costs to finish one unit of work at expected volume.
Design choices move that number. Extra agents, extra steps, huge context, and unbounded retries are cost decisions. So is a review rate of 100%.
A cheap model that fails and retries can cost more than a stronger model that finishes. A low review rate on a high-error-cost action is not savings. It is risk.
Price the agent like a shared service. If the unit cost is higher than the human path, do not scale it.
Common mistakes
What teams usually get wrong.
Budgeting only for the model
Review and operations are usually visible in the first month.
Ignoring retries and failures
The unhappy path is where the bill grows.
Comparing to zero
Compare to the current cost of the workflow, not to doing nothing in a vacuum.
A useful diagnostic
Five questions before you fund the work.
What is a completed case, and how many will you run per month?
Without volume, you cannot price the agent.How many model and tool calls does a successful run take?
If unknown, you are not ready to forecast.What is the human review rate and minutes per review?
That is often the largest operating cost.What is the retry and failure assumption?
Zero is not a serious input.Is unit cost below the current human path?
If not, change the design or do not scale.
Economic model
Cost per completed case(model + tools + infrastructure + monitoring + review + failure handling) ÷ completed cases
Use this number in the ROI model. Do not use list prices for tokens alone.
Three credible paths
How far should you go?
Do not force one solution. Choose the path the economics, the risk, and the organization can support.
Measure a supervised pilot
Instrument real runs before you forecast production cost.
You have not yet seen live volume.
Pilot mix may be cleaner than production mix.
Optimize the expensive drivers
Cut context, steps, retries, or review where the design allows it.
You know which line item dominates.
Cutting review on a high-risk action is not optimization.
Scale only under a unit-cost ceiling
Set a maximum cost per case and treat a breach as an incident.
The agent is moving to real volume.
Requires observability, not a spreadsheet once a quarter.
When this is the wrong next step
Do not fund an agent here.
- Nobody will share volume or review time.
- The team wants to scale because the demo was cheap on ten cases.
- Failure handling is assumed to be free.

A useful next step
Bring one workflow. Get guided into production.
We guide the implementation, go deep on the technical path, and stay hands-on through operations — or tell you when a simpler answer is better.
Discuss an AI opportunity
