Operate

An agent that nobody operates will be treated as a failed hire.

94%

of teams with agents in production have observability in place. Full step tracing is still only 72%.

LangChain, State of Agent Engineering 20261

Launch is the start of the cost. Someone has to watch volume, exceptions, model bills, and the cases that make customers or controllers unhappy.

Adnan Boz
Adnan Boz

How I see it

AI agent operations

Agent operations is the unglamorous work that makes leverage real. There is an owner. There is a queue for exceptions. There is a weekly picture of completion, cost, and quality. There is a way to roll back a change.

This is closer to running a shared service than to running a chatbot. The agent has capacity, failure modes, and a cost per completed case. If those numbers are not owned, the business will quietly put people back on the workflow.

Incidents will happen. A tool will change. A prompt will drift. A vendor model will behave differently. Operations means you notice that in hours, not after a quarter of bad records.

For a Miami company without an agent platform team, the first operating model can be one owner and a short review. It does not need a new department.

Common mistakes

What teams usually get wrong.

01

Project team leaves, nobody remains

The agent becomes an orphan with credentials.

02

Only watching uptime

The model can be up and the cases can still be wrong.

03

Treating exceptions as product bugs only

Many exceptions are process and data issues. Operations has to route both.

A useful diagnostic

Five questions before you fund the work.

  1. Who gets paged or emailed when the agent fails a class of cases?

    If the answer is nobody, you do not have operations.
  2. Is there a weekly review with volume, cost, and quality?

    Monthly storytelling is too slow.
  3. Can you turn off a change without turning off the workflow?

    If not, every fix is a business risk.
  4. Are exception SLAs defined?

    An agent that escalates into a black hole recreates the old queue.
  5. Is run cost compared to the original business case?

    Operations includes the economics, not only the incidents.

Economic model

Operating load

exception volume × handling time + monitoring + model cost + change review = cost to keep the agent alive

If this approaches the value created, the agent is not leverage yet.

Three credible paths

How far should you go?

Do not force one solution. Choose the path the economics, the risk, and the organization can support.

01

Named owner and weekly review

One person watches the cases, cost, and exceptions.

Best when

You have a single production agent.

Limitation

Does not scale to a portfolio.

02

Exception queue with SLAs

Treat escalations as an operated queue, not a side inbox.

Best when

Volume is high enough that exceptions are a real workload.

Limitation

Needs staffing math, or the queue becomes the new bottleneck.

03

Shared agent operations

Common monitoring, release, and incident patterns across agents.

Best when

Several agents share the same failure and cost problems.

Limitation

Easy to staff a team before the value is there.

When this is the wrong next step

Do not fund an agent here.

  • Nobody will own the agent after the build.
  • The company wants 24/7 automation with no one watching exceptions.
  • Leadership will not look at cost per case after launch.
Adnan Boz

A useful next step

Bring one workflow. Get guided into production.

We guide the implementation, go deep on the technical path, and stay hands-on through operations — or tell you when a simpler answer is better.

Discuss an AI opportunity