Build

Human-in-the-loop is not a brake on the agent. It is the operating design.

21%

of enterprises have a mature model for what an agent may do without a person.

Deloitte, State of AI in the Enterprise 20261

If people review everything, you have a slow copilot. If they review nothing, you have unmanaged risk. The work is choosing the control points.

Adnan Boz
Adnan Boz

How I see it

Human-in-the-loop agents

The point of an agent is to remove people from routine coordination, not from accountability. Humans should see the cases that require judgment, authority, customer nuance, or risk acceptance.

Good control points are specific: approve a discount over a threshold, accept a credit decision, send an external commitment, or close a case that changed a financial record. Bad control points are “please check this” on every run.

Reviewer design matters as much as model design. If the queue is noisy, people rubber-stamp. If the packet is incomplete, they re-do the research. The agent should present the recommendation, the evidence, and the action it wants to take.

This is how a COO gets a visible win without betting the operating model. The team keeps authority. The agent takes the grind.

Common mistakes

What teams usually get wrong.

01

Review as a formality

A checkbox at the end of a 20-step run trains people to approve without reading.

02

Review as a second job

If reviewers have to re-gather context, the agent did not prepare the work.

03

The same person in every loop

A single bottleneck reviewer recreates the original queue.

A useful diagnostic

Five questions before you fund the work.

  1. Which actions require authority, not just accuracy?

    Those are the first control points.
  2. Can a reviewer decide in minutes from the packet the agent prepares?

    If not, the loop will not scale.
  3. What share of cases should reach a human after the first release?

    If the answer is 100%, you are still in assist mode.
  4. Who owns the review SLA?

    An agent that waits on an unowned queue is not in production.
  5. Can you measure reviewer override rate?

    Overrides tell you whether the agent is learning the work or wasting time.

Economic model

Review load

cases × human review rate × minutes per review × loaded cost = annual review cost

If review cost erases the savings, the control points are too wide or the packet is too weak.

Three credible paths

How far should you go?

Do not force one solution. Choose the path the economics, the risk, and the organization can support.

01

Prepare and recommend

The agent assembles context and a recommended action. People complete the case.

Best when

Trust, access, or error cost is still high.

Limitation

People remain in the critical path for every case.

02

Approve the exceptions

The agent completes the common path. People see only threshold breaches and low-confidence cases.

Best when

You can define the common path and measure it.

Limitation

Thresholds need an owner or they drift.

03

Sample and audit

Humans review a sample and all high-risk actions after the agent is stable.

Best when

Failure modes are known and rollback exists.

Limitation

Sampling can miss a new failure mode if monitoring is weak.

When this is the wrong next step

Do not fund an agent here.

  • There is no one with authority to sit in the loop.
  • The company wants a fully unsupervised story for a high-risk process.
  • Reviewers cannot be given a complete packet and a clear decision.
Adnan Boz

A useful next step

Bring one workflow. Get guided into production.

We guide the implementation, go deep on the technical path, and stay hands-on through operations — or tell you when a simpler answer is better.

Discuss an AI opportunity