AI Strategy

When should businesses use AI agents?

Agents are the wrong default for most processes. Here is the decision test we use, the questions that disqualify a candidate, and where autonomy actually pays.

Perspective 6 min read All insights

Agents are the last rung, not the first

An agent decides which tools to call and in what order. That flexibility is exactly what makes it useful and exactly what makes it hard to verify. If a workflow can be written as a fixed sequence of steps, write it as a fixed sequence of steps.

Most production automations belong on the lower rungs: a deterministic pipeline, then a model inside one known step, then a bounded agent with a fixed toolset, and only then supervised autonomy that plans across systems. Each rung is cheaper to test, cheaper to explain and cheaper to switch off.

The common mistake is treating agent capability as the goal. The goal is a shorter cycle time on work that used to wait for a person. Autonomy is one route to it, and the most expensive one.

  • Deterministic workflow: fixed steps, fixed order, fully testable
  • Model inside a step: classification, extraction or drafting inside a known frame
  • Bounded agent: several tools, but an allowlist, a budget and a stopping rule
  • Supervised autonomy: plans across systems, acts, reviewed after the fact
level: 2
tools_allowed: [search_orders, read_policy, draft_refund]
max_steps: 8
budget_usd_per_run: 0.50
stop_when: [no_confident_action, budget_exceeded]
requires_human_approval: [refund_issued]

The level is a decision to revisit, not a permanent property of the system.

Five questions that disqualify a candidate

Before designing anything, run the candidate through these. Two or more no answers means the workflow stays deterministic, and that is a perfectly valid outcome.

The questions are deliberately blunt, because each one discovered late invalidates weeks of work.

  • Can you write an acceptance test for a good run?
  • Can you detect a bad run without a person reading the output?
  • Is the action reversible, or is a person able to compensate for it?
  • Is the blast radius one record rather than a whole portfolio?
  • Is the work frequent enough that failures surface within a week?

If quality can only be judged by reading the output, keep the person in the loop and automate the preparation instead.

Where agents hold up

Agents earn their place where the input is genuinely ambiguous, the path is not knowable in advance, and the work is reversible. Research across an internal corpus, triage across many queues, investigation of a failed pipeline, first-pass scoping of an incident. In each case the output is a proposal, not a commitment.

They also help when the work today is repeated reading and re-reading the same material to reach a conclusion someone could write down if asked clearly. The agent compresses that effort. The judgement still sits with the reviewer.

  • Open-ended questions over scattered documents, answered with citations
  • Triage where routing depends on content rather than on a fixed rule
  • Investigation across logs, traces and recent changes
  • Drafting a first artefact for a person to correct

Where agents are the wrong answer

Do not hand an agent a tool that moves money, changes a contract or files with an authority unless a human authorises each execution and the action is idempotent. The failure mode is not a wrong sentence. It is a correct-looking action taken against the wrong record.

Avoid them where success is genuinely subjective, where the tool surface is wide and hard to enumerate, or where the work happens once a quarter and offers too few repetitions to learn from. Occasional autonomous action on unfamiliar input is where confident mistakes are most likely.

Avoid them when the underlying data is unreliable. An agent reasoning over a source of truth that is already wrong will produce a confident answer about a broken fact, and it will do so consistently.

  • Irreversible or legally significant actions without per-execution approval
  • Objectives two competent people would score differently
  • Broad or unenumerated tool access
  • Low-frequency workflows with no evaluation data
  • Sources of truth with known accuracy problems

The boundaries that make an agent safe

Safety here is mostly a matter of making failure legible. An agent that stops and explains is an inconvenience. One that keeps going quietly is a liability. Design the stop conditions before the happy path.

Scope credentials to the specific tools the agent needs, cap what a single run may spend, limit how many actions it may take, and support a dry-run mode that resolves the plan without executing it. Every run keeps a trace of the calls it made, so a bad outcome can be explained rather than guessed at.

  • Explicit tool allowlist with no wildcard access
  • Per-run and per-day spend ceilings
  • Step limits so a loop cannot run indefinitely
  • Dry-run mode available for every agent before it is enabled
  • A trace retained per run and owned by somebody who reads it
policy:
  tools: read_only_except: [draft, propose]
  credentials: scoped_service_token
  max_steps: 12
  budget: { per_run_usd: 0.75, per_day_usd: 50 }
  idempotency: required_on_writes
  dry_run: supported
  on_stop: notify_owner_with_trace

A trace nobody reads is not an audit trail. Make sure someone owns reading it.

Pilot in shadow before you let it act

Run the agent against historical traffic with writes disabled, or alongside the current process, and compare. Disagreement with the existing outcome is the most informative signal available, and it costs nothing but compute.

Turn a sample of those disagreements into an evaluation set and track how the rate moves as you adjust the prompt, the toolset or the model. Treat escalation rate as the go or no-go metric, because it tells you how much human attention the design still requires.

Give the pilot an end date and a written rollback. Autonomy that grows by default is how a bounded agent becomes the thing you were trying to avoid.

  • Shadow mode against real historical cases before any write access
  • A fixed evaluation set that grows from real disagreements
  • Escalation rate as the primary operational metric
  • A written rollback and a named owner for any scope expansion

Takeaways

  • Climb the ladder of autonomy only as far as the problem requires
  • Two failed disqualifying questions means keep it deterministic
  • Reversibility and detectability matter more than capability
  • Trace every run and give somebody ownership of reading them
  • Pilot in shadow, and decide in advance when to expand scope

Facing something similar?

These articles describe how we think about a problem. A conversation is where we would apply it to yours, against your actual constraints.

Start a conversation All insights