Back to blogEngineering

ArchitectingAIAgentsforRealWorkflows,NotDemos

Marcus Lindqvist· Staff ML Engineer· February 25, 2026· 10 min read

Tool use as an API contract, not a suggestion

An agent's tools should be designed with the same rigor as a public API: strict input validation, clear error messages the model can actually parse and act on, and idempotency wherever an action might get retried. We have seen agents double-book a calendar slot or issue a duplicate refund because a tool call timed out and the agent, following its own reasoning, simply tried again without knowing the first attempt had actually succeeded.

Every tool should also return enough structured context for the agent to know whether it should continue, retry, or stop. A tool that returns a bare error string forces the model to guess at what went wrong, and guessing under uncertainty is exactly where agents make their worst decisions.

Guardrails on the plan, not just the output

Most teams add guardrails only at the final output: a content filter or a moderation check before the response reaches the user. For multi-step agents, that is too late, because damage can happen at step three of six even if the final message looks fine. We add checks at the planning stage too: before executing an irreversible action like sending an email or issuing a payment, the agent must produce an explicit plan that gets validated against a policy list, not just execute the next tool call it decided on.

We also cap the blast radius of any single action by default, limiting things like maximum refund amount, number of records touched, or number of messages sent per run, with anything above the cap requiring escalation regardless of how confident the agent's reasoning appears.

Human-in-the-loop escalation as a first-class path

Escalation to a human should not be an afterthought bolted onto error handling; it should be a designed outcome the agent can choose deliberately when confidence is low or the action is high-stakes. We build escalation as a first-class tool the agent can call, with a clear rubric for when to use it: ambiguous user intent, an action above the defined blast-radius cap, or a tool returning conflicting information.

In our production agents, an escalation rate between 5% and 15% of runs is usually healthy, not a failure signal. An agent that never escalates is either operating in an overly narrow domain or silently taking risks it should be flagging, and we treat a near-zero escalation rate as something to investigate rather than celebrate.

Observability: treating agent runs like distributed traces

Debugging a multi-step agent without full tracing is close to impossible, because the failure is often several steps removed from where it becomes visible. We log every planning step, tool call, tool response, and intermediate reasoning summary as a structured trace, similar to distributed tracing in microservices, so a failed run can be replayed step by step rather than reconstructed from a single final log line.

This tracing also feeds back into the eval harness: failed or escalated runs get triaged weekly, tagged by failure category (bad tool call, ambiguous instruction, incorrect plan), and the most common categories drive the next round of prompt or tool changes, closing the loop between production incidents and engineering priorities.

Failure modes we design for from the start

The most common agent failures we see are not dramatic; they are quiet ones like an agent looping on a tool call that keeps returning an ambiguous result, or drifting from the original goal after several steps of reasoning. We mitigate both with hard step limits, explicit goal restatement at fixed intervals during long runs, and a supervisor check that compares the agent's current action against the original task before allowing continuation past a set number of steps.

AI AgentsTool UseObservabilityAI Engineering

Wanthelpshippingsomethinglikethis?