Skip to content
PD
AI Automation 8 min read

Guardrails Before Autonomy: My Checklist for AI Agents That Touch Real Systems

Prompts are not guardrails. Here's the enforcement layer I build in code - tool allowlists, idempotency keys, budgets, dry-run mode, approval gates and replayable traces - before an AI agent gets access to a client's production data.

PD

Pavel Duglas

AI Automation & MVP Architect

Last quarter I shipped an outreach agent for a client. It had two rules in the system prompt: never message the same lead twice, and never send anything outside 9:00–18:00 local time. On day three it sent 41 duplicate messages at 2 a.m. The prompt was fine. The model was fine. The architecture was the problem: I had written rules where I should have written code.

This is the single most common mistake I see in agent projects right now. Teams spend two weeks tuning prompts and zero days building an enforcement layer. Then they discover that a probabilistic system asked politely to behave will, eventually, not behave. Below is the checklist I now run through before any agent gets a token that can write to a real system.

The core principle: the model proposes, the code decides

An agent’s output is a request, not an action. Everything the model produces should pass through deterministic code that can reject it. If a constraint matters - money, deduplication, rate limits, data deletion - it lives in a function with an if statement, not in a paragraph of English.

Said differently: assume the model is a talented, slightly drunk contractor with good intentions and no memory. You wouldn’t hand that person your production database credentials and a Slack message saying “please be careful.” You’d give them one tool at a time and check the work.

Layer 1: A tool allowlist scoped to the task, not the agent

Most frameworks let you register tools on the agent. That’s already too coarse. A support agent that can refund_order is a support agent that will refund an order during an unrelated conversation about shipping.

I scope tools per task type, resolved before the run starts:

TOOLSETS = {
    "answer_question": ["search_docs", "get_order_status"],
    "process_refund":  ["get_order", "create_refund_request"],  # request, not refund
    "update_address":  ["get_order", "propose_address_change"],
}

def tools_for(intent: str) -> list[Tool]:
    return [TOOLS[name] for name in TOOLSETS[intent]]

Note the naming: create_refund_request, not refund. Write tools that produce intents into a queue, and let a separate boring worker execute them. This one habit removes about 70% of agent blast radius, because now every write goes through a place you control, log and rate-limit.

Layer 2: Tool signatures narrow enough to be safe by construction

A tool with a free-text parameter is a hole in your system. Compare:

# bad: the model writes SQL, you pray
def query_db(sql: str) -> list[dict]: ...

# good: the model picks from a closed set
def get_orders(customer_id: int,
               status: Literal["pending", "paid", "shipped"],
               limit: int = 20) -> list[dict]: ...

Use enums, integer IDs, bounded ranges. Validate with Pydantic (or zod) and return a structured error on failure so the model can self-correct: {"error": "limit must be <= 50", "retryable": true}. Vague errors like "invalid input" cause models to hallucinate new parameter names and burn your step budget.

Layer 3: Idempotency keys and a write ledger

My 2 a.m. duplicate-message incident had a mundane cause: a network timeout, a retry, and no idempotency. The agent “knew” it had already messaged the lead, but the memory it consulted was its own context window, which had been trimmed.

Every write tool gets a deterministic key derived from business meaning, not from a UUID:

key = sha256(f"outreach:{lead_id}:{campaign_id}".encode()).hexdigest()

if ledger.exists(key):
    return {"status": "already_done", "at": ledger.get(key).created_at}

ledger.reserve(key)              # unique index in Postgres does the real work
try:
    result = send_message(...)
    ledger.commit(key, result)
except Exception:
    ledger.release(key)
    raise

The unique index is the guardrail. The prompt is a suggestion. If your ledger lives in the same transaction as the effect, duplicates become structurally impossible rather than statistically unlikely.

Layer 4: Four budgets, all enforced outside the model

Agents fail loudly when they crash and expensively when they loop. I cap four things per run:

BudgetTypical valueWhat it prevents
Steps (tool calls)12–20Infinite plan/replan loops
Wall clock90s interactive, 10 min batchStuck HTTP calls, hung browsers
Tokens60k–150k per runContext bloat from dumping raw HTML
Money$0.15 per support ticketDeath by a thousand retries

When a budget trips, the agent doesn’t get another chance - it hands off with a partial result and a reason code. budget_exceeded:steps in your logs is a product signal: it usually means a tool is returning garbage the model can’t parse.

On cost: I track spend per run, not per month. Monthly dashboards hide the one workflow that costs 40x the others. I’ve had a document-classification job where a single malformed PDF cost more than the previous 900 documents combined, because the extraction tool kept returning the same 30k-token wall of ligature noise and the model kept trying.

Layer 5: Dry-run mode as a first-class feature

Every write tool I build accepts a dry_run flag threaded through from the run context. In dry-run, the tool validates everything, logs the exact payload it would send, and returns a plausible fake response.

This is not a testing nicety, it’s the deployment ramp. New agent workflows run in dry-run against real production traffic for a few days. I read the proposed actions like a diff. Roughly one in three agents I build gets a design change during this phase - usually because I discover it wants to do something reasonable that I hadn’t built a tool for, or something insane that I hadn’t thought to forbid.

Layer 6: Approval gates on the irreversible

Sort actions into three buckets:

  • Reversible and cheap - drafting text, tagging, reading. Full autonomy.
  • Reversible but visible - sending a message, changing a status. Autonomy with rate limits and a kill switch.
  • Irreversible or financial - refunds, deletions, payouts, outbound email at scale. Human approval, always, until you have months of clean data.

The approval UI does not need to be fancy. My default is a Telegram message with inline Approve/Reject buttons carrying the queued action ID, plus a 30-minute expiry so nothing sits pending forever. Ten minutes of a human’s day is cheaper than one wrongly wired payout.

Layer 7: Traces you can replay, and checks that don’t need an LLM

Log every run as a structured trace: input, model version, each tool call with arguments and result, timings, cost, final outcome, reason code. Store it as JSONL in object storage; it’s cheap and grep-able. I use SQLite locally to slice it.

Then - and this is the part most teams skip - write deterministic assertions over traces. You don’t need an LLM judge to catch the majority of regressions:

  • Did any run call a tool outside its allowlist? (Should be impossible; assert it anyway.)
  • Did any run exceed 2 retries on the same tool with identical arguments?
  • Are more than 5% of runs ending in budget_exceeded?
  • Did the agent produce a final answer without calling the retrieval tool at all? (Classic silent hallucination signal.)
  • P95 step count trending up after a model or prompt change?

These run in CI over a fixture set of 50–100 recorded inputs, and in production as hourly aggregates. Cheap, fast, no flakiness. Save the LLM-as-judge evaluations for answer quality, where you genuinely need semantics - and keep them off the critical path.

Reconciliation: assume the agent’s report is fiction

An agent saying “I’ve updated the order” is a claim, not a fact. For any workflow with an external system, I run a reconciliation job that compares what the agent’s ledger claims with what the source of truth says. Payment status is the classic: your ledger says paid, the store says pending, and nobody notices for eleven days.

The job is boring and 60 lines long. It pulls both sides for the last N hours, diffs on a business key, and posts mismatches to a channel. Every non-trivial automation I’ve shipped has caught something real within the first month - usually a webhook that silently stopped, not an AI failure at all.

The rollout ladder

This is the sequence I use, and I don’t skip rungs:

  1. Shadow - agent runs on real inputs, all writes dry-run. Read the diffs daily.
  2. Assisted - agent proposes, human clicks approve. Measure approval rate; below 85% means go back to shadow.
  3. Autonomous, narrow - full autonomy for the cheapest reversible bucket only, hard rate limits, kill switch in a config value you can flip without a deploy.
  4. Autonomous, wide - expand the bucket one action type at a time, each with its own two weeks of metrics.

Building all seven layers costs me roughly a day and a half on a new project, and I reuse the same ledger/budget/trace module across clients. That’s the whole point: guardrails are infrastructure, not per-project cleverness. The prompt is the part you’ll rewrite fifty times. The enforcement layer is the part that lets you sleep while the agent runs.

If you’re about to hand an agent production credentials this week, do one thing first: add a dry_run flag and read a day of proposed actions. You will not regret it.

FAQ

Isn't a strict guardrail layer just reinventing traditional software and defeating the point of agents?

Partly, and that's fine. The agent's value is handling messy, unstructured input and deciding what should happen - that's genuinely hard to code. Execution of the decision is the easy, deterministic part, so keep it deterministic. In practice my agents are a thin reasoning layer over a set of boring, well-tested tools. When a workflow becomes so predictable that the agent adds nothing, I replace it with a plain function and save the tokens.

Do I need a full observability platform, or is logging enough for a small project?

For a solo project or an MVP, structured JSONL traces in object storage plus SQLite for querying will take you a long way - that's what I run for most client work under moderate volume. What matters is that traces are complete (every tool call with arguments and results) and replayable against a fixture set. Buy a platform when you need team-wide dashboards, per-user cost attribution or side-by-side eval runs, not before.

How do I pick sensible budget limits without guessing?

Run 50 real inputs in dry-run mode, record step count, tokens, latency and cost per run, then set each limit at roughly 1.5x the P95. Log budget trips with a reason code instead of failing silently. If a specific limit trips on more than about 5% of runs, that's usually a broken tool returning unparseable output rather than a limit that's too tight - fix the tool before you raise the ceiling.

Related articles