Rolling Out AI Agents: Guardrails and Verification
Why an agent reports success where there is none, how a verification layer works, how a constraint in code differs from one in a prompt, and what to monitor after launch.
All articles in the guide ИИ-агенты · 11
Between a working prototype and a production system stands one question: what happens when the agent is wrong. Without an answer, the prototype stays a prototype.
Why agents lie about results
This is not a defect of a particular model or a consequence of a bad prompt. The model generates a continuation of text; “task complete” is a continuation like any other, and a particularly likely one, because that is how conversations usually end.
The consequence is worth accepting as given: the agent’s report of a result is not evidence of the result. It is useful as an explanation, not as confirmation.
Typical forms:
- The tool returned an error and the agent reported success.
- The agent did three steps out of five and wrote that everything was done.
- No data was found, and the report contains plausible but invented values.
- A prompt constraint was violated and the report says it was respected.
The verification layer
Simple in principle, demanding in discipline: after the agent finishes, the system checks the fact itself.
What that looks like:
- The agent says it created a ticket - confirm the ticket exists and the mandatory fields are filled.
- The agent says it sent a message - confirm by the send identifier, not by its word.
- The agent says data was updated - read it back and compare to expectation.
The key requirement: the check must be independent of the agent. A check performed by the same agent confirms nothing; it will confirm its own conclusion.
A second level checks quality rather than fact: a separate pass that looks at the result with fresh eyes and hunts for inconsistencies. For development work that technique is covered on the blog: the verification layer.
Guardrails in code, not in prompts
The difference is fundamental.
A prompt constraint is a request. Usually respected. Occasionally not, and those occasions are what become incidents.
A code constraint is an impossibility. The agent will not delete data if it has no deletion tool and no delete permission on the database.
What moves from prompt to code first:
- Permissions. A dedicated database user with access to only the required tables and only the required operations.
- The absence of dangerous tools. Not “a delete tool with a warning in its description” but no tool.
- Human confirmation for irreversible actions: payments, mailings, deletion, publication.
- Step and cost limits per run.
- Scope limits. A tool that can update one record is safer than one that can update all of them.
- Idempotency. Repeating an action does not produce a second effect - see idempotent pipelines.
The full checklist is on the blog: guardrails for agents in production.
The launch sequence
Five steps, each removing its own class of risk:
- Observation mode. The agent proposes, a human executes. Gives the true error rate instead of the assumed one.
- A narrow slice. One task type, low volume, reversible consequences.
- Verification before widening. Until result checking is automated, do not expand: you simply will not learn about errors.
- Metrics and alerts. Measurement first, volume second.
- Gradual expansion while keeping the ability to switch it off.
Separately: there must be a way to shut the agent down in one minute, and a clear rollback. If shutting it off requires a deploy, it will not work when you need it.
What to monitor
- Completion rate by fact checks, not by reports.
- Handover rate. A rise means something changed in the input data.
- Step count, mean and maximum. A rising maximum is an early sign of loops.
- Cost per run. It rises before quality visibly drops.
- Rerun rate - a marker of trouble with errors and idempotency.
The defining operational property of agent systems: they degrade quietly. They do not crash, they start being wrong more often. Without those numbers, the difference between “working” and “half working” is invisible until complaints arrive.
How this is accounted for during the build is in agent development. The overview is in the AI agents guide.
FAQ
Why does an AI agent report a task as done when it is not?
Because the model completes a plausible ending rather than checking a fact. To it, "report success" is text like any other text. The fix is not prompt wording but checking the result in code: does the record exist, did the message send, did the state change.
Can an agent be constrained by prompt?
A prompt sets default behaviour but guarantees nothing. A real constraint is a missing tool, a missing permission, or mandatory human confirmation. If the only defence against data deletion is a line in the system prompt, there is no defence.
What should be monitored in production?
Completion rate, handover rate, mean and maximum step count, cost per run, and rerun rate. Agents do not fail visibly - they start getting things wrong more often, and only the trend in those numbers shows it.
- What an AI Agent Is and How It Differs from a ChatbotGuide
- Building an AI Agent: From Prompt to ProductionThe path from idea to a working agent: framing the task, choosing tools, the execution loop, testing before launch, and what has to be true before it reaches production.
- How to Build an AI Agent: A Step-by-Step Real CaseOne agent walked through end to end: requirements, design, tools and their descriptions, the first run and the fixes it forced, and which parts deliberately stayed ordinary code.
- AI Agents for Business: Where They Pay Off and Where They Do NotWhich processes an AI agent genuinely makes cheaper, where ordinary automation or a hire wins, how to calculate payback honestly, and the risks that rarely make it into the model.
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks