Skip to content
PD
AI Automation 8 min read

Your AI Agent Is Lying About Being Done: Build a Verification Layer

AI agents report success they never achieved. Here is the verification layer I put in front of every production agent - acceptance contracts, deterministic checks, evidence artifacts and retry budgets.

PD

Pavel Duglas

AI Automation & MVP Architect

Last month a client sent me a Slack screenshot of their agent’s final message: “Successfully updated 412 customer records in the CRM. All validations passed.” The CRM had 0 updated records. The agent had gotten a 403 on the first write, retried twice, decided the endpoint was “eventually consistent”, and wrote a triumphant summary. Nobody noticed for nine days.

This is the single most expensive failure mode in agent systems right now, and it is not a model problem. It is an architecture problem. An agent that produces text as its final artifact will always be able to produce text that looks like success. If you accept that text as proof, you have built a system with no feedback loop. Here is the layer I put between every agent and every consequence it can cause.

The failure mode: unverified success

LLM agents do not lie in the human sense. They optimize for a plausible continuation of a conversation where the task gets completed. When the tool call fails, the most plausible continuation is still “task completed” - because that is what 99% of the training data looks like at the end of a task thread.

In production I see three flavors:

Self-report as evidence. The agent says it wrote the file, sent the email, updated the row. There is no artifact. The only evidence is the sentence itself.

Partial completion, total claim. It processed 12 of 412 items, hit a rate limit, then summarized the batch as done. Numbers in the summary are frequently hallucinated round-ish figures.

Silent scope reduction. You asked it to parse 50 product pages. It found 6, decided the site structure changed, scraped the sitemap instead, got 6 URLs of category pages, and reported “6 products extracted, site appears to have reduced inventory.” Technically it did something. It just wasn’t your task.

All three are invisible if your pipeline ends with the agent’s own message.

Rule one: an agent never grades its own homework

The fix is boring and old: separate the actor from the verifier. Not “add a self-reflection step to the prompt” - that is the same model, in the same context window, with the same incentive to close the loop. I mean a physically separate process that has never seen the agent’s reasoning and only looks at the world state.

My standard stack has four layers, cheapest first. Anything that fails at layer N never reaches layer N+1.

Layer 1: schema

The agent’s final output is never prose. It is a structured object with a fixed schema, and the fields must be things that are checkable outside the agent.

{
  "status": "completed",
  "records_written": 412,
  "target_table": "crm_contacts",
  "write_ids": ["c_8812", "c_8813", "..."],
  "failures": [],
  "evidence": {
    "http_statuses": [200, 200, 200],
    "screenshot": "s3://runs/9f2a/final.png"
  }
}

Notice write_ids. A count is easy to hallucinate. A list of 412 IDs that must all exist in the database is not. Design your schema so that lying requires fabricating verifiable primary keys - then verify them.

Layer 2: deterministic checks

Plain code. No model. This is where 90% of fake successes die.

def verify(claim, db):
    if claim["status"] != "completed":
        return fail("agent did not claim completion")
    if len(claim["write_ids"]) != claim["records_written"]:
        return fail("count mismatch in claim itself")
    found = db.count_ids(claim["target_table"], claim["write_ids"])
    if found != claim["records_written"]:
        return fail(f"claimed {claim['records_written']}, found {found}")
    stale = db.count_untouched_since(run_started_at, claim["write_ids"])
    if stale:
        return fail(f"{stale} rows exist but were not modified in this run")
    return ok()

That last check catches my favorite failure: the agent “finds” existing records and reports them as writes. Always check updated_at, not just existence.

Deterministic checks are cheap, instant, and you can run them on every single run forever. Write them before you write the agent prompt. If you cannot express “done” as code, you do not yet know what you are automating.

Layer 3: an independent judge, on evidence only

Some tasks genuinely need judgment: “is this extracted product description accurate”, “does this generated reply answer the customer’s question”. Here I use a second model call, but with three hard constraints:

  1. It sees only the original task spec and the produced artifacts. It never sees the actor’s reasoning, plan, or summary. Reasoning is persuasive; that is the problem.
  2. It answers a closed question with a rubric, not “is this good?” - for example: does_output_contain_price: yes/no, price_matches_source_html: yes/no, language_matches_request: yes/no.
  3. It runs on a different model than the actor when possible. Same-model judges share the same blind spots, especially on formatting and instruction-following.

A judge that returns three booleans is worth ten judges that return a paragraph of praise.

Layer 4: human gate, but only on the diff

Humans are the most expensive verifier, so use them on the smallest possible surface. Not “review this run” - instead: “14 of 412 records failed layer 2, here they are with the raw input and the agent’s attempt.” Approving or rejecting 14 rows takes two minutes. Reviewing a 3000-token summary takes ten and teaches you nothing.

Evidence artifacts beat narration

The practical rule I now enforce on every project: an action is not done until it left a trace outside the agent’s context window.

  • Database write → return the IDs, then re-read them in a separate connection.
  • HTTP call → log status code, response body hash, and latency. Store them.
  • File written → hash it, record byte size and line count.
  • Email or Telegram message sent → store the provider’s message ID and fetch it back.
  • Browser action → a screenshot plus the post-action DOM snippet you asserted on.

In Browser Automation Studio work this matters more than anywhere else, because the UI lies too. A click that hit an invisible overlay, a form that reset silently, a modal that swallowed the submit - the script continues happily. So my BAS flows end each meaningful step with an assertion against a result element, not the element I clicked: the order number on the confirmation page, the row count in the table, the balance change. Then a screenshot named with the run ID goes to storage. When a client asks “did it really submit 200 forms”, I send 200 screenshots, not a log line.

Write the acceptance contract before the prompt

This single habit changed my success rate more than any model upgrade. Before I write a system prompt, I write a contract:

task: enrich_leads
inputs: csv with >= 1 row, columns [company, domain]
success:
  - every input row has an output row (no silent drops)
  - email field matches RFC 5322 or is explicitly null
  - domain of email == input domain OR flagged as third_party
  - source_url present and returns 200 on recheck
budget:
  max_tool_calls: 40
  max_usd: 0.35
  max_wall_clock_s: 180
failure_mode: partial_ok  # write good rows, queue the rest

The contract becomes the deterministic validator, the judge rubric, and the prompt, in that order. It also forces you to decide something founders love to avoid: what happens on partial success. partial_ok versus all_or_nothing is a business decision, not an engineering one, and the agent will pick for you if you don’t.

Retry budgets, not retry loops

Self-correction is good. Infinite self-correction is a way to spend $60 discovering an API key expired. My rules:

  • Max 2 corrective attempts per task, and each one must receive the validator’s error message as input, not just “try again”. A retry without new information is a coin flip.
  • Second attempt gets a reduced scope: process the failed subset only.
  • Hard caps on tool calls, tokens and wall clock, enforced by the runner, not by prompt instruction. Models ignore “do not use more than 40 calls”. Runners don’t.
  • Same error twice → stop, mark the run blocked, page a human. Three identical failures is a signal about the environment, not the agent.

And log a verification_failed metric separately from task_failed. When verification_failed spikes, either the world changed (site redesign, API deprecation) or your model provider quietly shifted behavior. It is the best early-warning signal in an agent system.

What to log for every run

Keep it flat and queryable, one row per run:

run ID, task type, input hash, contract version, model and version, tool calls used, tokens, USD cost, wall clock, agent claim, validator verdict, validator failure reasons, evidence URIs, retry count, final state.

With that table you can answer the questions that actually matter: which task types have the lowest first-pass verification rate, what a verified success costs, whether last week’s prompt change helped or just felt better.

A checklist you can implement this week

  1. Pick your riskiest agent. Write its acceptance contract in YAML. 20 minutes.
  2. Make its final output a JSON schema with checkable identifiers, not prose.
  3. Write the deterministic validator as separate code with its own DB connection.
  4. Store one evidence artifact per side effect.
  5. Cap tool calls, cost and time in the runner.
  6. Add the run log table and a dashboard with one number: first-pass verification rate.
  7. Only after all that, consider a model judge for the subjective parts.

None of this is glamorous and none of it requires a framework. It is the difference between a demo that impresses a client and a system they can leave running while they sleep. Trust in automation is not earned by the model - it is earned by the layer that refuses to believe the model.

FAQ

Isn't adding a verifier just doubling my costs?

No, because most verification is free. Layer 1 (schema validation) and layer 2 (deterministic checks against your database, HTTP statuses, file hashes) cost nothing in tokens - they are plain code. Only genuinely subjective checks need a second model call, and those should be tiny prompts returning booleans, not paragraphs. In practice my verified pipelines cost 5–15% more per run and eliminate the multi-day silent failures that actually cost real money.

Can I just tell the agent to double-check its own work in the prompt?

It helps marginally and it is not a substitute. Self-reflection happens in the same context window, with the same commitment to the story it already told, and it can only check what it believes it observed. If a tool call silently failed, the agent's internal model of the world is already wrong - reflecting on wrong state produces confident wrong conclusions. Verification must read the real world state from a separate process.

How do I verify agents doing browser automation where there is no API to check against?

You assert on result elements rather than actions. After a submit, wait for and extract something that only exists on success: an order number, a confirmation ID, a changed row count, a balance. Store that value plus a screenshot named with the run ID. Then, where possible, do an independent recheck later - reload the account page in a fresh session and confirm the record exists. That fresh-session recheck is the closest thing to a real API assertion in browser work.

Related articles