Skip to content
PD
AI Automation 8 min read

Cancel Is Not a UI State: Building a Real Kill Switch for AI Agents

Most agent systems have a Stop button that only stops the spinner. Here is how I build cancellation that actually halts loops, aborts provider calls, undoes side effects and caps spend.

PD

Pavel Duglas

AI Automation & MVP Architect

A client pinged me last month with a simple complaint: “I cancelled the run and the invoice still went up by four dollars.” That is not a billing bug. That is a system where the Stop button changes a row in the UI and nothing else. The agent kept looping, kept calling the model, kept writing to their CRM, and the only thing that stopped was the spinner.

Cancellation is one of those features that looks like a two-hour task and turns out to be an architectural decision. If you bolt it on after launch, you get the worst possible version: a button that lies to the user while your agent keeps spending money and mutating production data. I want to walk through how I actually build it.

”Cancel” means four different things

Before you write any code, separate the meanings. Every real system needs all four, and they fail independently.

  1. UI acknowledgement. The user sees that the request was received. Cheap. This is where most implementations stop.
  2. Loop termination. The agent stops planning new steps. This is the one that saves you money.
  3. In-flight abort. The HTTP request to the model provider or the tool API is actually torn down.
  4. Side effect reconciliation. Whatever the agent already did to the outside world is either finished properly or compensated.

If you only implement 1 and 2, cancellation takes up to one full step to take effect, which with a slow reasoning model can be 90 seconds of billed tokens. If you implement 1, 2, 3 but not 4, you get half-sent email campaigns and orders created without payment records. Number 4 is where the real damage lives.

The run record is the source of truth, not the socket

The first mistake I see: cancellation is implemented as “the websocket closed” or “the HTTP request was aborted”. That works until the user’s laptop sleeps, or a reverse proxy drops an idle connection, or the tab is refreshed. Now you have cancelled a run the user still wants, or worse, you think you cancelled it and the worker never heard.

Agent runs belong in a table. Every run gets an id, a status, a deadline, a budget and a cancel flag. The client asking to cancel is just an update to that row plus a fast signal for the worker.

create table agent_runs (
  id            uuid primary key,
  tenant_id     uuid not null,
  status        text not null, -- queued|running|cancelling|cancelled|done|failed
  cancel_reason text,
  deadline_at   timestamptz not null,
  budget_cents  integer not null,
  spent_cents   integer not null default 0,
  step_count    integer not null default 0,
  created_at    timestamptz not null default now()
);

Notice cancelling as a distinct state from cancelled. The user asked, the worker has not confirmed yet. Do not show “cancelled” until the worker writes it. Otherwise you are back to lying in the UI, just with more tables.

For the signal, I use whatever is already in the stack: a Redis key with a short TTL, or Postgres LISTEN/NOTIFY. The DB row is authority, the signal is latency optimisation.

Cooperative cancellation with real checkpoints

Agents are loops. That is good news: loops have natural seams. I put a checkpoint before every model call, after every model call, before every tool call and after every tool call. The checkpoint does three things: reads the cancel flag, checks the deadline, checks the budget.

class Cancelled(Exception):
    pass

async def checkpoint(run, phase: str):
    state = await store.load(run.id)  # cached, ~1ms
    if state.status == "cancelling":
        raise Cancelled(f"user cancel at {phase}")
    if state.spent_cents >= state.budget_cents:
        raise Cancelled(f"budget exceeded at {phase}")
    if now() >= state.deadline_at:
        raise Cancelled(f"deadline exceeded at {phase}")

async def run_agent(run):
    try:
        while True:
            await checkpoint(run, "pre-plan")
            step = await plan(run)           # LLM call, cancellable
            await checkpoint(run, "pre-tool")
            result = await execute(run, step) # tool call, cancellable
            await record(run, step, result)
            if step.is_final:
                return await finish(run)
    except Cancelled as e:
        await compensate(run)
        await store.mark_cancelled(run.id, reason=str(e))

Two details that matter more than they look. First, checkpoint must be cheap, because you call it a lot. Cache the run state for a second or two, but never cache the cancel flag longer than your acceptable stop latency. Second, cancellation is an exception, not a return value. If it is a return value, someone will forget to check it in a nested helper and you will have a loop that is cancelled at the top level and still running three frames down.

Abort the model call, and be honest about what you still pay for

A cancel that only lands between steps is not enough when a single reasoning step takes a minute. Pass a cancellation token into the HTTP layer and actually tear the socket down.

In Python, run the provider call as an asyncio.Task and race it against a watcher that polls the cancel flag. In Node, use AbortController and pass the signal to fetch or the SDK. Both work. What people get wrong is the accounting.

If you stream and abort mid-response, you still pay for the input tokens and for every output token that was already generated. Aborting saves the tail, not the head. So log usage on abort. I write a partial usage record with the tokens counted from the stream so far, which means my cost dashboard and the provider invoice agree at the end of the month. If you skip this, your spend attribution quietly drifts and you will never trust it again.

Also: if you are not streaming, you get nothing back on abort and still pay for the whole generation. That alone is a decent argument for streaming every call inside an agent loop, even when the UI does not show tokens.

Side effects are the part that actually hurts

Here is the failure that cost a client real money: the agent called a tool that enqueued a batch of outbound messages, then the user cancelled. Loop stopped, socket closed, and a completely separate worker happily sent 400 messages twenty seconds later. Cancelling the agent did not cancel the work the agent had started.

The fix is boring and it works. Every tool call that touches the outside world gets:

  • An idempotency key derived from run id plus step index, so retries after a partial failure do not duplicate.
  • An intent record written before execution. Row in a table: what we are about to do, with what arguments, at what time.
  • A completion or compensation path. After execution, mark it done. On cancellation, look at every intent that is written but not done and either finish it or reverse it.

Then classify your tools explicitly. I use three buckets:

  • Safe to abandon. Reads, searches, scrapes. Abort and forget.
  • Must complete. Anything where a half-state is worse than a full state, like a two-phase write or a payment capture. Let it finish, then stop. Cancellation is not permission to corrupt data.
  • Must compensate. Created a record, scheduled a job, published a message. On cancel, delete, unschedule, publish a retraction.

And propagate. If your tool triggers downstream work, the run id has to travel with it, and that downstream worker needs the same checkpoint before it starts. Otherwise you have a cancellation boundary that ends at your process edge, which is exactly where the expensive stuff lives.

Cancel yourself before the user does

The most useful kill switch is the one nobody presses. Every run I ship has a deadline and a budget assigned at creation, from the plan tier, not hardcoded in the worker. Same checkpoint code enforces all three stop conditions: user cancel, deadline, budget. One code path, three triggers.

I also add a loop guard: max steps, plus a repetition detector. If the last three tool calls have identical arguments, the agent is stuck, and stuck agents are the single largest source of surprise invoices I have seen. Cancel with reason loop_detected and surface it. That reason field is gold during debugging, because “cancelled” with no reason tells you nothing six weeks later.

Detect the orphans

Workers get OOM-killed. Pods get rescheduled mid-step. You need a sweeper that looks for runs in running or cancelling with no heartbeat in the last N seconds, and either resumes or fails them with a clear status. Have the worker write a heartbeat timestamp on every checkpoint. You already have the hook.

Then put two numbers on a dashboard: time from cancel request to confirmed cancelled, and tokens billed after cancel request. If p95 stop latency is 40 seconds, your Stop button is decoration. If post-cancel spend is not trending to near zero, your aborts are not landing.

Test it like a real path

Cancellation is not tested by clicking the button once. Write tests that cancel at each seam: before the first model call, mid-stream, between a tool call and its completion write, and after the final step has already committed. That last one should be a no-op, not a rollback of a successful run. Add a fault injector that cancels at a random checkpoint and assert that your side effect ledger has no rows stuck in pending. A fuzzer over five checkpoints finds the bugs that manual QA never will.

Quick checklist

  • Run state in a table, not in a socket.
  • Separate cancelling from cancelled.
  • Checkpoint before and after every model and tool call.
  • Cancellation raises, it does not return.
  • Abort in-flight HTTP with a token or signal, and log partial usage.
  • Classify every tool: abandon, must complete, must compensate.
  • Intent record plus idempotency key for every external write.
  • Propagate run id to downstream workers and check it there too.
  • Deadline, budget and max-steps enforced by the same checkpoint.
  • Heartbeat plus sweeper for orphaned runs.
  • Dashboard: stop latency p95, tokens billed after cancel.

A Stop button that works is one of the cheapest trust signals you can ship in an agent product. Users forgive slow. They do not forgive a system that keeps spending their money after they told it to stop.

FAQ

Does aborting a streaming LLM request actually save money?

Partially. You still pay for the full input context and for every output token already generated before the abort. Aborting saves the remaining generation, which on a long reasoning step can be most of the output. If you are not streaming, an abort usually saves nothing at all because the provider bills the completed generation. That is a good reason to stream every call inside an agent loop even if the UI never shows the tokens, and to log a partial usage record on abort so your cost reporting matches the provider invoice.

Should cancellation roll back everything the agent already did?

No, and trying to is how you corrupt data. Classify each tool. Reads and scrapes can simply be abandoned. Operations where a half-state is worse than a full state, like a two-phase write or a payment capture, must be allowed to complete before the loop stops. Only the third category, side effects that are reversible such as created records, scheduled jobs or published messages, should be compensated. Keep an intent ledger with a row written before each external call so the compensation step knows exactly what is outstanding.

How fast should a cancel take effect before users complain?

In my experience under about three seconds feels instant and under ten is acceptable if the UI shows a "stopping" state with a reason. Beyond that people press the button again and assume it is broken. Measure p95 time from cancel request to confirmed cancelled status and treat it as a product metric. The two main levers are how often you poll the cancel flag at checkpoints and whether you actually abort in-flight provider requests instead of waiting for the current step to finish.

Related articles