Skip to content
PD
AI Automation 8 min read

Stop Running Agents Inside HTTP Requests: The Job Queue Architecture I Use

An AI agent is a long-running background job, not a web request. Here is the queue, run-state and checkpointing architecture I use in production, with schema, retry rules and cancellation.

PD

Pavel Duglas

AI Automation & MVP Architect

Every AI product I have been called in to rescue in the last year had the same bug at the bottom of the stack. Not a bad prompt. Not the wrong model. The agent was running inside an HTTP request. Someone wired POST /api/agent straight to a loop of tool calls and hoped the browser would wait. It waits for about 30 seconds, then a load balancer somewhere kills the connection, the client retries, and now two copies of the agent are editing the same CRM record.

An agent is not a request. It is a batch job that happens to talk. Once you accept that, most of the hard problems (timeouts, retries, cost blowups, cancellation, progress UI) turn into solved infrastructure problems you already know how to handle.

Why the request path always loses

HTTP requests are optimised for the opposite of what an agent does. Requests are short, stateless, cheap to retry, and safe to drop. Agent runs are long, stateful, expensive to retry, and dangerous to drop halfway.

Concrete things that kill agents in the request path:

  • Proxy timeouts you do not control. Cloudflare, nginx, Heroku routers, API Gateway. Each has its own idle limit. Your 4-minute research agent will meet the strictest one.
  • Client retries you did not ask for. Fetch wrappers, mobile networks, users hitting the button twice. If the agent has side effects, every retry is a duplicate side effect.
  • Deploys. Any deploy during a 6-minute run kills the run with no record of where it stopped.
  • No backpressure. Ten users click at once and you have ten concurrent agents, each holding a web worker and burning tokens. Your web server dies of something that is not traffic.
  • No observability. When it fails you have one log line and a 502. You cannot answer “what step was it on?” because there were no steps, only a stack frame.

In serverless the failure is even sharper: the function hits its wall clock limit mid tool call, the platform reports success or a generic timeout, and the model provider still bills you for the tokens you already spent.

The two-layer model

Split the system in half and never blur the line.

Request layer. Validates input, creates a run row, enqueues it, returns 202 Accepted with a run_id. Target: under 100ms. It never calls a model. Ever.

Work layer. A separate worker process (different deploy unit, different scaling rules) that pulls runs off a queue, executes steps, writes state after each one, and emits events. It can run for 20 minutes and nobody cares.

The client then polls or subscribes to GET /api/runs/:id. That is the whole architecture. Everything below is detail on how to make the work layer survive contact with reality.

The minimum viable queue is Postgres

You do not need Kafka. For anything under a few thousand runs a day, Postgres is the right answer because your run state and your queue stay in the same transaction, which removes an entire class of “job enqueued but row missing” bugs.

Two tables. Runs and steps.

create table agent_runs (
  id            uuid primary key default gen_random_uuid(),
  tenant_id     uuid not null,
  kind          text not null,           -- 'research', 'enrich_lead', ...
  input         jsonb not null,
  status        text not null default 'queued',
                -- queued|running|waiting_human|done|failed|cancelled
  attempt       int  not null default 0,
  max_attempts  int  not null default 3,
  cursor        jsonb not null default '{}',  -- where to resume
  result        jsonb,
  error         text,
  idempotency_key text unique,
  cost_cents    int not null default 0,
  locked_by     text,
  locked_until  timestamptz,
  run_after     timestamptz not null default now(),
  created_at    timestamptz not null default now(),
  updated_at    timestamptz not null default now()
);

create table agent_steps (
  id         bigserial primary key,
  run_id     uuid not null references agent_runs(id) on delete cascade,
  seq        int not null,
  name       text not null,
  input      jsonb,
  output     jsonb,
  tokens_in  int,
  tokens_out int,
  status     text not null,
  started_at timestamptz not null default now(),
  ended_at   timestamptz,
  unique (run_id, seq)
);

Claiming work, with a lease so a dead worker does not block the run forever:

update agent_runs
set status = 'running',
    locked_by = $1,
    locked_until = now() + interval '5 minutes',
    attempt = attempt + 1,
    updated_at = now()
where id = (
  select id from agent_runs
  where status in ('queued')
    and run_after <= now()
  order by run_after
  for update skip locked
  limit 1
)
returning *;

for update skip locked is the one Postgres feature that makes this whole pattern work with many workers. The worker heartbeats locked_until forward every 30 seconds while it works. A separate reaper marks anything with locked_until < now() back to queued so crashed runs get picked up.

Checkpoint after every step, not at the end

The single biggest upgrade over “agent in a request” is that the run has a resume point. If your agent is one giant while (true) loop in memory, a restart loses everything, including the expensive parts.

So write the loop as explicit steps with persisted output:

async function runStep(run, step) {
  const existing = await findStep(run.id, step.seq);
  if (existing?.status === 'done') return existing.output; // replay, free

  const out = await execute(step);                          // model or tool call
  await saveStep(run.id, step.seq, out);
  await updateCursor(run.id, nextCursor(step, out));
  return out;
}

On retry the worker replays completed steps from the database instead of re-calling the model. A run that died on step 7 of 9 costs you two steps to finish, not nine. On a real client project this cut retry token spend by about 70% because most failures happened late in the run, in the write phase, not the reasoning phase.

This also gives you a free audit trail. When a client asks why the agent emailed the wrong contact, you open agent_steps and read the actual tool inputs.

Retry rules that do not make things worse

Blind retries on an agent are how you send four invoices. Classify errors before you decide.

  • Retryable, no state change: 429, 5xx from the model provider, connection resets, timeouts on read-only tools. Exponential backoff with jitter, set run_after, back to queued.
  • Not retryable: schema validation failures, 400s, refusals, budget exceeded. Fail fast, surface to the human.
  • Ambiguous side-effect failures: timeout on a POST to a payment or CRM endpoint. Never blind retry. Either the tool call carries an idempotency key the vendor honours, or the step goes to waiting_human. “Ambiguous write” is the one case where a human is cheaper than clever code.

Cap attempts at 3. A run that fails three times is a bug, not bad luck, and a poison run in an infinite retry loop will happily spend your monthly model budget over a weekend.

Concurrency, fairness and a hard cost ceiling

The queue is also where you enforce economics, which you cannot do in the request path.

  • Global worker concurrency. Start at 5. Not 50. Model rate limits will punish you long before your CPU does.
  • Per-tenant concurrency. One enterprise client uploading a 3,000-row CSV must not starve everyone else. Add and (select count(*) from agent_runs r2 where r2.tenant_id = agent_runs.tenant_id and r2.status = 'running') < 3 to the claim query, or keep a simple per-tenant counter table.
  • Cost ceiling per run. Increment cost_cents after every step. When it crosses the limit, stop and fail with a clear reason. This has saved me twice from a tool-loop that kept re-searching the same query.
  • Separate queues by shape. Cheap 5-second classifications and 8-minute research runs should not share a worker pool, or the long ones will block the short ones and your latency graph will look haunted.

Cancellation, pause and human-in-the-loop

Because the run is a row, cancellation is just update agent_runs set status = 'cancelled'. The worker checks the status between steps and bails. This is the cheapest correct implementation of a stop button and users care about it far more than you expect.

The same mechanism gives you approvals for free. A step that needs sign-off writes its proposed action into the step row and sets the run to waiting_human. The worker releases the lease and moves on to other work. When someone approves in the UI, you set the status back to queued and the run resumes at the cursor, possibly hours later. No connection held open, no state in memory. Try building that inside an HTTP handler.

Progress without a websocket cathedral

Founders ask for streaming. Users actually want to know it is alive and roughly where it is. Poll GET /api/runs/:id every 2 seconds and return status, current step name, completed step count and partial results. That is 20 lines of code and it works on flaky mobile networks, through corporate proxies, and across your next deploy.

Add server-sent events later if the product genuinely needs token-level streaming, and even then stream from the worker into a channel, not from a handler that owns the agent loop.

Do you need a workflow engine?

Eventually maybe. Durable execution frameworks solve exactly this problem more thoroughly, and if you already run one, use it. But I have shipped this Postgres pattern into production more times than anything else because it takes an afternoon, has no new infrastructure, and the state is in tables your team can already query with psql at 2am. Reach for the heavier tool when you have fan-out over hundreds of parallel branches or multi-day workflows with real compensation logic.

Migration checklist

If you have an agent in a request handler right now, here is the order I do it in:

  1. Add the two tables. Handler creates a run and returns 202 with the id.
  2. Extract the agent loop into a worker.ts you can run as a separate process.
  3. Add the claim query with skip locked plus a lease and a reaper.
  4. Split the loop into named steps and persist each output.
  5. Add the replay check so retries skip completed steps.
  6. Classify errors into retryable, fatal and ambiguous.
  7. Add per-tenant concurrency and a per-run cost ceiling.
  8. Add cancel and waiting_human.
  9. Point the UI at polling.

Steps 1 to 3 alone remove the timeout class of bugs. Step 5 is where the cost savings show up. Step 8 is where clients start trusting the thing enough to give it write access.

The request layer should be boring, fast and dumb. All the intelligence lives in a worker that is allowed to take its time, save its place, and be interrupted.

FAQ

Can I keep the agent in the request path if runs are only 10 seconds?

You can, until a model provider has a slow day and your 10 second run becomes a 90 second run. My rule: if any single run can plausibly exceed 15 seconds or has side effects outside your own database, move it to a worker. The migration costs an afternoon early on and a painful week after you have users.

How do I stop duplicate runs when a user double-clicks the button?

Put a unique idempotency key on the run row, derived from tenant plus input hash plus a client-generated token. The insert either creates a new run or conflicts, and on conflict you return the existing run id. The client gets the same run either way, and the queue never sees two copies.

What do I do about a run that died mid-write to an external system?

Treat it as ambiguous, not failed. Never blind retry an unconfirmed write. Either the vendor endpoint accepts an idempotency key you can safely resend, or you check current remote state before acting, or the run goes to waiting_human with the proposed action recorded so a person can decide in one click.

Related articles