Stop Running Agents Inside HTTP Requests: The Job Queue Architecture I Use
An AI agent is a long-running background job, not a web request. Here is the queue, run-state and checkpointing architecture I use in production, with schema, retry rules and cancellation.
Pavel Duglas
AI Automation & MVP Architect
Every AI product I have been called in to rescue in the last year had the same bug at the bottom of the stack. Not a bad prompt. Not the wrong model. The agent was running inside an HTTP request. Someone wired POST /api/agent straight to a loop of tool calls and hoped the browser would wait. It waits for about 30 seconds, then a load balancer somewhere kills the connection, the client retries, and now two copies of the agent are editing the same CRM record.
An agent is not a request. It is a batch job that happens to talk. Once you accept that, most of the hard problems (timeouts, retries, cost blowups, cancellation, progress UI) turn into solved infrastructure problems you already know how to handle.
Why the request path always loses
HTTP requests are optimised for the opposite of what an agent does. Requests are short, stateless, cheap to retry, and safe to drop. Agent runs are long, stateful, expensive to retry, and dangerous to drop halfway.
Concrete things that kill agents in the request path:
- Proxy timeouts you do not control. Cloudflare, nginx, Heroku routers, API Gateway. Each has its own idle limit. Your 4-minute research agent will meet the strictest one.
- Client retries you did not ask for. Fetch wrappers, mobile networks, users hitting the button twice. If the agent has side effects, every retry is a duplicate side effect.
- Deploys. Any deploy during a 6-minute run kills the run with no record of where it stopped.
- No backpressure. Ten users click at once and you have ten concurrent agents, each holding a web worker and burning tokens. Your web server dies of something that is not traffic.
- No observability. When it fails you have one log line and a 502. You cannot answer “what step was it on?” because there were no steps, only a stack frame.
In serverless the failure is even sharper: the function hits its wall clock limit mid tool call, the platform reports success or a generic timeout, and the model provider still bills you for the tokens you already spent.
The two-layer model
Split the system in half and never blur the line.
Request layer. Validates input, creates a run row, enqueues it, returns 202 Accepted with a run_id. Target: under 100ms. It never calls a model. Ever.
Work layer. A separate worker process (different deploy unit, different scaling rules) that pulls runs off a queue, executes steps, writes state after each one, and emits events. It can run for 20 minutes and nobody cares.
The client then polls or subscribes to GET /api/runs/:id. That is the whole architecture. Everything below is detail on how to make the work layer survive contact with reality.
The minimum viable queue is Postgres
You do not need Kafka. For anything under a few thousand runs a day, Postgres is the right answer because your run state and your queue stay in the same transaction, which removes an entire class of “job enqueued but row missing” bugs.
Two tables. Runs and steps.
create table agent_runs (
id uuid primary key default gen_random_uuid(),
tenant_id uuid not null,
kind text not null, -- 'research', 'enrich_lead', ...
input jsonb not null,
status text not null default 'queued',
-- queued|running|waiting_human|done|failed|cancelled
attempt int not null default 0,
max_attempts int not null default 3,
cursor jsonb not null default '{}', -- where to resume
result jsonb,
error text,
idempotency_key text unique,
cost_cents int not null default 0,
locked_by text,
locked_until timestamptz,
run_after timestamptz not null default now(),
created_at timestamptz not null default now(),
updated_at timestamptz not null default now()
);
create table agent_steps (
id bigserial primary key,
run_id uuid not null references agent_runs(id) on delete cascade,
seq int not null,
name text not null,
input jsonb,
output jsonb,
tokens_in int,
tokens_out int,
status text not null,
started_at timestamptz not null default now(),
ended_at timestamptz,
unique (run_id, seq)
);
Claiming work, with a lease so a dead worker does not block the run forever:
update agent_runs
set status = 'running',
locked_by = $1,
locked_until = now() + interval '5 minutes',
attempt = attempt + 1,
updated_at = now()
where id = (
select id from agent_runs
where status in ('queued')
and run_after <= now()
order by run_after
for update skip locked
limit 1
)
returning *;
for update skip locked is the one Postgres feature that makes this whole pattern work with many workers. The worker heartbeats locked_until forward every 30 seconds while it works. A separate reaper marks anything with locked_until < now() back to queued so crashed runs get picked up.
Checkpoint after every step, not at the end
The single biggest upgrade over “agent in a request” is that the run has a resume point. If your agent is one giant while (true) loop in memory, a restart loses everything, including the expensive parts.
So write the loop as explicit steps with persisted output:
async function runStep(run, step) {
const existing = await findStep(run.id, step.seq);
if (existing?.status === 'done') return existing.output; // replay, free
const out = await execute(step); // model or tool call
await saveStep(run.id, step.seq, out);
await updateCursor(run.id, nextCursor(step, out));
return out;
}
On retry the worker replays completed steps from the database instead of re-calling the model. A run that died on step 7 of 9 costs you two steps to finish, not nine. On a real client project this cut retry token spend by about 70% because most failures happened late in the run, in the write phase, not the reasoning phase.
This also gives you a free audit trail. When a client asks why the agent emailed the wrong contact, you open agent_steps and read the actual tool inputs.
Retry rules that do not make things worse
Blind retries on an agent are how you send four invoices. Classify errors before you decide.
- Retryable, no state change: 429, 5xx from the model provider, connection resets, timeouts on read-only tools. Exponential backoff with jitter, set
run_after, back toqueued. - Not retryable: schema validation failures, 400s, refusals, budget exceeded. Fail fast, surface to the human.
- Ambiguous side-effect failures: timeout on a POST to a payment or CRM endpoint. Never blind retry. Either the tool call carries an idempotency key the vendor honours, or the step goes to
waiting_human. “Ambiguous write” is the one case where a human is cheaper than clever code.
Cap attempts at 3. A run that fails three times is a bug, not bad luck, and a poison run in an infinite retry loop will happily spend your monthly model budget over a weekend.
Concurrency, fairness and a hard cost ceiling
The queue is also where you enforce economics, which you cannot do in the request path.
- Global worker concurrency. Start at 5. Not 50. Model rate limits will punish you long before your CPU does.
- Per-tenant concurrency. One enterprise client uploading a 3,000-row CSV must not starve everyone else. Add
and (select count(*) from agent_runs r2 where r2.tenant_id = agent_runs.tenant_id and r2.status = 'running') < 3to the claim query, or keep a simple per-tenant counter table. - Cost ceiling per run. Increment
cost_centsafter every step. When it crosses the limit, stop and fail with a clear reason. This has saved me twice from a tool-loop that kept re-searching the same query. - Separate queues by shape. Cheap 5-second classifications and 8-minute research runs should not share a worker pool, or the long ones will block the short ones and your latency graph will look haunted.
Cancellation, pause and human-in-the-loop
Because the run is a row, cancellation is just update agent_runs set status = 'cancelled'. The worker checks the status between steps and bails. This is the cheapest correct implementation of a stop button and users care about it far more than you expect.
The same mechanism gives you approvals for free. A step that needs sign-off writes its proposed action into the step row and sets the run to waiting_human. The worker releases the lease and moves on to other work. When someone approves in the UI, you set the status back to queued and the run resumes at the cursor, possibly hours later. No connection held open, no state in memory. Try building that inside an HTTP handler.
Progress without a websocket cathedral
Founders ask for streaming. Users actually want to know it is alive and roughly where it is. Poll GET /api/runs/:id every 2 seconds and return status, current step name, completed step count and partial results. That is 20 lines of code and it works on flaky mobile networks, through corporate proxies, and across your next deploy.
Add server-sent events later if the product genuinely needs token-level streaming, and even then stream from the worker into a channel, not from a handler that owns the agent loop.
Do you need a workflow engine?
Eventually maybe. Durable execution frameworks solve exactly this problem more thoroughly, and if you already run one, use it. But I have shipped this Postgres pattern into production more times than anything else because it takes an afternoon, has no new infrastructure, and the state is in tables your team can already query with psql at 2am. Reach for the heavier tool when you have fan-out over hundreds of parallel branches or multi-day workflows with real compensation logic.
Migration checklist
If you have an agent in a request handler right now, here is the order I do it in:
- Add the two tables. Handler creates a run and returns
202with the id. - Extract the agent loop into a
worker.tsyou can run as a separate process. - Add the claim query with
skip lockedplus a lease and a reaper. - Split the loop into named steps and persist each output.
- Add the replay check so retries skip completed steps.
- Classify errors into retryable, fatal and ambiguous.
- Add per-tenant concurrency and a per-run cost ceiling.
- Add cancel and
waiting_human. - Point the UI at polling.
Steps 1 to 3 alone remove the timeout class of bugs. Step 5 is where the cost savings show up. Step 8 is where clients start trusting the thing enough to give it write access.
The request layer should be boring, fast and dumb. All the intelligence lives in a worker that is allowed to take its time, save its place, and be interrupted.
FAQ
Can I keep the agent in the request path if runs are only 10 seconds?
You can, until a model provider has a slow day and your 10 second run becomes a 90 second run. My rule: if any single run can plausibly exceed 15 seconds or has side effects outside your own database, move it to a worker. The migration costs an afternoon early on and a painful week after you have users.
How do I stop duplicate runs when a user double-clicks the button?
Put a unique idempotency key on the run row, derived from tenant plus input hash plus a client-generated token. The insert either creates a new run or conflicts, and on conflict you return the existing run id. The client gets the same run either way, and the queue never sees two copies.
What do I do about a run that died mid-write to an external system?
Treat it as ambiguous, not failed. Never blind retry an unconfirmed write. Either the vendor endpoint accepts an idempotency key you can safely resend, or you check current remote state before acting, or the run goes to waiting_human with the proposed action recorded so a person can decide in one click.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks