Skip to content
PD
EN RU
AI Automation 7 min read

Your Provider Dashboard Is Not a Budget: Hard Spend Caps for AI Automations

Provider spending limits react too late, apply to the whole org and don't know your customers. Here is the reserve-then-settle budget layer I put in front of every LLM call, proxy and paid API.

PD

Pavel Duglas

AI Automation & MVP Architect

Every few months I get the same message from a founder: “We woke up to a bill four times bigger than last month. What happened?” The answer is almost always boring. An agent got stuck in a retry loop, a scheduled job processed the same batch twelve times, or one free-tier user found out your product calls a frontier model on every keystroke. Nobody did anything malicious. The system simply had no concept of “stop, you’ve spent enough.”

Most teams think they have that concept because they set a monthly limit in the provider dashboard. They don’t. A dashboard limit is a smoke detector in the next building. In this article I’ll show the budget layer I now build into every AI automation, agent and parser before it touches production.

Why provider limits don’t protect you

Provider-side spending limits are useful as a last line of defense, but they fail as a primary control for four reasons.

They lag. Usage reporting on most platforms is delayed by minutes, sometimes longer. An agent running 20 parallel calls with large contexts can burn through a lot of money in that window.

They are org-wide. When the cap hits, everything stops. Your production chatbot, your internal tools, your nightly enrichment job. You wanted to stop one runaway workflow and you took down the whole company.

They don’t know your business. The provider has no idea that customer A pays $29 a month and customer B pays $900. It can’t tell that a document summarization should cost cents and that one costing $14 is a bug.

They only cover one vendor. My automations rarely spend money on a single API. A typical scraping pipeline pays for residential proxies, a captcha solver, an LLM for extraction and sometimes an embedding model. A BAS project running 50 threads can quietly drain a captcha balance with no LLM involved at all.

So the cap has to live in your code, at the point where the money is about to be spent.

The budget hierarchy

I think about budgets as nested scopes. Every paid call has to fit inside all of them at once.

Per-call ceiling

The smallest scope. Set max_tokens on every request, always. Estimate input tokens before sending and refuse calls where input alone exceeds a sane size for that task. If a classification prompt suddenly arrives with 180k tokens of context, something upstream broke, and the right move is to fail loudly, not to pay for it.

Per-run budget

An agent run, a pipeline execution, one processed document. This is the scope that catches loops. I give each run a budget derived from the task type: “lead enrichment run: max $0.40”, “research agent run: max $3”. When the run hits its budget, it stops and reports what it got done so far.

Per-customer budget

Daily and monthly, tied to the plan. This is what protects margins and stops one heavy user from becoming your biggest cost center. It’s also the scope that turns cost into a product decision: what happens when a customer hits the cap is a UX question, not an infrastructure one.

Per-workflow budget

Each automation gets its own daily allowance. The nightly sync can’t eat the budget of the customer-facing assistant.

Global budget

Your own hard ceiling per day, set well below the provider limit. When it’s hit, non-critical workflows pause and you get paged. The provider limit stays as the backstop above it.

Reserve, then settle

The naive implementation checks “have we spent less than the cap?” and then makes the call. That breaks the moment you have concurrency. Twenty workers all read “$9.80 spent of $10”, all decide they are fine, and you end at $16.

The pattern that works is the same one card payments use: reserve first, settle after.

  1. Before the call, estimate the maximum it could cost: input tokens plus max_tokens at the model’s output price. For a proxy or captcha call, use the known unit price.
  2. Atomically reserve that amount against every scope the call belongs to. If any scope would go over its cap, the reservation fails and the call never happens.
  3. Make the call.
  4. Settle with the actual cost from the usage data in the response. Release the difference.
  5. If the call crashes and never settles, the reservation expires after a timeout and gets settled at the reserved amount. Pessimistic, but safe.

In Postgres, a minimal version looks like this:

CREATE TABLE budget_scope (
  scope_key   text PRIMARY KEY,  -- 'run:8f2a', 'customer:412:2026-09-14', 'global:2026-09-14'
  cap_micros  bigint NOT NULL,
  spent_micros bigint NOT NULL DEFAULT 0,
  reserved_micros bigint NOT NULL DEFAULT 0
);

-- reserve against one scope; run for every scope in a single transaction
UPDATE budget_scope
SET reserved_micros = reserved_micros + $1
WHERE scope_key = $2
  AND spent_micros + reserved_micros + $1 <= cap_micros
RETURNING scope_key;

If the update returns no row for any scope, roll back the transaction and deny the call. Store money as integer micro-dollars, never floats. For high-throughput systems I move the hot counters to Redis with a Lua script doing the same check-and-increment atomically, and write the settled entries to Postgres as an append-only ledger for reporting.

The important property: there is exactly one function in the codebase allowed to call a paid API, and it goes through this reservation. No direct SDK calls scattered across the project. If a developer, or a coding agent, adds a new feature that calls the model directly, code review should treat it like a raw SQL string concatenation.

What happens when the cap is hit

A cap without a defined behavior just becomes a new kind of outage. I decide the behavior per scope in advance.

Per-call ceiling hit: reject immediately, log the input size and source. This is almost always a bug upstream.

Per-run budget hit: stop the run gracefully. Save partial results, mark the run as budget_exhausted (a distinct status from failed), and surface it. Often the right fix is not a bigger budget but a smaller task.

Per-customer budget hit: degrade, don’t die. Route to a cheaper model, switch to queued processing instead of instant, or show a clear “you’ve used today’s allowance” message with an upgrade path. This is also a great pricing signal: customers who hit caps every week are telling you what your next plan should be.

Per-workflow budget hit: pause the workflow, alert the owner, let everything else continue.

Global budget hit: pause all non-critical workflows, keep the paid customer-facing paths alive on a reserved slice of the budget, page a human.

That reserved slice matters. I usually keep 20-30% of the global daily budget available only to user-facing requests from paying customers, so a background job going wild can’t starve the product.

Loops are a cost problem first

Most runaway bills come from agents and pipelines repeating themselves. Budgets catch loops eventually, but you can catch them earlier and cheaper.

  • Count steps, not just dollars. An agent run gets a max number of tool calls. A run that hits 40 steps on a task that usually takes 6 is broken, even if each step was cheap.
  • Fingerprint repeated calls. Hash the model, the last tool call and its arguments. If the same fingerprint appears three times in one run, stop. Agents that retry the same failing search query forever are the classic case.
  • Track cost velocity. Alert when a workflow spends more in 10 minutes than it usually spends in a day. That catches the case where every individual run is within budget but a scheduler bug launched a thousand of them.

Don’t forget the non-LLM costs

In my scraping and BAS projects, the LLM is often not the biggest line item. Proxies are billed per gigabyte, captcha solvers per thousand solves, some data APIs per request, and video or image generation APIs per output, often at prices that make text tokens look free.

The same budget layer handles all of them. Every paid resource gets a unit price in config and goes through the same reserve-and-settle function. In BAS I wrap paid actions in a function that calls a small internal HTTP endpoint for the reservation before proceeding. A thread that can’t get a reservation stops cleanly instead of hammering a target site with expensive proxy traffic that’s going to fail anyway.

If you run a Telegram bot on top of a generation API, this is the difference between a fun side project and a surprise invoice. Reserve per user per day, and tell the user exactly when their allowance resets.

Making estimates good enough

You don’t need perfect estimates. You need estimates that are never too low.

  • Use the provider’s tokenizer or a close approximation for input, then add 10%.
  • Always assume the full max_tokens will be used for the reservation. Settlement fixes it afterward.
  • Keep a price table in config, versioned, with a date. When prices change you update one file, not twelve services.
  • For tools with variable costs, reserve the worst plausible case and settle the real number.

The overestimation only ties up budget for a few seconds, and it guarantees the cap is truly hard.

A rollout plan you can start today

  1. Find every place in your codebase that calls a paid API. Route them all through one function.
  2. Add max_tokens and an input size check to every LLM call.
  3. Add per-run step limits and loop fingerprinting to your agents.
  4. Introduce the ledger in log-only mode for a week: record reservations, don’t block. Look at real distributions per workflow and per customer.
  5. Set caps at roughly 3x the observed p99 and switch to enforcing mode.
  6. Define the degrade behavior for each scope and build the user-facing messages.
  7. Set your own global cap below the provider limit, and keep the provider limit as the final backstop.

Once this is in place, you stop fearing the invoice. More importantly, you can let agents do more, because the worst case is now a number you chose, not a number you discover.

FAQ

Isn't setting a spending limit in the OpenAI or Anthropic dashboard enough?

It's a useful backstop, but not a control. Dashboard limits react with a delay, apply to your whole organization at once, cover only one vendor and know nothing about your customers or workflows. When they trigger, everything stops. Enforce caps in your own code per run, per customer and per workflow, and keep the provider limit set higher as a final safety net.

How do I set the right budget for an agent run if I don't know typical costs yet?

Run the budget layer in log-only mode for a week or two. Record the cost of every run per workflow, look at the p50 and p99, and set the cap around three times the p99. Revisit it monthly. If many runs hit the cap, the task is usually too big and should be split, rather than given a bigger budget.

Won't reserving the maximum possible cost block legitimate calls near the limit?

Occasionally, yes, and that's the point of a hard cap. Reservations are held only for the duration of the call and the difference is released on settlement, so the impact is small. If it happens often for a scope, the cap is too tight for real usage, and the log data will show that clearly.

Related articles