Skip to content
PD
EN RU
AI Automation 7 min read

Not Every 400 Is Your Fault: Classify LLM Provider Failures Before You Retry

Most AI pipelines treat provider errors as one bucket and retry everything. Here is the five-class failure taxonomy I use, and how each class gets its own response.

PD

Pavel Duglas

AI Automation & MVP Architect

A client once called me because their document pipeline had “stopped understanding invoices.” Nothing was wrong with the invoices. The prepaid API balance had hit zero on a Saturday night, the provider started returning a 400, and the pipeline did exactly what it was told: a 400 means bad input, so mark the document as unprocessable and move on. By Monday morning, 1,800 perfectly good invoices were flagged as garbage, and someone had already started manually re-keying them.

The bug was not the billing. The bug was that the system had one idea of what an error is. In this article I will show you the failure taxonomy I now put around every LLM call in production, and why the response to an error matters more than the retry count.

HTTP status codes are too coarse for LLM APIs

Status codes were designed for generic web resources. LLM providers squeeze very different situations into the same few numbers:

  • A 400 can mean your JSON schema is invalid, your context is too long, your account has no credits, or the model refused the content.
  • A 429 can mean “slow down for two seconds” or “you have used your monthly quota, see you in 19 days.”
  • A 500 or an overloaded response usually means retry, but sometimes means this model version is having a bad hour and you should go elsewhere.
  • A timeout tells you nothing about whether the provider processed the request and billed you.

If your retry logic branches only on the status code, you will retry things that can never succeed, give up on things that would succeed in ten seconds, and blame your data for problems that live in your billing dashboard.

The five failure classes I use

Every error from a model call gets mapped into exactly one of these classes before anything else happens.

1. Transient

The provider had a hiccup: a 500, a dropped connection before sending, an overloaded response. The same request will likely succeed soon. This is the only class where plain exponential backoff with jitter is the right answer.

2. Throttled

You are going too fast. Short-window rate limits on requests or tokens per minute. The request is fine, your pacing is not. The fix is to respect the retry-after hint and slow the whole worker pool, not just this one job.

3. Exhausted

The account, not the request, is the problem. Out of credits, monthly quota reached, key revoked, organization suspended. Retrying is pointless and every job on that key will fail the same way. This is the class that destroyed those invoices.

4. Rejected

The request itself is wrong. Context too long, malformed tool definition, invalid parameter, content policy refusal. Retrying the identical request will fail forever. Sending it to a fallback provider usually fails too, because a 400k-token prompt is too big almost everywhere.

5. Ambiguous

You do not know what happened. The request was sent, then the connection timed out, or a stream cut off halfway. The model may have run, may have billed you, may have triggered a tool call. This needs its own handling because a naive retry can double the cost or double the side effect.

I keep output problems, like invalid JSON or a hallucinated field, out of this taxonomy on purpose. Those are a different layer: validation. Here we only care about whether the call itself worked.

Classify on the error body, not the status

Every serious provider returns a structured error with a type or code field and a message. Classify on that, use the status code only as a tiebreaker, and keep the mapping in one place. Here is a trimmed version of what I use in Python pipelines:

from enum import Enum

class Failure(Enum):
    TRANSIENT = "transient"
    THROTTLED = "throttled"
    EXHAUSTED = "exhausted"
    REJECTED = "rejected"
    AMBIGUOUS = "ambiguous"

EXHAUSTED_HINTS = ("credit", "balance", "billing", "quota", "insufficient", "suspended")
REJECTED_HINTS = ("context length", "too long", "invalid", "schema", "policy")

def classify(status, body, sent=True):
    if status is None:
        return Failure.AMBIGUOUS if sent else Failure.TRANSIENT
    text = (str(body.get("type", "")) + " " + str(body.get("message", ""))).lower()
    if status in (401, 403) or any(h in text for h in EXHAUSTED_HINTS):
        return Failure.EXHAUSTED
    if status == 429:
        return Failure.THROTTLED
    if status >= 500:
        return Failure.TRANSIENT
    if any(h in text for h in REJECTED_HINTS):
        return Failure.REJECTED
    return Failure.REJECTED  # unknown 4xx: fail safe, but alert

Three details matter here. First, the exhausted check runs before the 429 check, because some providers report quota exhaustion as a 429 with a billing message. Second, sent tells the classifier whether bytes actually left your server, which is the difference between a safe retry and an ambiguous one. Third, unknown 4xx errors fall into rejected but fire an alert, because an error you have never seen is exactly the one you need to look at. Each time an alert fires, I add the new hint string to the mapping. After a month the unknown bucket is nearly empty.

String matching on messages feels fragile, and it is. That is why it lives in one function with tests, and not scattered across twelve except blocks.

Each class gets a different response

This is the actual point. Classification is useless unless the reaction differs:

  • Transient: retry this job with backoff and jitter, max 4 to 6 attempts, then route to the fallback model if you have one.
  • Throttled: pause the worker pool for the retry-after window, reduce concurrency, re-enqueue the job without counting it as a failed attempt.
  • Exhausted: open a circuit breaker for that API key, stop pulling jobs, page a human, leave jobs in the queue untouched.
  • Rejected: do not retry the same payload. Either transform it (truncate, split, simplify the tool schema) or send it to a dead letter queue with the reason attached.
  • Ambiguous: check whether the work already happened, using an idempotency key or your own ledger, before retrying.

Notice that only one of five classes is “retry.” Most retry libraries assume the opposite.

Exhausted means stop the line

The expensive mistake with account-level failures is letting every job discover the problem individually. If you have 2,000 queued jobs and the balance is zero, you get 2,000 failures, 2,000 log entries, maybe 2,000 customer-facing error emails.

Instead, the first exhausted error flips a flag in Redis or your database: provider:anthropic:key_main = open. Workers check the flag before taking a job. While it is open, jobs stay in the queue in their original state. A human tops up the balance or fixes the key, a single probe request succeeds, the flag closes, and the queue drains as if nothing happened. Nobody re-keys invoices by hand.

Two extras I always add:

  1. Balance alerts before zero. If the provider exposes usage, poll it. If not, track spend in your own cost ledger and alert at 70% of the prepaid amount. Exhaustion should be a scheduled top-up, not an incident.
  2. Separate keys per workload. A runaway batch job should not drain the key your customer-facing chat depends on. One key per workload means one breaker per workload.

Fallback providers only make sense for some classes

Everyone wants multi-provider fallback now, especially with open-weight models getting good enough for many tasks. It is a great idea, but only for transient and exhausted failures. Fail over on a rejected request and you will just collect the same rejection from a second vendor, plus a second bill for input tokens.

Before a job goes to a fallback, I check capability, not just availability: does the fallback model support the context size of this payload, the tool calling format, the structured output mode? I store these as a small capability table per model. If the fallback cannot handle the job, the job waits for the breaker rather than producing a worse result silently. A clearly delayed answer beats a quietly degraded one in most business workflows.

Ambiguous failures need a ledger, not a retry

When a request times out after sending, the honest answer is “I do not know.” Treat it that way:

  • Give every model call an idempotency key derived from the job ID and attempt purpose.
  • If the call can trigger side effects through tools, those tools must check the key before acting. A timeout should never send the same email twice.
  • Log the ambiguous call in your cost ledger as “possibly billed.” If you charge customers per AI action, do not charge for ambiguous attempts until you can confirm them.
  • For streaming, keep the partial output. Sometimes 90% of a long generation arrived, and you can resume or salvage it instead of paying for the full thing again.

Test the failures on purpose

You cannot wait for a real billing outage to find out if your breaker works. I keep a fake provider in the test suite, a tiny HTTP server that returns canned responses for each class: a 500, a 429 with retry-after, a 400 with a credit message, a 400 with a context length message, and a connection that hangs after accepting the body.

The test asserts behavior, not just status: exhausted opens the breaker and leaves jobs queued, rejected goes to the dead letter queue with a reason, ambiguous does not run the tool twice. When a provider changes its error format, I capture the new body, add it to the fixtures, and the test tells me if the classifier still holds.

The checklist

If you take one thing from this article, take this list and check your pipeline against it today:

  1. All model calls go through one wrapper with one classifier.
  2. Classification uses the error body first, status code second.
  3. Only transient errors get plain backoff retries.
  4. Throttling slows the pool and does not burn attempts.
  5. Exhaustion opens a breaker per key and pages a human.
  6. Rejected payloads are transformed or dead-lettered, never retried as-is.
  7. Ambiguous calls are checked against idempotency keys before retry.
  8. Fallback routing checks model capabilities, not just uptime.
  9. A fake provider test covers every class.

None of this is glamorous. It is about 200 lines of code and an afternoon of tests. But it is the difference between an AI pipeline that pauses politely when the money runs out and one that tells your client their invoices are unreadable.

FAQ

Can't I just use a generic retry library with exponential backoff?

You can use one for the transient class, and you should. The problem is applying it to everything. A backoff library will happily retry an out-of-credits error six times per job across thousands of jobs, and it will mark a context-too-long request as a temporary failure. Put your own classifier in front and only hand transient errors to the retry library.

How do I classify errors from providers whose messages keep changing?

Keep all matching logic in one function, default unknown errors to the safe rejected class, and alert on every unknown. Save the raw error bodies as test fixtures. In practice, providers change wording occasionally but the type or code fields are fairly stable, so match on those first and use message text only as a backup signal.

Is multi-provider fallback worth it for a small MVP?

Usually not on day one. A circuit breaker plus balance alerts covers most real incidents for an MVP, because the common outage is your own account, not the provider. Add a fallback once you have paying customers who notice delays, and when you do, route only transient and exhausted failures to it after checking the fallback model can actually handle the payload.

Related articles