"Up" Is Not "Ready": Health Checks That Actually Catch AI Pipeline Failures
A running process tells you nothing about an AI pipeline. Here is the three-layer health check system I install in every automation: liveness, readiness, and capability canaries with golden inputs.
Pavel Duglas
AI Automation & MVP Architect
Last spring a client pinged me at 9am: “Everything is green, but the sales team says the enrichment data has been garbage since yesterday.” Their monitoring checked one thing - whether the container answered HTTP 200 on /health. The handler returned the string ok. It had never touched the model provider, the vector store, or the proxy pool. Meanwhile the OpenAI-compatible endpoint they used through a gateway had started returning empty completions for a specific model alias, and their code happily wrote empty strings into Postgres for eleven hours.
A running process is not a working pipeline. This is doubly true for AI automation, because the failure mode is rarely a crash. It is plausible garbage. Here is the health check architecture I now install in every project before I ship any agent or parser to production.
Why the usual health check lies to you
A traditional web service has a short dependency chain: process, database, maybe a cache. If the process runs and the DB responds, you are mostly fine.
An AI automation has a long chain of things you do not control:
- a model provider (or three, if you route)
- a token quota and a rate limit that changes without notice
- a vector store that can be up but empty after a bad reindex
- a headless browser or BAS profile whose session has expired
- a proxy pool where 70% of IPs are now blocked
- prompt templates that someone edited in a web UI
- output schemas that the model silently stops honoring after a model version bump
Every single one of these can break while your process stays perfectly alive. And because LLM output is text, a broken pipeline produces confident nonsense instead of a stack trace. Your alerting sees zero errors and 100% uptime.
So stop asking “is it running?” and start asking three separate questions.
Layer 1: liveness (is the process alive?)
This one stays dumb on purpose. It answers only: should the orchestrator restart me?
@app.get("/livez")
def livez():
return {"status": "alive"}
No dependency calls. Ever. If you check Postgres in your liveness probe, a five-minute database blip will make Kubernetes restart every replica you have, turning a small outage into a large one. I have watched this happen. Keep liveness selfish.
Layer 2: readiness (can I do useful work right now?)
Readiness checks dependencies, with timeouts, in parallel, and cached. It answers: should traffic or queue jobs be routed to me?
The rules I follow:
- Every check has a hard timeout (300-800ms typical).
- Checks run concurrently, not in a loop.
- Results are cached for 5-15 seconds so a monitoring flood cannot DDoS your own provider.
- Checks are read-only. Never mutate production data to prove you are healthy.
- Distinguish
degradedfromdown.
import asyncio, time
CACHE = {"at": 0, "payload": None}
async def check(name, coro, timeout=0.8):
started = time.perf_counter()
try:
await asyncio.wait_for(coro, timeout=timeout)
ok = True
detail = None
except Exception as e:
ok = False
detail = f"{type(e).__name__}: {e}"[:200]
return name, {
"ok": ok,
"ms": round((time.perf_counter() - started) * 1000),
"detail": detail,
}
@app.get("/readyz")
async def readyz():
if time.time() - CACHE["at"] < 10 and CACHE["payload"]:
return CACHE["payload"]
results = await asyncio.gather(
check("db", db.execute("select 1")),
check("queue", redis.ping()),
check("vectors", qdrant.count("docs", exact=False)),
check("model_primary", probe_model(PRIMARY)),
check("model_fallback", probe_model(FALLBACK)),
check("proxies", proxy_pool.healthy_ratio()),
)
checks = dict(results)
critical = ["db", "queue"]
down = any(not checks[c]["ok"] for c in critical)
model_down = not checks["model_primary"]["ok"] and not checks["model_fallback"]["ok"]
status = "down" if (down or model_down) else (
"degraded" if any(not v["ok"] for v in checks.values()) else "ready"
)
payload = {"status": status, "checks": checks, "version": BUILD_SHA}
CACHE.update(at=time.time(), payload=payload)
return payload
The AI-specific checks people skip
Model probe, not model ping. Do not just check that the base URL returns 200. Send a one-token request and assert you get non-empty content back:
async def probe_model(model):
r = await client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "ping"}],
max_tokens=1,
temperature=0,
)
assert (r.choices[0].message.content or "").strip(), "empty completion"
That single assertion would have caught my client’s eleven-hour outage in the first minute. Cost: fractions of a cent per check.
Vector store emptiness. A collection that exists with 0 vectors is the classic post-reindex disaster. Check that the count is above a floor you know is sane (count > 0.8 * expected), not just that the connection works.
Prompt and schema versions. Expose the SHA-256 of every prompt template and the JSON schema you validate against. When someone edits a prompt in a UI and quality drops, you want the hash in your logs and in /readyz output so you can correlate in ten seconds instead of two hours.
Proxy and session health. For browser automation and BAS-driven jobs, readiness means: is my proxy pool above 50% healthy, and is at least one profile logged in? I run a cheap authenticated request to a stable endpoint on the target and check for the logged-in marker, not for HTTP 200. A login wall returns 200 all day long.
Quota headroom. If your provider exposes remaining quota or you track spend yourself, report degraded when you are past 85% of the daily budget. That is your chance to switch to a cheaper model before you switch to zero output.
Layer 3: capability canaries (is the output still correct?)
Readiness proves your dependencies answer. It does not prove your pipeline still produces correct results. For that you need golden inputs run on a schedule.
Pick 5 to 10 real inputs with known-good expectations. Not full ground truth - assertions.
GOLDEN = [
{"id": "inv-01", "file": "fixtures/invoice_de.pdf",
"assert": {"currency": "EUR", "total": 1240.50, "vat_present": True}},
{"id": "cls-03", "text": "Refund my last order please",
"assert": {"intent": "refund_request"}},
]
Run them every 15 minutes through the real pipeline (same prompts, same model routing, same parsing) and check:
- output validates against the JSON schema
- required fields are non-empty
- exact-match assertions hold for deterministic fields (currency, ID, enum labels)
- numeric fields are within tolerance
- latency and token usage are within 2x of the 30-day baseline
Two or more canary failures in a row is a page. One failure is a warning, because models are stochastic and you will burn out your team paging on single flakes.
This is the single highest-value monitoring I have added to AI projects in the past two years. It costs a few dollars a month and it catches model deprecations, silent provider routing changes, prompt edits, and parser regressions before customers do.
Layer 4: freshness heartbeats (did the work actually happen?)
Canaries test the code path. Heartbeats test that the pipeline ran at all. Every scheduled job writes a row when it finishes successfully:
create table pipeline_heartbeat (
pipeline text primary key,
last_success_at timestamptz not null,
rows_written int not null,
run_id text not null
);
Then one alert rule: for each pipeline, now() - last_success_at > expected_interval * 2 is an incident. Add a second rule on volume: rows_written < 0.3 * median_last_7_runs means the job ran but produced suspiciously little. A scraper that returns 4 items instead of 900 will never throw an exception. It will just quietly starve your product.
I put this in n8n workflows as a final HTTP node, in BAS scripts as a last POST before exit, and in Python jobs in a finally block that only fires on success.
Severity mapping: what actually pages a human
Not every red light deserves a 3am phone call. My default mapping:
| Signal | Action |
|---|---|
| Liveness fails | Auto-restart, no page |
| Primary model down, fallback ok | Route to fallback, log, no page |
| All models down | Page |
| Vector count collapsed | Freeze reindex, page |
| 1 canary fails | Warning in Slack |
| 2+ canaries fail twice in a row | Page |
| Heartbeat stale 2x interval | Page |
| Volume anomaly | Warning, auto-open ticket |
| Proxy pool below 30% healthy | Pause scraping, warning |
The important half of that table is the “no page” rows. Every automatic degradation you build (fallback model, cheaper model, pause instead of hammer) converts an incident into a line in a log. That is the real goal.
A one-day rollout plan
If you have an AI pipeline in production with a fake /health, here is the order I would fix it in:
- Morning: split
/livezand/readyz. Add DB, queue, and a real one-token model probe with assertions. Cache for 10 seconds. - Midday: add the heartbeat table plus one alert rule per scheduled job. This catches the most common real-world failure: the job silently stopped running.
- Afternoon: capture 5 golden inputs from real traffic, write assertions, schedule them every 15 minutes.
- Late afternoon: write the severity table above into your alerting config and delete every alert nobody has ever acted on.
That is one day of work that converts “the sales team told us” into “we knew in 90 seconds.” In AI automation, that gap is the whole difference between a system people trust and a system people quietly stop using.
Anti-patterns I keep removing from client code
- Health check that calls the LLM on every user request. You just doubled cost and latency. Cache it.
- Checks without timeouts. One slow dependency makes your readiness endpoint hang, and the orchestrator declares the whole fleet dead.
- Checks that write test rows into production tables. I have seen canary invoices reach a real customer’s dashboard. Use a dedicated tenant or dry-run flag.
status: okwith no detail. Return per-check latency and error strings. Debugging time drops by an order of magnitude.- Alerting on model latency percentiles only. Latency looks great when the model returns empty strings instantly.
FAQ
Will running canary tests every 15 minutes get expensive?
Almost never. Ten golden inputs at a few thousand tokens each, four times an hour, is on the order of a few dollars a month on mid-tier models. If your golden set includes something genuinely heavy like long-document extraction, run the cheap subset every 15 minutes and the heavy subset hourly. Compare that to the cost of eleven hours of empty output silently written to your database.
How do I health-check a browser automation or BAS script that has no HTTP endpoint?
Push instead of pull. At the start of a run, verify the essentials yourself: proxy pool healthy ratio, at least one profile with a valid session (check for a logged-in marker, not HTTP 200), and disk space for downloads. At the end of a successful run, POST a heartbeat with run ID, items processed, and captcha rate to a small collector endpoint. Then alert on stale heartbeats and on volume drops. Captcha rate is one of the most useful early warning signals in browser automation - it rises well before you get fully blocked.
What is the minimum version of this if I am a solo founder with an MVP?
Two things. First, a real model probe in your readiness endpoint that asserts non-empty output, so you never ship empty completions into your database. Second, a heartbeat row per scheduled job plus one alert on staleness. That is maybe two hours of work and it catches the two failure modes that actually hurt early products: silent empty output and jobs that stopped running. Add golden-input canaries as soon as you have paying users who depend on the quality of the output.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks