Build a Model Router Before You Build a Bigger Budget
Most AI products pay frontier prices for junk work. Here's the model router I put in front of every LLM call: routing tiers, a cheap classifier, caching, fallbacks, and per-feature cost accounting.
Pavel Duglas
AI Automation & MVP Architect
Every AI product I audit has the same line item problem: one hardcoded model name, used for everything, from “classify this message as spam” to “write a 3-page legal summary.” The bill grows linearly with users, margins shrink, and the founder’s reaction is to negotiate credits instead of fixing the architecture. The fix is boring and it works: put a router in front of your LLM calls.
I’ve done this for a support-automation SaaS, a document parser and a Telegram bot doing lead qualification. In all three cases the bill dropped 60–80% and quality stayed flat or improved, because routing forces you to actually define what “good enough” means per task.
The real reason your AI bill is insane
It’s rarely token volume. It’s three things:
- Uniform model choice. You picked the smartest model during the prototype, when correctness mattered more than cents, and never revisited it.
- Uniform context. You stuff the same 6k-token system prompt + full chat history into a call that needs 200 tokens of context to answer “yes/no”.
- Retry blindness. Something fails, you retry with the same expensive model, sometimes three times, sometimes in a loop nobody watches.
A router attacks all three, because routing decisions naturally come with context decisions.
Step 1: Inventory your calls, not your models
Before you write any routing code, list every LLM call site in your product. For each one write down four fields:
- Task type - classification, extraction, rewriting, generation, reasoning, tool selection.
- Failure cost - what happens if it’s wrong? A mislabeled tag is cheap. A wrong invoice total is expensive.
- Latency budget - does a human wait for this in real time?
- Volume - calls per day.
I do this in a spreadsheet in about an hour. The output is always the same shape: 70–85% of calls are high-volume, low-failure-cost tasks, and they’re eating 70–85% of the budget for no reason.
That spreadsheet is your routing policy. Everything after is implementation.
Step 2: Define three tiers, not twelve
I use exactly three tiers. More than that and nobody can reason about the system.
Tier 0 - no model at all
The cheapest LLM call is the one you don’t make. Regex, string matching, a lookup table, a small local embedding model, a deterministic parser. In the document-parsing project, 40% of “AI extraction” calls were replaced by a regex plus a currency normalizer, because the source documents came from three known templates. Cost went to zero, latency went to 2ms, and accuracy went up.
Always ask: is this actually a language problem, or did I just have an LLM in my hand?
Tier 1 - small / open-weight workhorse
Classification, tagging, routing, short extraction, intent detection, yes/no gates, summarizing a single paragraph, reformatting. Small hosted models or self-hosted open-weight models handle this now with no meaningful quality gap on constrained tasks. The important part is constrained: give it a tight schema, an enum of allowed outputs, and 3–5 few-shot examples. A small model with a great prompt beats a big model with a lazy one on this class of work.
Tier 2 - frontier model
Multi-step reasoning, ambiguous inputs, long documents, code generation, anything customer-facing where a wrong answer creates a support ticket or legal risk. This tier should be a minority of your traffic and you should be able to name exactly which features use it.
Step 3: Write the router
The router is a function, not a framework. Something like this shape:
def route(task: str, payload: dict) -> ModelChoice:
policy = POLICIES[task] # from your spreadsheet
if policy.deterministic_handler and policy.deterministic_handler.can_handle(payload):
return ModelChoice(tier=0)
tokens = estimate_tokens(payload)
complexity = policy.complexity_signal(payload) # cheap heuristic
if complexity == "low" and tokens < policy.small_model_limit:
return ModelChoice(tier=1, model=policy.small_model)
return ModelChoice(tier=2, model=policy.big_model)
Notice what’s not here: an LLM deciding which LLM to use. That’s the trap I see people fall into - a “router agent” that costs a frontier call to decide whether to make a frontier call. Use it only when heuristics genuinely can’t tell, and then use your Tier 1 model as the classifier with a hard token cap and an enum output.
Good complexity signals that cost nothing:
- Input length and structure (does it parse as a known template?)
- Number of entities/questions in the request
- Whether tool use is required
- User tier (paid users can get Tier 2 on borderline cases)
- Retry count (see below)
Step 4: Escalation, not retries
Blind retries are how you triple a bill during an incident. Replace them with escalation:
- Tier 1 call runs.
- Output goes through a validator - JSON schema check, enum check, business-rule check (does the total equal the sum of line items?), or a confidence field the model must emit.
- If validation fails, escalate to Tier 2 once, with the failure reason appended to the prompt.
- If Tier 2 fails validation, fail loudly to a human queue. Do not loop.
This is the single highest-leverage pattern in the whole article. It means your cheap tier only needs to be right most of the time, because failures are caught and upgraded instead of shipped. In practice a 92%-accurate small model plus a validator plus one escalation beats a 97%-accurate expensive model on both cost and end-to-end accuracy, because the expensive model’s 3% failures go out unchecked.
Hard cap total spend per request. I attach a token/cost budget to every incoming job and the router refuses to escalate past it. One runaway agent loop pays for a lot of engineering time.
Step 5: Cache like you mean it
Three layers, in order of value:
- Exact-match cache. Hash of (task, model, prompt, params) → response. Trivially correct, surprisingly effective. Support bots and product-description generators see huge hit rates.
- Prompt/prefix caching. Most providers charge less for repeated prefixes. Restructure prompts so the stable part (instructions, schema, examples) comes first and the variable part last. This is a 30-minute refactor with real savings.
- Semantic cache. Embed the request, look for a near-duplicate above a similarity threshold, reuse the answer. Powerful but dangerous: use it only where a slightly-off answer is acceptable, and never for anything with per-user data in the response. Set the threshold high and log every hit so you can audit it.
Step 6: An eval gate, or you’re guessing
You cannot downgrade a model responsibly without a test set. It doesn’t need to be fancy. For each task, collect 50–200 real inputs with expected outputs. For classification, that’s exact match. For extraction, field-level match. For generation, a rubric scored by your Tier 2 model plus 20 hand-checked examples to sanity-check the grader.
Run it as a CI job. Rules I enforce:
- A routing policy change is a code change with an eval diff in the PR.
- Every task has a minimum accuracy threshold; below it, the build fails.
- The eval reports cost per run, so you see the tradeoff in the same table as the accuracy.
This converts “I feel like the cheap model is worse” into a number. Half the time the number says the cheap model is fine and the argument ends.
Step 7: Cost accounting per feature, not per month
Log every call with: task name, tier, model, input tokens, output tokens, computed cost, cache hit flag, validation result, escalation flag, user/tenant id. One wide table, one dashboard.
What you want to answer in five seconds:
- Cost per feature per day
- Cost per active user, and cost per paying user
- Escalation rate by task (rising escalation = a prompt or upstream data change)
- Top 10 most expensive individual requests yesterday
That last one finds the pathological cases: the user who pasted a 200-page PDF, the loop that ran 47 times, the tenant whose usage alone destroys your unit economics. In the support SaaS, three accounts out of 400 were generating 31% of the AI cost. That’s not a model problem, that’s a pricing problem - and you can only see it if you log per tenant.
The migration order that avoids drama
Don’t rewrite everything. Sequence it:
- Add logging and cost attribution. Change nothing else. Wait a week.
- Build the eval set for your top 2 tasks by cost.
- Move those two tasks to Tier 1 with a validator and single escalation.
- Add exact-match caching and prompt-prefix restructuring.
- Replace the obvious Tier 0 candidates with deterministic code.
- Only then look at self-hosting anything.
Self-hosting open-weight models is last on purpose. It’s the option with the best headline savings and the worst hidden costs - GPU capacity, ops, evals, on-call. Do it when volume is stable and predictable, not because a benchmark score impressed you.
What this actually buys you
Beyond the bill: model independence. When you have a router, an eval harness and per-task policies, swapping in a new model is a config change and a CI run, not a two-week panic. New models ship every few weeks now. The teams that benefit are the ones who can test and adopt in an afternoon; the teams with a hardcoded model name in 40 files just watch.
Build the router before you need it. It’s a day of work and it makes every future AI decision cheaper.
FAQ
Won't routing to smaller models visibly hurt quality?
Not if you pair it with validation and escalation. A small model on a tightly-constrained task with a schema and few-shot examples is usually within a point or two of a frontier model. The failures that matter are the ones you ship silently - so add a validator (JSON schema, business rules, enum check) and escalate failures to your big model once. End-to-end accuracy often goes up, because the expensive model's own errors were never being checked before.
Should I use an LLM to decide which LLM to call?
Only as a last resort. Most routing decisions are answerable with free heuristics: input length, whether the input matches a known template, number of entities, whether tools are needed, user tier. If you genuinely need a semantic decision, use your cheap tier as the classifier with an enum output and a hard token cap - never a frontier model, or you've paid the price you were trying to avoid.
When does self-hosting an open-weight model actually make sense?
When your volume is high and predictable, the task is narrow, and you already have evals to prove quality. Do it after you've done deterministic replacement, caching, prompt-prefix restructuring and tier routing - those are cheaper wins with no ops burden. Self-hosting trades a variable API bill for fixed GPU cost plus on-call responsibility, which is a good trade only at scale or when data residency requires it.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks