Skip to content
PD
EN RU
AI Automation 8 min read

Switching to an Open Model Is a Migration, Not a Config Change

Moving an automation from a paid API to a self-hosted open-weight model looks like a one-line change. This is the playbook I use so the switch saves money without quietly breaking production.

PD

Pavel Duglas

AI Automation & MVP Architect

Every few weeks a client sends me the same message: “There is a new open model that runs on one GPU and benchmarks close to what we pay for. Can we just switch?” The honest answer is yes, often you can. But the way most teams do it is to change the base URL, swap the model name, run three test prompts, see nice output and ship. Two weeks later the extraction pipeline is dropping fields, the classifier is drifting, and nobody connects it to the switch because “the outputs looked fine”.

I treat moving a workload from a closed API to an open-weight model the same way I treat a database migration. There is a before state, an after state, a way to prove they are equivalent, and a way back. Here is the playbook I actually use.

Why teams are moving now

The reasons are real, and they are not only about price:

  • Cost at volume. A narrow task that runs millions of times a month is where API bills become painful.
  • Data control. Some clients cannot send contracts, medical notes or internal tickets to a third party, full stop.
  • Stability. A model you host does not get silently updated, deprecated or rate limited during your busiest hour.
  • Latency. A small model close to your workers can beat a large remote one for short requests.

The reasons that are not good enough: a leaderboard screenshot, a tweet about tokens per second, or a vague feeling that open is more virtuous. Those are signals to run an experiment, not to migrate.

Step 1: Pick the right workload, not the whole system

Do not migrate “the AI”. Migrate one workload. The best first candidates share three traits:

  1. High volume. If it runs 2,000 times a month, the savings will never pay for your time.
  2. Narrow task. Classification, field extraction, routing, deduplication, short rewrites. Things with a clear right answer.
  3. Measurable output. You can check correctness automatically or with a cheap human review.

The worst first candidates are open-ended agent loops with many tools, long-context reasoning over messy documents, and anything customer-facing where tone matters and nobody has defined “good”. Keep those on the frontier API for now. A model router can send the hard 10% there while the easy 90% goes to your own box.

Step 2: Build the eval set from production, not from your imagination

The single biggest mistake is testing with prompts you wrote by hand. Your hand-written prompts are clean. Production input is not.

What I do:

  • Pull 300 to 1,000 real requests from logs, stratified by type. Include the weird ones: empty fields, mixed languages, huge inputs, obvious junk.
  • Store the current model’s output next to each input. That is your baseline, not your ground truth.
  • Have a human label a subset of 100 to 200 as correct or incorrect for both models. This is where you find out the current model was also wrong 6% of the time.
  • Define the metric per workload: exact match on extracted fields, label accuracy for a classifier, schema validity rate for JSON output.

Then you get a real comparison table instead of a vibe. “New model: 94.1% field accuracy, old: 95.3%, schema valid: 99.8% vs 100%” is a decision you can make. “Looks about the same” is not.

Step 3: Expect the prompt to need porting

Prompts are not portable between model families. Things that change when you switch:

Chat templates and system prompts

Open models are trained with a specific chat template. If your serving layer applies the wrong one, quality drops in ways that look like the model is dumb. Use the template that ships with the model, and verify it by inspecting the raw token sequence once. Some smaller models also follow system prompts more loosely, so instructions you buried in the system message may need to move closer to the actual task.

Structured output

This is where most of my migrations break. A frontier API with a JSON mode almost never returns invalid JSON. A small open model will, maybe 1 to 3% of the time, which in a pipeline doing 50,000 calls a day is a lot of broken records.

The fix is not a stricter prompt. It is constrained decoding: serving engines like vLLM and llama.cpp can force output to match a JSON schema or grammar, so invalid output becomes impossible at the token level. I turn this on by default for every extraction workload. Keep your schema validator anyway, because valid JSON can still contain a hallucinated value.

Tokenizers

Different tokenizers mean the same text can be 10 to 30% more tokens. That breaks context budgets you tuned carefully, and if your chunking logic counts tokens with the old tokenizer, you will silently truncate input. Recount with the new tokenizer and adjust chunk sizes.

Few-shot examples

Smaller models lean on examples much more than frontier ones. Two or three good examples of tricky cases often close half the accuracy gap. They also cost context, so measure both.

Step 4: Measure serving under your real load

The “100 tokens per second on a consumer GPU” number is almost always single-request generation speed. Your pipeline does not run one request. It runs forty concurrent workers at 9am when the queue fills up.

What to measure before committing:

  • Throughput at your peak concurrency, in requests per minute, not tokens per second.
  • p95 latency, not average. Queues are killed by the tail.
  • Behavior at the context sizes you actually use. Long prompts eat memory and reduce how many requests fit on the card at once.
  • Quantization effect on accuracy. A 4-bit model fits on cheaper hardware, but run your eval set on the exact quantized build you will serve. I have seen 4-bit versions lose two or three points on extraction while being fine on classification.

Use a proper serving engine with batching. A naive script wrapping the model will give you a fraction of the throughput of the same GPU behind a batching server.

Step 5: Do the break-even math honestly

Here is an illustrative example from the kind of workload I see often. An extraction job runs 2 million requests a month, around 1,500 input tokens and 200 output tokens each. At mid-tier API pricing that might land somewhere around $2,000 a month. A rented GPU that can handle that load costs maybe $1,000 to $1,200 a month running 24/7.

Looks like a clear win. Now add what people forget:

  • Idle time. If your load is bursty, you pay for the GPU at 3am too, unless you build autoscaling, which is its own project.
  • Redundancy. One GPU is one point of failure. Two GPUs halve your savings.
  • Ops time. Driver updates, out-of-memory crashes, monitoring, someone getting paged. Even four hours a month of senior time has a price.
  • The fallback API bill. You will keep the API for overflow and outages.

In that example the real saving might be 20 to 30%, not 50%. Still worth it at scale, and the data control argument may matter more than the money. At a tenth of that volume, it is almost never worth it. My rough rule: if the workload costs less than $500 a month on an API, do not self-host for cost reasons alone.

Step 6: Roll out like a migration

Shadow first

Send production traffic to both models. Only the old model’s output is used. Log both, compare automatically on the metric from Step 2. Run this for a week, including your peak days. Shadow traffic finds the input types your eval set missed.

Percentage rollout

Move 5%, then 25%, then 100% of traffic to the new model, watching the same metrics plus downstream signals: how often humans correct output, how many records fail validation, how many retries you trigger.

Keep the way back

The API path stays in the code behind a flag for at least a month. Configure automatic fallback: if your server times out or returns invalid output after a constrained retry, the request goes to the API. Your error budget then becomes a cost line instead of an outage.

Pin everything

Pin the exact weights file by hash, the quantization, the serving engine version and the chat template. Treat them as one build artifact with your prompt. The whole point of self-hosting is that nothing changes unless you change it, so do not throw that away by pulling “latest” on deploy.

The checklist I run before saying yes

  • One workload chosen, high volume, narrow, measurable
  • Eval set built from real production logs with human labels
  • Prompt ported: correct chat template, examples added, tokenizer recounted
  • Constrained decoding enabled for structured output, validator still in place
  • Throughput and p95 measured at real concurrency on the exact quantized build
  • Break-even calculated with idle time, redundancy, ops and fallback included
  • Shadow run for a week, then staged rollout
  • API fallback wired in and tested by actually killing the GPU server
  • Weights, engine and template pinned and versioned

Open models are genuinely good now, and for a lot of automation work they are the right call. But the gap between “it answered my test prompt” and “it handles 50,000 messy production requests a day” is exactly where the migration work lives. Do that work once, properly, and the next switch becomes a routine change instead of a gamble.

FAQ

Can I run an open-weight model on my own machine for a production automation?

For low-volume internal tools, yes, a single consumer GPU can be enough. For anything customers depend on, I rent at least one server-grade GPU, put it behind a batching serving engine, and keep a paid API as automatic fallback. A desktop under someone's desk is not production infrastructure.

Which tasks should stay on a frontier API?

Open-ended agent loops with many tools, long-context reasoning over messy documents, and customer-facing writing where quality is subjective. Narrow, high-volume tasks like classification, extraction and routing are the best candidates to move first, often with a router sending only hard cases to the frontier model.

How do I stop a smaller model from returning broken JSON?

Use constrained decoding in your serving engine so output is forced to match your JSON schema at the token level. Prompting alone will not get you to zero failures. Keep a schema and value validator after it anyway, because valid JSON can still contain wrong or invented values.

Related articles