Ship the Agent You Evaluated: Treat Prompts, Tools and Models as One Build Artifact
Your eval scored 94% and production is a mess. The gap is almost never the model - it is untracked drift between what you tested and what you deployed. Here is the manifest-based workflow I use to close it.
Pavel Duglas
AI Automation & MVP Architect
Last month a client called me because their support agent had “gotten dumber overnight.” Nothing was deployed. No code changes, no prompt edits, no incidents. Their eval suite still passed at 94%. Production complaint volume had tripled.
The cause took forty minutes to find. They were calling a model alias instead of a pinned version, and the provider had rolled the alias forward. The new snapshot was better at reasoning and much worse at obeying their output format, which their downstream parser depended on. Their eval suite had been run three weeks earlier, against a different model, using a prompt file that had since had two words changed by a well-meaning intern.
They were not running the agent they evaluated. Almost nobody is.
The drift surface nobody inventories
When you deploy a normal service, you deploy a container with a digest. The artifact is the thing you tested. Agents break this because the behavior lives in a dozen places that are not in your build.
Here is the full drift surface I check on every audit:
- Model identity.
gpt-5,claude-sonnet-latest,openrouter/autoare not versions. They are subscriptions to someone else’s release schedule. - Sampling parameters. Temperature, top_p, max output tokens, reasoning effort, seed. Often hardcoded in one place for evals and another for prod.
- System prompt. Living in a
.mdfile that nobody treats as code, edited directly on the server, or stored in a prompt management SaaS where a teammate can hit “publish.” - Tool schemas. Change a parameter description from “ISO date” to “date” and the model starts sending
12/03/2026. Tool descriptions are prompt text. They are not documentation. - Tool implementations. The schema is identical, but
search_ordersnow returns 50 rows instead of 10 and blows the context window. - Retrieval state. The index you evaluated against had 4,000 documents. Prod has 40,000, including three contradictory policy PDFs somebody uploaded on Friday.
- Context assembly. Truncation order, history window, summarization thresholds. Usually implemented twice: once in the eval harness, once in the real runtime.
- Routing and fallbacks. Your failover path sends 2% of traffic to a completely different model that was never evaluated at all.
- Provider-side changes. Safety filters, caching behavior, tokenizer updates, silent server-side prompt injections for “helpfulness.”
Every item on that list has burned me at least once. The fix is not more evals. It is making the thing under evaluation identical to the thing in production.
Make the agent a single hashed artifact
I put everything behavior-affecting into one manifest, and I hash it. If the hash changes, it is a new release and it needs a new eval run. No exceptions.
# agents/support/manifest.yaml
id: support-agent
version: 7
model:
provider: anthropic
# pinned snapshot, never an alias
name: claude-sonnet-4-5-20250929
temperature: 0.2
max_output_tokens: 1200
fallback:
# evaluated separately, tracked separately
manifest: agents/support/manifest.fallback.yaml
prompts:
system: prompts/support/system.v7.md
refusal: prompts/support/refusal.v3.md
tools:
- name: search_orders
schema: tools/search_orders.v2.json
impl: app.tools.orders:search
max_rows: 10
- name: issue_refund
schema: tools/issue_refund.v1.json
impl: app.tools.billing:refund
requires_approval: true
retrieval:
index: support-kb
snapshot: 2026-02-14T09:00:00Z
top_k: 6
context:
history_turns: 8
truncation: oldest_first
hard_token_cap: 24000
The hash is computed over the manifest plus the contents of every referenced file:
import hashlib, json, yaml
from pathlib import Path
def artifact_hash(manifest_path: Path) -> str:
manifest = yaml.safe_load(manifest_path.read_text())
h = hashlib.sha256()
h.update(json.dumps(manifest, sort_keys=True).encode())
for rel in sorted(_referenced_files(manifest)):
h.update(rel.encode())
h.update(Path(rel).read_bytes())
return h.hexdigest()[:16]
That short hash becomes the agent’s real version number. It goes in the logs, in the eval report, in the deploy annotation, and in every response record. When someone says “the agent got worse,” the first question is no longer “what changed?” but “which artifact answered that request?”
Pin the model, ban aliases in CI
One regex in CI saves you a whole class of outage:
FORBIDDEN = ("latest", "auto", "preview")
def test_model_is_pinned(manifest):
name = manifest["model"]["name"]
assert not any(tag in name for tag in FORBIDDEN), \
f"model must be a pinned snapshot, got {name}"
Yes, pinning means you have to manually adopt new models. That is the point. Model upgrades are releases, not weather.
Version tool contracts like public APIs
search_orders.v2.json is a file that changes only by creating v3. If you edit a description in place, the hash changes, the eval gate fires, and you find out that your careful wording tweak cost you 6 points on date parsing before a customer does.
The same discipline applies to the implementation side. max_rows lives in the manifest, not in the function body, because response size is a behavior parameter, not an implementation detail.
Snapshot the retrieval layer
This is the one people skip, and it is the biggest silent drift source. If your knowledge base is mutable and your evals run against “whatever is in the index right now,” your eval numbers are noise. I use append-only indexes with a timestamp filter, so a manifest can say “documents as of 14 Feb 09:00” and both eval and prod see the same corpus. For small projects a versioned folder of markdown files in git is completely adequate.
Gate deploys on the hash, not on vibes
The CI rule is short: every artifact hash must have a passing eval report at or above the current baseline.
HASH=$(python -m agentkit hash agents/support/manifest.yaml)
if ! agentkit evals report --hash "$HASH" --exists; then
echo "No eval run for artifact $HASH. Run: agentkit evals run"
exit 1
fi
agentkit evals compare --hash "$HASH" --baseline main --max-regression 2
Two details that make this survivable in real life:
- The eval harness must load the manifest. If your harness builds its own prompt and calls its own client, you are back to testing a different agent. One loader, used by both paths, or the whole exercise is theater.
- Allow a small regression budget. Demanding zero regression on a noisy metric means people start disabling tests. A 2-point allowance with a hard floor on critical checks (tool-call validity, format compliance, refusal behavior) is workable.
Assert at boot, not at 3am
CI is not enough because production has its own environment. I add a startup check that fails the container if reality does not match the manifest.
def assert_runtime_parity(manifest):
# provider actually serves the pinned snapshot
models = {m.id for m in client.models.list()}
assert manifest["model"]["name"] in models
# every declared tool exists and its schema matches on disk
for tool in manifest["tools"]:
fn = import_impl(tool["impl"])
assert schema_of(fn) == load_json(tool["schema"])
# retrieval snapshot is reachable
assert index_has_snapshot(manifest["retrieval"]["snapshot"])
A container that refuses to start is an annoying Tuesday. An agent quietly running with a missing tool and hallucinating the results is a refund queue.
Canary, and keep the previous artifact warm
Because the artifact is a single hashed unit, rollout gets simple. I route by hash: 5% of sessions to the new artifact for a few hours, with both versions loaded in the same process. Compare four things between the two cohorts: tool-call error rate, average tokens per resolved session, escalation-to-human rate, and p95 latency. Quality scores from an LLM judge are useful but slow. Those four operational numbers catch most regressions within an hour.
Rollback is a config flag flipping traffic back to the previous hash. No rebuild, no prompt archaeology.
The solo-founder version of all this
If you are one person shipping an MVP, do not build a platform. Do these four things this week:
- Replace every model alias with a pinned snapshot. Ten minutes.
- Move prompts and tool schemas into git files, and make the runtime read them from disk rather than from string literals or a hosted prompt UI.
- Write the hash function above and log the hash with every agent run.
- Keep 30 to 50 real cases in a JSONL file, and a script that runs them through the same loader production uses. That is a real eval suite.
That is maybe half a day of work. It converts “the agent feels worse” from an unanswerable question into a diff between two hashes. Everything else in this article is an optimization on top of that foundation.
The uncomfortable truth is that most agent reliability work is not model work at all. It is release engineering applied to a system whose behavior happens to be written in English.
FAQ
FAQ
Doesn't pinning model versions mean I miss out on improvements?
You miss them until you deliberately adopt them, which is exactly what you want. Adopting a new snapshot becomes a normal release: bump the model in the manifest, get a new hash, run the eval suite, canary it at 5%, promote or roll back. In practice this takes under an hour and you learn whether the newer model actually helps your use case. Auto-upgrading via aliases gives you unvalidated behavior changes on the provider's schedule, usually discovered by a customer.
How do I version retrieval if my knowledge base changes every day?
Make the index append-only and filter by timestamp, so a manifest can reference documents as of a specific moment and both eval and production see the same corpus. If your stack does not support that, keep source documents in git and rebuild the index from a commit SHA that you record in the manifest. For a daily-changing KB, treat index rebuilds as their own release: new snapshot, re-run the eval suite, and watch whether retrieval precision moved before you point production at it.
Is this overkill for an MVP with fifty users?
The full setup is, but the core is not. Pin your model, keep prompts and tool schemas as files in git, hash them, and log the hash with every run. That is a few hours of work and it pays off the first time someone reports a regression, because you can immediately tell whether behavior changed or expectations did. Skip the canary routing, the registry and the automated comparison until you have enough traffic for the numbers to mean something.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks