Prompt Engineering: Layers Instead of Forks
Why a copied prompt per client becomes unmaintainable, how a layered prompt is structured, how to test a change, and why prompts need versioning like code.
All articles in the guide ИИ-агенты · 11
Prompt engineering stops being about wording the moment you have more than one prompt. After that it is architecture.
The forking problem
The story is almost always the same. There is a working prompt. A second client arrives with a small difference, so the prompt is copied and edited. A third arrives, and it is copied again.
Six months later:
- There are ten prompts, similar but not identical.
- An improvement has to be applied in ten files.
- Two of them were missed and still have the old behaviour.
- Nobody remembers why the fourth contains an odd sentence, so nobody touches it.
- Checking that a change broke nothing is impossible: there is no example set.
The real damage is not the manual work but that divergence is discovered through complaints. Weeks pass between introducing the error and noticing it.
Layers
The alternative is assembling the prompt from layers, each responsible for its level of generality.
The base layer. Shared by everyone: role, response format, constraints and prohibitions, how to handle ambiguity. Changes rarely, reaches everywhere.
The domain layer. What is common to a class of tasks: support terminology, document handling rules, domain specifics.
The case layer. Thin: a company name, a special term, an exception to a general rule, a category list.
The rule that settles most arguments: if something is true for more than one case, it belongs a layer down. Specifics that leak into the base break other clients; a general rule left in the top layer has to be duplicated.
A practical detail: layers must be assembled programmatically, not by copy-paste. Log the assembled prompt in full - during an incident you need to see what actually went to the model.
Testing changes
Without an example set, editing a prompt is guessing.
The minimum workable set:
- Twenty to thirty real cases with known correct answers.
- Edge cases without fail: empty input, ambiguous input, and one where the correct answer is a refusal.
- The cases that broke before. Every fixed bug adds an example. This is the most valuable part of the set.
The routine: run before the change, make it, run after, compare. Look not only at what improved but at what regressed: improving one case routinely breaks another, and without the set that goes unnoticed.
Separately: run each example several times. The model is non-deterministic, and run-to-run variation is sometimes larger than the effect of your edit.
Versioning
A prompt is code that governs system behaviour. Treat it accordingly.
- Keep it in the repository, not in a database or an admin UI with no history.
- Change it through review. A prompt edit changes product behaviour as much as a code edit.
- Stamp it with a version that appears in logs. Otherwise, during an incident, you cannot tell which prompt was live.
- Be able to roll back. Not “remember how it was” but restore the previous version.
One more field observation: tie the prompt version to the model version. Changing model changes behaviour under the same prompt, and a prompt polished for one model can perform worse than a simpler one on another.
What this buys
A base-layer fix reaches everyone. Specifics stay thin and legible. The example set shows the effect of a change before release rather than after. Logs make incidents explicable.
The long version on a real multi-client project is on the blog: prompt layers instead of forks. What to know about the prompt itself before building layers: prompt engineering basics. The overview is in the AI agents guide.
FAQ
What is wrong with copying the prompt per client?
Copies diverge. Six months later you have ten near-identical prompts, a shared fix has to land in ten places, and two of them will be missed. The problem is not the volume of work but that the divergence is discovered through complaints rather than at edit time.
How is a layered prompt structured?
A shared base defines the role, response format and prohibitions. A domain layer sits on top, then a thin per-client layer with its terminology and exceptions. A fix in the base reaches everywhere while the specifics stay in the thin top layer.
How do I test a prompt change?
Against a fixed set of examples with known correct answers, run before and after the change. Without such a set, any edit swaps one unknown accuracy for another, and you usually notice when something that used to work breaks.
- What an AI Agent Is and How It Differs from a ChatbotGuide
- Building an AI Agent: From Prompt to ProductionThe path from idea to a working agent: framing the task, choosing tools, the execution loop, testing before launch, and what has to be true before it reaches production.
- How to Build an AI Agent: A Step-by-Step Real CaseOne agent walked through end to end: requirements, design, tools and their descriptions, the first run and the fixes it forced, and which parts deliberately stayed ordinary code.
- AI Agents for Business: Where They Pay Off and Where They Do NotWhich processes an AI agent genuinely makes cheaper, where ordinary automation or a hire wins, how to calculate payback honestly, and the risks that rarely make it into the model.
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks