Multi-Agent Systems: When Several Agents Beat One
When one agent is not enough, how roles and a coordinator are arranged, how context is passed between agents, and the places where multi-agent designs fall apart.
All articles in the guide ИИ-агенты · 11
A multi-agent design looks reasonable: split a hard task among specialists. In practice it does not always win, and knowing the boundary beforehand is worth more than discovering it afterwards.
When one is not enough
Three situations where splitting is justified.
Parallel independent work. Check five sources, parse ten documents, walk twenty sections. Each agent takes its own piece, nobody blocks anybody, and wall-clock time drops by a multiple.
Too broad a tool set. An agent with twenty tools chooses worse than an agent with five: descriptions look alike, context is crowded, error probability rises. Splitting by domain restores precision.
Context economy. A worker agent read a hundred pages and returned a paragraph. Everything it read stayed in its window rather than the main one. This is the same technique as subagents in Claude Code.
When not to split: when steps are strictly sequential and each depends on the last. Then you get the same steps plus the cost of handing context between them.
Roles and the coordinator
The design that works best is a coordinator with workers.
The coordinator breaks the task down, issues briefs, collects results and decides what comes next. Workers do narrow jobs and know nothing about each other.
Why this rather than “agents negotiating among themselves”: flat designs debug badly. When a result is wrong, a hierarchy shows you whose brief was poor; in a free-conversation design you have to read the whole exchange, and the causal chain often cannot be reconstructed at all.
Practical rules for assigning roles:
- A role is defined by its tool set, not by a personality description. An “analyst agent” with no tools is just a prompt.
- Every worker has its own definition of done. Otherwise the coordinator cannot check the result.
- Workers do not call each other. All edges go through the coordinator, or the dependency graph stops being comprehensible.
Passing context
This is where most of the loss happens.
Pass structure, not conversation. The worker gets a brief with the data it needs and returns a result in an agreed format. Retelling the dialogue is both longer and worse: the important detail is exactly what gets lost in the retelling.
The brief must be self-contained. The worker did not participate in earlier steps and cannot ask. Anything not in the brief does not exist for it.
Failure must come back explicitly. A worker that could not do the job and returned plausible prose poisons the whole chain: the coordinator will treat invention as data. The result format must carry a failure flag, and the coordinator must check it.
Shared state goes in external storage. If agents need common data, it belongs in a database, not in messages between them.
Where the design falls apart
An honest list of the problems that make multi-agent systems worse than single agents:
Compounding errors. Three steps at ninety percent reliability give seventy-three percent. The longer the chain, the worse: that is arithmetic, not a property of the model.
Boundary losses. The first agent considered a detail obvious and left it out. The second does not know it. The error appears between agents rather than inside one, and no single agent’s log shows it.
Cost. Every agent has its own context and its own calls. A multi-agent design costs several times more, and it has to earn that back in time or quality.
Debugging. When the result is wrong you have to find which step. Without per-agent logs that is impossible.
Coordinator loops. It issues a task, gets an unsatisfactory result, issues it again. An attempt limit is mandatory, or the loop terminates only when the budget does.
The practical conclusion
Start with one agent and split only when you hit a specific constraint: too many tools, a need for parallelism, not enough context. Splitting for architectural elegance reliably produces a system that is more expensive and worse.
And the rule that holds at any agent count: every result is checked by fact, not by report - see rollout and verification. On what agents remember and how, see agent memory. The overview is in the AI agents guide.
FAQ
When do I need several agents instead of one?
When the task splits into independent parts that can run in parallel, or when one agent would need too broad a tool set. Twenty tools on one agent is a reliable sign it is time to split. If the steps are strictly sequential and depend on each other, several agents only add failure points.
How do agents pass context to each other?
Through an explicit structured result, not through a retelling of the conversation. Each agent receives a brief and returns data in an agreed format. Trying to pass "everything that was said" hits the context window and loses exactly the detail that mattered.
Why does a multi-agent system sometimes perform worse than one agent?
Usually because of losses at the boundaries: what the first agent considered obvious never made it into the result, so the second works from an incomplete picture. Errors also compound: three steps at ninety percent reliability give seventy-three percent at the end.
- What an AI Agent Is and How It Differs from a ChatbotGuide
- Building an AI Agent: From Prompt to ProductionThe path from idea to a working agent: framing the task, choosing tools, the execution loop, testing before launch, and what has to be true before it reaches production.
- How to Build an AI Agent: A Step-by-Step Real CaseOne agent walked through end to end: requirements, design, tools and their descriptions, the first run and the fixes it forced, and which parts deliberately stayed ordinary code.
- AI Agents for Business: Where They Pay Off and Where They Do NotWhich processes an AI agent genuinely makes cheaper, where ordinary automation or a hire wins, how to calculate payback honestly, and the risks that rarely make it into the model.
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks