ComparisonAgent Swarms
Single agent or multi-agent system: deciding on evidence, not enthusiasm
A multi-agent design earns its place when work splits into independent parts, needs distinct expertise or benefits from a reviewer that did not write the answer. It loses when steps depend tightly on one another, context is small or every second and token counts. This page gives the scoring criteria, the costs that demonstrations hide, a fair way to test both designs and a migration path that starts with one agent.
On this page
- What actually changes when one agent becomes several
- Scoring a task for multi-agent fit
- Costs that a demonstration rarely shows
- One agent and a coordinated team, side by side
- Running a fair head-to-head test
- A hypothetical market research brief, built both ways
- Which design to choose
- Splitting roles only where the evidence supports it
- Questions and answers
- Sources
What actually changes when one agent becomes several
A single agent holds the whole task in one context, decides its own next step and calls tools until it is done. A multi-agent system splits that work: a coordinator plans, specialists take bounded pieces with their own context windows and tools, and their outputs are combined and checked. The underlying model can be identical in both designs. What differs is how information is divided, passed on and verified.
Three things can improve with the split. Independent sub-tasks can run at the same time, each specialist works with a smaller and more focused context, and a reviewer can judge a result it did not produce. Anthropic's account of its research system describes this kind of design suiting breadth-first questions with many independent directions, while tasks with heavy shared context or many dependencies, including most coding work, fit less well1.
Three things get harder. Every handoff can lose or distort information, the same background material is often loaded into several contexts, and a failure must be traced across agents instead of read from one transcript.
Scoring a task for multi-agent fit
Rate the task against each criterion. A team of agents is worth prototyping only when most rows land in the right-hand column.
| Criterion | Points to one agent | Points to several agents |
|---|---|---|
| Decomposability | Steps form one chain in which each needs the full result of the last | The work splits into parts with clear inputs and outputs |
| Parallelism | Little can usefully happen at the same time | Several sources, documents or options can be worked on at once |
| Context pressure | Everything needed fits comfortably in one context | Material for the whole task would crowd out the instructions |
| Distinct expertise or tools | One set of instructions and tools covers the job | Parts need different instructions, tools or data access |
| Independent review | Errors are cheap and easy to spot afterwards | A separate check is needed before anything is used |
| Latency tolerance | Users wait on the answer in real time | The result can take minutes without harming anyone |
| Cost tolerance | Each task is low in value and high in volume | Each task is valuable enough to justify extra model usage |
| Error independence | Agents would share the same model, sources and blind spots | Roles can use different sources, models or deterministic checks |
No row decides alone: a decomposable but latency-critical task may still suit one agent making parallel tool calls.
Costs that a demonstration rarely shows
Coordination overhead
Early signalPlanning and status messages take a growing share of each run.
MitigationKeep coordinator instructions short, cap the number of sub-agents per task and measure the share of usage spent on coordination.
Duplicated context
Early signalSeveral agents load the same background documents on every run.
MitigationPass each specialist only the excerpts its contract needs and share stable material through retrieval or prompt caching.
Correlated errors
Early signalAgents agree with each other but disagree with a human reviewer.
MitigationGive the reviewer different inputs or a different model, and add deterministic checks wherever a rule can decide.
Harder debugging
Early signalA wrong final answer cannot be traced to the step that introduced it.
MitigationTrace each agent separately, with inputs, outputs and tool calls linked to one run identifier.
Higher cost per task
Early signalQuality improves slightly while model usage multiplies.
MitigationReport cost per completed task beside quality; Anthropic reported multi-agent systems using about fifteen times the tokens of a chat1.
One agent and a coordinated team, side by side
| Dimension | Single agent | Multi-agent system |
|---|---|---|
| Quality control | Self-checks inside one context; review happens afterwards | A separate reviewer role can block output before it is used |
| Cost profile | Lower and easier to predict per task | Higher, growing with the number of agents, handoffs and retries |
| Latency | Sequential, but with no coordination steps | Parallel work can shorten runs; planning and merging lengthen them |
| Observability | One transcript to read | Several traces that must be joined to explain one result |
| Typical failure | Loses track of a long task or runs short of context | Drops information at a handoff or duplicates work between agents |
| Change management | One prompt and toolset to version | Contracts between roles must stay compatible as each role changes |
Running a fair head-to-head test
Comparisons go wrong when the multi-agent version gets more time, better prompts or softer grading. Hold everything constant except the architecture.
Freeze the task set
Collect representative tasks, including awkward ones that went wrong in the past, and freeze them before either design is tuned.
Give both designs the same resources
Use the same model family, tools, source material and budget ceiling per task. If the team quietly gets a stronger model, you are testing the model, not the architecture.
Agree the grading in advance
Write acceptance criteria and a scoring rubric before running anything. Where people grade, hide which design produced each output.
Record more than quality
Capture cost, elapsed time, tool calls and escalations for every task, so a small quality gain can be weighed against what it took to achieve.
Read the failures, not only the scores
Classify where each design went wrong. Published research groups multi-agent failures into system design issues, misalignment between agents and weak verification3. A team failing at handoffs needs better contracts; a single agent failing on long tasks may need decomposition after all.
Decide per task type
Results often differ between categories of task. Keep the simpler design wherever it performs as well, and reserve the team for the categories where it clearly wins.
A hypothetical market research brief, built both ways
Which design to choose
- If
The task is a sequence in which each step needs the previous result, and users wait for the answer.
ThenUse one agent, with parallel tool calls where the platform supports them.
Coordination would add delay without creating any parallel work.
- If
The task splits into independent parts and each part needs different sources or tools.
ThenPrototype a coordinator with specialist roles and compare it against the single-agent baseline.
This is the situation in which separate contexts and parallel work tend to pay off.
- If
Outputs carry legal, financial or reputational weight and need checking before use.
ThenAdd an independent reviewer role, even if the drafting stays with one agent.
Separate review is often the main benefit, and it does not require splitting everything else.
- If
The multi-agent version wins on quality by a small margin at a much higher cost.
ThenKeep the single agent and revisit when models, prices or the value of the task change.
Extra handoffs also add maintenance work that a small gain rarely covers.
Splitting roles only where the evidence supports it
Anthropic's guidance on building agents recommends starting with the simplest solution that works and adding complexity only when it demonstrably improves results2. In practice that means starting with one agent and its evaluation set. When results show a specific weakness, such as missed sources on broad questions or unchecked claims, split out the one role that addresses it and test again. A reviewer is often the first role worth separating, because it changes what reaches users without restructuring the rest of the work.
Each split should leave a measurable trace in the evaluation. If adding a role does not improve results on the frozen task set, remove it. The coordination options for each split are compared in our guide to multi-agent orchestration patterns.
Questions and answers
Can a multi-agent system ever cost less than a single agent?
Sometimes. If one agent needs a large, expensive model to handle everything, a team in which most roles run on smaller models, with only planning or review on a larger one, can cost less per task. It depends on how much context each role loads and how many retries the handoffs cause, so measure cost per completed task in the head-to-head test rather than estimating it.
How many agents is too many?
There is no fixed number. A useful test is whether you can state what each agent contributes that the evaluation would miss without it. Warning signs include agents passing work back and forth, a coordinator spending more effort on status than on the task, and two roles producing overlapping output. Effort rules, such as one agent for simple requests and more only for broad ones, keep the count proportional.
Should every agent in a team use the same model?
Not necessarily. Mixing models lets you match cost to difficulty, for example a smaller model for extraction and a larger one for planning or synthesis. A different model in the reviewer role can also reduce shared blind spots. The trade-off is more to evaluate and maintain, because each model change can alter how that role's output fits the next contract.
Should we choose a multi-agent framework before designing the task?
It is better to settle the task architecture first. Frameworks make some coordination patterns easy and others awkward, so choosing one early can push a design toward more agents than the work needs. Build the baseline agent, identify which roles the evaluation justifies, then pick the framework that fits those roles and your team's ability to maintain it.
Sources
- How we built our multi-agent research system — Anthropic · checked 10 October 2026
- Building effective agents — Anthropic · checked 10 October 2026
- Why Do Multi-Agent LLM Systems Fail? (Cemri et al.) — arXiv · checked 10 October 2026