ComparisonAgent Swarms

Single agent or multi-agent system: deciding on evidence, not enthusiasm

A multi-agent design earns its place when work splits into independent parts, needs distinct expertise or benefits from a reviewer that did not write the answer. It loses when steps depend tightly on one another, context is small or every second and token counts. This page gives the scoring criteria, the costs that demonstrations hide, a fair way to test both designs and a migration path that starts with one agent.

Reviewed 8 min read

On this page
  1. What actually changes when one agent becomes several
  2. Scoring a task for multi-agent fit
  3. Costs that a demonstration rarely shows
  4. One agent and a coordinated team, side by side
  5. Running a fair head-to-head test
  6. A hypothetical market research brief, built both ways
  7. Which design to choose
  8. Splitting roles only where the evidence supports it
  9. Questions and answers
  10. Sources

What actually changes when one agent becomes several

A single agent holds the whole task in one context, decides its own next step and calls tools until it is done. A multi-agent system splits that work: a coordinator plans, specialists take bounded pieces with their own context windows and tools, and their outputs are combined and checked. The underlying model can be identical in both designs. What differs is how information is divided, passed on and verified.

Three things can improve with the split. Independent sub-tasks can run at the same time, each specialist works with a smaller and more focused context, and a reviewer can judge a result it did not produce. Anthropic's account of its research system describes this kind of design suiting breadth-first questions with many independent directions, while tasks with heavy shared context or many dependencies, including most coding work, fit less well1.

Three things get harder. Every handoff can lose or distort information, the same background material is often loaded into several contexts, and a failure must be traced across agents instead of read from one transcript.

Scoring a task for multi-agent fit

Rate the task against each criterion. A team of agents is worth prototyping only when most rows land in the right-hand column.

CriterionPoints to one agentPoints to several agents
DecomposabilitySteps form one chain in which each needs the full result of the lastThe work splits into parts with clear inputs and outputs
ParallelismLittle can usefully happen at the same timeSeveral sources, documents or options can be worked on at once
Context pressureEverything needed fits comfortably in one contextMaterial for the whole task would crowd out the instructions
Distinct expertise or toolsOne set of instructions and tools covers the jobParts need different instructions, tools or data access
Independent reviewErrors are cheap and easy to spot afterwardsA separate check is needed before anything is used
Latency toleranceUsers wait on the answer in real timeThe result can take minutes without harming anyone
Cost toleranceEach task is low in value and high in volumeEach task is valuable enough to justify extra model usage
Error independenceAgents would share the same model, sources and blind spotsRoles can use different sources, models or deterministic checks

No row decides alone: a decomposable but latency-critical task may still suit one agent making parallel tool calls.

Costs that a demonstration rarely shows

Coordination overhead

Early signalPlanning and status messages take a growing share of each run.

MitigationKeep coordinator instructions short, cap the number of sub-agents per task and measure the share of usage spent on coordination.

Duplicated context

Early signalSeveral agents load the same background documents on every run.

MitigationPass each specialist only the excerpts its contract needs and share stable material through retrieval or prompt caching.

Correlated errors

Early signalAgents agree with each other but disagree with a human reviewer.

MitigationGive the reviewer different inputs or a different model, and add deterministic checks wherever a rule can decide.

Harder debugging

Early signalA wrong final answer cannot be traced to the step that introduced it.

MitigationTrace each agent separately, with inputs, outputs and tool calls linked to one run identifier.

Higher cost per task

Early signalQuality improves slightly while model usage multiplies.

MitigationReport cost per completed task beside quality; Anthropic reported multi-agent systems using about fifteen times the tokens of a chat1.

One agent and a coordinated team, side by side

DimensionSingle agentMulti-agent system
Quality controlSelf-checks inside one context; review happens afterwardsA separate reviewer role can block output before it is used
Cost profileLower and easier to predict per taskHigher, growing with the number of agents, handoffs and retries
LatencySequential, but with no coordination stepsParallel work can shorten runs; planning and merging lengthen them
ObservabilityOne transcript to readSeveral traces that must be joined to explain one result
Typical failureLoses track of a long task or runs short of contextDrops information at a handoff or duplicates work between agents
Change managementOne prompt and toolset to versionContracts between roles must stay compatible as each role changes

Running a fair head-to-head test

Comparisons go wrong when the multi-agent version gets more time, better prompts or softer grading. Hold everything constant except the architecture.

  1. Freeze the task set

    Collect representative tasks, including awkward ones that went wrong in the past, and freeze them before either design is tuned.

    Output
    Frozen evaluation set
  2. Give both designs the same resources

    Use the same model family, tools, source material and budget ceiling per task. If the team quietly gets a stronger model, you are testing the model, not the architecture.

  3. Agree the grading in advance

    Write acceptance criteria and a scoring rubric before running anything. Where people grade, hide which design produced each output.

  4. Record more than quality

    Capture cost, elapsed time, tool calls and escalations for every task, so a small quality gain can be weighed against what it took to achieve.

  5. Read the failures, not only the scores

    Classify where each design went wrong. Published research groups multi-agent failures into system design issues, misalignment between agents and weak verification3. A team failing at handoffs needs better contracts; a single agent failing on long tasks may need decomposition after all.

  6. Decide per task type

    Results often differ between categories of task. Keep the simpler design wherever it performs as well, and reserve the team for the categories where it clearly wins.

    Output
    Design choice per task category

A hypothetical market research brief, built both ways

Which design to choose

  • If

    The task is a sequence in which each step needs the previous result, and users wait for the answer.

    Then

    Use one agent, with parallel tool calls where the platform supports them.

    Coordination would add delay without creating any parallel work.

  • If

    The task splits into independent parts and each part needs different sources or tools.

    Then

    Prototype a coordinator with specialist roles and compare it against the single-agent baseline.

    This is the situation in which separate contexts and parallel work tend to pay off.

  • If

    Outputs carry legal, financial or reputational weight and need checking before use.

    Then

    Add an independent reviewer role, even if the drafting stays with one agent.

    Separate review is often the main benefit, and it does not require splitting everything else.

  • If

    The multi-agent version wins on quality by a small margin at a much higher cost.

    Then

    Keep the single agent and revisit when models, prices or the value of the task change.

    Extra handoffs also add maintenance work that a small gain rarely covers.

Splitting roles only where the evidence supports it

Anthropic's guidance on building agents recommends starting with the simplest solution that works and adding complexity only when it demonstrably improves results2. In practice that means starting with one agent and its evaluation set. When results show a specific weakness, such as missed sources on broad questions or unchecked claims, split out the one role that addresses it and test again. A reviewer is often the first role worth separating, because it changes what reaches users without restructuring the rest of the work.

Each split should leave a measurable trace in the evaluation. If adding a role does not improve results on the frozen task set, remove it. The coordination options for each split are compared in our guide to multi-agent orchestration patterns.

Questions and answers

Can a multi-agent system ever cost less than a single agent?

Sometimes. If one agent needs a large, expensive model to handle everything, a team in which most roles run on smaller models, with only planning or review on a larger one, can cost less per task. It depends on how much context each role loads and how many retries the handoffs cause, so measure cost per completed task in the head-to-head test rather than estimating it.

How many agents is too many?

There is no fixed number. A useful test is whether you can state what each agent contributes that the evaluation would miss without it. Warning signs include agents passing work back and forth, a coordinator spending more effort on status than on the task, and two roles producing overlapping output. Effort rules, such as one agent for simple requests and more only for broad ones, keep the count proportional.

Should every agent in a team use the same model?

Not necessarily. Mixing models lets you match cost to difficulty, for example a smaller model for extraction and a larger one for planning or synthesis. A different model in the reviewer role can also reduce shared blind spots. The trade-off is more to evaluate and maintain, because each model change can alter how that role's output fits the next contract.

Should we choose a multi-agent framework before designing the task?

It is better to settle the task architecture first. Frameworks make some coordination patterns easy and others awkward, so choosing one early can push a design toward more agents than the work needs. Build the baseline agent, identify which roles the evaluation justifies, then pick the framework that fits those roles and your team's ability to maintain it.

Sources

  1. How we built our multi-agent research system — Anthropic · checked 10 October 2026
  2. Building effective agents — Anthropic · checked 10 October 2026
  3. Why Do Multi-Agent LLM Systems Fail? (Cemri et al.) — arXiv · checked 10 October 2026

More in Agent Swarms

Back to Agent Swarms

Next step

Send us the task you are thinking of splitting across agents

Describe the task, how often it runs and what a good result looks like. We will suggest how to baseline it with one agent and which roles, if any, are worth testing as a team.

Discuss a multi-agent design