Open Challenges in Networked Agent Collaboration
A research agenda for coordinating capable, tool-using agents across information, ownership, and control boundaries.
Over the past three years, the multi-agent design space has widened. In 2023, one prominent pattern was a small cast of role-prompted models talking through a fixed script. By 2025, production reports described systems that dispatched parallel researchers; agents were operating tools and shared artifacts; and open protocols were standardizing how remote capabilities could be described and invoked. By 2026, research agendas included swarms, adaptive topologies, persistent experience, shared software environments, and agents whose interests do not fully align. These developments do not form a clean succession, and many prominent systems remain centrally owned and tightly orchestrated. They do, however, make a broader set of coordination questions empirically accessible and increasingly consequential.
We are a young lab working close to this transition: implementing runtimes, reading coordination traces, and trying to separate genuine coordination gains from extra inference, privileged context, and hand-built scaffolding. That proximity has made us less impressed by the fact that agents can talk and more concerned with the conditions under which their work composes. We offer this situated view not to endorse a preferred architecture, but to state the shared questions that any approach to runtime coordination must answer.
Many ingredients are older than language models. The Contract Net protocol used announcement, bidding, award, and recursive subcontracting in 1980. Distributed coalition formation was studied for open agent systems in the 1990s (Shehory and Kraus, 1998). Contextual trust, witness reports, and certified reputation were combined in FIRE. Researchers also explored bounded institutional adaptation to heterogeneous agent populations (Bou et al., 2009). Recursion, team formation, trust, and governance are not inventions of the LLM era.
What changed is the available substrate. Some systems now let probabilistic natural-language actors choose tools, modify external state, sustain longer tasks, recruit or instantiate collaborators at runtime, and communicate across products. Their competence remains uneven, their failures can be semantically ambiguous, and copying an agent does not create an independent source of judgment. As these capabilities improve, collective behavior can become both more useful and more consequential; coordination does not become automatic.
We can already put ten model instances in a room, give them roles, and add a voting round. That is not yet a theory, or a reliable system, of networked collaboration. The field has crossed the orchestration threshold. It has not yet crossed the network threshold.
Why the timing changed
The evidence points to several shifts rather than a single benchmark jump. Instruction following and model-directed tool use supported richer action loops. Context-window reports showed how much more state a model could retrieve within one interaction, while separate task-horizon measurements captured longer work on specific software and research suites. Standardized interfaces described connections across tools and agents. As systems began acting on external resources, some effects persisted beyond the conversation.
| Period | What became practical | What the change exposed |
|---|---|---|
| 2023: action | ReAct interleaved reasoning and environmental action; Toolformer studied model-selected API calls; CAMEL and AutoGen made role-based multi-agent interaction easier to program. | Tool choice, role adherence, turn termination, and the agent-computer interface became part of system behavior rather than fixed application logic. |
| 2024: context and inference effort | The Gemini 1.5 report tested retrieval across context windows up to 10 million tokens; test-time compute research studied adaptive allocation of inference effort. | Context capacity and inference effort became explicit design variables. Retrieval still did not guarantee synthesis, and additional compute created allocation and stopping questions. |
| 2024: environment and topology | SWE-agent showed that an agent-computer interface could materially affect repository work; OSWorld evaluated agents in real computer environments; GPTSwarm and Sparse Debate treated communication topology as a design variable. | Interfaces and communication graphs became part of the algorithm rather than incidental scaffolding. |
| 2025: task horizon | On a suite of software and research tasks, Kwa et al. estimated a roughly 110-minute 50%-completion horizon for o3, while stressing limits on external validity. | Reliability, error recovery, reasoning, and tool use became relevant to how long agents could work, but suite-specific horizons could not be treated as universal autonomy. |
| Late 2024–2025: connectivity and scale | The Model Context Protocol standardized connections to tools and data, while A2A introduced agent discovery and cross-vendor task exchange. IoA and AgentNet explored heterogeneous or evolving organizations. | Interfaces for interoperability became more standardized, while system-level failure, communication cost, and assumptions of central control became explicit research targets. MAST documented failures spanning system design, inter-agent misalignment, and task verification. |
| 2026: shared worlds | An Anthropic research report describes experiments with agents in shared codebases, finite queues, markets, hidden-profile tasks, and conflict settings (Anthropic, 2026). | A failure can now be a conflicting write, an exhausted shared resource, a misleading reputation update, an unsafe delegation, or a group that becomes confidently homogeneous. |
This timeline should not be read as a story of smooth progress. Interface design contributed substantially to the improvement from early repository benchmarks to later coding agents. Long-context retrieval does not guarantee correct synthesis. A protocol connection does not establish trust. A stronger individual agent can discover a vulnerability more quickly and can also execute a conflicting plan more effectively.
The safety and reliability evidence is similarly mixed. In τ-bench, leading tool agents at the time often failed both to complete tasks and to follow domain policies consistently across repeated runs. AgentDojo showed that untrusted content returned by tools can redirect an agent through prompt injection. In one production research system, parallel agents improved an internal breadth-oriented evaluation, but the developers also reported sharply higher token use and poor fit for tasks with tightly coupled dependencies (Anthropic, 2025). These results do not prove that coordination has replaced base capability as the bottleneck. They show that coordination failures have become measurable and consequential enough to deserve first-class study.
Old problems on a new substrate
Several research communities already have pieces of the answer.
Classical multi-agent systems studied allocation, coalition formation, negotiation, reputation, institutions, and organizational adaptation. Distributed systems separated message delivery from consistency, ordering, failure detection, and recovery. Security research distinguished possessing information from holding authority, and treated least privilege, revocation, and provenance as system properties. Economics studied private information, strategic reporting, auctions, congestion, and externalities. Organizational science showed that groups can fail to pool uniquely held information even when every member is individually competent, as in the classic hidden-profile experiments of Stasser and Titus. Human-computer interaction has long warned that useful automation requires appropriate reliance, legible state, and meaningful human intervention, not merely high average accuracy (Lee and See, 2004).
LLM agents combine these concerns in an unusual package. They communicate in an expressive but ambiguous medium. Their policies are partially encoded in prompts and partially inherited from opaque model behavior. They can synthesize plans across domains, yet small framing changes can alter which rules they follow. Their state may be distributed across model context, memory stores, tool backends, shared artifacts, and human conversations. A copied checkpoint can produce many agents cheaply, but those agents may share the same blind spots. The result is neither a conventional distributed service nor a conventional human organization.
That combination changes the failure surface. An agent may complete the requested semantic task while exceeding its authority. A verifier may agree because it shares the worker's model, evidence, or prompt-induced bias. A team may look productive because four agents consumed four times the compute. A coordination protocol may transmit every message correctly while the organization selects the wrong collaborator, duplicates work, or leaves a shared world inconsistent.
The opportunity is to treat these as connected design questions. A network is an organization formed under uncertainty. Its topology shapes what evidence travels, its authority system bounds what actions are possible, its economics determine whether participation remains viable, and its learning rule decides which past outcomes affect future choices.
What counts as networked collaboration?
“Multi-agent,” “distributed,” and “decentralized” are often used as if they named the same property. They do not. A useful description separates four axes:
| Axis | Less distributed | More distributed |
|---|---|---|
| Information | One shared context or global transcript | Private observations, memories, and selectively disclosed evidence |
| Ownership | One operator, budget, and policy boundary | Independently controlled agents with different incentives or obligations |
| Execution | One process and one tool environment | Remote runtimes acting in multiple or shared external systems |
| Control | A central scheduler chooses membership and stopping | Local decisions about participation, delegation, verification, and exit |
A system can be distributed along one axis and centralized along another. A supervisor may coordinate remote workers while retaining complete control. A decentralized protocol may run among agents that all share one owner and one model. A group may make local decisions but write into a centrally managed database. These are legitimate designs; the point is to state which boundaries are real.
We use networked agent collaboration for the problem of composing partially informed, independently situated, tool-using agents into a temporary organization whose authority, evidence, cost, and learning remain bounded and auditable. Independence comes in degrees. It can arise from different evidence, tools, models, memories, owners, or action rights. Merely sampling the same model several times provides weaker independence than these structural differences.
Open protocols are important infrastructure. MCP can expose resources and tools; A2A can exchange agent messages and tasks. A2A's Agent Cards, task lifecycle, artifacts, and security schemes lower the cost of cross-agent interoperability (A2A specification). Both provide interoperability primitives, but neither chooses a deployment's application-level coordination policy, such as whether a particular delegation should be admitted, whether an advertised skill is reliable in context, whether its expected value justifies the traffic, or how concurrent effects should commit. Connectivity is a prerequisite for networked collaboration, not its policy.
This definition also keeps strong baselines in view. If one capable agent with the right tools solves a task more cheaply and safely, adding a network is a regression. Fixed workflows and central supervisors remain appropriate when tasks are stable, information is already pooled, and one owner can specify the process. Network mechanisms must earn their cost where private information, heterogeneous capabilities, changing membership, multiple authority domains, or tightly coupled shared state make static orchestration insufficient.
Ten open challenges
Forming the network
1. Finding collaborators without revealing the task or flooding the network
Contract Net assumed that a manager could announce a task to potential contractors. Modern registries and Agent Cards make richer capability descriptions possible. In an open network, however, discovery is simultaneously a retrieval, privacy, identity, and congestion problem.
A good query should find rare expertise without broadcasting sensitive intent. A good advertisement should be specific enough to route work without exposing private data or becoming an invitation to spam. Capability claims may be stale, strategically inflated, or copied across many identities. Semantic similarity can identify a plausible collaborator, but it cannot establish current availability, competence, permission, or independence.
Progress requires joint measurement of recall, disclosure, traffic, freshness, and identity abuse. A discovery method that finds the best agent only by sending the full task to the whole network has not solved discovery. Neither has a private registry that misses the one collaborator with decisive local evidence. The open question is how to search under an explicit information budget while preserving enough provenance to challenge deceptive or obsolete claims.
2. Forming teams when capability and complementarity are uncertain
Coalition-formation research already asked how self-interested agents could form useful groups under computational constraints. Current LLM systems add uncertainty about what a natural-language capability claim means and whether two individually strong agents provide complementary evidence.
The Captain Agent preprint explores dynamic, task-specific team formation; GPTSwarm optimizes graph components and connectivity; AgentNet learns evolving directed organizations. These are valuable partial answers, generally within a known candidate pool and an experiment owner who can observe outcomes. Open networks must also handle cold start, private costs, heterogeneous tools, and candidates that can decline or misrepresent a role.
The objective is not to recruit the four highest-scoring agents. It is to estimate the verified marginal value of a candidate to this team, given what the current members already know and what they are allowed to reveal. Evaluation should therefore include deliberately redundant experts, privately informed specialists, misleading self-descriptions, and tasks where the optimal team changes as evidence arrives.
3. Delegating recursively without runaway growth or premature stopping
Recursive subcontracting appeared in Contract Net decades ago. LLM agents make it unusually easy because a worker can generate a subtask in natural language and launch another worker at low organizational friction. The same ease creates two symmetric hazards: explosive recruitment and premature closure.
Sparse Debate shows that fewer communication edges can match or outperform denser debate on the tasks it studied. AgentPrune removes redundant or harmful communication. These results show that topology and message selection are part of the algorithm, although they do not establish a universal sparse optimum. Stopping is a separate control problem. The right breadth depends on task decomposition, error correlation, communication cost, and how much useful work remains.
A practical network needs budgets over depth, width, concurrency, information disclosure, and verification. More importantly, it needs an observable stopping rule. “The agents stopped” is not evidence that additional work had low value. Experiments should compare free stopping with effort floors, marginal-value stopping, and fixed-budget controls while preserving failed and unfinished trajectories. Otherwise a seemingly efficient design may simply be leaving early.
Cooperating across boundaries
4. Preserving authority and information boundaries across owners
Messages carry requests and evidence. They should not silently carry authority. An agent that can describe a file, account, or deployment does not thereby gain permission to read or modify it. Cross-owner collaboration therefore needs explicit, attenuable capabilities with resource scope, purpose, expiry, provenance, and revocation.
The hard part is composition. A parent may delegate only part of its authority; a child may recruit another agent; a result may return through a different route; and an emergency revocation may race with work already in flight. Tool discovery itself can leak information, so authorization must apply both when tools are listed and when an action is invoked. AgentDojo adds another constraint: data received through an authorized tool can still contain untrusted instructions.
Authority, competence, and cooperation must remain separate. A good reputation cannot grant permission. A valid capability cannot prove competence. A cooperative prompt cannot eliminate conflicting interests. A convincing system should demonstrate fail-closed behavior under expired, narrowed, revoked, replayed, and cross-owner delegations, not merely describe access rules in the system prompt.
5. Aligning incentives while pricing congestion and shared risk
Agents in one product can often be treated as components of a single objective. Across owners, participation has a cost and truthful reporting cannot be assumed. An agent may exaggerate competence, hide load, free-ride on verification, split into multiple identities, or prefer a locally beneficial action that imposes traffic and risk on everyone else.
Auctions and reputation systems offer ingredients, but semantic work is difficult to price. The value of an answer may appear only after integration; verification is costly; failures can be correlated; and a cheap low-quality contribution can consume more downstream attention than an expensive good one. Congestion is also endogenous. Broadcasting to find expertise increases the load that makes expertise less available.
The central design question is which signals make truthful and useful participation sustainable without turning every delegation into a heavy market transaction. Useful experiments should expose private costs and conflicting preferences, then measure social utility, individual utility, traffic, delay, manipulation, and distribution of risk. A system that improves aggregate task score by externalizing failures onto one owner has not established cooperative value.
6. Combining and verifying semantic work when evidence and errors are correlated
Combining evidence is harder than voting. In hidden-profile settings, a team can fail because uniquely held evidence is never surfaced or is discarded during synthesis. Even when evidence is pooled, multiple agents may agree because they share a checkpoint, training distribution, retrieval source, prompt template, or mistaken premise. A central evaluator may introduce its own systematic bias. Adding more copies of the same judge can increase confidence without increasing epistemic independence.
MAST's failure taxonomy places task verification alongside system design and inter-agent misalignment. Anthropic's analysis of evaluation awareness in BrowseComp also illustrates how stronger search and coordination can exploit unintended paths in an evaluation rather than solve the intended problem. The lesson is broader than benchmark contamination: accepted evidence must be tied to the claim and to the rules under which it was produced.
Verification should expose both coverage and independence. Which unique evidence was surfaced? Which claims or dissenting observations were lost during synthesis? Did reviewers use different evidence, tools, models, owners, or methods? Can a critic reproduce the decisive step? What happens when the verifier abstains or its own authority is limited? Progress should be measured with hidden-profile recall, retention of minority evidence, calibrated false acceptance and false rejection, correlated-error stress tests, evidence lineage, and the full cost of review. Consensus is an observation; it is not a proof obligation discharged.
Learning and acting over time
7. Learning from outcomes without assigning the wrong credit
A network that never learns repeats expensive routing and trust mistakes. A network that updates too eagerly turns one lucky outcome, one poisoned trace, or one fashionable evaluator into durable policy. Credit assignment and experience management are therefore one loop.
FIRE treated reputation as contextual rather than a single global score. Modern adaptive systems can update prompts, memories, routing weights, topology, skill claims, or reusable artifacts. Each carrier has a different blast radius. A posterior over one agent's performance on code review is safer than a global “trust score.” A prompt patch copied across the network can spread an error much faster than the trace that motivated it.
The unresolved questions are causal and institutional. Which participant changed the outcome? Was the team successful because of coordination or because it spent more compute? Who may publish an experience, who may inherit it, and when should it expire? Useful records need scope, provenance, confidence, counterevidence, and rollback. Forgetting is not a defect here. It is part of keeping a stochastic organization adaptable and limiting the reach of bad updates.
8. Making concurrent work compose in a shared world
Most multi-agent benchmarks end when an answer is emitted. Real agents edit repositories, update tickets, reserve resources, send messages, and alter records. Once several agents act on the same world, correctness depends on coordination over state as well as language.
Distributed systems distinguishes ordering from delivery and defines consistency at the level of observable operations. Lamport's logical clocks and linearizability are reminders that an event log alone does not give operations a valid shared meaning. Agent systems need analogous choices: optimistic or pessimistic concurrency, leases, idempotency keys, conflict detection, merge rules, compensating actions, and human escalation.
A message ledger can explain who requested a change; it cannot by itself make concurrent changes compose. Evaluation must include overlapping work, stale reads, retries after partial success, reordered messages, and irreversible side effects. The outcome metric should inspect final world state and causal history, not only each agent's stated completion. This is where apparently successful independent work often becomes a systems failure.
9. Recovering under churn, drift, and partial failure
Open networks do not hold membership, models, tools, latency, or policy constant. Agents join and leave. A provider silently changes a model alias. A skill advertisement outlives the deployment behind it. A worker times out after committing an external action. A network partition makes two coordinators believe they own the same task.
These are not all crash failures. An LLM worker can return a fluent but incomplete result, lose track of a constraint, stop too early, or continue consuming resources while producing little new evidence. Such soft failures are difficult to detect because the process is alive and the output is plausible.
Resilience requires versioned identities, leases, heartbeats, attempt lineage, resumable state, bounded retries, circuit breakers, and graceful degradation. It also requires semantic health signals that do not rely solely on the worker's self-report. Benchmarks should inject churn, stale claims, latency spikes, tool-version changes, partial commits, and correlated provider failures. A network that succeeds only when every member and dependency remains stable is an orchestration demo, not a robust collaboration substrate.
Retaining human control
10. Keeping autonomous organizations legible and governable
As delegation becomes recursive, the human cannot approve every low-level action. Yet “human on the loop” is empty if the system presents only a final answer or an overwhelming transcript. Meaningful control requires a compact account of what organization formed, which authority moved, what evidence supported consequential decisions, what resources were spent, and which effects can still be reversed.
The interface should distinguish proposal, authorization, execution, verification, and settlement. It should show uncertainty and disagreement without asking the human to replay every token. It should support intervention at multiple levels: revoke a capability, pause a branch, replace a verifier, cap a budget, undo a change, or appeal an automated refusal. Appropriate reliance depends on knowing when the system is outside its competence, not merely seeing a confidence score.
Governance also extends beyond one operator. When agents from several organizations collaborate, responsibility for harmful action, false evidence, or leaked information cannot disappear into the network. The open problem is to combine local autonomy with inspectable decisions and enforceable recourse. Auditability should reduce the work of assigning responsibility after a failure and improve the chance of preventing one.
Partial answers, not a winning architecture
Current research supplies mechanisms at different layers:
- Contract and coalition protocols address task announcement, bidding, and team formation.
- Conversational frameworks make roles, tools, humans, and control flow programmable.
- Learned or pruned graphs search over prompts, agents, and communication topology.
- Dynamic team builders compile a task into a temporary organization.
- Interoperability protocols expose capabilities and transport tasks across runtimes.
- Trust and institutional models update beliefs and rules from interaction history.
- Shared-artifact systems give agents a medium for parallel work and integration.
These mechanisms are complementary. A graph optimizer generally assumes a candidate pool and an evaluator. A protocol assumes identities and a security model supplied by deployments. A trust update assumes an observable outcome and a defensible attribution rule. A shared workspace assumes a concurrency policy. A supervisor assumes that central control and global context are acceptable.
The research goal should not be to crown one universal topology. It should be to identify the regions in which a mechanism earns its complexity. Sparse communication may be best when bandwidth dominates. A hub may be best when one actor has legitimate authority and global context. Local coordination may be necessary when information or ownership cannot be centralized. Recursive recruitment may matter only when task decomposition exposes genuinely new capability. The interesting output is a phase diagram over task coupling, agent capability, information distribution, bandwidth, churn, authority, and adversarial participation.
What would count as progress?
Networked collaboration needs evaluations designed around causal comparisons rather than impressive traces.
Match opportunities, not labels. Comparators should have access to the same models, tools, information, authority, verifier quality, and total compute unless one of those is the treatment. If a network gets four workers and a solo baseline gets one quarter of the tokens, the result estimates a bundle of coordination and test-time compute. That may describe a product choice, but it does not isolate a coordination mechanism.
Separate closed and open worlds. A benchmark with a fixed roster, shared owner, reliable links, and complete observability can establish an optimization result inside that setting. Claims about networks need additional treatments: private information, heterogeneous runtimes, independent owners, churn, strategic reporting, limited bandwidth, and shared mutable resources. Simulating those boundaries in one process is useful only when the simulation preserves the information and authority constraints under study.
Use native, conformance-tested comparators. A reimplementation that imitates another system's headline mechanism may omit the very state, verification, or failure semantics that matter. Comparisons should run the native implementation where possible, record the version and configuration, and verify that each treatment actually changes agent-visible behavior. Labels such as “decentralized,” “forum,” or “runtime coordination” are not experimental manipulations by themselves.
Account for the whole system. Measure task utility alongside unauthorized effects, disclosure, verifier errors, token and monetary cost, latency, discovery traffic, duplicate work, state conflicts, retries, recovery, and human audit time. Preserve failures, refusals, unfinished trajectories, and side effects in the denominator. A mechanism should not improve its score by hiding work that failed after consuming resources or changing the world.
Vary base capability and task structure. The value of coordination is conditional. A weak worker may need decomposition yet fail to integrate; a strong worker may make recruitment unnecessary; highly parallel research may benefit from breadth; tightly coupled coding may suffer from synchronous dependencies. Sweeping agent capability, tool reliability, task coupling, and effort reveals when a mechanism helps instead of averaging across incompatible regimes.
Test independence and adaptation directly. Introduce shared blind spots, duplicated evidence, poisoned experience, misleading skill claims, and model drift. Compare same-model votes with reviewers that use different tools or evidence. Evaluate an adaptive system both before and after controlled successes and failures, then inspect whether updates remain scoped and reversible.
Pre-register stopping and preserve provenance. Team size, effort floors, retry rules, evaluator prompts, exclusion criteria, and primary metrics should be fixed before outcomes are inspected. Each result should bind the task, world state, model and provider version, policy, authority, attempts, tool effects, and scoring code. Same-machine replay is useful; independent reconstruction from durable evidence is stronger.
No single score will summarize this agenda. The purpose of a benchmark suite should be to reveal tradeoffs and failure boundaries. A good result tells us under which assumptions a mechanism works, how much it costs, what it exposes, and how it fails when those assumptions break.
A shared agenda
The next generation of agent systems will be shaped by more than model intelligence. It will depend on whether partially informed actors can find one another, form bounded teams, delegate safely, verify claims, share state, learn without poisoning the future, survive churn, and remain answerable to people.
This is a meeting point for several fields. Multi-agent systems contribute allocation, negotiation, trust, and institutions. Distributed systems contribute consistency, recovery, and fault models. Security contributes authority and information boundaries. Economics contributes incentives and externalities. Organizational science contributes evidence about group judgment and hidden information. HCI contributes the design of reliance, intervention, and recourse. LLM-agent research contributes a new substrate on which all of these questions interact.
We should neither present these problems as entirely new nor imply that any current approach resolves them as a system. What would help now is shared terminology, credible comparators, traceable evidence, and experiments that expose rather than hide boundary conditions. The network threshold is not crossed when agents merely exchange tasks; it is crossed when independently situated agents create additional verified value together while their authority, evidence, costs, updates, and failure modes remain bounded and accountable.
Selected references
- Bou, E., López-Sánchez, M., Rodríguez-Aguilar, J. A., and Sichman, J. S. (2009). Adapting autonomic electronic institutions to heterogeneous agent societies.
- Herlihy, M. P., and Wing, J. M. (1990). Linearizability: A correctness condition for concurrent objects.
- Huynh, T. D., Jennings, N. R., and Shadbolt, N. R. (2006). An integrated trust and reputation model for open multi-agent systems.
- Lamport, L. (1978). Time, clocks, and the ordering of events in a distributed system.
- Lee, J. D., and See, K. A. (2004). Trust in automation: Designing for appropriate reliance.
- Shehory, O., and Kraus, S. (1998). Methods for task allocation via agent coalition formation.
- Smith, R. G. (1980). The Contract Net Protocol: High-level communication and control in a distributed problem solver.
- Stasser, G., and Titus, W. (1985). Pooling of unshared information in group decision making.
Recent agent systems and evaluations are linked at first mention in the text. Corporate research reports are described as reports from their authors, not as independently reproduced estimates.