TL;DR
Concurrent coding agents do not fail the way concurrent programs fail. There is no race to find and no deadlock to break. Given a task large enough, a single agent writes one stub file and stops, and on the hardest of our four tasks, a double-entry financial ledger, it does that up to half the time across Sonnet 4.6, Haiku 4.5, Codex GPT-5.4 and Gemini 3. The obvious fix: run a second agent in parallel and merge the outputs, the way Claude Code (Anthropic, 2025), Cursor (Cursor, 2026) and JetBrains (JetBrains, 2026) now ship it. That makes things worse: the merged result scores below a single agent.
Prior systems sidestep coordination. Large language models are single-stream, and the multi-agent systems built on top of them inherit that serial limit: they either sequence agents through a fixed design, implement, and review pipeline, or pool independent samples with no signalling between them (Shinn et al., 2023; Wang et al., 2023).
The thirteen systems this describes
AgentCoder (Huang et al., 2024), CodeCoR (Pan et al., 2025), CodeSim (Islam et al., 2025), SEMAG (Peng et al., 2026), Self-Organized Agents (Ishibashi & Nishimura, 2024), AgentVerse (Chen et al., 2024), Mixture-of-Agents (J. Wang et al., 2025), ChatDev (Qian et al., 2024), MetaGPT (Hong et al., 2024), AgileCoder (Nguyen et al., 2024), AutoGen (Wu et al., 2024), OpenHands (X. Wang et al., 2025) and CodeCRDT (Pugachev, 2025).
For human teams that coordination problem is already solved by realtime collaborative editing, built on CRDTs: Yjs (Jahns, 2024) and Automerge (Automerge Project, 2023) in production, over the sequence CRDTs Logoot (Weiss et al., 2009), Treedoc (Preguiça et al., 2009), RGA (Roh et al., 2011) and WOOT (Oster et al., 2006), all of them the same convergence guarantee (Kleppmann & Beresford, 2017; Shapiro et al., 2011).
The race at the top of the page sets the two runs side by side. On the left, one agent stubs and exits at about eighty seconds. On the right, two agents in a room claim files, broadcast, merge, and finish all twenty-eight files with every test passing. AgentRoom gives concurrent agents a shared CRDT workspace (Shapiro et al., 2011) plus a small coordination room. At matched compute across five frontier coding command-line interface (CLI) models:
- the give-up failure mode collapses, pooled odds ratio 13.7 (95% confidence interval [3.9, 48], Cochran-Mantel-Haenszel (Mantel & Haenszel, 1959)),
- AgentRoom beats naive parallel-merge by +0.213 in large language model (LLM) judge quality,
- stripping out the MCP layer shows the coordination tools carry most of that gain, not the CRDT substrate,
- and the sweet spot is N = 2.
The problem
Hand a hard backend task to today’s coding agents and two things go wrong:
- A lone agent gives up. Facing a fifteen-file financial ledger with multi-currency accounts and a hash-chained audit log (Haber & Stornetta, 1991), it performs what we call a stub-and-exit: one source skeleton, at most one test passing, then it quits within about eighty seconds, having judged the task too large to finish.
- Naive parallelism makes it worse. A second uncoordinated agent overwrites the first’s entry file in a shared directory; in separate directories, a post-hoc merge (Shinn et al., 2023; X. Wang et al., 2023) amplifies rather than averages the lone-agent failure — and the merged result lands below a single agent.
How AgentRoom works
AgentRoom is one primitive, and three co-designed parts make it up. A shared workspace makes every agent’s writes immediately visible, and concurrent edits merge automatically through a
room_claim, room_release, room_broadcast, room_state and room_read. And a short advisory protocol asks you to claim a file before you write it, respect files others have claimed, and broadcast your progress.
The claim is an
Coordination, not concurrency
Naive concurrency (parallel-merge) lands below a single agent, and the CRDT substrate alone helps only a little. That asymmetry separates AgentRoom from implicit-coordination systems that share a workspace but never signal intent (Pugachev, 2025).
Results
A lone agent abandons; two agents do not
We define abandonment as a scorer-independent binary: a sub-threshold run that leaves at most two source files or exits early. Pooling twelve model-by-task strata under a Cochran-Mantel-Haenszel (CMH) test (Mantel & Haenszel, 1959), the common odds ratio (OR) for Solo versus AgentRoom abandonment is 13.7, with a 95% confidence interval (CI) of [3.9, 48], and every stratum that records abandonment lines up the same way.
Coordination tightens the variance
And beyond eliminating the give-up mode, a room cuts run-to-run variance by 30 to 45 percent across all three CLI-stable models: sigma falls from 0.22 to 0.14 on Sonnet 4.6, 0.25 to 0.17 on Haiku 4.5 and 0.26 to 0.14 on Codex GPT-5.4.
The ablation ladder
A six-condition ablation at matched compute on the headline task comes out monotonic. So the two-agent parallel-merge baseline sits below a single agent — adding agents without coordination is destructive. AgentRoom sits on top, +0.213 over parallel-merge (Welch t = 3.35, p = 0.003, two-sided (Welch, 1947)).
What carries the gain: the bundle probe
Strip away the MCP tools, keeping the CRDT and the collaboration prompt, and the gain over shared-only is only +0.013. And add the tools back and you recover +0.081 more, most of the gain. The small remainder (the prompt-only step over the bare substrate) is within noise at this sample size (n = 7) under a Welch two-sample test (Welch, 1947). So we read the split as an ordering rather than a precise percentage, and the MCP layer carries most of the gain.
The 86 / 14 split is a point estimate from a small same-budget pool (n = 7). So the precise fraction is noisy — the load-bearing claim is the ordering: the coordination layer carries the larger share.
Sequential pipelines amplify the failure
Against a same-model ChatDev-style sequential pipeline (Qian et al., 2024) at matched compute, AgentRoom wins by +0.336 (on the six genuine ChatDev runs, after excluding seven that were orchestrator crashes). But the pipeline stalls at the implementation boundary in four of six runs; AgentRoom abandons in none of its ten.
How many agents? N = 2
Per-run cost grows linearly and the single broadcast channel saturates past three, the coordination ceiling that AgentsNet (Grötschla et al., 2026) reports for larger agent graphs. So two is the operational sweet spot.
Agents start talking to each other
Given a room, agents produce coordination language no one prompted for: claiming a module before working it, announcing a planned interface, and occasionally apologizing for a cross-agent fix (“I touched your file for a one-line fix to unblock the tests, sorry”). Shared-only agents, with the same CRDT but no signalling channel, produce none of this — the channel affordance activates the behavior (Park et al., 2023; Wu et al., 2024).
The full collaboration taxonomy (and how we labelled it)
Across six financial-ledger Sonnet AgentRoom (2 agents) runs we observe module-claiming in 6/6, plan-adjustment in 5/6, and cross-agent bug-fix in 2/6. These are author labels over the room logs, not an automated metric — the representative apology quote is drawn from a three-agent run where the pattern is more frequent. None of the patterns appear in shared-only.
What we got wrong first
We began with a much larger claim, an emergence score of 3.82x and talk of grokking (Power et al., 2022) — our own repeated validity audits killed it: that number was an artifact of a scorer that rewarded raw file and line count, a single bad baseline run, and a CRDT bridge bug under which the agents never actually shared edits. AgentRoom buys reliability: lower abandonment and tighter variance, not a large mean-quality lift.
Limitations and conclusion
We score with an LLM-judge composite (Kim et al., 2024; Li et al., 2026; Verga et al., 2024; Y. Wang et al., 2024; Wataoka et al., 2025) cross-validated against regex and abstract syntax tree (AST) scorers — not a held-out execution oracle (Liu et al., 2023) (tasks ship only an agent-authored test suite), so we make no execution-verified correctness claim. The core result covers two-agent TypeScript backend tasks across three models, with cross-language checks in Python DevBench and Rust+axum and small samples per group, and not the standing multi-agent coding benchmarks (Ashrafi et al., 2025; Grötschla et al., 2026; B. Li et al., 2024; Shafin et al., 2025; Zhu et al., 2025). The abandonment finding therefore rests on pooling twelve strata rather than any single run group.
Within that scope the picture is consistent. Prior work has treated concurrent multi-agent coding as a parallelism problem or a merge-correctness problem. It’s mostly a coordination problem — and a small set of state-management operations turns a crowd of agents tripping over each other into one that ships working code: claim files, announce progress, query the room state.
Give two agents a room and a way to claim the work, and they stop overwriting each other and start finishing the job.
Resources
- Paper and reviews: OpenReview 0aGLZqKJjt (Workshop on Failure Modes of Agentic AI at ICML 2026, FAGEN).
- Preprint: arXiv:2608.23740.
- Related systems referenced above: ChatDev (Qian et al., 2024), MetaGPT (Hong et al., 2024), AgileCoder (Nguyen et al., 2024), CodeCRDT (Pugachev, 2025), AutoGen (Wu et al., 2024), OpenHands (X. Wang et al., 2025), and the MCP specification (Anthropic, 2024).
- Anthropic. (2024). Introducing the Model Context Protocol. https://www.anthropic.com/news/model-context-protocol back: 1, 2
- Anthropic. (2025). Create Custom Subagents: Claude Code Documentation. https://code.claude.com/docs/en/sub-agents
- Ashrafi, N., Bouktif, S., & Mediani, M. (2025). Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency. https://arxiv.org/abs/2505.02133
- Automerge Project. (2023). Automerge: A library of data structures for building collaborative applications. https://automerge.org
- Burrows, M. (2006). The Chubby lock service for loosely-coupled distributed systems. 7th USENIX Symposium on Operating Systems Design and Implementation (OSDI). https://www.usenix.org/conference/osdi-06/chubby-lock-service-loosely-coupled-distributed-systems
- Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., Qin, Y., Cong, X., Xie, R., Liu, Z., Sun, M., & Zhou, J. (2024). AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=EHg5GDnyq1
- Cursor. (2026). New Cursor Interface (Cursor 3.0): Run Many Agents in Parallel Across Repos and Environments. https://cursor.com/changelog/3-0
- Ellis, C. A., & Gibbs, S. J. (1989). Concurrency Control in Groupware Systems. Proceedings of the 1989 ACM SIGMOD International Conference on Management of Data, 399–407. 10.1145/67544.66963
- Gray, C. G., & Cheriton, D. R. (1989). Leases: An Efficient Fault-Tolerant Mechanism for Distributed File Cache Consistency. Proceedings of the Twelfth ACM Symposium on Operating Systems Principles (SOSP), 202–210. 10.1145/74850.74870
- Grötschla, F., Müller, L., Tönshoff, J., Galkin, M., & Perozzi, B. (2026). AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs. https://openreview.net/forum?id=gsSIH0mZ0Y back: 1, 2
- Haber, S., & Stornetta, W. S. (1991). How to Time-Stamp a Digital Document. Journal of Cryptology, 3(2), 99–111. 10.1007/BF00196791
- Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Zhang, C., Wang, J., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., & Schmidhuber, J. (2024). MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=VtmBAGCN7o back: 1, 2
- Huang, D., Zhang, J. M., Luck, M., Bu, Q., Qing, Y., & Cui, H. (2024). AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation. https://arxiv.org/abs/2312.13010
- Ishibashi, Y., & Nishimura, Y. (2024). Self-Organized Agents: A LLM Multi-Agent Framework toward Ultra Large-Scale Code Generation and Optimization. CoRR, abs/2404.02183. 10.48550/arXiv.2404.02183
- Islam, Md. A., Ali, M. E., & Parvez, M. R. (2025). CodeSim: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging. In L. Chiruzzo, A. Ritter, & L. Wang (Eds.), Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 5128–5154). Association for Computational Linguistics. 10.18653/v1/2025.findings-naacl.285
- Jahns, K. (2024). Yjs: A CRDT framework for shared editing. https://github.com/yjs/yjs
- JetBrains. (2026). Air Launches as Public Preview: A New Wave of Dev Tooling. https://blog.jetbrains.com/air/2026/03/air-launches-as-public-preview-a-new-wave-of-dev-tooling-built-on-26-years-of-experience/
- Kim, S., Suk, J., Longpre, S., Lin, B. Y., Shin, J., Welleck, S., Neubig, G., Lee, M., Lee, K., & Seo, M. (2024). Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. https://arxiv.org/abs/2405.01535
- Kleppmann, M., & Beresford, A. R. (2017). A Conflict-Free Replicated JSON Datatype. IEEE Transactions on Parallel and Distributed Systems, 28(10), 2733–2746. 10.1109/tpds.2017.2697382
- Lamport, L. (1978). Time, Clocks, and the Ordering of Events in a Distributed System. Communications of the ACM, 21(7), 558–565. 10.1145/359545.359563
- Li, B., Wu, W., Tang, Z., Shi, L., Yang, J., Li, J., Yao, S., Qian, C., Hui, B., Zhang, Q., Yu, Z., Du, H., Yang, P., Lin, D., Peng, C., & Chen, K. (2024). Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study. https://arxiv.org/abs/2403.08604
- Li, D., Sun, R., Huang, Y., Zhong, M., Jiang, B., Han, J., Zhang, X., Wang, W., & huan liu. (2026). Preference Leakage: A Contamination Problem in LLM-as-a-judge. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=grIvSXVJ65
- Liu, J., Xia, C. S., Wang, Y., & Zhang, L. (2023). Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. https://arxiv.org/abs/2305.01210
- Mantel, N., & Haenszel, W. (1959). Statistical Aspects of the Analysis of Data From Retrospective Studies of Disease. Journal of the National Cancer Institute, 22(4), 719–748. 10.1093/jnci/22.4.719 back: 1, 2
- Nguyen, M. H., Chau, T. P., Nguyen, P. X., & Bui, N. D. Q. (2024). AgileCoder: Dynamic Collaborative Agents for Software Development based on Agile Methodology. https://arxiv.org/abs/2406.11912 back: 1, 2
- Ongaro, D., & Ousterhout, J. (2014). In Search of an Understandable Consensus Algorithm. 2014 USENIX Annual Technical Conference (USENIX ATC 14), 305–319. https://www.usenix.org/conference/atc14/technical-sessions/presentation/ongaro
- Oster, G., Urso, P., Molli, P., & Imine, A. (2006). Data Consistency for P2P Collaborative Editing. Proceedings of the 2006 20th Anniversary Conference on Computer Supported Cooperative Work (CSCW), 259–268. 10.1145/1180875.1180916
- Pan, R., Zhang, H., & Liu, C. (2025). CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation. https://arxiv.org/abs/2501.07811
- Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. https://arxiv.org/abs/2304.03442
- Peng, Y., Hou, H., Zhu, X., He, Y. T., & Yu, F. R. (2026). SEMAG: Self-Evolutionary Multi-Agent Code Generation. https://arxiv.org/abs/2603.15707
- Power, A., Burda, Y., Edwards, H., Babuschkin, I., & Misra, V. (2022). Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets. https://arxiv.org/abs/2201.02177
- Preguiça, N., Marquès, J. M., Shapiro, M., & Letia̧, M. (2009). A Commutative Replicated Data Type for Cooperative Editing. 29th IEEE International Conference on Distributed Computing Systems (ICDCS), 395–403. 10.1109/ICDCS.2009.20
- Pugachev, S. (2025). CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation. https://arxiv.org/abs/2510.18893 back: 1, 2, 3
- Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z., & Sun, M. (2024). ChatDev: Communicative Agents for Software Development. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 15174–15186). Association for Computational Linguistics. 10.18653/v1/2024.acl-long.810 back: 1, 2, 3
- Roh, H.-G., Jeon, M., Kim, J.-S., & Lee, J. (2011). Replicated Abstract Data Types: Building Blocks for Collaborative Applications. Journal of Parallel and Distributed Computing, 71(3), 354–368. 10.1016/j.jpdc.2010.12.006
- Shafin, W. I., Rafi, M. N., Li, Z., & Chen, T.-H. (2025). Evaluating Software Process Models for Multi-Agent Class-Level Code Generation. https://arxiv.org/abs/2511.09794
- Shapiro, M., Preguiça, N., Baquero, C., & Zawirski, M. (2011). Conflict-Free Replicated Data Types. Symposium on Self-Stabilizing Systems (SSS), 386–400. 10.1007/978-3-642-24550-3_29 back: 1, 2
- Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: language agents with verbal reinforcement learning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Advances in Neural Information Processing Systems (Vol. 36, pp. 8634–8652). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2023/file/1b44b878bb782e6954cd888628510e90-Paper-Conference.pdf back: 1, 2
- Sun, C., Jia, X., Zhang, Y., Yang, Y., & Chen, D. (1998). Achieving Convergence, Causality Preservation, and Intention Preservation in Real-Time Cooperative Editing Systems. ACM Transactions on Computer-Human Interaction, 5(1), 63–108. 10.1145/274444.274447
- Verga, P., Hofstatter, S., Althammer, S., Su, Y., Piktus, A., Arkhangorodsky, A., Xu, M., White, N., & Lewis, P. (2024). Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. https://arxiv.org/abs/2404.18796
- Wang, J., WANG, J., Athiwaratkun, B., Zhang, C., & Zou, J. (2025). Mixture-of-Agents Enhances Large Language Model Capabilities. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=h0ZfDIrj7T
- Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., Tran, H. H., Li, F., Ma, R., Zheng, M., Qian, B., Shao, Y., Muennighoff, N., Zhang, Y., Hui, B., … Neubig, G. (2025). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=OJd3ayDDoF back: 1, 2
- Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=1PL1NIMMrw back: 1, 2
- Wang, Y., Yu, Z., Zeng, Z., Yang, L., Wang, C., Chen, H., Jiang, C., Xie, R., Wang, J., Xie, X., Ye, W., Zhang, S., & Zhang, Y. (2024). PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization. https://arxiv.org/abs/2306.05087
- Wataoka, K., Takahashi, T., & Ri, R. (2025). Self-Preference Bias in LLM-as-a-Judge. https://arxiv.org/abs/2410.21819
- Weiss, S., Urso, P., & Molli, P. (2009). Logoot: A Scalable Optimistic Replication Algorithm for Collaborative Editing on P2P Networks. 29th IEEE International Conference on Distributed Computing Systems (ICDCS), 404–412. 10.1109/ICDCS.2009.75
- Welch, B. L. (1947). The Generalization of `Student’s’ Problem when Several Different Population Variances are Involved. Biometrika, 34(1–2), 28–35. 10.1093/biomet/34.1-2.28 back: 1, 2
- Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2024). AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations. First Conference on Language Modeling. https://openreview.net/forum?id=BAakY1hNKS back: 1, 2, 3
- Zhu, K., Du, H., Hong, Z., Yang, X., Guo, S., Wang, Z., Wang, Z., Qian, C., Tang, X., Ji, H., & You, J. (2025). MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 8580–8622). Association for Computational Linguistics. 10.18653/v1/2025.acl-long.421