AgentRoom

Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

loading task...
Durations, file counts, test outcome and conflict count are from the two runs. The event text is representative.
Solo 0selapsed 0files 0 / 18tests
AgentRoom 0selapsed 0files 0 / 18tests
Dev A
🔒 Shared Room
room log
Dev B
0s / 605s

Authors

Published

Jun. 2026

Paper

Checking access...

Table of Contents

TL;DR

Concurrent coding agents do not fail the way concurrent programs fail. There is no race to find and no deadlock to break. Given a task large enough, a single agent writes one stub file and stops, and on the hardest of our four tasks, a double-entry financial ledger, it does that up to half the time across Sonnet 4.6, Haiku 4.5, Codex GPT-5.4 and Gemini 3. The obvious fix: run a second agent in parallel and merge the outputs, the way Claude Code (Anthropic, 2025), Cursor (Cursor, 2026) and JetBrains (JetBrains, 2026) now ship it. That makes things worse: the merged result scores below a single agent.

Prior systems sidestep coordination. Large language models are single-stream, and the multi-agent systems built on top of them inherit that serial limit: they either sequence agents through a fixed design, implement, and review pipeline, or pool independent samples with no signalling between them (Shinn et al., 2023; Wang et al., 2023).

The thirteen systems this describes

AgentCoder (Huang et al., 2024), CodeCoR (Pan et al., 2025), CodeSim (Islam et al., 2025), SEMAG (Peng et al., 2026), Self-Organized Agents (Ishibashi & Nishimura, 2024), AgentVerse (Chen et al., 2024), Mixture-of-Agents (J. Wang et al., 2025), ChatDev (Qian et al., 2024), MetaGPT (Hong et al., 2024), AgileCoder (Nguyen et al., 2024), AutoGen (Wu et al., 2024), OpenHands (X. Wang et al., 2025) and CodeCRDT (Pugachev, 2025).

For human teams that coordination problem is already solved by realtime collaborative editing, built on CRDTs: Yjs (Jahns, 2024) and Automerge (Automerge Project, 2023) in production, over the sequence CRDTs Logoot (Weiss et al., 2009), Treedoc (Preguiça et al., 2009), RGA (Roh et al., 2011) and WOOT (Oster et al., 2006), all of them the same convergence guarantee (Kleppmann & Beresford, 2017; Shapiro et al., 2011).

The race at the top of the page sets the two runs side by side. On the left, one agent stubs and exits at about eighty seconds. On the right, two agents in a room claim files, broadcast, merge, and finish all twenty-eight files with every test passing. AgentRoom gives concurrent agents a shared CRDT workspace (Shapiro et al., 2011) plus a small coordination room. At matched compute across five frontier coding command-line interface (CLI) models:

The problem

Hand a hard backend task to today’s coding agents and two things go wrong:

How AgentRoom works

AgentRoom is one primitive, and three co-designed parts make it up. A shared workspace makes every agent’s writes immediately visible, and concurrent edits merge automatically through a

CRDT
. A coordination room lets an agent claim a file, release it, broadcast a note, and query who holds what, all exposed as five MCP tools (Anthropic, 2024): room_claim, room_release, room_broadcast, room_state and room_read. And a short advisory protocol asks you to claim a file before you write it, respect files others have claimed, and broadcast your progress.

Architecture
Figure 2. N agents share a CRDT-merged workspace; an MCP room exposes claim / release / broadcast / state / read. Hover a component to see its role.

The claim is an

advisory lock
(Burrows, 2006; Lamport, 1978) (one that agents agree to honor rather than a heavyweight consensus protocol such as Raft (Ongaro & Ousterhout, 2014), and closer in spirit to a lease (Gray & Cheriton, 1989) than to a mutex). And it doesn’t need to be heavier — the CRDT already guarantees the bytes converge.

CRDT merge
Figure 3. Two agents insert into the same line concurrently; the CRDT interleaves both insertions deterministically with no conflict.

Coordination, not concurrency

Naive concurrency (parallel-merge) lands below a single agent, and the CRDT substrate alone helps only a little. That asymmetry separates AgentRoom from implicit-coordination systems that share a workspace but never signal intent (Pugachev, 2025).

Results

A lone agent abandons; two agents do not

We define abandonment as a scorer-independent binary: a sub-threshold run that leaves at most two source files or exits early. Pooling twelve model-by-task strata under a Cochran-Mantel-Haenszel (CMH) test (Mantel & Haenszel, 1959), the common odds ratio (OR) for Solo versus AgentRoom abandonment is 13.7, with a 95% confidence interval (CI) of [3.9, 48], and every stratum that records abandonment lines up the same way.

Abandonment forest plot
Figure 4. Per-stratum abandonment (Solo vs AgentRoom (2 agents)), with the pooled CMH odds ratio. Toggle between by-task and by-model views.

Coordination tightens the variance

And beyond eliminating the give-up mode, a room cuts run-to-run variance by 30 to 45 percent across all three CLI-stable models: sigma falls from 0.22 to 0.14 on Sonnet 4.6, 0.25 to 0.17 on Haiku 4.5 and 0.26 to 0.14 on Codex GPT-5.4.

Variance contraction
Figure 5. Run-to-run standard deviation of the quality score, Solo vs AgentRoom (2 agents), on the financial-ledger task. Sigma contracts 30 to 45 percent per model.

The ablation ladder

A six-condition ablation at matched compute on the headline task comes out monotonic. So the two-agent parallel-merge baseline sits below a single agent — adding agents without coordination is destructive. AgentRoom sits on top, +0.213 over parallel-merge (Welch t = 3.35, p = 0.003, two-sided (Welch, 1947)).

Six-condition ablation
Figure 6. Mean quality score for six conditions on the financial-ledger task. The gap from parallel-merge to AgentRoom is +0.213. Toggle the model.

What carries the gain: the bundle probe

Strip away the MCP tools, keeping the CRDT and the collaboration prompt, and the gain over shared-only is only +0.013. And add the tools back and you recover +0.081 more, most of the gain. The small remainder (the prompt-only step over the bare substrate) is within noise at this sample size (n = 7) under a Welch two-sample test (Welch, 1947). So we read the split as an ordering rather than a precise percentage, and the MCP layer carries most of the gain.

Bundle decomposition
Figure 8. Decomposing the AgentRoom bundle: substrate (CRDT) vs coordination (MCP). Drag to see the attribution.
A caveat we report honestly

The 86 / 14 split is a point estimate from a small same-budget pool (n = 7). So the precise fraction is noisy — the load-bearing claim is the ordering: the coordination layer carries the larger share.

Sequential pipelines amplify the failure

Against a same-model ChatDev-style sequential pipeline (Qian et al., 2024) at matched compute, AgentRoom wins by +0.336 (on the six genuine ChatDev runs, after excluding seven that were orchestrator crashes). But the pipeline stalls at the implementation boundary in four of six runs; AgentRoom abandons in none of its ten.

Paradigm contrast
Figure 7. Concurrent AgentRoom vs a sequential 3-phase pipeline (same model, matched 1200s budget).

How many agents? N = 2

Per-run cost grows linearly and the single broadcast channel saturates past three, the coordination ceiling that AgentsNet (Grötschla et al., 2026) reports for larger agent graphs. So two is the operational sweet spot.

Agent-count scaling
Figure 9. Agent count on the financial-ledger task (Sonnet). Quality peaks at N=2; max tests peak at N=3; the broadcast channel saturates beyond N=3. Drag the slider.

Agents start talking to each other

Given a room, agents produce coordination language no one prompted for: claiming a module before working it, announcing a planned interface, and occasionally apologizing for a cross-agent fix (“I touched your file for a one-line fix to unblock the tests, sorry”). Shared-only agents, with the same CRDT but no signalling channel, produce none of this — the channel affordance activates the behavior (Park et al., 2023; Wu et al., 2024).

The full collaboration taxonomy (and how we labelled it)

Across six financial-ledger Sonnet AgentRoom (2 agents) runs we observe module-claiming in 6/6, plan-adjustment in 5/6, and cross-agent bug-fix in 2/6. These are author labels over the room logs, not an automated metric — the representative apology quote is drawn from a three-agent run where the pattern is more frequent. None of the patterns appear in shared-only.

What we got wrong first

We began with a much larger claim, an emergence score of 3.82x and talk of grokking (Power et al., 2022) — our own repeated validity audits killed it: that number was an artifact of a scorer that rewarded raw file and line count, a single bad baseline run, and a CRDT bridge bug under which the agents never actually shared edits. AgentRoom buys reliability: lower abandonment and tighter variance, not a large mean-quality lift.

Limitations and conclusion

We score with an LLM-judge composite (Kim et al., 2024; Li et al., 2026; Verga et al., 2024; Y. Wang et al., 2024; Wataoka et al., 2025) cross-validated against regex and abstract syntax tree (AST) scorers — not a held-out execution oracle (Liu et al., 2023) (tasks ship only an agent-authored test suite), so we make no execution-verified correctness claim. The core result covers two-agent TypeScript backend tasks across three models, with cross-language checks in Python DevBench and Rust+axum and small samples per group, and not the standing multi-agent coding benchmarks (Ashrafi et al., 2025; Grötschla et al., 2026; B. Li et al., 2024; Shafin et al., 2025; Zhu et al., 2025). The abandonment finding therefore rests on pooling twelve strata rather than any single run group.

Within that scope the picture is consistent. Prior work has treated concurrent multi-agent coding as a parallelism problem or a merge-correctness problem. It’s mostly a coordination problem — and a small set of state-management operations turns a crowd of agents tripping over each other into one that ships working code: claim files, announce progress, query the room state.

Give two agents a room and a way to claim the work, and they stop overwriting each other and start finishing the job.


Resources

  1. Anthropic. (2024). Introducing the Model Context Protocol. https://www.anthropic.com/news/model-context-protocol back: 1, 2
  2. Anthropic. (2025). Create Custom Subagents: Claude Code Documentation. https://code.claude.com/docs/en/sub-agents
  3. Ashrafi, N., Bouktif, S., & Mediani, M. (2025). Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency. https://arxiv.org/abs/2505.02133
  4. Automerge Project. (2023). Automerge: A library of data structures for building collaborative applications. https://automerge.org
  5. Burrows, M. (2006). The Chubby lock service for loosely-coupled distributed systems. 7th USENIX Symposium on Operating Systems Design and Implementation (OSDI). https://www.usenix.org/conference/osdi-06/chubby-lock-service-loosely-coupled-distributed-systems
  6. Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., Qin, Y., Cong, X., Xie, R., Liu, Z., Sun, M., & Zhou, J. (2024). AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=EHg5GDnyq1
  7. Cursor. (2026). New Cursor Interface (Cursor 3.0): Run Many Agents in Parallel Across Repos and Environments. https://cursor.com/changelog/3-0
  8. Ellis, C. A., & Gibbs, S. J. (1989). Concurrency Control in Groupware Systems. Proceedings of the 1989 ACM SIGMOD International Conference on Management of Data, 399–407. 10.1145/67544.66963
  9. Gray, C. G., & Cheriton, D. R. (1989). Leases: An Efficient Fault-Tolerant Mechanism for Distributed File Cache Consistency. Proceedings of the Twelfth ACM Symposium on Operating Systems Principles (SOSP), 202–210. 10.1145/74850.74870
  10. Grötschla, F., Müller, L., Tönshoff, J., Galkin, M., & Perozzi, B. (2026). AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs. https://openreview.net/forum?id=gsSIH0mZ0Y back: 1, 2
  11. Haber, S., & Stornetta, W. S. (1991). How to Time-Stamp a Digital Document. Journal of Cryptology, 3(2), 99–111. 10.1007/BF00196791
  12. Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Zhang, C., Wang, J., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., & Schmidhuber, J. (2024). MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=VtmBAGCN7o back: 1, 2
  13. Huang, D., Zhang, J. M., Luck, M., Bu, Q., Qing, Y., & Cui, H. (2024). AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation. https://arxiv.org/abs/2312.13010
  14. Ishibashi, Y., & Nishimura, Y. (2024). Self-Organized Agents: A LLM Multi-Agent Framework toward Ultra Large-Scale Code Generation and Optimization. CoRR, abs/2404.02183. 10.48550/arXiv.2404.02183
  15. Islam, Md. A., Ali, M. E., & Parvez, M. R. (2025). CodeSim: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging. In L. Chiruzzo, A. Ritter, & L. Wang (Eds.), Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 5128–5154). Association for Computational Linguistics. 10.18653/v1/2025.findings-naacl.285
  16. Jahns, K. (2024). Yjs: A CRDT framework for shared editing. https://github.com/yjs/yjs
  17. JetBrains. (2026). Air Launches as Public Preview: A New Wave of Dev Tooling. https://blog.jetbrains.com/air/2026/03/air-launches-as-public-preview-a-new-wave-of-dev-tooling-built-on-26-years-of-experience/
  18. Kim, S., Suk, J., Longpre, S., Lin, B. Y., Shin, J., Welleck, S., Neubig, G., Lee, M., Lee, K., & Seo, M. (2024). Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. https://arxiv.org/abs/2405.01535
  19. Kleppmann, M., & Beresford, A. R. (2017). A Conflict-Free Replicated JSON Datatype. IEEE Transactions on Parallel and Distributed Systems, 28(10), 2733–2746. 10.1109/tpds.2017.2697382
  20. Lamport, L. (1978). Time, Clocks, and the Ordering of Events in a Distributed System. Communications of the ACM, 21(7), 558–565. 10.1145/359545.359563
  21. Li, B., Wu, W., Tang, Z., Shi, L., Yang, J., Li, J., Yao, S., Qian, C., Hui, B., Zhang, Q., Yu, Z., Du, H., Yang, P., Lin, D., Peng, C., & Chen, K. (2024). Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study. https://arxiv.org/abs/2403.08604
  22. Li, D., Sun, R., Huang, Y., Zhong, M., Jiang, B., Han, J., Zhang, X., Wang, W., & huan liu. (2026). Preference Leakage: A Contamination Problem in LLM-as-a-judge. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=grIvSXVJ65
  23. Liu, J., Xia, C. S., Wang, Y., & Zhang, L. (2023). Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. https://arxiv.org/abs/2305.01210
  24. Mantel, N., & Haenszel, W. (1959). Statistical Aspects of the Analysis of Data From Retrospective Studies of Disease. Journal of the National Cancer Institute, 22(4), 719–748. 10.1093/jnci/22.4.719 back: 1, 2
  25. Nguyen, M. H., Chau, T. P., Nguyen, P. X., & Bui, N. D. Q. (2024). AgileCoder: Dynamic Collaborative Agents for Software Development based on Agile Methodology. https://arxiv.org/abs/2406.11912 back: 1, 2
  26. Ongaro, D., & Ousterhout, J. (2014). In Search of an Understandable Consensus Algorithm. 2014 USENIX Annual Technical Conference (USENIX ATC 14), 305–319. https://www.usenix.org/conference/atc14/technical-sessions/presentation/ongaro
  27. Oster, G., Urso, P., Molli, P., & Imine, A. (2006). Data Consistency for P2P Collaborative Editing. Proceedings of the 2006 20th Anniversary Conference on Computer Supported Cooperative Work (CSCW), 259–268. 10.1145/1180875.1180916
  28. Pan, R., Zhang, H., & Liu, C. (2025). CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation. https://arxiv.org/abs/2501.07811
  29. Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. https://arxiv.org/abs/2304.03442
  30. Peng, Y., Hou, H., Zhu, X., He, Y. T., & Yu, F. R. (2026). SEMAG: Self-Evolutionary Multi-Agent Code Generation. https://arxiv.org/abs/2603.15707
  31. Power, A., Burda, Y., Edwards, H., Babuschkin, I., & Misra, V. (2022). Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets. https://arxiv.org/abs/2201.02177
  32. Preguiça, N., Marquès, J. M., Shapiro, M., & Letia̧, M. (2009). A Commutative Replicated Data Type for Cooperative Editing. 29th IEEE International Conference on Distributed Computing Systems (ICDCS), 395–403. 10.1109/ICDCS.2009.20
  33. Pugachev, S. (2025). CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation. https://arxiv.org/abs/2510.18893 back: 1, 2, 3
  34. Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z., & Sun, M. (2024). ChatDev: Communicative Agents for Software Development. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 15174–15186). Association for Computational Linguistics. 10.18653/v1/2024.acl-long.810 back: 1, 2, 3
  35. Roh, H.-G., Jeon, M., Kim, J.-S., & Lee, J. (2011). Replicated Abstract Data Types: Building Blocks for Collaborative Applications. Journal of Parallel and Distributed Computing, 71(3), 354–368. 10.1016/j.jpdc.2010.12.006
  36. Shafin, W. I., Rafi, M. N., Li, Z., & Chen, T.-H. (2025). Evaluating Software Process Models for Multi-Agent Class-Level Code Generation. https://arxiv.org/abs/2511.09794
  37. Shapiro, M., Preguiça, N., Baquero, C., & Zawirski, M. (2011). Conflict-Free Replicated Data Types. Symposium on Self-Stabilizing Systems (SSS), 386–400. 10.1007/978-3-642-24550-3_29 back: 1, 2
  38. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: language agents with verbal reinforcement learning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Advances in Neural Information Processing Systems (Vol. 36, pp. 8634–8652). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2023/file/1b44b878bb782e6954cd888628510e90-Paper-Conference.pdf back: 1, 2
  39. Sun, C., Jia, X., Zhang, Y., Yang, Y., & Chen, D. (1998). Achieving Convergence, Causality Preservation, and Intention Preservation in Real-Time Cooperative Editing Systems. ACM Transactions on Computer-Human Interaction, 5(1), 63–108. 10.1145/274444.274447
  40. Verga, P., Hofstatter, S., Althammer, S., Su, Y., Piktus, A., Arkhangorodsky, A., Xu, M., White, N., & Lewis, P. (2024). Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. https://arxiv.org/abs/2404.18796
  41. Wang, J., WANG, J., Athiwaratkun, B., Zhang, C., & Zou, J. (2025). Mixture-of-Agents Enhances Large Language Model Capabilities. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=h0ZfDIrj7T
  42. Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., Tran, H. H., Li, F., Ma, R., Zheng, M., Qian, B., Shao, Y., Muennighoff, N., Zhang, Y., Hui, B., … Neubig, G. (2025). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=OJd3ayDDoF back: 1, 2
  43. Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=1PL1NIMMrw back: 1, 2
  44. Wang, Y., Yu, Z., Zeng, Z., Yang, L., Wang, C., Chen, H., Jiang, C., Xie, R., Wang, J., Xie, X., Ye, W., Zhang, S., & Zhang, Y. (2024). PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization. https://arxiv.org/abs/2306.05087
  45. Wataoka, K., Takahashi, T., & Ri, R. (2025). Self-Preference Bias in LLM-as-a-Judge. https://arxiv.org/abs/2410.21819
  46. Weiss, S., Urso, P., & Molli, P. (2009). Logoot: A Scalable Optimistic Replication Algorithm for Collaborative Editing on P2P Networks. 29th IEEE International Conference on Distributed Computing Systems (ICDCS), 404–412. 10.1109/ICDCS.2009.75
  47. Welch, B. L. (1947). The Generalization of `Student’s’ Problem when Several Different Population Variances are Involved. Biometrika, 34(1–2), 28–35. 10.1093/biomet/34.1-2.28 back: 1, 2
  48. Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2024). AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations. First Conference on Language Modeling. https://openreview.net/forum?id=BAakY1hNKS back: 1, 2, 3
  49. Zhu, K., Du, H., Hong, Z., Yang, X., Guo, S., Wang, Z., Wang, Z., Qian, C., Tang, X., Ji, H., & You, J. (2025). MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 8580–8622). Association for Computational Linguistics. 10.18653/v1/2025.acl-long.421