Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

arXiv:2609.25396v1 Announce Type: cross Abstract: Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for semantic coordination…

aiscience

Sources