MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents
Abstract
Developers increasingly run not one coding agent but a fleet of them side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern memory. When ten agents each spawn language servers, test runners, and browsers, no standard tool in that stack can say how much memory belongs to which agent, confirm that a terminated agent's descendants are gone, notice a child that has escaped its agent, or keep the machine off the swap cliff when an OOM kill would silently discard uncommitted work. We treat these as runtime-verification problems. An agent-hosting substrate should continuously emit observable signals that an operator or auditor can check while agents run. We present MemMux, a local runtime that turns resource governance into a set of checkable signals (per-agent attribution, complete reclamation, escaped-process visibility, bounded footprint under overcommit, and monitoring overhead) and a claims-disciplined benchmark that measures each one against tmux, a purpose-built agent multiplexer, and a raw-process baseline on identical workloads. Under a binding memory budget on a Linux host, MemMux keeps the fleet under budget (7.5 GiB) with zero swap by admitting a subset and reclaiming under pressure. The ungoverned tools instead run every agent, pin the machine at its RAM ceiling (2× over budget), and spill about 2 GiB into swap. MemMux reclaims 100% of a terminated agent's process subtree where the raw baseline strands half of it (ten orphaned processes, 303 MiB), and it is the only system that surfaces an escaped child, detecting all 10 of 10 injected escapes. The cost is real and we report it: the always-on 1 Hz attribution scan runs near 0.6% CPU for one agent but reaches 2.7% at ten, above our 2% target. Running the same harness on real Claude Code sessions shows that 100% attribution and low overhead carry over to live, multi-process agent trees. We release the engine, the benchmark, and a one-command reproducer.