Collusion Does Not Localize to Jacobian-Lens Transport Subspaces
Abstract
As LLM agents are deployed in multi-agent systems that make consequential decisions, oversight has to cover coordination the agents are concealing from the transcript. Collusion is detectable from internal activations, but a probe over the full residual stream localises nothing, and a monitor that reads every dimension at every layer is too expensive to run online and too opaque to audit --- neither of which is acceptable for oversight in deployment. A cheap, auditable alternative has been proposed: watch a small structurally-defined subspace, the directions that most influence the model's output, on the theory that concealed intent surfaces there. We test whether that shortcut exists and find it does not. Across three model families, four multi-agent environments and eight depths, the subspace performs like a random subspace of equal dimension under a permutation null, and the result holds under two further constructions of the target. The monitor is not blind: the same procedure does detect the stance an agent states, the largest excess in all three models. It detects what agents say, and not what they withhold. No individual comparison survives correction over 72 tests, so the claim rests on the consistency of the ordering. Practical oversight of multi-agent systems will not get this cheaply, and any monitor that reduces the observed surface owes a capacity-matched control before it is trusted.