Marauder's Map: Mapping the reachable argument space of LLM-generated tool calls
Felix Mächtle ⋅ Giulio Zizzo ⋅ Anisa Halimi ⋅ Thomas Eisenbarth ⋅ Mark Purcell
Abstract
Deployed agents can issue tool calls nobody asked them to perform. A coding agent can search the web on a task whose prompt names no URL, fetch a page it noticed in a comment, or install a dependency it decided was missing. No specification fixes the arguments those calls carry: they emerge from the model, the scaffold and the task, over a space that is effectively unbounded. An operator therefore has no measured account of what a LLM agent could ask the world out of its own initiative, which leaves security tool development without a target. We present MARAUDER, a method that seeks to map this space. MARAUDER is a quality-diversity fuzzer whose reward is the observed emission of a single nominated tool call, read out of an instrumented sandbox (rather than from a judge model or a human label), and whose archive of task regions grows from the data instead of being partitioned in advance. The output is a map of the tool's argument space: argument families with a canonical argument, a measured trigger rate, a stability score and coverage per deployment. As a case study, we map the web search of coding agents over twelve deployments, $8{,}280$ tasks and $35{,}314$ sandboxed runs. It resolves into $87$ families. Unseen tasks written from inside a family trigger a search in $0.69$ of draws against $0.12$ for a matched control. In public traces of production coding agents, $77.2\%$ of the queries those agents issue fall into families the map already contains. It is also actionable: ten benign probe pages published for ten families on an ordinary-authority domain reach the top ten of the agent-facing index for eight, and agents handed a page returned that way execute the instructions it carries in $49\%$ of runs.
Chat is not available.
Successful Page Load