Where Should a Diffusion-LM Agent Start Its Tools?
Abstract
Waiting for tools is most of the time a small tool-calling agent spends on a task. A diffusion language model (dLLM) offers a way to hide it that autoregressive models do not: at every decoding step it holds a provisional draft of its whole answer, so a complete tool call can be readable while generation is still running. We record every draft of DiffusionGemma 26B-A4B (4B active parameters, one desktop machine, no training) and build CrystalCall, a runtime that launches only side-effect-free calls and hands a result to the agent only if the final program makes exactly that call. On 300 unseen prompts CrystalCall cuts median task time on tool-calling tasks by 18.1% at 500 ms tool delay and more than doubles the tasks that finish correctly inside a 2 s budget (from 23 to 57), without changing any program, tool result or trace. Run live, every early result it handed to the agent was identical on re-execution. The gain does not come from reading uncommitted drafts. A call becomes readable only one decoding step (about 219.5 ms) before the decoder accepts it, so CrystalCall gains exactly what launching at first full acceptance would, and a learned launch rule closes the remaining gap to the oracle only by launching less accurately. Against a same-size autoregressive model on the same prompts, runtime and tool scheduling, the diffusion agent finishes 2.0× faster at equal task success. Our advice to builders: start tools at first full acceptance, hold results back until the final call matches, and measure the acceptance rule before building a predictor.