Execution-First Synthetic Tool-Use Traces for Small Language Model Agents
Abstract
Small language models (SLMs) are attractive backbones for agentic systems, but their limited parametric capacity increases their reliance on reliable tool use and, consequently, on training data containing tool interactions that actually execute. Query-first synthesis can produce plausible requests that do not correspond to valid tool sequences, compatible parameters, or available data. We propose \textsc{SyntheticAgentTraceQA}, an execution-first framework that constructs workflow templates, assigns tools under data-flow constraints, executes and validates traces, and only then synthesizes user tasks, teacher reasoning, and reference answers. Across four tool ecosystems, fine-tuning Qwen3.5-4B and 9B on the resulting data improves tool execution, reference-trace agreement, and answer production. The recipe that worked best in our experiments is to train only on trajectories that executed successfully and to supervise tool calls and final answers while masking teacher \texttt{} tokens: masked supervision outperformed full supervision in answer production, and full supervision dropped 9B answer production below the base 9B thinking baseline.