How Much Skill Orchestration Can LLM Agents Automatically Compile Away? Towards Token-Efficient Skill Use
Abstract
Reusable skills provide large language model (LLM) agents with procedural guidance, but each invocation still requires the model to interpret the skill and reconstruct its orchestration. Can agents automatically compile this orchestration into an executable fast path without human intervention? We first study zero-shot skill specialization, which incurs a one-time cost to compile a skill and request into a reusable fast path without execution evidence. Despite the procedural guidance provided by the skill, the compiler must speculate about runtime bindings and observations, making generated plans brittle and recovery costly. This motivates SkillCache, a one-trace skill specialization method that compiles a completed invocation into a parameterized capsule of executable regions and residual instructions, while retaining the source skill for recovery. Across all 87 SkillsBench tasks with the Claude Code harness, warm execution reduces aggregate provider-token traffic and model calls by 53.2\% and 43.1\% with Claude Sonnet 4.6 and by 45.6\% and 38.0\% with Claude Opus 4.8, while mean reward changes from 0.555 to 0.534 with Sonnet 4.6 and remains at 0.669 with Opus 4.8. Capsules generated by Opus 4.8 also transfer across executor models, raising DeepSeek-V4-Flash mean reward from 0.376 to 0.648 while reducing aggregate provider-token traffic by 51.8\%. These results support execution-grounded compilation as a promising approach to efficient repeated skill use and strong-to-small execution transfer.