Where Natural Language Ends in Agents: Executable Skill Composition with MetaSkill
Abstract
Agent Skills express open-ended task knowledge in natural language, but composing them requires persistent execution order and upstream-output handoffs. We propose \emph{selective formalization at Skill boundaries}: coordination relations with stable references and machine-checkable constraints become executable, while substantive task judgments remain within Skills. We introduce \emph{\mbox{MetaSkill}}, a hybrid representation that preserves complete Skill packages and local instructions while making invocation dependencies and predecessor-text bindings executable through a directed acyclic graph (DAG) runtime. This interface supports explicit coordination without prescribing how each Skill performs its local task. Across 20 multi-Skill artifact tasks and two LLMs, we compare natural-language procedures, prompt-visible graphs, and \mbox{MetaSkill} derived from a common frozen plan per task, using the same designated Skill packages. Additional comparisons include matched-Skills \mbox{AgentSkillOS} and native agent stacks. Blinded artifact judgments yield favorable preference point estimates across all reported comparisons, with 95\% task-bootstrap intervals excluding parity against Prompt~Graph under \mbox{DeepSeek-V4-Flash} and \mbox{OpenClaw} under \mbox{GLM-5.2}. In audited GLM execution traces, all planned nodes complete and all declared precedence constraints are satisfied. Together, these findings support \mbox{MetaSkill} as a practical interface for composing open-ended Skills, combining local linguistic flexibility with executable dependencies, explicit upstream-text bindings, and auditable coordination.