RAOP: Step-Level Resource Orchestration for LLM Agents across Edge and Cloud
Abstract
LLM agents do not present a fixed inference job to an edge and cloud system. At each step, a placement decision chooses where to generate the next message, where to execute any resulting tool call, which prefix must be replayed, which cache state is updated, and whether the episode continues. As a result, the workload to be scheduled is partly produced by the scheduling decisions themselves. We formalize this problem as an endogenous, sequentially revealed workload (ESRW) controlled trajectory MDP. We then introduce Resource Aware Orchestration Policy (RAOP), a step level orchestration policy that first places language generation and then, after observing any tool call, places tool execution. RAOP predicts candidate action workloads from prefix replay, prefill, KV cache residency, queue, network, and tool payload features, and is trained with GRPO to optimize measured full trajectory success and latency rather than a one step latency surrogate. To make controlled evaluation auditable, the main simulator replays language and tool continuations produced by real model and tool execution while recomputing physical timing under the sampled placement sequence. As an analytical lens, we derive a Bellman regret decomposition showing that myopic placement policies incur error from workload prediction, success priors, and a controlled endogeneity term. On GSM8K, HotpotQA, WebShop, and a calibrated three tier deployment, RAOP matches the strong cloud only success reference while reducing P99 latency by about 53\% and communication by about 73\%. Against the strongest learned step level baseline, RAOP improves average success by 2.5 points, indicating that full trajectory optimization is central for agent orchestration.