Beyond Binary Benchmarks: Telemetry-Grounded Evaluation of On-Device Computer-Using Agents
Samreen Anjum ⋅ Muhammad Saad ⋅ Ibrahim Abu Alhaol ⋅ Gwenael Poitau
Abstract
Small Language Models (SLMs) are a natural substrate for on-device Computer-Using Agents (CUAs), keeping data local, working offline, and avoiding per-token cloud costs. Standard pass/fail benchmarks, measured on datacenter GPUs, say little about whether such agents can actually run on client hardware. We characterize the UI-TARS CUA family on OSWorld across three hardware tiers, from H100 servers to workstation and consumer AI PCs, using a telemetry harness that logs per-step latency, token growth, VRAM composition, and GPU utilization. Three bottlenecks emerge. Dynamic context state, not static model weights, dominates VRAM. Re-encoding full-frame screenshots inflates prompts by roughly 11.9K tokens per step, consuming over half the context window within six steps. Compute is duty-cycled, with brief saturation bursts separated by multi-second idle gaps. We also derive a trace-level taxonomy of agent failure modes. Motivated by these findings, we build AgOS, an OS-integrated system pairing a 2B-parameter SLM agent with NPU intent routing and vector task memory, and evaluate it on a separate 30-task suite (300 trials) on a consumer AI PC. AgOS raises task success rate from 57$\%$ to 71$\%$ while cutting VLM calls by 2.0×, tokens by 2.3×, and latency by 1.8×. These results advocate for evaluating on-device agents on memory, context, and latency traces rather than success rate alone.
Chat is not available.
Successful Page Load