Evaluating and Improving Large Language Models for OpenCode Agents
Abstract
Modern coding assistants are deployed as agents inside harnesses where success depends not only on writing correct code but on operating the harness: emitting schema-valid tool calls, following project-level instructions, delegating to subagents, invoking packaged skills, and respecting permission constraints. Because existing code benchmarks largely grade only the final artifact, a low end-to-end score cannot distinguish weak software-engineering reasoning from failure to operate the harness. We introduce OPENCODEBENCH, a benchmark of 150 tasks graded inside an unmodified opencode environment across five capability categories (code editing, code localization, orchestration, skill invocation, tool restriction), entirely through deterministic rule-based checks including native tool-schema validation and recursive scoring over delegated subagent traces. On OPENCODEBENCH we observe a substantial gap between strong reference models and open-weight base checkpoints, concentrated in categories that explicitly require schema-valid tool use, delegation, skill invocation, and permission compliance. To support adapting open-weight models to this surface, we additionally release OPENCODEDATA, a supervised fine-tuning corpus of ~570K examples organised around the same taxonomy of harness mechanisms. Fine-tuning on OPENCODEDATA substantially improves OPENCODEBENCH accuracy for two open-weight models, and we also observe gains on SWE-bench Verified when both benchmarks are run through the same OpenCode harness.