OfficeInstruct: A Dataset for Tool-Augmented Office Agents on Long-Trajectory Generative Tasks
Abstract
The recent trend of integrating autonomous agents into operating systems, supported by the Model Context Protocol (MCP) tools and Deep Searching, has driven research on efficient assistants for daily office tasks like document writing, spreadsheet analysis, and presentation generation. However, training these agents still faces several challenges: the overemphasis on question-answering tasks at the expense of multi-turn generative tasks; the absence of a training paradigm that jointly enhances coding and tool-calling capabilities; the difficulty of evaluating tool efficiency and artifact quality in the presence of drifting user intents. To address these problems, we deploy a unified multi-turn agent framework to collect real-world user interactions, creating \textit{OfficeInstruct}: a comprehensive dataset of 199k high-quality trajectories across Doc, Spreadsheet, and PPT, with an average of 56.2 steps. To ensure data quality, we propose a graph-based evidence evaluation method that models interaction histories as directed graphs, filtering out redundant tool cycles, execution errors, and hallucinations. Experiments demonstrate that our \textit{OfficeAgent} 8B and 14B models surpass open-source baselines on in-domain and out-of-distribution benchmarks, with the 32B variant attaining a 96.19\% success rate on the newly proposed \textit{OfficeAgentBench}, surpassing proprietary models like Claude-4.5-Sonnet and Gemini-3-Pro.