Closed by Construction: Bounded Algorithmic Data Processing Without Model-Authored Code
Abstract
Data agents need to perform algorithmic work across heterogeneous datasets and intermediate results. Placing records in the prompt consumes context in proportion to dataset size. Generating code that reads data by reference preserves algorithmic breadth and keeps prompts small, but turns untrusted model output into executable instructions. Sandboxing can constrain that execution; it does not remove the interpreter and runtime from the attack surface. We present CANON, an architecture in which the LLM receives a task, references to data, and a closed catalog of 17 common data operations. It responds with typed steps; the host rejects a request if it violates the schema or an invoker-set resource limit, and otherwise executes a fixed implementation. The model can decline a task when the catalog lacks a required operation. Operation results remain in memory and are referenced by name. The model can inspect bounded views of those results, use its observations to choose the next operation, and iterate toward a final result without writing code or repeatedly serializing the underlying data. While the LLM cannot create new instructions at runtime, developers can add reviewed operations to support additional classes of tasks. We evaluate representative data tasks such as inventory reconciliation, credential-stuffing detection, inactive-subscriber detection, and permission inheritance, which span relational tables, temporal logs, set difference, and trees. Across three models, CANON completes all 120 runs and repairs all 27 rejected calls. Its prompt size remains nearly constant as data grows, while raw-data prompting becomes consistently wrong far below the tested context limit. Comparisons with generated Pandas show similar performance on mechanical tasks and identify a separate requirement for value discovery: the agent must observe an execution result before choosing its next operation. These results demonstrate useful algorithmic breadth, not computational universality. Closure prevents model output from adding executable instructions, but it does not prevent a valid plan from implementing the wrong computation.