Mapping LLM-Agent Engineering Decisions from the Literature with an Autonomous Research System
Abstract
Building a system around a large language model means making engineering decisions — how control flows, what memory persists, which tools the agent may call — that current survey coverage does not organize. We present a decision-level map of how these systems are engineered: 1,141 leaf engineering decisions organized into 35 themes over a corpus of 12,060 distinct works, each carrying its options, its placed works and a consensus status — 227 settled, 226 contested, 317 open and 371 thin. The named examples include settled interface conventions and open architecture questions, with counterexamples to both patterns. An autonomous agent ran the investigation end to end, from one OpenAlex query (28,886 records over the window; 22,701 under the academic filter, all fetched) through deduplication, screening of 14,533 distinct works, taxonomy induction, synthesis of all 1,141 decision nodes and 219 sealed full-text readings, to a manuscript draft. An audit of the sealed abstract-level extractions against the full texts finds pooled agreement of 41.2% (77 of 187 eligible readings), a finding about how papers in this corpus frame their contributions. For builders, the map names the conventions to inspect and the works behind them; for researchers, it names the comparisons that would settle what the corpus leaves open.