Lookup Is Not All You Need
Abstract
This study measures whether large language models acquire decisive evidence before they commit to an operational decision. Five models are evaluated on 40 scenarios across military, cybersecurity, healthcare, and governance, with five independent episodes per scenario, producing 1,000 multi-turn episodes. In each episode, evidence is available through investigation tools, and the model submits a terminal action. Execution logs record evidence lookup, decisive evidence acquisition, and premature commitment, and an independent judge scores decision defensibility, evidence alignment, and reasoning consistency. Information-seeking lookups occur in 89.9\% to 97.4\% of episodes. Decisive evidence acquisition falls to 30.7\% for Qwen3.7 Flash and 41.2\% for GPT-OSS 120B, and partial premature decisions exceed 50\% in both models. Stated rationales remain consistent with submitted actions across all models (95.6\% to 98.0\%).