Execution Fragility in LLM Inference: Margin–Displacement Geometry and Strong Baselines for Diversity-Driven Robustness Search
Abstract
Execution-only changes can alter greedy LLM decisions even when model weights and the logical prefix are fixed. We use margin--displacement geometry as an organizing framework: a competitor can replace the reference winner only when its execution-induced relative shift clears its reference gap; empirically, low reference margin localizes where this happens in our workloads. Across 168 parent prompts, only 37 unique parents exhibit an observed positive-reference-margin flip. A discovery-derived raw screen retains 5.9--11.2\% of evaluated pairs while containing every observed flip, but clean held-out, unseen-prompt, and cross-family conditions already contain all flips at half that threshold, so we do not treat the raw constant as transferable. Decision-directed displacement alone is less predictive than margin on matched diagnostics, and the necessary-condition indicator has perfect recall but low precision, confirming that the elementary geometry is permissive rather than a new predictive law. Search results lead to a second finding: after the first observed flip for a state, simply expanding its remaining execution variants matches BC-QD on discovery and has higher AUC on all five out-of-sample matrices. This is a deliberately narrow, single-token A100 study (models less than or equal to 1.7B; eleven framework-level regimes), but it yields a useful evaluation rule for diversity-driven robustness search: control for susceptibility and adaptive within-state exploitation before crediting diversity pressure.