SPLICE: Structured Prompt Local Iterative Combinatorial Evolution
Dr. Anish Acharya ⋅ Phillip Studans ⋅ Amit Dhanda ⋅ Ninad V Rao ⋅ Vishal A Khatri ⋅ Sina A Niaki ⋅ Brian Verkhovsky
Abstract
Prompt engineering has become a central lever for deploying large language models (LLMs), yet the prompts that actually ship to production bear little resemblance to the short instructions studied in most of the prompt-optimization literature. A deployed prompt is a structured multi-section artifact---role framing, reasoning directives, examples, constraints, tool schemas, and output specifications woven into a single token stream---in which some sections encode hard-won safety and format contracts while others are the legitimate targets of optimization. Existing iterative prompt optimizers treat the prompt as a monolithic string and the search as an unstructured rewrite loop, with no mechanism to respect this partition and no theoretical account of when their updates improve the prompt, preserve diversity, or control length. We address this gap by casting iterative prompt optimization as combinatorial search over an edit graph of block-structured, selectively-mutable prompt states. Under this view, a broad slice of the recent literature collapses to selector--proposal--state-space instances of a single \emph{batched iterative prompt search} template that differ only in the selector. Within this framework we introduce \textsc{Splice}, which contributes four coupled design choices, each paired with a formal guarantee: (i) per-block mutability, with a price-of-freezing bound quantifying the safety--optimality trade-off; (ii) dual section-local textual gradients with momentum, yielding an edit-graph reachability bound and a ${(1-e^{-\gamma})}$ sub-modular approximation; (iii) elitist beam search with lineage-aware diversity, giving deterministic monotonicity, a noisy no regression tail bound, and conditional no-collapse of lineages; and (iv) a ratio regularizer for length control, with an explicit accuracy--length bias bound. Together these constitute, to our knowledge, the first theoretical analysis of iterative LLM-driven prompt optimization. Empirically, \textsc{Splice} consistently outperforms prior prompt-optimization methods and single-step sampling baselines across public benchmarks, system models spanning three capability tiers, and tasks ranging from classification to LLM-as-judge rubric optimization, while preserving frozen sections verbatim and exhibiting narrower cross-model variance than single-step sampling.
Chat is not available.
Successful Page Load