Test-Time Scaling for Quantum Code Assistants
Abstract
Large language models (LLMs) have demonstrated strong performance in general software and code generation tasks, yet they often struggle with niche domains such quantum code assistants, which require specialized reasoning and domain-specific libraries. Prior approaches address this gap through fine-tuning, but such methods are resource-intensive and can degrade performance on out-of-domain tasks. In this work, we explore test-time scaling (TTS) as an alternative paradigm, focusing on iterative refinement via in-context Thompson sampling (ICTS), which enhances model performance without modifying the underlying model. We show that ICTS with self-generated rich symbolic feedback significantly improves quantum code generation across diverse LLM families and scales, yielding gains of up to 22%, and, when combined with post-selection heuristics, even up to 25%. To improve its stability, we identify excessive exploration in later iteration rounds as a key failure mode and introduce a calibration strategy combining temperature decay and improved prompt structuring. This modification enhances both consistency and overall performance. Through extensive evaluation on 12 models spanning 4 families and 3 benchmarks, we demonstrate that stronger base models benefit most from ICTS, while weaker models require improved task formulations and richer examples to realize gains. Additionally, we highlight instability in fine-tuned models, showing that domain-specific fine-tuning can yield inconsistent improvements even within the same evaluation domain. In contrast, ICTS provides robust, model-agnostic improvements without sacrificing general-purpose capabilities. Overall, our results establish test-time scaling as a practical and effective alternative to fine-tuning, offering an accessible method for enhancing quantum code generation.