STParOpt: An End-to-End Framework for Execution-Aware Parallelism Inference and Optimized CUDA Migration
Abstract
Automatically migrating serial C programs to parallel CUDA is essential for high-performance computing, yet remains challenging due to the partial observability of data dependencies: critical memory access patterns cannot be reliably inferred from source code alone. Existing static analysis methods often fail on irregular and pointer-intensive loops, while LLM-based approaches cannot infer memory-access-dependent optimization strategies from source text alone. To address these limitations, we propose STParOpt, which formulates parallelism and optimization inference as an uncertainty-aware learning problem under partial observability, where ambiguity in static analysis is explicitly quantified and resolved using execution signals. STParOpt augments static program graphs with stride-level execution features via cross-attention, and employs an entropy-guided fusion mechanism that up-weights dynamic evidence when the static branch's predictive entropy over parallelizability is high. Experiments on standard HPC benchmarks demonstrate that STParOpt significantly outperforms both compiler-based and LLM-based baselines on parallelism detection and CUDA code generation. Notably, STParOpt achieves 96.4% parallelism detection accuracy and 86.1% end-to-end generation correctness, surpassing the strongest baseline by +5.0 pp and +11.4 pp respectively.