Zero-Order Optimization via Learnable Direction Sampling
Valery Parfenov ⋅ Grigoriy Evseev ⋅ Andrey Veprikov ⋅ Nikolay Bushkov ⋅ Stanislav Moiseev ⋅ Aleksandr Beznosikov
Abstract
Directional and zero-order (ZO) methods optimize objectives when first-order information is unavailable or inefficient, but their classical convergence guarantees deteriorate with the ambient dimension because randomly sampled directions are poorly aligned with the gradient. We introduce a plug-and-play policy-driven sampling framework that treats the distribution over directions as a learnable policy. Our nonstandard analysis yields, to the best of our knowledge, the first dimension-free convergence guarantees in this setting. Experiments on LLM fine-tuning tasks further show consistent gain over established baselines accross all setups.
Chat is not available.
Successful Page Load