Lightweight Neural ODEs for LLM Safety–Helpfulness Steering
Kevin Chen ⋅ Eric Hanchen Jiang ⋅ Ying Nian Wu
Abstract
Language-model agents must refuse harmful requests without rejecting benign ones, but conventional fine-tuning is costly and hard-codes a single safety--helpfulness operating point into the model weights. We introduce a neural ODE steering field, a 306K-parameter bottleneck module that learns nonlinear corrections from cached mid-layer activations, trains in under two minutes, and provides continuous inference-time control. On Llama-3.2-1B-Instruct, steering increases harmful-prompt refusal from 0.41 to 0.94 without significantly affecting MMLU performance, while moderate steering nearly doubles harmful-prompt refusal without increasing false refusals on XSTest.
Chat is not available.
Successful Page Load