What Makes a Good Function Vector?
Abstract
Function vectors (FVs) are task representations elicited during in-context learning that can be used to steer Large Language Models. Yet there is little consensus on how FVs should be defined, and design choices in their construction remain poorly understood. We hypothesize that faithful task representations depend on the intensional frame (IF): the most salient subset of instruction tokens that is essential for task execution. We test this by comparing activation-based and gradient-based head selection for FV extraction. Restricting the computation of the attribution scores to the IF improves steering accuracy by 18% on average, confirming that task-defining information is concentrated in few tokens. Gradient-based attribution via Layer-wise Relevance Propagation recovers the most optimal FV at a ~500x speedup, and injecting selected head outputs in a distributed manner considerably outperforms aggregating them into a single vector.