SAM3RL: Memory Control via Reinforcement Learning for Visual Object Tracking
Abstract
Segment Anything Model 3 (SAM 3) is a foundation model for image and video segmentation that sets the state of the art in visual object tracking. Its tracker, inherited from SAM 2, stores information from past frames in a memory bank, allowing temporal consistency in complex video sequences. Several recent methods improve SAM 2 with hand-crafted memory update rules, each targeting a specific failure mode, such as distractors, occlusions, or object motion. We instead propose a more general, data-driven approach that frames memory control as a sequential decision-making task. We present SAM3RL, an extension of the SAM 3 tracker with a lightweight memory controller trained by reinforcement learning that decides which frames enter the memory bank in order to maximize tracking performance. SAM3RL sets a new state of the art on 10 of 12 evaluated tracking and segmentation benchmarks without any runtime overhead. It outperforms hand-crafted memory update rules, which are effective on SAM 2 but yield limited gains on the stronger SAM 3. The same approach applied to SAM 2 gives SAM2RL, which improves the baseline on every evaluated benchmark, showing the generality of the introduced method across model generations.