Where to Hitch? Finding Small Parts Across Long Spatiotemporal Contexts
Cheng-Lin Hsieh ⋅ Zilin Wang ⋅ Stella X. Yu ⋅ Enrique Corona ⋅ Sandeep Sasidharan ⋅ Erol D Sumer
Abstract
Detecting a trailer's coupler across an approach video requires identifying a small attachment point at changing scales. Text-prompted foundation models often identify the surroundings instead of the coupler. Video provides clearer views, but searching every frame with every visual example is expensive. We introduce HitchIt, a training-free method for trailer-coupler localization using a frozen video foundation model. It selects useful views and annotated examples from other videos, then searches from the trailer through its tongue to the coupler. The identified coupler is tracked backward and forward through the recording. No target annotation is supplied in the new video. On 20 densely annotated development videos, HitchIt reaches $0.627$ mean box overlap (mIoU). The tested text-prompted foundation models reach at most $0.067$. Controlled comparisons show that access to useful views improves localization without changing the finder or tracker. The complete method also improves on direct search in a separate test on sparsely labeled videos. These results show how selective use of spatial and temporal context helps locate a small functional part across a video. We will release dense coupler annotations for the 20 videos.
Chat is not available.
Successful Page Load