WovenAnchor Matcher: Specialized Intra- and Inter-Image Context Modeling for Feature Matching
Zhiyang Li ⋅ Ruijiang Jin ⋅ Thibaut Klenke ⋅ Yusuke Sekikawa ⋅ Nakamasa Inoue
Abstract
Semi-dense local feature matching commonly aggregates contextual information before coarse-to-fine correspondence estimation. However, intra-image and inter-image aggregation serve different roles: the former should propagate spatial evidence within each image while retaining structured two-dimensional dependencies, whereas the latter should gather information from plausible corresponding regions while limiting noise from unrelated locations. We propose WovenAnchor Matcher (WAM), a semi-dense matching framework based on role-specialized context aggregation. WAM introduces two complementary operators. WovenMamba performs intra-image aggregation through horizontal-then-vertical state-space propagation, allowing vertical updates to operate on horizontally contextualized features and thereby encouraging a structured two-dimensional receptive-field bias. AnchorCrossAttention performs inter-image aggregation by using cross-attention to gather context from plausible corresponding regions across the image pair. To make this retrieval robust, it operates on local representative anchors rather than dense point-wise tokens, reducing sensitivity to noisy affinities that can otherwise lead to incorrect correspondence propagation. On MegaDepth, WAM achieves 66.1 pose AUC@5$^\circ$ with a runtime of 35.4 ms on a single H100 GPU, improving over JamMa, which achieves 64.1 pose AUC@5$^\circ$ at 43.8 ms under the same setup. Ablations show that removing or replacing either specialized component reduces accuracy, providing evidence that intra-image propagation and inter-image retrieval benefit from role-specific aggregation designs.
Chat is not available.
Successful Page Load