anon-reranker: Hybrid Attention for Efficient Long-Context Listwise Reranking
Abstract
Listwise rerankers encode a query and many candidates in one sequence, so the concatenated context is a long-context problem: full self-attention grows quadratically with list length and document length. We present anon-reranker, a 0.6B-parameter listwise reranker that keeps the last-but-not-late (LBNL) interaction of anon-reranker-v3 while making that joint context cheaper to run. It replaces uniform global attention with a 3L2G hybrid schedule—three sliding-window layers (w=1024) followed by two global layers—and pins the terminal layer to global (G*) so the trailing query embedding can still observe the full candidate list. A multi-domain mixture and a three-stage self-distillation recipe transfer quality from a full-attention teacher into this sparse student. The model reaches 63.20 nDCG@10 on BEIR, matching a 4B model at roughly 7× fewer parameters, and improves over anon-reranker-v3 on MIRACL, RTEB, and semi-structured retrieval (+9.6 nDCG@10). Latency drops by 1.22× on short web lists and 1.56× on long legal documents (AILACasedocs). We release the weights on Hugging Face under a non-commercial license.