MuRoK: Multi-level Routing for Knowledge Graph Retrieval
Abstract
Prevailing multimodal knowledge graph retrievers typically apply a single, fixed aggregation rule to all candidates, such as additive fusion, image captioning, or two-stage reranking. This approach ignores two critical signals inherently available in the graph, the candidate relation, which determines whether textual or visual evidence is more trustworthy, and the query modality, which determines whether a cross-modal or same-modality comparison is appropriate. To address this limitation, we introduce Multi-level Routing for Knowledge Graph Retrieval (MuRoK), a lightweight two-level routing framework that dynamically adapts retrieval scoring using frozen Contrastive Language–Image Pre-training (CLIP) and Bootstrapping Language Image Pre-training (BLIP) representations, without costly representation learning. At the candidate level, Candidate-Level Expert Routing uses candidate relations to combine five complementary scoring experts, including same-modality text and image matching to mitigate the cross-modal similarity gap. At the query level, Query-Level Modality Dispatch routes queries to modality-specific router instances, addressing the limitations of a single router under imbalanced query-modality distributions and maintaining robust performance across different query regimes. With fewer than 150K trainable parameters, MuRoK provides an efficient routing solution within a frozen-encoder pipeline. Evaluated on the large-scale FB15k-237-IMG benchmark containing 310K triplets, MuRoK consistently outperforms strong baselines across all reported NDCG@K, Precision@K, and Recall@K metrics under different query and candidate modality settings.