Zero-Write Subgraph Overlay for Low-Latency Inductive GNN Inference on Streaming Graphs
Shivangi Khare
Abstract
\documentclass[11pt,letterpaper]{article} \usepackage[margin=0.75in]{geometry} \usepackage{times} \usepackage{microtype} \usepackage{amsmath,amssymb} \usepackage{booktabs} \usepackage{hyperref} \usepackage{enumitem} \setlist[itemize]{noitemsep, topsep=1pt, parsep=0pt, partopsep=0pt, leftmargin=12pt} \setlength{\parskip}{2pt} \setlength{\parindent}{0pt} \pagestyle{empty} \hypersetup{ colorlinks=true, linkcolor=blue, citecolor=blue, urlcolor=blue } \begin{document} \begin{center} {\Large \textbf{Zero-Write Subgraph Overlay for Low-Latency Inductive GNN Inference on Streaming Graphs}} \\[0.25em] {\normalsize \textit{Anonymous WiML 2026 Submission}} \end{center} \vspace{-0.5em} \noindent\textbf{Abstract.} \textit{Graph Neural Networks (GNNs) effectively detect coordinated fraud across web platforms. However, real-time inference on streaming contributions faces a core bottleneck: persisting unverified entities adds prohibitive write latency and data pollution, while evaluating entities in isolation discards multi-hop relational context. We propose a \textbf{Zero-Write Subgraph Overlay} framework for low-latency inductive GNN inference. By superimposing transient in-flight subgraphs over a persistent historical graph at query time, our system executes temporal sampling and message passing with sub-second latency ($P99 < 300\text{ ms}$) without database writes. Deployed on a large-scale platform, this reduces data staleness from 3 days to minutes, while unified edge annotations cut graph compute by 65.05\% with zero score degradation.} \vspace{0.3em} \noindent\textbf{1. Introduction \& Motivation.} Moderation platforms rely on heterogeneous GNNs to detect collusive abuse across accounts, behavioral signatures, and user contributions. Deploying GNNs on streaming data faces an operational trilemma: (1)~\textbf{Write Latency}: Ingesting events into transactional storage before scoring incurs indexing delays ($O(\text{minutes})$ to $O(\text{hours})$). (2)~\textbf{Topology Loss}: Evaluating unpersisted entities in isolation discards relational structures that reveal abuse rings. (3)~\textbf{Data Pollution}: Writing unverified contributions risks corrupting historical aggregations with adversarial noise. Our \textbf{In-Flight Subgraph Overlay} resolves this by fusing transient candidate subgraphs with persistent topology at query execution time. \vspace{0.3em} \noindent\textbf{2. System Architecture \& Methodology.} The overlay framework decouples feature emission, persistence, and inference into three stages: \begin{itemize} \item \textbf{Transient Subgraph Projection}: Candidate contributions are parsed into atomic entity and relational signals ($\Delta \mathcal{G} = (\mathcal{V}_{\text{transient}}, \mathcal{E}_{\text{transient}})$). To prevent edge overwrites during sampling, relations are strictly directed from the unpersisted seed node $u \in \mathcal{V}_{\text{transient}}$ to existing entities $v \in \mathcal{V}_{\text{persistent}}$. \item \textbf{Dynamic In-Memory Overlay Fusion}: Rather than persisting $\Delta \mathcal{G}$, the inference engine evaluates $\mathcal{G}_{\text{infer}} = \mathcal{G}_{\text{persistent}} \oplus \Delta \mathcal{G}$ at query time. The engine grafts ephemeral seed nodes and directed edges onto the retrieved $k$-hop neighborhood in memory, bypassing database presence checks and write transactions. \item \textbf{Unified Multi-Model Graph via Edge Annotations}: To serve multiple concurrent GNN versions without duplicating pipelines, we introduce semantic edge annotations. Samplers extract model-specific subgraphs on-the-fly via predicate filtering over a single unified graph, decoupling schema evolution from deployment. \end{itemize} \vspace{0.3em} \noindent\textbf{3. Empirical Results \& Production Impact.} The architecture is deployed in production across a web platform serving hundreds of millions of user contributions: \begin{itemize} \item \textbf{Sub-Second Serving Latency}: Delivers end-to-end GNN scoring with $P99 < 300\text{ ms}$ (including 2-hop expansion and neural scoring), reducing freshness latency from \textbf{3 days to $<10\text{ minutes}$}. \item \textbf{65.05\% Compute Reduction}: Consolidating per-model extraction pipelines into a single annotated graph cut graph compute by \textbf{65.05\%} with 100\% score parity. \item \textbf{Abuse Prevention}: In-flight scoring intercepts adversarial edits (e.g., attribute hijacking, fraudulent redirection, synthetic entities) prior to publication. \end{itemize} \vspace{0.3em} \noindent\textbf{4. Conclusion.} The Zero-Write Subgraph Overlay demonstrates that low-latency inductive GNN inference is achievable on streaming data without sacrificing multi-hop depth or incurring database write overhead. \end{document}
Chat is not available.
Successful Page Load