Beyond Tool Accuracy: End-to-End Reliability of On-Device SMS Agents Under Homoglyph Attacks
Abstract
On-device SMS security agents must operate under strict latency constraints to intercept phishing links in real time, and therefore rely on lightweight classifiers—rather than large foundation models—as their decision-making tool. However, existing lightweight classifiers remain highly vulnerable to homoglyph attacks, constituting a known bottleneck on the tool side. We trace this vulnerability to a Representation-Decision Paradox: discriminative signals induced by homoglyph perturbations are detectable at the embedding layer, yet dissipate at the decision layer due to architecture-specific inductive biases (over-smoothing in CNNs, decision inertia in RNNs). To address this, we propose Contrastive Homoglyph Tolerance (CHT) loss, which substantially improves classifier sensitivity to homoglyph attacks on sequential architectures while preserving tolerance to random perturbations; relative to a 2B-parameter general-purpose guardrail model, our lightweight tool also reduces end-to-end inference latency by up to 22%. Whether mitigating this tool-side bottleneck suffices to ensure reliable behavior of the overall agent system, however, has not been systematically examined. We integrate the repaired tool into a real agent-tool pipeline, evaluating three small language models from distinct model families (Llama-3.2, Gemma-4, Qwen2.5) as hosts, and cross-validate against three off-the-shelf industry guardrail models. Results show that improving tool quality alone is insufficient to guarantee end-to-end agent reliability: on Qwen2.5, agent safety improves markedly with tool quality, whereas on Llama-3.2 and Gemma-4, the agent exhibits saturated, conservative blocking behavior irrespective of the tool backend in use. This finding reveals substantial, systematic differences in how host models incorporate tool outputs into their final decisions, exposing a blind spot in existing agent evaluation practices concerning end-to-end reliability.