On-Device Financial Claim Verification at Cloud-Level Accuracy
Abstract
Financial claim verification often requires grounding short claims in long regulatory filings, making cloud-only inference costly and potentially unsuitable for sensitive documents. We build a routed edge-cloud pipeline that runs two small models on device and escalates only when a disagreement gate or an arithmetic trigger fires, with no training or fine-tuning. On 1,700 held-out claims from FINDVER, the routed system reaches 75.8% accuracy versus 77.4% for a cloud-only baseline, with no significant overall difference, while using 56.6% of the cloud tokens and settling 46.6% of claims entirely on device. The aggregate tie, however, hides a significant loss on one subset: agreement between two models of the same family becomes false confidence because their errors are correlated, costing 8.7 points on information-extraction claims. Arithmetic claims show the complementary failure mode, where both local models fail in the same direction and an unconditional trigger succeeds where disagreement cannot. These results show that effective edge-cloud verification depends not only on local model accuracy, but on whether the routing signal can observe the failures that matter.