FinVerify: A Verification Framework for Financial Agents
Abstract
LLM-based financial agents read filings, select evidence, and report numbers. These numbers feed investment and reporting decisions. Verifying them remains an open problem. General-purpose claim verifiers ignore the numeric, temporal, and accounting structure of financial claims. Finance-specific verifiers each target a single task, evidence source, or failure mode. Self-critique relies on the very model under test. We first study how financial agents actually fail. We examine all 190 wrong trajectories in the Snorkel Agent Finance Reasoning task. The errors are recurrent: citing stale or unreleased data, citing values absent from the source, combining non-comparable quantities, breaking accounting identities, and applying the wrong formula. From these we derive a financial error taxonomy with five categories: Temporal Integrity, Source Fidelity, Comparability, Accounting Integrity, and Financial-Model Fidelity. We then introduce FinVerify, an agent-agnostic verification gate built on this taxonomy. Each category becomes a protocol. One evidence-bound language-model call structures the question, the frozen evidence, and the agent trace. Deterministic validators then check every extracted quote and run the five protocols. A proved violation yields Reject; missing proof yields Abstain. We evaluate FinVerify under a fixed-candidate-pool Accept/Reject/Abstain evaluation on five pools: Agent Finance Reasoning, FinQA, TAT-QA, FinanceBench-67, and XBRLFiling. We compare against a Direct baseline and an LLM judge. Every verifier sees the same candidate and evidence. FinVerify raises accepted-claim precision on every pool. On the two filing-native pools it admits one false accept in 650 decisions, and reaches 100.0\% and 99.8\% precision. On Agent Finance Reasoning it raises precision from 46.8\% (Direct) and 54.5\% (LLM judge) to 75.2\%. Gains on FinQA and TAT-QA are smaller and come with more abstentions.