Grounding Long Financial Text with Semantically-Enriched Tables via Multimodal Learning
Abstract
We propose a multimodal transformer framework for efficiently modeling long financial documents that combine structured tabular data with noisy, unstructured textual disclosures. The framework represents financial tables as semantically enriched token sequences and integrates them with long-form text through a cross-attention mechanism that explicitly captures cross-modal dependencies. In particular, we augment each numerical variable with a semantic embedding derived from its textual definition, enabling the model to encode both numerical magnitude and economic meaning. This design yields a more expressive and parameter-efficient representation of tabular data compared to value-only embeddings. Using a large-scale dataset of financial statements and MD&A disclosures, we show that structured numerical features remain the dominant source of predictive signal; however, textual information provides consistent incremental gains when incorporated via cross-attention rather than simple concatenation. Our results further demonstrate that effective multimodal learning in this domain requires grounding noisy textual representations in stable, semantically structured tabular embeddings, particularly when the tabular encoder is pretrained or partially frozen. This highlights the importance of modeling fine-grained interactions between modalities and suggests that cross-attention serves as an efficient mechanism for aligning long-form narrative text with structured financial signals.