Survey-to-Text Representation Learning for Cross-Instrument Predictive Modelling with Bootstrap-Regularised Uncertainty Quantification
Abstract
Survey instruments measuring the same underlying construct are often incompatible as they differ in their questions, response scales and scoring systems. As a result, predictive models trained on one instrument cannot be applied to another, and cohorts measured with different questionnaires cannot be pooled for model development. We propose a semantic harmonisation framework that converts structured survey responses into natural-language narratives using a large language model, embeds them into a shared semantic space, and trains predictive models on the resulting representations. Because narrative generation is stochastic, we explicitly model this variability through a variance-regularised objective and an extension based on embedding perturbations. Across two health surveys and two language-model backends, pooling data across incompatible instruments consistently improved discrimination (AUC 0.762 vs. 0.723), while substantially reducing narrative-induced prediction variability (change in predicted risk across alternative narratives: 2.6 vs. 4.4 per- centage points). In a real deployment setting based on the temporary withdrawal of the SF-12 questionnaire during the COVID-19 pandemic, the pooled model maintained substantially more stable predictions than a model trained on a single instrument. These results demonstrate that semantic harmonisation enables predictive modelling across incompatible survey instruments without requiring shared questionnaire items or respondents.