Interventional Analysis of Evidence Faithfulness in Generative Cloud Removal
Abstract
Cloud-free reconstruction from heavily obscured satellite imagery is inherently ambiguous. Task-specific restoration models may produce conservative or incomplete outputs, whereas a strong pretrained generative prior enables sharper and more coherent completion but can also introduce visually plausible structures that are inconsistent with the observed scene. We present \method{}, a generative cloud-removal framework that constrains a pretrained BAGEL backbone using explicit, source-specific optical, spectral, SAR, and temporal condition-token blocks. After establishing strong performance across multichannel, multimodal, and multitemporal benchmarks, we analyze how the model uses this heterogeneous evidence. A structural audit shows that strong-prior restoration changes the failure profile: it reduces missing structures relative to specialist baselines, but may introduce more prediction-only structures. Scene-correspondence interventions further show that \method{} functionally depends on matched NIR, spectral, SAR, and temporal observations rather than merely the presence of additional tokens. At the spatial level, intervention-derived patch attribution identifies spatial evidence groups whose replacement causes substantially greater degradation than random replacement, validating the faithfulness of the attribution ranking. Increasing heterogeneous evidence also consistently reduces prediction-only and missing structures. Nevertheless, residual reference-inconsistent content remains possible under severe ambiguity. These results show that explicit heterogeneous evidence can constrain, but cannot fully eliminate, unreliable completion by a strong generative prior.