SemISP: Semantic-Consistent Diffusion for Cross-Camera RAW-to-sRGB Generation
Abstract
Cross-camera RAW-to-sRGB mapping seeks to convert smartphone RAW captures into DSLR-quality sRGB images, yet the large domain gap between linear RAW and nonlinear sRGB spaces makes this task highly challenging. Existing methods mainly condition on pixel-level supervision and local RAW cues, where content preservation and domain translation remain entangled in low-level measurements. We propose SemISP, a semantic-guided diffusion framework that uses scene semantics as a domain-invariant anchor for cross-domain generation. Specifically, we adopt a frozen DINOv3 as the sRGB semantic backbone and develop a Domain Alignment Adapter to extract sRGB-consistent semantic representations from RAW inputs through parameter-efficient adaptation. Because semantic priors and RAW signals operate at fundamentally different information levels, and the diffusion denoising process follows a coarse-to-fine trajectory, we further design a Time-aware Dual-stream Adaptive Modulation (TDAM) module that separately encodes both condition streams and uses the diffusion timestep to dynamically balance their contributions---emphasizing semantic anchoring at early stages and fine-grained RAW details at later ones. Experiments on two established cross-camera benchmarks show consistent improvements over prior methods in both fidelity and perceptual quality.