Tokenizer-Generator Coupling in Medical Image Generation
Abstract
Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST study at 64×64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The results show that tokenizer, generator, and sampler cannot be ranked independently. LFQ is the strongest default discrete tokenizer in the matched VQ/LFQ/FSQ grid, but generator rankings change with quantizer family and categorical rate. On LFQ-1024, retuning D3PM and SEDD moves them from default FID-192 0.44/0.41 to 0.09/0.11 at lower NFE, approaching the strongest continuous references. Reconstruction quality is not a reliable selection criterion: the highest-PSNR tokenizers are not the best generation substrates, even when each tokenizer is paired with its best observed generator. We interpret these results as a rate-distortion-modelability tradeoff, where modelability is conditional on the generator, sampler, and inference budget. The study covers 54 discrete tokenizer-generator cells and 16 continuous reference cells, and its scope is deliberately limited: low resolution, one training seed, unconditional generation, and non-clinical FID-based evaluation.