When Does a VAE Mix? Aggregate-Posterior KL, Mutual Information, and Encode--Decode Gibbs Chains
Ahmad Ayaz Amin ⋅ Guanghui Wang
Abstract
We study the Markov chain obtained by alternating a variational autoencoder (VAE) decoder and encoder, $z_t\to x_t\sim p_\theta(x\mid z_t)\to z_{t+1}\sim q_\phi(z\mid x_t)$, initialized at the reference prior. The \textbf{qualifying autonomous result} is a decomposition of its sampling behavior: when encoder and decoder are compatible with the variational joint $Q_\phi(x,z)=p_{\rm data}(x)q_\phi(z\mid x)$, the aggregate-posterior term $KL(q_\phi(z)\Vert p(z))$ is exactly the target-to-initialization KL, while $X$--$Z$ dependence controls chain memory. Zero mutual information gives exact one-sweep latent mixing. For linear-Gaussian VAEs, canonical correlations $r_i$ give latent transition eigenvalues $\lambda_i=r_i^2$, autocorrelation $\lambda_i^\ell$, and $MI=-\frac12\sum_i\log(1-\lambda_i)$. Controlled experiments show that the same one nat of MI can yield worst-direction integrated autocorrelation time (IAT) $1.57$ or $13.78$, depending on how information is distributed across latent modes. A population-trained linear $\beta$-VAE further shows that low MI can make the chain fast while incompatibility makes it converge to the wrong stationary law. The resulting diagnostic separates \emph{initialization error}, \emph{memory}, and \emph{target correctness}.
Chat is not available.
Successful Page Load