🤖 AI Summary
This work analyzes the theoretical error of the widely used analytic approximation to the entropy of a Gaussian mixture model (GMM)—namely, the weighted sum of component entropies. Focusing on high-dimensional settings, such as uncertainty quantification in deep neural networks, we derive the first rigorous, explicit upper bound on this approximation error. We prove that the error monotonically decreases with the normalized inter-component separation—the ratio of mean distance to summed variances—and converges to zero as this ratio tends to infinity. Moreover, we characterize an inherent separation-amplification phenomenon in high dimensions, which ensures high accuracy of the approximation in such regimes. Our analysis bridges a fundamental theoretical gap in GMM entropy approximation and provides a solid information-theoretic foundation for efficient and reliable entropy estimation in large-scale models.
📝 Abstract
Gaussian mixture distributions are commonly employed to represent general probability distributions. Despite the importance of using Gaussian mixtures for uncertainty estimation, the entropy of a Gaussian mixture cannot be calculated analytically. In this paper, we study the approximate entropy represented as the sum of the entropies of unimodal Gaussian distributions with mixing coefficients. This approximation is easy to calculate analytically regardless of dimension, but there is a lack of theoretical guarantees. We theoretically analyze the approximation error between the true and the approximate entropy to reveal when this approximation works effectively. This error is essentially controlled by how far apart each Gaussian component of the Gaussian mixture is. To measure such separation, we introduce the ratios of the distances between the means to the sum of the variances of each Gaussian component of the Gaussian mixture, and we reveal that the error converges to zero as the ratios tend to infinity. In addition, the probabilistic estimate indicates that this convergence situation is more likely to occur in higher-dimensional spaces. Therefore, our results provide a guarantee that this approximation works well for high-dimensional problems, such as neural networks that involve a large number of parameters.