🤖 AI Summary
This study addresses the inherent trade-off between covariate dependence and latent structure in disentangled representation learning by proposing a unified supervised framework that elucidates the constraints linking latent independence with covariate alignment. By establishing disentanglement orderliness and deriving a closed-form transformation for realignment, combined with informed Factor Analysis (iFA), this work enables precise regulation of structured representations in pretrained models. Extensive experiments on both simulated and real-world multi-omics datasets validate the method’s effectiveness, demonstrating significant improvements in the controllability and interpretability of learned representations. Ultimately, this research establishes a novel paradigm for disentangling complex data structures, offering a theoretically grounded solution to balance statistical independence with semantic alignment in high-dimensional representation learning.
📝 Abstract
Disentangled representation learning seeks latent representations whose indicidual dimensions each align with a distinct covariate. Unsupervised approaches typically target latent dimension independence, yet this gives no guarantee that the resulting dimensions align with semantically meaningful covariates. Supervised approaches structure the latent space using observed covariates, but under correlated covariates they cannot simultaneously control one-to-one latent-covariate alignment and latent independence. We introduce a unified, supervised framework that couples latent dimension-covariate dependence with constraints on the latent structure. Within this framework, we show an inherent trade-off, where enforcing latent independence or exclusive one-to-one latent-covariate dependence comes at a provable cost in latent-covariate alignment. We prove that the resulting disentanglement regimes are ordered by the strength of that alignment. Each regime admits a closed-form transformation of the latent space. We apply these transformations post-hoc to realign the representations of pretrained models such as CLIP, DINOv2, and ViT, and we fold them into the inference of informed factor analysis (iFA), a probabilistic model with covariate-informed factors. On simulated and real multi-omics data, we show that both post-hoc alignment and iFA enable controllability of structured latent representations.