Institution profile

Glam AI

Industry research
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Lightweight Optimal-Transport Harmonization on Edge Devices

Nov 16, 2025

To address the unnatural appearance of augmented reality (AR) scenes caused by color inconsistency between virtual objects and real-world backgrounds, this paper proposes a lightweight, real-time color harmonization method. Grounded in optimal transport theory, our approach employs a compact encoder to directly predict the Monge–Kantorovich transport map for pixel-level color transfer. Notably, this is the first work to adapt optimal transport to on-device AR color harmonization, enabling efficient inference on edge devices. Our key contributions are: (1) the first pixel-accurate, manually annotated dataset specifically designed for AR color harmonization, along with an open-source toolkit for data acquisition; and (2) state-of-the-art performance on real AR composite images—achieving superior visual quality while maintaining real-time efficiency, thus attaining an optimal trade-off between fidelity and computational cost.

0 citationsRead paper

Hessian Geometry of Latent Space in Generative Models

Jun 12, 2025

Understanding the geometric structure of latent spaces in generative models—particularly diffusion models and statistical physics–inspired models—remains a fundamental challenge. Method: We propose learning the logarithm of the partition function via posterior approximation to construct an exponential-family Fisher information metric, thereby linking latent-space Hessian geometry to thermodynamic phase transition theory. Contribution/Results: We theoretically and empirically demonstrate the existence of fractal phase boundaries in latent space, where the Lipschitz constant diverges and the metric exhibits discontinuous transitions. Geodesic interpolation remains approximately linear within phases but fails across phase boundaries. Our approach outperforms baselines on Ising and TASEP models and successfully identifies and characterizes phase-transition structures in diffusion models. This work establishes a novel geometric framework for analyzing the intrinsic dynamics and generalization mechanisms of generative models, bridging statistical physics, differential geometry, and deep generative modeling.

0 citationsRead paper

Training-Free Voice Conversion with Factorized Optimal Transport

Jun 11, 2025

This paper addresses zero-shot, arbitrary-to-arbitrary cross-lingual voice conversion with only 5 seconds of reference speech, requiring no training. The method leverages unsupervised cross-lingual alignment and WavLM representations to achieve robust content–acoustic disentanglement. Key contributions include: (1) replacing conventional kNN regression with a factorized optimal transport mapping for speaker identity transfer; (2) introducing the Monge–Kantorovich linear solution (MKL) within a WavLM feature subspace to mitigate anisotropic variance in high-dimensional embeddings; and (3) enabling fully zero-shot, language-agnostic conversion via geometrically principled feature transport. Experiments on LibriSpeech and FLEURS demonstrate substantial improvements in content fidelity and robustness to short-duration references. The proposed approach matches or exceeds the cross-lingual performance of supervised methods such as FACodec, while eliminating reliance on parallel data, speaker-specific fine-tuning, or language-pair specialization.

0 citationsRead paper
Recent publications

Latest Papers

Lightweight Optimal-Transport Harmonization on Edge Devices

Nov 16, 2025

To address the unnatural appearance of augmented reality (AR) scenes caused by color inconsistency between virtual objects and real-world backgrounds, this paper proposes a lightweight, real-time color harmonization method. Grounded in optimal transport theory, our approach employs a compact encoder to directly predict the Monge–Kantorovich transport map for pixel-level color transfer. Notably, this is the first work to adapt optimal transport to on-device AR color harmonization, enabling efficient inference on edge devices. Our key contributions are: (1) the first pixel-accurate, manually annotated dataset specifically designed for AR color harmonization, along with an open-source toolkit for data acquisition; and (2) state-of-the-art performance on real AR composite images—achieving superior visual quality while maintaining real-time efficiency, thus attaining an optimal trade-off between fidelity and computational cost.

0 citationsRead paper

Hessian Geometry of Latent Space in Generative Models

Jun 12, 2025

Understanding the geometric structure of latent spaces in generative models—particularly diffusion models and statistical physics–inspired models—remains a fundamental challenge. Method: We propose learning the logarithm of the partition function via posterior approximation to construct an exponential-family Fisher information metric, thereby linking latent-space Hessian geometry to thermodynamic phase transition theory. Contribution/Results: We theoretically and empirically demonstrate the existence of fractal phase boundaries in latent space, where the Lipschitz constant diverges and the metric exhibits discontinuous transitions. Geodesic interpolation remains approximately linear within phases but fails across phase boundaries. Our approach outperforms baselines on Ising and TASEP models and successfully identifies and characterizes phase-transition structures in diffusion models. This work establishes a novel geometric framework for analyzing the intrinsic dynamics and generalization mechanisms of generative models, bridging statistical physics, differential geometry, and deep generative modeling.

0 citationsRead paper

Training-Free Voice Conversion with Factorized Optimal Transport

Jun 11, 2025

This paper addresses zero-shot, arbitrary-to-arbitrary cross-lingual voice conversion with only 5 seconds of reference speech, requiring no training. The method leverages unsupervised cross-lingual alignment and WavLM representations to achieve robust content–acoustic disentanglement. Key contributions include: (1) replacing conventional kNN regression with a factorized optimal transport mapping for speaker identity transfer; (2) introducing the Monge–Kantorovich linear solution (MKL) within a WavLM feature subspace to mitigate anisotropic variance in high-dimensional embeddings; and (3) enabling fully zero-shot, language-agnostic conversion via geometrically principled feature transport. Experiments on LibriSpeech and FLEURS demonstrate substantial improvements in content fidelity and robustness to short-duration references. The proposed approach matches or exceeds the cross-lingual performance of supervised methods such as FACodec, while eliminating reliance on parallel data, speaker-specific fine-tuning, or language-pair specialization.

0 citationsRead paper