Institution profile

Geely

Industry researchasia · cn
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

Jul 29, 2026

Existing prompt inversion methods for text-to-image diffusion models struggle to simultaneously preserve image fidelity and semantic interpretability, often overlooking the critical role of latent noise in structural consistency. This work proposes Dualin, a two-stage joint inversion framework: the first stage integrates vision-language and large language models to generate faithful and human-readable hard prompts, while the second stage employs unconditional DDIM inversion to accurately recover latent noise, achieving dual alignment in both semantics and structure. Dualin is the first approach to unify prompt inversion with latent noise reconstruction, and we theoretically demonstrate that the recovered noise enables flexible editing without re-optimization, overcoming the limitations of prompt-only paradigms. Experiments show that Dualin achieves high-quality prompt inversion and state-of-the-art image fidelity across multiple datasets, laying a foundation for precise and controllable image editing.

0 citationsRead paper

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

Jun 29, 2026

Existing approaches treat controllable generation with diffusion models and dense prediction as disjoint tasks, overlooking the potential of jointly modeling their heterogeneous output distributions. This work proposes the Dense Unified Generation and Perception (DUGP) module and a unified dataset training strategy built upon the MMDiT architecture, enabling non-RGB dense outputs to be modeled simply by duplicating the image branch—thereby achieving mutual enhancement between generative and perceptual tasks. Adhering to Occam’s razor, the method requires no task-specific designs and instead learns the joint distribution of image-geometry pairs through multi-task co-training. The resulting unified model surpasses prior unified approaches and matches the performance of specialized models: generative priors enrich perceptual detail, while perception signals improve structural alignment in generation.

0 citationsRead paper

J-LAW: Joint Localization and Actionable World Modeling via Coupled Latent Factor Graphs

Jun 26, 2026

This work addresses the limitations of traditional SLAM, which lacks predictive planning capabilities, and action-conditioned world models, which often disregard metric consistency and suffer from open-loop drift. The paper presents the first unified formulation of SLAM and action-conditioned world modeling as a joint estimation problem, introducing a coupled factor graph framework. This framework enables bidirectional optimization between localization and actionable world representation through pose-latent coupling factors and supports latent loop closure. By integrating multi-source probabilistic constraints—including observations, action predictions, odometry, latent landmarks, and loop closures—the approach significantly reduces latent prediction RMSE and end-point drift on real-world PushT and WildGS datasets, while generating globally consistent maps that preserve both metric accuracy and actionability.

0 citationsRead paper

Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM

May 03, 2026

Traditional large language models exhibit limited generalization under corpus heterogeneity and fine-grained conditional shifts, while fine-tuning often incurs catastrophic forgetting, and existing meta-learning approaches struggle to scale to large models. This work proposes a novel meta-control paradigm that introduces a learnable β parameter within the SwiGLU module as a meta-signal, enabling an adaptive meta-gating mechanism to modulate the nonlinearity of the feedforward network. A hypernetwork is further designed to dynamically generate β parameters conditioned on arbitrary textual attributes—such as task, domain, persona, or style—thereby unifying multidimensional textual conditions into a single, generalizable control framework. The method achieves state-of-the-art performance across both seen and unseen conditions, significantly outperforming standard fine-tuning and existing meta-learning baselines.

0 citationsRead paper
Recent publications

Latest Papers

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

Jul 29, 2026

Existing prompt inversion methods for text-to-image diffusion models struggle to simultaneously preserve image fidelity and semantic interpretability, often overlooking the critical role of latent noise in structural consistency. This work proposes Dualin, a two-stage joint inversion framework: the first stage integrates vision-language and large language models to generate faithful and human-readable hard prompts, while the second stage employs unconditional DDIM inversion to accurately recover latent noise, achieving dual alignment in both semantics and structure. Dualin is the first approach to unify prompt inversion with latent noise reconstruction, and we theoretically demonstrate that the recovered noise enables flexible editing without re-optimization, overcoming the limitations of prompt-only paradigms. Experiments show that Dualin achieves high-quality prompt inversion and state-of-the-art image fidelity across multiple datasets, laying a foundation for precise and controllable image editing.

0 citationsRead paper

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

Jun 29, 2026

Existing approaches treat controllable generation with diffusion models and dense prediction as disjoint tasks, overlooking the potential of jointly modeling their heterogeneous output distributions. This work proposes the Dense Unified Generation and Perception (DUGP) module and a unified dataset training strategy built upon the MMDiT architecture, enabling non-RGB dense outputs to be modeled simply by duplicating the image branch—thereby achieving mutual enhancement between generative and perceptual tasks. Adhering to Occam’s razor, the method requires no task-specific designs and instead learns the joint distribution of image-geometry pairs through multi-task co-training. The resulting unified model surpasses prior unified approaches and matches the performance of specialized models: generative priors enrich perceptual detail, while perception signals improve structural alignment in generation.

0 citationsRead paper

J-LAW: Joint Localization and Actionable World Modeling via Coupled Latent Factor Graphs

Jun 26, 2026

This work addresses the limitations of traditional SLAM, which lacks predictive planning capabilities, and action-conditioned world models, which often disregard metric consistency and suffer from open-loop drift. The paper presents the first unified formulation of SLAM and action-conditioned world modeling as a joint estimation problem, introducing a coupled factor graph framework. This framework enables bidirectional optimization between localization and actionable world representation through pose-latent coupling factors and supports latent loop closure. By integrating multi-source probabilistic constraints—including observations, action predictions, odometry, latent landmarks, and loop closures—the approach significantly reduces latent prediction RMSE and end-point drift on real-world PushT and WildGS datasets, while generating globally consistent maps that preserve both metric accuracy and actionability.

0 citationsRead paper

Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM

May 03, 2026

Traditional large language models exhibit limited generalization under corpus heterogeneity and fine-grained conditional shifts, while fine-tuning often incurs catastrophic forgetting, and existing meta-learning approaches struggle to scale to large models. This work proposes a novel meta-control paradigm that introduces a learnable β parameter within the SwiGLU module as a meta-signal, enabling an adaptive meta-gating mechanism to modulate the nonlinearity of the feedforward network. A hypernetwork is further designed to dynamically generate β parameters conditioned on arbitrary textual attributes—such as task, domain, persona, or style—thereby unifying multidimensional textual conditions into a single, generalizable control framework. The method achieves state-of-the-art performance across both seen and unseen conditions, significantly outperforming standard fine-tuning and existing meta-learning baselines.

0 citationsRead paper