Institution profile

International Digital Economy Academy

Academic institutionasia · cn
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models

Jan 04, 2026arXiv.org

This work addresses the limitations of current generative speech models in zero-shot multilingual synthesis and editing, which stem from the scarcity of large-scale, high-quality multilingual speech data with word-level timestamps. To overcome this, the authors introduce LEMAS-Dataset, an open-source corpus spanning 10 languages and 150,000 hours of speech, uniquely annotated with word-level alignment timestamps. Leveraging this dataset, they propose LEMAS-TTS, a non-autoregressive model for zero-shot multilingual text-to-speech synthesis, and LEMAS-Edit, an autoregressive model that formulates speech editing as a masked token infilling task. Through accent adversarial training, CTC loss, and adaptive decoding strategies, the models achieve substantial improvements in cross-lingual accent robustness and naturalness at edit boundaries, demonstrating the efficacy of both the dataset and the proposed methodologies.

1 citationsRead paper

GRAPE: Graduated Routing for Articulated Portrait mesh Estimation

Jul 26, 2026

This work addresses the challenges of high-fidelity 3D human portrait reconstruction from monocular images, particularly the coupling between head and torso motion and the difficulty in disentangling jaw articulation from facial expressions. To tackle these issues, the authors propose the Portrait Parametric Model (PPM), which explicitly models the kinematic chain from torso to head and integrates the canonical spaces of FLAME and SMPL-X. They further introduce a Progressive Anatomical Alignment (PAA) network featuring a graduated-mask router—a coarse-to-fine expert routing mechanism guided by anatomical priors—and multi-source supervision from sparse keypoints, feature distillation, foreground masks, and relative geometric constraints. The method significantly outperforms existing approaches in reconstruction fidelity, pose alignment accuracy, and jaw-expression disentanglement, leading to notable improvements in downstream tasks such as speech-driven talking-head animation and 3D portrait generation.

0 citationsRead paper

Reflective VLA: In-Context Action Consequences Make VLAs Generalize

Jun 23, 2026

Existing vision-language-action (VLA) models are predominantly reactive and struggle to infer robot embodiment characteristics—such as camera geometry or actuation biases—from a single observation, limiting their generalization across environments. This work proposes Reflective VLA, the first approach to explicitly model the environmental impact of actions by treating action consequences as critical contextual information within observation–action–consequence triplets. The method employs a shared-attention architecture grounded in vision-language models, combined with blockwise causal masking to enable efficient parallel training while supporting real-time inference via KV caching. Experiments on LIBERO-Plus and its Hard variant demonstrate consistent improvements under distribution shifts, with average success rates increasing by 5.4 and 4.2 percentage points, respectively, without compromising performance on the original data distribution.

0 citationsRead paper

Mozi: Governed Autonomy for Drug Discovery LLM Agents

Mar 03, 2026

This work addresses the challenge of unreliable and irreproducible reasoning trajectories in large language model (LLM) agents applied to drug discovery, which often stem from unconstrained tool usage and fragile long-horizon reasoning. To mitigate these issues, the authors propose Mozi, a dual-layer architecture that integrates the flexibility of generative AI with the rigor of computational biology through hierarchical supervision–workflow control and a stateful skill graph. Mozi incorporates role isolation, constrained action spaces, reflective replanning, and human-in-the-loop checkpoints to prevent error propagation and ensure scientific validity. Evaluated on the PharmaBench benchmark, Mozi substantially outperforms existing approaches and demonstrates end-to-end therapeutic case completion by efficiently identifying low-toxicity candidate molecules and generating competitive virtual compounds.

0 citationsRead paper

PEAR: Pixel-aligned Expressive humAn mesh Recovery

Jan 30, 2026

This work addresses the challenges of high-fidelity 3D human reconstruction from a single in-the-wild image, where existing methods suffer from slow inference, coarse pose estimation, and inaccurate hand and facial details. The authors propose PEAR, a lightweight, unified Vision Transformer (ViT) framework that jointly estimates SMPL-X and scaled-FLAME parameters in an end-to-end manner without requiring any preprocessing, enabling real-time inference at over 100 FPS. Pixel-level geometric supervision is introduced to refine fine-grained details, while a modular annotation strategy enhances data robustness. Extensive experiments demonstrate that PEAR significantly outperforms state-of-the-art approaches across multiple benchmarks, achieving substantial improvements in both body pose and facial expression reconstruction accuracy.

0 citationsRead paper
Recent publications

Latest Papers

GRAPE: Graduated Routing for Articulated Portrait mesh Estimation

Jul 26, 2026

This work addresses the challenges of high-fidelity 3D human portrait reconstruction from monocular images, particularly the coupling between head and torso motion and the difficulty in disentangling jaw articulation from facial expressions. To tackle these issues, the authors propose the Portrait Parametric Model (PPM), which explicitly models the kinematic chain from torso to head and integrates the canonical spaces of FLAME and SMPL-X. They further introduce a Progressive Anatomical Alignment (PAA) network featuring a graduated-mask router—a coarse-to-fine expert routing mechanism guided by anatomical priors—and multi-source supervision from sparse keypoints, feature distillation, foreground masks, and relative geometric constraints. The method significantly outperforms existing approaches in reconstruction fidelity, pose alignment accuracy, and jaw-expression disentanglement, leading to notable improvements in downstream tasks such as speech-driven talking-head animation and 3D portrait generation.

0 citationsRead paper

Reflective VLA: In-Context Action Consequences Make VLAs Generalize

Jun 23, 2026

Existing vision-language-action (VLA) models are predominantly reactive and struggle to infer robot embodiment characteristics—such as camera geometry or actuation biases—from a single observation, limiting their generalization across environments. This work proposes Reflective VLA, the first approach to explicitly model the environmental impact of actions by treating action consequences as critical contextual information within observation–action–consequence triplets. The method employs a shared-attention architecture grounded in vision-language models, combined with blockwise causal masking to enable efficient parallel training while supporting real-time inference via KV caching. Experiments on LIBERO-Plus and its Hard variant demonstrate consistent improvements under distribution shifts, with average success rates increasing by 5.4 and 4.2 percentage points, respectively, without compromising performance on the original data distribution.

0 citationsRead paper

Mozi: Governed Autonomy for Drug Discovery LLM Agents

Mar 03, 2026

This work addresses the challenge of unreliable and irreproducible reasoning trajectories in large language model (LLM) agents applied to drug discovery, which often stem from unconstrained tool usage and fragile long-horizon reasoning. To mitigate these issues, the authors propose Mozi, a dual-layer architecture that integrates the flexibility of generative AI with the rigor of computational biology through hierarchical supervision–workflow control and a stateful skill graph. Mozi incorporates role isolation, constrained action spaces, reflective replanning, and human-in-the-loop checkpoints to prevent error propagation and ensure scientific validity. Evaluated on the PharmaBench benchmark, Mozi substantially outperforms existing approaches and demonstrates end-to-end therapeutic case completion by efficiently identifying low-toxicity candidate molecules and generating competitive virtual compounds.

0 citationsRead paper

PEAR: Pixel-aligned Expressive humAn mesh Recovery

Jan 30, 2026

This work addresses the challenges of high-fidelity 3D human reconstruction from a single in-the-wild image, where existing methods suffer from slow inference, coarse pose estimation, and inaccurate hand and facial details. The authors propose PEAR, a lightweight, unified Vision Transformer (ViT) framework that jointly estimates SMPL-X and scaled-FLAME parameters in an end-to-end manner without requiring any preprocessing, enabling real-time inference at over 100 FPS. Pixel-level geometric supervision is introduced to refine fine-grained details, while a modular annotation strategy enhances data robustness. Extensive experiments demonstrate that PEAR significantly outperforms state-of-the-art approaches across multiple benchmarks, achieving substantial improvements in both body pose and facial expression reconstruction accuracy.

0 citationsRead paper

LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models

Jan 04, 2026arXiv.org

This work addresses the limitations of current generative speech models in zero-shot multilingual synthesis and editing, which stem from the scarcity of large-scale, high-quality multilingual speech data with word-level timestamps. To overcome this, the authors introduce LEMAS-Dataset, an open-source corpus spanning 10 languages and 150,000 hours of speech, uniquely annotated with word-level alignment timestamps. Leveraging this dataset, they propose LEMAS-TTS, a non-autoregressive model for zero-shot multilingual text-to-speech synthesis, and LEMAS-Edit, an autoregressive model that formulates speech editing as a masked token infilling task. Through accent adversarial training, CTC loss, and adaptive decoding strategies, the models achieve substantial improvements in cross-lingual accent robustness and naturalness at edit boundaries, demonstrating the efficacy of both the dataset and the proposed methodologies.

1 citationsRead paper