Flexible Motion Generation from Language and Style References
为解决文本描述难以捕捉精细动作风格的问题,提出FlexMoGen框架,结合文本和风格示例生成高质量、符合意图且具目标风格的人体动作。
为解决文本描述难以捕捉精细动作风格的问题,提出FlexMoGen框架,结合文本和风格示例生成高质量、符合意图且具目标风格的人体动作。
Existing single-agent world models struggle to capture the distinct influence of each agent’s actions on scene evolution in multi-player dynamic environments. This work proposes the first world model capable of supporting multi-player interaction by explicitly conditioning on multiple agents’ action streams, enabling long-term consistent scene prediction in highly dynamic physical settings. Built upon a 5-billion-parameter latent diffusion architecture, the model integrates autoencoded representations, a dedicated video codec, and a multi-player conditioning mechanism, trained and validated in Rocket League. It generates four-player gameplay videos at 20 FPS in real time on a single NVIDIA B200 GPU, maintaining simulation stability beyond five minutes—with some scenarios remaining coherent for several hours—significantly advancing physical consistency and temporal stability in multi-agent world modeling.
This work addresses the challenge that existing vision-language models struggle to perform moral reasoning aligned with human judgment in multimodal and socially ambiguous contexts. To this end, we introduce MM-SCALE, a large-scale multimodal dataset for moral alignment that uniquely incorporates 5-point scalar ratings and an explicit modality alignment mechanism. Using a custom annotation interface, we collect fine-grained moral acceptability scores and justifications for image-scenario pairs. Leveraging this dataset, we replace conventional binary or pairwise supervision with scalar-based supervision and integrate listwise preference optimization with model fine-tuning to achieve more continuous, nuanced, and granular moral alignment. Experimental results demonstrate that our fine-tuned models significantly outperform baselines in both ranking fidelity and moral safety calibration.
This work addresses the challenging problem of reconstructing high-fidelity, renderable 3D facial models from only a few uncalibrated face images. We propose the first co-optimization framework integrating Gaussian splatting with explicit triangular meshes. Methodologically, we introduce semantic-segmentation-guided geometric alignment and soft mesh constraints to ensure accurate neutral-pose modeling; design a view-dependent mapping from Gaussian points to texture space for generating 4K neural textures; and achieve illumination-decoupled albedo extraction with cross-illumination robust training. Our contributions are threefold: (1) high-quality meshes and textures are generated from Gaussian representations without modifying standard graphics pipelines; (2) fine-grained, animation- and relighting-ready facial assets are produced from merely 11 input images; and (3) strong generalization and practical utility are demonstrated in text-driven 3D face generation tasks.
To address the challenge of autonomous hair styling under fine-grained structural complexity and highly dynamic deformations, this paper proposes the first robot-based hairstyle shaping framework for open-world scenarios. Methodologically: (1) we introduce an action-conditioned latent-space state editing mechanism, integrating a compact, large-scale pre-trained 3D hairstyle latent space with a learned latent dynamics model; (2) leveraging our in-house hair physics simulator for synthetic data generation, we deploy an MPPI-based planner to enable vision-guided closed-loop control. Our contributions include the first high-generalization dynamical modeling of unseen hairstyles and zero-shot transfer capability. In simulation, our method reduces local deformation error by 22% and improves task success rate by 42%. On real synthetic wigs, it robustly achieves complex styling tasks, outperforming the current state-of-the-art system.
为解决文本描述难以捕捉精细动作风格的问题,提出FlexMoGen框架,结合文本和风格示例生成高质量、符合意图且具目标风格的人体动作。
Existing single-agent world models struggle to capture the distinct influence of each agent’s actions on scene evolution in multi-player dynamic environments. This work proposes the first world model capable of supporting multi-player interaction by explicitly conditioning on multiple agents’ action streams, enabling long-term consistent scene prediction in highly dynamic physical settings. Built upon a 5-billion-parameter latent diffusion architecture, the model integrates autoencoded representations, a dedicated video codec, and a multi-player conditioning mechanism, trained and validated in Rocket League. It generates four-player gameplay videos at 20 FPS in real time on a single NVIDIA B200 GPU, maintaining simulation stability beyond five minutes—with some scenarios remaining coherent for several hours—significantly advancing physical consistency and temporal stability in multi-agent world modeling.
This work addresses the challenge that existing vision-language models struggle to perform moral reasoning aligned with human judgment in multimodal and socially ambiguous contexts. To this end, we introduce MM-SCALE, a large-scale multimodal dataset for moral alignment that uniquely incorporates 5-point scalar ratings and an explicit modality alignment mechanism. Using a custom annotation interface, we collect fine-grained moral acceptability scores and justifications for image-scenario pairs. Leveraging this dataset, we replace conventional binary or pairwise supervision with scalar-based supervision and integrate listwise preference optimization with model fine-tuning to achieve more continuous, nuanced, and granular moral alignment. Experimental results demonstrate that our fine-tuned models significantly outperform baselines in both ranking fidelity and moral safety calibration.
This work addresses the challenging problem of reconstructing high-fidelity, renderable 3D facial models from only a few uncalibrated face images. We propose the first co-optimization framework integrating Gaussian splatting with explicit triangular meshes. Methodologically, we introduce semantic-segmentation-guided geometric alignment and soft mesh constraints to ensure accurate neutral-pose modeling; design a view-dependent mapping from Gaussian points to texture space for generating 4K neural textures; and achieve illumination-decoupled albedo extraction with cross-illumination robust training. Our contributions are threefold: (1) high-quality meshes and textures are generated from Gaussian representations without modifying standard graphics pipelines; (2) fine-grained, animation- and relighting-ready facial assets are produced from merely 11 input images; and (3) strong generalization and practical utility are demonstrated in text-driven 3D face generation tasks.
To address the challenge of autonomous hair styling under fine-grained structural complexity and highly dynamic deformations, this paper proposes the first robot-based hairstyle shaping framework for open-world scenarios. Methodologically: (1) we introduce an action-conditioned latent-space state editing mechanism, integrating a compact, large-scale pre-trained 3D hairstyle latent space with a learned latent dynamics model; (2) leveraging our in-house hair physics simulator for synthetic data generation, we deploy an MPPI-based planner to enable vision-guided closed-loop control. Our contributions include the first high-generalization dynamical modeling of unseen hairstyles and zero-shot transfer capability. In simulation, our method reduces local deformation error by 22% and improves task success rate by 42%. On real synthetic wigs, it robustly achieves complex styling tasks, outperforming the current state-of-the-art system.