Institution profile

Academy of Military Sciences

Academic institutionasia · cn
Research library61linked papers
Opportunities0open roles
Selected work

Representative Papers

LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering

Aug 07, 2026

This work addresses the limitations of existing AI-generated image detection methods, which often suffer from the loss of low-level texture details due to downsampling and exhibit limited generalization capability. To overcome these challenges, the study formulates high-resolution generated image detection as a visual question answering task and introduces a novel three-branch multimodal architecture. This architecture integrates a custom-designed low-level texture branch, a high-level semantic branch based on SigLIP2, and textual descriptions generated by BLIP-2, enabling an end-to-end differentiable detection model. A hierarchical fusion mechanism is further devised to effectively combine features across low-level, high-level, and semantic modalities. The proposed approach achieves high accuracy and strong robustness on high-resolution images synthesized by diverse diffusion and autoregressive models, significantly enhancing generalization to unseen generative models.

0 citationsRead paper

3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view Synthesis

Jun 23, 2026

Existing single-view 3D vehicle generation methods suffer from limited viewpoint coverage and cross-view geometric inconsistencies, resulting in low reconstruction fidelity. This work proposes a scalable generative framework that, for the first time, synthesizes an arbitrary number of geometrically consistent multi-view images from a single real-world input image by integrating explicit 3D priors into a diffusion model. The approach leverages a fast mesh reconstruction algorithm combining 3D Gaussian splatting with joint color-normal optimization, enabling high-fidelity geometry recovery. Evaluated on both synthetic and real-world datasets, the method significantly outperforms current state-of-the-art approaches, achieving detailed, view-coherent, and geometrically consistent 3D vehicle reconstructions.

0 citationsRead paper

MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving

Jun 23, 2026

Existing methods struggle to fuse multi-view images and LiDAR point clouds for generating geometrically accurate and high-fidelity 3D vehicle models in real-world driving scenarios. This work proposes MM-TRELLIS, the first approach to incorporate LiDAR point clouds as test-time guidance within a native 3D diffusion generative model. By conditioning on multi-view images and enforcing geometric alignment during the denoising process, MM-TRELLIS achieves highly consistent generation through multimodal fusion. Additionally, it introduces a voxel filtering strategy based on 3D Gaussian Splatting opacities to effectively suppress floating artifacts. Evaluated on the Waymo dataset, the method significantly outperforms existing approaches, achieving state-of-the-art performance in fidelity, geometric accuracy, and cross-view consistency of the generated vehicle models.

0 citationsRead paper

DynaMOMA: Instantaneous Prediction of Grasp Poses for Mobile Manipulation of Dynamic Objects

Jun 23, 2026

This work proposes a dynamic mobile manipulation framework to address the challenges of coordinating a mobile base with an arm and generating temporally consistent grasping trajectories in real time when interacting with dynamic objects. The framework innovatively couples an anchor-based diffusion model with a whole-body reinforcement learning policy: the diffusion model predicts instantaneous grasp trajectories and encodes them into compact features, while the reinforcement learning policy leverages a forward-looking reward mechanism to anticipate dynamic targets and coordinate precise grasps. Evaluated in Isaac Gym simulations across diverse dynamic scenarios and multiple grasping metrics, the system demonstrates strong generalization capabilities and successfully transfers to real-world robotic platforms.

0 citationsRead paper

HiFiVe: High-Fidelity Vehicle Generation Leveraging Auto-Regressive 2D Generative Priors

Jun 23, 2026

Existing 3D vehicle generation methods suffer from low geometric fidelity and blurry textures, limiting their utility for downstream tasks. This work proposes a training-free optimization framework that leverages 2D generative priors anchored by 3D geometric constraints to jointly refine geometry and texture. A novel autoregressive texture refinement pipeline is introduced, integrating depth-guided multi-view fusion and vehicle symmetry priors to enforce cross-view consistency and mitigate error accumulation. Furthermore, high-frequency geometric details are recovered via normal map inversion, enabling mesh refinement. Evaluated on both synthetic and real-world vehicle datasets, the proposed method significantly outperforms current state-of-the-art approaches, achieving notable improvements in both geometric accuracy and texture quality.

0 citationsRead paper
Recent publications

Latest Papers

LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering

Aug 07, 2026

This work addresses the limitations of existing AI-generated image detection methods, which often suffer from the loss of low-level texture details due to downsampling and exhibit limited generalization capability. To overcome these challenges, the study formulates high-resolution generated image detection as a visual question answering task and introduces a novel three-branch multimodal architecture. This architecture integrates a custom-designed low-level texture branch, a high-level semantic branch based on SigLIP2, and textual descriptions generated by BLIP-2, enabling an end-to-end differentiable detection model. A hierarchical fusion mechanism is further devised to effectively combine features across low-level, high-level, and semantic modalities. The proposed approach achieves high accuracy and strong robustness on high-resolution images synthesized by diverse diffusion and autoregressive models, significantly enhancing generalization to unseen generative models.

0 citationsRead paper

3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view Synthesis

Jun 23, 2026

Existing single-view 3D vehicle generation methods suffer from limited viewpoint coverage and cross-view geometric inconsistencies, resulting in low reconstruction fidelity. This work proposes a scalable generative framework that, for the first time, synthesizes an arbitrary number of geometrically consistent multi-view images from a single real-world input image by integrating explicit 3D priors into a diffusion model. The approach leverages a fast mesh reconstruction algorithm combining 3D Gaussian splatting with joint color-normal optimization, enabling high-fidelity geometry recovery. Evaluated on both synthetic and real-world datasets, the method significantly outperforms current state-of-the-art approaches, achieving detailed, view-coherent, and geometrically consistent 3D vehicle reconstructions.

0 citationsRead paper

MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving

Jun 23, 2026

Existing methods struggle to fuse multi-view images and LiDAR point clouds for generating geometrically accurate and high-fidelity 3D vehicle models in real-world driving scenarios. This work proposes MM-TRELLIS, the first approach to incorporate LiDAR point clouds as test-time guidance within a native 3D diffusion generative model. By conditioning on multi-view images and enforcing geometric alignment during the denoising process, MM-TRELLIS achieves highly consistent generation through multimodal fusion. Additionally, it introduces a voxel filtering strategy based on 3D Gaussian Splatting opacities to effectively suppress floating artifacts. Evaluated on the Waymo dataset, the method significantly outperforms existing approaches, achieving state-of-the-art performance in fidelity, geometric accuracy, and cross-view consistency of the generated vehicle models.

0 citationsRead paper

DynaMOMA: Instantaneous Prediction of Grasp Poses for Mobile Manipulation of Dynamic Objects

Jun 23, 2026

This work proposes a dynamic mobile manipulation framework to address the challenges of coordinating a mobile base with an arm and generating temporally consistent grasping trajectories in real time when interacting with dynamic objects. The framework innovatively couples an anchor-based diffusion model with a whole-body reinforcement learning policy: the diffusion model predicts instantaneous grasp trajectories and encodes them into compact features, while the reinforcement learning policy leverages a forward-looking reward mechanism to anticipate dynamic targets and coordinate precise grasps. Evaluated in Isaac Gym simulations across diverse dynamic scenarios and multiple grasping metrics, the system demonstrates strong generalization capabilities and successfully transfers to real-world robotic platforms.

0 citationsRead paper

HiFiVe: High-Fidelity Vehicle Generation Leveraging Auto-Regressive 2D Generative Priors

Jun 23, 2026

Existing 3D vehicle generation methods suffer from low geometric fidelity and blurry textures, limiting their utility for downstream tasks. This work proposes a training-free optimization framework that leverages 2D generative priors anchored by 3D geometric constraints to jointly refine geometry and texture. A novel autoregressive texture refinement pipeline is introduced, integrating depth-guided multi-view fusion and vehicle symmetry priors to enforce cross-view consistency and mitigate error accumulation. Furthermore, high-frequency geometric details are recovered via normal map inversion, enabling mesh refinement. Evaluated on both synthetic and real-world vehicle datasets, the proposed method significantly outperforms current state-of-the-art approaches, achieving notable improvements in both geometric accuracy and texture quality.

0 citationsRead paper