Institution profile

Sany Heavy Industry Co.,Ltd

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder

Apr 02, 2026

This work addresses the limitations of existing image-to-3D multimodal retrieval methods, which predominantly rely on single-view images and struggle to handle the multi-view observations typical of real-world scenarios. To overcome this, the authors propose FusionBERT, a novel framework that first introduces a multi-view visual aggregator leveraging cross-attention mechanisms to adaptively fuse complementary features from multiple viewpoints. Additionally, a normal-aware 3D encoder is incorporated to jointly encode point coordinates and surface normals, thereby enhancing geometric representation for models lacking texture or suffering from color degradation. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches under both single-view and multi-view settings, establishing a strong baseline for image-to-3D cross-modal retrieval.

0 citationsRead paper

AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models

Mar 18, 2026

This work proposes AcceRL, a novel framework that addresses the computational inefficiency and high data demands of large-scale vision-language-action (VLA) models in reinforcement learning. AcceRL introduces, for the first time, a pluggable and trainable world model within a distributed asynchronous reinforcement learning setting. By physically decoupling training, inference, and environment interaction, the framework generates synthetic experiences to dramatically improve sample efficiency. This design overcomes the synchronization bottleneck inherent in conventional approaches, achieving state-of-the-art performance on the LIBERO benchmark. At the algorithmic level, AcceRL significantly enhances training stability and sample efficiency; at the system level, it enables superlinear throughput scaling and high hardware utilization, demonstrating both methodological and engineering advances.

0 citationsRead paper

GIPO: Gaussian Importance Sampling Policy Optimization

Mar 04, 2026

This work addresses the data inefficiency commonly encountered in reinforcement learning during post-training phases, where interaction data are scarce and prone to becoming outdated. The authors propose a novel policy optimization objective based on truncated importance sampling, innovatively incorporating log-ratio Gaussian trust weights to softly suppress extreme importance ratios while preserving non-zero gradients. By replacing hard truncation with an adjustable implicit update constraint, the method achieves a favorable balance between stability and robustness under limited sample budgets. Theoretical analysis grounded in concentration inequalities demonstrates improved bias-variance trade-offs, and empirical evaluations across varying replay buffer sizes consistently show enhanced training stability and sample efficiency.

0 citationsRead paper

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

Nov 23, 2025

Current multimodal large language models (MLLMs) excel at high-level planning in embodied intelligence but exhibit severe deficiencies in fine-grained action understanding. To address this gap, we propose CFG-Bench—the first systematic benchmark for evaluating embodied agents’ fine-grained action cognition, spanning four dimensions: physical interaction, temporal causality, intent comprehension, and evaluative judgment. Built upon 1,368 videos and 19,562 tri-modal question-answer pairs, it establishes the first multimodal fine-grained action evaluation framework. Our empirical analysis reveals significant shortcomings of state-of-the-art MLLMs in generating physically grounded instructions and performing higher-order action reasoning. Furthermore, we demonstrate that supervised fine-tuning substantially improves model performance on real-world embodied tasks. This work advances the translation of visual perception into executable action knowledge and introduces a novel paradigm for action-cognitive modeling in embodied intelligence.

0 citationsRead paper
Recent publications

Latest Papers

FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder

Apr 02, 2026

This work addresses the limitations of existing image-to-3D multimodal retrieval methods, which predominantly rely on single-view images and struggle to handle the multi-view observations typical of real-world scenarios. To overcome this, the authors propose FusionBERT, a novel framework that first introduces a multi-view visual aggregator leveraging cross-attention mechanisms to adaptively fuse complementary features from multiple viewpoints. Additionally, a normal-aware 3D encoder is incorporated to jointly encode point coordinates and surface normals, thereby enhancing geometric representation for models lacking texture or suffering from color degradation. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches under both single-view and multi-view settings, establishing a strong baseline for image-to-3D cross-modal retrieval.

0 citationsRead paper

AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models

Mar 18, 2026

This work proposes AcceRL, a novel framework that addresses the computational inefficiency and high data demands of large-scale vision-language-action (VLA) models in reinforcement learning. AcceRL introduces, for the first time, a pluggable and trainable world model within a distributed asynchronous reinforcement learning setting. By physically decoupling training, inference, and environment interaction, the framework generates synthetic experiences to dramatically improve sample efficiency. This design overcomes the synchronization bottleneck inherent in conventional approaches, achieving state-of-the-art performance on the LIBERO benchmark. At the algorithmic level, AcceRL significantly enhances training stability and sample efficiency; at the system level, it enables superlinear throughput scaling and high hardware utilization, demonstrating both methodological and engineering advances.

0 citationsRead paper

GIPO: Gaussian Importance Sampling Policy Optimization

Mar 04, 2026

This work addresses the data inefficiency commonly encountered in reinforcement learning during post-training phases, where interaction data are scarce and prone to becoming outdated. The authors propose a novel policy optimization objective based on truncated importance sampling, innovatively incorporating log-ratio Gaussian trust weights to softly suppress extreme importance ratios while preserving non-zero gradients. By replacing hard truncation with an adjustable implicit update constraint, the method achieves a favorable balance between stability and robustness under limited sample budgets. Theoretical analysis grounded in concentration inequalities demonstrates improved bias-variance trade-offs, and empirical evaluations across varying replay buffer sizes consistently show enhanced training stability and sample efficiency.

0 citationsRead paper

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

Nov 23, 2025

Current multimodal large language models (MLLMs) excel at high-level planning in embodied intelligence but exhibit severe deficiencies in fine-grained action understanding. To address this gap, we propose CFG-Bench—the first systematic benchmark for evaluating embodied agents’ fine-grained action cognition, spanning four dimensions: physical interaction, temporal causality, intent comprehension, and evaluative judgment. Built upon 1,368 videos and 19,562 tri-modal question-answer pairs, it establishes the first multimodal fine-grained action evaluation framework. Our empirical analysis reveals significant shortcomings of state-of-the-art MLLMs in generating physically grounded instructions and performing higher-order action reasoning. Furthermore, we demonstrate that supervised fine-tuning substantially improves model performance on real-world embodied tasks. This work advances the translation of visual perception into executable action knowledge and introduces a novel paradigm for action-cognitive modeling in embodied intelligence.

0 citationsRead paper