Institution profile

SIASUN Robot & Automation CO., Ltd

Industry researchasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Triple-S: A Collaborative Multi-LLM Framework for Solving Long-Horizon Implicative Tasks in Robotics

Aug 10, 2025

LLMs frequently misconfigure API parameters, omit critical annotations, and violate execution ordering constraints when executing long-horizon implicit robotic tasks. To address these challenges, this paper proposes a multi-LLM collaborative closed-loop framework. Our method introduces: (1) a role-based, three-stage collaboration mechanism—Simplification (task decomposition and understanding), Solution (code generation), and Summary (result abstraction)—to decouple and specialize each functional phase; and (2) a success-driven dynamic demonstration library that enables generalized repair of failed tasks via context-aware, real-time demonstration updates. Integrating in-context learning with adaptive demonstration retrieval, our approach achieves 89% task success on the LDIP benchmark across both fully and partially observable settings. Extensive validation confirms robust performance in both simulation and real-world robotic platforms.

0 citationsRead paper

UGNA-VPR: A Novel Training Paradigm for Visual Place Recognition Based on Uncertainty-Guided NeRF Augmentation

Mar 27, 2025IEEE Robotics and Automation Letters

Existing visual place recognition (VPR) datasets are predominantly captured from single viewpoints, leading to poor generalization under multi-directional driving or in feature-sparse environments; acquiring additional real-world multi-view data is prohibitively expensive. To address this, we propose an uncertainty-guided, NeRF-based self-supervised data augmentation paradigm: a NeRF scene model is reconstructed from single-view images, and an uncertainty estimation network identifies reconstruction-weak regions to guide the synthesis of high-informativeness multi-view observations. A hybrid memory storage strategy is further introduced to improve training efficiency. Our method requires no additional real-world acquisition yet significantly enhances the generalization of VPR models—including NetVLAD, CosPlace, and PittsBERT—under multi-view and sparse-scene conditions. Evaluated on three standard benchmarks and a newly constructed indoor-outdoor dataset, our approach achieves up to 12.7% improvement in recall, consistently outperforming state-of-the-art methods. Code and data are publicly available.

0 citationsRead paper
Recent publications

Latest Papers

Triple-S: A Collaborative Multi-LLM Framework for Solving Long-Horizon Implicative Tasks in Robotics

Aug 10, 2025

LLMs frequently misconfigure API parameters, omit critical annotations, and violate execution ordering constraints when executing long-horizon implicit robotic tasks. To address these challenges, this paper proposes a multi-LLM collaborative closed-loop framework. Our method introduces: (1) a role-based, three-stage collaboration mechanism—Simplification (task decomposition and understanding), Solution (code generation), and Summary (result abstraction)—to decouple and specialize each functional phase; and (2) a success-driven dynamic demonstration library that enables generalized repair of failed tasks via context-aware, real-time demonstration updates. Integrating in-context learning with adaptive demonstration retrieval, our approach achieves 89% task success on the LDIP benchmark across both fully and partially observable settings. Extensive validation confirms robust performance in both simulation and real-world robotic platforms.

0 citationsRead paper

UGNA-VPR: A Novel Training Paradigm for Visual Place Recognition Based on Uncertainty-Guided NeRF Augmentation

Mar 27, 2025IEEE Robotics and Automation Letters

Existing visual place recognition (VPR) datasets are predominantly captured from single viewpoints, leading to poor generalization under multi-directional driving or in feature-sparse environments; acquiring additional real-world multi-view data is prohibitively expensive. To address this, we propose an uncertainty-guided, NeRF-based self-supervised data augmentation paradigm: a NeRF scene model is reconstructed from single-view images, and an uncertainty estimation network identifies reconstruction-weak regions to guide the synthesis of high-informativeness multi-view observations. A hybrid memory storage strategy is further introduced to improve training efficiency. Our method requires no additional real-world acquisition yet significantly enhances the generalization of VPR models—including NetVLAD, CosPlace, and PittsBERT—under multi-view and sparse-scene conditions. Evaluated on three standard benchmarks and a newly constructed indoor-outdoor dataset, our approach achieves up to 12.7% improvement in recall, consistently outperforming state-of-the-art methods. Code and data are publicly available.

0 citationsRead paper