Institution profile

Agency for Defense Development

Academic institutionasia · kr
Official website
Research library22linked papers
Opportunities0open roles
Selected work

Representative Papers

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction

Jul 22, 2026

Existing diffusion-based approaches for robotic sequential prediction rely on unidirectional denoising, which struggles to maintain global consistency in long-horizon tasks. This work proposes a reversible denoising mechanism that incorporates a structured re-noising strategy, selectively re-adding noise to temporally stable local regions during the diffusion process. This enables iterative refinement by integrating cross-temporal contextual information, facilitating mutual correction between early and late segments of the predicted sequence. Implemented within a unified video-action diffusion framework, the method combines context-aware denoising with selective re-noising, achieving up to a 56.5% improvement in average success rate on the OGBench and LIBERO-10 benchmarks. Moreover, it demonstrates enhanced robustness and stronger action-video consistency in out-of-distribution scenarios.

0 citationsRead paper

GARAGE: Characterizing the Automation Boundary in LLM-based Attack Graph Generation

Jul 20, 2026

Existing automated tools struggle to effectively process unstructured cyber threat intelligence (CTI) and the unique architectures of automotive systems, hindering vehicle-specific attack graph generation. This work proposes GARAGE, a novel framework that integrates retrieval-augmented generation (RAG) with domain-specific automotive security knowledge. Built upon 12,786 CVE entries and 140 incident reports, GARAGE constructs a knowledge base compliant with STIX 2.1 and Auto-ISAC ATM standards, enabling fine-grained kill-chain analysis for tactical-level attack scenario modeling. The approach supports generalization to unseen vehicle architectures and demonstrates accurate knowledge transfer across 320 leave-one-out experiments. Furthermore, it delineates the capability boundaries of large language models in threat analysis and provides cost-performance deployment strategies, thereby enhancing human-machine collaborative TARA processes.

0 citationsRead paper

Agile perceptive multi-skill locomotion for quadrupedal robots in the wild

Jul 15, 2026

This work addresses the challenge of achieving high-speed, multi-skilled, perception-driven agile locomotion and smooth gait transitions for quadrupedal robots in complex natural environments. The authors propose the APT-RL framework, which integrates an Action Pre-trained Transformer with reinforcement learning. By leveraging a large-scale 2D motion dataset generated through trajectory optimization, the framework pre-trains a transferable, high-quality motion prior. This prior is then combined with onboard perception and a simplified dynamics model to enable efficient policy learning and deployment on 3D rough terrain. Using only a single policy, the method robustly navigates diverse obstacles—including stairs, steps, and gaps—and achieves a peak speed of 6 m/s, significantly enhancing the robot’s agility and autonomy in both indoor and outdoor complex settings.

0 citationsRead paper

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

Jun 11, 2026

This work addresses the challenge of sustained reasoning about localization, environmental dynamics, and task progress in long-horizon mobile manipulation, where image observations alone are insufficient. The authors propose an online-updatable neural point map that jointly models the environment and robot embodiment as neural points within a shared latent space. By integrating object-level rigid-body tracking with forward kinematics, the method achieves an efficient spatiotemporal representation. The map is dynamically updated using first-person visual observations and proprioceptive states, providing multiscale, multi-view contextual information to vision–language–action policies. Evaluated on the BEHAVIOR-1K benchmark, the approach yields more direct trajectories, faster subgoal completion, and greater robustness to scene changes compared to image-only baselines, and demonstrates the ability to recover from failures such as object drops.

0 citationsRead paper
Recent publications

Latest Papers

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction

Jul 22, 2026

Existing diffusion-based approaches for robotic sequential prediction rely on unidirectional denoising, which struggles to maintain global consistency in long-horizon tasks. This work proposes a reversible denoising mechanism that incorporates a structured re-noising strategy, selectively re-adding noise to temporally stable local regions during the diffusion process. This enables iterative refinement by integrating cross-temporal contextual information, facilitating mutual correction between early and late segments of the predicted sequence. Implemented within a unified video-action diffusion framework, the method combines context-aware denoising with selective re-noising, achieving up to a 56.5% improvement in average success rate on the OGBench and LIBERO-10 benchmarks. Moreover, it demonstrates enhanced robustness and stronger action-video consistency in out-of-distribution scenarios.

0 citationsRead paper

GARAGE: Characterizing the Automation Boundary in LLM-based Attack Graph Generation

Jul 20, 2026

Existing automated tools struggle to effectively process unstructured cyber threat intelligence (CTI) and the unique architectures of automotive systems, hindering vehicle-specific attack graph generation. This work proposes GARAGE, a novel framework that integrates retrieval-augmented generation (RAG) with domain-specific automotive security knowledge. Built upon 12,786 CVE entries and 140 incident reports, GARAGE constructs a knowledge base compliant with STIX 2.1 and Auto-ISAC ATM standards, enabling fine-grained kill-chain analysis for tactical-level attack scenario modeling. The approach supports generalization to unseen vehicle architectures and demonstrates accurate knowledge transfer across 320 leave-one-out experiments. Furthermore, it delineates the capability boundaries of large language models in threat analysis and provides cost-performance deployment strategies, thereby enhancing human-machine collaborative TARA processes.

0 citationsRead paper

Agile perceptive multi-skill locomotion for quadrupedal robots in the wild

Jul 15, 2026

This work addresses the challenge of achieving high-speed, multi-skilled, perception-driven agile locomotion and smooth gait transitions for quadrupedal robots in complex natural environments. The authors propose the APT-RL framework, which integrates an Action Pre-trained Transformer with reinforcement learning. By leveraging a large-scale 2D motion dataset generated through trajectory optimization, the framework pre-trains a transferable, high-quality motion prior. This prior is then combined with onboard perception and a simplified dynamics model to enable efficient policy learning and deployment on 3D rough terrain. Using only a single policy, the method robustly navigates diverse obstacles—including stairs, steps, and gaps—and achieves a peak speed of 6 m/s, significantly enhancing the robot’s agility and autonomy in both indoor and outdoor complex settings.

0 citationsRead paper

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

Jun 11, 2026

This work addresses the challenge of sustained reasoning about localization, environmental dynamics, and task progress in long-horizon mobile manipulation, where image observations alone are insufficient. The authors propose an online-updatable neural point map that jointly models the environment and robot embodiment as neural points within a shared latent space. By integrating object-level rigid-body tracking with forward kinematics, the method achieves an efficient spatiotemporal representation. The map is dynamically updated using first-person visual observations and proprioceptive states, providing multiscale, multi-view contextual information to vision–language–action policies. Evaluated on the BEHAVIOR-1K benchmark, the approach yields more direct trajectories, faster subgoal completion, and greater robustness to scene changes compared to image-only baselines, and demonstrates the ability to recover from failures such as object drops.

0 citationsRead paper