Institution profile

Guangdong Laboratory of Artificial Intelligence and Digital Economy

Academic institutionasia · cn
Official website
Research library106linked papers
Opportunities0open roles
Selected work

Representative Papers

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Aug 13, 2026

Although JEPA-based world models perform prediction in latent space, they remain susceptible to visual perturbations, leading to distorted state representations and biased action-conditioned predictions. This work proposes Action-Conditioned Prediction Consistency (ACPC), a diagnostic framework that, for the first time, operationalizes the twin-simulation concept into a computable metric by quantifying the divergence between multi-step forward rollouts from clean and perturbed observations under identical action sequences. To enable robustness evaluation across tasks and architectures, we introduce two complementary metrics: Invariance Radius (IR) and Separation Rate (SR). Empirical results demonstrate that ACPC effectively predicts perturbation-induced errors in prediction and planning costs, and the IR–SR pair exhibits strong generalization and diagnostic capability across models such as LeWM and PLDM.

0 citationsRead paper

Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution

Aug 11, 2026

This work addresses the limitations of existing approaches in modeling the evolution of research problems and solutions within academic literature, particularly their neglect of interactions across distinct research trajectories. To overcome this, the authors propose the Tree-of-Ideas framework, which reconstructs branching scholarly trajectories via EvoTrace and introduces EvoAgent—the first cross-trajectory evolutionary reasoning mechanism—to identify common challenges and synthesize complementary solutions, thereby generating well-grounded novel ideas. Integrating citation network analysis, trajectory reconstruction algorithms, and large language model–based reasoning, the method transcends the constraints of traditional isolated citation-chain modeling. Automatic evaluation across six AI research topics demonstrates that the generated ideas achieve near-human performance in novelty (6.36), grounding (7.00), and overall score (6.27), closely matching human-level results (6.29).

0 citationsRead paper

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking

Aug 04, 2026

This work addresses the vulnerability of vision-language-action (VLA) robotic systems to physical adversarial patch attacks, which hijack policy-critical action-visual attention and lead to task failure. The study is the first to identify and formally name this “attention hijacking” mechanism. To mitigate it, the authors propose SARF (Structure-Aware Robust Fine-tuning), a zero-inference-overhead method that fine-tunes only the visual encoder through feature anchoring, critical attention correction, and language-guided geometric consistency constraints. Evaluated on the LIBERO benchmark, SARF reduces the attack-induced failure rate of OpenVLA from 100% to 28.6%. On a real PiPER robotic arm, it improves task success under attack from 23.0% to 65.0% without degrading performance on clean inputs, demonstrating strong cross-task and cross-architecture transferability.

0 citationsRead paper

Fused Bayesian Flow Networks for Dual-Target Molecular Design

Aug 02, 2026

This work addresses the challenge of generating dual-target molecules in polypharmacology by proposing a distribution-fusion-based 3D molecular generation framework. The approach formulates dual-target binding as a distribution fusion problem within a unified continuous space, dynamically integrating information from both targets via a product-of-experts mechanism and employing a pretrained target-aware Bayesian Flow Network (BFN) as a shared backbone. To mitigate the scarcity of structural data for dual targets, the method introduces chemically aware prior alignment and a prior-free pocket alignment strategy. Experimental results demonstrate that the generated molecules exhibit high binding affinity for both targets while maintaining favorable physicochemical properties, thereby validating the efficacy and novelty of the proposed framework.

0 citationsRead paper

VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks

Aug 02, 2026

This work addresses the vulnerability of Vision-Language-Action (VLA) robotic systems in wireless sensor networks to physical adversarial attacks, with a particular focus on motion-guided visual attention hijacking as a critical threat. The authors propose VLAGuard, the first framework to systematically evaluate and effectively mitigate this vulnerability. It introduces Visual-motor Attention-guided Semantic Attack (VASA), a printable patch-based stress test, and Attention-Protected Fine-Tuning (APFT), a defense method that stabilizes spatiotemporal attention and enforces geometric consistency without incurring any inference overhead. Experimental results demonstrate that VLAGuard reduces the failure rate of OpenVLA from 100.0% to 25.9% in the LIBERO simulation benchmark and improves task success under severe attacks from 23.0% to 67.4% across 2,000 real-world trials.

0 citationsRead paper
Recent publications

Latest Papers

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Aug 13, 2026

Although JEPA-based world models perform prediction in latent space, they remain susceptible to visual perturbations, leading to distorted state representations and biased action-conditioned predictions. This work proposes Action-Conditioned Prediction Consistency (ACPC), a diagnostic framework that, for the first time, operationalizes the twin-simulation concept into a computable metric by quantifying the divergence between multi-step forward rollouts from clean and perturbed observations under identical action sequences. To enable robustness evaluation across tasks and architectures, we introduce two complementary metrics: Invariance Radius (IR) and Separation Rate (SR). Empirical results demonstrate that ACPC effectively predicts perturbation-induced errors in prediction and planning costs, and the IR–SR pair exhibits strong generalization and diagnostic capability across models such as LeWM and PLDM.

0 citationsRead paper

Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution

Aug 11, 2026

This work addresses the limitations of existing approaches in modeling the evolution of research problems and solutions within academic literature, particularly their neglect of interactions across distinct research trajectories. To overcome this, the authors propose the Tree-of-Ideas framework, which reconstructs branching scholarly trajectories via EvoTrace and introduces EvoAgent—the first cross-trajectory evolutionary reasoning mechanism—to identify common challenges and synthesize complementary solutions, thereby generating well-grounded novel ideas. Integrating citation network analysis, trajectory reconstruction algorithms, and large language model–based reasoning, the method transcends the constraints of traditional isolated citation-chain modeling. Automatic evaluation across six AI research topics demonstrates that the generated ideas achieve near-human performance in novelty (6.36), grounding (7.00), and overall score (6.27), closely matching human-level results (6.29).

0 citationsRead paper

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking

Aug 04, 2026

This work addresses the vulnerability of vision-language-action (VLA) robotic systems to physical adversarial patch attacks, which hijack policy-critical action-visual attention and lead to task failure. The study is the first to identify and formally name this “attention hijacking” mechanism. To mitigate it, the authors propose SARF (Structure-Aware Robust Fine-tuning), a zero-inference-overhead method that fine-tunes only the visual encoder through feature anchoring, critical attention correction, and language-guided geometric consistency constraints. Evaluated on the LIBERO benchmark, SARF reduces the attack-induced failure rate of OpenVLA from 100% to 28.6%. On a real PiPER robotic arm, it improves task success under attack from 23.0% to 65.0% without degrading performance on clean inputs, demonstrating strong cross-task and cross-architecture transferability.

0 citationsRead paper

Fused Bayesian Flow Networks for Dual-Target Molecular Design

Aug 02, 2026

This work addresses the challenge of generating dual-target molecules in polypharmacology by proposing a distribution-fusion-based 3D molecular generation framework. The approach formulates dual-target binding as a distribution fusion problem within a unified continuous space, dynamically integrating information from both targets via a product-of-experts mechanism and employing a pretrained target-aware Bayesian Flow Network (BFN) as a shared backbone. To mitigate the scarcity of structural data for dual targets, the method introduces chemically aware prior alignment and a prior-free pocket alignment strategy. Experimental results demonstrate that the generated molecules exhibit high binding affinity for both targets while maintaining favorable physicochemical properties, thereby validating the efficacy and novelty of the proposed framework.

0 citationsRead paper

VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks

Aug 02, 2026

This work addresses the vulnerability of Vision-Language-Action (VLA) robotic systems in wireless sensor networks to physical adversarial attacks, with a particular focus on motion-guided visual attention hijacking as a critical threat. The authors propose VLAGuard, the first framework to systematically evaluate and effectively mitigate this vulnerability. It introduces Visual-motor Attention-guided Semantic Attack (VASA), a printable patch-based stress test, and Attention-Protected Fine-Tuning (APFT), a defense method that stabilizes spatiotemporal attention and enforces geometric consistency without incurring any inference overhead. Experimental results demonstrate that VLAGuard reduces the failure rate of OpenVLA from 100.0% to 25.9% in the LIBERO simulation benchmark and improves task success under severe attacks from 23.0% to 67.4% across 2,000 real-world trials.

0 citationsRead paper