Belief-Adaptive Online Autonomy for Quadrotor UAV Navigation under GNSS Degradation in Urban Environments
本文提出一种信念自适应在线自主框架,通过增强扩展卡尔曼滤波器来解决城市环境中GNSS信号退化问题,提高无人机导航的可靠性。
本文提出一种信念自适应在线自主框架,通过增强扩展卡尔曼滤波器来解决城市环境中GNSS信号退化问题,提高无人机导航的可靠性。
This work addresses the vulnerability of cross-camera face recognition systems to adversarial evasion and impersonation attacks by proposing a conditional encoder-decoder framework for generating adversarial patches. By fusing multi-scale features from both source and target images, the method simultaneously achieves efficient evasion and impersonation in a single forward pass, while leveraging a pre-trained latent diffusion model to enhance the visual realism of the patches for physical-world deployment. The approach innovatively incorporates a push-pull dual-objective optimization mechanism and employs activation map clustering to uncover the critical facial features exploited by the attack. Experimental results demonstrate that the proposed method reduces mean average precision (mAP) to 0.4% under both white-box and black-box settings, exhibits strong cross-model generalization, and achieves a 27% impersonation success rate on CelebA-HQ, significantly outperforming existing approaches.
This work addresses the challenges of coordinated management between AI-native Radio Access Networks (AI-RAN) and edge AI in the 6G era, particularly the lack of human-in-the-loop interaction mechanisms and the scarcity of on-site domain experts in enterprise settings. To this end, it proposes the first turn-based conversational agent framework for hierarchical collaborative AI-RAN management. The framework integrates a retrieval-augmented generation (RAG)-enhanced large language model within a three-tier architecture—comprising a user interface, an AI-RAN intelligent interface layer, and a knowledge layer—to enable intent understanding and dynamic decision-making across design planning, tool operation, and performance tuning. Experimental results demonstrate an average system response time of 13 seconds, with task accuracy rates of 78%, 89%, and 67% in service design, tool operation, and performance tuning, respectively, significantly reducing operational costs for small enterprises.
This study addresses the limitations of current autonomous driving perception systems—particularly their constrained cost-efficiency, robustness, and performance under adverse environmental conditions—by introducing the Calyo Pulse solid-state 3D ultrasonic sensor into the autonomous driving domain for the first time. The authors propose a semantic segmentation framework based on a 3D U-Net architecture, trained on voxelized ultrasonic data and enhanced with a weighted loss function to optimize segmentation accuracy. Experimental results on real-world ultrasonic data demonstrate robust 3D semantic segmentation performance, validating the potential of 3D ultrasound as a complementary sensing modality to LiDAR and cameras. This work thus offers a novel pathway toward enhancing perception robustness in challenging driving conditions.
This work addresses the absence of turn-level observability in existing autonomous information-gathering dialogue systems, which hinders real-time monitoring of information acquisition efficiency and detection of unproductive queries. The authors propose a Dialogue Telemetry (DT) framework that, after each interaction turn, generates two model-agnostic signals: a Progress Estimator (PE) quantifying remaining information potential and a Stagnation Index (SI) identifying repetitive, low-yield questioning. DT introduces, for the first time, an interpretable, turn-level stagnation detection mechanism that requires no causal diagnosis, integrating information-theoretic measures (in bits), semantic similarity, and marginal utility analysis to enable real-time quantification and intervention in dialogue efficiency. In simulated search-and-rescue scenarios, DT effectively discriminates between efficient and stagnant dialogues, and when incorporated into reinforcement learning policies, significantly enhances performance in settings with operational costs.
本文提出一种信念自适应在线自主框架,通过增强扩展卡尔曼滤波器来解决城市环境中GNSS信号退化问题,提高无人机导航的可靠性。
This work addresses the vulnerability of cross-camera face recognition systems to adversarial evasion and impersonation attacks by proposing a conditional encoder-decoder framework for generating adversarial patches. By fusing multi-scale features from both source and target images, the method simultaneously achieves efficient evasion and impersonation in a single forward pass, while leveraging a pre-trained latent diffusion model to enhance the visual realism of the patches for physical-world deployment. The approach innovatively incorporates a push-pull dual-objective optimization mechanism and employs activation map clustering to uncover the critical facial features exploited by the attack. Experimental results demonstrate that the proposed method reduces mean average precision (mAP) to 0.4% under both white-box and black-box settings, exhibits strong cross-model generalization, and achieves a 27% impersonation success rate on CelebA-HQ, significantly outperforming existing approaches.
This work addresses the challenges of coordinated management between AI-native Radio Access Networks (AI-RAN) and edge AI in the 6G era, particularly the lack of human-in-the-loop interaction mechanisms and the scarcity of on-site domain experts in enterprise settings. To this end, it proposes the first turn-based conversational agent framework for hierarchical collaborative AI-RAN management. The framework integrates a retrieval-augmented generation (RAG)-enhanced large language model within a three-tier architecture—comprising a user interface, an AI-RAN intelligent interface layer, and a knowledge layer—to enable intent understanding and dynamic decision-making across design planning, tool operation, and performance tuning. Experimental results demonstrate an average system response time of 13 seconds, with task accuracy rates of 78%, 89%, and 67% in service design, tool operation, and performance tuning, respectively, significantly reducing operational costs for small enterprises.
This study addresses the limitations of current autonomous driving perception systems—particularly their constrained cost-efficiency, robustness, and performance under adverse environmental conditions—by introducing the Calyo Pulse solid-state 3D ultrasonic sensor into the autonomous driving domain for the first time. The authors propose a semantic segmentation framework based on a 3D U-Net architecture, trained on voxelized ultrasonic data and enhanced with a weighted loss function to optimize segmentation accuracy. Experimental results on real-world ultrasonic data demonstrate robust 3D semantic segmentation performance, validating the potential of 3D ultrasound as a complementary sensing modality to LiDAR and cameras. This work thus offers a novel pathway toward enhancing perception robustness in challenging driving conditions.
This work addresses the absence of turn-level observability in existing autonomous information-gathering dialogue systems, which hinders real-time monitoring of information acquisition efficiency and detection of unproductive queries. The authors propose a Dialogue Telemetry (DT) framework that, after each interaction turn, generates two model-agnostic signals: a Progress Estimator (PE) quantifying remaining information potential and a Stagnation Index (SI) identifying repetitive, low-yield questioning. DT introduces, for the first time, an interpretable, turn-level stagnation detection mechanism that requires no causal diagnosis, integrating information-theoretic measures (in bits), semantic similarity, and marginal utility analysis to enable real-time quantification and intervention in dialogue efficiency. In simulated search-and-rescue scenarios, DT effectively discriminates between efficient and stagnant dialogues, and when incorporated into reinforcement learning policies, significantly enhances performance in settings with operational costs.