HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving
为解决远程传感任务中单体决策框架的不稳定性和错误传播问题,提出HiRS-Agent系统,采用分层多代理架构和强化学习策略优化任务执行。
为解决远程传感任务中单体决策框架的不稳定性和错误传播问题,提出HiRS-Agent系统,采用分层多代理架构和强化学习策略优化任务执行。
This study addresses the unknown applicability of traditional cartographic color principles to spatial reasoning in foundation models. By constructing a controlled benchmark and integrating multimodal evaluation, LoRA fine-tuning, and factorial experiments, this work systematically quantifies the impact of color variables on model reasoning for the first time. Results indicate that disordered color sequences and low contrast significantly impair performance, and notably, fine-tuning fails to eliminate this sensitivity. Highlighting the critical roles of sequential color ordering and contrast, this research proposes AI-friendly cartographic design guidelines. These findings provide empirical evidence and methodological guidance for optimizing map understanding capabilities in artificial intelligence systems, bridging the gap between classical cartography and modern vision-language models.
Existing point cloud scene generation methods rely on partial scans as conditioning inputs, leading to a mismatch between training and inference, poor handling of sparsity in distant regions and occluded areas, and limited flexibility in generating scenes without LiDAR observations. To address these limitations, this work proposes a unified generation framework that dispenses with partial scans by predicting density, height, and occupancy masks in bird’s-eye view (BEV) to construct structured point sources. Furthermore, it introduces a teacher–student approximate optimal transport mechanism that learns straighter transport paths for efficient single-step point generation. The approach supports both unconditional and multi-cue conditional generation, achieving state-of-the-art Jensen–Shannon divergence (JSD) and voxel IoU on SemanticKITTI, and the best Coverage score on KITTI-360 under unconditional generation.
This work addresses the challenge that existing hyperspectral image fusion methods often suffer from spatial structural degradation and spectral distortion due to their inability to effectively model geometric constraints. To overcome this limitation, the paper proposes a dual-domain manifold modeling framework that jointly introduces manifold structures in both spatial and spectral domains. A topology-aware Transformer is developed to simultaneously capture spatial topology and pixel-level manifold relationships. Furthermore, a frequency-domain decoupled fusion module is designed to separate high- and low-frequency components, thereby enhancing high-frequency geometric details and spectral reconstruction fidelity. Integrating discrete cosine transform, low-rank priors, and a spectral-driven spatial enhancement strategy, the proposed method achieves state-of-the-art performance on multiple benchmark datasets, demonstrating superior spatial fidelity and spectral accuracy compared to current approaches.
This work addresses the challenge of achieving finite-time accurate pose control for nonholonomic mobile robots under curvature constraints and actuator limitations. Existing vector field approaches typically guarantee only asymptotic convergence and rely on input saturation, which can compromise stability. To overcome these issues, this paper proposes a framework combining a finite-time curvature-constrained vector field (FT-C²VF) with a saturation-free, smooth control law. The approach introduces, for the first time, a vector field that is curvature-continuous, bounded, and monotonically decreasing with respect to the radial ratio, ensuring finite-time convergence. A nearly globally C¹-smooth, Jacobian-free, saturation-free controller is designed to enforce actuator constraints while achieving almost global finite-time stability. Simulations and real-world experiments with an Ackermann-steered vehicle demonstrate superior performance over representative vector field methods, confirming the method’s effectiveness and robustness.
为解决远程传感任务中单体决策框架的不稳定性和错误传播问题,提出HiRS-Agent系统,采用分层多代理架构和强化学习策略优化任务执行。
This study addresses the unknown applicability of traditional cartographic color principles to spatial reasoning in foundation models. By constructing a controlled benchmark and integrating multimodal evaluation, LoRA fine-tuning, and factorial experiments, this work systematically quantifies the impact of color variables on model reasoning for the first time. Results indicate that disordered color sequences and low contrast significantly impair performance, and notably, fine-tuning fails to eliminate this sensitivity. Highlighting the critical roles of sequential color ordering and contrast, this research proposes AI-friendly cartographic design guidelines. These findings provide empirical evidence and methodological guidance for optimizing map understanding capabilities in artificial intelligence systems, bridging the gap between classical cartography and modern vision-language models.
Existing point cloud scene generation methods rely on partial scans as conditioning inputs, leading to a mismatch between training and inference, poor handling of sparsity in distant regions and occluded areas, and limited flexibility in generating scenes without LiDAR observations. To address these limitations, this work proposes a unified generation framework that dispenses with partial scans by predicting density, height, and occupancy masks in bird’s-eye view (BEV) to construct structured point sources. Furthermore, it introduces a teacher–student approximate optimal transport mechanism that learns straighter transport paths for efficient single-step point generation. The approach supports both unconditional and multi-cue conditional generation, achieving state-of-the-art Jensen–Shannon divergence (JSD) and voxel IoU on SemanticKITTI, and the best Coverage score on KITTI-360 under unconditional generation.
This work addresses the challenge that existing hyperspectral image fusion methods often suffer from spatial structural degradation and spectral distortion due to their inability to effectively model geometric constraints. To overcome this limitation, the paper proposes a dual-domain manifold modeling framework that jointly introduces manifold structures in both spatial and spectral domains. A topology-aware Transformer is developed to simultaneously capture spatial topology and pixel-level manifold relationships. Furthermore, a frequency-domain decoupled fusion module is designed to separate high- and low-frequency components, thereby enhancing high-frequency geometric details and spectral reconstruction fidelity. Integrating discrete cosine transform, low-rank priors, and a spectral-driven spatial enhancement strategy, the proposed method achieves state-of-the-art performance on multiple benchmark datasets, demonstrating superior spatial fidelity and spectral accuracy compared to current approaches.
This work addresses the challenge of achieving finite-time accurate pose control for nonholonomic mobile robots under curvature constraints and actuator limitations. Existing vector field approaches typically guarantee only asymptotic convergence and rely on input saturation, which can compromise stability. To overcome these issues, this paper proposes a framework combining a finite-time curvature-constrained vector field (FT-C²VF) with a saturation-free, smooth control law. The approach introduces, for the first time, a vector field that is curvature-continuous, bounded, and monotonically decreasing with respect to the radial ratio, ensuring finite-time convergence. A nearly globally C¹-smooth, Jacobian-free, saturation-free controller is designed to enforce actuator constraints while achieving almost global finite-time stability. Simulations and real-world experiments with an Ackermann-steered vehicle demonstrate superior performance over representative vector field methods, confirming the method’s effectiveness and robustness.