HarvestPoint-ACT: Explicit Target Selection and Harvest-Point Conditioning for Robotic Fruit Harvesting under Occlusion
本文提出HarvestPoint-ACT方法,通过显式目标选择和收获点条件设置解决机器人在遮挡情况下采摘水果的问题,提高了成功率。
本文提出HarvestPoint-ACT方法,通过显式目标选择和收获点条件设置解决机器人在遮挡情况下采摘水果的问题,提高了成功率。
为解决机器人装配布局规划中的碰撞问题,提出了一种基于反向搜索的方法BLS,通过逆序分配初始部件姿态并进行多种检查,有效减少了评估次数和搜索时间。
Existing multiple instance learning (MIL) approaches treat whole-slide images as unstructured collections of image patches, thereby neglecting the morphological semantics and spatial geometric relationships inherent in tissue architecture. This limitation renders them susceptible to background noise and misaligned with clinical diagnostic reasoning. To address this, this work proposes the HPDP framework, which introduces a Morphology-Anchored Prototype System (MAPS) to explicitly model histological structural semantics, incorporates sinusoidal positional encoding (SPE) to capture spatial geometry, and designs a Hierarchical Cross-Modal Alignment (HCMA) module that leverages pathology descriptions generated by large language models to achieve image–text semantic alignment. Evaluated across seven cancer cohorts, the proposed method significantly improves diagnostic accuracy, robustness, and interpretability, outperforming current state-of-the-art approaches.
To address the challenges of effectively fusing local and global features and robustly identifying high-quality correspondences in point cloud registration, this paper proposes a Gestalt-inspired parallel interaction network. Our method introduces three key innovations: (1) a Gestalt Feature Attention module that models structural completeness at the perceptual level; (2) a dual-path, multi-granularity parallel interaction architecture that jointly leverages self-attention and cross-attention, augmented by an orthogonal geometric consistency constraint to strengthen global structural representation; and (3) an orthogonal feature fusion strategy to enhance complementarity across granularities. Extensive experiments on standard benchmarks—including ModelNet40 and 3DMatch—demonstrate significant improvements over state-of-the-art methods in both matching accuracy and noise robustness. The source code is publicly available.
本文提出HarvestPoint-ACT方法,通过显式目标选择和收获点条件设置解决机器人在遮挡情况下采摘水果的问题,提高了成功率。
为解决机器人装配布局规划中的碰撞问题,提出了一种基于反向搜索的方法BLS,通过逆序分配初始部件姿态并进行多种检查,有效减少了评估次数和搜索时间。
Existing multiple instance learning (MIL) approaches treat whole-slide images as unstructured collections of image patches, thereby neglecting the morphological semantics and spatial geometric relationships inherent in tissue architecture. This limitation renders them susceptible to background noise and misaligned with clinical diagnostic reasoning. To address this, this work proposes the HPDP framework, which introduces a Morphology-Anchored Prototype System (MAPS) to explicitly model histological structural semantics, incorporates sinusoidal positional encoding (SPE) to capture spatial geometry, and designs a Hierarchical Cross-Modal Alignment (HCMA) module that leverages pathology descriptions generated by large language models to achieve image–text semantic alignment. Evaluated across seven cancer cohorts, the proposed method significantly improves diagnostic accuracy, robustness, and interpretability, outperforming current state-of-the-art approaches.
To address the challenges of effectively fusing local and global features and robustly identifying high-quality correspondences in point cloud registration, this paper proposes a Gestalt-inspired parallel interaction network. Our method introduces three key innovations: (1) a Gestalt Feature Attention module that models structural completeness at the perceptual level; (2) a dual-path, multi-granularity parallel interaction architecture that jointly leverages self-attention and cross-attention, augmented by an orthogonal geometric consistency constraint to strengthen global structural representation; and (3) an orthogonal feature fusion strategy to enhance complementarity across granularities. Extensive experiments on standard benchmarks—including ModelNet40 and 3DMatch—demonstrate significant improvements over state-of-the-art methods in both matching accuracy and noise robustness. The source code is publicly available.