Institution profile

Konica Minolta, Inc.

Industry researchasia · jp
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Intuitive Surgical SurgToolLoc Challenge Results: 2022-2023

May 11, 2023

To address the challenge of real-time, robust surgical instrument localization in minimally invasive robotic-assisted surgery (RAS) video streams, this work introduces SurgToolLoc—the first large-scale, multi-view, multi-scenario benchmark dataset with pixel-level mask annotations. We further propose a novel evaluation protocol emphasizing both cross-center generalizability and real-time inference (≥30 FPS). Methodologically, we integrate instance segmentation and keypoint detection with temporal modeling (ConvLSTM/Transformer), domain adaptation, and weakly supervised learning. Our best-performing model achieves 92.4% mAP@0.5 on the test set while maintaining an inference speed of 36 FPS—substantially outperforming conventional template matching and early CNN-based approaches. The solution has undergone rigorous preclinical validation across multiple surgical scenarios. By providing a reproducible, scalable, end-to-end framework for visual instrument localization in RAS, this work establishes a new standard for benchmarking and advancing vision-based surgical navigation systems.

17 citationsRead paper

Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge

Jul 22, 2025

Robust instrument recognition and localization in minimally invasive surgical endoscopic videos remains challenging under real-world conditions due to complex backgrounds, occlusions, and inter-center variability. Method: We introduce the first unified, multi-center cholecystectomy video dataset with synchronized annotations for surgical phase classification, instrument instance segmentation, and anatomical keypoint localization—enabling contextual and temporal modeling. We propose a temporal-aware multi-task learning framework that jointly optimizes all three tasks while explicitly incorporating surgical workflow priors. Evaluation strictly follows the BIAS guidelines to establish a high-quality benchmark for robot-assisted surgery. Contribution/Results: Our method significantly improves instrument localization accuracy under cluttered backgrounds and enhances cross-center generalization. It also advances intraoperative scene understanding by improving interpretability and clinical utility, setting a new standard for vision-based surgical intelligence.

0 citationsRead paper

SICNav-Diffusion: Safe and Interactive Crowd Navigation with Diffusion Trajectory Predictions

Mar 11, 2025

This work addresses safe interactive navigation for a single robot operating amid multiple pedestrians. We propose a unified trajectory prediction and navigation framework that synergistically integrates a denoising diffusion probabilistic model (DDPM) with bilevel model predictive control (MPC). The upper-level MPC optimizes the robot’s reference trajectory, while the lower-level MPC dynamically filters and refines this trajectory in real time by enforcing safety constraints derived from multimodal pedestrian trajectories generated by the DDPM—thereby enabling tight coupling between prediction and planning and end-to-end collision avoidance. To our knowledge, this is the first approach to embed a diffusion model within a bilevel MPC architecture. Evaluated on the ETH/UCY benchmark, it reduces trajectory prediction error by 12.7%. Extensive simulations and real-robot experiments demonstrate zero collisions in dense pedestrian crowds, a 23% improvement in passage efficiency, and sub-0.3-second system response latency.

0 citationsRead paper
Recent publications

Latest Papers

Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge

Jul 22, 2025

Robust instrument recognition and localization in minimally invasive surgical endoscopic videos remains challenging under real-world conditions due to complex backgrounds, occlusions, and inter-center variability. Method: We introduce the first unified, multi-center cholecystectomy video dataset with synchronized annotations for surgical phase classification, instrument instance segmentation, and anatomical keypoint localization—enabling contextual and temporal modeling. We propose a temporal-aware multi-task learning framework that jointly optimizes all three tasks while explicitly incorporating surgical workflow priors. Evaluation strictly follows the BIAS guidelines to establish a high-quality benchmark for robot-assisted surgery. Contribution/Results: Our method significantly improves instrument localization accuracy under cluttered backgrounds and enhances cross-center generalization. It also advances intraoperative scene understanding by improving interpretability and clinical utility, setting a new standard for vision-based surgical intelligence.

0 citationsRead paper

SICNav-Diffusion: Safe and Interactive Crowd Navigation with Diffusion Trajectory Predictions

Mar 11, 2025

This work addresses safe interactive navigation for a single robot operating amid multiple pedestrians. We propose a unified trajectory prediction and navigation framework that synergistically integrates a denoising diffusion probabilistic model (DDPM) with bilevel model predictive control (MPC). The upper-level MPC optimizes the robot’s reference trajectory, while the lower-level MPC dynamically filters and refines this trajectory in real time by enforcing safety constraints derived from multimodal pedestrian trajectories generated by the DDPM—thereby enabling tight coupling between prediction and planning and end-to-end collision avoidance. To our knowledge, this is the first approach to embed a diffusion model within a bilevel MPC architecture. Evaluated on the ETH/UCY benchmark, it reduces trajectory prediction error by 12.7%. Extensive simulations and real-robot experiments demonstrate zero collisions in dense pedestrian crowds, a 23% improvement in passage efficiency, and sub-0.3-second system response latency.

0 citationsRead paper

Intuitive Surgical SurgToolLoc Challenge Results: 2022-2023

May 11, 2023

To address the challenge of real-time, robust surgical instrument localization in minimally invasive robotic-assisted surgery (RAS) video streams, this work introduces SurgToolLoc—the first large-scale, multi-view, multi-scenario benchmark dataset with pixel-level mask annotations. We further propose a novel evaluation protocol emphasizing both cross-center generalizability and real-time inference (≥30 FPS). Methodologically, we integrate instance segmentation and keypoint detection with temporal modeling (ConvLSTM/Transformer), domain adaptation, and weakly supervised learning. Our best-performing model achieves 92.4% mAP@0.5 on the test set while maintaining an inference speed of 36 FPS—substantially outperforming conventional template matching and early CNN-based approaches. The solution has undergone rigorous preclinical validation across multiple surgical scenarios. By providing a reproducible, scalable, end-to-end framework for visual instrument localization in RAS, this work establishes a new standard for benchmarking and advancing vision-based surgical navigation systems.

17 citationsRead paper