Institution profile

OM-RON SINIC X Corporation

Industry researchasia · jp
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Aug 11, 2026

This work addresses the common oversight in existing hand pose estimation methods—namely, the neglect of joint visibility—which hinders reliable assessment of estimation quality under occlusion. The study introduces joint visibility estimation as a standalone task and proposes a visibility detector built upon a large-scale pretrained hand pose model. Furthermore, it integrates a visibility-weighted multi-view triangulation strategy to refine 3D pose reconstruction. The proposed approach substantially improves visibility prediction accuracy and effectively reduces reprojection error in 3D hand pose annotation, thereby demonstrating the practical utility of explicit visibility estimation. To facilitate adoption and further research, the authors release a ready-to-use toolkit alongside their findings.

0 citationsRead paper

GAVEL: Grounded Caption Error Verification and Localization

Jun 25, 2026

Existing vision-language models often generate descriptions inconsistent with input images and lack effective mechanisms for verifying and explaining such mismatches. This work introduces the GAVEL task, which unifies image-text consistency verification, natural language explanation generation, and fine-grained visual evidence localization into a single end-to-end learnable framework, accompanied by the first high-quality, human-annotated dataset for this purpose. The authors propose a supervised baseline model that leverages vision-language alignment and multi-task joint optimization, significantly outperforming strong closed-source models in both localization accuracy and explanation quality. These results underscore the challenge posed by the GAVEL task and demonstrate the effectiveness of the proposed approach.

0 citationsRead paper

Stable In-hand Manipulation for a Lightweight Four-motor Prosthetic Hand

Jan 12, 2026

This study addresses the limited stability of lightweight four-motor prosthetic hands when grasping objects of varying sizes and weights, particularly heavy items. The authors propose a closed-loop control strategy based on motor current feedback, integrated with an optimized single-axis thumb mechanism, to enable real-time estimation of object width and dynamic adjustment of the index finger position—facilitating adaptive in-hand manipulation without prior knowledge of object dimensions. This work presents the first application of current feedback for in-hand manipulation in lightweight prosthetic hands, significantly enhancing stability during heavy-object handling. The approach achieves 100% success in manipulating lightweight objects 5–30 mm wide and maintains ≥80% success with a 289 g aluminum prism—doubling the performance of non-coordinated control—and successfully executes daily tasks such as unscrewing bottle caps and holding a pen.

0 citationsRead paper

A Flexible Funnel-Shaped Robotic Hand with an Integrated Single-Sheet Valve for Milligram-Scale Powder Handling

Dec 07, 2025

Addressing the challenges of low dispensing accuracy and poor adaptability in laboratory automation for milligram-scale powders—stemming from complex flow behavior and task diversity—this work proposes a flexible funnel-shaped robotic gripper integrating a conical-tip monolithic valve with a model predictive feedback control framework. A closed-loop system is established by coupling a physics-based powder flow model with an online parameter identification algorithm, and interfacing with an external analytical balance. This enables adaptive, high-precision gravimetric dispensing across diverse powder types. Experiments demonstrate that, for target masses ranging from 20 mg to 3 g, 80% of dispenses achieve ≤2 mg absolute error, with a maximum error of approximately 20 mg; both convergence speed and accuracy significantly surpass those of conventional PID control. The core innovation lies in the synergistic integration of a compliant, controllable discharge mechanism with a model-based, real-time adaptive control strategy.

0 citationsRead paper

SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters

Nov 23, 2025

Scientific poster layout analysis has long been hindered by the scarcity of annotated datasets and dedicated models. To address this, we propose the first structured analysis framework for scientific posters. We introduce SciPostLayoutTree, a large-scale dataset comprising 8,000 posters meticulously annotated with reading order and hierarchical parent–child relationships. We further design Layout Tree Decoder, a Transformer-based model that jointly encodes visual features, bounding box coordinates, and semantic class labels to capture both spatial and semantic dependencies; it employs beam search to optimize sequential decoding of tree-structured layouts. Experiments demonstrate that our method significantly outperforms existing baselines in predicting complex spatial relationships, establishing a robust benchmark for poster content understanding. All components—including the dataset, model, and code—are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Aug 11, 2026

This work addresses the common oversight in existing hand pose estimation methods—namely, the neglect of joint visibility—which hinders reliable assessment of estimation quality under occlusion. The study introduces joint visibility estimation as a standalone task and proposes a visibility detector built upon a large-scale pretrained hand pose model. Furthermore, it integrates a visibility-weighted multi-view triangulation strategy to refine 3D pose reconstruction. The proposed approach substantially improves visibility prediction accuracy and effectively reduces reprojection error in 3D hand pose annotation, thereby demonstrating the practical utility of explicit visibility estimation. To facilitate adoption and further research, the authors release a ready-to-use toolkit alongside their findings.

0 citationsRead paper

GAVEL: Grounded Caption Error Verification and Localization

Jun 25, 2026

Existing vision-language models often generate descriptions inconsistent with input images and lack effective mechanisms for verifying and explaining such mismatches. This work introduces the GAVEL task, which unifies image-text consistency verification, natural language explanation generation, and fine-grained visual evidence localization into a single end-to-end learnable framework, accompanied by the first high-quality, human-annotated dataset for this purpose. The authors propose a supervised baseline model that leverages vision-language alignment and multi-task joint optimization, significantly outperforming strong closed-source models in both localization accuracy and explanation quality. These results underscore the challenge posed by the GAVEL task and demonstrate the effectiveness of the proposed approach.

0 citationsRead paper

Stable In-hand Manipulation for a Lightweight Four-motor Prosthetic Hand

Jan 12, 2026

This study addresses the limited stability of lightweight four-motor prosthetic hands when grasping objects of varying sizes and weights, particularly heavy items. The authors propose a closed-loop control strategy based on motor current feedback, integrated with an optimized single-axis thumb mechanism, to enable real-time estimation of object width and dynamic adjustment of the index finger position—facilitating adaptive in-hand manipulation without prior knowledge of object dimensions. This work presents the first application of current feedback for in-hand manipulation in lightweight prosthetic hands, significantly enhancing stability during heavy-object handling. The approach achieves 100% success in manipulating lightweight objects 5–30 mm wide and maintains ≥80% success with a 289 g aluminum prism—doubling the performance of non-coordinated control—and successfully executes daily tasks such as unscrewing bottle caps and holding a pen.

0 citationsRead paper

A Flexible Funnel-Shaped Robotic Hand with an Integrated Single-Sheet Valve for Milligram-Scale Powder Handling

Dec 07, 2025

Addressing the challenges of low dispensing accuracy and poor adaptability in laboratory automation for milligram-scale powders—stemming from complex flow behavior and task diversity—this work proposes a flexible funnel-shaped robotic gripper integrating a conical-tip monolithic valve with a model predictive feedback control framework. A closed-loop system is established by coupling a physics-based powder flow model with an online parameter identification algorithm, and interfacing with an external analytical balance. This enables adaptive, high-precision gravimetric dispensing across diverse powder types. Experiments demonstrate that, for target masses ranging from 20 mg to 3 g, 80% of dispenses achieve ≤2 mg absolute error, with a maximum error of approximately 20 mg; both convergence speed and accuracy significantly surpass those of conventional PID control. The core innovation lies in the synergistic integration of a compliant, controllable discharge mechanism with a model-based, real-time adaptive control strategy.

0 citationsRead paper

SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters

Nov 23, 2025

Scientific poster layout analysis has long been hindered by the scarcity of annotated datasets and dedicated models. To address this, we propose the first structured analysis framework for scientific posters. We introduce SciPostLayoutTree, a large-scale dataset comprising 8,000 posters meticulously annotated with reading order and hierarchical parent–child relationships. We further design Layout Tree Decoder, a Transformer-based model that jointly encodes visual features, bounding box coordinates, and semantic class labels to capture both spatial and semantic dependencies; it employs beam search to optimize sequential decoding of tree-structured layouts. Experiments demonstrate that our method significantly outperforms existing baselines in predicting complex spatial relationships, establishing a robust benchmark for poster content understanding. All components—including the dataset, model, and code—are publicly released.

0 citationsRead paper