Institution profile

Panasonic Corporation

Industry researchasia · jp
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

T-Rex: Tactile-Reactive Dexterous Manipulation

Jun 15, 2026

This work addresses the limitation of existing vision–language–action (VLA) models in robotic dexterous manipulation, which typically neglect tactile feedback or rely solely on static tactile representations, thereby failing to support dynamic tactile responses. To overcome this, the study introduces dynamic tactile perception into the VLA framework for the first time, proposing three core innovations: a large-scale, motion-primitive-based dataset enriched with high-frequency tactile signals, a temporal tactile VQ-VAE encoder that captures time-varying tactile features, and a variable-rate Mixture-of-Transformers architecture. The proposed method effectively leverages rich tactile dynamics, achieving an average success rate improvement of over 30% compared to the strongest baseline across twelve fine-grained force-control and deformable object manipulation tasks.

0 citationsRead paper

Contrastive Action-Image Pre-training for Visuomotor Control

Jun 15, 2026

Existing robotic vision encoders struggle to achieve effective pretraining due to the scarcity of large-scale, action-annotated datasets and insufficient alignment with the visuomotor signals required by downstream control tasks. This work proposes a novel paradigm that leverages 3D hand keypoints extracted from human egocentric videos as proxies for robotic end-effector actions, enabling a unified contrastive learning objective tailored for visuomotor control. By combining large-scale self-supervised pretraining with fine-tuning on limited real robot data, the approach significantly improves success rates—by over 30%—on dexterous manipulation tasks such as folding and pouring when deployed on the Dexmate Vega and Sharpa Wave robotic hands. The method consistently outperforms state-of-the-art vision encoders like DINOv2 and SigLIP in downstream control performance.

0 citationsRead paper

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

May 27, 2026

This work addresses the coarse credit assignment inherent in existing sample-level reward-based reinforcement learning methods for discrete policy optimization, which struggle to discern fine-grained contributions of individual tokens. To overcome this limitation, the paper proposes Guided Contrastive Policy Optimization (GCPO), a novel algorithm that introduces, for the first time, a prediction contrast mechanism guided by positive and negative prompts to enable token-level advantage estimation and credit assignment. By integrating contrastive learning with discrete policy gradients, GCPO generates more precise learning signals. Experimental results demonstrate that GCPO significantly outperforms baseline methods such as GRPO and DAPO on both text-to-image generation and chain-of-thought reasoning tasks, confirming its effectiveness and broad applicability.

0 citationsRead paper

Portable Active Learning for Object Detection

May 11, 2026

This work addresses the high annotation cost and limited scalability in object detection by proposing a detector-agnostic active learning framework that maintains high detection accuracy while substantially reducing human labeling effort. The approach requires no modifications to model architecture or training procedures, relying solely on inference outputs to select informative samples through a joint consideration of image-level diversity, category distribution imbalance, and instance-level uncertainty. Specifically, a lightweight category-specific classifier generates entropy-based uncertainty scores, which are combined with global image entropy and image similarity metrics to identify high-value unlabeled instances. Extensive experiments on COCO, PASCAL VOC, and BDD100K demonstrate that the proposed method significantly outperforms existing active learning strategies, offering strong efficiency, broad compatibility across detectors, and practical potential for real-world deployment.

0 citationsRead paper

Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

May 08, 2026

Existing vision-language models often struggle with 3D spatial reasoning due to inefficient geometric representations or insufficient spatial consistency, making it challenging to balance efficiency and performance. This work proposes Proxy3D, a method that operates solely on video frames by employing semantic and geometric encoders to extract features and introducing semantic-aware clustering to generate compact yet spatially consistent 3D proxy representations. Through a multi-stage training strategy and a newly curated SpaceSpan dataset, Proxy3D effectively aligns with vision-language models. Notably, it is the first approach to integrate semantic clustering with 3D proxy representations, achieving competitive or state-of-the-art performance on 3D visual question answering, visual grounding, and general spatial reasoning tasks—using significantly shorter visual sequences and thereby surpassing the limitations of conventional 2D pipelines.

0 citationsRead paper
Recent publications

Latest Papers

T-Rex: Tactile-Reactive Dexterous Manipulation

Jun 15, 2026

This work addresses the limitation of existing vision–language–action (VLA) models in robotic dexterous manipulation, which typically neglect tactile feedback or rely solely on static tactile representations, thereby failing to support dynamic tactile responses. To overcome this, the study introduces dynamic tactile perception into the VLA framework for the first time, proposing three core innovations: a large-scale, motion-primitive-based dataset enriched with high-frequency tactile signals, a temporal tactile VQ-VAE encoder that captures time-varying tactile features, and a variable-rate Mixture-of-Transformers architecture. The proposed method effectively leverages rich tactile dynamics, achieving an average success rate improvement of over 30% compared to the strongest baseline across twelve fine-grained force-control and deformable object manipulation tasks.

0 citationsRead paper

Contrastive Action-Image Pre-training for Visuomotor Control

Jun 15, 2026

Existing robotic vision encoders struggle to achieve effective pretraining due to the scarcity of large-scale, action-annotated datasets and insufficient alignment with the visuomotor signals required by downstream control tasks. This work proposes a novel paradigm that leverages 3D hand keypoints extracted from human egocentric videos as proxies for robotic end-effector actions, enabling a unified contrastive learning objective tailored for visuomotor control. By combining large-scale self-supervised pretraining with fine-tuning on limited real robot data, the approach significantly improves success rates—by over 30%—on dexterous manipulation tasks such as folding and pouring when deployed on the Dexmate Vega and Sharpa Wave robotic hands. The method consistently outperforms state-of-the-art vision encoders like DINOv2 and SigLIP in downstream control performance.

0 citationsRead paper

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

May 27, 2026

This work addresses the coarse credit assignment inherent in existing sample-level reward-based reinforcement learning methods for discrete policy optimization, which struggle to discern fine-grained contributions of individual tokens. To overcome this limitation, the paper proposes Guided Contrastive Policy Optimization (GCPO), a novel algorithm that introduces, for the first time, a prediction contrast mechanism guided by positive and negative prompts to enable token-level advantage estimation and credit assignment. By integrating contrastive learning with discrete policy gradients, GCPO generates more precise learning signals. Experimental results demonstrate that GCPO significantly outperforms baseline methods such as GRPO and DAPO on both text-to-image generation and chain-of-thought reasoning tasks, confirming its effectiveness and broad applicability.

0 citationsRead paper

Portable Active Learning for Object Detection

May 11, 2026

This work addresses the high annotation cost and limited scalability in object detection by proposing a detector-agnostic active learning framework that maintains high detection accuracy while substantially reducing human labeling effort. The approach requires no modifications to model architecture or training procedures, relying solely on inference outputs to select informative samples through a joint consideration of image-level diversity, category distribution imbalance, and instance-level uncertainty. Specifically, a lightweight category-specific classifier generates entropy-based uncertainty scores, which are combined with global image entropy and image similarity metrics to identify high-value unlabeled instances. Extensive experiments on COCO, PASCAL VOC, and BDD100K demonstrate that the proposed method significantly outperforms existing active learning strategies, offering strong efficiency, broad compatibility across detectors, and practical potential for real-world deployment.

0 citationsRead paper

Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

May 08, 2026

Existing vision-language models often struggle with 3D spatial reasoning due to inefficient geometric representations or insufficient spatial consistency, making it challenging to balance efficiency and performance. This work proposes Proxy3D, a method that operates solely on video frames by employing semantic and geometric encoders to extract features and introducing semantic-aware clustering to generate compact yet spatially consistent 3D proxy representations. Through a multi-stage training strategy and a newly curated SpaceSpan dataset, Proxy3D effectively aligns with vision-language models. Notably, it is the first approach to integrate semantic clustering with 3D proxy representations, achieving competitive or state-of-the-art performance on 3D visual question answering, visual grounding, and general spatial reasoning tasks—using significantly shorter visual sequences and thereby surpassing the limitations of conventional 2D pipelines.

0 citationsRead paper