Institution profile

KDDI Research

Industry researchasia · jp
Official website
Research library25linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs

Jan 17, 2026

This work addresses the challenges of spurious negative samples and missing cross-modal semantic associations in audio-visual embedding learning caused by sparse annotations. To mitigate these issues, we propose a novel learning framework that leverages soft-label prediction and an implicit interaction graph. Our approach employs a teacher–student architecture to generate reliable soft supervision signals and utilizes the GRaSP algorithm to construct a directed inter-class dependency graph. By incorporating graph-guided regularization and semantic alignment losses, the model effectively captures latent semantic dependencies among unannotated co-occurring events. Experiments on the AVE and VEGAS benchmarks demonstrate that the proposed method significantly improves mean average precision (mAP), enhancing both semantic consistency and robustness in cross-modal embeddings.

1 citationsRead paper

Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic Labeling

Aug 15, 2026

This study addresses the degradation in recognition performance and annotation budget allocation challenges caused by concept drift in IoT traffic. We propose an adaptive strategy that decouples annotation volume from sample selection decisions. Specifically, an interpretable consistency-based drift detector is designed to dynamically adjust labeling budgets, while uniform sampling is employed to mitigate retraining bias. Extensive validation on 3.8 million real-world records spanning two years demonstrates that this approach effectively sustains classifier performance under limited budgets. The proposed method significantly outperforms traditional detector-guided selection strategies and achieves performance comparable to confidence-based adaptive methods. Consequently, this work provides an efficient and robust active learning solution for IoT device identification in dynamic environments.

0 citationsRead paper

MuST-VAD: Mutual Structured Learning for Video Anomaly Detection

Aug 07, 2026

This work addresses the performance limitations imposed by fixed features in weakly supervised video anomaly detection by proposing a bidirectional mutual learning framework. For the first time, it extends unidirectional feature transfer to a reciprocal knowledge exchange between large vision-language models and anomaly detectors. The framework enables efficient co-training through alternating updates, critical clip selection, and a confidence-weighted reliable supervision mechanism. Integrating weakly supervised learning, bidirectional knowledge distillation, and an annotation-guided question-answering strategy, the method achieves an AUROC of 88.63% and an average precision of 42.46% on the UCF-Crime dataset, outperforming the current state-of-the-art approach by 4.13 percentage points.

0 citationsRead paper

Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

Aug 03, 2026

This study addresses the limitations of existing post-disaster building damage assessment methods, which typically rely on large amounts of labeled data, exhibit poor cross-regional generalization, and offer limited task adaptability. To overcome these challenges, the authors propose a decoupled hybrid framework that first leverages high-precision computer vision models—such as Grounding DINO—to accurately localize buildings, followed by a large vision-language model (LVLM) for damage classification and contextual reasoning. This approach uniquely integrates the precise detection capabilities of computer vision with the semantic reasoning strengths of LVLMs, achieving significantly improved assessment performance with only minimal annotated data. Evaluated on real-world disaster datasets RescueNet and FloodNet, the method outperforms single-model baselines by up to 2.1 R² points and accurately quantifies the counts of undamaged, partially damaged, and completely destroyed buildings, demonstrating strong generalization and practical utility.

0 citationsRead paper
Recent publications

Latest Papers

Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic Labeling

Aug 15, 2026

This study addresses the degradation in recognition performance and annotation budget allocation challenges caused by concept drift in IoT traffic. We propose an adaptive strategy that decouples annotation volume from sample selection decisions. Specifically, an interpretable consistency-based drift detector is designed to dynamically adjust labeling budgets, while uniform sampling is employed to mitigate retraining bias. Extensive validation on 3.8 million real-world records spanning two years demonstrates that this approach effectively sustains classifier performance under limited budgets. The proposed method significantly outperforms traditional detector-guided selection strategies and achieves performance comparable to confidence-based adaptive methods. Consequently, this work provides an efficient and robust active learning solution for IoT device identification in dynamic environments.

0 citationsRead paper

MuST-VAD: Mutual Structured Learning for Video Anomaly Detection

Aug 07, 2026

This work addresses the performance limitations imposed by fixed features in weakly supervised video anomaly detection by proposing a bidirectional mutual learning framework. For the first time, it extends unidirectional feature transfer to a reciprocal knowledge exchange between large vision-language models and anomaly detectors. The framework enables efficient co-training through alternating updates, critical clip selection, and a confidence-weighted reliable supervision mechanism. Integrating weakly supervised learning, bidirectional knowledge distillation, and an annotation-guided question-answering strategy, the method achieves an AUROC of 88.63% and an average precision of 42.46% on the UCF-Crime dataset, outperforming the current state-of-the-art approach by 4.13 percentage points.

0 citationsRead paper

Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

Aug 03, 2026

This study addresses the limitations of existing post-disaster building damage assessment methods, which typically rely on large amounts of labeled data, exhibit poor cross-regional generalization, and offer limited task adaptability. To overcome these challenges, the authors propose a decoupled hybrid framework that first leverages high-precision computer vision models—such as Grounding DINO—to accurately localize buildings, followed by a large vision-language model (LVLM) for damage classification and contextual reasoning. This approach uniquely integrates the precise detection capabilities of computer vision with the semantic reasoning strengths of LVLMs, achieving significantly improved assessment performance with only minimal annotated data. Evaluated on real-world disaster datasets RescueNet and FloodNet, the method outperforms single-model baselines by up to 2.1 R² points and accurately quantifies the counts of undamaged, partially damaged, and completely destroyed buildings, demonstrating strong generalization and practical utility.

0 citationsRead paper

Motion Estimation Techniques for Volumetric Video Attribute Compression

Jul 03, 2026

This work addresses the challenge of inefficient temporal redundancy removal in dynamic point cloud attribute compression and the limited support of existing motion estimation methods for attribute coding. To overcome these limitations, the authors propose a geometry-guided inter-frame coding framework that, for the first time, integrates geometric information with graph signal processing. They design a graph-structured motion estimation algorithm and introduce a sub-voxel-level motion compensation mechanism that operates without interpolation. Implemented within the G-PCC framework, the proposed method significantly improves compression efficiency. Experimental results on MPEG standard datasets demonstrate substantial bitrate savings under lossy geometry conditions, achieving average reductions of 55.3%, 42.3%, and 16.5% compared to G-PCC, GeS-TM, and V-PCC, respectively.

0 citationsRead paper