Institution profile

LINE Corporation

Industry researchasia · jp
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio Representation

May 23, 2022IEEE International Conference on Acoustics, Speech, and Signal Processing

Conventional self-supervised audio representation learning relies solely on clip-level sampling, leading to insufficient frame-level modeling capability. Method: This paper proposes a multi-granularity contrastive learning framework that jointly leverages clip-level, frame-level, and task-guided sampling to construct multi-perspective contrastive losses, enabling collaborative optimization of general-purpose audio representations. Contribution/Results: To our knowledge, this is the first work to incorporate both frame-level and task-specific sampling into self-supervised pre-training, overcoming the limitations of single-granularity representation learning. Pre-trained on a subset of AudioSet and evaluated via frozen-feature transfer to downstream tasks, our method achieves 25%, 20%, and 3.6% absolute improvements in clip classification, sound event detection, and pitch detection, respectively—demonstrating significantly enhanced fine-grained frame-level perception.

1 citationsRead paper

MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

Jun 18, 2026

This work addresses the inherent many-to-many ambiguity in audio–text alignment, which arises from the superposition of multiple sound events in acoustic scenes and the diversity of textual descriptions. To tackle this challenge, the authors propose a probabilistic language–audio pretraining framework that represents each modality as a distribution rather than a deterministic embedding. By simulating realistic acoustic mixtures through mixed audio–text pairs, the model achieves uncertainty-aware cross-modal alignment. The approach innovatively incorporates mixture-based sound event modeling to capture semantic inclusion relationships and introduces a multi-level inclusion loss function, thereby overcoming the limitations of conventional contrastive learning paradigms that rely on deterministic representations. Experimental results demonstrate that the proposed method significantly outperforms existing deterministic baselines on standard audio–text retrieval benchmarks.

0 citationsRead paper

WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing

Jun 02, 2025

End-to-end CTC-based ASR models suffer from poor recognition accuracy for rare words—especially proper nouns—due to their strong dependence on the lexical distribution observed during training. To address this, we propose a novel inference-time context bias correction method that requires neither model retraining nor text-to-speech (TTS) synthesis. Our approach leverages intermediate acoustic features to implement a robust, real-time keyword detection mechanism via Wildcard CTC—a wildcard-aware extension of the CTC loss. Detected keywords then dynamically modulate subsequent network layers through context-aware acoustic bias injection. This cross-layer bias injection scheme is the first of its kind and enables plug-and-play integration with large-scale pre-trained ASR models. Experiments on Japanese ASR demonstrate a 29% improvement in F1 score for out-of-vocabulary words, substantially enhancing open-vocabulary recognition capability.

0 citationsRead paper
Recent publications

Latest Papers

MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

Jun 18, 2026

This work addresses the inherent many-to-many ambiguity in audio–text alignment, which arises from the superposition of multiple sound events in acoustic scenes and the diversity of textual descriptions. To tackle this challenge, the authors propose a probabilistic language–audio pretraining framework that represents each modality as a distribution rather than a deterministic embedding. By simulating realistic acoustic mixtures through mixed audio–text pairs, the model achieves uncertainty-aware cross-modal alignment. The approach innovatively incorporates mixture-based sound event modeling to capture semantic inclusion relationships and introduces a multi-level inclusion loss function, thereby overcoming the limitations of conventional contrastive learning paradigms that rely on deterministic representations. Experimental results demonstrate that the proposed method significantly outperforms existing deterministic baselines on standard audio–text retrieval benchmarks.

0 citationsRead paper

WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing

Jun 02, 2025

End-to-end CTC-based ASR models suffer from poor recognition accuracy for rare words—especially proper nouns—due to their strong dependence on the lexical distribution observed during training. To address this, we propose a novel inference-time context bias correction method that requires neither model retraining nor text-to-speech (TTS) synthesis. Our approach leverages intermediate acoustic features to implement a robust, real-time keyword detection mechanism via Wildcard CTC—a wildcard-aware extension of the CTC loss. Detected keywords then dynamically modulate subsequent network layers through context-aware acoustic bias injection. This cross-layer bias injection scheme is the first of its kind and enables plug-and-play integration with large-scale pre-trained ASR models. Experiments on Japanese ASR demonstrate a 29% improvement in F1 score for out-of-vocabulary words, substantially enhancing open-vocabulary recognition capability.

0 citationsRead paper

Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio Representation

May 23, 2022IEEE International Conference on Acoustics, Speech, and Signal Processing

Conventional self-supervised audio representation learning relies solely on clip-level sampling, leading to insufficient frame-level modeling capability. Method: This paper proposes a multi-granularity contrastive learning framework that jointly leverages clip-level, frame-level, and task-guided sampling to construct multi-perspective contrastive losses, enabling collaborative optimization of general-purpose audio representations. Contribution/Results: To our knowledge, this is the first work to incorporate both frame-level and task-specific sampling into self-supervised pre-training, overcoming the limitations of single-granularity representation learning. Pre-trained on a subset of AudioSet and evaluated via frozen-feature transfer to downstream tasks, our method achieves 25%, 20%, and 3.6% absolute improvements in clip classification, sound event detection, and pitch detection, respectively—demonstrating significantly enhanced fine-grained frame-level perception.

1 citationsRead paper