Institution profile

VK

Industry researcheurope · ru
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks

Aug 16, 2025

In conventional fine-tuning of self-supervised speech models, fixed-layer aggregation—such as using only the top layer or static weighted summation—introduces an information bottleneck and limits cross-sample generalization. To address this, we propose VARAN, a novel framework featuring input-adaptive dynamic layer aggregation. VARAN employs layer-specific probe heads and data-dependent weights to dynamically allocate contributions from different transformer layers per sample, thereby preserving layer-specific characteristics while enhancing representational flexibility. The aggregation process is optimized via variational inference, and VARAN integrates LoRA for parameter-efficient fine-tuning. Evaluated on automatic speech recognition and speech emotion recognition tasks, VARAN consistently outperforms strong baselines; its combination with LoRA yields particularly substantial gains. These results demonstrate VARAN’s superior downstream adaptability and robust generalization capability across diverse speech understanding tasks.

0 citationsRead paper
Recent publications

Latest Papers

VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks

Aug 16, 2025

In conventional fine-tuning of self-supervised speech models, fixed-layer aggregation—such as using only the top layer or static weighted summation—introduces an information bottleneck and limits cross-sample generalization. To address this, we propose VARAN, a novel framework featuring input-adaptive dynamic layer aggregation. VARAN employs layer-specific probe heads and data-dependent weights to dynamically allocate contributions from different transformer layers per sample, thereby preserving layer-specific characteristics while enhancing representational flexibility. The aggregation process is optimized via variational inference, and VARAN integrates LoRA for parameter-efficient fine-tuning. Evaluated on automatic speech recognition and speech emotion recognition tasks, VARAN consistently outperforms strong baselines; its combination with LoRA yields particularly substantial gains. These results demonstrate VARAN’s superior downstream adaptability and robust generalization capability across diverse speech understanding tasks.

0 citationsRead paper