Institution profile

Fano Labs

Industry researchasia · hk
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Pushing the Limits of End-to-End Diarization

Sep 18, 2025

This study addresses the limited accuracy of joint segmentation and clustering in end-to-end speaker diarization. We propose a unified solution based on an extended non-autoregressive EEND-TA architecture. To enhance modeling capability for complex overlapping speech and long-tail scenarios, we construct a large-scale synthetic dataset covering up to eight simultaneous speakers and adopt a pretraining-finetuning paradigm to improve generalization. Our method directly outputs speaker label sequences without requiring post-processing steps such as clustering or voice activity detection (VAD), thereby simplifying the pipeline and improving robustness. Evaluated on standard benchmarks—including AliMeeting, AMI, and DIHARD III—our approach achieves state-of-the-art performance, with a DER of 14.49% on DIHARD III, demonstrating strong cross-domain adaptability and practical deployability.

0 citationsRead paper

VP-SelDoA: Visual-prompted Selective DoA Estimation of Target Sound via Semantic-Spatial Matching

Jul 09, 2025

To address three key challenges in audio-visual sound source localization (AVSSL) under multi-source scenarios—poor target-source selectivity, semantic-visual–spatial-auditory feature mismatch, and excessive reliance on paired audio-visual data—this paper proposes Visual-Prompted Selective Direction-of-Arrival (VP-SelDoA). We introduce a novel cross-instance audio-visual learning (CI-AVL) paradigm, featuring semantic-level modality fusion and Semantic-Spatial Matching to substantially reduce dependence on paired data. Furthermore, we design a Frequency-Temporal ConMamba architecture to generate selective masks, integrating cross- and self-attention for heterogeneous feature alignment. Evaluated on our large-scale spatial audio dataset VGG-SSL, VP-SelDoA achieves a mean absolute error of 12.04° and a localization accuracy of 78.23%, outperforming state-of-the-art methods and demonstrating superior selectivity and generalization.

0 citationsRead paper
Recent publications

Latest Papers

Pushing the Limits of End-to-End Diarization

Sep 18, 2025

This study addresses the limited accuracy of joint segmentation and clustering in end-to-end speaker diarization. We propose a unified solution based on an extended non-autoregressive EEND-TA architecture. To enhance modeling capability for complex overlapping speech and long-tail scenarios, we construct a large-scale synthetic dataset covering up to eight simultaneous speakers and adopt a pretraining-finetuning paradigm to improve generalization. Our method directly outputs speaker label sequences without requiring post-processing steps such as clustering or voice activity detection (VAD), thereby simplifying the pipeline and improving robustness. Evaluated on standard benchmarks—including AliMeeting, AMI, and DIHARD III—our approach achieves state-of-the-art performance, with a DER of 14.49% on DIHARD III, demonstrating strong cross-domain adaptability and practical deployability.

0 citationsRead paper

VP-SelDoA: Visual-prompted Selective DoA Estimation of Target Sound via Semantic-Spatial Matching

Jul 09, 2025

To address three key challenges in audio-visual sound source localization (AVSSL) under multi-source scenarios—poor target-source selectivity, semantic-visual–spatial-auditory feature mismatch, and excessive reliance on paired audio-visual data—this paper proposes Visual-Prompted Selective Direction-of-Arrival (VP-SelDoA). We introduce a novel cross-instance audio-visual learning (CI-AVL) paradigm, featuring semantic-level modality fusion and Semantic-Spatial Matching to substantially reduce dependence on paired data. Furthermore, we design a Frequency-Temporal ConMamba architecture to generate selective masks, integrating cross- and self-attention for heterogeneous feature alignment. Evaluated on our large-scale spatial audio dataset VGG-SSL, VP-SelDoA achieves a mean absolute error of 12.04° and a localization accuracy of 78.23%, outperforming state-of-the-art methods and demonstrating superior selectivity and generalization.

0 citationsRead paper