Institution profile

Beijing University of Civil Engineering and Architecture

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

May 12, 2026

This work addresses the challenge that existing singing voice conversion (SVC) methods struggle to reliably extract clean vocal melodies from accompanied recordings due to harmonic interference. To overcome this limitation, we propose a zero-shot, cross-lingual SVC system that explicitly models both the main melody and residual harmonics—a first in SVC—enabling effective processing of polyphonic audio. The architecture integrates a CQT-based pitch extractor, a stochastic sampler, and a conditional flow-matching diffusion decoder, jointly optimizing pitch, linguistic content, and time–frequency features. Experimental results demonstrate that our approach consistently outperforms current baselines on both harmonically rich and monophonic datasets, achieving superior performance in terms of naturalness, timbre similarity, and harmonic reconstruction fidelity.

0 citationsRead paper

LLM-YOLOMS: Large Language Model-based Semantic Interpretation and Fault Diagnosis for Wind Turbine Components

Nov 13, 2025

To address the lack of semantic interpretability in visual inspection results for wind turbine (WT) component fault diagnosis—hindering operational and maintenance decision-making—this paper proposes a vision–language collaborative diagnostic framework. First, YOLOMS with multi-scale detection and sliding-window cropping enhances small-object recognition accuracy. Second, a lightweight key-value (KV) mapping module automatically converts structured detection outputs—including bounding box coordinates, class labels, and confidence scores—into qualitative and quantitative natural-language descriptions. Third, a domain-adapted large language model (LLM) performs semantic reasoning to generate human-understandable fault analyses and actionable maintenance recommendations. Evaluated on a real-world dataset, the framework achieves 90.6% fault detection accuracy and 89% accuracy in generating maintenance reports, significantly improving both the interpretability and engineering applicability of diagnostic outcomes.

0 citationsRead paper

MedBuild AI: An Agent-Based Hybrid Intelligence Framework for Reshaping Agency in Healthcare Infrastructure Planning through Generative Design for Medical Architecture

Oct 17, 2025

Global healthcare infrastructure is severely unevenly distributed, with remote and underserved regions lacking access to basic medical services; conventional planning methods fail to address the scale and urgency of this challenge. This study proposes an agent-based hybrid intelligence framework that integrates large language models (LLMs) with deterministic rule engines, implementing a tri-agent collaborative system—comprising demand elicitation, rule-based inference, and 3D generative design—to enable natural-language health-need articulation, automated functional layout translation, and climate- and resource-constrained adaptation in a closed-loop design workflow. A lightweight web platform, deployed via satellite internet, supports real-time, multilingual solution generation under low-bandwidth conditions. Empirical evaluation demonstrates that the system generates code-compliant healthcare facility designs within minutes, reducing preliminary planning time by over 80%. Feasibility and improved accessibility have been validated across multiple low-resource settings.

0 citationsRead paper

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis

Sep 04, 2025

Current conversational TTS systems suffer from a scarcity of natural, interactive bilingual speech data, hindering effective modeling of authentic dialogue phenomena—such as overlapping speech, backchannel responses, and laughter. To address this, we introduce the first high-quality, bilingual (Chinese–English), full-duplex, spontaneous dialogue speech corpus (15 hours of multi-track recordings), covering everyday topics and genuine interactive behaviors. We propose an open-source dual-channel acquisition protocol and a fine-grained transcription annotation schema explicitly designed to capture overlapping utterances, nonverbal vocalizations, and feedback responses. Fine-tuning TTS models on this dataset yields statistically significant improvements over strong baselines in both objective metrics and subjective evaluations—particularly in speech naturalness and dialogue realism. This work establishes a foundational bilingual resource and methodological framework for advancing conversational speech synthesis.

0 citationsRead paper
Recent publications

Latest Papers

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

May 12, 2026

This work addresses the challenge that existing singing voice conversion (SVC) methods struggle to reliably extract clean vocal melodies from accompanied recordings due to harmonic interference. To overcome this limitation, we propose a zero-shot, cross-lingual SVC system that explicitly models both the main melody and residual harmonics—a first in SVC—enabling effective processing of polyphonic audio. The architecture integrates a CQT-based pitch extractor, a stochastic sampler, and a conditional flow-matching diffusion decoder, jointly optimizing pitch, linguistic content, and time–frequency features. Experimental results demonstrate that our approach consistently outperforms current baselines on both harmonically rich and monophonic datasets, achieving superior performance in terms of naturalness, timbre similarity, and harmonic reconstruction fidelity.

0 citationsRead paper

LLM-YOLOMS: Large Language Model-based Semantic Interpretation and Fault Diagnosis for Wind Turbine Components

Nov 13, 2025

To address the lack of semantic interpretability in visual inspection results for wind turbine (WT) component fault diagnosis—hindering operational and maintenance decision-making—this paper proposes a vision–language collaborative diagnostic framework. First, YOLOMS with multi-scale detection and sliding-window cropping enhances small-object recognition accuracy. Second, a lightweight key-value (KV) mapping module automatically converts structured detection outputs—including bounding box coordinates, class labels, and confidence scores—into qualitative and quantitative natural-language descriptions. Third, a domain-adapted large language model (LLM) performs semantic reasoning to generate human-understandable fault analyses and actionable maintenance recommendations. Evaluated on a real-world dataset, the framework achieves 90.6% fault detection accuracy and 89% accuracy in generating maintenance reports, significantly improving both the interpretability and engineering applicability of diagnostic outcomes.

0 citationsRead paper

MedBuild AI: An Agent-Based Hybrid Intelligence Framework for Reshaping Agency in Healthcare Infrastructure Planning through Generative Design for Medical Architecture

Oct 17, 2025

Global healthcare infrastructure is severely unevenly distributed, with remote and underserved regions lacking access to basic medical services; conventional planning methods fail to address the scale and urgency of this challenge. This study proposes an agent-based hybrid intelligence framework that integrates large language models (LLMs) with deterministic rule engines, implementing a tri-agent collaborative system—comprising demand elicitation, rule-based inference, and 3D generative design—to enable natural-language health-need articulation, automated functional layout translation, and climate- and resource-constrained adaptation in a closed-loop design workflow. A lightweight web platform, deployed via satellite internet, supports real-time, multilingual solution generation under low-bandwidth conditions. Empirical evaluation demonstrates that the system generates code-compliant healthcare facility designs within minutes, reducing preliminary planning time by over 80%. Feasibility and improved accessibility have been validated across multiple low-resource settings.

0 citationsRead paper

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis

Sep 04, 2025

Current conversational TTS systems suffer from a scarcity of natural, interactive bilingual speech data, hindering effective modeling of authentic dialogue phenomena—such as overlapping speech, backchannel responses, and laughter. To address this, we introduce the first high-quality, bilingual (Chinese–English), full-duplex, spontaneous dialogue speech corpus (15 hours of multi-track recordings), covering everyday topics and genuine interactive behaviors. We propose an open-source dual-channel acquisition protocol and a fine-grained transcription annotation schema explicitly designed to capture overlapping utterances, nonverbal vocalizations, and feedback responses. Fine-tuning TTS models on this dataset yields statistically significant improvements over strong baselines in both objective metrics and subjective evaluations—particularly in speech naturalness and dialogue realism. This work establishes a foundational bilingual resource and methodological framework for advancing conversational speech synthesis.

0 citationsRead paper