Institution profile

Chuzhou University

Academic institutionasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

U$^2$Mamba: A Two-level Nested U-structure Mamba for Salient Object Detection

Jun 18, 2026

This work addresses the limitations of existing Mamba-based salient object detection methods, which struggle to adequately model long-range contextual dependencies and are constrained by network depth. To overcome these issues, we propose a dual-nested U-shaped architecture that enhances local feature extraction through multi-scale Mamba U-blocks and effectively integrates shallow and deep features with diverse receptive fields to capture richer contextual relationships. Furthermore, we introduce a novel hierarchical training supervision mechanism that applies loss constraints at multiple network levels, departing from conventional top-layer-only supervision. Extensive experiments demonstrate that the proposed method achieves state-of-the-art or superior performance on salient object detection benchmarks, validating its effectiveness and architectural advantages.

0 citationsRead paper

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

May 20, 2026

This work addresses the challenges of fine-grained fruit recognition—namely, the scarcity of high-quality labeled data and the high visual similarity among categories—by introducing a large-scale dataset encompassing 306 fruit classes. The authors propose a two-stage dynamic inference framework: in the first stage, a verification-calibrated ensemble of heterogeneous models generates a Top-3 candidate set; for low-confidence samples, the second stage employs a novel chain-of-thought arbitration mechanism guided by a multimodal large language model (MLLM). Coupled with a hard-sample-aware joint loss, this approach significantly enhances generalization. Evaluated on the newly curated dataset, the method achieves a classification accuracy of 70.49%, outperforming current state-of-the-art approaches and demonstrating strong potential for real-world deployment in agricultural visual sorting and quality inspection systems.

0 citationsRead paper

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

Mar 31, 2026

This work addresses the high computational complexity, substantial memory footprint, and deployment challenges in diffusion-based text-to-speech synthesis caused by reliance on attention mechanisms and recurrent structures. The authors propose the first fully state space model (SSM)-based conditional pathway, entirely eliminating attention and RNN modules. Their approach employs a gated bidirectional Mamba text encoder, a temporally bidirectional Mamba alignment module, and an expressive Mamba modulated via AdaLN, trained under the supervision of a lightweight alignment teacher. This design achieves high-quality speech generation while significantly improving memory efficiency, inference stability, and streaming capability. Experiments demonstrate consistent performance gains over strong baselines such as StyleTTS2 and VITS across multiple datasets, with improvements in MOS/CMOS scores, F0 RMSE, MCD, and WER. The encoder size is reduced to 21M parameters, and throughput increases by 1.6×.

0 citationsRead paper
Recent publications

Latest Papers

U$^2$Mamba: A Two-level Nested U-structure Mamba for Salient Object Detection

Jun 18, 2026

This work addresses the limitations of existing Mamba-based salient object detection methods, which struggle to adequately model long-range contextual dependencies and are constrained by network depth. To overcome these issues, we propose a dual-nested U-shaped architecture that enhances local feature extraction through multi-scale Mamba U-blocks and effectively integrates shallow and deep features with diverse receptive fields to capture richer contextual relationships. Furthermore, we introduce a novel hierarchical training supervision mechanism that applies loss constraints at multiple network levels, departing from conventional top-layer-only supervision. Extensive experiments demonstrate that the proposed method achieves state-of-the-art or superior performance on salient object detection benchmarks, validating its effectiveness and architectural advantages.

0 citationsRead paper

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

May 20, 2026

This work addresses the challenges of fine-grained fruit recognition—namely, the scarcity of high-quality labeled data and the high visual similarity among categories—by introducing a large-scale dataset encompassing 306 fruit classes. The authors propose a two-stage dynamic inference framework: in the first stage, a verification-calibrated ensemble of heterogeneous models generates a Top-3 candidate set; for low-confidence samples, the second stage employs a novel chain-of-thought arbitration mechanism guided by a multimodal large language model (MLLM). Coupled with a hard-sample-aware joint loss, this approach significantly enhances generalization. Evaluated on the newly curated dataset, the method achieves a classification accuracy of 70.49%, outperforming current state-of-the-art approaches and demonstrating strong potential for real-world deployment in agricultural visual sorting and quality inspection systems.

0 citationsRead paper

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

Mar 31, 2026

This work addresses the high computational complexity, substantial memory footprint, and deployment challenges in diffusion-based text-to-speech synthesis caused by reliance on attention mechanisms and recurrent structures. The authors propose the first fully state space model (SSM)-based conditional pathway, entirely eliminating attention and RNN modules. Their approach employs a gated bidirectional Mamba text encoder, a temporally bidirectional Mamba alignment module, and an expressive Mamba modulated via AdaLN, trained under the supervision of a lightweight alignment teacher. This design achieves high-quality speech generation while significantly improving memory efficiency, inference stability, and streaming capability. Experiments demonstrate consistent performance gains over strong baselines such as StyleTTS2 and VITS across multiple datasets, with improvements in MOS/CMOS scores, F0 RMSE, MCD, and WER. The encoder size is reduced to 21M parameters, and throughput increases by 1.6×.

0 citationsRead paper