CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing automatic spoken language assessment methods struggle to disentangle acoustic and content-related information effectively, often suffering from model complexity and insufficient stability. To overcome these limitations, the authors propose a lightweight architecture that integrates the Whisper-medium speech encoder with the Qwen2-7B large language model, enabling—without any additional training—the first demonstration of content validation purely based on pre-trained components. The approach further enhances interpretability by incorporating three handcrafted fluency features. Through multimodal fusion and comprehensive ablation studies, the method achieves a root mean square error (RMSE) of 0.358 on the Speak & Improve Corpus 2025, outperforming current state-of-the-art systems while reducing inference parameters by nearly half, thereby significantly improving both performance and deployability.
📝 Abstract
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.
Problem

Research questions and friction points this paper is trying to address.

automatic speaking assessment
acoustic-content contribution
performance stability
multimodal speech evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

automatic speaking assessment
multimodal large language model
speech-content disentanglement
interpretable ASA
training-free content validation
🔎 Similar Papers
No similar papers found.