Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了长音频临床文档自动生成问题,通过构建多任务语料库和使用端到端模型直接从音频生成SOAP笔记,优于传统级联系统。
📝 Abstract
Automating clinical documentation from long-form doctor-patient conversations remains challenging for modern audio-language models. While cascaded ASR systems perform well, end-to-end (E2E) models often struggle with information loss and hallucinations on extended audio. For the BeTraC 2026 challenge, the ASLP team presents a fully E2E multimodal system that generates structured SOAP notes directly from audio, bypassing intermediate transcripts. We constructed a 1.41-million-sample multi-task corpus and applied a multi-stage pipeline: domain pre-training, supervised fine-tuning, and reward optimization. Evaluating the architecture under both Lightweight (3B) and Heavyweight (30B) constraints reveals that each training stage progressively enhances performance. Furthermore, scaling to 30B parameters substantially boosts concept extraction and summarization quality. Ultimately, our E2E systems consistently outperform representative cascaded ASR+LLM baselines, proving the efficacy of direct multimodal optimization for clinical documentation.
Problem

Research questions and friction points this paper is trying to address.

End-to-End
Clinical Documentation
Long-Form Audio
Information Loss
Hallucinations
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-End
Multimodal System
SOAP Generation
Multi-Stage Training
Concept Extraction
🔎 Similar Papers
2024-09-272024 6th International Conference on Artificial Intelligence and Computer Applications (ICAICA)Citations: 6
💼 Related Jobs
No related jobs found.
Z
Ziyu Zhang
Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University, China
M
Mingchen Shao
Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University, China
Wenjie Tian
Wenjie Tian
Northwest Polytechnical University
speech generation
T
Tianlun Zuo
Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University, China
L
Longhao Li
Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University, China
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence