Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation
研究通过调整Whisper模型规模、多任务训练和两阶段适应,解决希腊歌曲自动歌词转录中的旋律变化、节奏不规则和伴奏干扰问题。
研究通过调整Whisper模型规模、多任务训练和两阶段适应,解决希腊歌曲自动歌词转录中的旋律变化、节奏不规则和伴奏干扰问题。
本文针对低资源希腊语TTS质量问题,通过数据处理和模型微调方法,利用确定性提示解决了说话人一致性问题。
This study addresses the inefficiency of traditional literature retrieval methods, which rely heavily on manually crafted queries and screening, in the face of exponentially growing academic publications. The authors present a systematic review of 34 peer-reviewed studies that apply generative large language models (LLMs) to scientific literature retrieval and screening, offering the first comprehensive mapping of their application paradigms in scientific knowledge discovery. By leveraging Boolean queries on the OpenAIRE Graph and analyzing approaches through the lenses of prompt engineering, model adaptation, and architectural design, the work identifies key technical pathways and evaluation frameworks. Beyond structuring the current landscape, the study highlights the substantial potential of LLMs to enhance the automation and efficiency of scientific discovery, providing a systematic reference for future research.
This work addresses the limitations of Recall@k as the dominant evaluation metric in approximate nearest neighbor (ANN) search, which often overestimates retrieval quality and incurs redundant computation. The authors propose 1/Ratio@k—an inverse approximation ratio that is hyperparameter-free and directly computable from benchmark data—as a more principled alternative. Through extensive evaluation of state-of-the-art ANN algorithms across diverse high-dimensional datasets, coupled with efficiency analysis and validation on downstream tasks such as classification and retrieval-augmented generation, the study demonstrates that optimizing 1/Ratio@k significantly reduces computational overhead while preserving practical utility. Moreover, 1/Ratio@k exhibits substantially stronger correlation with real-world effectiveness—measured by label accuracy and semantic similarity—than Recall@k.
This study addresses the linguistic barriers impeding the global dissemination of scientific research, which generic machine translation systems struggle to overcome due to their inability to accurately handle domain-specific terminology and complex syntactic structures in scholarly texts. To bridge this gap, the authors present the first systematic construction of Spanish–English, French–English, and Portuguese–English parallel and monolingual corpora spanning four scientific subfields: cancer, energy, neuroscience, and transportation. Leveraging these resources, they perform domain-adaptive fine-tuning of neural machine translation models. Experimental results demonstrate that the fine-tuned systems significantly outperform generic baselines in translation quality for scientific content, thereby confirming the critical role of multilingual, multidisciplinary specialized corpora in enhancing the accuracy and fluency of research literature translation.
研究通过调整Whisper模型规模、多任务训练和两阶段适应,解决希腊歌曲自动歌词转录中的旋律变化、节奏不规则和伴奏干扰问题。
本文针对低资源希腊语TTS质量问题,通过数据处理和模型微调方法,利用确定性提示解决了说话人一致性问题。
This study addresses the inefficiency of traditional literature retrieval methods, which rely heavily on manually crafted queries and screening, in the face of exponentially growing academic publications. The authors present a systematic review of 34 peer-reviewed studies that apply generative large language models (LLMs) to scientific literature retrieval and screening, offering the first comprehensive mapping of their application paradigms in scientific knowledge discovery. By leveraging Boolean queries on the OpenAIRE Graph and analyzing approaches through the lenses of prompt engineering, model adaptation, and architectural design, the work identifies key technical pathways and evaluation frameworks. Beyond structuring the current landscape, the study highlights the substantial potential of LLMs to enhance the automation and efficiency of scientific discovery, providing a systematic reference for future research.
This work addresses the limitations of Recall@k as the dominant evaluation metric in approximate nearest neighbor (ANN) search, which often overestimates retrieval quality and incurs redundant computation. The authors propose 1/Ratio@k—an inverse approximation ratio that is hyperparameter-free and directly computable from benchmark data—as a more principled alternative. Through extensive evaluation of state-of-the-art ANN algorithms across diverse high-dimensional datasets, coupled with efficiency analysis and validation on downstream tasks such as classification and retrieval-augmented generation, the study demonstrates that optimizing 1/Ratio@k significantly reduces computational overhead while preserving practical utility. Moreover, 1/Ratio@k exhibits substantially stronger correlation with real-world effectiveness—measured by label accuracy and semantic similarity—than Recall@k.
This study addresses the linguistic barriers impeding the global dissemination of scientific research, which generic machine translation systems struggle to overcome due to their inability to accurately handle domain-specific terminology and complex syntactic structures in scholarly texts. To bridge this gap, the authors present the first systematic construction of Spanish–English, French–English, and Portuguese–English parallel and monolingual corpora spanning four scientific subfields: cancer, energy, neuroscience, and transportation. Leveraging these resources, they perform domain-adaptive fine-tuning of neural machine translation models. Experimental results demonstrate that the fine-tuned systems significantly outperform generic baselines in translation quality for scientific content, thereby confirming the critical role of multilingual, multidisciplinary specialized corpora in enhancing the accuracy and fluency of research literature translation.