What Will This Copper Look Like Later? Forecasting Surface Appearance and Rendering It as a PBR Material
本文提出了一种预测铜表面氧化外观的方法,并将其渲染为PBR材质,通过无参数的全局颜色外推法实现跨样本预测。
本文提出了一种预测铜表面氧化外观的方法,并将其渲染为PBR材质,通过无参数的全局颜色外推法实现跨样本预测。
This study addresses the low speech recognition accuracy of Burmese medical audio in existing large models, primarily due to data scarcity in clinical settings and environmental noise. To tackle this challenge, we present the first high-quality Burmese medical speech corpus comprising 28 hours of native-speaker-recorded and validated utterances. Leveraging this dataset, we fine-tune the Whisper model using both full-parameter fine-tuning (FFT) and parameter-efficient LoRA-based fine-tuning (PEFT), further enhancing robustness through waveform- and spectrogram-level data augmentation to mitigate the effects of noise and reverberation. Our best-performing system, myMediWhisper-Medium, achieves a word error rate (WER) of 23.44% without data augmentation, significantly outperforming larger general-domain fine-tuned models and establishing a new state-of-the-art result for Burmese medical speech recognition.
This study addresses the lack of systematic research on automatic speech recognition (ASR) error correction for low-resource Burmese. We propose the first Transformer-based correction framework that jointly models phonemes and alignment information. Our approach innovatively incorporates International Phonetic Alphabet (IPA) representations and forced-alignment features, while fusing outputs from multiple ASR backends to enable joint word- and character-level optimization within a sequence-to-sequence architecture. Experiments show that our method significantly reduces the average word error rate (WER) from 51.56% to 39.82% on unaugmented data, and improves the chrF++ score to 0.627 (+0.0406), substantially outperforming baselines. It maintains robust performance even with data augmentation (WER = 43.59%). This work establishes a transferable, phoneme-aware modeling paradigm for ASR error correction in low-resource languages.
Large language models (LLMs) face deployment challenges due to their massive parameter counts and high computational overhead. Conventional low-rank compression methods—e.g., singular value decomposition (SVD)—optimize only for matrix reconstruction error, neglecting functional information loss, which leads to substantial performance degradation. To address this, we propose CALR, a layer-wise low-rank compression framework tailored for LLMs: its primary path employs SVD for efficient parameter reduction, while a novel, learnable, parallel low-rank correction module explicitly models functional information loss as a trainable signal and recovers functional residuals via end-to-end optimization. CALR is architecture-agnostic and achieves 26.93%–51.77% parameter compression on multiple mainstream small-scale LLMs, retaining 59.45%–90.42% of original task performance—significantly outperforming baselines including LaCo, ShortGPT, and LoSparse.
This paper addresses three challenging sentence classification tasks for Burmese: imbalanced binary hate speech detection, balanced multi-class news categorization, and imbalanced multi-class ethnic language identification. We propose KAConvText, the first text classification model to incorporate Kolmogorov–Arnold convolution (KA-Conv), integrated with fine-tuned fastText embeddings (CBOW/Skip-gram) and an interpretable Kolmogorov–Arnold Network (KAN)-based classifier head; an MLP-based baseline is also supported. On the three tasks, KAConvText achieves accuracies of 91.23% (F1 = 0.9109), 92.66% (F1 = 0.9267), and 99.82% (F1 = 0.9982), respectively—outperforming all existing baselines. Our core contribution is the first architectural integration of KA networks into NLP text classification, uniquely balancing performance gains with intrinsic model interpretability. This work establishes a novel paradigm for imbalanced text classification in low-resource languages.
本文提出了一种预测铜表面氧化外观的方法,并将其渲染为PBR材质,通过无参数的全局颜色外推法实现跨样本预测。
This study addresses the low speech recognition accuracy of Burmese medical audio in existing large models, primarily due to data scarcity in clinical settings and environmental noise. To tackle this challenge, we present the first high-quality Burmese medical speech corpus comprising 28 hours of native-speaker-recorded and validated utterances. Leveraging this dataset, we fine-tune the Whisper model using both full-parameter fine-tuning (FFT) and parameter-efficient LoRA-based fine-tuning (PEFT), further enhancing robustness through waveform- and spectrogram-level data augmentation to mitigate the effects of noise and reverberation. Our best-performing system, myMediWhisper-Medium, achieves a word error rate (WER) of 23.44% without data augmentation, significantly outperforming larger general-domain fine-tuned models and establishing a new state-of-the-art result for Burmese medical speech recognition.
This study addresses the lack of systematic research on automatic speech recognition (ASR) error correction for low-resource Burmese. We propose the first Transformer-based correction framework that jointly models phonemes and alignment information. Our approach innovatively incorporates International Phonetic Alphabet (IPA) representations and forced-alignment features, while fusing outputs from multiple ASR backends to enable joint word- and character-level optimization within a sequence-to-sequence architecture. Experiments show that our method significantly reduces the average word error rate (WER) from 51.56% to 39.82% on unaugmented data, and improves the chrF++ score to 0.627 (+0.0406), substantially outperforming baselines. It maintains robust performance even with data augmentation (WER = 43.59%). This work establishes a transferable, phoneme-aware modeling paradigm for ASR error correction in low-resource languages.
Large language models (LLMs) face deployment challenges due to their massive parameter counts and high computational overhead. Conventional low-rank compression methods—e.g., singular value decomposition (SVD)—optimize only for matrix reconstruction error, neglecting functional information loss, which leads to substantial performance degradation. To address this, we propose CALR, a layer-wise low-rank compression framework tailored for LLMs: its primary path employs SVD for efficient parameter reduction, while a novel, learnable, parallel low-rank correction module explicitly models functional information loss as a trainable signal and recovers functional residuals via end-to-end optimization. CALR is architecture-agnostic and achieves 26.93%–51.77% parameter compression on multiple mainstream small-scale LLMs, retaining 59.45%–90.42% of original task performance—significantly outperforming baselines including LaCo, ShortGPT, and LoSparse.
This paper addresses three challenging sentence classification tasks for Burmese: imbalanced binary hate speech detection, balanced multi-class news categorization, and imbalanced multi-class ethnic language identification. We propose KAConvText, the first text classification model to incorporate Kolmogorov–Arnold convolution (KA-Conv), integrated with fine-tuned fastText embeddings (CBOW/Skip-gram) and an interpretable Kolmogorov–Arnold Network (KAN)-based classifier head; an MLP-based baseline is also supported. On the three tasks, KAConvText achieves accuracies of 91.23% (F1 = 0.9109), 92.66% (F1 = 0.9267), and 99.82% (F1 = 0.9982), respectively—outperforming all existing baselines. Our core contribution is the first architectural integration of KA networks into NLP text classification, uniquely balancing performance gains with intrinsic model interpretability. This work establishes a novel paradigm for imbalanced text classification in low-resource languages.