CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education
CRITICS项目通过结合大型语言模型的机器翻译与教育技术,解决科学教育中的语言障碍问题,提高科学材料的可访问性和理解度,促进批判性思维。
CRITICS项目通过结合大型语言模型的机器翻译与教育技术,解决科学教育中的语言障碍问题,提高科学材料的可访问性和理解度,促进批判性思维。
This study addresses the challenge of securely deploying large language models (LLMs) for low-resource Baltic languages—Lithuanian, Latvian, and Estonian—in privacy- and security-sensitive domains such as government and defense. We systematically evaluate open-source, locally deployable multilingual LLMs—including Llama 3, Gemma 2, Phi, and NeMo—across machine translation, multiple-choice question answering, and free-text generation. We identify pervasive token-level hallucinations (average error rate ≥5%, i.e., one error per 20 tokens), demonstrating that high translation accuracy does not guarantee semantic reliability. Through FP16/INT4 precision analysis and a custom evaluation benchmark, we find Gemma 2 approaches commercial-model performance, yet all models exhibit critical deficiencies requiring language-specific optimization. Our work establishes the first empirical benchmark for LLM deployment in privacy-critical, low-resource language settings and provides concrete, actionable pathways for improving linguistic fidelity and trustworthiness.
CRITICS项目通过结合大型语言模型的机器翻译与教育技术,解决科学教育中的语言障碍问题,提高科学材料的可访问性和理解度,促进批判性思维。
This study addresses the challenge of securely deploying large language models (LLMs) for low-resource Baltic languages—Lithuanian, Latvian, and Estonian—in privacy- and security-sensitive domains such as government and defense. We systematically evaluate open-source, locally deployable multilingual LLMs—including Llama 3, Gemma 2, Phi, and NeMo—across machine translation, multiple-choice question answering, and free-text generation. We identify pervasive token-level hallucinations (average error rate ≥5%, i.e., one error per 20 tokens), demonstrating that high translation accuracy does not guarantee semantic reliability. Through FP16/INT4 precision analysis and a custom evaluation benchmark, we find Gemma 2 approaches commercial-model performance, yet all models exhibit critical deficiencies requiring language-specific optimization. Our work establishes the first empirical benchmark for LLM deployment in privacy-critical, low-resource language settings and provides concrete, actionable pathways for improving linguistic fidelity and trustworthiness.