From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports
This study addresses the low credibility of liver MRI reports generated by Chinese large language models (LLMs). We propose the first multidimensional credibility assessment framework specifically designed for medical imaging reporting and introduce a clinical-context-driven, institution-level prompt optimization methodology. Leveraging the SiliconFlow platform, we systematically evaluate leading open-weight Chinese LLMs—including Kimi-K2, Qwen3-235B, DeepSeek-V3, and ByteDance-Seed-OSS—across diverse clinical scenarios, uncovering how prompt design differentially impacts diagnostic accuracy, terminology standardization, logical coherence, and clinical interpretability. Results demonstrate that our framework significantly improves report accuracy (+18.7%) and inter-model consistency (Cohen’s κ = 0.82). It establishes a reproducible, verifiable evaluation paradigm and an engineering-oriented optimization pathway for radiology AI-assisted report generation.