PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对3D脑MRI报告生成,通过上游3D分割和分类获取感知衍生事实,并用这些事实提示微调的视觉-语言模型来提高报告质量。
📝 Abstract
Radiology report generation has matured almost entirely on 2D chest radiographs, where the default route to better reports is a larger backbone or a pre-training one on medical data. We revisit that assumption on 3D multi-sequence brain MRI, a volumetric multi-disease regime, and find that the model is not the lever. Zero-shot medical and radiology vision-language models transfer poorly to brain MRI, with chest radiograph specialists failing most conspicuously, and five backbones fine-tuned identically across three model families and an order of magnitude in scale differ only marginally. What determines the quality of the report is the information injected into the prompt. We delegate perception to upstream 3D segmentation and classification, serialize their outputs into a structured fact sentence, and prompt a LoRA-adapted vision-language model with it; we call this \textbf{PerFact}. In a controlled study that fixes the backbone, data split, target reports, and adaptation while varying only the injected grounding, perception-derived facts outperform retrieved prior reports, retrieval becomes redundant once facts are present, and end-to-end predicted facts remain effective without any ground-truth annotation at inference. The residual gap between predicted and oracle facts is explained by the granularity of the facts rather than by the generator. Closed-ended visual question answering comes at no measurable cost to report quality, though the grounding source has little effect on it. On 3D brain MRI, grounding information, not model choice, is the dominant controllable factor in report quality.
Problem

Research questions and friction points this paper is trying to address.

3D Brain MRI
Report Generation
Perception-Derived Facts
Innovation

Methods, ideas, or system contributions that make the work stand out.

PerFact
3D segmentation and classification
structured fact sentence
LoRA-adapted vision-language model
information injection
💼 Related Jobs
No related jobs found.
J
Jianyu Sun
Department of Bioengineering, Imperial College London, London, UK
Zhenxuan Zhang
Zhenxuan Zhang
Georgia Institute of Technology
G
Guang Yang
Department of Bioengineering, Imperial College London, London, UK; Imperial-X, Imperial College London, London, UK; School of Biomedical Engineering & Imaging Sciences, King’s College London, London, UK
P
Peter J. Lally
Department of Bioengineering, Imperial College London, London, UK