VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决非人类病理学领域缺乏基准问题,通过创建包含1251个问题和419张大鼠组织切片图像的VIPER数据集,评估视觉-语言模型性能。
📝 Abstract
Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E-stained rat histology images across seven organ systems, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.
Problem

Research questions and friction points this paper is trying to address.

non-human pathology
toxicologic pathology
vision-language models
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

VIPER
vision-language models
toxicologic pathology
domain-specific training
veterinary pathology