V-REX: Efficient Specialist VLM Training for Veterinary X-Rays

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对兽医X光片诊断问题,通过重新设计VLM训练流程,提出更高效的训练方法,不依赖大规模基础模型即实现优越性能。
📝 Abstract
While generalist VLMs are expensive to train, creating domain experts is widely assumed to require fine-tuning increasingly large foundation models. We show that, in veterinary radiology, this assumption is misguided. By rethinking the entire VLM pipeline - from text tokenisation and pre-training to grounding and inference - we demonstrate that careful engineering can yield models that outperform much larger foundation models from scratch, without relying on any other data. Our approach introduces new strategies for generative pre-training and grounding that improve training efficiency, increasing data utilisation and downstream performance. Using only a fraction of the parameters, data, and compute of contemporary generalist models, we develop the first VLM capable of generating diagnostic reports for veterinary radiographs, surpassing open foundation models on this task by significant margin.
Problem

Research questions and friction points this paper is trying to address.

VLM
veterinary radiology
fine-tuning
foundation models
training efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Efficient VLM Training
Veterinary Radiology
Generative Pre-training
Data Utilization
Diagnostic Report Generation