VET-DINO: Learning Anatomical Understanding Through Multi-View Distillation in Veterinary Imaging

๐Ÿ“… 2025-05-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Veterinary medical imaging suffers from severe scarcity of expert annotations. Method: This paper proposes a self-supervised learning framework leveraging multi-view X-ray images (e.g., ventrodorsal and lateral views) from the same clinical case, exploiting naturally occurring standardized anatomical correspondences to implicitly learn view-invariant representations and 3D spatial anatomyโ€”without synthetic data augmentation. Contribution/Results: It pioneers the integration of clinical multi-view anatomical priors into medical self-supervised learning, establishing an anatomy-consistency-driven paradigm. The method synergistically combines the DINO architecture, multi-view knowledge distillation, cross-view feature alignment, and contrastive learning. Trained on 5 million canine radiographs, it achieves state-of-the-art performance across multiple downstream tasks, significantly enhancing anatomical understanding and generalization to synthetic or out-of-distribution data.

Technology Category

Application Category

๐Ÿ“ Abstract
Self-supervised learning has emerged as a powerful paradigm for training deep neural networks, particularly in medical imaging where labeled data is scarce. While current approaches typically rely on synthetic augmentations of single images, we propose VET-DINO, a framework that leverages a unique characteristic of medical imaging: the availability of multiple standardized views from the same study. Using a series of clinical veterinary radiographs from the same patient study, we enable models to learn view-invariant anatomical structures and develop an implied 3D understanding from 2D projections. We demonstrate our approach on a dataset of 5 million veterinary radiographs from 668,000 canine studies. Through extensive experimentation, including view synthesis and downstream task performance, we show that learning from real multi-view pairs leads to superior anatomical understanding compared to purely synthetic augmentations. VET-DINO achieves state-of-the-art performance on various veterinary imaging tasks. Our work establishes a new paradigm for self-supervised learning in medical imaging that leverages domain-specific properties rather than merely adapting natural image techniques.
Problem

Research questions and friction points this paper is trying to address.

Leveraging multi-view veterinary images for anatomical learning
Improving 3D understanding from 2D projections via self-supervision
Overcoming labeled data scarcity in medical imaging tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leverages multi-view veterinary radiographs for learning
Develops 3D understanding from 2D projections
Uses self-supervised learning with real multi-view pairs
๐Ÿ’ผ Related Jobs
No related jobs found.
A
Andre Dourson
Mars Petcare
K
Kylie Taylor
Mars Petcare
Xiaoli Qiao
Xiaoli Qiao
Mars Petcare
M
Michael Fitzke
Mars Petcare