Visual Instruction Pretraining for Domain-Specific Foundation Models

📅 2025-09-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing vision foundation models lack a backward guidance mechanism whereby high-level reasoning informs low-level perceptual feature learning, resulting in an incomplete perception–reasoning–generation loop. Method: We propose Vision Instruction Pre-training (ViTP), a novel paradigm built upon Vision Transformers and trained end-to-end on large-scale, domain-specific vision–language instruction data via instruction following and cross-modal alignment. Crucially, we introduce Vision Robustness Learning (VRL), which explicitly encourages the model to extract domain-relevant, interference-resistant, discriminative features from sparse visual tokens. Contribution/Results: ViTP establishes the first end-to-end, domain-specialized foundation model capable of reasoning-driven perceptual optimization. Evaluated on 16 remote sensing and medical imaging benchmarks, it achieves state-of-the-art performance, significantly improving generalization and semantic understanding across downstream tasks.

Technology Category

Application Category

📝 Abstract
Modern computer vision is converging on a closed loop in which perception, reasoning and generation mutually reinforce each other. However, this loop remains incomplete: the top-down influence of high-level reasoning on the foundational learning of low-level perceptual features is not yet underexplored. This paper addresses this gap by proposing a new paradigm for pretraining foundation models in downstream domains. We introduce Visual insTruction Pretraining (ViTP), a novel approach that directly leverages reasoning to enhance perception. ViTP embeds a Vision Transformer (ViT) backbone within a Vision-Language Model and pretrains it end-to-end using a rich corpus of visual instruction data curated from target downstream domains. ViTP is powered by our proposed Visual Robustness Learning (VRL), which compels the ViT to learn robust and domain-relevant features from a sparse set of visual tokens. Extensive experiments on 16 challenging remote sensing and medical imaging benchmarks demonstrate that ViTP establishes new state-of-the-art performance across a diverse range of downstream tasks. The code is available at github.com/zcablii/ViTP.
Problem

Research questions and friction points this paper is trying to address.

Enhancing low-level perception through high-level reasoning integration
Addressing incomplete top-down influence in vision learning loops
Pretraining domain-specific foundation models with visual instruction data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pretrains Vision Transformer with visual instruction data
Uses Visual Robustness Learning for domain-relevant features
Embeds ViT backbone within Vision-Language Model end-to-end
🔎 Similar Papers
No similar papers found.
Y
Yuxuan Li
PCA Lab, VCIP, Computer Science, NKU
Y
Yicheng Zhang
PCA Lab, VCIP, Computer Science, NKU
W
Wenhao Tang
PCA Lab, VCIP, Computer Science, NKU
Y
Yimian Dai
PCA Lab, VCIP, Computer Science, NKU
Ming-Ming Cheng
Ming-Ming Cheng
Professor of Computer Science, Nankai University
Computer VisionComputer GraphicsVisual AttentionSaliency
X
Xiang Li
PCA Lab, VCIP, Computer Science, NKU; NKIARI, Futian, Shenzhen
J
Jian Yang
PCA Lab, VCIP, Computer Science, NKU