Institution profile

NKIARI

Research institutionasia · jp
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Visual Instruction Pretraining for Domain-Specific Foundation Models

Sep 22, 2025

Existing vision foundation models lack a backward guidance mechanism whereby high-level reasoning informs low-level perceptual feature learning, resulting in an incomplete perception–reasoning–generation loop. Method: We propose Vision Instruction Pre-training (ViTP), a novel paradigm built upon Vision Transformers and trained end-to-end on large-scale, domain-specific vision–language instruction data via instruction following and cross-modal alignment. Crucially, we introduce Vision Robustness Learning (VRL), which explicitly encourages the model to extract domain-relevant, interference-resistant, discriminative features from sparse visual tokens. Contribution/Results: ViTP establishes the first end-to-end, domain-specialized foundation model capable of reasoning-driven perceptual optimization. Evaluated on 16 remote sensing and medical imaging benchmarks, it achieves state-of-the-art performance, significantly improving generalization and semantic understanding across downstream tasks.

0 citationsRead paper
Recent publications

Latest Papers

Visual Instruction Pretraining for Domain-Specific Foundation Models

Sep 22, 2025

Existing vision foundation models lack a backward guidance mechanism whereby high-level reasoning informs low-level perceptual feature learning, resulting in an incomplete perception–reasoning–generation loop. Method: We propose Vision Instruction Pre-training (ViTP), a novel paradigm built upon Vision Transformers and trained end-to-end on large-scale, domain-specific vision–language instruction data via instruction following and cross-modal alignment. Crucially, we introduce Vision Robustness Learning (VRL), which explicitly encourages the model to extract domain-relevant, interference-resistant, discriminative features from sparse visual tokens. Contribution/Results: ViTP establishes the first end-to-end, domain-specialized foundation model capable of reasoning-driven perceptual optimization. Evaluated on 16 remote sensing and medical imaging benchmarks, it achieves state-of-the-art performance, significantly improving generalization and semantic understanding across downstream tasks.

0 citationsRead paper