MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing medical vision-language models struggle with precise localization, while segmentation models rely on explicit class labels or spatial prompts, leading to a mismatch in supervisory signals between the two tasks. To address this, this work proposes MedPixel, a unified medical pixel-language model that jointly learns multiple tasks through a shared language-to-mask interface. The approach introduces a novel clinical-driven synthetic data generation strategy that eliminates the need for annotations from external large models. Furthermore, it designs a mask-grounded, pixel-level preference optimization mechanism to unify explicit localization and implicit reasoning. MedPixel achieves strong performance on both pixel prediction and text generation tasks, demonstrates zero-shot transferability to external localization benchmarks, and exhibits robustness against imperfect spatial prompts.
📝 Abstract
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language--mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at https://github.com/yhy-whu/Medpixel.
Problem

Research questions and friction points this paper is trying to address.

medical vision-language models
pixel-level grounding
segmentation
language supervision
spatial annotations
Innovation

Methods, ideas, or system contributions that make the work stand out.

pixel-language model
medical image segmentation
vision-language grounding
preference optimization
zero-shot transfer
🔎 Similar Papers
No similar papers found.
H
Haoyu Yang
Zhejiang University
M
Meixing Shi
Zhejiang University
Z
Zengjie Chen
Zhejiang University
H
Haoran Sun
Fudan University
H
Haitao Leng
Kuaishou
Xiaoming Shi
Xiaoming Shi
East China Normal University
Large Language ModelDialogue SystemsNatural Language Processing
Y
Yuxiang Cai
Zhejiang University
Yankai Jiang
Yankai Jiang
Shanghai AI Laboratory
Multimodal LLMVision-Language PretrainingAI for Science