Institution profile

Qilu University of Technology

Academic institutionasia · cn
Official website
Research library143linked papers
Opportunities0open roles
Selected work

Representative Papers

Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model

Jan 20, 2026

This work addresses the limitations of existing out-of-distribution (OOD) detection methods, which overly rely on textual features and struggle with distributional shifts in the visual domain—particularly under both near- and far-OOD scenarios. To overcome this, the authors propose MM-OOD, a novel framework that systematically leverages the multimodal reasoning and generative capabilities of large multimodal language models (MLLMs). For near-OOD samples, MM-OOD performs zero-shot inference by jointly utilizing image and text prompts. For far-OOD cases, it introduces a three-stage “sketch–generate–refine” pipeline that enhances multimodal prompting through generated visual exemplars. By moving beyond conventional unimodal or text-only paradigms, MM-OOD achieves state-of-the-art performance on multimodal benchmarks such as Food-101 and demonstrates strong scalability on ImageNet-1K.

1 citationsRead paper

Imaging foundation model for universal enhancement of non-ideal measurement CT

Oct 02, 2024arXiv.org

Non-ideal computed tomography (NICT)—characterized by suboptimal acquisition protocols—suffers from degraded image quality and low clinical acceptance, while existing deep learning methods rely heavily on large-scale annotated datasets and exhibit poor generalizability. To address these challenges, we propose TAMP, the first foundation model tailored for NICT. Its core contributions are: (1) a multi-scale integrated Transformer amplifier that jointly incorporates physical priors and data-driven modeling; (2) a physics-informed large-scale synthetic pretraining paradigm, leveraging 10.8 million simulated scans to learn robust representations across diverse protocols, anatomical regions, and noise levels; and (3) parameter-efficient fine-tuning via LoRA-style adaptation requiring only a few slices. Extensive evaluation demonstrates significant PSNR/SSIM improvements across multiple NICT tasks. Clinical validation—including radiologist-blinded assessment and real-world deployment—confirms markedly enhanced diagnostic acceptability, underscoring TAMP’s readiness for clinical translation.

1 citationsRead paper
Recent publications

Latest Papers

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

Aug 17, 2026

This study addresses the challenge of attribute misbinding in large vision-language models within dense homogeneous scenes, where existing metrics prove inadequate. We formally define the DSCAM task and construct InstaBind-Lite, a controlled benchmark accompanied by a specialized evaluation framework. Through fine-grained instance annotation and multi-level question-answering design, this work enables quantitative assessment of attribute transfer while exposing critical blind spots in traditional evaluations. Experiments reveal misbinding rates of 19.84% for open-source models and 7.55% for API-based models, precisely localizing error sources. Ultimately, this research establishes a novel, traceable evaluation paradigm for assessing fine-grained attribute binding capabilities in large multimodal models.

0 citationsRead paper