Institution profile

State Key Laboratory of Multimodal Artificial Intelligence Systems

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

DOME: Learning Transferable Domain Variables from Sparse Supervision for Test-Time Adaptation

Jun 01, 2026

This work addresses the limitations of existing test-time adaptation (TTA) methods, which typically assume a single global domain distribution and overlook sample-level, multi-dimensional domain shifts, leading to fragile adaptation. To overcome this, the authors propose DOME, a novel framework that, for the first time, explicitly models continuous sample-level domain variables in a zero-shot manner. DOME leverages vision-language pretraining to extract dense domain representations and employs a momentum-updated sparse domain bank to provide decoupled supervision, subsequently injecting explicit domain information into the downstream model. By moving beyond implicit global domain assumptions, DOME enables structured domain representation. It achieves state-of-the-art performance on ImageNet-C, ImageNet-R, and ImageNet-Sketch, significantly outperforming existing TTA approaches—including more complex ones—and demonstrates the efficacy and robustness of explicit domain modeling.

0 citationsRead paper

Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views

Nov 11, 2025

Existing zero-shot large language model (LLM) approaches for 3D scene understanding suffer from low accuracy and poor efficiency. To address this, we propose a training-free hierarchical 3D scene understanding framework that operates solely on sparse RGB views. Our method constructs a plane-augmented scene graph—using dominant planar surfaces as spatial anchors—and introduces a task-adaptive dynamic subgraph extraction mechanism to enable open-vocabulary hierarchical parsing and robust reasoning. Technically, it integrates pre-trained LLMs with multi-view geometric awareness, spatial relation modeling, and dynamic context filtering. On Space3D-Bench, our method achieves a 28.7% improvement in EM@1 and 78.2% inference speedup; on ScanQA, it matches supervised methods, demonstrating strong generalization and robustness. Key innovations include a plane-guided hierarchical graph structure and a task-driven lightweight inference paradigm.

0 citationsRead paper

MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction

Aug 09, 2025

Current AI-based weather forecasting systems face three key limitations: scarcity of extreme-event samples, insufficient alignment between meteorological data and textual warning messages, and the inability of multimodal models to effectively capture four-dimensional (4D) spatiotemporal–vertical meteorological dependencies. To address these challenges, we propose the first 4D meteorological multimodal large language model specifically designed for severe weather prediction. We introduce MP-Bench, a large-scale spatiotemporal multimodal benchmark, and design three plug-and-play adaptive fusion modules enabling end-to-end joint modeling of meteorological fields and textual warnings. Our model processes raw 4D meteorological data directly and supports full-parameter fine-tuning and end-to-end training. On MP-Bench, our approach achieves significant improvements in extreme weather detection, localization, and warning generation—advancing the development of fully automated AI-powered weather forecasting systems.

0 citationsRead paper

EvoVLMA: Evolutionary Vision-Language Model Adaptation

Aug 02, 2025

Existing vision-language model (VLM) adaptation methods—such as prompt tuning and adapter-based approaches—rely on manual design, suffering from low efficiency and lacking automated architecture search mechanisms. Method: This paper proposes the first automated algorithm design framework for zero-training VLM adaptation, featuring a two-stage LLM-assisted evolutionary algorithm that jointly optimizes feature selection and logits computation structure. To enable scalable exploration of large search spaces, it incorporates low-precision code translation and web-based execution monitoring. Contribution/Results: It is the first work to apply evolutionary algorithms to automatically construct VLM adaptation pipelines. Evaluated on 8-shot image classification, the discovered algorithm outperforms the manually designed APE method by +1.91 percentage points, demonstrating both effectiveness and state-of-the-art performance.

0 citationsRead paper
Recent publications

Latest Papers

DOME: Learning Transferable Domain Variables from Sparse Supervision for Test-Time Adaptation

Jun 01, 2026

This work addresses the limitations of existing test-time adaptation (TTA) methods, which typically assume a single global domain distribution and overlook sample-level, multi-dimensional domain shifts, leading to fragile adaptation. To overcome this, the authors propose DOME, a novel framework that, for the first time, explicitly models continuous sample-level domain variables in a zero-shot manner. DOME leverages vision-language pretraining to extract dense domain representations and employs a momentum-updated sparse domain bank to provide decoupled supervision, subsequently injecting explicit domain information into the downstream model. By moving beyond implicit global domain assumptions, DOME enables structured domain representation. It achieves state-of-the-art performance on ImageNet-C, ImageNet-R, and ImageNet-Sketch, significantly outperforming existing TTA approaches—including more complex ones—and demonstrates the efficacy and robustness of explicit domain modeling.

0 citationsRead paper

Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views

Nov 11, 2025

Existing zero-shot large language model (LLM) approaches for 3D scene understanding suffer from low accuracy and poor efficiency. To address this, we propose a training-free hierarchical 3D scene understanding framework that operates solely on sparse RGB views. Our method constructs a plane-augmented scene graph—using dominant planar surfaces as spatial anchors—and introduces a task-adaptive dynamic subgraph extraction mechanism to enable open-vocabulary hierarchical parsing and robust reasoning. Technically, it integrates pre-trained LLMs with multi-view geometric awareness, spatial relation modeling, and dynamic context filtering. On Space3D-Bench, our method achieves a 28.7% improvement in EM@1 and 78.2% inference speedup; on ScanQA, it matches supervised methods, demonstrating strong generalization and robustness. Key innovations include a plane-guided hierarchical graph structure and a task-driven lightweight inference paradigm.

0 citationsRead paper

MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction

Aug 09, 2025

Current AI-based weather forecasting systems face three key limitations: scarcity of extreme-event samples, insufficient alignment between meteorological data and textual warning messages, and the inability of multimodal models to effectively capture four-dimensional (4D) spatiotemporal–vertical meteorological dependencies. To address these challenges, we propose the first 4D meteorological multimodal large language model specifically designed for severe weather prediction. We introduce MP-Bench, a large-scale spatiotemporal multimodal benchmark, and design three plug-and-play adaptive fusion modules enabling end-to-end joint modeling of meteorological fields and textual warnings. Our model processes raw 4D meteorological data directly and supports full-parameter fine-tuning and end-to-end training. On MP-Bench, our approach achieves significant improvements in extreme weather detection, localization, and warning generation—advancing the development of fully automated AI-powered weather forecasting systems.

0 citationsRead paper

EvoVLMA: Evolutionary Vision-Language Model Adaptation

Aug 02, 2025

Existing vision-language model (VLM) adaptation methods—such as prompt tuning and adapter-based approaches—rely on manual design, suffering from low efficiency and lacking automated architecture search mechanisms. Method: This paper proposes the first automated algorithm design framework for zero-training VLM adaptation, featuring a two-stage LLM-assisted evolutionary algorithm that jointly optimizes feature selection and logits computation structure. To enable scalable exploration of large search spaces, it incorporates low-precision code translation and web-based execution monitoring. Contribution/Results: It is the first work to apply evolutionary algorithms to automatically construct VLM adaptation pipelines. Evaluated on 8-shot image classification, the discovered algorithm outperforms the manually designed APE method by +1.91 percentage points, demonstrating both effectiveness and state-of-the-art performance.

0 citationsRead paper