Institution profile

Shell

Industry researcheurope · nl
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Multimodal Language Models Cannot Spot Spatial Inconsistencies

Apr 01, 2026

Current multimodal large language models struggle to recognize three-dimensional spatial inconsistencies across different viewpoints of the same scene. This work introduces a novel task: given a pair of images depicting the same scene from two distinct viewpoints, detect objects that violate 3D motion consistency. To facilitate research on this task, we develop a scalable synthetic framework capable of generating multiview image pairs with controllable spatial inconsistencies, and establish an evaluation protocol integrating human comparative experiments with model assessments. This study presents the first systematic evaluation of multimodal large language models’ ability to reason about 3D spatial consistency, revealing significant limitations in their understanding of physical world dynamics—state-of-the-art models perform substantially worse than humans and exhibit unstable performance across varying scene attributes.

0 citationsRead paper

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective

Jun 10, 2025

To address the challenge of vision modeling under few-shot settings (≤10 images per class) where transfer learning is inadmissible, this work proposes a data-efficient training paradigm that explicitly embeds visual inductive priors. Methodologically, it employs end-to-end training without pretraining, leveraging a hybrid Transformer-CNN architecture, aggressive data augmentation, and large-scale ensembling, while systematically designing prior-guided model architectures and regularization mechanisms. Evaluated on the VIPriors benchmark—where transfer learning is strictly prohibited across four editions—the approach achieves significant gains over baselines under extremely low-data regimes. Key contributions include: (i) establishing the first reproducible, “from-scratch” few-shot vision benchmark; (ii) empirically demonstrating that explicit modeling of inductive priors critically alleviates data hunger; and (iii) providing a principled, theoretically grounded, and engineering-practical framework for data-constrained vision learning.

0 citationsRead paper
Recent publications

Latest Papers

Multimodal Language Models Cannot Spot Spatial Inconsistencies

Apr 01, 2026

Current multimodal large language models struggle to recognize three-dimensional spatial inconsistencies across different viewpoints of the same scene. This work introduces a novel task: given a pair of images depicting the same scene from two distinct viewpoints, detect objects that violate 3D motion consistency. To facilitate research on this task, we develop a scalable synthetic framework capable of generating multiview image pairs with controllable spatial inconsistencies, and establish an evaluation protocol integrating human comparative experiments with model assessments. This study presents the first systematic evaluation of multimodal large language models’ ability to reason about 3D spatial consistency, revealing significant limitations in their understanding of physical world dynamics—state-of-the-art models perform substantially worse than humans and exhibit unstable performance across varying scene attributes.

0 citationsRead paper

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective

Jun 10, 2025

To address the challenge of vision modeling under few-shot settings (≤10 images per class) where transfer learning is inadmissible, this work proposes a data-efficient training paradigm that explicitly embeds visual inductive priors. Methodologically, it employs end-to-end training without pretraining, leveraging a hybrid Transformer-CNN architecture, aggressive data augmentation, and large-scale ensembling, while systematically designing prior-guided model architectures and regularization mechanisms. Evaluated on the VIPriors benchmark—where transfer learning is strictly prohibited across four editions—the approach achieves significant gains over baselines under extremely low-data regimes. Key contributions include: (i) establishing the first reproducible, “from-scratch” few-shot vision benchmark; (ii) empirically demonstrating that explicit modeling of inductive priors critically alleviates data hunger; and (iii) providing a principled, theoretically grounded, and engineering-practical framework for data-constrained vision learning.

0 citationsRead paper