INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

๐Ÿ“… 2026-04-20
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the high annotation cost in compositional image retrieval, which induces two types of noise: cross-modal correspondence noise and modality-intrinsic noise. To tackle this challenge, the paper proposes the INTENT framework, which is the first to systematically distinguish and jointly mitigate both noise sources. Specifically, INTENT employs fast Fourier transform (FFT) to implement visual causal intervention, constructing modality-invariant representations that suppress intrinsic noise. Concurrently, it introduces a sample fidelityโ€“based dynamic decision boundary to co-optimize a dual-objective learning scheme, thereby enhancing robustness in cross-modal alignment. Experiments on two mainstream benchmarks demonstrate that INTENT significantly outperforms existing methods, achieving notable advances in both retrieval performance and noise robustness.

Technology Category

Application Category

๐Ÿ“ Abstract
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are correctly matched. However, in real-world scenarios, due to high triplet annotation costs, CIR datasets inevitably contain annotation errors, resulting in incorrectly matched triplets. To address this issue, the problem of Noisy Triplet Correspondence (NTC) has attracted growing attention. We argue that noise in CIR can be categorized into two types: cross-modal correspondence noise and modality-inherent noise. The former arises from mismatches across modalities, whereas the latter originates from intra-modal background interference or visual factors irrelevant to the coarse-grained modification annotations. However, modality-inherent noise is often overlooked, and research on cross-modal correspondence noise remains nascent. To tackle above issues, we propose the Invariance and discrimiNaTion-awarE Noise neTwork (INTENT), comprising two components: Visual Invariant Composition and Bi-Objective Discriminative Learning, specifically designed to handle the two-aspect noise. The former applies causal intervention on the visual side via Fast Fourier Transform (FFT) to generate intervened composed features, enforcing visual invariance and enabling the model to ignore modality-inherent noise during composition. The latter adopts collaborative optimization with both positive and negative samples, and constructs a scalable decision boundary that dynamically adjusts decisions based on the loyalty degree, enabling robust correspondence discrimination. Extensive experiments on two widely used benchmark datasets demonstrate the superiority and robustness of INTENT.
Problem

Research questions and friction points this paper is trying to address.

Composed Image Retrieval
Noisy Triplet Correspondence
Cross-modal Correspondence Noise
Modality-inherent Noise
Annotation Errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Composed Image Retrieval
Noisy Triplet Correspondence
Visual Invariance
Discriminative Learning
Causal Intervention
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Z
Zhiwei Chen
School of Software, Shandong University
Yupeng Hu
Yupeng Hu
Shandong University
Multimedia Information RetrievalData Mining and Knowledge Discovery
Z
Zhiheng Fu
School of Software, Shandong University
Z
Zixu Li
School of Software, Shandong University
J
Jiale Huang
School of Software, Shandong University
Q
Qinlei Huang
School of Software, Shandong University
Yinwei Wei
Yinwei Wei
Shandong University | National University of Singapore
Multimedia ComputingInformation RetrievalRecommender System