Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in zero-shot image captioning where synthetic training data generated by text-to-image models often suffers from fine-grained entity misalignment—such as missing objects or mislocalized attributes—leading to distorted supervision signals. To mitigate this, the authors propose ReCap, a framework that explicitly detects image entities and guides caption rewriting to achieve fine-grained image-text alignment. ReCap further incorporates an adaptive dynamic weighting strategy to downweight unreliable synthetic samples during training. By shifting data refinement from implicit global matching to explicit entity-level realignment, the method introduces a plug-and-play mechanism for fine-grained correction. Experiments demonstrate that ReCap achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks, significantly improving caption consistency and supervision fidelity.
📝 Abstract
Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to synthesize training data from text-only corpora, but most focus on improving overall data quality. In contrast, we observe that synthetic image-text misalignment is often structured and fine-grained: pairs may remain globally plausible while containing missing entities or misgrounded attributes, thereby degrading supervision fidelity. As a result, methods based on global similarity for image rematching or regeneration may improve apparent plausibility, but cannot systematically repair entity-level misalignment. To address this issue, we propose ReCap, a plug-and-play framework that shifts synthetic data refinement from implicit global matching to explicit fine-grained realignment. Specifically, ReCap enforces entity-level correspondence by using detected image-supported entities to guide caption rewriting, yielding more faithful synthetic supervision. In addition, we introduce an adaptive dynamic weighted learning strategy to downweight unreliable synthetic pairs during training. As a general framework, ReCap can be integrated into existing synthetic-data pipelines. Extensive experiments show that ReCap consistently improves image-text consistency and achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks.
Problem

Research questions and friction points this paper is trying to address.

zero-shot image captioning
synthetic supervision
entity-level misalignment
image-text alignment
fine-grained realignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

entity-faithful repair
zero-shot image captioning
synthetic supervision
fine-grained realignment
adaptive weighted learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhiyue Liu
School of Computer, Electronics and Information, Guangxi Key Laboratory of Multimedia Communications and Network Technology, Guangxi University, Nanning, Guangxi, China
W
Wenkai Zhou
School of Computer, Electronics and Information, Guangxi University, Nanning, Guangxi, China
J
Jian Qin
School of Computer, Electronics and Information, Guangxi University, Nanning, Guangxi, China
Q
Qipeng Jiang
School of Computer, Electronics and Information, Guangxi University, Nanning, Guangxi, China