Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
This work addresses the performance bottleneck in strict zero-shot image captioning caused by the absence of visual feedback during inference. The authors propose a multi-agent framework that enhances vision-language alignment without retraining existing captioners, leveraging multi-stage alignment scoring and unsupervised consensus distillation. A novel multi-checkpoint visual feedback mechanism is introduced during decoding, accompanied by a lightweight learnable reranking module that integrates TriFuse and MemAttend architectures. The approach further incorporates Borda count–based consensus distillation to refine caption selection. Experimental results demonstrate significant improvements: the method achieves a CIDEr score of 117.6 on the COCO Karpathy test set, surpassing the baseline by 9.6 points, and yields gains of 8.1 and 5.7 on Flickr30k and NoCaps, respectively.