Aesthetic Image Captioning with Saliency Enhanced MLLMs

📅 2025-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing aesthetic image captioning (AIC) methods primarily rely on fine-tuning general-purpose multimodal large language models (MLLMs), but they lack explicit modeling of aesthetic saliency, leading to insufficient attention to aesthetic content. Method: We propose the first end-to-end framework that explicitly integrates aesthetic saliency into MLLMs. It introduces an Image Aesthetic Saliency Module (IASM) to extract fine-grained aesthetic features and constructs an IAS-ViT encoder based on cross-attention to deeply fuse these features with visual representations. Contribution/Results: Our method requires no task-agnostic pretraining or additional annotations. It achieves state-of-the-art performance on major AIC benchmarks, significantly outperforming prior approaches. Experimental results validate that explicit aesthetic modeling is critical for improving both the accuracy and expressiveness of generated captions.

Technology Category

Application Category

📝 Abstract
Aesthetic Image Captioning (AIC) aims to generate textual descriptions of image aesthetics, becoming a key research direction in the field of computational aesthetics. In recent years, pretrained Multimodal Large Language Models (MLLMs) have advanced rapidly, leading to a significant increase in image aesthetics research that integrates both visual and textual modalities. However, most existing studies on image aesthetics primarily focus on predicting aesthetic ratings and have shown limited application in AIC. Existing AIC works leveraging MLLMs predominantly rely on fine-tuning methods without specifically adapting MLLMs to focus on target aesthetic content. To address this limitation, we propose the Aesthetic Saliency Enhanced Multimodal Large Language Model (ASE-MLLM), an end-to-end framework that explicitly incorporates aesthetic saliency into MLLMs. Within this framework, we introduce the Image Aesthetic Saliency Module (IASM), which efficiently and effectively extracts aesthetic saliency features from images. Additionally, we design IAS-ViT as the image encoder for MLLMs, this module fuses aesthetic saliency features with original image features via a cross-attention mechanism. To the best of our knowledge, ASE-MLLM is the first framework to integrate image aesthetic saliency into MLLMs specifically for AIC tasks. Extensive experiments demonstrated that our approach significantly outperformed traditional methods and generic MLLMs on current mainstream AIC benchmarks, achieving state-of-the-art (SOTA) performance.
Problem

Research questions and friction points this paper is trying to address.

Generating aesthetic textual descriptions for images
Integrating aesthetic saliency into multimodal language models
Improving aesthetic captioning beyond rating prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Aesthetic Saliency Enhanced MLLM framework
Image Aesthetic Saliency Module extracts features
Cross-attention fuses aesthetic and original features
Y
Yilin Tao
UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, Beijing 100049, China
J
Jiashui Huang
AI Lab, Terminus International, Terminus Group, Chongqing 400000, China
H
Huaze Xu
AI Lab, Terminus International, Terminus Group, Chongqing 400000, China
L
Ling Shao
UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, Beijing 100049, China