From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Image-Grounded Review Assistance

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了电商平台上用户上传图片缺乏详细文字反馈的问题,通过多智能体视觉-语言框架生成基于图像的产品评价草稿。
📝 Abstract
Visual feedback in the form of user-uploaded images and videos is becoming increasingly common in e-commerce platforms because it provides authentic evidence of product quality, defects, packaging conditions, and real-world usage. However, visual feedback alone often lacks the contextual explanations and subjective opinions necessary for informed decision-making, while many users provide limited textual feedback due to the effort required to compose detailed reviews. To bridge this gap, we introduce image-grounded review assistance, a novel task that aims to generate editable review drafts from user-uploaded product images. Unlike conventional image captioning, which focuses on objective visual description, the proposed task requires product-specific understanding, sentiment estimation, and evidence-driven review composition under challenging real-world conditions, including degraded image quality, excessive zoom-in, target ambiguity, and partial product visibility. We propose a multi-agent vision-language framework consisting of four specialised roles: product grounding, visual sentiment estimation, visual evidence generation, and review synthesis. The framework employs explicit intermediate representations, including product entities, predicted ratings, and evidence summaries, to improve interpretability and visual grounding. Experiments on a curated subset of the Amazon Reviews Electronics dataset demonstrate the feasibility of generating coherent, product-aware, and sentiment-aware review drafts from visual feedback. To the best of our knowledge, this is the first study to formulate image-grounded review assistance as a multi-agent vision-language reasoning problem, providing a practical step toward AI-assisted review authoring in e-commerce systems.
Problem

Research questions and friction points this paper is trying to address.

Visual Feedback
Textual Reviews
E-commerce
Product Images
Review Assistance
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent vision-language framework
image-grounded review assistance
visual sentiment estimation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
U
Utsav Kumar Nareti
Dept. of CSE, Indian Institute of Technology Patna, India-801106
Ayush Bansal
Ayush Bansal
Dept. of CSE, Indian Institute of Technology Patna, India-801106
K
Kumari Priya
Dept. of CSE, Indian Institute of Technology Patna, India-801106
Chandranath Adak
Chandranath Adak
Indian Institute of Technology Patna
Computer VisionDeep LearningBiometricsData Analytics
S
Soumi Chattopadhyay
Dept. of CSE, Indian Institute of Technology Indore, India-453552
M
Muhammad Saqib
NCMI, CSIRO, Australia-2601
Saeed Anwar
Saeed Anwar
University of Western Australia; Australian National University
Computer Vision3D VisionMachine learningGenerative AI