Vision-Guided Text Prompt Tuning for Multimodal Sentiment Analysis

๐Ÿ“… 2026-09-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้’ˆๅฏนๅคšๆจกๆ€ๆƒ…ๆ„Ÿๅˆ†ๆžไธญ็š„ๆ–‡ๆœฌไธŽ่ง†่ง‰ไฟกๆฏ่žๅˆ้—ฎ้ข˜๏ผŒๆๅ‡บไบ†ไธ€็ง่ง†่ง‰ๅผ•ๅฏผ็š„ๆ–‡ๆœฌๆ็คบ่ฐƒไผ˜ๆ–นๆณ•(VG-TPT)๏ผŒ้€š่ฟ‡ๅฏๆŽง็š„่ง†่ง‰ๆ กๅ‡†ๅ†ป็ป“ๆ–‡ๆœฌ่กจ็คบ๏ผŒๆœ‰ๆ•ˆๆ้ซ˜ไบ†ๆƒ…ๆ„Ÿๅˆ†ๆžๆ€ง่ƒฝใ€‚
๐Ÿ“ Abstract
Multimodal sentiment analysis requires effective modeling of both verbal semantics and non-verbal affective cues. A central challenge is to calibrate text-centered sentiment understanding with visual facial evidence in a controlled, adaptive, and parameter-efficient manner. Text usually serves as the semantic anchor, whereas visual cues provide complementary evidence for ambiguous or implicit expressions; however, indiscriminate fusion may introduce visual noise and distort textual semantics. Moreover, fully fine-tuning large visual and textual encoders is costly and prone to overfitting on limited and scenario-dependent MSA benchmarks. To address these issues, we propose Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations. VG-TPT injects visual affective cues into a frozen BERT encoder through layer-wise adaptive prompts, rather than relying on late-stage feature fusion or full backbone tuning. A co-guided router composes prompts from a trainable prompt bank according to both the evolving text state and the visual guidance feature, enabling sample-specific and layer-specific modulation. Experiments on CMU-MOSEI and CMU-MOSI show that VG-TPT consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods, while updating only 2.4M trainable parameters. The code is available at https://github.com/ma-tubu/VG-TPT.
Problem

Research questions and friction points this paper is trying to address.

multimodal sentiment analysis
textual semantics
visual cues
parameter-efficient
overfitting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Guided Text Prompt Tuning
layer-wise adaptive prompts
co-guided router
frozen BERT encoder
multimodal sentiment analysis
๐Ÿ”Ž Similar Papers
X
Xiaoran Kou
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
J
Jingyi Wu
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai 200433, China
Peng Sun
Peng Sun
Duke Kunshan University
Computer NetworksAlgorithmsWireless Sensor NetworksUnderwater NetworksIntelligent Transportation Systems
Y
Yang Liu
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
H
Hong Chen
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China