HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of inaccurate target identification and absent evidence localization in harmful meme detection by constructing the Meme3W dataset and introducing the JRA metric. We propose an Anchor-Calibrated Decoupled Optimization framework that integrates entity-aware fine-tuning, conditional policy optimization, and a virtual positive anchor mechanism to jointly enhance fine-grained target recognition and harm classification. Experimental results demonstrate that this approach significantly improves performance on the Qwen3-VL-8B model, with the JRA metric increasing from 17.58% to 52.51% alongside substantial gains in harm detection accuracy. These findings effectively advance research in multimodal safety alignment by providing a robust solution for identifying and localizing harmful content within internet memes.
📝 Abstract
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked target or its supporting evidence. We therefore extend harmful meme detection with fine-grained target identification, asking what type of target is attacked, who is targeted, and where the target appears in the meme. The model predicts harmfulness for every meme and, for harmful memes, outputs the target category, target entity, textual mention, and visual region. To support this task, we introduce Meme3W, which unifies multiple public harmful meme datasets and provides human-verified annotations for harmful instances. We further introduce Joint Record Accuracy (JRA), a strict record-level metric requiring the harmfulness label and all target-identification fields to be jointly correct. Experiments with representative multimodal large language models reveal a substantial gap between harmfulness accuracy and JRA. To narrow this gap, we propose HarmTrace, an anchor-calibrated decoupled optimization framework. HarmTrace strengthens target-entity supervision through entity-aware supervised fine-tuning. It then applies Conditional Target-identification Policy Optimization (CTPO) to decouple harmfulness and target-identification advantages, restricting target-identification optimization to label-correct responses for harmful examples. CTPO uses a Virtual Positive Anchor (VPA) as a fully correct reference for target-identification advantage normalization. HarmTrace improves both JRA and harmfulness accuracy across the evaluated backbones, with JRA on the Qwen3-VL-8B backbone increasing from 17.58\% to 52.51\%. Our code is publicly available at https://github.com/llly1234/HarmTrace-for-Harmful-Memes.
Problem

Research questions and friction points this paper is trying to address.

Harmful Meme Detection
Fine-Grained Target Identification
Multimodal Learning
Joint Record Accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Anchor-Calibrated Decoupled Optimization
Conditional Target-identification Policy Optimization
Virtual Positive Anchor
Entity-Aware Supervised Fine-Tuning
Fine-Grained Target Identification
Yujia Li
Yujia Li
Research Scientist, Google DeepMind
Machine LearningComputer VisionNatural Language ProcessingOptimization
Y
Yiqun Zhang
Apple Inc.
Z
Zihan Cheng
School of Computer Science and Engineering, Northeastern University
Y
Yijie Huang
School of Computer Science and Engineering, Northeastern University
T
Tenglong Ye
School of Computer Science and Engineering, Northeastern University
Z
Zihan Wang
School of Computer Science and Engineering, Northeastern University
Xiaocui Yang
Xiaocui Yang
Lecturer, Northeastern University (China)
Multimodal Sentiment AnalysisData MiningMultimodal Large Language Models
S
Shi Feng
School of Computer Science and Engineering, Northeastern University
Yifei Zhang
Yifei Zhang
Institute of Information Engineering, Chinese Academy of Sciences
Computer VisionUnsupervised Learning
D
Daling Wang
School of Computer Science and Engineering, Northeastern University