Reasoning-Guided Claim Normalization for Noisy Multilingual Social Media Posts

📅 2025-11-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenge of claim normalization in multilingual social media posts—transforming noisy, unstructured user-generated content into clear, verifiable statements to support multilingual misinformation detection—without requiring multilingual annotated data. Methodologically, it introduces a cross-lingual claim standardization framework based on a question-answering decomposition schema (Who/What/Where/When/Why/How), trained exclusively on English data. The approach integrates Qwen3-14B with LoRA fine-tuning and incorporates intra-post deduplication, token-level recall filtering, and retrieval-augmented few-shot inference. Experiments span 20 languages, achieving a peak METEOR score of 41.16—outperforming baseline methods by an average of 41.3%. Results demonstrate strong cross-lingual generalization and practical efficacy, significantly advancing state-of-the-art performance in zero-shot multilingual claim normalization.

Technology Category

Application Category

📝 Abstract
We address claim normalization for multilingual misinformation detection - transforming noisy social media posts into clear, verifiable statements across 20 languages. The key contribution demonstrates how systematic decomposition of posts using Who, What, Where, When, Why and How questions enables robust cross-lingual transfer despite training exclusively on English data. Our methodology incorporates finetuning Qwen3-14B using LoRA with the provided dataset after intra-post deduplication, token-level recall filtering for semantic alignment and retrieval-augmented few-shot learning with contextual examples during inference. Our system achieves METEOR scores ranging from 41.16 (English) to 15.21 (Marathi), securing third rank on the English leaderboard and fourth rank for Dutch and Punjabi. The approach shows 41.3% relative improvement in METEOR over baseline configurations and substantial gains over existing methods. Results demonstrate effective cross-lingual generalization for Romance and Germanic languages while maintaining semantic coherence across diverse linguistic structures.
Problem

Research questions and friction points this paper is trying to address.

Transforming noisy social media posts into clear verifiable statements across 20 languages
Enabling robust cross-lingual transfer using systematic post decomposition questions
Achieving effective multilingual misinformation detection despite English-only training data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematic decomposition using six W-questions for cross-lingual transfer
Fine-tuning Qwen3-14B with LoRA after deduplication and filtering
Retrieval-augmented few-shot learning with contextual examples during inference
🔎 Similar Papers
No similar papers found.
Manan Sharma
Manan Sharma
TIFIN
Arya Suneesh
Arya Suneesh
NLP, TIFIN India
natural language processingquantum computing
M
Manish Jain
TIFIN
P
Pawan Rajpoot
TIFIN
Prasanna Devadiga
Prasanna Devadiga
R&D, TIFIN India
MLStatistics
B
Bharatdeep Hazarika
TIFIN
Ashish Shrivastava
Ashish Shrivastava
TIFIN
K
Kishan Gurumurthy
TIFIN
A
Anshuman B Suresh
TIFIN
A
Aditya U Baliga
TIFIN