Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders
This work investigates whether user-specific modality weighting mechanisms widely adopted in multimodal recommendation systems genuinely capture individual user preferences. To this end, the authors propose an auditing framework comprising two metrics—real-GM and real-shuf—that evaluate personalization efficacy by comparing six personalized weighting methods against global weights and shuffled user-weight assignments, all under a unified collaborative filtering backbone. Experimental results across three short-video and one cross-domain e-commerce dataset reveal that performance gains from most methods stem primarily from increased model capacity rather than authentic user signals, with gating mechanisms often inducing spurious personalization due to their reliance on shared embeddings. Notably, global modality weights already achieve nearly all attainable gains, while personalized weighting shows no consistent improvement; the proposed audit framework effectively identifies architectures that truly encode user-specific patterns.