Mixture of Experts based Multi-task Supervise Learning from Crowds

📅 2024-07-18
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing truth inference methods in crowdsourcing treat true labels as latent variables and model worker behavior via statistical or deep learning approaches, yet they neglect workers’ heterogeneous responses to item features—leading to modeling inaccuracies and suboptimal inference performance. This paper proposes a novel multi-task supervised learning paradigm that dispenses with the latent-truth assumption and explicitly models worker competence at the item-feature level. We introduce a dual-path framework: (i) MMLC-owf, an end-to-end truth inference model; and (ii) MMLC-df, a plug-and-play module that enhances existing methods. Technically, our approach integrates a Mixture-of-Experts (MoE) architecture, spectral clustering of workers in the embedding space, and feature-aware supervised modeling. Experiments demonstrate that MMLC-owf achieves significant improvements over state-of-the-art methods across multiple benchmarks; meanwhile, MMLC-df consistently boosts the performance of mainstream approaches, validating the effectiveness and generalizability of feature-aware modeling.

Technology Category

Application Category

📝 Abstract
Existing truth inference methods in crowdsourcing aim to map redundant labels and items to the ground truth. They treat the ground truth as hidden variables and use statistical or deep learning-based worker behavior models to infer the ground truth. However, worker behavior models that rely on ground truth hidden variables overlook workers' behavior at the item feature level, leading to imprecise characterizations and negatively impacting the quality of truth inference. This paper proposes a new paradigm of multi-task supervised learning from crowds, which eliminates the need for modeling of items's ground truth in worker behavior models. Within this paradigm, we propose a worker behavior model at the item feature level called Mixture of Experts based Multi-task Supervised Learning from Crowds (MMLC). Two truth inference strategies are proposed within MMLC. The first strategy, named MMLC-owf, utilizes clustering methods in the worker spectral space to identify the projection vector of the oracle worker. Subsequently, the labels generated based on this vector are considered as the inferred truth. The second strategy, called MMLC-df, employs the MMLC model to fill the crowdsourced data, which can enhance the effectiveness of existing truth inference methods. Experimental results demonstrate that MMLC-owf outperforms state-of-the-art methods and MMLC-df enhances the quality of existing truth inference methods.
Problem

Research questions and friction points this paper is trying to address.

Improves truth inference by modeling worker behavior at item feature level
Eliminates need for ground truth modeling in worker behavior analysis
Proposes two strategies to enhance accuracy and effectiveness of truth inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Experts for multi-task learning
Worker behavior modeling at feature level
Clustering in worker spectral space
🔎 Similar Papers
No similar papers found.
Zhejiang Gongshang University
T
Tao Han
School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310018, China
H
Huaixuan Shi
School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310018, China
X
Xinyi Ding
School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310018, China
X
Xiao Ma
School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310018, China
H
Huamao Gu
School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310018, China
Y
Yili Fang
School of Computer Science and Technology, Zhejiang Gongshang University, Hangzhou 310018, China