PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对CTR模型中子群优化竞争问题,提出PRIME方法,通过基于输入条件的专家混合和低秩残差专家组合,在保持原有预测功能的同时增强了模型性能。
📝 Abstract
Click-through rate (CTR) models vary in feature-interaction design, yet their top networks usually remain a single multilayer perceptron shared by all examples. Heterogeneous user, item, and context subgroups therefore update the same parameters; weakly aligned learning signals make the aggregate gradient a compromise among competing directions. We study the competition on Avazu with 4 models and 4 semantic fields. Across all architectures, semantic subgroups show lower Top-NN gradient cosine similarity than random groups matched by sample size and label ratio, with reductions of 0.23-0.37. This competition motivates input-conditioned experts, but directly replacing an established Dense mapping changes its initial function, sharing pattern, and capacity, obscuring the source of gains. We introduce PRIME (Plug-in Residual Input-conditioned Mixture of Experts), a Dense-anchored mixture of low-rank residual experts. PRIME anchors the original prediction and uses zero-residual initialization to match the Dense baseline exactly at training onset. Input-dependent routing weights low-rank experts for example-specific logit corrections; multi-bag aggregation and EMA load biases stabilize conditional estimation. We evaluate PRIME on held-out Avazu and Criteo test sets across 13 CTR architectures and five paired seeds. Median paired AUC gains are +0.0022 and +0.0066, with LogLoss reductions of 0.0011 and 0.0081, respectively. On FiBiNET and DCNv2, PRIME outperforms APG in all ten seed-level AUC comparisons while using fewer parameters and lower inference latency on both backbones. These results show that function-preserving conditional residuals add input-dependent capacity while preserving the Dense path and its optimization stability. Code is available at https://github.com/YH-learning/PRIME.
Problem

Research questions and friction points this paper is trying to address.

CTR models
subgroup optimization competition
shared parameters
learning signals
gradient compromise
Innovation

Methods, ideas, or system contributions that make the work stand out.

PRIME
Input-Conditioned Mixture of Experts
Zero-Residual Initialization
Conditional Estimation Stabilization
💼 Related Jobs
No related jobs found.
Heng Yao
Heng Yao
University of Shanghai for Science and Technology
Digital image forensicsInformation HidingImage Processing
S
Siyun Hou
Henan Polytechnic University
T
Tianying Liu
Independent Researcher
Y
Yulou Shu
Alibaba Inc.
Y
Yong He
Ant Group
C
Chuan Yuan
Ant Group
K
Kaibin Qiu
Ant Group
G
Guowei Chen
Ant Group
J
Jiayu Zhao
Ant Group
C
Chao Yu
Ant Group
K
Ke Ding
Ant Group