Institution profile

DreamSports

Industry research
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts

May 22, 2026

This study addresses the challenge of person identity matching across heterogeneous records characterized by linguistic and cultural complexity, diverse naming conventions, and high data noise. To tackle this problem, the authors propose the Structure-Guided Entity Resolution (SGER) framework, which introduces a novel two-stage curriculum fine-tuning strategy: first guiding a large language model to learn the syntactic and semantic structures of personal names, followed by optimizing it for binary entity matching. Evaluated on 50,000 real-world Indian identity record pairs, SGER achieves 99.02% accuracy and an F1 score of 0.994, significantly outperforming few-shot prompting with GPT-4o and single-stage fine-tuning baselines. The method has been deployed on the Dream11 platform, serving over 250 million users, and demonstrates enhanced robustness and precision in multilingual, high-noise entity resolution scenarios.

0 citationsRead paper

LUMOS: Large User MOdels for User Behavior Prediction

Nov 28, 2025

To address poor generalizability and scalability in large-scale user behavior prediction on online B2C platforms—caused by heavy reliance on manual feature engineering, task-specific models, and domain expertise—this paper proposes the first unified large model for massive-user behavioral forecasting. The model operates solely on raw user behavioral sequences, employing multimodal tokenization to jointly encode temporal actions, event contexts, and static demographic attributes. It introduces a novel conditional cross-attention mechanism explicitly designed for future events (e.g., promotions, holidays), enabling causal reasoning. Built upon the Transformer architecture, it adopts end-to-end joint multi-task learning. Evaluated on real-world data comprising 275 billion tokens and 250 million users, the model achieves an average ROC-AUC improvement of 0.025 and a 4.6% reduction in MAPE across five prediction tasks. Online A/B testing demonstrates a 3.15% increase in daily active users.

0 citationsRead paper
Recent publications

Latest Papers

Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts

May 22, 2026

This study addresses the challenge of person identity matching across heterogeneous records characterized by linguistic and cultural complexity, diverse naming conventions, and high data noise. To tackle this problem, the authors propose the Structure-Guided Entity Resolution (SGER) framework, which introduces a novel two-stage curriculum fine-tuning strategy: first guiding a large language model to learn the syntactic and semantic structures of personal names, followed by optimizing it for binary entity matching. Evaluated on 50,000 real-world Indian identity record pairs, SGER achieves 99.02% accuracy and an F1 score of 0.994, significantly outperforming few-shot prompting with GPT-4o and single-stage fine-tuning baselines. The method has been deployed on the Dream11 platform, serving over 250 million users, and demonstrates enhanced robustness and precision in multilingual, high-noise entity resolution scenarios.

0 citationsRead paper

LUMOS: Large User MOdels for User Behavior Prediction

Nov 28, 2025

To address poor generalizability and scalability in large-scale user behavior prediction on online B2C platforms—caused by heavy reliance on manual feature engineering, task-specific models, and domain expertise—this paper proposes the first unified large model for massive-user behavioral forecasting. The model operates solely on raw user behavioral sequences, employing multimodal tokenization to jointly encode temporal actions, event contexts, and static demographic attributes. It introduces a novel conditional cross-attention mechanism explicitly designed for future events (e.g., promotions, holidays), enabling causal reasoning. Built upon the Transformer architecture, it adopts end-to-end joint multi-task learning. Evaluated on real-world data comprising 275 billion tokens and 250 million users, the model achieves an average ROC-AUC improvement of 0.025 and a 4.6% reduction in MAPE across five prediction tasks. Online A/B testing demonstrates a 3.15% increase in daily active users.

0 citationsRead paper