Institution profile

Meituan

Industry researchasia · cn
Official website
Research library630linked papers
Opportunities0open roles
Selected work

Representative Papers

Vertical Semi-Federated Learning for Efficient Online Advertising

Sep 30, 2022arXiv.org

Traditional vertical federated learning (VFL) is constrained by the sample-overlap assumption and incurs high real-time inference overhead, rendering it ill-suited for online advertising. To address these limitations, this paper proposes Semi-VFL—a novel vertical semi-federated learning paradigm that eliminates the requirement for sample overlap and enables modeling over the full sample space. We design a Joint Privileged Learning (JPL) framework coupled with cross-party representation distillation to jointly train on both overlapping and non-overlapping data. Furthermore, we introduce a lightweight single-party student model and cross-party feature correlation modeling to balance prediction accuracy, inference efficiency, and feasibility of localized deployment. Evaluated on real-world advertising datasets, Semi-VFL consistently outperforms state-of-the-art baselines, achieving significant improvements in AUC, queries-per-second (QPS), and deployment flexibility.

19 citations1 influentialRead paper

Your Group-Relative Advantage Is Biased

Jan 13, 2026

This work addresses a systematic bias in advantage estimation within population-based reinforcement learning, where difficult prompts are consistently underestimated while easy ones are overestimated, thereby disrupting the balance between exploration and exploitation. The study is the first to uncover the underlying mechanism of this bias and proposes a novel method—History-Aware Adaptive Difficulty Weighting (HA-DW)—which dynamically corrects advantage estimates by leveraging training dynamics and difficulty anchors. Through theoretical analysis grounded in GRPO and its variants, the effectiveness of HA-DW is empirically validated across five mathematical reasoning benchmarks, demonstrating significant performance improvements. These results underscore that correcting advantage bias is crucial for effective reinforcement learning from verifiable feedback (RLVR).

6 citationsRead paper

Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text

Jan 15, 2026

This work addresses the limitation of current large language models in autonomous tool use, which stems from a scarcity of diverse and realistic multi-turn tool interaction data. The authors propose a novel paradigm that automatically synthesizes multi-turn tool-use trajectories from general-purpose text corpora, treating natural text as a scalable source of behavioral traces for the first time. Their approach employs a four-stage pipeline—comprising relevance filtering, workflow and tool extraction, trajectory embodiment, and complexity optimization—alongside a dedicated trajectory synthesis model fine-tuned with supervised learning to enable efficient and generalizable data generation. Evaluated on the BFCL V3 multi-turn benchmark, the resulting GEM-32B model achieves a 16.5% performance gain, surpassing certain models trained on domain-specific τ-bench data while significantly reducing inference latency and computational cost.

3 citationsRead paper

MobileDreamer: Generative Sketch World Model for GUI Agent

Jan 07, 2026arXiv.org

This work addresses the limitations of existing mobile GUI agents, which predominantly rely on reactive decision-making and struggle with long-horizon tasks. To overcome this, we propose a world model framework grounded in textual sketches that predicts post-action GUI states by generating task-relevant textual descriptions. The framework incorporates an imagination-based planning mechanism to refine action selection and introduces a permutation-invariant learning strategy that preserves spatial awareness while enabling efficient state prediction. Evaluated on the Android World benchmark, our method achieves state-of-the-art performance, improving task success rate by 5.25% and accurately forecasting key GUI elements, thereby significantly enhancing the agent’s capacity for foresighted planning.

2 citationsRead paper

Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws

Feb 15, 2026

This work addresses the problem of designing optimal batch size schedules under a fixed data budget to balance optimization dynamics and computational efficiency. Building upon the function scaling law (FSL) framework and incorporating gradient noise forgetting dynamics, the study systematically investigates how task difficulty influences batch size scheduling. The analysis reveals that easy tasks benefit from steadily increasing batch sizes throughout training, whereas difficult tasks achieve better performance by switching to large batch sizes only in later stages—a strategy termed “late switching.” This approach leverages a “fast catch-up” mechanism to substantially reduce data consumption without compromising model performance. Empirical validation on dense and mixture-of-experts (MoE) large language models—trained with 1.1B parameters and 1T tokens—consistently demonstrates the superiority of the late-switching strategy over both constant batch size and early-switching baselines.

1 citationsRead paper
Recent publications

Latest Papers