Institution profile

Renmin University of China

Academic institutionasia · cn
Official website
Research library1,483linked papers
Opportunities0open roles
Selected work

Representative Papers

Quantifying Individual Risk for Binary Outcome

Feb 16, 2024

This paper addresses the challenge of quantifying individual-level treatment risk in binary-outcome settings. We propose the Fraction of Negative Average (FNA) metric—the proportion of individuals whose outcomes deteriorate upon treatment—thereby complementing the Conditional Average Treatment Effect (CATE), which captures only subgroup-level averages and obscures individual harm. Under the ignorability assumption, we introduce the Pearson correlation coefficient between potential outcomes as a sensitivity parameter and derive tight, feasible theoretical bounds for FNA—substantially improving upon the classical Fréchet–Hoeffding bounds. We establish an analytical relationship among FNA, CATE, and the correlation coefficient, revealing the counterintuitive phenomenon that positive CATE can coexist with substantial individual harm. We further propose principled guidelines for selecting plausible correlation ranges and develop a nonparametric estimator for FNA that is consistent and asymptotically normal.

10 citations2 influentialRead paper

Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Jun 06, 2024arXiv.org

This work addresses the inefficiency of estimating marginal probability ratios—termed “concrete scores”—in absorption-based discrete diffusion models. We establish, for the first time, an analytical equivalence between concrete scores and clean-data conditional probabilities: concrete scores decompose into the product of the conditional probability and a closed-form time-dependent factor. Leveraging this insight, we propose RADD (Time-Invariant Reparameterized Absorption Diffusion), a reparameterization that eliminates explicit dependence on timestep indices and enables NFE (number-of-function-evaluations) caching for accelerated sampling. Theoretically, our framework unifies performance bounds for absorption diffusion and arbitrary-order autoregressive models. Empirically, RADD achieves state-of-the-art perplexity among diffusion-based language models at the GPT-2 scale across five zero-shot language modeling benchmarks. Code is publicly available.

9 citations2 influentialRead paper

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

Jan 09, 2026arXiv.org

This work addresses the scarcity of scalable, high-quality tool-interaction environments for training large language model (LLM) agents, a limitation imposed by restricted access to real-world systems, simulation hallucinations, and prohibitive human annotation costs. To overcome this, the authors propose EnvScaler, a novel framework featuring a two-stage automated synthesis pipeline: first constructing structured environment skeletons through topic mining and logical modeling, then generating diverse task scenarios and verifiable execution trajectories via rule-based methods. These synthetic data are leveraged for supervised fine-tuning and reinforcement learning to train the Qwen3 series of models. The approach automatically produces 191 distinct environments and approximately 7,000 task scenarios, yielding significant performance gains across three benchmarks in complex, multi-turn, multi-tool interaction tasks.

5 citations1 influentialRead paper

Controlled LLM Training on Spectral Sphere

Jan 13, 2026

Existing large-model optimizers struggle to simultaneously stabilize both weights and their updates, often leading to issues such as activation blowup, slow convergence, and imbalanced expert utilization in Mixture-of-Experts (MoE) architectures. This work proposes the Spectral Sphere Optimizer (SSO), which introduces, for the first time, module-wise joint spectral constraints on both weights and their updates, rigorously aligning with Maximal Update Parametrization (μP) to ensure optimization stability. SSO is derived from the steepest descent direction on the spectral sphere and integrates seamlessly into the Megatron framework, supporting Dense, MoE, and DeepNet architectures. Experiments demonstrate that SSO consistently outperforms AdamW and Muon across a 1.7B Dense model, an 8B-A1B MoE model, and a 200-layer DeepNet, effectively suppressing anomalous activations, improving routing balance, and enhancing training stability and scalability.

5 citationsRead paper

Semiparametric Efficient Fusion of Individual Data and Summary Statistics

Oct 01, 2022

This study addresses the problem of efficiently integrating internal individual-level data with external aggregate statistics to improve estimation accuracy for population-level functional parameters under weak transportability assumptions. To overcome bias and efficiency loss arising from model misspecification in conventional approaches, we first derive the semiparametric efficiency bound for this fusion setting. Methodologically, we propose two novel estimators: (i) an efficient semiparametric estimator achieving the derived bound, and (ii) an adaptive fusion estimator with oracle properties—ensuring double robustness and automatic bias correction. We establish its asymptotic efficiency and unbiasedness theoretically. Simulation studies and real-data analysis of *Helicobacter pylori* infection demonstrate that our methods significantly enhance estimation precision and statistical power compared to using internal data alone or simple weighted aggregation.

4 citations1 influentialRead paper
Recent publications

Latest Papers