Institution profile

Xiamen University

Academic institutionasia · cn
Official website
Research library823linked papers
Opportunities0open roles
Selected work

Representative Papers

How to Detect Network Dependence in Latent Factor Models? A Bias-Corrected CD Test

Sep 01, 2021

This paper addresses the failure of residual cross-sectional dependence tests in latent factor panel models. We propose a bias-corrected CD* test statistic. Theoretically, we first establish the asymptotic validity of the standard CD test under weak factors and rigorously derive the asymptotic standard normality of CD* under the null hypothesis, while demonstrating its high local power against network-type alternatives. Methodologically, the CD* test integrates factor estimation, residual extraction, and analytical bias correction, accommodating both strong and weak factors as well as serially correlated errors. Monte Carlo simulations show that CD* achieves accurate size and superior power in small samples, consistently outperforming the JR test. Empirically, applying CD* to a housing price dynamics model across 377 U.S. metropolitan statistical areas reveals statistically significant spatial dependence in residuals.

49 citations1 influentialRead paper

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

Jan 28, 2026

This work addresses the limitation of existing reinforcement learning approaches in mathematical reasoning, which often neglect challenging problems and lack a systematic mechanism for progressive difficulty escalation, thereby constraining model performance on complex tasks. To overcome this, the authors propose MathForge, a novel framework that jointly emphasizes high-difficulty problems from both algorithmic and data perspectives. Algorithmically, they introduce Difficulty-aware Grouped Policy Optimization (DGPO) with a difficulty-balanced advantage estimator. On the data side, they develop Multi-dimensional Question Rewriting (MQR), which enables controllable difficulty enhancement while preserving answer consistency. Extensive experiments demonstrate that MathForge significantly outperforms current methods across multiple mathematical reasoning benchmarks, validating the efficacy of a difficulty-centric training paradigm for enhancing large language models’ reasoning capabilities.

4 citationsRead paper

Federated modality-specific encoders and partially personalized fusion decoder for multimodal brain tumor segmentation

Aug 18, 2025Medical Image Anal.

This work addresses the challenge of simultaneously handling missing modalities and personalization demands in federated multi-modal medical image segmentation. To this end, the authors propose FedMEPD, a novel framework that assigns a dedicated encoder to each modality to accommodate heterogeneous modality availability across clients, and introduces a partially personalized fusion decoder. This decoder leverages global multi-modal representation anchors and cross-attention mechanisms to effectively compensate for missing modality information. As the first approach to jointly tackle modality heterogeneity and personalization under the federated learning paradigm, FedMEPD demonstrates significant performance gains over existing methods on the BraTS 2018 and 2020 datasets, validating its effectiveness and superiority in personalized federated multi-modal learning.

4 citationsRead paper

Beyond the Black Box: Theory and Mechanism of Large Language Models

Jan 06, 2026arXiv.org

While large language models have demonstrated remarkable engineering success, they remain theoretically underdeveloped and mechanistically opaque—essentially operating as “black boxes.” This work proposes the first unified theoretical framework encompassing the entire lifecycle of large language models, systematically analyzing the core mechanisms across six stages: data preparation, model construction, training, alignment, inference, and evaluation. By integrating information theory, optimization theory, and representation learning, the framework elucidates the mathematical principles underlying critical issues such as data mixing strategies, architectural expressivity, and alignment optimization. Furthermore, it identifies forward-looking challenges including self-improving synthetic data generation, safety boundaries, and the origins of emergent intelligence. This study provides a structured roadmap toward transforming large language models from empirical engineering artifacts into an explainable, predictable, and verifiable scientific discipline.

3 citations1 influentialRead paper

Learning Unbiased Cluster Descriptors for Interpretable Imbalanced Concept Drift Detection

Feb 01, 2026IEEE Transactions on Emerging Topics in Computational Intelligence

This work addresses the challenge of detecting small concept drifts in data streams that are often obscured by dominant, larger concepts—a phenomenon known as the "masking effect" caused by concept imbalance. To overcome this limitation, the authors propose ICD3, a novel method that employs multi-granularity distribution search to identify concepts of varying scales and constructs an individual one-class classifier (OCC) for each concept to monitor its drift independently, thereby preventing larger concepts from dominating the detection process. ICD3 is the first approach to enable unbiased and interpretable detection of imbalanced concept drifts, accurately pinpointing the specific drifting concepts while remaining robust to variations in inter-concept imbalance ratios. Extensive experiments on multiple benchmark datasets demonstrate that ICD3 consistently outperforms state-of-the-art methods in both detection accuracy and interpretability.

3 citationsRead paper
Recent publications

Latest Papers