Institution profile

Beijing University of Posts and Telecommunications

Academic institutionasia · cn
Official website
Research library1,765linked papers
Opportunities0open roles
Selected work

Representative Papers

SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Mar 24, 2025arXiv.org

This work investigates the universality and training dynamics of zero-shot reinforcement learning (Zero RL) across heterogeneous foundation models. Method: We systematically evaluate whether chain-of-thought (CoT) reasoning emerges directly from base models—without explicit CoT supervision—across ten open-source models spanning diverse architectures and scales. We introduce two key design principles: format reward shaping and query difficulty control, and integrate rule-based RL, implicit CoT supervision, joint monitoring of response length and verification behavior, and a cross-model training dynamics analysis framework. Contribution/Results: We observe, for the first time, a “reasoning insight moment” in non-Qwen small-scale models. Our analysis reveals a non-monotonic relationship between model scale and training dynamics. Experiments demonstrate significant improvements in reasoning accuracy and response length across most models. To foster reproducibility, we open-source all code, fine-tuned models, and analytical tools.

25 citations4 influentialRead paper

Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting

Jan 05, 2026arXiv.org

This work addresses the issue of catastrophic forgetting in supervised fine-tuning (SFT) for domain adaptation, which stems from destructive gradient updates caused by low-entropy yet low-probability “confidently conflicting” samples. To mitigate this, the authors propose Entropy-Adaptive Fine-Tuning (EAFT), a novel approach that introduces token-level entropy as a gradient gating mechanism. EAFT effectively distinguishes between epistemic uncertainty and knowledge conflict, dynamically suppressing harmful updates from conflicting samples while preserving the model’s ability to learn from uncertain ones. Experiments on large language models—including Qwen and GLM—demonstrate that EAFT significantly alleviates degradation in general capabilities across mathematical, medical, and agent-based tasks, while maintaining downstream performance comparable to standard SFT.

4 citations2 influentialRead paper

Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

Jan 28, 2026

This work addresses the limited capability of current text-to-image models in modeling complex spatial relationships—such as positional arrangements, occlusion, and causality—highlighting that existing evaluation benchmarks fall short due to their reliance on short, information-sparse prompts. To bridge this gap, the authors introduce SpatialGenEval, a fine-grained evaluation framework spanning ten spatial subdomains, along with the SpatialT2I dataset comprising 1,230 information-dense long prompts and 15,400 human-verified image–text pairs. Through structured multiple-choice question-style prompts and fine-tuning experiments on models like Stable Diffusion-XL, they demonstrate that training on information-rich data significantly enhances spatial reasoning. Evaluations across 21 state-of-the-art models reveal persistent difficulties with higher-order spatial relations; however, fine-tuning on SpatialT2I yields consistent performance gains of 4.2%–5.7%, producing images more aligned with real-world spatial logic.

1 citationsRead paper

CovertComBench: The First Domain-Specific Testbed for LLMs in Wireless Covert Communication

Jan 26, 2026

This work addresses the lack of suitable benchmarks for evaluating large language models (LLMs) in the domain of wireless covert communication, where stringent security constraints—such as Kullback–Leibler (KL) divergence limits—are critical. To bridge this gap, the authors propose the first LLM-specific evaluation benchmark tailored to this field, encompassing tasks in conceptual understanding, optimization derivation, and code generation. They further introduce a novel automatic scoring mechanism grounded in detection theory, implementing an “LLM-as-Judge” framework. Experimental results reveal that while LLMs achieve strong performance in concept identification (81%) and code generation (83%), their accuracy drops significantly in security-critical mathematical derivations, ranging from 18% to 55%. These findings underscore the models’ limitations in high-order reasoning and affirm their role as assistive tools rather than autonomous solvers in safety-sensitive applications.

1 citationsRead paper
Recent publications

Latest Papers