Institution profile

Zhongguancun Academy

Academic institutionasia · cn
Research library334linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards a Theoretical Understanding to the Generalization of RLHF

Jan 23, 2026

This work investigates the generalization ability of Reinforcement Learning from Human Feedback (RLHF) in high-dimensional large language models. Departing from conventional analyses that rely on the consistency of maximum likelihood estimation, the study establishes, for the first time, a generalization bound under an end-to-end RLHF framework by leveraging the algorithmic stability perspective. By introducing a linear reward model and a feature coverage condition, and analyzing both Gradient Ascent (GA) and Stochastic Gradient Ascent (SGA) algorithms, the authors prove that when the feature coverage condition holds, the generalization error of the empirical optimal policy converges at a rate of $O(n^{-1/2})$. This result further extends to policies obtained via gradient-based optimization, thereby providing theoretical justification for the generalization performance of practical RLHF implementations.

3 citationsRead paper

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering

Jan 07, 2026arXiv.org

This work addresses the challenge that large language models often follow inefficient or erroneous reasoning paths due to uncontrolled autonomous strategy selection, lacking fine-grained steering over their reasoning processes. The authors introduce sparse autoencoders (SAEs) to disentangle hidden states and construct an interpretable feature space, proposing a two-stage SAE-Steering framework: first identifying sparse features associated with specific reasoning strategies, then precisely modulating model behavior through vector-based interventions. This approach overcomes the limitations of existing methods in exerting fine-grained control over reasoning dynamics, achieving a greater than 15% improvement in steering effectiveness and enabling reliable correction of erroneous reasoning trajectories, which yields a 7% absolute gain in task accuracy.

3 citationsRead paper

VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory

Jan 13, 2026

This work addresses the limitations of existing vision-language-action (VLA) models in complex, long-horizon navigation tasks, where the absence of explicit reasoning and persistent memory hinders performance in dynamic environments with strong spatial dependencies. Inspired by human dual-process cognition, the authors propose VLingNav, a novel framework featuring an adaptive chain-of-thought mechanism that dynamically triggers deliberative reasoning and a vision-augmented linguistic memory module enabling cross-modal semantic retention and long-term spatial inference. The contributions include the first VLA architecture capable of zero-shot transfer to real-world robots, the Nav-AdaCoT-2.9M dataset—a large-scale navigation benchmark with reasoning annotations—and an online expert-guided reinforcement learning strategy. VLingNav achieves state-of-the-art performance across multiple embodied navigation benchmarks and demonstrates exceptional generalization across domains and tasks.

2 citationsRead paper

SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More

Jan 09, 2026arXiv.org

Existing vision-language models struggle with precise credit assignment in complex visual reasoning tasks—such as chart understanding—leading to unreliable multi-step reasoning. To address this, this work proposes SketchVL, a framework that constructs traceable multi-step reasoning trajectories by iteratively drawing intermediate reasoning markers on the image and feeding them back into the model. SketchVL introduces, for the first time, a fine-grained process reward mechanism (FinePRM) coupled with a novel reinforcement learning algorithm, FinePO, which enables credit assignment at the level of each individual drawing action. This approach substantially enhances the model’s controllability over complex reasoning pathways, achieving an average performance gain of 7.23% across chart understanding, natural image reasoning, and mathematical reasoning benchmarks.

2 citationsRead paper

SciNetBench: A Relation-Aware Benchmark for Scientific Literature Retrieval Agents

Dec 16, 2025arXiv.org

Existing scientific literature retrieval agents struggle to model complex inter-paper relationships—such as support, contradiction, or technological evolution—leading to fragmented knowledge and misinterpretation of research landscapes. This work introduces SciNet, the first systematically constructed benchmark dataset for relation-aware retrieval, encompassing 269 million papers and 8,940 tasks, and evaluates relational understanding at three granularities: ego-centric, pairwise, and path-wise. Built upon large-scale cross-disciplinary metadata and knowledge graphs, SciNet exposes a fundamental limitation of current mainstream retrieval methods, whose accuracy in modeling such relationships consistently falls below 20%. Incorporating this benchmark substantially improves the quality of downstream literature reviews by 25.3%, demonstrating the critical value of relation-awareness for scientific intelligence.

2 citationsRead paper
Recent publications

Latest Papers