CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning
本文提出CAST方法,通过结合规划引导行为和当前策略来改进值学习,以解决模型基础强化学习中的动作选择问题。
本文提出CAST方法,通过结合规划引导行为和当前策略来改进值学习,以解决模型基础强化学习中的动作选择问题。
本文提出了一种两级域分解AdaGrad方法(2DD-AG2m),以解决图神经网络在分布式环境下的高效训练问题,通过交替优化全局和分区图,降低了计算成本并提高了预测性能。
This study investigates whether reinforcement learning with verifiable rewards (RLVR) genuinely enhances the reasoning capabilities of large language models or merely improves sampling efficiency. To this end, we introduce BODHI-Trees—a novel tree-based representation that extracts semantically equivalent structures from mathematical reasoning trajectories—and propose semantic branching entropy as a new metric to quantify reasoning diversity. Through controlled maze experiments and trajectory analyses, we find that while RLVR strengthens constraint adherence and backtracking abilities, it substantially contracts the semantic reasoning space, leading to a concurrent collapse in both policy entropy and semantic branching entropy. These findings suggest that the efficiency gains conferred by RLVR may come at the cost of reduced reasoning diversity.
Existing 3D morphable face models (3DMMs) exhibit shape biases due to limited training data, compromising their generalization and fairness across diverse populations. This work proposes the first evaluation framework integrating curvature-aware and spectral geometric analysis, leveraging the Laplace–Beltrami operator to generate high-resolution curvature error maps that enable precise localization, quantification, and visualization of reconstruction biases. Experimental results demonstrate that the proposed error metric aligns closely with human perception and significantly outperforms conventional Euclidean distance measures. Systematic evaluation across multiple state-of-the-art 3DMMs reveals pronounced reconstruction biases correlated with age, gender, and ethnicity, with age-related discrepancies being particularly substantial.
This study investigates whether human visual representations align more closely with discriminative or generative learning objectives and examines how these objectives influence alignment with human perception. By leveraging Joint Energy Models (JEMs) to continuously interpolate between the two paradigms—adjusting only a single mixing coefficient—while holding architecture, model scale, and training data constant, the authors systematically evaluate performance across six human-alignment benchmarks. The results reveal that optimal alignment with human vision is achieved not at either pure discriminative or pure generative extremes, but at an intermediate training objective, challenging the conventional dichotomy. This finding is robustly corroborated across multiple dimensions, including perceptual similarity, gloss modeling, response uncertainty, robustness evaluations, shape–texture conflict resolution, and feature attribution, collectively demonstrating the superior capacity of hybrid JEMs to emulate the multifaceted nature of human visual behavior.
本文提出CAST方法,通过结合规划引导行为和当前策略来改进值学习,以解决模型基础强化学习中的动作选择问题。
本文提出了一种两级域分解AdaGrad方法(2DD-AG2m),以解决图神经网络在分布式环境下的高效训练问题,通过交替优化全局和分区图,降低了计算成本并提高了预测性能。
This study investigates whether reinforcement learning with verifiable rewards (RLVR) genuinely enhances the reasoning capabilities of large language models or merely improves sampling efficiency. To this end, we introduce BODHI-Trees—a novel tree-based representation that extracts semantically equivalent structures from mathematical reasoning trajectories—and propose semantic branching entropy as a new metric to quantify reasoning diversity. Through controlled maze experiments and trajectory analyses, we find that while RLVR strengthens constraint adherence and backtracking abilities, it substantially contracts the semantic reasoning space, leading to a concurrent collapse in both policy entropy and semantic branching entropy. These findings suggest that the efficiency gains conferred by RLVR may come at the cost of reduced reasoning diversity.
Existing 3D morphable face models (3DMMs) exhibit shape biases due to limited training data, compromising their generalization and fairness across diverse populations. This work proposes the first evaluation framework integrating curvature-aware and spectral geometric analysis, leveraging the Laplace–Beltrami operator to generate high-resolution curvature error maps that enable precise localization, quantification, and visualization of reconstruction biases. Experimental results demonstrate that the proposed error metric aligns closely with human perception and significantly outperforms conventional Euclidean distance measures. Systematic evaluation across multiple state-of-the-art 3DMMs reveals pronounced reconstruction biases correlated with age, gender, and ethnicity, with age-related discrepancies being particularly substantial.
This study investigates whether human visual representations align more closely with discriminative or generative learning objectives and examines how these objectives influence alignment with human perception. By leveraging Joint Energy Models (JEMs) to continuously interpolate between the two paradigms—adjusting only a single mixing coefficient—while holding architecture, model scale, and training data constant, the authors systematically evaluate performance across six human-alignment benchmarks. The results reveal that optimal alignment with human vision is achieved not at either pure discriminative or pure generative extremes, but at an intermediate training objective, challenging the conventional dichotomy. This finding is robustly corroborated across multiple dimensions, including perceptual similarity, gloss modeling, response uncertainty, robustness evaluations, shape–texture conflict resolution, and feature attribution, collectively demonstrating the superior capacity of hybrid JEMs to emulate the multifaceted nature of human visual behavior.