BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
This study investigates whether reinforcement learning with verifiable rewards (RLVR) genuinely enhances the reasoning capabilities of large language models or merely improves sampling efficiency. To this end, we introduce BODHI-Trees—a novel tree-based representation that extracts semantically equivalent structures from mathematical reasoning trajectories—and propose semantic branching entropy as a new metric to quantify reasoning diversity. Through controlled maze experiments and trajectory analyses, we find that while RLVR strengthens constraint adherence and backtracking abilities, it substantially contracts the semantic reasoning space, leading to a concurrent collapse in both policy entropy and semantic branching entropy. These findings suggest that the efficiency gains conferred by RLVR may come at the cost of reduced reasoning diversity.