Problem
Research questions and friction points this paper is trying to address.
Reinforcement learning
verifiable rewards
large language models
reasoning capabilities
Research questions and friction points this paper is trying to address.
Methods, ideas, or system contributions that make the work stand out.