AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
This work proposes a post-training framework that integrates synthetic data generation with reinforcement learning to enable large language models to automatically translate natural language descriptions of operations research (OR) problems into formal optimization models—including linear, mixed-integer, and nonlinear formulations—thereby substantially reducing reliance on specialized OR expertise. The approach innovatively leverages solver feedback as a reward signal and introduces a curriculum-based reinforcement learning strategy tailored for nonlinear dynamical problems, achieving, for the first time, effective automated modeling in this challenging domain. Evaluated on an 8B-parameter model, the method matches or exceeds state-of-the-art performance across six standard OR benchmarks, rivaling results from significantly larger models, and boosts solution accuracy on nonlinear dynamics tasks from near 0% to within solvable ranges.