Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于GPU的批处理Levenberg-Marquardt求解器,用于树型遗传编程中的常数优化问题,显著提高了符号回归中数值系数优化的速度和质量。
📝 Abstract
Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated frameworks to omit it or restrict it to lightweight forms. We present a GPU-resident, batched Levenberg--Marquardt solver that optimizes constants across a structurally heterogeneous population of expression trees using a fixed number of population-wide CUDA launches per iteration. Reverse-mode automatic differentiation assembles the per-tree Jacobian in one backward sweep, making the dominant per-iteration cost independent of the number of constants per tree, and a double-precision delivery guard guarantees that returned constants are never worse than their initial values. On early-generation populations, the solver sustains up to $5.1{\times}10^{5}$ trees per second on an NVIDIA A100; at a GPU-saturated benchmark configuration it delivers roughly $9.9{\times}$ the throughput of Operon running on a 64-core EPYC 7763, while matching fp64-reference quality. Integrated in-process into EvoGP, the solver enables end-to-end search to recover governing equations on $10$ of $18$ constructed problems versus 0 for stock EvoGP. Our code is at https://github.com/TensorConv/CuSR.
Problem

Research questions and friction points this paper is trying to address.

constant optimization
symbolic regression
tree-based genetic programming
GPU acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU-accelerated
Levenberg--Marquardt solver
reverse-mode automatic differentiation
constant optimization