🤖 AI Summary
This work demonstrates that neural reasoners such as the Lattice Deduction Transformer, despite appearing to perform search and backtracking on constraint-rich tasks like Sudoku, actually rely solely on a single forward pass for prediction, with their apparent “search” serving only to reduce computational redundancy. The study introduces the concept of “first-pass poisoning,” showing that solution accuracy is governed by model calibration and symmetry handling. It further reveals that constraint-graph attention mechanisms substantially outperform positional encodings and proposes two effective interventions: symmetry-aware data augmentation and test-time ensembling. By integrating digit permutation augmentation, symmetric ensembling, and CoLT’s branching heuristics with shared unsatisfiable core techniques, the model achieves a dramatic improvement—from under 1% to 96.5 ± 0.3%—in accuracy on a 9×9 Sudoku symmetric holdout set, attaining 100% success on the most challenging instances.
📝 Abstract
Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly that. In clue-rich Sudoku, it does not: one forward pass commits essentially the entire grid (every blank cell on standard 6x6, 94-96% on augmented 9x9), turning the iterative solver into a one-shot predictor wrapped in an exact verifier. All hard-slice failures are decided before search begins, when the first pass confidently deletes a value required by the true solution. We call this first-pass poisoning. Adding learned branching, MRV, backtracking, value exclusion, and shared nogoods (CoLT) does not change which Sudoku instances are solved; it cuts repeated invalid derivations 1,497-fold. At the frozen training budget, constraint-graph attention alone matches full-CoLT accuracy, while positional tables recover only under substantially longer training, indicating an optimization and sample-efficiency advantage rather than an absolute capacity difference. The diagnosis predicts two effective interventions. Digit-permutation augmentation raises 9x9 accuracy from below 1% to 96.5 +/- 0.3 across three training seeds on a symmetry-disjoint split. Test-time union over symmetry-transformed passes raises all three hard-slice checkpoints from 72.8-78.9% to 100% without retraining. On from-scratch graph coloring, one-shot behavior disappears and search changes accuracy. In clue-rich completion, LDT-like systems are one-shot amortized predictors rather than learned search procedures: accuracy is determined by calibration and symmetry, while search primarily removes computational waste.