🤖 AI Summary
This study addresses feedback bias and the lack of convergence in counterfactual regret minimization for persistent partial evaluation by proposing a deterministic convergence theory that eliminates the need for unbiased estimators. By establishing an objective transfer theorem, bounding public chance slicing, and formulating a component decomposition theorem, this work reveals the coupling mechanism between coverage discrepancy and policy paths, thereby providing convergence guarantees and numerical certificates for persistent scheduling. Integrating additive sign regret matching with the RM+ algorithm, experiments demonstrate that the proposed method significantly outperforms reshuffling strategies in Texas Hold’em endgames. Furthermore, this research identifies public chance width as a critical learning variable, facilitating the auditing of execution trajectories and offering a robust theoretical foundation for imperfect-information game solving.
📝 Abstract
At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update. Exact evaluation processes the full cut at one strategy profile; persistent partial evaluation processes a fixed without-replacement order across evolving profiles. The latter covers every outcome once per epoch, yet its feedback is generally conditionally biased because earlier batches influence the profiles seen by later batches. We establish a deterministic target-transfer theorem for uniform, nonnested additive public cuts. The theorem bounds full-cut exploitability by regret on the delivered feedback and a public-debit term that couples prefix coverage discrepancy with motion along the realized strategy path. Consecutively balanced schedules consequently converge for additive signed regret matching (RM) and RM+ under predetermined averaging weights, while a fixed RM+ construction proves that the discrepancy--path product is necessary in general. A component-resolved form of the theorem converts an execution trace into a numerical exploitability certificate. On two released heads-up no-limit hold'em turn endgames, persistent order improves substantially over fresh reshuffling despite identical epochwise coverage, and partial coverage wins every registered shallow matched-budget comparison. A depth study locates a crossover between 32 and 64 full-cut outcome budgets, after which complete coverage dominates. These results characterize public-chance width and order as learning variables and provide a deterministic basis for designing and auditing persistent CFR schedules.