🤖 AI Summary
This study addresses the systematic bias and coverage failure in opponent modeling for poker-like games, where folding actions induce missingness in private hand information. The authors introduce the concept of “safe observation capacity,” characterize its concave piecewise-linear frontier, and establish a theoretical relationship between reveal rate and estimation accuracy. They propose the first Safe Active De-censoring (SAD) framework, which employs floor-safe probes to guide play toward showdowns and leverages sequence-form flow recovery to reconstruct censored fold quality, enabling unbiased estimation and robust exploitation of opponent strategies. Experiments demonstrate that SAD significantly outperforms public-data-only baselines at the million-hand scale (Holm-corrected p ≤ 0.012), improving certified river over-fold exploitation payoff from 0.485 to 0.692 and correctly identifying all 30 simulated seeds in twin tests (certified absolute value 0.655, 95% CI half-width 0.008).
📝 Abstract
In poker-like games, folds hide private cards, so showdown data are missing not at random. The usual per-card estimator converges to a selected distribution, and its confidence sets can lose coverage as the sample grows. A floor-safe probe changes the monitoring process: it drives a chosen line to showdown, reveals every non-fold continuation, and uses sequence-form flow to recover censored fold mass on reveal-certified histories. We price this repair through safe observation capacity $\kappa_\rho(I)$, the largest floor-safe reach rate at safety budget $\rho$. Its frontier is concave and piecewise linear, with origin slope equal to the floor's shadow price. When one-hand reveal mass factorizes into safe reach and opponent continuation, matching bounds give conditional per-target cost $N=\widetilde{\Theta}(1/(\kappa_\rho(I)\pi\varepsilon^2))$ for a local censored-fiber direction. Safe Active De-censoring (SAD) combines capacity with public-anomaly routing and robust deployment; routing across targets remains heuristic. Evidence spans bucketed turn-river endgames, controlled instances, and a fixed-board unbucketed river subgame. In the latter, a constructed public twin admits a floor-safe response of value $V=0.815$; the audited public channel certifies only the blueprint floor, while population reveal evidence certifies at least $96\%$ of $V$. Across a broader synthetic opponent population, public and solved grouped reveal fibers certify median shares of $73\%$ and $91\%$ of the safe-exploitable gap. Independent floor audits cover every evaluated probe and response. The results separate unconditional safety from the conditional statistical value of active reveal.