🤖 AI Summary
This study addresses the challenge of feedback control under noisy continuous measurements where quantum states are not directly accessible. We propose a Kraus-parameterized belief reinforcement learning method that integrates quantum state geometry into the learning loop. By imposing Stiefel manifold constraints on the encoder, the approach generates physically valid and interpretable density matrix estimates, while employing the PPO algorithm to achieve continuous control mapping. Experimental results demonstrate that this method attains a belief fidelity of 0.77–0.80 with significantly lower return variance compared to LSTM baselines. Consequently, it enables more stable quantum feedback control under non-ideal conditions, effectively resolving the lack of physical constraints in belief representations inherent to traditional approaches.
📝 Abstract
Quantum feedback control requires acting on noisy continuous measurement records without direct access to the underlying quantum state. We propose Kraus-Parameterized Belief Reinforcement Learning, a pipeline in which a recurrent encoder, constrained to the Stiefel manifold, produces density-matrix estimates that are guaranteed positive-semidefinite and trace-normalized by construction, embedding quantum state geometry directly into the learning loop. A Proximal Policy Optimization (PPO) actor then maps these physically valid belief states to continuous control actions. On a simulated continuously monitored qubit, the resulting policy achieves stable feedback control, maintaining a measurement-conditioned belief fidelity of approximately 0.77-0.80 and exhibiting substantially lower return variance than a parameter-matched LSTM-history baseline across both nominal and out-of-distribution conditions. Although gains in raw target fidelity are modest, the geometric constraint guarantees a physically valid, interpretable belief representation and yields markedly more stable control under measurement inefficiency and abrupt dynamics switches. These results indicate that physics-informed neural memory is a practical inductive bias for reliable quantum feedback control.