๐ค AI Summary
This study addresses the lack of a unified evaluation framework for generative AI in quantum circuit and code generation, particularly the absence of validation on real quantum hardware. Through a structured scoping review, the authors systematically analyze 13 generative systems and five datasets, proposing the first taxonomy based on output modality and training paradigm. They further introduce a three-tiered evaluation framework encompassing syntactic validity, semantic correctness, and hardware executability. The findings reveal that while existing methods generally satisfy syntactic requirements and partially meet semantic criteria, none demonstrate end-to-end validation on actual quantum devices. This critical gap underscores the fieldโs current limitation in real-device verification and provides a clear direction for future research toward practical, hardware-aware quantum program synthesis.
๐ Abstract
We review thirteen generative systems and five supporting datasets for quantum circuit and quantum code generation, identified through a structured scoping review of Hugging Face, arXiv, and provenance tracing (January-February 2026). We organize the field along two axes: artifact type (Qiskit code, OpenQASM programs, circuit graphs); crossed with training regime (supervised fine-tuning, verifier-in-the-loop RL, diffusion/graph generation, agentic optimization); and systematically apply a three-layer evaluation framework covering syntactic validity, semantic correctness, and hardware executability. The central finding is that while all reviewed systems address syntax and most address semantics to some degree, none reports end-to-end evaluation on quantum hardware (Layer 3b), leaving a significant gap between generated circuits and practical deployment. Scope note: quantum code refers throughout to quantum program artifacts (QASM, Qiskit); we do not cover generation of quantum error-correcting codes (QEC).