Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文介绍了Backline框架,通过Python接口和MLIR编译,解决量子工作负载在CPU、GPU和FPGA上的低延迟执行问题。
📝 Abstract
Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time quantum error correction (QEC) demand performance that traditional interpreted environments cannot provide. FPGAs and ASICs play a central role at these layers, but their specialized programming models make development rigid and time-consuming. CPUs, GPUs, and other accelerators introduce a different challenge: as infrastructure becomes increasingly heterogeneous, programming across different devices and their associated abstractions becomes more complex. Allowing researchers to write workloads in high-level languages that map to low-latency execution across diverse distributed target platforms will enable the development of key infrastructure for utility-scale quantum systems. For this, we introduce $\textit{Backline}$, a heterogeneous compilation and runtime framework built within PennyLane and Catalyst. Backline allows us to design and build quantum-classical workloads for high-performance and low-latency devices, with compilation directly from a Python interface through MLIR. We demonstrate the compilation and execution of several quantum workloads with low-latency data movement across a mix of CPUs, GPUs, and FPGAs, for both local and distributed remote hardware targets, all from a vendor-agnostic Python frontend. With an AMD VPK120 FPGA board as the controller, issuing each round from its hardware-handshake engine, we measured median steady-state round-trip latencies over RoCE v2 of $2.305~μ$s to an AMD Ryzen Threadripper PRO CPU and $4.5~μ$s to an AMD Instinct MI210 GPU across $10^6-1$ rounds per path, demonstrating microsecond-scale synchronous co-processing.
Problem

Research questions and friction points this paper is trying to address.

quantum workloads
low-latency
heterogeneous infrastructure
quantum error correction
Python frameworks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous Compilation
Low-latency Execution
Quantum Workloads
Python Interface
Distributed Hardware Targets
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Joseph K. L. Lee
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
M
Mehrdad Malekmohammadi
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
H
Hong-Sheng Zheng
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
S
Shuli Shu
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
C
Cheick Doumbia
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
K
Kalman Szenes
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
M
Mehran Zamani Abnili
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
T
Thomas Ainsworth
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
M
Matthew Seymour
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
T
Thomas Germain
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
L
Leonhard Neuhaus
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada
Josh Izaac
Josh Izaac
Xanadu Quantum Technologies, Inc.
L
Lee J. O'Riordan
Xanadu Quantum Technologies Inc., Toronto, Ontario M5G 2C8, Canada