Ziting Wang, Qiuhao Chen, Yuxuan Du, Zhihu Yang, Xiaoxia Cai, Kaixuan Huang, Jing-Ning Zhang, Kai Xu, Jun Du, Yinan Li, Yuling Jiao, Xingyao Wu, Xiliang Lu, Xiliang Lu, Yirong Jin, Ruixia Wang, Haifeng Yu, S. P. Zhao
Abstract Effectively implementing quantum algorithms on noisy intermediate-scale quantum (NISQ) processors is a central task in modern quantum technology. NISQ processors feature tens to a few hundreds of noisy qubits with limited coherence times and gate operations with errors, so NISQ algorithms naturally require employing circuits of short lengths via quantum compilation. Here, we evaluate a reinforcement learning (RL)-based quantum compiler on a superconducting processor. Our experiments reveal that for two-qubit circuits, the RL-based compiler surpasses conventional methods, demonstrating its ability to discover hardware-amenable circuits with near-optimal lengths. However, for three-qubit circuits, the RL-based compiler does not achieve unity theoretical fidelity. To address this limitation, we integrate a variational strategy with the RL-based compiler, highlighting their complementary strengths. Systematic experiments show that this variational RL-based compiler consistently identifies near-optimal circuits, even under stringent hardware constraints, outperforming conventional techniques. Furthermore, we analyze the impact of decoherence and gate errors, providing critical insights into the practical performance of RL-based compilers on quantum hardware. These findings exemplify the codesign of the software with hardware for efficient quantum compilation, offering valuable insights for the advancement of RL-based compilers.