跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Bisma Majid, Shabir Ahmed Sofi, Mir Mohammad Yousuf·· 15 小时前AI 评分32

APGEM:面向 NISQ 混合量子强化学习的上下文感知错误缓解编排框架

Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems

AI 导读

研究提出自适应策略引导错误缓解(APGEM),作为混合量子-经典训练循环中的上下文感知编排层,在量子强化学习训练中动态选择 ZNE、PEC、CDR 和 REM 中最合适的缓解策略。在 CVRP 问题上,APGEM 持续优于传统静态缓解方法,达到 oracle 策略约 94% 的效用,并在噪声增大时保持更高的量子态保真度。消融实验显示该框架能学习适应不同噪声环境与电路执行条件的上下文感知缓解策略。

正文

View PDF HTML (experimental)

Abstract:Quantum Reinforcement Learning (QRL) integrates reinforcement learning with parameterized quantum circuits and is a promising approach to combinatorial optimization. On Noisy Intermediate-Scale Quantum (NISQ) devices, however, decoherence, gate imperfections, and measurement errors reduce policy quality and make learning less reliable. Existing error mitigation techniques are generally applied as fixed corrections that do not adapt to changing noise conditions or to the evolving state of training. This work presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer of the hybrid quantum-classical training loop that dynamically selects the most suitable mitigation strategy during QRL training. APGEM evaluates Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) using policy-level indicators, including quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio, and integrates the selected strategy directly into the reinforcement learning loop. The framework is evaluated on the Capacitated Vehicle Routing Problem (CVRP), a representative NP-hard problem in urban logistics, under a range of NISQ noise models and noise levels. APGEM consistently outperforms conventional static mitigation methods, reaches approximately 94% of the utility of an oracle strategy, maintains higher quantum-state fidelity as noise increases, and produces more stable learning behaviour throughout training. Ablation studies show that the framework learns context-aware mitigation policies that adapt to different noise environments and circuit execution conditions. These findings demonstrate that integrating adaptive error mitigation into the learning process substantially improves the robustness and reliability of QRL on NISQ hardware.
Comments: 9 pages, 18 figures, 13 tables
Subjects: Machine Learning (cs.LG); Quantum Physics (quant-ph)
Cite as: arXiv:2610.01253 [cs.LG]
  (or arXiv:2610.01253v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01253

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mir Mohammad Yousuf [view email]
[v1] Thu, 1 Oct 2026 07:45:12 UTC (359 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org