Jiaming Liu, Zhijian Hao, Qi Zheng, Chen Lu, Honglei Chen, Wenzhong Bao, Hongkai Xiong
Effective black-box image signal processor (ISP) tuning is critical for transforming RAW sensor data into high-quality RGB images across applications like autonomous driving and mobile photography, yet manual tuning is labor-intensive and impractical for diverse conditions. While deep reinforcement learning (DRL) offers the most promising learnable solution for automating this process, its training inefficiencies, driven by slow black-box computation, limit scalability. This paper proposes Q-Accel, an off-policy DRL framework that accelerates the training of black-box ISP tuning by decoupling computation from sampling and optimizing resource usage. Q-Accel integrates a RAW-Input Agent for low-complexity processing, an Index Replay Buffer for memory efficiency, an Auxiliary Proxy for rapid sample generation, and a Confidence-based Hybrid Reward to balance real and synthetic data. Experiments demonstrate a 59.94% reduction in training time, an 88.89% decrease in memory increment, and state-of-the-art performance. By addressing DRL inefficiencies, Q-Accel enables scalable, efficient training for ISP tuning, advancing practical deployment in diverse real-world scenarios and aligning with sustainable AI principles.