Yu-Hao Liu, Yu-Chun Chang, Yanhua Ma
The rapid growth of short video platforms, along with advances in big data and the Internet of Things (IoT), has significantly increased the volume of data being generated, providing a strong foundation for the development of artificial intelligence (AI). Among AI technologies, deep learning based on neural networks has achieved notable success in fields such as speech recognition, natural language processing, and image analysis. However, as these models become more complex, traditional hardware architectures face growing limitations. The slowdown of Moore's Law and increasing concerns about power consumption highlight the urgent need for more efficient hardware solutions. In resource-constrained environments like real-time and edge computing, achieving a balance between performance, power, and latency is especially important. This review addresses these challenges through three main contributions: (1) it categorizes and analyzes key optimization techniques at both the algorithm and hardware levels, offering a clear theoretical framework; (2) it summarizes recent advancements in accelerator design, with a focus on technologies such as collaborative acceleration and in-memory computing; and (3) it explores future trends and challenges, offering insights into the evolution of neural network accelerators and potential solutions to emerging technical bottlenecks.