科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Very Large Scale Integration (VLSI) Systems2026-01-22· Bitwise operation

TABv2: A Faster Ternary and Binary Neural Network Inference Library on the Edge

Guanshujie Fu, Olivier Fischer, Shien Zhu, Gustavo Alonso

原始摘要(英文原文)· Original abstract
Modern deep neural networks and large language models (LLMs) are trained and inferred using low-bitwidth representations like FP8 and FP4. Ternary and binary quantization can further reduce the computation cost by extreme 1/2-bit data representation and bitwise operations. As ternary and binary neural networks (TNNs, BNNs, and the mixed-precision ternary-activation binary-weight and binary-activation ternary-weight networks (TBNs and BTNs)) achieve different trade-off points between the speed, model size, and accuracy, they are suitable for certain applications on the edge. However, existing works mainly focus on accelerating BNN and DoReFa-Net-style bit-serial operations, leaving TNNs and mixed-precision ones under-optimized. Although some related works such as ternary and binary (TAB) provide reference implementation for ternary and binary networks, they encounter high data access overhead during quantization and suffer from slow scalar popcount on advanced vector extension 2 (AVX2) central processing unit (CPU) which have no single-instruction multiple-data (SIMD) popcount instructions. In this article, we propose TABv2 to achieve faster inference for TNNs, BNNs, TBNs, and BTNs on CPU and graphic processing unit (GPU). First, we optimize the quantization and image-to-row (img2row) by operator fusion and better address calculation to improve the data locality. We also propose warp-cooperative quantization on GPU by utilizing data sync intrinsics. Second, we replace the scalar popcount on AVX2 with equivalent SIMD operations and reduce the complexity of the SIMD popcount utilizing ternary encoding, which can reduce approximately 15% total instruction in ternary bitwise general matrix-matrix multiplication (GEMM). Third, we propose new bitwise GEMM algorithms for TNNs and TBNs by utilizing GPU-native bitwise matrix multiplication intrinsics on tensor cores for higher efficiency. We further apply a four-degree pipeline for bitwise GEMM on GPU to hide the memory access latency. Finally, we implement these methods in C++ and combine them into an open-source library. Evaluation results show that we achieve layer-level speedup of up to 2.7× on AVX2 CPU, 2.3× on ARM CPU, and 8.7× on Nvidia GPU over the TAB baseline for TNNs, TBNs, BTNs, and BNNs. Moreover, we achieve 1.3× - 1.9× end-to-end speedup and 1.2× - 1.8× energy efficiency on CPUs and 1.3× - 3.5× end-to-end speedup and 1.2× - 3.0× energy efficiency on GPU compared to TAB on ResNet, Darknet, and visual geometry group (VGG) models.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

TABv2: A Faster Ternary and Binary Neural Network Inference Library on the Edge — 科研速览 Science Skim