James Read, Ming-Yen Lee, Wei-Hsing Huang, Yuan-Chun Luo, Anni Lu, Shimeng Yu
The exponential growth of artificial intelligence (AI) applications has strained conventional von Neumann architectures, where frequent data transfers between compute units and memory create significant energy and latency bottlenecks. Compute-in-Memory (CIM) addresses this challenge by performing multiply-accumulate (MAC) operations directly in memory arrays, substantially reducing data movement. As transformers increasingly dominate AI workloads, extending CIM frameworks to support these architectures becomes critical. In this work, we present NeuroSim V1.5, an integrated framework for cooptimizing accuracy and hardware efficiency in CIM accelerator design. Key contributions include: (1) a hybrid ACIM/DCIM architecture enabling transformer acceleration, with a case study comparing vision transformer (ViT) and CNN performance; (2) an integrated co-optimization framework combining TensorRT-based quantization, flexible noise modeling, and circuit-level PPA estimation; (3) expanded device support including non-volatile capacitive memories; and (4) up to 6.5× faster runtime through GPU-accelerated simulation. Our ViT case study reveals that ViTs exhibit lower noise tolerance and area efficiency than CNNs of comparable size, suggesting CNNs remain well-suited for vision applications on CIM hardware. The hybrid architecture provides the foundation for future transformer workloads including large language models. All versions are available open-source at https://github.com/neurosim/NeuroSim.