Adriano Caprara, Ying Yu, Fei Teng, Adrià Junyent‐Ferré, Eduard Bullich‐Massagué, Mònica Aragüés-Peñalba
The growing energy demand of data centers, particularly those offering machine learning services, poses significant challenges to power system stability. As electricity grids become more dynamic due to the increasing penetration of renewable generation, data centers represent a promising but underutilized source of demand-side flexibility. This paper presents a quantitative framework to assess and activate the flexibility potential of data centers by deferring non-critical, high-latency IT workloads. Using real-world workload traces from the Alibaba Cluster Trace dataset, tasks are classified based on observed queuing latency, and a power model is developed to estimate energy consumption associated with different latency groups. A Latency-Aware Deferral for Flexibility (LAD-Flex) strategy is proposed to temporarily reduce modeled IT-side power demand in response to aggregator requests, while maintaining quality of service. The strategy accounts for estimated task duration, latency thresholds, and flexibility window constraints. Results from a one-week case study show that up to 22% of load can be deferred during flexibility windows, and that more than 20% of estimated GPU-side power is attributable to tasks with deferrable latency. Additionally, an optimization framework is introduced to identify the notification period that maximizes the value of flexibility under a time-sensitive pricing scheme. The findings demonstrate that data centers, coupled with an aggregator, can reliably participate in short-notice demand response by leveraging workload-aware deferral strategies, offering both operational and economic benefits.