Haoming Qu, Mohsen Aldaadi, Gregor Verbič
As electric-vehicle (EV) adoption accelerates, unmanaged charging can increase peak demand and net-load volatility, making demand shaping an important operational task. Dynamic retail pricing by an EV aggregator can guide user charging and discharging while preserving aggregator profitability. This paper studies day-ahead EV aggregator pricing under a bilevel leader–follower structure, where the aggregator sets retail prices and EV users respond with charging and discharging schedules. The price–response relation is challenging because discrete charge/discharge/idle decisions, plug-in availability, state-of-charge dynamics, and vehicle-to-grid (V2G) operation must be represented explicitly. To address this issue, we develop a mixed-integer linear programming (MILP)-in-the-loop reinforcement learning (RL) framework. The lower-level MILP models user cost minimization subject to EV operating constraints, while the upper-level learner generates bounded multi-period retail price trajectories with average-price shaping and updates the policy from repeated MILP-induced feedback. The learner uses a prioritized-experience-replay-enhanced Parallel Twin Delayed Deep Deterministic Policy Gradient (PER–PTD3)-style actor–critic implementation to improve sample reuse and training stability. Benchmark simulations under grid-to-vehicle (G2V) and V2G operation across multiple flexibility penetration levels show that the proposed method is competitive with PDDPG, PTD3, and PER–TD3. In a representative 30% V2G setting, it achieves about 5% higher final-window profit and 15% lower cross-seed variability relative to PDDPG, while inducing load shifting. Case studies further show that V2G flexibility strengthens peak shaving and yields smoother intra-day retail price trajectories.