E-Mobility Track
Submission 166
Preference-Conditioned Multi-Agent Reinforcement Learning for Real-Time EV Charging Coordination Under Grid Constraints
03 GIW26-166
Presented by: Syed Irtaza Haider
Syed Irtaza Haider 1, Bingyi Jin 1, Shiwei Shen 1, Razan Habeeb 1, Rico Radeke 1, Frank H.P. Fitzek 1, 2
1 Deutsche Telekom Chair of Communication Networks, TUD Dresden University of Technology, Dresden, Germany, Germany
2 Centre for Tactile Internet with Human-in-the-Loop (CeTI), TUD Dresden University of Technology, Dresden, Germany, Germany

Public electric vehicle (EV) charging stations must make real-time scheduling decisions under stochastic arrivals, heterogeneous charging demands, and time-varying electricity prices. In practical deployments, operator priorities such as profit maximisation and user satisfaction vary with congestion levels and grid conditions, requiring scheduling policies that adapt to changing system states and operational objectives in real time. Optimisation-based scheduling methods rely on fixed objective weights and require repeated re-optimisation when preferences change, limiting their applicability under operational constraints. Existing learning-based approaches often employ value-based action selection across the full preference space, which can lead to abrupt policy changes when priorities shift and reduced Pareto front coverage due to conflicting gradients during training. To address these limitations, we develop a preference-conditioned multi-agent reinforcement learning framework based on a value-decomposition actor-critic architecture. The objective preference space is partitioned into subspaces, and policies are trained sequentially with parameter transfer to improve coverage across competing objectives. At runtime, a rule-based controller switches among the trained policies according to observed queue congestion and grid capacity, enabling adaptation without retraining. Simulation results on real-world ACN-Data, averaged over 50 Monte Carlo runs, show that our approach improves Pareto front coverage compared to baseline methods, reduces peak-hour queue length by up to 25%, improves charging completion rates, and maintains millisecond-level inference latency. The approach can be integrated into existing charging station management systems without requiring changes to hardware infrastructure, enabling practical deployment under real-world operating conditions.