综合智慧能源 ›› 2026, Vol. 48 ›› Issue (7): 68-73.doi: 10.3969/j.issn.2097-0706.2026.07.007

• 电力系统智能化与控制 • 上一篇    下一篇

基于DDPG算法的V2G用户特征识别及定价策略

李明扬(), 包金龙*()   

  1. 华北电力大学 控制与计算机工程学院北京 102206
  • 收稿日期:2025-11-14 修回日期:2026-01-13 出版日期:2026-07-25
  • 通讯作者: * 包金龙(2001),男,硕士生,从事电网安全运行与优化调度方面的研究,120232227281@ncepu.edu.cn
  • 作者简介:李明扬(1983),男,讲师,博士,从事电网安全运行与优化调度方面的研究,limy@ncepu.edu.cn
  • 基金资助:
    青岛特来电新能源科技有限公司技术咨询项目(2206003785-FW02)

V2G user feature recognition and pricing strategy using DDPG algorithm

LI Mingyang(), BAO Jinlong*()   

  1. School of Control and Computer EngineeringNorth China Electric Power UniversityBeijing 102206, China
  • Received:2025-11-14 Revised:2026-01-13 Published:2026-07-25
  • Supported by:
    Qingdao TELD New Energy Technology Company Limited Technical Consulting Project(2206003785-FW02)

摘要:

分布式光伏等可再生能源的大规模并网以及电动汽车数量的快速增长导致电力系统运行面临的不确定性问题日趋明显。随着电动汽车入网(V2G)技术的逐步发展和推广,计及V2G策略的电动汽车充放电站运行优化逐步受到关注。然而电动汽车用户群体的充放电行为是大量个体用户行为特征的聚合,常难以准确获知用户的充放电偏好并据此优化制定V2G回购电价。针对此问题,对电动汽车充放电站与分布式光伏、区域负荷聚合的虚拟电厂(VPP),构建了一种考虑车网互动的优化调度模型,以VPP综合效益最大化为目标,计及电动汽车充放电功率与新能源发电出力等约束。考虑到大量电动汽车用户个体的充放电行为及分布式光伏发电的出力特性难以精确建模,采用基于强化学习的交互式优化调度方法进行求解,并对比分析了若干常用深度强化学习算法的求解性能。进一步,提出了一种电动汽车用户行为特征的识别方法,并在此基础上建立了充放电站运营商对用户V2G放电的定价策略。采用放电荷电阈值与放电差价阈值两个参数,分别描述电动汽车用户对续航里程的焦虑程度和对放电收益的经济性偏好,并提出了辨识这些阈值的系统化流程。将辨识所得阈值引入VPP优化调度模型,提出了电动汽车充放电站中V2G放电电价的定价策略。基于深度确定性策略梯度(DDPG)算法,通过若干典型场景下的仿真算例对电动汽车用户充放电行为的荷电阈值、差价阈值等特征参数进行识别,进而优化充放电站的V2G电价。算例结果表明,电动汽车用户行为的特征差异对VPP总效益具有显著影响,所提出V2G定价策略能更好地兼顾电动汽车用户的响应意愿与系统经济性,显著提升VPP的整体运行效益。

关键词: 电动汽车入网, 分布式能源, 虚拟电厂, 定价策略, 强化学习

Abstract:

The large-scale grid integration of renewable energy sources such as distributed photovoltaic (PV) generation and the rapid growth in the number of electric vehicles (EV) have led to increasingly prominent uncertainties in the operation of power systems. With the gradual development and promotion of vehicle-to-grid (V2G) technology, the optimization of EV charging and discharging station operation considering V2G strategies has gradually attracted increasing attention. However, the charging and discharging behaviors of EV users represent the aggregation of the behavioral characteristics of a large number of individual users, making it difficult to accurately identify users' charging and discharging preferences and optimize V2G electricity purchase prices accordingly. To address this issue, an optimization scheduling model considering vehicle-grid interaction was constructed for a virtual power plant (VPP) consisting of EV charging and discharging stations, distributed PV generation, and aggregated regional loads. This model aims to maximize the comprehensive benefits of the VPP while considering constraints such as EV charging and discharging power and renewable energy generation output. Considering the difficulty in accurately modeling the charging and discharging behaviors of individual EV users and the output characteristics of distributed PV generation, an interactive optimization scheduling method based on reinforcement learning was employed for solving the problem, and the solution performance of several commonly used deep reinforcement learning algorithms was compared and analyzed. Furthermore, a method for identifying EV user behavioral characteristics was proposed, and a pricing strategy for EV charging and discharging station operators regarding users' V2G discharging was developed. Two parameters, namely the discharge state-of-charge threshold and the discharge price difference threshold, were used to respectively characterize EV users' range anxiety and economic preference for discharge benefits, and a systematic procedure for identifying these thresholds was proposed. The identified thresholds were incorporated into the VPP optimization scheduling model, and a pricing strategy for V2G discharge electricity prices at EV charging and discharging stations was proposed. Based on the deep deterministic policy gradient (DDPG) algorithm, the characteristic parameters of EV users' charging and discharging behaviors, including the state-of-charge threshold and discharge price difference threshold, were identified through simulation cases under several typical scenarios, thereby optimizing V2G electricity prices at charging and discharging stations. The results of the case studies indicated that differences in EV users' behavioral characteristics significantly affected the overall benefits of the VPP. The proposed V2G pricing strategy better balanced EV users' willingness to respond and system economic efficiency, thereby significantly improving the overall operational benefits of the VPP.

Key words: vehicle to grid, distributed energy, virtual power plant, pricing strategy, reinforcement learning

中图分类号: