The rapid development of the transportation industry has been accompanied by increasingly severe energy consumption and environmental pollution, particularly in the road transport sector, which accounts for 21.7% of global transportation energy consumption. Therefore, developing effective energy-saving and emission-reduction technologies holds significant importance for environmental protection and energy security. Eco-driving, a technology that reduces energy consumption by intervening in vehicle behavior, has garnered widespread attention [1]. Specifically, methods that apply modern technology to intervene in vehicle behavior and reduce energy consumption during operation have demonstrated significant effectiveness [2, 3]. Based on this, this paper focuses on formulating a longitudinal motion control strategy for eco-driving.
The theoretical foundation of eco-driving speed guidance lies in recognizing that vehicle fuel consumption and emissions are closely linked to driving behavior, particularly cruising speed and the aggressiveness of acceleration and deceleration. Practice demonstrates that energy consumption can be effectively reduced through smooth acceleration, avoiding frequent braking, and maintaining optimal fuel-efficient speeds [4]. Early eco-driving strategies primarily relied on expert experience and intuitive insights to establish a set of driving rules, cultivating habits through driver training. However, due to the inherent unpredictability and forgetfulness of human driving behavior, this training-based approach proved difficult to sustain and yielded limited overall effectiveness. Fortunately, the rise of autonomous driving technology has opened new avenues for achieving long-term, efficient eco-driving [5, 6]. Currently, eco-driving strategies are broadly categorized into optimization-based and learning-based approaches. Optimization-based strategies frequently employ methods like dynamic programming [7] and model predictive control (MPC) [8] to optimize driving behaviors such as acceleration, deceleration, and cruising, aiming to reduce fuel or energy consumption [9]. For instance, one paper integrated deep reinforcement learning into the MPC framework and combined it with a vehicle speed prediction model. This approach not only optimized hydrogen consumption and delayed system aging but also enhanced adaptability and decision robustness across diverse driving environments [10]. However, such approaches generally require precise driving environment models and often involve substantial computational demands. In contrast, learning-based strategies, particularly deep reinforcement learning (DRL), demonstrate strong optimization capabilities and high computational efficiency when handling complex environments. They effectively alleviate performance pressure on onboard computers, making them a current research hotspot in this field [11, 12].
Notably, machine learning—particularly DRL—can flexibly utilize multi-source traffic information and establish a direct mapping from raw data to control actions [13, 14]. It is precisely due to DRL's strong adaptability to eco-driving problems that numerous scholars have conducted related research on eco-driving strategies from diverse perspectives [15]. For instance, Tong et al. [16] treated DRL as a high-dimensional function calculator, integrating information such as traffic signals, preceding vehicles, and traffic conditions to optimize vehicle longitudinal speed profiles, thereby achieving energy-saving and emission-reduction goals. Considering the complexity of real-world environments, He et al. [17] further designed a DRL-based eco-driving method for urban scenarios, combining low energy consumption with high computational efficiency. Addressing the limited adaptability of existing eco-driving methods, Shi et al. [18] constructed a human driving experience dataset by collecting human driving behavior trajectories. They employed imitation learning techniques to learn eco-driving strategies from these trajectories, effectively enhancing the method's applicability. Furthermore, Zhang et al. [19] addressed the core requirements of eco-driving by designing a reward function that integrates safety, efficiency, and energy consumption using DRL technology, significantly reducing the probability of unnecessary vehicle stops.
DRL combines the perceptual capabilities of deep learning with the decision-making abilities of reinforcement learning, enabling agents to learn optimal policies through interaction with their environment [20, 21]. In the field of eco-driving, DRL has been widely applied to tasks such as vehicle speed optimization, energy consumption management, and handling complex traffic scenarios [22]. The evolution of DRL algorithms has provided a rich toolkit for developing eco-driving strategies: Deep Q-Network (DQN), a landmark achievement in DRL, first enabled direct policy learning from high-dimensional inputs [23]. Subsequently, algorithms like Deterministic Policy Gradient (DPG), Trust Region Policy Optimization (TRPO), and Deep Deterministic Policy Gradient (DDPG) were developed to address control problems in continuous action spaces [24]. Furthermore, extension techniques such as hierarchical DRL, transfer DRL, offline DRL, and safe DRL provide effective solutions for tackling key challenges like sparse rewards, enhancing learning efficiency, and ensuring driving safety.
Despite significant progress in the field of eco-driving, DRL still faces numerous challenges, particularly in highly complex, dynamically uncertain real-world urban traffic environments where existing methods often lack universality and adaptability [25]. Simultaneously, the training process for DRL strategies is typically time-consuming and prone to issues such as convergence difficulties, poor stability, and complex parameter tuning [26].
Rule-based reinforcement learning methods typically use model outputs as constraints or rewards, which limits the freedom of exploration. Imitation learning relies on large amounts of high-quality human driving data; it struggles to cover all scenarios, is prone to learning undesirable driving behaviors, and cannot be optimized online. Traditional model-based reinforcement learning methods typically incorporate reference values as additional state inputs but fail to establish a closed-loop safety mechanism. The method proposed in this paper does not require a dataset; instead, it directly generates safety priors from the model, allowing the agent to autonomously optimize within a safe range, resulting in stronger generalization and greater stability. In contrast, this paper establishes a unified Intelligent Driver Model (IDM) speed guidance model applicable to both following and intersection scenarios, overcoming the limitation of traditional IDM models that are only suitable for following scenarios. It achieves unified speed guidance for both following and signal-controlled intersections, automatically switches strategies based on distance and traffic light phases, and provides a safe and conservative reference speed as a learning benchmark—features not found in existing guided DRL methods; Simultaneously, a triple guidance mechanism is constructed: the IDM reference speed serves as a prior to narrow the strategy search space; IDM safety rules act as hard constraints to ensure driving safety; and IDM trajectories limit the agent's exploration scope, thereby avoiding ineffective exploration and accelerating convergence.
To address these challenges, this paper proposes a model-guided deep reinforcement learning framework aimed at enhancing the practicality and reliability of eco-driving strategies. Specifically, we introduce a speed-guiding algorithm based on the IDM model [27], integrating it as prior knowledge with the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. This approach does not replace DRL but serves as its guidance module, providing a safe and reasonable speed reference benchmark. This significantly shortens training cycles, stabilizes the learning process, and safeguards the baseline of driving safety. By deeply integrating IDM's rule-based knowledge with DRL's environmental adaptability and training/validating in real traffic scenarios, this approach effectively enhances the energy-saving and emission-reduction performance of autonomous vehicles' eco-driving strategies while ensuring driving safety and efficiency [28]. This framework achieves vehicle energy conservation and emission reduction targets under strict safety guarantees and minimal impact on traffic flow efficiency. Training and validation in real-world traffic environments further improve the algorithm's adaptability.
This paper is conducted in an intelligent connected environment, assuming the ego vehicle can perceive key information such as the status of the preceding vehicle and traffic signals in real time. The research focuses on longitudinal speed planning, a critical factor directly impacting energy consumption [29]. Under single-vehicle operation without lane changes, driving behavior is primarily constrained by the preceding vehicle (in following scenarios) and traffic signals (at intersections). Therefore, establishing a unified model for following and intersection passage is fundamental to formulating rational speed strategies [30]. However, traditional modeling approaches often prove complex and inflexible when integrating these two scenarios. To address this, this paper innovatively combines the guidance signals generated by the IDM model with the adaptive optimization capabilities of DRL. This creates a unified speed guidance mechanism that handles both following and intersection scenarios. The speed guidance algorithm manages the fundamental logic of safe following and intersection passage, while the DRL agent learns refined energy-saving driving strategies on this foundation. The ultimate goal is to achieve significant optimization of vehicle energy consumption while ensuring acceptable trip efficiency.
The overall framework of the eco-driving strategy is illustrated in Figure 1. Specifically, after receiving the necessary information, the vehicle trains a vehicle speed planning agent through extensive data interactions. This agent achieves the vehicle's energy-saving optimization control objectives without unduly compromising driving efficiency. Once eco-driving control actions are generated, they interact with the environment in real time and provide guidance for generating subsequent decisions. Throughout this process, the vehicle's navigation tasks are accomplished through strategic lane-changing decisions.
To achieve precise and stable control of vehicle speed, this paper employs a DRL algorithm with a continuous action space. After a comprehensive comparison, the TD3 algorithm was selected as the foundation for the longitudinal speed control strategy due to its advantages in accurate evaluation and stable convergence. The core objective of training the TD3 algorithm is to develop an agent responsible for speed decision-making.
This paper focuses on pure electric vehicles, establishing their dynamic models and conducting detailed modeling of the powertrain system. Vehicle parameters are shown in Table 1. This modeling work lays a solid foundation for subsequent optimization of eco-driving and energy management strategies, thereby supporting systematic evaluation and performance validation of the proposed methodology [31].
The implementation of energy management primarily relies on the vehicle's longitudinal dynamics model. It calculates instantaneous power requirements based on real-time speed and acceleration data, and employs energy management strategies to allocate power efficiently. The driving force and power demand can be calculated based on the following equations:(1)(2)(3)where Ft represents the driving force of the vehicle, m represents the mass of the fuel cell hybrid bus, θ represents the slope angle of the road surface, Cd represents the vehicle's aerodynamic drag coefficient, A represents the frontal area of the vehicle, v represents the vehicle's speed, δ represents the rotational mass conversion factor of the entire vehicle. Pdemand represents the vehicle's power demand, Twheel represents the wheel torque, and rwheel renpresents wheel radius.
Based on the drive chain of the powertrain system, the required torque of the motor can be calculated using the following equation according to the wheel torque of the fuel cell hybrid bus:(4)where Tmotor represents the motor's required torque, ig represents the total transmission ratio from the motor to the wheel.
Based on the relationship between power and torque, the required power can be calculated using the following equation:(5)where nwheel represents the wheel speed.
The efficiency of the motor was measured based on experimental data, with the equation being:(6)where ηmoto represents engine efficiency, fmotor represents the fitting function for motor torque versus speed based on measured data, and nmotor renpresents engine speed.
Therefore, the required power of the motor can be calculated as follows:(7)
Based on the actual operating conditions of the vehicle, the motor power should be limited. Therefore, we impose boundary constraints on it:(8)where represents the minimum power of the restricted motor, and represents the maximum power of the restricted motor.
Based on the vehicle dynamics and powertrain models established above, the following section presents the proposed eco-driving control methodology.
This chapter details the overall framework of an eco-driving control strategy based on DRL. It proposes a hybrid control strategy combining a speed guidance algorithm with deep reinforcement learning, achieving more efficient and stable eco-driving control by integrating the deterministic advantages of model prediction with the adaptive optimization capabilities of reinforcement learning. Specific components include: defining the state space and action space for longitudinal speed control; designing a speed guidance algorithm based on an improved IDM; constructing a multi-objective reward function; and establishing a driving behavior decision mechanism. This framework aims to optimize vehicle trajectories through intelligent decision-making, minimizing energy consumption in pure electric vehicles while ensuring driving safety and traffic efficiency.
The decision-making of intelligent agents relies on a precise perception of the traffic environment. To comprehensively characterize key factors influencing vehicle speed planning, the TD3 policy state space proposed in this paper incorporates the reference speed from the speed guidance algorithm, ego vehicle-related states, the driving state of the preceding vehicle, and the status of the traffic signal ahead. The specific configuration is as follows:
Among these, vref represents the reference speed, calculated by the speed guidance algorithm proposed in this paper, providing prior knowledge guidance for the convergence of the DRL agent. Using the model-predicted reference speed as a state input effectively narrows the policy search space, preventing the DRL agent from blindly exploring during the early training phase and significantly accelerating convergence. vego, aego, and i represent the ego vehicle's velocity, acceleration, and road gradient, respectively, enabling the agent to perceive its own driving state. vpre, apre, and dpre represent the speed, acceleration, and distance to the nearest preceding vehicle relative to the ego vehicle, forming the basis for formulating the following strategy. Considering variations in traffic intersections, the intersection passage state is characterized by the phase of the preceding traffic signal tlight, the remaining time in the current state Tlight, and the ego vehicle's distance to the stop line dlight, aiding in the formulation of intersection passage strategies.
In the action space design, this paper selects the longitudinal acceleration of the vehicle as the continuous output action of the agent, with its variation range constrained to adrl∈ [‒2.6, 2.6] m s−2. Compared to directly controlling speed, controlling acceleration effectively avoids violent oscillations in the speed trajectory, reduces reliance on additional constraints, and better aligns with the characteristics of real vehicle actuators. The agent outputs acceleration commands at each control step interval ΔT, updating the vehicle speed while ensuring it does not exceed the road speed limit vlimit. Vehicle speed can be calculated based on the following equations:(9)where ΔT denotes the control step size for the intelligent agent, which is set to 0.1 s in this paper to achieve real-time precise control of the ego vehicle.
To guide the learning process of the DRL agent, provide a high-quality initial policy, and serve as a performance benchmark, this paper designs an improved IDM applicable to both following scenarios and intersection scenarios, as shown in Figure 2. This algorithm not only functions as the core speed planning module but also serves as a steering signal input to the state space of the DRL agent, effectively accelerating the convergence process of reinforcement learning.
The speed guidance algorithm accelerates reinforcement learning convergence through three mechanisms. First, it provides goal orientation—the reference speed vref output by the algorithm offers the DRL agent a clear optimization direction, preventing aimless random exploration during early training. Second, it reduces the action space—agents need not learn basic following and stopping behaviors from scratch, but instead perform optimization fine-tuning based on model predictions. It mitigates reward sparsity by serving as an intermediate reward signal, alleviating potential sparsity issues in the original reward function.
In pure follow-and-drive scenarios, vehicle acceleration is calculated using the classical IDM model:(10)(11)where amax represents the maximum acceleration of the ego vehicle, represents the desired speed of the ego vehicle, S and Δv represent the distance and speed difference between the leading and following vehicles, respectively. represents the desired following distance, while s0 represents the minimum spacing during stationary conditions. T and bcomf are parameters ensuring vehicle safety and comfort, set to 1.5 and 2, respectively, in this paper.
The above Equation (10) describes only the calculation of vehicle acceleration under the following conditions. When approaching a signalized intersection, the vehicle must determine its passage strategy based on the traffic signal status. Converting this to an intersection passage requires considering two key parameters: maximum travel distance dlight_max and minimum stopping distance dstop_min. The maximum travel distance dlight_max is the farthest distance the vehicle can travel within the remaining time of the current signal phase Tlimit. The minimum stopping distance dstop_min is the shortest distance required for the vehicle to decelerate from its current speed vego to a stop using maximum deceleration amax. Both parameters are calculated using the following equations:(12)(13)(14)where vlimit represents the road speed limit, while dstop denotes the minimum distance requirement from the stop line when parking. Additionally, the minimum following distance dgap is defined. To enhance decision robustness, a distance margin dmargin is introduced; a larger value indicates a more conservative driving strategy. These parameters are intuitively illustrated in Figure 2. Combined Equations (10, 11) can further yield the vehicle following reference speed vfollow and the intersection stopping reference speed vstop. The specific calculation process is detailed in algorithm 1 (lines 1–8) of Table 2.
The activation of coasting mode and intersection passage mode is determined by the relative magnitude of the distance between the vehicle and the preceding vehicle dpre and the distance to the stop line dlight: When, the intersection passage mode is initiated, otherwise, the coasting mode is adopted.(15)
In intersection passage mode, the reference speed vlight is determined according to the traffic signal phase tlight using the following rules:(16)
The above equation defines the intersection passage rules, determining whether the vehicle can safely pass through the intersection based on the traffic light status. When the ego vehicle can proceed through the intersection, it adopts the follow-vehicle reference speed vfollow; when it cannot proceed, it adopts the stop reference speed vstop. Based on the phase of the forward traffic signal, the following three scenarios can be summarized:
When the signal is green, calculate the maximum travel distance dlight_max the ego vehicle can cover within the remaining phase time. If this distance exceeds the required travel distance (dlight + dmargin), the vehicle can pass through the intersection; otherwise, it cannot.
When the traffic light ahead is yellow, this phase is typically followed immediately by a full red light phase. Vehicles should come to a complete stop before the stop line at this time.
When the traffic light ahead is red, calculate the ego vehicle's minimum stopping distance dstop_min. If it exceeds the required stopping distance (dlight - dmargin), the vehicle must immediately execute a stopping strategy. Otherwise, it continues following the preceding vehicle.
Ultimately, the comprehensive reference speed output by the speed guidance algorithm is:(17)
The specific calculation and determination process of the reference speed is illustrated in Table 3: reference speed calculation and determination algorithm pseudocode. The derived reference speed evolves from the IDM model while accounting for safety constraints, fully meeting traffic safety requirements. Consequently, it can serve as a driving safety constraint for subsequent DRL rewards. Furthermore, this algorithm is defined conservatively, permitting the ego vehicle to moderately exceed the set reference speed.
The design objective for longitudinal speed control proposed in this paper to achieve eco-driving tasks is to reduce vehicle energy consumption while ensuring safe operation and minimizing impacts on traffic efficiency. To meet this objective, the reward function design should also incorporate the aforementioned three aspects.
To achieve eco-driving objectives, ensuring driving safety is paramount. The designed speed guidance algorithm fully complies with safe driving standards, making it a reliable safety benchmark. Therefore, this reward is set as:(18)
The purpose of setting an upper limit for the reward is to enhance the stability of algorithm convergence.
To ensure traffic efficiency, the ego vehicle's speed should not be excessively slow. Therefore, the following reward is designed:(19)
For pure electric vehicles, reducing energy consumption essentially means lowering the output power of the power battery. Therefore, with energy-saving design objectives in mind, the following incentives have been established:(20)where Preq represents the power demand calculated based on vehicle longitudinal dynamics, η denotes the overall efficiency during the power transmission process.
DRL-based eco-driving algorithms typically face challenges in training and convergence. Previous studies achieved convergence only after extensive parameter tuning and experimentation. In this research, we accelerate convergence through a velocity-guided strategy. The specific design of the guiding reward is as follows:(21)
Ultimately, the reward function is designed as a weighted sum of the aforementioned rewards to ensure that the designed longitudinal velocity planning strategy meets the intended objectives:(22)where w1–w4 represent the weighting coefficient for each reward, used to measure the degree of bias toward various consideration factors. This composite reward function, through the adjustment of weighting coefficients, can flexibly adapt to the optimization priorities in different scenarios. It serves as the core guiding mechanism for the DRL agent to learn eco-driving strategies that satisfy multi-objective constraints.
The strategic objective of strategic lane changes is to ensure the ego vehicle remains in the correct lane without deviating from its intended route. This paper employs such a lane-changing strategy to achieve vehicle navigation. Specifically, it utilizes the strategic lane change strategy defined in Simulation of Urban Mobility (SUMO), which is a lane-changing strategy determined by the degree of lane change urgency. A strategic lane change is deemed urgent when the following relationship holds:(23)where d denotes the distance from the intersection to the ego vehicle on a road requiring strategic lane changes. o represents the discount factor applied to account for target lane occupancy. lookAheadSpee is the ego vehicle's expected average speed from its current position to the intersection. bestLaneOffset indicates the number of lane changes required for the ego vehicle to reach the target lane, where a value less than 0 signifies a left lane change and greater than 0 indicates a right lane change. f* denotes the time required for the ego vehicle to complete a lane change, set to 5 s in this paper.
When the driver has the motivation to change lanes, it is also necessary to consider whether the current traffic conditions meet the requirements for lane changing. Specifically, both the vehicle ahead and the vehicle behind on the target lane must satisfy the lane-changing criteria, meaning they must be farther away than the lane-changing distance requirement. This distance is defined by Equation (11).
This chapter aims to comprehensively evaluate the performance of the proposed model-guided deep reinforcement learning (MG-DRL) eco-driving strategy. First, the setup of the training and simulation environment, along with key hyperparameters, is introduced. Subsequently, multiple experiments determine the optimal value of the energy-saving weighting coefficient in the reward function. Subsequently, a systematic comparison and in-depth analysis is conducted between the MG-DRL strategy (proposed), the baseline method (improved IDM), and the contrastive strategy (contrast) based on DDPG. This evaluation spans multiple dimensions: convergence, overall energy efficiency and traffic efficiency, traffic capacity and trajectory smoothness in typical scenarios, and control domain robustness. The goal is to validate the superiority and generalization capability of the proposed method.
To comprehensively evaluate the performance of the proposed eco-driving strategy, this section details the training and testing environments, parameter configurations, and hardware platforms. All experiments were conducted in a unified computational environment to ensure comparability and reproducibility of results. All experiments in this paper were performed on a single computer equipped with a 13th Gen Intel® Core™ i7-13700F CPU and an NVIDIA GeForce RTX 3050 OEM graphics card.
The experiment employed the TAVF-Hamburg traffic simulation environment, developed based on Hamburg's real-world road network. This environment incorporates complex urban road networks, signalized intersections, and dynamic traffic flows, enabling high-fidelity simulation of real-world driving scenarios. As shown in Figure 3, a representative route was selected for training and validation. The experimental section spans 2,700 meters, featuring seven signal-controlled intersections and encompassing diverse road types and varying traffic flow densities. Road speed limits vary between 50 and 60 km h−1, with gradient changes ranging from –5% to 5%. During training, to enhance the strategy's generalization capability, each training iteration randomly resets the traffic environment using a seed. This includes resetting the initial position and speed of the vehicle under test and randomly generating background traffic flow.

To ensure the autonomous vehicle can complete its designated navigation tasks, SUMO's built-in strategic lane-changing policy is enabled throughout both the training and testing phases of the longitudinal velocity control strategy. The training process comprised 300 episodes. At the start of each episode, both the SUMO simulation environment and the autonomous vehicle's state were reset. The vehicle had to complete a 2,700-meter route within a simulation time not exceeding 1,000 s or reach the destination early. At the end of each episode, metrics such as cumulative reward, total energy consumption, and travel time were collected. To evaluate the final performance of the trained strategy, network parameters were frozen after training. The strategy was then tested across 10 independent, unseen random traffic scenarios, with average performance metrics recorded to assess its generalization capability. All experiments were repeated five times, and the average result was taken as the final outcome to mitigate random effects.
Training consists of 300 episodes. At the beginning of each episode, the simulation environment is reset with random seeds, including the ego vehicle's initial position and velocity, background traffic flow, and signal phase offsets. After training, the actor network is frozen, and the policy is evaluated on 10 unseen random scenarios. Each test case is run five times, and the results are averaged. All experiments are repeated five times with different random seeds.
The specific parameters are shown in Table 4. The learning rate 1 × 10−5 is selected after grid search to balance convergence speed and stability. The discount factor 0.95 aligns with the time scale of driving tasks. The replay buffer capacity 1 × 105 stores approximately 100 episodes of transitions. The batch size 256 balances computational efficiency and gradient estimation stability.
To validate the effectiveness of the proposed MG-DRL framework, this section conducts an in-depth analysis of the algorithm's convergence process. Figure 4 illustrates the training curve comparison between the MG-DRL policy (proposed) based on TD3 and the contrast policy (contrast) based on DDPG under the same environment.

In terms of algorithmic convergence stability, the proposed MG-DRL strategy (proposed) exhibits a more stable convergence trend. Although divergence occurred around 150 iterations, convergence was ultimately achieved by the 193rd iteration, whereas the comparison strategy only converged at the 252nd iteration. The MG-DRL reward curve rises steadily overall, with only a slight fluctuation in the middle of training, followed by a rapid recovery; in contrast, the DDPG curve shows significant fluctuations and an unstable learning process. This acceleration effect is primarily attributed to the reference velocity vref provided by the velocity guidance algorithm. By supplying high-quality behavioral priors, shaping the reward function, and guiding exploration directions, it effectively mitigates the sparse reward problem. This narrows the policy search space from learning fundamental driving behaviors to optimizing energy consumption within safety boundaries. Despite the increased computational complexity of the TD3 algorithm due to its dual-critic architecture, its training time was only 9.48 hours—nearly equivalent to the 9.43 hours of the DDPG algorithm—ensuring computational efficiency. Algorithm convergence results demonstrate that the proposed method achieves higher terminal rewards. The average round reward of the MG-DRL policy significantly exceeds that of the DDPG baseline policy. This superior terminal reward value indicates that the MG-DRL policy strikes a better balance among the three objectives of safety, efficiency, and economy. This finding aligns with the actual driving performance test results presented in subsequent sections.
It is important to clarify that the TD3 agent does not simply track the IDM reference speed. Unlike the IDM, which outputs a deterministic speed based only on instantaneous following gaps and signal status, the TD3 agent learns a forward-looking acceleration policy that anticipates phase changes and preceding vehicle trends. This enables a "pulse-and-glide" eco-driving pattern that cannot be achieved by minor adjustments to the IDM. The performance gains—7.96% energy reduction with only 0.44% travel time increase—directly demonstrate that the agent learns a superior policy beyond mere reference tracking. Furthermore, the training curve in Figure 4 shows that the cumulative reward continues to rise after convergence, indicating active optimization rather than passive fitting.
The weighting of reward components decisively shapes the behavioral preferences of deep reinforcement learning agents. To achieve an optimal balance among competing objectives such as safety, efficiency, and energy conservation, this paper systematically determined the final weight configuration through experimental validation, as illustrated in Figure 5. Given the importance of driving safety, the safety reward weight w1 is assigned the highest priority. Ensuring that the vehicle does not collide or violate traffic signal rules is a prerequisite for all decisions, and its ratio to the efficiency reward weight w2 is set at 3 : 1. The guidance reward w4 is primarily used to accelerate convergence during the early stages of training. To avoid unduly influencing the final policy, its value is set to a very small constant. On this basis, the energy-saving reward weight w3 becomes the key parameter for adjusting the policy's bias.
To determine the optimal value of w3, we conducted extensive comparative experiments within the interval w3∈ [2, 6] and evaluated the performance of each configuration across 10 distinct traffic scenarios. The results are shown in Figure 5. Experiments indicate that when w3∈ [2, 3.5], energy savings significantly improve as the energy efficiency coefficient increases. However, travel time exhibits substantial variation, as the strategy overly emphasizes travel efficiency, resulting in unstable energy savings and insufficient generalization capability. When w3∈ [5, 6], the strategy becomes overly energy-focused, leading to significantly increased travel times and insufficient driving efficiency. However, within the range w3∈ [3.5, 5], the strategy achieves a favorable balance between energy savings and efficiency. Specifically, at w3 = 4, the strategy demonstrates optimal overall performance, with the efficiency reward and energy savings reward exhibiting an approximate ratio of 10 : 1. Moreover, the sensitivity analysis reveals potential failure modes. When the guidance weight w4 exceeds 0.5, the agent tends to merely replicate the IDM reference, losing the ability to optimize beyond the model. When the safety weight w1 is reduced below 2, the agent occasionally violates traffic signals or follows too aggressively, which is unacceptable for real-world deployment. Detailed results are presented in Table 5. Therefore, the chosen weights (w1 = 3, w2 = 1, w3 = 4, w4 = 0.1) provide a robust balance across all scenarios tested.
When the weights deviate from the optimal range, the following trade-offs and failure modes may occur.
Consequently, the reward function weights finalized in this paper ensure that the agent can simultaneously pursue the optimization objectives of high travel efficiency and low energy consumption while strictly maintaining safety.
To comprehensively evaluate the overall performance of the proposed MG-DRL eco-driving strategy, this section systematically compares it with the baseline algorithm (improved IDM) and a DDPG-based competitor across 10 independent random traffic scenarios. The evaluation encompasses two core metrics: energy consumption and travel time. Detailed results are presented in Table 6.
This paper focuses on validating the effectiveness of the proposed framework in the typical scenario of urban arterial roads. Furthermore, the proposed MG-DRL strategy does not rely on any fixed network topology. The state space only includes local observations (ego vehicle state, preceding vehicle state, signal phase, distance to stop line, and road slope), excluding global map geometry or fixed signal timing. During training, each episode randomly resets initial positions, initial speeds, background traffic density, and signal phase offsets.
Overall, the baseline algorithm exhibits the highest average energy consumption due to its lack of global optimization capabilities. Although it achieves the shortest travel time, its efficiency advantage is extremely limited. Both the comparison strategy and the proposed strategy effectively reduce energy consumption, with average energy savings rates of 7.29% and 7.96%, respectively. Regarding travel time, the proposed strategy exhibits an average increase of only 0.44%, which is nearly negligible. The average travel time increase of 0.44% is practically negligible, falling within normal driving variability such as traffic light waiting times or random queue lengths. This minor increase is a reasonable trade-off for achieving 7.96% energy savings, as the strategy prioritizes smooth acceleration, coasting, and reduced harsh braking. In contrast, the DDPG baseline incurs a 2.80% travel time increase with lower energy savings, confirming that the MG-DRL achieves a superior balance among safety, efficiency, and energy optimization. In contrast, the comparison strategy shows an average increase of 2.80% with significant fluctuations in certain scenarios. Notably, both learning-based strategies demonstrate performance fluctuations in Case 6 and Case 8: the comparison strategy experiences increased energy consumption and significantly prolonged travel time in Case 6, while the proposed strategy also exhibited a slight increase in energy consumption in Case 8. These anomalies primarily stemmed from complex, sudden changes in traffic flow under specific conditions, indicating that while pursuing overall generalization, the strategies still have room for improvement in adapting to individual extreme scenarios. In contrast, the proposed strategy achieved synergistic optimization of energy consumption and efficiency in the vast majority of cases, demonstrating superior balancing capabilities and generalization robustness.
The proposed MG-DRL strategy achieves nearly 8% energy savings with minimal impact on traffic flow efficiency. Its overall performance significantly outperforms both the baseline algorithm and the DDPG comparison strategy, validating its effectiveness and advanced capabilities in addressing multi-objective optimization challenges in eco-driving.
All ablations were run under the same environment and hyperparameters, each repeated 5 times with different random seeds. The results are shown in Table 7.
The full model outperforms both ablations: compared to Ablation A, the speed-guiding mechanism accelerates convergence by 28.5% and reduces energy consumption by 6.0%, confirming its critical role in achieving faster convergence and better final performance. Compared to Ablation B, where both variants incorporate speed guidance, TD3 converges faster and achieves lower energy consumption than DDPG, demonstrating that TD3 is superior for the eco-driving task. The full model achieves the best overall performance, thereby revealing a clear synergy between the guiding mechanism and the TD3 algorithm.
Based on the TAVF-Hamburg simulation platform, this paper constructed three typical test scenarios by configuring traffic flows and signal timing, The heterogeneous traffic flow scenario (Case 1) includes background vehicles with three driving styles and speed perturbations; the dense urban interaction scenario (Case 2) uses a test section with seven consecutive intersections, a minimum spacing of 150 meters, and random phase differences with cycle times ranging from 60 to 90 s; and the non-ideal signal scenario (Case 3) features non-standard signals with a 15-second green phase, a 5-second yellow phase, and randomly varying cycle times. Detailed results are presented in Table 8.
In heterogeneous traffic flow scenarios, the average energy consumption of IDM was 0.324 kWh; the DDPG strategy exhibited performance degradation due to disturbances, while MG-DRL achieved a 4.32% reduction in energy consumption and a 2.10% reduction in travel time. This indicates that while IDM can maintain basic safety, its conservative nature limits its optimization potential when faced with heterogeneous disturbances; conversely, MG-DRL, guided by IDM, effectively suppresses the volatility of the learning-based strategy. In dense urban interaction scenarios, due to its lack of global foresight, IDM's average energy consumption reached 0.371 kWh, significantly higher than MG-DRL's 0.334 kWh. By learning early coasting and phase-coordination strategies, MG-DRL achieved smoother driving trajectories in multi-intersection continuous scenarios, validating its global optimization capability within complex road networks. In scenarios with non-ideal traffic signal behavior, signal non-ideality compressed the energy-saving optimization space, causing MG-DRL's energy-saving rate to drop to 6.83%; however, it still maintained a positive benefit, demonstrating the stabilizing role of IDM guidance under uncertain signal conditions.
Although IDM has obvious limitations in complex scenarios, within the MG-DRL framework, it serves as a state reference and safety constraint, providing a stable learning benchmark for the DRL agent. Under IDM guidance, the agent can learn optimization strategies that deviate from conventional rules, thereby achieving the coordinated optimization of energy consumption and travel efficiency.
The background traffic constructed in this paper was generated using random seeds, with average traffic densities ranging from 120 to 280 vehicles per hour per lane. These scenarios cover a wide range of traffic conditions, from free-flowing traffic to near-saturation, and are categorized into three types based on density: low-density scenarios (average density < 150 vehicles/hour/lane), medium-density scenarios (150–220 vehicles/hour/lane), and high-density scenarios(> 220 vehicles/hour/lane). Detailed results are presented in Table 9.
Analysis of the above data shows that the MG-DRL strategy achieves positive energy-saving effects across all traffic densities, with a particularly high energy-saving rate of 11.95% in high-density scenarios. Furthermore, travel time is shorter than that of the IDM, demonstrating its significant advantages under congested conditions and proving its generalizability.
All experiments in this paper were conducted on the TAVF-Hamburg simulation platform. The test cases introduced systematic variations in traffic density, covering typical scenarios found in real-world traffic environments. The MG-DRL strategy achieved fuel savings in all cases, with an average fuel savings rate of 7.96% and a travel time increase of only 0.44%, fully demonstrating its robustness and universality under various traffic conditions.
To thoroughly investigate the micro-level behavioral differences among various strategies in typical traffic scenarios, this section focuses on Case 7 to conduct a detailed analysis of the spatio-temporal trajectories of the three strategies. Figure 6 illustrates the spatio-temporal distribution relationships of the self-vehicle under the baseline method, the DDPG comparison strategy, and the proposed MG-DRL strategy.

Overall, all three strategies enabled the vehicle to complete the entire journey, but the proposed MG-DRL strategy demonstrated the smoothest and most continuous driving characteristics. Figure 6A reveals that both the comparison strategy and the proposed strategy respond promptly to the red light at the upcoming intersection. However, they adopt fundamentally different control approaches: the comparison strategy, due to inadequate initial speed planning, fails to decelerate smoothly within a safe distance and is ultimately forced to stop and wait before the stop line. In contrast, the proposed MG-DRL strategy achieves seamless passage without stopping by precisely adjusting speed while ensuring safety, demonstrating outstanding continuity of travel.
Further analysis of the baseline method's performance in Figure 6B, C reveals that this rule-based approach exhibits clear limitations due to its lack of holistic perception and prediction capabilities for the traffic system. In Figure 6B, the baseline method barely clears the intersection during the final moments of the green light, with decision boundaries stretched to the limit. In Figure 6C, it misses the opportunity to proceed due to an inability to adjust speed in time, resulting in an unnecessary stop. Both scenarios highlight the insufficient adaptability of purely model-based approaches in dynamic environments.
The superiority of the proposed MG-DRL strategy stems from its hybrid architecture that combines model guidance with data-driven learning. The speed-guided algorithm provides a decision-making framework compliant with safety regulations, while the TD3 agent learns more refined and forward-looking speed adjustment strategies on this foundation. This collaborative mechanism enables the agent to anticipate traffic state changes in advance. Through smooth acceleration or deceleration, it avoids energy loss and discomfort caused by emergency braking while maximizing utilization of green light time windows. Consequently, it significantly enhances traffic flow and efficiency while ensuring safety.
Beyond individual vehicle performance, the MG-DRL strategy contributes to overall traffic stability. By producing smoother speed profiles and gentler acceleration/deceleration, it reduces frequent stop-and-go oscillations that typically propagate shockwaves upstream. Compared to the DDPG baseline, the proposed method reduces the standard deviation of acceleration by approximately, thereby mitigating disturbance propagation and enhancing traffic continuity. This system-level benefit is particularly valuable in dense urban networks where local perturbations can amplify into widespread congestion.
In summary, the proposed MG-DRL strategy achieves intelligent, smooth, and efficient passage through intersections by integrating the rule-based nature of the fusion model with the adaptability of the learning component. This approach not only reduces vehicle energy consumption but also contributes positively to enhancing the overall operational efficiency of traffic flow.
The vehicle's speed and acceleration trajectory directly determine driving energy consumption and ride comfort. Figure 7 illustrates the speed-time curves and corresponding acceleration distributions for three strategies under Case 7, revealing the core differences in dynamics between different control strategies.

During the initial phase of the journey, the acceleration behavior of the three strategies was largely consistent. However, as the vehicle approached the second signal-controlled intersection, differences between the strategies began to emerge. The baseline method exhibited the most aggressive acceleration and deceleration behavior, with the highest absolute acceleration values throughout the entire trip. These frequent and abrupt speed changes directly contributed to its high energy consumption and inevitably negatively impacted ride comfort. The speed curve of the DDPG comparison strategy exhibited noticeable fluctuations and multiple unnecessary acceleration-deceleration transitions. This indicates that the strategy failed to fully learn a smooth, forward-looking driving approach. Its control actions remained, to some extent, reactive responses to environmental changes rather than proactive planning based on long-term benefits. Consequently, both its energy consumption optimization and driving smoothness were unstable.
In contrast, the proposed MG-DRL strategy exhibits optimal dynamic characteristics. Its velocity curve is smooth and continuous, with significantly reduced and more gradual acceleration variations. This "Pulse-and-Glide" velocity profile is a hallmark of eco-driving: minimizing unnecessary braking, utilizing vehicle inertia for coasting, and employing gentle yet effective acceleration when required. This control mode not only directly reduces the instantaneous power demand defined by Equation (2), thereby conserving energy; more importantly, the gradual deceleration process provides the optimal operating range for the regenerative braking system of pure electric vehicles, maximizing the recovery of braking energy.
The underlying mechanism by which the strategy proposed in this paper achieves energy savings is primarily the reduction of unnecessary braking. Figure 7 shows that the baseline method frequently exhibits an "accelerate-to-brake" pattern before intersections, whereas the proposed strategy adopts a "predict-coast-light-brake" pattern. Taking the second intersection in Case 7 as an example, the baseline method maintains a relatively high speed 50 meters from the stop line and then performs an emergency stop; in contrast, the proposed strategy begins coasting 100 meters away and passes through the green light window at approximately 25 km h−1.
The MG-DRL strategy successfully acquired a smooth, forward-looking, and optimized speed control law through deep reinforcement learning. While significantly reducing energy consumption, it also delivers enhanced ride comfort, achieving the dual objectives of eco-driving and premium driving.
The control time domain is a critical parameter affecting the performance and feasibility of real-time control strategies. To validate the robustness of the proposed algorithm and investigate the impact of control granularity on performance, this section tests how the performance of each strategy varies relative to the baseline algorithm as the control time domain changes within the range of 0.1 to 1.0 s. The results are shown in Figure 8.
When the control time window is 0.1–0.3 s, the strategy performs stably with fine-grained control. In the 0.4–0.7 s range, performance declines slightly but remains acceptable; moreover, the longer time step encourages the agent to learn more forward-looking planning strategies, further expanding the energy-saving advantage over the comparison method. In the 0.8–1.0 s range, performance drops significantly, as the control update frequency struggles to adapt to dynamic traffic changes. The results indicate that the proposed strategy maintains excellent performance across a wide range of 0.1–0.7 s and exhibits good robustness.
As the control time domain expands, the algorithm's control command update frequency decreases, resulting in coarser adjustments to vehicle speed. This leads to varying degrees of reduced tracking accuracy for all strategies against the predetermined optimization target. Specifically, the baseline algorithm, lacking forward-looking energy-saving awareness, exhibits control coarsening manifested as more aggressive acceleration behavior. This leads to a significant increase in energy consumption without a corresponding reduction in travel time, resulting in negligible efficiency gains. In contrast, the DDPG comparison strategy and the proposed MG-DRL strategy, which incorporate energy-saving optimization objectives, do not exhibit worsening energy consumption with increasing control time horizons. Instead, they show a trend of relatively improved energy efficiency in certain time windows (e.g., 0.3 to 0.7 s). This occurs because extended control horizons compel the agent to learn more foresighted "pulse-and-glide" strategies. By planning velocity curves over longer windows with a single decision, it avoids the frequent, minor, and inefficient velocity adjustments common in short horizons. However, this trend is not monotonically increasing. When the control time horizon extends to 0.9 s, all strategy performances exhibit abnormal fluctuations: the baseline algorithm's performance occasionally improves, weakening the relative energy savings of the learning-based strategies and increasing travel time. This reveals a complex nonlinear relationship between control granularity, environmental randomness, and strategy performance. Excessively long control time horizons amplify disturbances caused by environmental uncertainty, potentially causing strategies to fail in specific scenarios.
Overall, the MG-DRL strategy proposed in this paper demonstrates stable performance advantages across a broad control time domain (0.1 to 0.7 s), exhibiting excellent robustness. A longer control time domain typically indicates greater potential for improving speed control strategies. Until vehicle controllers and related components achieve sufficient responsiveness for precise control, strategies operating within larger time domains remain critically important. Additionally, refining the segmentation of the control time domain may contribute to further enhancing the strategy's performance.
This paper proposes a MG-DRL framework to optimize eco-driving for pure electric vehicles. An IDM-based speed guidance algorithm is designed as a baseline. The reference speed generated by the IDM is embedded into the state space, reward function, and exploration process, forming a triple-guidance mechanism that accelerates convergence, ensures safety, and allows the agent to learn optimized policies beyond the reference trajectory.
The framework contributes three innovations: a unified IDM-based speed guidance model for both car-following and signalized intersection scenarios; a triple-guidance mechanism (state embedding, reward constraint, and exploration guidance); and a safety-bounded architecture that retains the IDM as a fallback. Experiments in a realistic urban network show that the proposed strategy achieves an average energy saving of 7.96% with only a 0.44% increase in travel time. Microscopic analysis reveals smooth speed profiles, pause-free intersection passage, and robust performance across a wide range of control intervals (0.1–0.7 s).
This paper has several limitations. The framework assumes perfect real-time perception of the leading vehicle and traffic signal states, and does not account for sensor noise or communication delays. Under heterogeneous traffic conditions or suboptimal signal conditions, IDM guidance may provide suboptimal reference speeds. The validation was conducted in a simulation environment; real-world deployment may present additional challenges. Future research will focus on three areas: enhancing robustness through perception noise modeling and domain randomization; extending the framework to energy-efficient driving scenarios involving multi-agent collaboration; and conducting real-world vehicle experiments to validate real-time performance.
Mei Yan: Writing–original draft; visualization; methodology; funding acquisition; formal analysis; data curation. Shengjie Chen: Writing–original draft; visualization; validation; methodology; investigation; data curation. Hongwen He: Software; resources; investigation; data curation. Yunlong Wang: Validation; software; resources; formal analysis. Menglin Li: Writing–original draft; supervision; resources; project administration; funding acquisition; conceptualization.
This work is supported by the National Natural Science Foundation of China (Grant No. 52402483), Hebei Natural Science Foundation (Grant No. F2024203111), Hebei Province Foundation for Returnees (Grant No. A2024002), the Science Research Project of Hebei Education Department (Grant No. BJ2026369), and Yanshan University Fundamental Creative Research Project (Grant No. 2023LGZD004).
The data that support the findings of this study are available from the corresponding author upon reasonable request.
J. J. Cai, Y. G. Liu, Y. J. Zhang, S. Q. Shen, Z. Z. Lei, and Z. Chen, "Hierarchical Cooperative Eco-Driving Optimization for Multidimensional Mixed Vehicles at Signalized Intersections," Energy 325 (2025): 136174, https://doi.org/10.1016/j.energy.2025.136174.
Y. J. Pan, Y. Xi, W. P. Fang, Y. S. Liu, Y. L. Zhang, and W. S. Zhang, "An Eco-driving Strategy for Electric Buses at Signalized Intersection with a Bus Stop Based on Energy Consumption Prediction," Energy 317 (2025): 134672, https://doi.org/10.1016/j.energy.2025.134672.
C. X. Liu, Y. Chen, R. Z. Xu, H. J. Ruan, C. Wang, and X. Y. Li, "Co-Optimization of Energy Management and Eco-Driving Considering Fuel Cell Degradation via Improved Hierarchical Model Predictive Control," Green Energy and Intelligent Transportation 3, no. 6 (2024): 100176, https://doi.org/10.1016/j.geits.2024.100176.
Y. H. Hu, Y. B. Wang, J. Q. Guo, et al., "Eco-Driving of Connected Autonomous Vehicles in Urban Traffic Networks of Mixed Autonomy with Cut-in and Escape Lane-Changes of Manually-Driven Vehicles," Transportation research Part C: Emerging technologies 169 (2024): 104889, https://doi.org/10.1016/j.trc.2024.104889.
J. Li, X. D. Wu, M. Xu, and Y. G. Liu, "Deep Reinforcement Learning and Reward Shaping Based Eco-Driving Control for Automated HEVs among Signalized Intersections," Energy 251 (2022):123924, https://doi.org/10.1016/j.energy.2022.123924.
Z. X. Zhu, S. Gupta, A. Gupta, and M. Canova, "A Deep Reinforcement Learning Framework for Eco-Driving in Connected and Automated Hybrid Electric Vehicles," IEEE Transactions on Vehicular Technology 73, no. 2 (2024): 1713–1725, https://doi.org/10.1109/TVT.2023.3318552.
S. Y. Bao, P. Sun, J. X. Zhu, Q. Ji, and J. H. Liu, "Improved Multi-Dimensional Dynamic Programming Energy Management Strategy for a Vehicle Power-Split Hybrid Powertrain," Energy 256 (2022): 124682, https://doi.org/10.1016/j.energy.2022.124682.
A. Irshayyid, J. Chen, and G. J. Xiong, "A Review on Reinforcement Learning-Based Highway Autonomous Vehicle Control," Green Energy and Intelligent Transportation 3, no. 4 (2024): 100156, https://doi.org/10.1016/j.geits.2024.100156.
Y. T. Wu, Y. G. Liu, J. Peng, Z. Chen, Y. Liu, and Y. Zhang, "Enhanced Hierarchical Eco-Driving Control for Electric Vehicles via Global Velocity Optimization With Road Slope Adaptation," IEEE Transactions on Transportation Electrification 11, no. 4 (2025): 10322–10335, https://doi.org/10.1109/TTE.2025.3562715.
W. W. Xin, E. Y. Xu, W. G. Zheng, H. B. Feng, and J. R. Qin, "Optimal Energy Management of Fuel Cell Hybrid Electric Vehicle Based on Model Predictive Control and Online Mass Estimation," Energy Reports 8 (2022): 4964–4974, https://doi.org/10.1016/j.egyr.2022.03.194.
Z. W. Yang, Z. D. Zheng, J. W. Kim, and H. Rakha, "Eco-Driving Strategies Using Reinforcement Learning for Mixed Traffic in the Vicinity of Signalized Intersections," Transportation Research Part C: Emerging Technologies 165 (2024): 104683, https://doi.org/10.1016/j.trc.2024.104683.
S. H. Chen, Y. Huang, J. Zhang, X. S. Yu, Y. F. Lu, and D. J. Xuan, "Research on a Novel Multi-Agent Deep Reinforcement Learning Eco-Driving Framework," Energy 326 (2025): 136308, https://doi.org/10.1016/j.energy.2025.136308.
D. W. Li, F. Zhu, J. P. Wu, Y. D. Wong, and T. Y. Chen, "Managing Mixed Traffic at Signalized Intersections: An Adaptive Signal Control and CAV Coordination System Based on Deep Reinforcement Learning," Expert Systems with Applications 238, Part C (2024): 121959, https://doi.org/10.1016/j.eswa.2023.121959.
M. Y. Chen, B. B. Li, Y. G. Bian, et al., "Intersection Signal-Vehicle Coupled Coordination with Mixed Autonomy Vehicles," IEEE Transactions on Transportation Electrification 11, no. 1 (2025): 668–681, https://doi.org/10.1109/TTE.2024.3394595.
Y. N. Ye, Z. J. Xu, C. Y. Wang, and W. Z. Zhao, "Energy-Efficient Roundabout Crossing Strategy for Autonomous Electric Vehicles Based on Reinforcement Learning with Expert Demonstration," IEEE Transactions on Transportation Electrification 11, no. 6 (2025): 13802–13814, https://doi.org/10.1109/TTE.2025.3602960.
H. Tong, L. Chu, Z. Chen, Y. G. Liu, Y. J. Zhang, and J. C. Hu, "Multi-Objective Autonomous Eco-Driving Strategy: A Pathway to Future Green Mobility," Green Energy and Intelligent Transportation 4, no. 4 (2025): 100279, https://doi.org/10.1016/j.geits.2025.100279.
H. W. He, X. F. Meng, Y. Wang, et al., "Deep Reinforcement Learning Based Energy Management Strategies for Electrified Vehicles: Recent Advances and Perspectives," Renewable and Sustainable Energy Reviews 192 (2024): 114248, https://doi.org/10.1016/j.rser.2023.114248.
X. Y. Shi, J. Zhang, X. Jiang, J. Chen, W. Hao, and B. Wang, "Learning Eco-Driving Strategies from Human Driving Trajectories," Physica A: Statistical Mechanics and its Applications 633 (2024): 129353, https://doi.org/10.1016/j.physa.2023.129353.
X. L. Zhang, X. Jiang, N. Li, Z. Y. Yang, Z. Xiong, and J. Zhang, "Eco-Driving for Intelligent Electric Vehicles at Signalized Intersection: A Proximal Policy Optimization Approach," in Proceedings of ISCTT 2021-6th International Conference on Information Science, Computer Technology and Transportation, 1–7, VDE VERLAG, 2021.
W. Q. Chen, J. K. Peng, T. H. Ren, H. L. Zhang, H. W. He, and C. Y. Ma, "Integrated Velocity Optimization and Energy Management for FCHEV: An Eco-Driving Approach Based on Deep Reinforcement Learning," Energy Conversion and Management 296 (2023): 117685, https://doi.org/10.1016/j.enconman.2023.117685.
Y. Huang, H. Q. Hu, J. Q. Tan, C. L. Lu, and D. J. Xuan, "Deep Reinforcement Learning Based Energy Management Strategy for Range Extend Fuel Cell Hybrid Electric Vehicle," Energy Conversion and Management 277 (2023): 116678, https://doi.org/10.1016/j.enconman.2023.116678.
X. D. Wu, J. Li, C. R. Su, J. W. Fan, and M. Xu, "A Deep Reinforcement Learning Based Hierarchical Eco-Driving Strategy for Connected and Automated HEVs," IEEE Transactions on Vehicular Technology 72, no. 11 (2023): 13901–13916, https://doi.org/10.1109/TVT.2023.3283617.
H. Lee, K. Kim, N. Kim, and S. W. Cha, "Energy Efficient Speed Planning of Electric Vehicles for Car-Following Scenario Using Model-Based Reinforcement Learning," Applied Energy 313 (2022): 118460, https://doi.org/10.1016/j.apenergy.2021.118460.
Q. C. Su, R. C. Huang, and H. W. He, "Heterogeneous Multi-Agent Deep Reinforcement Learning for Eco-Driving of Hybrid Electric Tracked Vehicles: A Heuristic Training Framework," Journal of Power Sources 601 (2024): 234292, https://doi.org/10.1016/j.jpowsour.2024.234292.
J. Li, X. D. Wu, J. W. Fan, Y. G. Liu, and M. Xu, "Overcoming Driving Challenges in Complex Urban Traffic: A Multi-Objective Eco-Driving Strategy via Safety Model Based Reinforcement Learning," Energy 284 (2023): 128517, https://doi.org/10.1016/j.energy.2023.128517.
D. Coraci, S. Brandi, T. Z. Hong, and A. Capozzoli, "Online Transfer Learning Strategy for Enhancing the Scalability and Deployment of Deep Reinforcement Learning Control in Smart Buildings," Applied Energy 333 (2023): 120598, https://doi.org/10.1016/j.apenergy.2022.120598.
C. X. Ling, J. K. Peng, Y. Fan, Z. X. Wang, S. C. Yu, and C. C. Wu, "Safety-Awareness Enhanced Eco-Driving Strategy for Dual-Motor Electric Vehicle in Highway Scenarios Based on Improved Proximal Policy Optimization Algorithm," Energy 340 (2025): 139177, https://doi.org/10.1016/j.energy.2025.139177.
K. H. Mo, P. G. Ye, X. J. Ren, S. W. Wang, W. J. Li, and J. Li, "Security and Privacy Issues in Deep Reinforcement Learning: Threats and Countermeasures," ACM Computing Surveys 56, no. 6 (2024): 1–39, https://doi.org/10.1145/3640312.
C. T. Zhang, W. H. Huang, X. Y. Zhou, C. Lv, and C. Sun, "Expert-Demonstration-Augmented Reinforcement Learning for Lane-Change-Aware Eco-Driving Traversing Consecutive Traffic Lights," Energy 286 (2024): 129472, https://doi.org/10.1016/j.energy.2023.129472.
J. Q. Liu, C. Y. Wang, W. Z. Zhao, "An Eco-driving Strategy for Autonomous Electric Vehicles Crossing Continuous Speed-Limit Signalized Intersections," Energy 294 (2024): 130829, https://doi.org/10.1016/j.energy.2024.130829.
M. L. Li, L. Yin, M. Yan, J. D. Wu, H. W. He, and C. C. Jia, "Hierarchical Intelligent Energy-Saving Control Strategy for Fuel Cell Hybrid Electric Buses Based on Traffic Flow Predictions," Energy 304 (2024): 132144, https://doi.org/10.1016/j.energy.2024.132144.