Feiye Zhang

dblp:73/1183 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-7015-9768ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 An Attack-Defense Game-Based Reinforcement Learning Privacy-Preserving Method Against Inference Attack in Double Auction Market
abstract
Auction mechanism, as a fair and efficient resource allocation method, has been widely used in varieties trading scenarios, such as advertising, crowdsensoring and spectrum. However, in addition to obtaining higher profits and satisfaction, the privacy concerns have attracted researchers’ attention. In this paper, we mainly study the privacy preserving issue in the double auction market against the indirect inference attack. Most of the existing works apply differential privacy theory to defend against the inference attack, but there exists two problems. First, ‘indistinguishability’ of differential privacy (DP) cannot prevent the disclosure of continuous valuations in the auction market. Second, the privacy-utility trade-off (PUT) in differential privacy deployment has not been resolved. To this end, we proposed an attack-defense game-based reinforcement learning privacy-preserving method to provide practically privacy protection in double auction. First, the auctioneer acts as defender, adds noise to the bidders’ valuations, and then acts as adversary to launch inference attack. After that the auctioneer uses the attack results and auction results as a reference to guide the next deployment. The above process can be regarded as a Markov Decision Process (MDP). The state is the valuations of each bidders under the current steps. The action is the noise added to each bidders. The reward is composed of privacy, utility and training speed, in which attack success rate and social welfare are taken as measures of privacy and utility, a delay penalty term is used to reduce the training time. Utilizing the deep deterministic policy gradient (DDPG) algorithm, we establish an actor-critic network to solve the problem of MDP. Finally, we conducted extensive evaluations to verify the performance of our proposed method. The results show that compared with other existing DP-based double auction privacy preserving mechanisms, our method can achieve better results in both privacy and utility. We can reduce the attack success rate from nearly 100% to less than 20%, and the utility deviation is less than 5%. Note to Practitioners—Privacy protection in trading markets, such as advertising, crowdsensing, and spectrum, is crucial. Traditional approaches like differential privacy have been unable to entirely guard sensitive data against inference attacks. To address this, we introduce a novel privacy-preserving mechanism for double auction markets. Our approach employs an attack-defense game model, where noise is added to bidders’ valuations and then used to launch an inference attack. This process allows for the evaluation of the noise’s effectiveness and iteratively refines the privacy protection method. Transformed into a reinforcement learning model and optimized through a DDPG network, our mechanism reduces computational complexity. It has been shown to significantly diminish the success rate of inference attacks, while maintaining a minimal utility deviation. Practitioners in auction-based markets can leverage our approach to enhance privacy protection without negatively impacting market performance. By integrating our mechanism into their operations, auctioneers can foster a safer and more efficient trading environment.
Donghe Li, Chunlin Hu, Qingyu Yang 0003, Feiye Zhang, Dou An
IEEE Trans Autom. Sci. Eng.5
2025 Charge or Pick Up? Optimizing E-Taxi Management: A Dual-Stage Heuristic Coordinated Reinforcement Learning Approach
abstract
In recent years, the rapid adoption of electric vehicles (EVs) in the taxi industry has transformed traditional taxi-hailing systems into electric taxi (E-taxi) hailing systems. As a result, it is crucial to develop effective strategies for optimizing E-taxi management by considering both passenger-taxi matching and charging planning. In this paper, we first formalize the E-taxi management optimization problem as a Markov decision problem with dynamic state and heterogeneous action. We then propose a dual-stage heuristic coordinated reinforcement learning (RL) approach that incorporates advanced feature selection and heuristic allocation strategies. Our approach consists of two main stages. In the first stage, we introduce the feature-guided state dimensionality stabilization proximal policy optimization (PPO) method to address dynamic state dimensions by a feature selection method, and enabling E-taxis to decide whether to charge or pick up passengers. In the second stage, we propose a heuristic coordinated assignment method to further allocate charging stations and passengers for the E-taxis, and provide the RL network in the first stage with rewards based on the results. This approach effectively tackles the challenge of RL processing of heterogeneous action spaces (charge and pick up). We evaluate our proposed method in a real-world E-taxi environment and find that it significantly enhances the experience for both E-taxis and passengers. Specifically, due to our method’s rational planning for passenger pick-up and charging, E-taxis can increase their revenue by 20% compared to traditional RL methods or random scheduling approaches. As for passengers, since the taxis have more efficiently planned their charging behavior, the probability of their orders being answered increases by 15%, while their waiting time is reduced by 55%. These achievements contribute to the advancement of E-taxi management strategies and promote the widespread adoption of electric vehicles, ultimately supporting the transition to a more sustainable transportation system. Note to Practitioners—The increasing adoption of electric vehicles in the taxi industry has led to the need for effective E-taxi management strategies that consider both passenger-taxi matching and charging planning. In this study, we introduce a dual-stage heuristic coordinated reinforcement learning approach that addresses these challenges by integrating a feature-guided state dimensionality stabilization proximal policy optimization method and a heuristic coordinated assignment method. Our approach offers several practical benefits for E-taxi service providers, drivers, and passengers. For E-taxi service providers, the proposed method improves E-taxi dispatch efficiency, resulting in a more effective use of available resources and potentially increasing overall revenue. For E-taxi drivers, our approach leads to better planning of charging and passenger pick-up decisions, increasing their earnings by 20% compared to traditional methods, and reducing the average occurrence of low battery status from more than 4 times every 10 hours to less than 1 time. Passengers, on the other hand, experience improved service quality due to the more efficient E-taxi management. The probability of their orders being answered increases by 15%, and their waiting time is reduced by 100%. These improvements contribute to an enhanced user experience and may encourage further adoption of E-taxis as a sustainable transportation solution. The proposed method can be integrated into existing E-taxi hailing platforms, such as DiDi and Uber, to enhance their dispatch and charging management capabilities. As the global trend towards sustainable transportation continues to grow, our approach provides valuable insights and a practical solution for the efficient management of E-taxi fleets in modern urban environments.
Donghe Li, Chunlin Hu, Qingyu Yang 0003, Pengtao Song, Feiye Zhang, Dou An
IEEE Trans Autom. Sci. Eng.5
2024 Toward Data Integrity Attacks Against Distributed Dynamic State Estimation in Smart Grid
abstract
With the continuous expansion of the power grid nodes scale, traditional centralized state estimation method shows certain limitations in estimation efficiency and accuracy. Recently, some power grids adopt a distributed state estimation method, in which each partition independently estimates the partial state information by partitioning the entire power system. However, the deviation of the state estimation in certain partition will result in the deviation of the estimation results in the entire power grid system. In this paper, we propose the attack strategy against the distributed state estimation in smart grids from two perspectives, i.e. the attack against local physical measurement value of the power system partition and the attack against the measurement value of coordination center. Moreover, the theoretical analysis of the state estimation deviation caused by the proposed data integrity attack and the propagation processes of proposed attack vectors in measurement calculations are formalized. The effectiveness of the proposed attack strategy is verified in the IEEE-30 bus and IEEE-118 bus systems. Simulation results show that attacking against a certain partition of a distributed system can indirectly affect the state estimation results of other partitions and attacking against the measurement value of coordination center can directly threaten the state estimation results of the entire power grid. Note to Practitioners— This paper proposes two attack strategies against the distributed state estimation of power grid from two perspectives, i.e. the attack against local physical measurement value of the power system partition and the attack against the measurement value of coordination center. Most of the previous works fail to formalize the state estimation deviation of both partial and entire state estimation results of power grid after the attacker launches the attack against the distributed state estimation. We formalize the state estimation deviation caused by the proposed data integrity attack and the propagation processes of proposed attack vectors in measurement calculations. The effectiveness of the proposed attack strategy against the state estimation of power grid is verified in the IEEE-30 bus and IEEE-118 bus systems. Simulation results show that attacking against a certain partition of power grid can indirectly cause the deviation in the state estimation results of entire power grid and attacking against the measurement value of the coordination center can directly threaten the state estimation results of the entire power grid. In conclusion, the proposed attack strategies are helpful for the research community to design detection strategies in a targeted manner and can be conveniently applied to the real-world security management system of smart grid.
Dou An, Feiye Zhang, Feifei Cui, Qingyu Yang 0003
IEEE Trans Autom. Sci. Eng.2
2024 Research on Privacy Issues in Smart Metering System: An Improved TCN-Based NILM Attack Method and Practical DRL-Based Rechargeable Battery Assisted Privacy Preserving Method
abstract
Smart meters, as a key component of Advanced Metering Infrastructure (AMI), collect fine-grained electricity consumption data for demand response in smart grids. While this data improves the grid’s accuracy, it also poses significant threats to users’ privacy. In this paper, we study the privacy issue in smart metering systems from both attacker and defender perspectives. First, we propose an improved Temporal Convolutional Network (TCN) based Non-Intrusive Load Monitoring (NILM) attack method, which infers electrical appliance usage from public load curves, addressing the gradient vanishing, gradient exploding, and other problems while improving attack accuracy. Second, we develop a rechargeable battery-assisted energy management system to hide load characteristics of electrical appliances by adding physical noise, thus resisting NILM attacks. To address the privacy-cost trade-off optimization problem, we propose a Practical Deep Reinforcement Learning-based Rechargeable Battery assist Privacy Preserving Method (PRoP) that learns optimal battery charging/discharging policies. We design a novel privacy measurement method and constraints to ensure the feasibility of system deployment and prove PRoP’s effectiveness in resisting NILM attacks. Comprehensive evaluations demonstrate that our improved TCN-based NILM method achieves an attack success rate of over 80% on various electrical appliances, improving attack performance (MAE, RMSE) by 20% compared to existing methods while reducing model training time. Moreover, our proposed PRoP achieves a better trade-off between privacy protection and electricity cost than existing battery-assisted methods, reducing costs by 5% and the attack success ratio to 36%, while increasing the MAE and RMAE obtained by NILM by 3 times.Note to Practitioners—Smart grid, which can support bidirectional information transmission, has a series of advantages, such as high efficiency and high stability. However, it also brings a significant threat to users’ electricity privacy. Although encryption-based privacy protection methods have been deployed on terminal devices of smart grids to prevent privacy leaks, this method can often only defend against intrusive attacks and has little effect on non-intrusive attacks. To this end, this paper studies the privacy issues caused by non-intrusive attacks. Specifically, to better study the protection method, we first investigate the attack mechanism and design an improved TCN-based Non-Intrusive Load Monitoring method. Then, we propose a Practical Reinforcement learning-based rechargeable battery-assisted Privacy preserving method (PRoP) to defend against this attack physically. The most practical contribution of this paper is that, compared with existing battery-assisted privacy protection methods, we do not blindly pursue algorithm performance but fully consider the practical factors of deployment, such as limiting battery capacity and constraining battery charging and discharging behavior. This can guide practitioners to better apply this technology in practice.
Donghe Li, Qingyu Yang 0003, Feiye Zhang, Yingchen Qian, Dou An
IEEE Trans Autom. Sci. Eng.3
2022 A leader-following paradigm based deep reinforcement learning method for multi-agent cooperation games
Feiye Zhang, Qingyu Yang 0003, Dou An
Neural Networks1
2022 Data Integrity Attack in Dynamic State Estimation of Smart Grid: Attack Model and Countermeasures
abstract
A smart grid integrates advanced sensors, efficient measurement methods, progressive control technologies, and other techniques and devices to achieve safe, efficient and economical operation of the grid system. However, the diversified and open environment of a smart grid makes energy and information of the smart grid vulnerable to malicious attacks. As a representative cyber-physical attack, the data integrity attack has an extremely severe impact on the grid operation for it can bypass the traditional detection mechanisms by adjusting the attack vector. In this paper, we first present the attack strategy against dynamic state estimation of power grid in the perspective of adversary and formulate the data integrity attack detection problem that has the characteristic of sequential decision making as a partially observable Markov decision process. Then, a deep reinforcement learning-based approach is proposed to detect against data integrity attacks, which utilizes the Long Short-Term Memory layer to extract the state features of previous time steps in determining whether the system is currently under attack. Moreover, the noisy networks are employed to ensure effective agent exploration, which prevents the agent from sticking to the non-optimal policy. The principle of a multi-step learning is adopted to increase the estimation accuracy of Q value. To address the sparse rewards problem, the prioritized experience replay is proposed to increase training efficiency. Simulation results demonstrated that the proposed detection approach surpasses the benchmarks in the comparison metrics: delay error rate and false rate.Note to Practitioners—In this paper, we present a deep reinforcement learning-based algorithm to defend against the data integrity attacks of smart grid. Most of the previous works discretized the system states and utilized the current state information to identify whether the system is under attack. For this reason, the detection policy may totally ignored the continuously changing characteristics of the grid states, which will lead to poor detection performance. Moreover, the attacked system states only accounts for a small part of the entire grid operation states, the probability of sampling the experience containing the attack state is extremely small, which limits the learning efficiency of previous RL-based detection approaches. In order to increase the accuracy of detection, we first present the attack strategy against power grid’s dynamic state estimation in the perspective of adversary and formulate the partially observable Markov decision process model of attack detection problem. Moreover, we propose a deep reinforcement learning-based detection approach combining the LSTM network to extract the system state features of the previous time steps to determine whether the system is currently being attacked. To address the sparse rewards problem, the prioritized experience replay is used to increase learning efficiency. The experiments demonstrate the effectiveness of proposed detection scheme compared with benchmarks in terms of detection delay as well as accuracy. In conclusion, the proposed detection scheme is helpful in defending against the data integrity attacks without obtaining the opponent’s strategy in advance and can be conveniently applied to the real-world security management system of smart grid.
Dou An, Feiye Zhang, Qingyu Yang 0003, Chengwei Zhang 0001
IEEE Trans Autom. Sci. Eng.2
2021 CDDPG: A Deep-Reinforcement-Learning-Based Approach for Electric Vehicle Charging Control
abstract
Electric vehicle (EV) has become one of the most critical components in the smart grid with the applications of the Internet-of-Things (IoT) technologies. Real-time charging control is pivotal to ensure the efficient operation of EVs. However, the charging control performance is limited by the uncertainty of the environment. On the other hand, it is challenging to determine a charging control strategy that is able to optimize multiple objectives simultaneously. In this article, we formulate the EV charging control model as a Markov decision process (MDP) by constructing state, action, transition function, and reward. Then, we propose a deep-reinforcement-learning-based approach: charging control deep deterministic policy gradient (CDDPG) to learn the optimal charging control strategy for satisfying the user's requirement of battery energy while minimizing the user's charging expense. We utilize the long short-term memory (LSTM) network that extracts the information of previous energy price to determine the current charging control strategy. Moreover, Gaussian noise is added to the output of the actor network to prevent the agent from sticking into the nonoptimal strategy. In addition, we address the limitation of sparse rewards by using two replay buffers, of which one is used to store the rewards during the charging phase and another is used to store the rewards after charging is completed. The simulation results prove that the CDDPG-based approach outperforms the deep-$Q$ -learning-based approach (DQL) and the deep-deterministic-policy-gradient-based approach (DDPG) in satisfying the user's requirement for the battery energy and reducing the charging cost.
Feiye Zhang, Qingyu Yang 0003, Dou An
IEEE Internet Things J.1
1999 Visual speech analysis and synthesis with application to Mandarin speech training
abstract
This paper presents a novel vision-based speech analysis system STODE which is used in spoken Chinese training of oral deaf children. Its design goal is to help oral deaf children overcome two major difficulties in speech learning: the confusion of intonations for spoken Chinese characters and timing errors within different words and characters. It integrates such capabilities as real-time lip tracking and feature extraction, multi-state lip modeling, Time-delay Neural Network (TDNN) for visual speech analysis. A desk-mounted camera tracks users in real-time. At each frame, region of interest is identified and key information is extracted. The preprocessed acoustic and visual information are then fed into a modular TDNN and combined for visual speech analysis. Confusion of intonations for spoken Chinese characters can be easily identified, and timing error within words and characters also can be detected using a DTW (Dynamic Time Warping) algorithm. For visual feedback we have created an artificial talking head directly cloned from user's own images to generate correct outputs showing both correct and wrong ways of pronunciation. This system has been successfully used for spoken Chinese training of oral deaf children in cooperation with Nanjing Oral School under grants from National Natural Science Foundation of China.
Xiaodong Jiang, Yunlai Wang, Feiye Zhang
VRST3