Fei Miao

dblp:143/6002 · DBLP profile ↗
← Back
33ranked-venue papers
2as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 17 since 2021Systems, architecture and hardware · 15 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Computer networks · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust Detection of Cyberattacks on Distribution System Volt-VAR Control via Adversarial and Uncertain Analysis
Alaa Selim, Junbo Zhao 0001, Fei Miao, Sung-Yeul Park, Shan Zuo, Georgios Fragkos, Meng Yue 0001
IEEE Internet Things J.3
2025 CUQDS: Conformal Uncertainty Quantification Under Distribution Shift for Trajectory Prediction
abstract
Trajectory prediction models that can infer both future trajectories and their associated uncertainties of the target vehicles is crucial for safe and robust navigation and path planning of autonomous vehicles. However, the majority of existing trajectory prediction models have neither considered reducing the uncertainty as one objective during the training stage nor provided reliable uncertainty quantification during inference stage, especially under potential distribution shift. Therefore, in this paper, we propose the Conformal Uncertainty Quantification under Distribution Shift framework, CUQDS, to quantify the uncertainty of the predicted trajectories of existing trajectory prediction models under potential data distribution shift, while improving the prediction accuracy of the models and reducing the estimated uncertainty during the training stage. Specifically, CUQDS includes 1) a learning-based Gaussian process regression module that models the output distribution of the base model (any existing trajectory prediction neural networks) and reduces the estimated uncertainty by an additional loss term, and 2) a statistical-based Conformal P control module to calibrate the estimated uncertainty from the Gaussian process regression module in an online setting under potential distribution shift between training and testing data. Experimental results on various state-of-the-art methods using benchmark motion forecasting datasets demonstrate the effectiveness of our proposed design.
Huiqun Huang, Sihong He, Fei Miao
AAAI3
2025 Safety Guaranteed Robust Multi-Agent Reinforcement Learning with Hierarchical Control for Connected and Automated Vehicles
abstract
We address the problem of coordination and control of Connected and Automated Vehicles (CAVs) in the presence of imperfect observations in mixed traffic environment. A commonly used approach is learning-based decision-making, such as reinforcement learning (RL). However, most existing safe RL methods suffer from two limitations: (i) they assume accurate state information, and (ii) safety is generally defined over the expectation of the trajectories. It remains challenging to design optimal coordination between multi-agents while ensuring hard safety constraints under system state uncertainties (e.g., those that arise from noisy sensor measurements, communication, or state estimation methods) at every time step. We propose a safety guaranteed hierarchical coordination and control scheme called Safe-RMM to address the challenge. Specifically, the high-level coordination policy of CAVs in mixed traffic environment is trained by the Robust Multi-Agent Proximal Policy Optimization (RMAPPO) method. Though trained without uncertainty, our method leverages a worst-case Q network to ensure the model's robust performances when state uncertainties are present during testing. The low-level controller is implemented using model predictive control (MPC) with robust Control Barrier Functions (CBFs) to guarantee safety through their forward invariance property. We compare our method with baselines in different road networks in the CARLA simulator. Results show that our method provides the best evaluated safety and efficiency in challenging mixed traffic environments with uncertainties.
H. M. Sabbir Ahmad, Ehsan Sabouni, Yanchao Sun, Furong Huang, Wenchao Li 0001, Fei Miao
ICRA7
2025 Multi-Agent Reinforcement Learning Guided by Signal Temporal Logic Specifications
abstract
Reward design is a key component of deep reinforcement learning (DRL), yet some tasks and designer’s objectives may be unnatural to define as a scalar cost function. Among the various techniques, formal methods integrated with DRL have garnered considerable attention due to their expressiveness and flexibility in defining the reward and requirements for different states and actions of the agent. Nevertheless, the exploration of leveraging Signal Temporal Logic (STL) for guiding multi-agent reinforcement learning (MARL) reward design is still limited. The presence of complex interactions, heterogeneous goals, and critical safety requirements in multi-agent systems exacerbates this challenge. In this paper, we propose a novel STL-guided multi-agent reinforcement learning framework. The STL requirements are designed to include both task specifications according to the objective of each agent and safety specifications. The robustness values from checking the states against STL specifications are leveraged to generate rewards. We validate our approach by conducting experiments across various testbeds. The experimental results demonstrate significant performance improvements compared to MARL without STL guidance, along with a remarkable increase in the overall safety rate of the multi-agent systems.
Jiangwei Wang, Shuo Yang 0007, Ziyan An, Songyang Han, Rahul Mangharam, Meiyi Ma, Fei Miao
IROS8
2025 YOLO-MARL: You Only LLM Once for Multi-Agent Reinforcement Learning
abstract
Advancements in deep multi-agent reinforcement learning (MARL) have positioned it as a promising approach for decision-making in cooperative games. However, it still remains challenging for MARL agents to learn cooperative strategies for some game environments. Recently, large language models (LLMs) have demonstrated emergent reasoning capabilities, making them promising candidates for enhancing coordination among the agents. However, due to the model size of LLMs, it can be expensive to frequently infer LLMs for actions that agents can take. In this work, we propose You Only LLM Once for MARL (YOLO-MARL), a novel framework that leverages the high-level task planning capabilities of LLMs to improve the policy learning process of multi-agents in cooperative games. Notably, for each game environment, YOLO-MARL only requires one time interaction with LLMs in the proposed strategy generation, state interpretation and planning function generation modules, before the MARL policy training process. This avoids the ongoing costs and computational time associated with frequent LLMs API calls during training. Moreover, trained decentralized policies based on normal-sized neural networks operate independently of the LLM. We evaluate our method across two different environments and demonstrate that YOLO-MARL outperforms traditional MARL algorithms. The Github repository of our code can be found at https://github.com/paulzyzy/YOLO-MARL.
Yuxiao Chen 0008, Fei Miao
IROS5
2025 UCS-YOLO: a novel approach for multi-scale small target detection in UAVs
Guang-Ling Sun, Fen-Qi Zhang, Yu-Min Zhu, Fei Miao
J. Supercomput.4
2024 MetaAT: Active Testing for Label-Efficient Evaluation of Dense Recognition Tasks
Sanbao Su, Thang Long Doan, Sima Behpour, Liang Gou, Fei Miao, Liu Ren 0001
ECCV (78)7
2024 Momentum for the Win: Collaborative Federated Reinforcement Learning across Heterogeneous Environments
abstract
We explore a Federated Reinforcement Learning (FRL) problem where $N$ agents collaboratively learn a common policy without sharing their trajectory data. To date, existing FRL work has primarily focused on agents operating in the same or “similar" environments. In contrast, our problem setup allows for arbitrarily large levels of environment heterogeneity. To obtain the optimal policy which maximizes the average performance across all potentially completely different environments, we propose two algorithms: FedSVRPG-M and FedHAPG-M. In contrast to existing results, we demonstrate that both FedSVRPG-M and FedHAPG-M, both of which leverage momentum mechanisms, can exactly converge to a stationary point of the average performance function, regardless of the magnitude of environment heterogeneity. Furthermore, by incorporating the benefits of variance-reduction techniques or Hessian approximation, both algorithms achieve state-of-the-art convergence results, characterized by a sample complexity of $\mathcal{O}\left(\epsilon^{-\frac{3}{2}}/N\right)$. Notably, our algorithms enjoy linear convergence speedups with respect to the number of agents, highlighting the benefit of collaboration among agents in finding a common policy.
Han Wang 0016, Sihong He, Fei Miao, James Anderson 0001
ICML4
2024 Constrained Reinforcement Learning Under Model Mismatch
abstract
Existing studies on constrained reinforcement learning (RL) may obtain a well-performing policy in the training environment. However, when deployed in a real environment, it may easily violate constraints that were originally satisfied during training because there might be model mismatch between the training and real environments. To address this challenge, we formulate the problem as constrained RL under model uncertainty, where the goal is to learn a policy that optimizes the reward and at the same time satisfies the constraint under model mismatch. We develop a Robust Constrained Policy Optimization (RCPO) algorithm, which is the first algorithm that applies to large/continuous state space and has theoretical guarantees on worst-case reward improvement and constraint violation at each iteration during the training. We show the effectiveness of our algorithm on a set of RL tasks with constraints.
Zhongchang Sun, Sihong He, Fei Miao, Shaofeng Zou
ICML3
2024 Policy Optimization for Robust Average Reward MDPs
abstract
This paper studies first-order policy optimization for robust average cost Markov decision processes (MDPs). Specifically, we focus on ergodic Markov chains. For robust average cost MDPs, the goal is to optimize the worst-case average cost over an uncertainty set of transition kernels. We first develop a sub-gradient of the robust average cost. Based on the sub-gradient, a robust policy mirror descent approach is further proposed. To characterize its iteration complexity, we develop a lower bound on the difference of robust average cost between two policies and further show that the robust average cost satisfies the PL-condition. We then show that with increasing step size, our robust policy mirror descent achieves a linear convergence rate in the optimality gap, and with constant step size, our algorithm converges to an $\epsilon$-optimal policy with an iteration complexity of $\mathcal{O}(1/\epsilon)$. The convergence rate of our algorithm matches with the best convergence rate of policy-based algorithms for robust MDPs. Moreover, our algorithm is the first algorithm that converges to the global optimum with general uncertainty sets for robust average cost MDPs. We provide simulation results to demonstrate the performance of our algorithm.
Zhongchang Sun, Sihong He, Fei Miao, Shaofeng Zou
NeurIPS3
2024 Towards Safe Autonomy in Hybrid Traffic: Detecting Unpredictable Abnormal Behaviors of Human Drivers via Information Sharing
abstract
Hybrid traffic which involves both autonomous and human-driven vehicles would be the norm of the autonomous vehicles’ practice for a while. On the one hand, unlike autonomous vehicles, human-driven vehicles could exhibit sudden abnormal behaviors such as unpredictably switching to dangerous driving modes—putting its neighboring vehicles under risks; such undesired mode switching could arise from numbers of human driver factors, including fatigue, drunkenness, distraction, aggressiveness, and so on. On the other hand, modern vehicle-to-vehicle (V2V) communication technologies enable the autonomous vehicles to efficiently and reliably share the scarce run-time information with each other [ 1 ]. In this article, we propose, to the best of our knowledge, the first efficient algorithm that can (1) significantly improve trajectory prediction by effectively fusing the run-time information shared by surrounding autonomous vehicles, and can (2) accurately and quickly detect abnormal human driving mode switches or abnormal driving behavior with formal assurance without hurting human drivers’ privacy. To validate our proposed algorithm, we first evaluate our proposed trajectory predictor on NGSIM and Argoverse datasets and show that our proposed predictor outperforms the baseline methods. Then through extensive experiments on SUMO simulator, we show that our proposed algorithm has great detection performance in both highway and urban traffic. The best performance achieves detection rate of 97.3%, average detection delay of 1.2 s, and 0 false alarm.
Jiangwei Wang, Lili Su, Songyang Han, Dongjin Song, Fei Miao
ACM Trans. Cyber Phys. Syst.5
2024 A Multi-Agent Reinforcement Learning Approach for Safe and Efficient Behavior Planning of Connected Autonomous Vehicles
abstract
The recent advancements in wireless technology enable connected autonomous vehicles (CAVs) to gather information about their environment by vehicle-to-vehicle (V2V) communication. In this work, we design an information-sharing-based multi-agent reinforcement learning (MARL) framework for CAVs, to take advantage of the extra information when making decisions to improve traffic efficiency and safety. The safe actor-critic algorithm we propose has two new techniques: the truncated$\mathcal{Q}$-function and safe action mapping. The truncated$\mathcal{Q}$-function utilizes the shared information from neighboring CAVs such that the joint state and action spaces of the$\mathcal{Q}$-function do not grow in our algorithm for a large-scale CAV system. We prove the bound of the approximation error between the truncated-$\mathcal{Q}$and global$Q$-functions. The safe action mapping provides a provable safety guarantee for both the training and execution based on control barrier functions. Using the CARLA simulator for experiments, we show that our approach improves the CAV system’s efficiency in terms of average velocity and comfort under different CAV ratios and different traffic densities. We also show that our approach avoids the execution of unsafe actions and always maintains a safe distance from other vehicles. We construct an obstacle-at-corner scenario to show that the shared vision can help CAVs to observe obstacles earlier and take action to avoid traffic jams. The experiment video is on https://songyanghan.github.io/cavmarl/.
Songyang Han, Shanglin Zhou, Jiangwei Wang, Lynn Pepin, Caiwen Ding, Jie Fu 0002, Fei Miao
IEEE Trans. Intell. Transp. Syst.7
2024 FairMove: A Data-Driven Vehicle Displacement System for Jointly Optimizing Profit Efficiency and Fairness of Electric For-Hire Vehicles
abstract
With the worldwide mobility electrification initiative to reduce air pollution and energy security, more and more for-hire vehicles are being replaced with electric ones. A key difference between gas for-hire vehicles and electric for-hire vehicles (EFHV) is their energy replenishment mechanisms, i.e., refueling or charging, which is reflected in two aspects: (i) much longer charging processes vs. much shorter refueling processes and (ii) time-varying electricity prices vs. time-invariant gasoline prices during a day. The complicated charging issues (e.g., long charging time and dynamic charging pricing) potentially reduce the daily operation time and profits of EFHVs, and also cause overcrowded charging stations during some off-peak charging pricing periods. Motivated by a set of findings obtained from a data-driven investigation and field studies, in this paper, we design a fairness-aware vehicle displacement system calledFairMoveto jointly optimize the overall profit efficiency and profit fairness of EFHV drivers by considering both the passenger travel demand and vehicle charging demand. We first formulate the EFHV displacement problem as a Markov decision problem, and then we present a fairness-aware multi-agent actor-critic approach to tackle this problem. More importantly, we implement and evaluateFairMovewith real-world streaming data from the Chinese city Shenzhen, including GPS data and transaction data from over 20,100 EFHVs, coupled with the data of 123 charging stations, which constitute, to our knowledge, the largest EFHV network in the world. Extensive experimental results show that our fairness-awareFairMoveeffectively improves the profit efficiency and profit fairness of the EFHV fleet by 26.9% and 54.8%, respectively. It also improves the charging station utilization fairness by 38.4%.
Guang Wang 0001, Sihong He, Lin Jiang 0007, Shuai Wang 0008, Fei Miao, Fan Zhang 0019, Zheng Dong 0002, Desheng Zhang 0002
IEEE Trans. Mob. Comput.5
2023 Uncertainty Quantification of Collaborative Detection for Self-Driving
abstract
Sharing information between connected and autonomous vehicles (CAVs) fundamentally improves the performance of collaborative object detection for self-driving. However, CAVs still have uncertainties on object detection due to practical challenges, which will affect the later modules in self-driving such as planning and control. Hence, uncertainty quantification is crucial for safety-critical systems such as CAVs. Our work is the first to estimate the uncertainty of collaborative object detection. We propose a novel uncertainty quantification method, called Double- M Quantification, which tailors a moving block bootstrap (MBB) algorithm with direct modeling of the multivariant Gaussian distribution of each corner of the bounding box. Our method captures both the epistemic uncertainty and aleatoric uncertainty with one inference pass based on the offline Double- M training process. And it can be used with different collaborative object detectors. Through experiments on the comprehensive collaborative perception dataset, we show that our Double-M method achieves more than 4× improvement on uncertainty score and more than 3% accuracy improvement, compared with the state-of-the-art uncertainty quantification methods. Our code is public on https://coperception.github.io/double-m-quantification/.
Sanbao Su, Yiming Li 0003, Sihong He, Songyang Han, Chen Feng 0002, Caiwen Ding, Fei Miao
ICRA7
2023 Spatial-Temporal-Aware Safe Multi-Agent Reinforcement Learning of Connected Autonomous Vehicles in Challenging Scenarios
abstract
Communication technologies enable coordination among connected and autonomous vehicles (CAVs). However, it remains unclear how to utilize shared information to improve the safety and efficiency of the CAV system in dynamic and complicated driving scenarios. In this work, we propose a framework of constrained multi-agent reinforcement learning (MARL) with a parallel Safety Shield for CAVs in challenging driving scenarios that includes unconnected hazard vehicles. The coordination mechanisms of the proposed MARL include information sharing and cooperative policy learning, with Graph Convolutional Network (GCN)-Transformer as a spatial-temporal encoder that enhances the agent's environment awareness. The Safety Shield module with Control Barrier Functions (CBF)-based safety checking protects the agents from taking unsafe actions. We design a constrained multi-agent advantage actor-critic (CMAA2C) algorithm to train safe and cooperative policies for CAVs. With the experiment deployed in the CARLA simulator, we verify the performance of the safety checking, spatial-temporal encoder, and coordination mechanisms designed in our method by comparative experiments in several challenging scenarios with unconnected hazard vehicles. Results show that our proposed methodology significantly increases system safety and efficiency in challenging scenarios.
Songyang Han, Jiangwei Wang, Fei Miao
ICRA4
2023 A Robust and Constrained Multi-Agent Reinforcement Learning Electric Vehicle Rebalancing Method in AMoD Systems
abstract
Electric vehicles (EVs) play critical roles in autonomous mobility-on-demand (AMoD) systems, but their unique charging patterns increase the model uncertainties in AMoD systems (e.g. state transition probability). Since there usually exists a mismatch between the training and test/true environments, incorporating model uncertainty into system design is of critical importance in real-world applications. However, model uncertainties have not been considered explicitly in EV AMoD system rebalancing by existing literature yet, and the coexistence of model uncertainties and constraints that the decision should satisfy makes the problem even more challenging. In this work, we design a robust and constrained multi-agent reinforcement learning (MARL) framework with state transition kernel uncertainty for EV AMoD systems. We then propose a robust and constrained MARL algorithm (ROCOMA) with robust natural policy gradients (RNPG) that trains a robust EV rebalancing policy to balance the supply-demand ratio and the charging utilization rate across the city under model uncertainty. Experiments show that the ROCOMA can learn an effective and robust rebalancing policy. It outperforms non-robust MARL methods in the presence of model uncertainties. It increases the system fairness by 19.6% and decreases the rebalancing costs by 75.8%.
Sihong He, Yue Wang 0068, Shuo Han 0002, Shaofeng Zou, Fei Miao
IROS5
2023 Robust Electric Vehicle Balancing of Autonomous Mobility-on-Demand System: A Multi-Agent Reinforcement Learning Approach
abstract
Electric autonomous vehicles (EAVs) are getting attention in future autonomous mobility-on-demand (AMoD) systems due to their economic and societal benefits. However, EAVs' unique charging patterns (long charging time, high charging frequency, unpredictable charging behaviors, etc.) make it challenging to accurately predict the EAVs supply in E-AMoD systems. Furthermore, the mobility demand's prediction uncertainty makes it an urgent and challenging task to design an integrated vehicle balancing solution under supply and demand uncertainties. Despite the success of reinforcement learning-based E-AMoD balancing algorithms, state uncertainties under the EV supply or mobility demand remain unexplored. In this work, we design a multi-agent reinforcement learning (MARL)-based framework for EAVs balancing in E-AMoD systems, with adversarial agents to model both the EAVs supply and mobility demand uncertainties that may undermine the vehicle balancing solutions. We then propose a robust E-AMoD Balancing MARL (REBAMA) algorithm to train a robust EAVs balancing policy to balance both the supply-demand ratio and charging utilization rate across the whole city. Experiments show that our proposed robust method performs better compared with a non-robust MARL method that does not consider state uncertainties; it improves the reward, charging utilization fairness, and supply-demand fairness by 19.28%, 28.18%, and 3.97%, respectively. Compared with a robust optimization-based method, the proposed MARL algorithm can improve the reward, charging utilization fairness, and supply-demand fairness by 8.21%, 8.29%, and 9.42%, respectively.
Sihong He, Shuo Han 0002, Fei Miao
IROS3
2023 Privacy-Preserving and Uncertainty-Aware Federated Trajectory Prediction for Connected Autonomous Vehicles
abstract
Deep learning is the method of choice for trajectory prediction for autonomous vehicles. Unfortunately, its data-hungry nature implicitly requires the availability of sufficiently rich and high-quality centralized datasets, which easily leads to privacy leakage. Besides, uncertainty-awareness becomes increasingly important for safety-crucial cyber physical systems whose prediction module heavily relies on machine learning tools. In this paper, we relax the data collection requirement and enhance uncertainty-awareness by using Federated Learning on Connected Autonomous Vehicles with an uncertainty-aware global objective. We name our algorithm as FLTP. We further introduce ALFLTP which boosts FLTP via using active learning techniques in adaptatively selecting participating clients. We consider two different metrics negative log-likelihood (NLL) and aleatoric uncertainty (AU) for client selection. Experiments on Argoverse dataset show that FLTP significantly outperforms the model trained on local data. In addition, ALFLTP-AU converges faster in training regression loss and performs better in terms of Miss Rate (MR) than FLTP in most rounds, and has more stable round-wise performance than ALFLTP-NLL.
Muzi Peng, Jiangwei Wang, Dongjin Song, Fei Miao, Lili Su
IROS4
2023 Data-Driven Distributionally Robust Electric Vehicle Balancing for Autonomous Mobility-on-Demand Systems Under Demand and Supply Uncertainties
abstract
Electric vehicles (EVs) are being rapidly adopted due to their economic and societal benefits. Autonomous mobility-on-demand (AMoD) systems also embrace this trend. However, the long charging time and high recharging frequency of EVs pose challenges to efficiently managing EV AMoD systems. The complicated dynamic charging and mobility process of EV AMoD systems makes the demand and supply uncertainties significant when designing vehicle balancing algorithms. In this work, we design a data-driven distributionally robust optimization (DRO) approach to balance EVs for both the mobility service and the charging process. The optimization goal is to minimize the worst-case expected cost under both passenger mobility demand uncertainties and EV supply uncertainties. We then propose a novel distributional uncertainty sets construction algorithm that guarantees the produced parameters are contained in desired confidence regions with a given probability. To solve the proposed DRO AMoD EV balancing problem, we derive an equivalent computationally tractable convex optimization problem. Based on real-world EV data of a taxi system, we show that with our solution the average total balancing cost is reduced by 14.49%, and the average mobility fairness and charging fairness are improved by 15.78% and 34.51%, respectively, compared to solutions that do not consider uncertainties.
Sihong He, Shuo Han 0002, Lynn Pepin, Guang Wang 0001, Desheng Zhang 0002, John A. Stankovic, Fei Miao
IEEE Trans. Intell. Transp. Syst.8
2023 : Mobility-Driven Integration of Heterogeneous Urban Cyber-Physical Systems Under Disruptive Events
abstract
With the rapid development of cities, heterogeneous urban cyber-physical systems are designed to improve citizens’ experience, e.g., navigation and delivery service. However, the integration of services is not designed for disruptive events, an oversight that has rippling effects on service quality. For example, urban transportation systems consist of multiple transport modes that have complementary characteristics of capacities, speeds, and costs, facilitating smooth passenger transfers by planned schedules. Such integration may experience significantly increased delays during disruptions. Current solutions rely on a substitute service to transport passengers from and to affected areas using ad-hoc schedules and static routes, which are inefficient and do not utilize mobility patterns of mobile systems, e.g., dynamic passenger demand. To coordinate heterogeneous transportation systems under disruptions, we design a service to automatically select and integrate part of three systems (subway, bus, and taxi) using systems’ mobility patterns, e.g., predicted supply and demand. The service is presented in a normal version, eRoute, considering both subway and bus, and in a version taking taxis into account, called enhanced eRoute. We implement and evaluate eRoute with datasets including subway, bus and taxi, and a fare collection system. The data-driven evaluation results show that eRoute improves the ratio of served passengers per time interval by up to 11.5 times and reduces the average traveling time by up to 82.1 percent compared with existing solutions.
Yukun Yuan 0001, Desheng Zhang 0002, Fei Miao, John A. Stankovic, Tian He 0001, George J. Pappas, Shan Lin 0001
IEEE Trans. Mob. Comput.3
2023 Surrogate Lagrangian Relaxation: A Path to Retrain-Free Deep Neural Network Pruning
abstract
Network pruning is a widely used technique to reduce computation cost and model size for deep neural networks. However, the typical three-stage pipeline (i.e., training, pruning, and retraining (fine-tuning)) significantly increases the overall training time. In this article, we develop a systematic weight-pruning optimization approach based on surrogate Lagrangian relaxation (SLR), which is tailored to overcome difficulties caused by the discrete nature of the weight-pruning problem. We further prove that our method ensures fast convergence of the model compression problem, and the convergence of the SLR is accelerated by using quadratic penalties. Model parameters obtained by SLR during the training phase are much closer to their optimal values as compared to those obtained by other state-of-the-art methods. We evaluate our method on image classification tasks using CIFAR-10 and ImageNet with state-of-the-art multi-layer perceptron based networks such as MLP-Mixer; attention-based networks such as Swin Transformer; and convolutional neural network based models such as VGG-16, ResNet-18, ResNet-50, ResNet-110, and MobileNetV2. We also evaluate object detection and segmentation tasks on COCO, the KITTI benchmark, and the TuSimple lane detection dataset using a variety of models. Experimental results demonstrate that our SLR-based weight-pruning optimization approach achieves a higher compression rate than state-of-the-art methods under the same accuracy requirement and also can achieve higher accuracy under the same compression rate requirement. Under classification tasks, our SLR approach converges to the desired accuracy × faster on both of the datasets. Under object detection and segmentation tasks, SLR also converges 2× faster to the desired accuracy. Further, our SLR achieves high model accuracy even at the hardpruning stage without retraining, which reduces the traditional three-stage pruning into a two-stage process. Given a limited budget of retraining epochs, our approach quickly recovers the model’s accuracy.
Shanglin Zhou, Mikhail A. Bragin, Deniz Gurevin, Lynn Pepin, Fei Miao, Caiwen Ding
ACM Trans. Design Autom. Electr. Syst.5
2022 Stable and Efficient Shapley Value-Based Reward Reallocation for Multi-Agent Reinforcement Learning of Autonomous Vehicles
abstract
With the development of sensing and communication technologies in networked cyber-physical systems (CPSs), multi-agent reinforcement learning (MARL)-based methodologies are integrated into the control process of physical systems and demonstrate prominent performance in a wide array of CPS domains, such as connected autonomous vehicles (CAVs). However, it remains challenging to mathematically characterize the improvement of the performance of CAVs with communication and cooperation capability. When each individual autonomous vehicle is originally self-interest, we can not assume that all agents would cooperate naturally during the training process. In this work, we propose to reallocate the system's total reward efficiently to motivate stable cooperation among autonomous vehicles. We formally define and quantify how to reallocate the system's total reward to each agent under the proposed transferable utility game, such that communication-based cooperation among multi-agents increases the system's total reward. We prove that Shapley value-based reward reallocation of MARL locates in the core if the transferable utility game is a convex game. Hence, the cooperation is stable and efficient and the agents should stay in the coalition or the cooperating group. We then propose a cooperative policy learning algorithm with Shapley value reward reallocation. In experiments, compared with several literature algorithms, we show the improvement of the mean episode system reward of CAV systems using our proposed algorithm.
Songyang Han, Sanbao Su, Fei Miao
ICRA5
2022 DeResolver: A Decentralized Conflict Resolution Framework with Autonomous Negotiation for Smart City Services
abstract
As various smart services are increasingly deployed in modern cities, many unexpected conflicts arise due to various physical world couplings. Existing solutions for conflict resolution often rely on centralized control to enforce predetermined and fixed priorities of different services, which is challenging due to the inconsistent and private objectives of the services. Also, the centralized solutions miss opportunities to more effectively resolve conflicts according to their spatiotemporal locality of the conflicts. To address this issue, we design a decentralized negotiation and conflict resolution framework named DeResolver, which allows services to resolve conflicts by communicating and negotiating with each other to reach a Pareto-optimal agreement autonomously and efficiently. Our design features a two-step self-supervised learning-based algorithm to predict acceptable proposals and their rankings of each opponent through the negotiation. Our design is evaluated with a smart city case study of three services: intelligent traffic light control, pedestrian service, and environmental control. In this case study, a data-driven evaluation is conducted using a large dataset consisting of the GPS locations of 246 surveillance cameras and an automatic traffic monitoring system with more than 3 million records per day to extract real-world vehicle routes. The evaluation results show that our solution achieves much more balanced results, i.e., only increasing the average waiting time of vehicles, the measurement metric of intelligent traffic light control service, by 6.8% while reducing the weighted sum of air pollutant emission, measured for environment control service, by 12.1%, and the pedestrian waiting time, the measurement metric of pedestrian service, by 33.1%, compared to priority-based solution.
Yukun Yuan 0001, Meiyi Ma, Songyang Han, Desheng Zhang 0002, Fei Miao, John A. Stankovic, Shan Lin 0001
ACM Trans. Cyber Phys. Syst.5
2021 A Secure and Efficient Federated Learning Framework for NLP
abstract
Chenghong Wang, Jieren Deng, Xianrui Meng, Yijue Wang, Ji Li, Sheng Lin, Shuo Han, Fei Miao, Sanguthevar Rajasekaran, Caiwen Ding. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Chenghong Wang, Jieren Deng, Xianrui Meng, Yijue Wang, Ji Li 0006, Sheng Lin 0001, Shuo Han 0002, Fei Miao, Sanguthevar Rajasekaran, Caiwen Ding
EMNLP (1)8
2021 Data-Driven Fairness-Aware Vehicle Displacement for Large-Scale Electric Taxi Fleets
abstract
We are witnessing a rapid taxi electrification process due to the ever-increasing concern about urban air quality and energy security. A key difference between conventional gas taxis and electric taxis is their energy replenishment mechanisms, i.e., refueling or charging, which is reflected in two aspects: (i) much longer charging processes vs. short refueling processes and (ii) time-varying electricity prices vs. time-invariant gasoline prices during a day. The complicated charging issues (e.g., long charging time and dynamic charging pricing) potentially reduce electric taxis' daily operation time and profits, and also cause overcrowded charging stations during some off-peak charging pricing periods. Motivated by a set of findings obtained from a data-driven investigation, in this paper, we design a fairness-aware vehicle displacement system called FairMove to improve the overall profit efficiency and profit fairness of electric taxi fleets by considering both the passenger travel demand and taxi charging demand. We first formulate the electric taxi displacement problem as multi-agent deep reinforcement learning, and then we propose a centralized multi-agent actor-critic approach to tackle this problem. More importantly, we implement and evaluate FairMove with real-world streaming data from the Chinese city Shenzhen, including GPS data and transaction data from more than 20,100 electric taxis, coupled with the data of 123 charging stations, which constitute, to our knowledge, the largest all-electric taxi network in the world. The extensive experimental results show that our fairness-aware FairMove effectively improves the profit efficiency and profit fairness of the Shenzhen electric taxi fleet by 25.2% and 54.7%, respectively.
Guang Wang 0001, Shuxin Zhong, Shuai Wang 0008, Fei Miao, Zheng Dong 0002, Desheng Zhang 0002
ICDE4
2021 Enabling Retrain-free Deep Neural Network Pruning Using Surrogate Lagrangian Relaxation
abstract
Network pruning is a widely used technique to reduce computation cost and model size for deep neural networks. However, the typical three-stage pipeline, i.e., training, pruning and retraining (fine-tuning) significantly increases the overall training trails. In this paper, we develop a systematic weight-pruning optimization approach based on Surrogate Lagrangian relaxation (SLR), which is tailored to overcome difficulties caused by the discrete nature of the weight-pruning problem while ensuring fast convergence. We further accelerate the convergence of the SLR by using quadratic penalties. Model parameters obtained by SLR during the training phase are much closer to their optimal values as compared to those obtained by other state-of-the-art methods. We evaluate the proposed method on image classification tasks using CIFAR-10 and ImageNet, as well as object detection tasks using COCO 2014 and Ultra-Fast-Lane-Detection using TuSimple lane detection dataset. Experimental results demonstrate that our SLR-based weight-pruning optimization approach achieves higher compression rate than state-of-the-arts under the same accuracy requirement. It also achieves a high model accuracy even at the hard-pruning stage without retraining (reduces the traditional three-stage pruning to two-stage). Given a limited budget of retraining epochs, our approach quickly recovers the model accuracy.
Deniz Gurevin, Mikhail A. Bragin, Caiwen Ding, Shanglin Zhou, Lynn Pepin, Fei Miao
IJCAI7
2021 Automated Type-Aware Traffic Speed Prediction based on Sparse Intelligent Camera System
abstract
Many essential services for autonomous vehicles, e.g., navigation on high-quality maps, are designed based on the understanding of traffic conditions, e.g., travel time/speed on road segments, traffic flow, etc. However, most existing traffic condition models lack the consideration of the differentiation for vehicles with different types (e.g., personal vehicles or trucks) and thus they cannot satisfy some type-specific services, e.g., traffic-condition-based routing for autonomous vehicles with different types. To address this challenge, we design a novel vehicular mobility based sensing model called mDrive to predict the travel speed on the road segments, which is targeted for different types of vehicles by utilizing the camera data obtained from the traffic cameras equipped in the road intersections only, without any in-vehicle GPS devices. mDrive addresses the type-aware traffic speed prediction problem with sparse sensors based on three correlations: (1) the spatial correlation of travel speed on the connected road segments; (2) the temporal correlation of travel speed on the consecutive time slots; (3) the type correlation of different vehicular types’ speed on the same road segment. We implement mDrive on traffic camera data from the Chinese city Suzhou and evaluate it by using the detailed GPS data from personal vehicles, taxis, and trucks, with road contextual data as ground truth. The experiment show mDrive outperforms state-of-the-art methods by reducing 6.2% mean relative error on average for all types of vehicles.
Xiaoyang Xie, Kangjia Shao, Yang Wang 0015, Fei Miao, Desheng Zhang 0002
IROS4
2021 Data-driven Distributionally Robust Optimization For Vehicle Balancing of Mobility-on-Demand Systems
abstract
With the transformation to smarter cities and the development of technologies, a large amount of data is collected from sensors in real time. Services provided by ride-sharing systems such as taxis, mobility-on-demand autonomous vehicles, and bike sharing systems are popular. This paradigm provides opportunities for improving transportation systems’ performance by allocating ride-sharing vehicles toward predicted demand proactively. However, how to deal with uncertainties in the predicted demand probability distribution for improving the average system performance is still a challenging and unsolved task. Considering this problem, in this work, we develop a data-driven distributionally robust vehicle balancing method to minimize the worst-case expected cost. We design efficient algorithms for constructing uncertainty sets of demand probability distributions for different prediction methods and leverage a quad-tree dynamic region partition method for better capturing the dynamic spatial-temporal properties of the uncertain demand. We then derive an equivalent computationally tractable form for numerically solving the distributionally robust problem. We evaluate the performance of the data-driven vehicle balancing algorithm under different demand prediction and region partition methods based on four years of taxi trip data for New York City (NYC). We show that the average total idle driving distance is reduced by 30% with the distributionally robust vehicle balancing method using quad-tree dynamic region partitions, compared with vehicle balancing methods based on static region partitions without considering demand uncertainties. This is about a 60-million-mile or a 8-million-dollar cost reduction annually in NYC.
Fei Miao, Sihong He, Lynn Pepin, Shuo Han 0002, Abdeltawab M. Hendawi, Mohamed E. Khalefa, John A. Stankovic, George J. Pappas
ACM Trans. Cyber Phys. Syst.1
2020 Data-Driven Distributionally Robust Electric Vehicle Balancing for Mobility-on-Demand Systems under Demand and Supply Uncertainties
abstract
As electric vehicle (EV) technologies become mature, EV has been rapidly adopted in modern transportation systems, and is expected to provide future autonomous mobility-on-demand (AMoD) service with economic and societal benefits. However, EVs require frequent recharges due to their limited and unpredictable cruising ranges, and they have to be managed efficiently given the dynamic charging process. It is urgent and challenging to investigate a computationally efficient algorithm that provide EV AMoD system performance guarantees under model uncertainties, instead of using heuristic demand or charging models. To accomplish this goal, this work designs a data-driven distributionally robust optimization approach for vehicle supply-demand ratio and charging station utilization balancing, while minimizing the worst-case expected cost considering both passenger mobility demand uncertainties and EV supply uncertainties. We then derive an equivalent computationally tractable form for solving the distributionally robust problem in a computationally efficient way under ellipsoid uncertainty sets constructed from data. Based on E-taxi system data of Shenzhen city, we show that the average total balancing cost is reduced by 14.49%, the average unfairness of supply-demand ratio and utilization is reduced by 15.78% and 34.51% respectively with the distributionally robust vehicle balancing method, compared with solutions which do not consider model uncertainties.
Sihong He, Lynn Pepin, Guang Wang 0001, Desheng Zhang 0002, Fei Miao
IROS5
2019 p^2Charging: Proactive Partial Charging for Electric Taxi Systems
abstract
Electric taxis (e-taxis) have been increasingly deployed in metropolitan cities due to low operating cost and reduced emissions. Compared to conventional taxis, e-taxis require frequent recharging and each charge takes half an hour to several hours, which may result in unpredictable number of working taxis on the street. In current systems, E-taxi drivers usually charge their vehicles when the battery level is below a certain threshold, and then make a full charge. Although this charging strategy directly decreases the number of charges and the time to visit charging stations, our study reveals that it also significantly reduces the availability of number of taxis during busy hours with our data driven analysis. To meet dynamic passenger demand, we propose a new charging strategy: proactive partial charging (p2Charging), which allows an e-taxi to get partially charged before its remaining battery level is running too low. Based on this strategy, we propose a charging scheduling framework for e-taxis to meet dynamic passenger demand in spatial-temporal dimensions as much as possible while minimizing idle time to travel to charging stations and waiting time at charging stations. This work implements and evaluate our solution with large datasets that consist of (i) 7,228 regular internal combustion engine taxis and 726 e-taxis, (ii) an automatic taxi payment transaction collection system with total 62,100 records per day, (iii) charging station system, including 37 working charging stations over the city. The evaluation results show that p2Charging improves the ratio of unserved passengers by up to 83.2% on average and increases e-taxi utilization by up to 34.6% compared with ground truth and existing charging strategies.
Yukun Yuan 0001, Desheng Zhang 0002, Fei Miao, Jimin Chen, Tian He 0001, Shan Lin 0001
ICDCS3
2017 Artificial invariant subspace with potential functions for humanoid robot balancing
abstract
Existing trajectory planning based locomotion algorithms lack the analytic tools to fully comprehend energy based movements that would allow for full stability and mobility. Such drawbacks make humanoid robots' locomotion sensitive to external disturbances and compromise robots' agility in unstructured environment. In this work, we specifically focus on the push recovery problem for humanoid robots. We propose an approach to design a nonlinear controller that is robust to external disturbances. It allows the state of the rigid body dynamics to asymptotically converge to the subspace that meet the criteria of balancing, based on the properties of artificial invariant subspace and potential functions. Our algorithm is completely adaptive in real time without requiring trajectory planning in advance. We demonstrate the robustness of the proposed algorithm base on extensive push recovery experiments on the DARWIN-OP robot platform on flat terrains.
Fei Miao, Daniel D. Lee
IROS2
2016 Taxi Dispatch With Real-Time Sensing Data in Metropolitan Areas: A Receding Horizon Control Approach
abstract
Traditional taxi systems in metropolitan areas often suffer from inefficiencies due to uncoordinated actions as system capacity and customer demand change. With the pervasive deployment of networked sensors in modern vehicles, large amounts of information regarding customer demand and system status can be collected in real time. This information provides opportunities to perform various types of control and coordination for large-scale intelligent transportation systems. In this paper, we present a receding horizon control (RHC) framework to dispatch taxis, which incorporates highly spatiotemporally correlated demand/supply models and real-time Global Positioning System (GPS) location and occupancy information. The objectives include matching spatiotemporal ratio between demand and supply for service quality with minimum current and anticipated future taxi idle driving distance. Extensive trace-driven analysis with a data set containing taxi operational records in San Francisco, CA, USA, shows that our solution reduces the average total idle distance by 52%, and reduces the supply demand ratio error across the city during one experimental time slot by 45%. Moreover, our RHC framework is compatible with a wide variety of predictive models and optimization problem formulations. This compatibility property allows us to solve robust optimization problems with corresponding demand uncertainty models that provide disruptive event information.
Fei Miao, Shuo Han 0002, Shan Lin 0001, John A. Stankovic, Desheng Zhang 0002, Sirajum Munir, Hua Huang 0003, Tian He 0001, George J. Pappas
IEEE Trans Autom. Sci. Eng.1
2016 ATPC: Adaptive Transmission Power Control for Wireless Sensor Networks
abstract
Extensive empirical studies presented in this article confirm that the quality of radio communication between low-power sensor devices varies significantly with time and environment. This phenomenon indicates that the previous topology control solutions, which use static transmission power, transmission range, and link quality, might not be effective in the physical world. To address this issue, online transmission power control that adapts to external changes is necessary. This article presents ATPC, a lightweight algorithm for Adaptive Transmission Power Control in wireless sensor networks. In ATPC, each node builds a model for each of its neighbors, describing the correlation between transmission power and link quality. With this model, we employ a feedback-based transmission power control algorithm to dynamically maintain individual link quality over time. The intellectual contribution of this work lies in a novel pairwise transmission power control, which is significantly different from existing node-level or network-level power control methods. Also different from most existing simulation work, the ATPC design is guided by extensive field experiments of link quality dynamics at various locations over a long period of time. The results from the real-world experiments demonstrate that (1) with pairwise adjustment, ATPC achieves more energy savings with a finer tuning capability, and (2) with online control, ATPC is robust even with environmental changes over time.
Shan Lin 0001, Fei Miao, Gang Zhou 0002, Lin Gu 0001, Tian He 0001, John A. Stankovic, Sang Hyuk Son, George J. Pappas
ACM Trans. Sens. Networks2