Peiyuan Guan

dblp:221/7867 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-1248-487XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 1 first-author · 10 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
Truong Thanh Le, Hoang-Loc La, Amirhosein Taherkordi, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan
CCGrid6
2026 Personalized Federated Learning-Driven Beamforming Optimization for Integrated Sensing and Communication Systems
abstract
In this paper, we propose an Expectation-Maximization-based (EM) Personalized Federated Learning (PFL) framework for multi-objective optimization (MOO) in Integrated Sensing and Communication (ISAC) systems. In contrast to standard federated learning (FL) methods that handle all clients uniformly, the proposed approach enables each base station (BS) to adaptively determine its aggregation weight with the EM algorithm. Specifically, an EM posterior is computed at each BS to quantify the relative suitability between the global and each local model, based on the losses of models on their respective datasets. The proposed method is especially valuable in scenarios with competing communication and sensing objectives, as it enables BSs to dynamically adapt to application-specific trade-offs. To assess the effectiveness of the proposed approach, we conduct simulation studies under both objective-wise homogeneous and heterogeneous conditions. The results demonstrate that our approach outperforms existing PFL baselines, such as FedPer and pFedMe, achieving faster convergence and better multi-objective performance.
Zhou Ni, Sravan Reddy Chintareddy, Peiyuan Guan, Morteza Hashemi
CCNC3
2025 Cost Optimization for Serverless Edge Computing with Budget Constraints Using Deep Reinforcement Learning
Chen Chen 0073, Peiyuan Guan, Ziru Chen, Amirhosein Taherkordi, Fen Hou, Lin X. Cai
ICC2
2025 Opportunistic routing for mobile edge computing: A community detected and task priority aware approach
Jia Wu 0002, Tingyi Dai, Peiyuan Guan, Ziru Chen, Fangfang Gou, Amirhosein Taherkordi
Comput. Networks3
2025 Joint V2V Group Association and Transmission Power Control in Federated Vehicular Networks
abstract
With growing awareness of privacy protection, federated learning (FL) in vehicular network scenarios effectively addresses privacy concerns, leading to the development of federated vehicular networks (FVNs). In FVN, vehicles maintain a global model by transmitting local models and iterative processing, resulting in significant communication overhead. In the Internet of Vehicles (IoV), as the majority of the spectrum resources are allocated to Vehicle-to-Vehicle (V2V) communication, vehicles engaged in FL encounter uplink interference when these resources are reused. This compromises the efficacy of FL vehicle communication. To address this issue, we propose FedCDC, a new federated learning framework with adaptive weight clustering with knowledge distillation and channel sharing-based resource allocation. First, we implement a communication compression strategy based on clustering and distillation to alleviate transmission load. Then, we develop a resource allocation strategy to maximize the signal-to-interference-plus-noise ratio (SINR) for both FL vehicles and V2V groups, which investigates two critical components: 1) channel pairing and 2) power coordination. We formulate the optimization issue as a multiobjective optimization problem, that is, solved offline using the MIDACO solver. Due to the dynamic characteristics of FVN, we employ a multiagent deep deterministic policy gradient (MADDPG) to enhance efficiency. Finally, extensive experiments demonstrate that our approach significantly reduces data transmission and offers an efficient resource coordination plan, thus improving FL communication efficiency and ensuring the Quality of Service (QoS) for both V2V groups and FL vehicles.
Peiyuan Guan, Yushuai Li, Tianyi Li 0005, Zolaikha Zolfagharian, Tingwen Huang
IEEE Internet Things J.2
2024 Energy-Aware IoT Deployment Planning
abstract
Increasingly, the Internet of Things (IoT) is evolving toward an architecture consisting of sensing and actuation devices communicating with edge computers and storage systems. These "edge deployments" localize communication, computation, and storage for security, increased efficiencies (e.g. lower latency response), and reliability. In settings where electrical power infrastructure is lacking, however, these deployments typically rely on renewable energy and battery storage for power.
Peiyuan Guan, Animesh Dangwal, Amirhosein Taherkordi, Richard Wolski, Chandra Krintz
CF1
2024 Optimal Distribution of ML Models Over Edge for Applications with High Input Frequency
abstract
The rise of complex and sizeable Machine Learning (ML) models challenges traditional cloud computing models with respect to the high volume of incoming data which results in increased bandwidth usage and network congestion, as well as delays in inference. Such ML models are being rapidly developed thanks to advances in computing platforms and the real-time computing demands of ML-driven applications such as autonomous vehicles and video processing. To mitigate these challenges, ML model distribution and inference offloading to computing devices close to data sources have been explored, especially through partitioning the models across the IoT-Edge-Cloud continuum. Existing efforts in this area have not successfully mastered the fully automatic and efficient determination of optimal partition points. Additionally, they have not effectively integrated Early-Exit layers that allow for early termination of model inference at earlier stages when feasible. In this paper, we introduce a novel partitioning algorithm designed to distribute ML models across edge devices with the goal of reducing response time when facing high-rate input data streams. Our proposed approach leverages the principles of Dynamic Programming to determine optimal partition points and establish appropriate exit thresholds for Early-Exit layers. Our evaluation results reveal that, in the context of continuous, high-rate input data, our method consistently lowers the maximum round-trip time for processing inference requests compared to state-of-the-art methods such as NeuroSurgeon and Genetics.
Truong Thanh Le, Amirhosein Taherkordi, Frank Eliassen, Peiyuan Guan
CloudCom4
2024 Optimizing NOMA Transmissions to Advance Federated Learning in Vehicular Networks
abstract
Diverse critical data, such as location information and driving patterns, can be collected by IoT devices in vehicular networks to improve driving experiences and road safety. However, drivers are often reluctant to share their data due to privacy concerns. The Federated Vehicular Network (FVN) is a promising technology that tackles these concerns by transmitting model parameters instead of raw data, thereby protecting the privacy of drivers. Nevertheless, the performance of Federated Learning (FL) in a vehicular network depends on the joining ratio, which is restricted by the limited available wireless resources. To address these challenges, this paper proposes to apply Non-Orthogonal Multiple Access (NOMA) to improve the joining ratio in a FVN. Specifically, a vehicle selection and transmission power control algorithm is developed to exploit the power domain differences in the received signal to ensure the maximum number of vehicles capable of joining the FVN. Our simulation results demonstrate that the proposed NOMA-based strategy increases the joining ratio and significantly enhances the performance of the FVN. Index Terms—Federated Vehicular Network, NOMA
Ziru Chen, Zhou Ni, Peiyuan Guan, Lin X. Cai, Morteza Hashemi, Zongzhi Li
GLOBECOM3
2024 Context-aware Container Orchestration in Serverless Edge Computing
abstract
Adopting serverless computing to edge networks benefits end-users from the pay-as-you-use billing model and flexible scaling of applications. This paradigm extends the boundaries of edge computing and remarkably improves the quality of services. However, due to the heterogeneous nature of computing and bandwidth resources in edge networks, it is challenging to dynamically allocate different resources while adapting to the burstiness and high concurrency in serverless workloads. This article focuses on serverless function provisioning in edge networks to optimize end-to-end latency, where the challenge lies in jointly allocating wireless bandwidth and computing resources among heterogeneous computing nodes. To address this challenge, We devised a context-aware learning framework that adaptively orchestrates a wide spectrum of resources and jointly considers them to avoid resource fragmentation. Extensive simulation results justified that the proposed algorithm reduces over 95% of converge time while the end-to-end delay is comparable to the state of the art.
Peiyuan Guan, Chen Chen 0073, Ziru Chen, Lin X. Cai, Xing Hao, Amirhosein Taherkordi
GLOBECOM1
2024 AMbit: An Efficient Pruning Technique in Federated Learning for Edge Computing Systems
abstract
The Industrial Internet of Things (IIoT) has revolutionized industrial sectors with enhanced connectivity, data exchange, and predictive maintenance. However, it faces various challenges from non-IID data distributions and communication overheads, to the consistency and privacy of prediction models for maintenance. Federated Learning (FL) has been considered as a prevalent technique to address privacy concerns. On the other hand, Edge computing (EC) is being increasingly introduced to ensure low-latency data processing in IIoT systems, especially those with time-critical requirements, e.g., industrial robotics and motion control systems. Moreover, the complexity and design dimensions of today’s IIoT systems has lead to the development of large Machine Learning (ML) models with millions of parameters, e.g., using computer vision for field management. This introduces computational and privacy challenges in IIoT scenarios. Innovative FL optimization approaches such as pruning are aimed to tackle these challenges by reducing the number of training parameters, while they may impact accuracy. In this paper, we propose a new technique, called Adaptive Mean aBsolute devIaTion (AMbit), which is an innovative pruning approach optimizing data transmission without compromising model accuracy and inducing additional computation overhead. By dynamically comparing the difference between the current and the previous weight value, AMbit adapts better to parameter fluctuations at different stages, thereby accurately locating those parameters that have less impact on convergence. AMbit’s generality and efficiency are illustrated using MNIST and CIFAR-10 datasets, outperforming traditional Magnitude pruning in FL. For MNIST, AMbit reduces data uploads up to 43.75%, with an increase of 0.62% in accuracy. While for For CIFAR-10, AMbit achieved a 63.44% decrease with a 5.22% drop in accuracy.
Emad Hammami, Peiyuan Guan, Amirhosein Taherkordi, Amin Shahraki, Dapeng Lan
ICFEC2
2024 FedAPT: Joint Adaptive Parameter Freezing and Resource Allocation for Communication-Efficient Federated Vehicular Networks
abstract
Telematics technology development offers vehicles a range of intelligent and convenient functions, including navigation and mapping services, intelligent driving assistance, and intelligent traffic management. However, since these functions deal with sensitive information like vehicle location and driving habits, it is crucial to address concerns regarding information security and privacy protection. Federated learning (FL) is highly suitable for addressing such problems due to its characteristics, in which a client does not need to share private data and upload model parameters to a parameter server via the network. This results in the establishment of a federated vehicle network (FVN). As a distributed paradigm, the efficiency of communication is crucial in federated learning as it impacts all aspects of the FVN. This paper introduces a parameter freezing algorithm based on historical information to reduce the data transferred between the client and the parameter server in each round of communication, thus minimizing the communication overhead of federated learning. Additionally, we propose using a particle swarm algorithm to allocate network bandwidth to each vehicle based on the packet sizes sent by each vehicle (i.e., the non-freezing parameters) to minimize the communication latency in each FL round. Furthermore, due to the high time complexity of the particle swarm algorithm, we employ it to generate training data for training a transformer model with fast response and sufficient accuracy, thereby accelerating the bandwidth allocation process. Through extensive experiments, we prove the feasibility of our approach and its efficiency in improving communication in federated learning.
Jia Wu 0002, Tingyi Dai, Peiyuan Guan, Fangfang Gou, Amirhosein Taherkordi, Yushuai Li, Tianyi Li 0005
IEEE Internet Things J.3
2024 Big Data Analytics on Lung Cancer Diagnosis Framework With Deep Learning
abstract
As the segment of diseased tissue in PET images is time-consuming, laborious and low accuracy, this work proposes an automated framework for PET image screening, denoising and diseased tissue segmentation. First, taking into account the characteristics of PET images, the framework uses a differential activation filter to select whole-body images containing lesion tissue. Second, a new neural network containing residual connections which has powerful generalization performance compared with normal FCN network is proposed for PET image reconstruction and denoising. Finally, in the segmentation of lesion tissues, a custom clustering algorithm based on the density is used to distinguishe the lesion tissue part from the normal tissue. Tests on real medical PET images show that the whole automated framework has good performance and time cost in PET lesion image screening, image denoising and lesion tissue segmentation compared with other algorithms. The framework shows promising scientific study and application prospects.
Peiyuan Guan, Keping Yu, Wei Wei 0006, Yanlin Tan, Jia Wu 0002
IEEE Trans. Comput. Biol. Bioinform.1
2023 FedSSC: Joint client selection and resource management for communication-efficient federated vehicular networks
Peiyuan Guan, Amirhosein Taherkordi
Comput. Networks2
2023 Allocation of edge computing tasks for UAV-aided target tracking
Xiaoheng Deng, Jun Li 0084, Peiyuan Guan, Haichuan Ding
Comput. Commun.4
2023 Intelligent Delay-Aware Partial Computing Task Offloading for Multiuser Industrial Internet of Things Through Edge Computing
abstract
The development of Industrial Internet of Things (IIoT) and Industry 4.0 has completely changed the traditional manufacturing industry. Intelligent IIoT technology usually involves a large number of intensive computing tasks. Resource-constrained IIoT devices often cannot meet the real-time requirements of these tasks. As a promising paradigm, the mobile-edge computing (MEC) system migrates the computation intensive tasks from resource-constrained IIoT devices to nearby MEC servers, thereby obtaining lower delay and energy consumption. However, considering the varying channel conditions as well as the distinct delay requirements for various computing tasks, it is challenging to coordinate the computing task offloading among multiple users. In this article, we propose an autonomous partial offloading system for delay-sensitive computation tasks in multiuser IIoT MEC systems. Our goal is to provide offloading services with minimum delay for better Quality of Service (QoS). Enlighten by the recent advancement of reinforcement learning (RL), we propose two RL-based offloading strategies to automatically optimize the delay performance. Specifically, we first implement the$Q$-learning algorithm to provide a discrete partial offloading decision. Then, to further optimize the system performance with more flexible task offloading, the offloading decisions are given as continuous based on deep deterministic policy gradient (DDPG). The simulation results show that the$Q$-learning scheme reduces the delay by 23%, and the DDPG scheme reduces the delay by 30%.
Xiaoheng Deng, Jian Yin 0022, Peiyuan Guan, Naixue Xiong, Lan Zhang 0005, Shahid Mumtaz
IEEE Internet Things J.3
2022 Energy-Efficient UAV-Aided Target Tracking Systems Based on Edge Computing
abstract
Unmanned-aerial-vehicle (UAV)-aided target tracking has been applied in many important practical scenarios such as target vehicle tracking missions. However, the limited computation capability of UAVs can hardly support computation-intensive tasks, like the target tracking with real-time video processing. Inspired by the strong computation capabilities of edge computing servers nowadays, this article develops an energy-efficient UAV-aided target tracking system, where the video processing tasks can be offloaded from a UAV to the edge nodes (ENs) along its flight trajectory. To select appropriate offloading ENs for efficient task processing and energy saving, we formulate a cost minimization problem by jointly optimizing the task execution time and the offloading energy consumption. To devise a practical offloading strategy, we propose an energy-efficient UAV’s task distribution (EUTD) algorithm by jointly taking the different computation capabilities among ENs, time and energy requirements for different tasks, and fast-changing wireless channel conditions into account. Extensive experimental results demonstrate that our proposed algorithm can achieve significantly higher energy efficiency and lower latency in UAV-aided target tracking as compared with existing methods.
Xiaoheng Deng, Jun Li 0084, Peiyuan Guan, Lan Zhang 0005
IEEE Internet Things J.3
2020 Maximize Potential Reserved Task Scheduling for URLLC Transmission and Edge Computing
abstract
Emerging Internet of vehicles systems brings interesting new applications, such as VR entertainment systems in a car. These applications frequently generate data processing requirements and require a rapid response to ensure user experience. The combination of edge computing mode and ultra-low-latency communications (URLLC) traffic can better meet the requirements of the above scenarios. All requests for signal transmission and data processing can be considered a latency-limited task. We study scheduling strategies of these tasks intending to maximize overall utility for all users. We show that finding an optimal schedule for at least N tasks is NP-hard in the utility-maximizing issue. We propose a heuristic algorithm to maximize the overall utility of all users from the perspective of residual utility. To simulate the tolerance of delay for different tasks, we designed three utility curves: exponential, linear, and step. Simulation results show that the proposed algorithm outperforms the benchmark.
Peiyuan Guan, Xiaoheng Deng
VTC Fall1