Yufeng Zhan

dblp:173/1777 · DBLP profile ↗
← Back
61ranked-venue papers
13as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 3 first-author · 18 since 2021Computer networks · 18 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Security and privacy · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MCFL: A federated learning framework for network traffic classification in dynamic networks
Yongping He, Tijin Yan, Yufeng Zhan, Yuanqing Xia
Future Gener. Comput. Syst.3
2026 Patch-of-Interest ViT Inference Acceleration System for Edge-Assisted Video Analytics
abstract
The advent of edge computing has made real-time intelligent video analytics feasible. Previous works, based on traditional model architecture (e.g., CNN, RNN, etc.), employ various strategies to filter out non-region-of-interest content to minimize bandwidth and computation consumption but show inferior performance in adverse environments. Recently, visual foundation models based on transformers have shown great performance in adverse environments due to their amazing generalization capability. However, they require a large amount of computation power, which limits their applications in realtime intelligent video analytics. In this paper, we find visual foundation models like Vision Transformer (ViT) also have a dedicated acceleration mechanism for video analytics. To this end, we introduce Arena, an end-to-end edge-assisted video inference acceleration system based on ViT. We leverage the capability of ViT that can be accelerated through token pruning by only offloading and feeding Patches-of-Interest to the downstream models. Additionally, we design an adaptive keyframe inference switching algorithm tailored to different videos, capable of adapting to the current video content to jointly optimize accuracy and bandwidth. Through extensive experiments, our findings reveal that Arena can boost inference speeds by up to 1.58×, 1.82× and 1.98× on average while consuming only 47%, 31% and 27% of the bandwidth, respectively, all with high inference accuracy.
Haosong Peng, Hao Li 0075, Yufeng Zhan, Ren Jin, Yuanqing Xia
IEEE Trans. Computers4
2026 Scaling Blockchain via Dynamic Sharding
abstract
Sharding is considered a promising solution for scaling blockchain systems. However, most existing sharding systems have not considered the dynamics of the environment when making a sharding strategy, including the change of pending transactions, the leaving and joining of participants, and malicious attacks, which could cause performance instability and security issues. To address it, in this paper, we propose an intelligent and efficient dynamic sharding technology to advance the blockchain system performance and security. We first propose a formal and general evaluation framework for blockchain sharding in a dynamic environment, and conclude an optimization target for the system performance and security. To achieve a long-term benefit for the optimization target, a deep reinforcement learning (DRL)-based sharding approach has been proposed to intelligently make optimal sharding strategies. Next, we propose an adaptive resharding protocol to efficiently reduce the overhead introduced by dynamic sharding. Our experimental results illustrate that our proposed dynamic sharding in a simulation testbed can achieve 2.8 times transactions per second compared to traditional static sharding systems, and guarantee high security in a dynamic environment.
Zicong Hong, Xiaoyu Qiu, Wuhui Chen, Yufeng Zhan, Song Guo 0001
IEEE Trans. Dependable Secur. Comput.5
2026 Radiant: Efficient Timely Large-Scale Scene Analytics Based on Hierarchical Framework
abstract
With the advancement of computer vision, the recently emerged 3D Gaussian Splatting (3DGS) has increasingly become a popular scene analytics algorithm due to its outstanding performance. Existing cloud-based 3DGS architectures overlook the challenges in real-world environments when handling large-scale scene analysis. This exposes issues such as inefficiency, low security, lack of privacy, and limited scalability. In this paper, we propose Radiant, a hierarchical framework for large scene analytics in a heterogeneous cloud-edge-device system, which jointly considers high efficiency, privacy and security, and scalability. Via extensive empirical study, we find that it is crucial to partition the regions for each edge appropriately and allocate varying camera positions to each device for image collection and training. The core of Radiant is partitioning regions based on heterogeneous environment information and allocating workloads to each device accordingly. Furthermore, we provide a 3DGS model aggregation algorithm that enhances the quality and ensures the continuity of models' boundaries. Finally, we develop a testbed, and experiments demonstrate that Radiant improved reconstruction quality by up to 25.7% and reduced up to 79.6% end-to-end latency.
Haosong Peng, Tianyu Qi, Yufeng Zhan, Ren Jin, Hao Li 0075, Yalun Dai, Yuanqing Xia
IEEE Trans. Serv. Comput.3
2025 Sylva: Tailoring Personalized Adversarial Defense in Pre-trained Models via Collaborative Fine-tuning
abstract
The growing adoption of large pre-trained models in edge computing has made deploying model inference on mobile clients both practical and popular. These devices are inherently vulnerable to direct adversarial attacks, which pose a substantial threat to the robustness and security of deployed models. Federated adversarial training (FAT) has emerged as an effective solution to enhance model robustness while preserving client privacy. However, FAT frequently produces a generalized global model, which struggles to address the diverse and heterogeneous data distributions across clients, resulting in insufficiently personalized performance, while also encountering substantial communication challenges during the training process. In this paper, we propose Sylva, a personalized collaborative adversarial training framework designed to deliver customized defense models for each client through a two-phase process. In Phase 1, Sylva employs LoRA for local adversarial fine-tuning, enabling clients to personalize model robustness while drastically reducing communication costs by uploading only LoRA parameters during federated aggregation. In Phase 2, a game-based layer selection strategy is introduced to enhance accuracy on benign data, further refining the personalized model. This approach ensures that each client receives a tailored defense model that balances robustness and accuracy effectively. Extensive experiments on benchmark datasets demonstrate that Sylva can achieve up to 50× improvements in communication efficiency compared to state-of-the-art algorithms, while achieving up to 29.5% and 50.4% enhancements in adversarial robustness and benign accuracy, respectively.
Tianyu Qi, Lei Xue 0001, Yufeng Zhan, Xiaobo Ma 0001
CCS3
2025 DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
abstract
Novel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconstruction quality in vast environments. This paper presents DGTR, a novel distributed framework for efficient Gaussian reconstruction for sparse-view vast scenes. Our approach divides the scene into regions, processed independently by drones with sparse image inputs. Using a feed-forward Gaussian model, we predict high-quality Gaussian primitives, followed by a global alignment algorithm to ensure geometric consistency. Depth priors is incorporated to further enhance training, while a distillation-based model aggregation mechanism enables efficient reconstruction. Our method achieves high-quality large-scale scene reconstruction and novel-view synthesis in significantly reduced training times, outperforming existing approaches in both speed and scalability. We demonstrate the effectiveness of our framework on vast aerial scenes, achieving high-quality results within minutes. Code will released on our project page https://3d-aigc.github.io/DGTR.
Hao Li 0075, Haosong Peng, Chenming Wu, Weicai Ye, Yufeng Zhan, Chen Zhao 0011, Dingwen Zhang, Jingdong Wang 0001, Junwei Han 0001
ICRA6
2025 Uni-IL: Unified Incremental Learning of Vision-Language Models via Mixture of Attribute-Guided Experts
abstract
With the advent of parameter-efficient fine-tuning techniques for pre-trained vision-language models, interest in adapting them for various incremental learning scenarios has grown, i.e., sequential increments on task, class, and domain. However, no high-performance incremental learning framework has integrated these three incremental scenarios to achieve Unified Incremental Learning (Uni-IL) in complex settings. In this work, we propose an incremental learning framework called Mixture of Attribute-Guided Experts (MAGE) to alleviate the long-term forgetting in vision-language model incremental learning. Our approach involves acquiring image attribute knowledge via LLMs to form an attribute pool. We match the most relevant attributes as inputs to the Mixture of Experts (MoE) to fine-tune the pre-trained CLIP. Then the expert routers learn to select specific expert combinations based on the data and attribute features, alleviating catastrophic forgetting. The attribute pool incorporates both domain and class knowledge, enabling our approach to adapt to the three types of incremental learning scenarios and thus facilitating unified incremental learning. Through extensive experiments on our newly proposed benchmark and existing incremental learning scenarios, the results demonstrate that our proposed method not only performs well on the new Uni-IL tasks but also consistently outperforms previous state-of-the-art methods. Source code is available at https://github.com/ElectricField/Uni-IL.
Yufeng Zhan, Jie Zhang 0076, Yuanqing Xia
MMAsia2
2025 Meta-CAD: Few-shot anomaly detection for online social networks with meta-learning
Yongping He, Zihang Feng, Tijin Yan, Yufeng Zhan, Yuanqing Xia
Comput. Networks4
2025 A Kubernetes-based scheme for efficient resource allocation in containerized workflow
Yuanqing Xia, Chenggang Shan, Ke Tian, Yufeng Zhan
Future Gener. Comput. Syst.5
2025 Learning stabilizable symplectic ODE-net-based MPC for autonomous vehicle trajectory tracking
Hengheng Gong, Tijin Yan, Huahui Xie, Runze Gao, Yufeng Zhan, Yuanqing Xia
Neurocomputing5
2025 MVTC: Data and Knowledge-Based Distributed Multiview Information Mixing Network for Traffic Classification in Internet of Unmanned Agents
abstract
In the industrial IoT scenario, where massive data generation occurs, network traffic classification is crucial for operational security. The Internet of Unmanned Agents (IUA) is an emerging concept within the IoT framework. It focuses on the connectivity and interaction of various unmanned agents, such as drones, autonomous robots, and smart sensors. These unmanned agents collect and transmit large amounts of data in real-time, further contributing to the complexity of data in the IoT environment. The IUA aims to enable seamless cooperation and coordination among these agents, enhancing the overall efficiency and intelligence of industrial operations. The primary challenges in the IUA scenario lie in developing effective models and meeting real-time processing demands. Traditional methods struggle with large, high-dimensional data, while transformer-based models, although achieving good results, are difficult to deploy due to their size, training times, and complex tuning. In this article, we introduce a simple distributed architecture MVTC, which incorporates prior domain knowledge and delivers comparable results to transformer-based models but with shorter processing times and easier deployment. And it does not require large-scale unlabeled data for pretraining, which makes it highly suitable for real-world network traffic classification. The experiments demonstrate that the proposed method outperforms most existing approaches by up to 1.53% while using only 15.26% of the parameters.
Yang Liu 0038, Zhenkun Fu, Yufeng Zhan, Yuanqing Xia
IEEE Internet Things J.3
2025 Qora: Neural-Enhanced Interference-Aware Resource Provisioning for Serverless Computing
abstract
Serverless is an emerging cloud paradigm that offers fine-grained resource sharing through serverless functions. However, this resource sharing can cause interference, leading to performance degradation and QoS violations. Existing white box-based approaches for serverless resource provision often demand extensive expert knowledge, which is challenging to obtain due to the complexity of interference sources. This paper proposes Qora, a neural-enhanced interference-aware resource provisioning system for serverless computing. We model the resource provisioning of serverless functions as a novel combinatorial optimization problem, wherein the constraints on the queries per second are derived from neural network performance model. By leveraging neural networks to model the nonlinear performance fluctuations under various interference sources, our approach better captures the real-world behavior of serverless functions. To solve the formulated problem efficiently, rather than adopting commercial optimizer solvers like Gurobi, we propose a two-stage-VNS algorithm that searches discrete variables more efficiently and supports Sigmoid activations, avoiding introducing redundant discrete variables. Unlike pure machine learning methods lacking theoretical optimal guarantees, our approach is rigorously proven globally optimal based on optimization theory. We implement Qora on Kubernetes as a serverless system automating resource provisioning. Experimental results demonstrate that Qora reduces the QoS violation rate by 98% while reducing up to 35% resource costs compared with the state-of-the-arts. Note to Practitioners—From the perspective of cloud service providers, this paper considers the automatic resource provisioning for serverless functions. To improve hardware utilization, cloud providers tend to co-locate serverless functions on the same server. However, co-located functions compete for shared resources (memory bandwidth, L3 cache, etc.), which causes interference and leads to performance degradation and QoS violations. We use neural networks to build the performance models of interference-prone serverless functions and form the resource allocation optimization problem with neural network performance models as constraints. Compared to white box modeling methods, our neural network modeling adapts to complex and variable interference. Compared to deep reinforcement learning methods, our combinatorial optimization methods have stronger interpretability. In order to solve this optimization problem efficiently, we design the two-stage-VNS solution algorithm. We implement Qora on Kubernetes as a serverless system, which can automatically allocate computing resources. Experiments with small-scale real clusters and large-scale simulations demonstrate the effectiveness of Qora.
Ruifeng Ma, Yufeng Zhan, Chuge Wu, Zicong Hong, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.2
2025 Collaborative Neural Architecture Search for Personalized Federated Learning
abstract
Personalized federated learning (pFL) is a promising approach to train customized models for multiple clients over heterogeneous data distributions. However, existing works on pFL often rely on the optimization of model parameters and ignore the personalization demand on neural network architecture, which can greatly affect the model performance in practice. Therefore, generating personalized models with different neural architectures for different clients is a key issue in implementing pFL in a heterogeneous environment. Motivated by Neural Architecture Search (NAS), a model architecture searching methodology, this paper aims to automate the model design in a collaborative manner while achieving good training performance for each client. Specifically, we reconstruct the centralized searching of NAS into the distributed scheme called Personalized Architecture Search (PAS), where differentiable architecture fine-tuning is achieved via gradient-descent optimization, thus making each client obtain the most appropriate model. Furthermore, to aggregate knowledge from heterogeneous neural architectures, a knowledge distillation-based training framework is proposed to achieve a good trade-off between generalization and personalization in federated learning. Extensive experiments demonstrate that our architecture-level personalization method achieves higher accuracy under the non-iid settings, while not aggravating model complexity over state-of-the-art benchmarks.
Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Zicong Hong, Yufeng Zhan, Qihua Zhou
IEEE Trans. Computers5
2025 Robin: An Efficient Hierarchical Federated Learning Framework via a Learning-Based Synchronization Scheme
abstract
Hierarchical federated learning (HFL) extends traditional federated learning by introducing a cloud-edge-device framework to enhance scalability. However, the challenge of determining when devices and edges should aggregate models remains unresolved, making the design of an effective synchronization scheme crucial. Additionally, the heterogeneity in computing and communication capabilities, coupled with non-independent and identically distributed ( non-IID) data distributions, makes synchronization particularly complex. In this paper, we proposeRobin, a learning-based synchronization scheme for HFL systems. By collecting data such as models' parameters, CPU usage, communication time,etc., we design a deep reinforcement learning-based approach to decide the frequencies of cloud aggregation and edge aggregation, respectively. The proposed scheme well considers device heterogeneity, non-IID data and device mobility, to maximize the training model accuracy while minimizing the energy overhead. Meanwhile, we prove the convergence ofRobin's synchronization scheme. And we build an HFL testbed and conduct the experiments with real data obtained from Raspberry Pi and Alibaba Cloud. Extensive experiments under various settings are conducted to confirm the effectiveness ofRobin, which can improve 31.2% in model accuracy while reducing energy consumption by 36.4%.
Tianyu Qi, Yufeng Zhan, Peng Li 0017, Yuanqing Xia
IEEE Trans. Cloud Comput.2
2025 Multimodal Dual-Embedding Networks for Malware Open-Set Recognition
abstract
Malware open-set recognition (MOSR) is an emerging research domain that aims at jointly classifying malware samples from known families and detecting the ones from novel unknown families, respectively. Existing works mostly rely on a well-trained classifier considering the predicted probabilities of each known family with a threshold-based detection to achieve the MOSR. However, our observation reveals that the feature distributions of malware samples are extremely similar to each other even between known and unknown families. Thus, the obtained classifier may produce overly high probabilities of testing unknown samples toward known families and degrade the model performance. In this article, we propose the multi\modal dual-embedding networks, dubbed MDENet, to take advantage of comprehensive malware features from different modalities to enhance the diversity of malware feature space, which is more representative and discriminative for down-stream recognition. Concretely, we first generate a malware image for each observed sample based on their numeric features using our proposed numeric encoder with a re- designed multiscale CNN structure, which can better explore their statistical and spatial correlations. Besides, we propose to organize tokenized malware features into a sentence for each sample considering its behaviors and dynamics, and utilize language models as the textual encoder to transform it into a representable and computable textual vector. Such parallel multimodal encoders can fuse the above two components to enhance the feature diversity. Last, to further guarantee the open-set recognition (OSR), we dually embed the fused multimodal representation into one primary space and an associated sub-space, i.e., discriminative and exclusive spaces, with contrastive sampling and -bounded enclosing sphere regularizations, which resort to classification and detection, respectively. Moreover, we also enrich our previously proposed large-scaled malware dataset MAL-100 with multimodal characteristics and contribute an improved version dubbed MAL-100+. Experimental results on the widely used malware dataset Mailing and the proposed MAL-100+ demonstrate the effectiveness of our method.
Jingcai Guo, Yuanyuan Xu 0004, Wenchao Xu 0001, Yufeng Zhan, Yuxia Sun, Song Guo 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Feature Correlation-Guided Knowledge Transfer for Federated Self-Supervised Learning
abstract
Extensive attention has been paid to the application of self-supervised learning (SSL) approaches on federated learning (FL) to tackle the label scarcity problem. Previous works on federated SSL (FedSSL) generally fall into two categories: parameter-based model aggregation or data-based feature sharing to achieve knowledge transfer among multiple unlabeled clients. Despite the progress, they inevitably rely on some assumptions, such as homogeneous models or the existence of an additional public dataset, which hinder the universality of the training frameworks for more general scenarios (e.g., unlabeled clients with heterogeneous models). Therefore, in this article, we propose a novel and general method named federated self-supervised learning with feature-correlation-based aggregation (FedFoA) to tackle the above limitations. By exchanging feature correlation instead of model parameters or feature mappings, our approach reduces the discrepancies of local representations learning processes, thus promoting collaboration between heterogeneous clients. A factorization-based method is designed to extract the cross-feature relation matrix from local representations, which serves as a knowledge medium for the aggregation phase. We demonstrate that FedFoA is a heterogeneity-supportive and privacy-preserving training framework and can be easily compatible with state-of-the-art FedSSL methods. Extensive empirical experiments demonstrate our proposed approach outperforms the state-of-the-art methods by a significant margin.
Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Yufeng Zhan, Qihua Zhou, Yingchun Wang 0002
IEEE Trans. Neural Networks Learn. Syst.4
2025 6-D Object Pose Estimation Based on Point Pair Matching for Robotic Grasp Detection
abstract
The 6-D pose estimation is a critical work essential to achieve reliable robotic grasping. Currently, the prevalent method is reliant on keypoint correspondence. However, this approach hinges on the determination of object keypoint locations, alongside their detection and localization in real scenes. It also employs the random sample consensus (RANSAC)-based perspective-n-point (PnP) algorithm to solve the pose. Yet, it is nondifferentiable and incapable of backpropagation with loss during the training phase. Alternatively, the direct regression method, while speedy and differentiable, falls short in terms of pose estimation performance, and thus needs enhancement. In view of these gaps, we investigate PPM6D, a new method for 6-D object pose estimation based on regression and point pair matching. Our methodology begins with a proposed cross-fusion module, designed to achieve the fusion and complementation of RGB features and point cloud features. Subsequently, an attention module adjusts the features of the object's 3-D model. Finally, we design a point pair matching module for effective matching of points and characteristics, resulting in an integral matching and fusion. PPM6D is extensively trained and tested utilizing benchmark datasets like LINEMOD, occlusion LINEMOD (LINEMOD-occ), YCB-Video, and T-LESS dataset. Experimental results prove that PPM6D can outperform many keypoint-based pose estimation methods, given its relatively rapid speed, thereby offering novel regression-based pose estimation ideas. When applied to real-world scenarios of object pose estimation tasks and grasp tasks of an actual Baxter robot, PPM6D demonstrates superior performance as compared to most alternatives.
Sheng Yu 0009, Dihua Zhai, Yufeng Zhan, Wencai Wang, Yuyin Guan, Yuanqing Xia
IEEE Trans. Neural Networks Learn. Syst.3
2024 Tangram: High-Resolution Video Analytics on Serverless Platform with SLO-Aware Batching
abstract
Cloud-edge collaborative computing paradigm is a promising solution to high-resolution video analytics systems. The key lies in reducing redundant data and managing fluctuating inference workloads effectively. Previous work has focused on extracting regions of interest (RoIs) from videos and transmitting them to the cloud for processing. However, a naive Infrastructure as a Service (IaaS) resource configuration falls short in handling highly fluctuating workloads, leading to violations of Service Level Objectives (SLOs) and inefficient resource utilization. Besides, these methods neglect the potential benefits of RoIs batching to leverage parallel processing. In this work, we introduce Tangram, an efficient serverless cloud-edge video analytics system fully optimized for both communication and computation. Tangram adaptively aligns the RoIs into patches and transmits them to the scheduler in the cloud. The system employs a unique “stitching” method to batch the patches with various sizes from the edge cameras. Additionally, we develop an online SLO-aware batching algorithm that judiciously determines the optimal invoking time of the serverless function. Experiments on our prototype reveal that Tangram can reduce bandwidth consumption and computation cost up to 74.30 % and 66.35 %, respectively, while maintaining SLO violations within 5 % and the accuracy loss negligible.
Haosong Peng, Yufeng Zhan, Peng Li 0017, Yuanqing Xia
ICDCS2
2024 Probabilistic Time Series Modeling with Decomposable Denoising Diffusion Model
abstract
Probabilistic time series modeling based on generative models has attracted lots of attention because of its wide applications and excellent performance. However, existing state-of-the-art models, based on stochastic differential equation, not only struggle to determine the drift and diffusion coefficients during the design process but also have slow generation speed. To tackle this challenge, we firstly propose decomposable denoising diffusion model ($\text{D}^3\text{M}$) and prove it is a general framework unifying denoising diffusion models and continuous flow models. Based on the new framework, we propose some simple but efficient probability paths with high generation speed. Furthermore, we design a module that combines a special state space model with linear gated attention modules for sequence modeling. It preserves inductive bias and simultaneously models both local and global dependencies. Experimental results on 8 real-world datasets show that $\text{D}^3\text{M}$ reduces RMSE and CRPS by up to 4.6% and 4.3% compared with state-of-the-arts on imputation tasks, and achieves comparable results with state-of-the-arts on forecasting tasks with only 10 steps.
Tijin Yan, Hengheng Gong, Yongping He, Yufeng Zhan, Yuanqing Xia
ICML4
2024 Tomtit: Hierarchical Federated Fine-Tuning of Giant Models based on Autonomous Synchronization
abstract
With the quick evolution of giant models, the paradigm of pre-training models and then fine-tuning them for downstream tasks has become increasingly popular. The adapter has been recognized as an efficient fine-tuning technique and attracts much research attention. However, adapter-based fine-tuning still faces the challenge of lacking sufficient data. Federated fine-tuning has been recently proposed to fill this gap, but existing solutions suffer from a serious scalability issue, and they are inflexible in handling dynamic edge environments. In this paper, we propose Tomtit, a hierarchical federated fine-tuning system that can significantly accelerate fine-tuning and improve the energy efficiency of devices. Via extensive empirical study, we find that model synchronization schemes (i.e., when edge servers and devices should synchronize their models) play a critical role in federated fine-tuning. The core of Tomtit is a distributed design that allows each edge and device to have a unique synchronization scheme with respect to their heterogeneity in model structure, data distribution and computing capability. Furthermore, we provide a theoretical guarantee about the convergence of Tomtit. Finally, we develop a prototype of Tomtit and evaluate it on a testbed. Experimental results show that it can significantly outperform the state-of-the-art.
Tianyu Qi, Yufeng Zhan, Peng Li 0017, Yuanqing Xia
INFOCOM2
2024 Sonnet: A control-theoretic approach for resource allocation in cluster management
Ruifeng Ma, Yufeng Zhan, Yuanqing Xia, Chuge Wu, Liwen Yang, Runze Gao
Future Gener. Comput. Syst.2
2024 CAG-NSPDE: Continuous adaptive graph neural stochastic partial differential equations for traffic flow forecasting
Tijin Yan, Hengheng Gong, Yufeng Zhan, Yuanqing Xia
Neurocomputing3
2024 SGFM: Conditional Flow Matching for Time Series Anomaly Detection With State Space Models
abstract
The industrial Internet of Things (IoT) landscape is enriched with a diverse array of sensors, which are configured for the real-time monitoring and data collection to improve the production efficiency and optimizing industrial processes. The continual monitoring of IoT systems allows for promptly detection of anomalies in these time series data, thereby minimizing economic losses and ensuring the safe operation of the overall system. Existing deep learning-based methods for time series anomaly detection often rely on the generative models to learn the normal behavior of the data. However, these methods face challenges related to the speed and quality of data generation, ultimately impacting the overall detection performance of the models. In response to these challenges, this article proposes a new unsupervised anomaly detection method named SGFM. It combines state space models and graph neural networks to extract complex spatiotemporal dependencies in time series, which then serve as guidance to facilitate the learning process of a flow matching model, aiming for more refined predictions and consequently enhancing the effectiveness of anomaly detection. The effectiveness of SGFM is validated through the experiments on the three classic time series data sets. The results exhibits a notable improvement of up to 4% compared to the existing methods.
Yongping He, Tijin Yan, Yufeng Zhan, Zihang Feng, Yuanqing Xia
IEEE Internet Things J.3
2024 Workflow-Based Fast Data-Driven Predictive Control With Disturbance Observer in Cloud-Edge Collaborative Architecture
abstract
Data-driven predictive control (DPC) has been studied and applied in various scenarios. However, the challenge of computational efficiency remains. With the development of cloud computing, it provides potential solutions to the computational problem. Hence, this paper proposes a workflow-based fast DPC method in cloud-edge collaborative architecture. First, a workflow construction method of DPC is designed to make full use of the distributed computing ability of cloud computing. Next, to tackle the uncertainty in the cloud workflow processing, we design a cloud-edge collaborative scheme. In this scheme, a edge data-driven disturbance observer is proposed to estimate and compensate the uncertain event with guaranteed UUB stability. Then, An autonomous cloud control experimental system based on container technology is designed and implemented to execute the workflow-based DPC controller. Evaluations demonstrate that computation times are reduced by 45.19$\%$and 74.35$\%$for two real-time control examples, and by a maximum of 85.10$\%$for a high-dimensional control example.Note to Practitioners—This work is motivated by the challenge of the further combination of cloud computing and control system such as DPC. In the existing cloud-based control system, the computation mission of native control algorithm is deployed in a single cloud server directly. However, the structure of cloud environment is distributed, and the existing computation mode could not make full use of the parallel computing of cloud computing. Thus, the computation time could not be reduced significantly, and would still have serious effects on the control quality. In this work, a novel workflow-based DPC system in cloud-edge collaborative architecture is proposed, which decompose the native control mission into distributed cloud workflow with multiple smaller computation tasks. Then, an edge disturbance observer is designed to compensate the uncertainty occurring in the cloud workflow processing. As the evaluations show, the proposed workflow-based control system in cloud-edge collaborative architecture could be applied in the control mission with real-time requirement and the high-dimension control mission with intensive data.
Runze Gao, Qiwen Li, Li Dai 0001, Yufeng Zhan, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.4
2024 Fast Subspace Identification Method Based on Containerised Cloud Workflow Processing System
abstract
Subspace identification (SID) has been widely used in system identification and control fields, since it can estimate system models while only relying on the input and output data using reliable numerical operations. However, the high-dimension Hankel matrices are involved to store these data and used to obtain the system models, which increases the computation amount of SID and makes SID unsuitable for the large-scale or real-time identification tasks. In this paper, a novel fast SID method based on cloud workflow processing approach and container technology is proposed to accelerate the traditional algorithm. First, a workflow establishment method of SID is designed to match the distributed cloud environment, based on the computational feature of each calculation stage. Second, a containerised cloud workflow processing system is established to execute the logic-and data-dependent SID workflow mission based on the Kubernetes system. Finally, the experiments show that the computation time is reduced by at most$91.6\%$for the large-scale SID mission and decreased to within 20 ms for the real-time mission parameter.Note to Practitioners—Subspace identification has became a widely used method in various fields, including power grids, chemical processing, data-driven control, and fault detection. However, as systems become larger and more complex, the computational challenges increase. To address this issue, this paper proposes a workflow-based method for subspace identification that can be executed in a cloud environment to accelerate the process. This note outlines the steps that practitioners can take to apply this method. The first step is to design a workflow structure as proposed method in this paper. This structure should be customized to fit the specific needs of the practitioner’s application. The second step is to build a containerized cloud workflow processing system that can execute the workflow. This system should be based on the Kubernetes system and designed to handle the specific computational requirements of the workflow. Practitioners who work in fields where computational efficiency is crucial for system identification operations can benefit from the proposed method. By following the steps outlined above, practitioners can streamline the process of subspace identification and achieve improvements in computational efficiency.
Runze Gao, Yuanqing Xia, Liwen Yang, Yufeng Zhan
IEEE Trans Autom. Sci. Eng.5
2024 Classification-Based Diverse Workflows Scheduling in Clouds
abstract
Cloud workflow scheduling is a typical combinatorial optimization problem and becomes more challenging due to the increasing diversity of workflows. However, current research employs the same scheduling strategy on diverse workflows. In fact, a scheduling strategy may perform well on one workflow but poorly on other workflows owning to the unique characteristics of each workflow. Therefore, in practical applications, selecting suitable scheduling strategies for diverse workflows is a critical issue. To solve it, this paper investigates a diverse workflows scheduling problem and presents a classification-based workflow scheduling framework, which includes workflow parser, workflow classifier, workflow scheduler, resource manager and workflow status tracker, to manage and schedule diverse workflows using suitable strategies. Based on the framework, we propose a classification-based workflow scheduling algorithm (CWSA) to optimize the economic cost of workflow execution under deadline constraints. We conduct the experiments using diverse workflow instances randomly generated from five types of real-world workflows to evaluate the proposed CWSA approach. The results demonstrate the superiority of CWSA compared with the state-of-the-art approaches. Note to Practitioners—Diverse workflows (i.e., many workflows with various types, such as Montage, LIGO and Cybernetics) in clouds are widespread. How to efficiently schedule them in cloud is very important. This paper formulates the diverse workflows scheduling problem and proposes a CWSA to solve it. The basic idea of CWSA is to select a suitable scheduling strategy for each workflow. Specifically, in CWSA, we design a classification neural network architecture that consists of a graph neural network and a fully connected neural network to classify each workflow to its suitable deadline distribute strategy by its characteristics and deadline constraint. Then CWSA obtains the sub-deadlines of tasks and assigns tasks to appropriate VMs (Virtual Machines). Furthermore, as an important factor in workflow scheduling, the transmission time between dependent tasks is introduced into the graph neural network, which improves the classification accuracy.
Liwen Yang, Yuanqing Xia, Xiaopu Zhang, Lingjuan Ye, Yufeng Zhan
IEEE Trans Autom. Sci. Eng.5
2024 Zero-Norm Distance to Controllability of Linear Dynamic Networks
abstract
In this article, we consider the "nearest distance" from a given uncontrollable dynamical network to the set of controllable ones. We consider networks whose behaviors are represented via linear dynamical systems. The problem of interest is then finding the smallest number of entries/parameters in the system matrices, corresponding to the smallest number of edges of the networks, that need to be perturbed to achieve controllability. Such a value is called the zero-norm distance to controllability (ZNDC). We show genericity exists in this problem, so that other matrix norms (such as the 2-norm or the Frobenius norm) adopted in this notion are nonsense. For ZNDC, we show it is NP-hard to compute, even when only the state matrices can be perturbed. We then provide some nontrivial lower and upper bounds for it. For its computation, we provide two heuristic algorithms. The first one is by transforming the ZNDC into a problem of structural controllability of linearly parameterized systems, and then greedily selecting the candidate links according to a suitable objective function. The second one is based on the weighted -norm relaxation and the convex-concave procedure, which is tailored for ZNDC when additional structural constraints are involved in the perturbed parameters. Finally, we examine the performance of our proposed algorithms on several typical uncontrollable networks arising in multiagent systems.
Yuan Zhang 0016, Yuanqing Xia, Yufeng Zhan, Zhongqi Sun
IEEE Trans. Cybern.3
2024 Chiron: A Robustness-Aware Incentive Scheme for Edge Learning via Hierarchical Reinforcement Learning
abstract
Over the past few years, edge learning has achieved significant success in mobile edge networks. Few works have designed incentive mechanism that motivates edge nodes to participate in edge learning. However, most existing works only consider myopic optimization and assume that all edge nodes are honest, which lacks long-term sustainability and the final performance assurance. In this paper, we propose Chiron, an incentive-driven Byzantine-resistant long-term mechanism based on hierarchical reinforcement learning (HRL). First, our optimization goal includes both learning-algorithm performance criteria (i.e., global accuracy) and systematical criteria (i.e., resource consumption), which aim to improve the edge learning performance under a given resource budget. Second, we propose a three-layer HRL architecture to handle long-term optimization, short-term optimization, and byzantine resistance, respectively. Finally, we conduct experiments on various edge learning tasks to demonstrate the superiority of the proposed approach. Specifically, our system can successfully exclude malicious nodes and lazy nodes out of the edge learning participation and achieves 14.96% higher accuracy and 12.66% higher total utility than the state-of-the-art methods under the same budget limit.
Yi Liu 0057, Song Guo 0001, Yufeng Zhan, Leijie Wu, Zicong Hong, Qihua Zhou
IEEE Trans. Mob. Comput.3
2024 Rethinking Personalized Client Collaboration in Federated Learning
abstract
Federated Learning (FL) has gained considerable attention recently, as it allows clients to cooperatively train a global machine learning model without sharing raw data. However, its performance can be compromised due to the high heterogeneity in clients' local data distributions, commonly known as Non-IID (non-independent and identically distributed). Moreover, collaboration among highly dissimilar clients exacerbates this performance degradation. Personalized FL seeks to mitigate this by enabling clients to collaborate primarily with others who have similar data characteristics, thereby producing personalized models. We noticed that existing methods for assessing model similarity often do not capture the genuine relevance of client domains. In response, our paper enhances personalized client collaboration in FL by introducing a metric for domain relevance between clients. Specifically, to facilitate optimal coalition formation, we measure the marginal contributions of client models using coalition game theory, providing a more accurate representation of potential client domain relevance within the FL privacy-preserving framework. Based on this metric, we then adjust each client's coalition membership and implement a personalized FL aggregation algorithm that is robust to Non-IID data domain. We provide a theoretical analysis of the algorithm's convergence and generalization capabilities. Our extensive evaluations on multiple datasets, including MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100, and under varying Non-IID data distributions (Pathological and Dirichlet), demonstrate that our personalized collaboration approach consistently outperforms contemporary benchmarks in terms of accuracy for individual clients.
Leijie Wu, Song Guo 0001, Yaohong Ding, Wenchao Xu 0001, Yufeng Zhan, Anne-Marie Kermarrec
IEEE Trans. Mob. Comput.6
2024 Long-Term Adaptive VCG Auction Mechanism for Sustainable Federated Learning With Periodical Client Shifting
abstract
Federated Learning (FL) system needs to incentivize clients since they may be reluctant to participate in the resource consuming process. Existing incentive mechanisms fail to construct a sustainable environment for the long-term development of FL system: 1) They seldom focus on system economic properties (e.g., social welfare, individual rationality, and incentive compatibility) to guarantee client attraction. 2) Current online auction modeling methods divide the whole continual process into multiple independent rounds and solve them one-by-one, which breaks the correlation between each round. Besides, the inherent characteristics of FL system (model-agnostic and privacy-sensitive) also prevent it from the optimal strategy by precise mathematical analysis. 3) Current system modelings ignore the practical problem of periodical client shifting, which cannot adaptively update its strategy to handle system dynamics. To overcome the above challenges, this paper proposes a long-term adaptive Vickrey–Clarke–Groves (VCG) auction mechanism for FL system, which incorporate a multi-branch deep reinforcement learning (DRL) algorithm. First, VCG auction is the only one that can simultaneously guarantee all crucial economic properties. Second, we extend the economic properties to long-term forms and apply the experience-driven DRL algorithm to directly obtain long-term optimal strategy, without any prior system knowledge. Third, we reconstruct a multi-branch DRL network to accommodate periodical client shifting by adaptive decision head switching for different time periods. Finally, we theoretically prove he extended economic properties (i.e., IC) and conduct extensive experiments on several real-world datasets. Compared with state-of-the-art approaches, the long-term social welfare of FL system increases by 36% with a 37% reduction in payment. Besides, the multi-branch network can adaptively handle periodical client shifting on the timeline.
Leijie Wu, Song Guo 0001, Zicong Hong, Yi Liu 0057, Wenchao Xu 0001, Yufeng Zhan
IEEE Trans. Mob. Comput.6
2024 Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
abstract
As an emerging computing paradigm, edge computing offers computational resources closer to the data sources, helping to improve the service quality of many real-time applications. A crucial problem is designing a rational pricing mechanism to maximize the revenue of the edge computing service provider (ECSP). However, prior works have considerable limitations: clients are static and are required to disclose their preferences, which is impractical. To address this issue, we propose a novel sequential computation offloading mechanism, where the ECSP posts prices of computational resources with different configurations to clients in turn. Clients independently choose which computational resources to rent and how to offload based on their prices. Then Egret, a deep reinforcement learning-based approach that achieves maximum revenue, is proposed. Egret determines the optimal price and visiting orders online without infringing on clients’ privacy. Experimental results show that the revenue of ECSP in Egret is only 1.29% lower than Oracle and 23.43% better than the state-of-the-art when the client arrives dynamically.
Haosong Peng, Yufeng Zhan, Dihua Zhai, Xiaopu Zhang, Yuanqing Xia
IEEE Trans. Serv. Comput.2
2023 DaFKD: Domain-aware Federated Knowledge Distillation
abstract
Federated Distillation (FD) has recently attracted increasing attention for its efficiency in aggregating multiple diverse local models trained from statistically heterogeneous data of distributed clients. Existing FD methods generally treat these models equally by merely computing the average of their output soft predictions for some given input distillation sample, which does not take the diversity across all local models into account, thus leading to degraded performance of the aggregated model, especially when some local models learn little knowledge about the sample. In this paper, we propose a new perspective that treats the local data in each client as a specific domain and design a novel domain knowledge aware federated distillation method, dubbed DaFKD, that can discern the importance of each model to the distillation sample, and thus is able to optimize the ensemble of soft predictions from diverse models. Specifically, we employ a domain discriminator for each client, which is trained to identify the correlation factor between the sample and the corresponding domain. Then, to facilitate the training of the domain discriminator while saving communication costs, we propose sharing its partial parameters with the classification model. Extensive experiments on various datasets and settings show that the proposed method can improve the model accuracy by up to 6.02% compared to state-of-the-art baselines.
Haozhao Wang, Yichen Li 0006, Wenchao Xu 0001, Ruixuan Li 0001, Yufeng Zhan, Zhigang Zeng
CVPR5
2023 Hwamei: A Learning-Based Synchronization Scheme for Hierarchical Federated Learning
abstract
Federated learning (FL) enables collaborative model training among distributed devices without data sharing, but existing FL suffers from poor scalability because of global model synchronization. To address this issue, hierarchical federated learning (HFL) has been recently proposed to let edge servers aggregate models of devices in proximity, while synchronizing via the cloud periodically. However, a critical open challenge about how to design a good synchronization scheme (when devices and edges should be synchronized) is still unsolved. Devices are heterogeneous in computing and communication capability, and their data could be non-IID. No existing work can well synchronize various roles (e.g., devices and edge) in HFL to guarantee high learning efficiency and accuracy. In this paper, we propose a learning-based synchronization scheme for HFL systems. By collecting data such as edge models, CPU usage, communication time, etc., we design a deep reinforcement learning-based approach to decide the frequencies of cloud aggregation and edge aggregation, respectively. The proposed scheme well considers device heterogeneity, non-IID data and device mobility, to maximize the training model accuracy while minimizing the energy overhead. We build an HFL testbed and conduct experiments using real data obtained from Raspberry Pi and Alibaba Cloud. Extensive experimental results have confirmed the effectiveness of Hwamei.
Tianyu Qi, Yufeng Zhan, Peng Li 0017, Jingcai Guo, Yuanqing Xia
ICDCS2
2023 KubeAdaptor: A docking framework for workflow containerization on Kubernetes
Chenggang Shan, Yuanqing Xia, Yufeng Zhan, Jinhui Zhang 0003
Future Gener. Comput. Syst.3
2023 Look-ahead workflow scheduling with width changing trend in clouds
Liwen Yang, Lingjuan Ye, Yuanqing Xia, Yufeng Zhan
Future Gener. Comput. Syst.4
2023 Reliability-Aware and Energy-Efficient Workflow Scheduling in IaaS Clouds
abstract
Nowadays, more and more workflow applications with different computing requirements are migrated to clouds and executed with cloud resources. Workflow scheduling becomes a critical problem in the cloud environment, which focuses on meeting various quality of service (QoS) constraints. Workflow reliability and energy consumption are two essential parts in clouds and minimizing energy consumption for scheduling workflow with the reliability constraint is a challenging issue. In response to the challenge, we propose a workflow scheduling algorithm named REWS to reduce energy consumption and satisfy workflow reliability constraints. In REWS, a new sub-reliability constraint prediction strategy is adopted to break down the workflow reliability constraint to task sub-reliability constraints and the effectiveness of this strategy is proved. Moreover, an update method is adopted to adjust the task sub-reliability constraint for reducing energy consumption. In addition, a brief system framework which consists of five parts: workflow analyzer, reliability decomposer, resource manager, workflow scheduler and feedback processer is built to support the algorithm implementation of REWS. We conduct the experiments using both synthetic data and real-world data to evaluate the proposed REWS approach. The results demonstrate the superiority of REWS as compared with the state-of-the-art algorithms.Note to Practitioners—Workflow scheduling is a challenging issue in emerging trends of the cloud environment that focuses on satisfying various QoS constraints. In this paper, we investigate a reliability-aware and energy-efficient workflow scheduling problem in cloud computing. A novel workflow scheduling algorithm called REWS, is designed to reduce the energy consumption and meet the workfolw reliability constraint. The basic idea of REWS is to divide the workflow reliability constraint into task sub-reliability constraints and schedule tasks with an energy-efficient scheduling strategy. We conduct the experiments to evaluate the proposed REWS and the results demonstrate that REWS outperforms the state-of-the-art algorithms.
Lingjuan Ye, Yuanqing Xia, Siyuan Tao, Ce Yan, Runze Gao, Yufeng Zhan
IEEE Trans Autom. Sci. Eng.6
2023 Dynamic Scheduling Stochastic Multiworkflows With Deadline Constraints in Clouds
abstract
Nowadays, more and more workflows with different computing requirements are migrated to clouds and executed with cloud resources. In this work, we study the problem of stochastic multi-workflows scheduling in clouds and formalize this problem as an optimization problem that is NP-hard. To solve this problem, an efficient stochastic multi-workflows dynamic scheduling algorithm called SMWDSA is designed to schedule multi-workflows with deadline constraints for optimizing multi-workflows scheduling cost. The proposed SMWDSA consists of three stages including multi-workflows preprocessing, multi-workflow scheduling and scheduling feedback. In SMWDSA, a novel task sub-deadlines assignment stretagy is design to assign the task sub-deadlines to each task of multi-workflows for meeting workflow deadline constraints. Then, we propose a task scheduling method based on the minimal time slot availability to execution task for minimizing workflow scheduling cost while meetingt workflow deadlines. Finally, a scheduling feedback strategy is adopted to update the priorities and sub-deadlines of unscheduled tasks, for further minimizing workflow scheduling cost. We conduct the experiments using both synthetic data and real-world data to evaluate SMWDSA. The results demonstrate the superiority of SMWDSA as compared with the state-of-the-art algorithms. Note to Practitioners—Workflow scheduling in clouds is significantly challenging due to not only the large scale of workflows but also the elasticity and heterogeneity of cloud resources. Moreover, minimizing workflow scheduling cost and satisfying workflow deadlines are two critical issues in scheduling with cloud resources, especially the uncertainty of workflow arrive time and task execution time are considered. To meet workflow deadlines, it is an effective strategy to decompose workflow deadline constraints into task sub-deadline constraints. To minimize the workflow scheduling cost, each task in a workflow needs to be assigned to their most suitable VMs for execution. This article presents a novel workflow scheduling algorithm to schedule stochastic multi-workflows in clouds for optimizing multi-workflows scheduling cost and meeting workflows deadlines. This algorithm obtains the task sub-deadline constraints based on the characteristics of workflows for meeting the worklfow deadline constraint. Under the premise of meeting task deadlines, it schedules tasks to a VM with minimum the slot time, for minimizing the cost. Case studies based on well-known real-world workflows data sets suggest that it outperforms traditional ones in terms of success and cost of multi-workflows scheduling. It can thus aid the design and optimization of multi-workflows scheduling in a cloud environment. It can help practitioners better manage the scheduling cost and performance of real-world applications built upon cloud services.
Lingjuan Ye, Yuanqing Xia, Liwen Yang, Yufeng Zhan
IEEE Trans Autom. Sci. Eng.4
2023 A Fully Hybrid Algorithm for Deadline Constrained Workflow Scheduling in Clouds
abstract
With the migration of more and more workflows to clouds, the workflow scheduling in clouds (WSC) becomes a critical problem. Although many algorithms have been presented for WSC, there is still room and need for improvement. This paper formulates WSC as a constrained optimization problem that optimizes workflow execution cost within a workflow deadline constraint and proposes a fully hybrid workflow scheduling algorithm, called HPCP-PSO to solve it. Unlike previous works, HPCP-PSO is based on the repeated and alternated execution of two different methods, namely, the heuristic IaaS Cloud Partial Critical Paths (IC-PCP) and meta-heuristic Particle Swarm Optimization (PSO). Moreover, HPCP-PSO incorporates with two novel designs: 1) a new solution encoding strategy not only to sufficiently embody the elasticity of cloud resources, but also to reflect the scheduling relationship between assigned and unassigned tasks; 2) a solution repair strategy on each infeasible lease process to utilize a user-defined deadline more effectively and enhance the solution efficiency of the algorithm. Extensive experiments are conducted on four real-world scientific workflows and the results show that compared with IC-PCP, PSO, and HGSA, the proposed algorithm outperforms them on average by 35.83%, 70.53%, and 87.71% in terms of workflow execution cost.
Liwen Yang, Yuanqing Xia, Lingjuan Ye, Runze Gao, Yufeng Zhan
IEEE Trans. Cloud Comput.5
2023 A cost and makespan aware scheduling algorithm for dynamic multi-workflow in cloud environment
Yuanqing Xia, Yufeng Zhan, Li Dai 0001, Yuehong Chen
J. Supercomput.2
2022 Cycle: Sustainable Off-Chain Payment Channel Network with Asynchronous Rebalancing
abstract
Payment channel network (PCN) is a promising off-chain technology for blockchain scalability, but it suffers from poor sustainability in practice. In other words, due to the imbalanced transfer in channels, the balance in one direction of channels gradually becomes exhausted until the PCN is rebalanced via a consensus-based rebalancing protocol, during which the involved channels must be suspended. This paper presents Cycle, the first off-chain protocol for a sustainable PCN. It not only keeps the PCN at a balanced level consistently but also avoids the channel freeze incurred by the rebalancing protocol, leading to minimum failed payments and sustained PCN service, respectively. Cycle achieves these benefits based on a novel idea of asynchronous rebalancing. During the normal off-chain running, the participants share the information about their payments and asynchronously rebalance the PCN following the principle that payments along circular channels can cancel each other out. To guarantee security, the protocol resolves the disputes resulting from network latency or malicious participants by a message mechanism for synchronization and a smart contract for arbitration. Moreover, to address the privacy concern during the information sharing, a truncated Laplace mechanism is designed to achieve differential privacy. Finally, we provide a proof-of-concept implementation in Ethereum, over which a real data-based simulation shows that Cycle satisfies 31% more payments than the state-of-the-art technique.
Zicong Hong, Song Guo 0001, Rui Zhang 0080, Peng Li 0017, Yufeng Zhan, Wuhui Chen
DSN5
2022 Sustainable Federated Learning with Long-term Online VCG Auction Mechanism
abstract
Federated learning (FL) clients may be reluctant to participate in the energy-consuming FL unless they are incentivized. Existing incentive mechanisms seldom consider the economic properties, e.g., social welfare, individual rationality and incentive compatibility, which significantly limits the sustainability of FL to attract more clients. The Vickrey–Clarke–Groves (VCG) auction is an ideal mechanism for simultaneously guaranteeing all crucial economic properties to maximize social welfare. However, VCG auction cannot be applied directly to FL scenarios due to the following challenges: 1) It requires precise analytical derivation of the optimal strategy, which is unavailable due to the inherent model-unknown and privacy-sensitive characteristics of FL. 2) Current auction modeling decomposes the entire process into multiple independent rounds and solves them one-by-one, which breaks the successive correlation between rounds in the long-term training process of FL. To overcome these challenges, this paper presents a long-term online VCG auction mechanism for FL that employs an experience-driven deep reinforcement learning algorithm to obtain the optimal strategy. Besides, we extend long-term forms of the crucial economic properties for the successive FL process. Furthermore, knowledge transfer is applied to reduce the excessive training overhead arising from the VCG payment rules. By exploiting the environmental similarity among sub-auctions, we develop the strategy sharing to significantly cut the training time by half. Finally, we theoretically prove the extended economic properties and conduct extensive experiments on multiple real-world datasets. Compared with state-of-the-art approaches, the long-term social welfare of FL increases by 36% with a 37% reduction in payment.
Leijie Wu, Song Guo 0001, Yi Liu 0057, Zicong Hong, Yufeng Zhan, Wenchao Xu 0001
ICDCS5
2022 TFDPM: Attack detection for cyber-physical systems with diffusion probabilistic models
Tijin Yan, Yufeng Zhan, Yuanqing Xia
Knowl. Based Syst.3
2022 L4L: Experience-Driven Computational Resource Control in Federated Learning
abstract
As the large-scale deployment of machine learning applications, there is much research attention on exploiting a vast amount of data stored on mobile clients. To preserve data privacy, federated learning has been proposed to enable large-scale machine learning by massive clients without exposing raw data. Existing works of federated learning struggle for accelerating the learning process, but ignore the energy efficiency that is critical for resource-constrained clients. In this article, we propose to improve the energy efficiency of federated learning by lowering CPU cycle frequencies of clients who are faster in the training group. Based on this idea, we formulate an optimization problem aiming to minimize the total system cost defined as a weighted sum of learning time and energy consumption. Due to the hardness of the formulated optimization problem and unpredictability of network quality, we propose L4L (Learning for Learning), an experience-driven computational resource control approach based on the deep reinforcement learning, which can derive the near-optimal solution with only the clients’ bandwidth information in the previous training rounds. We conduct the experiments using both real-world traces and synthetic traces to evaluate the proposed L4L approach. The results demonstrate the superiority of L4L as compared with the state-of-the-art solutions.
Yufeng Zhan, Peng Li 0017, Leijie Wu, Song Guo 0001
IEEE Trans. Computers1
2022 Adaptive Federated Learning on Non-IID Data With Resource Constraint
abstract
Federated learning (FL) has been widely recognized as a promising approach by enabling individual end-devices to cooperatively train a global model without exposing their own data. One of the key challenges in FL is the non-independent and identically distributed (Non-IID) data across the clients, which decreases the efficiency of stochastic gradient descent (SGD) based training process. Moreover, clients with different data distributions may cause bias to the global model update, resulting in a degraded model accuracy. To tackle the Non-IID problem in FL, we aim to optimize the local training process and global aggregation simultaneously. For local training, we analyze the effect of hyperparameters (e.g., the batch size, the number of local updates) on the training performance of FL. Guided by the toy example and theoretical analysis, we are motivated to mitigate the negative impacts incurred by Non-IID data via selecting a subset of participants and adaptively adjust their batch size. A deep reinforcement learning based approach has been proposed to adaptively control the training of local models and the phase of global aggregation. Extensive experiments on different datasets show that our method can improve the model accuracy by up to 30 percent, as compared to the state-of-the-art approaches.
Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Deze Zeng, Yufeng Zhan, Rajendra Akerkar
IEEE Trans. Computers5
2021 Incentive-Driven Long-term Optimization for Edge Learning by Hierarchical Reinforcement Mechanism
abstract
Edge Learning is an emerging distributed machine learning in mobile edge network. Limited works have designed mechanisms to incentivize edge nodes to participate in edge learning. However, their mechanisms only consider myopia optimization on resource consumption, which results in the lack of learning algorithm performance guarantee and longterm sustainability. In this paper, we propose Chiron, an incentive-driven long-term mechanism for edge learning based on hierarchical deep reinforcement learning. First, our optimization goal combines learning-algorithms metric (i.e., model accuracy) with system metric (i.e., learning time, and resource consumption), which can improve edge learning quality under a fixed training budget. Second, we present a two-layer H-DRL design with exterior and inner agents to achieve both long-term and short-term optimization for edge learning, respectively. Finally, experiments on three different real-world datasets are conducted to demonstrate the superiority of our proposed approach. In particular, compared with the state-of-the-art methods under the same budget constraint, the final global model accuracy and time efficiency can be increased by 6.5 % and 39 %, respectively. Our implementation is available at https://github.com/Joey61Liuyi/Chiron.
Yi Liu 0057, Leijie Wu, Yufeng Zhan, Song Guo 0001, Zicong Hong
ICDCS3
2020 SkyChain: A Deep Reinforcement Learning-Empowered Dynamic Blockchain Sharding System
abstract
To overcome the limitations on the scalability of current blockchain systems, sharding is widely considered as a promising solution that divides the network into multiple disjoint groups processing transactions in parallel to improve throughput while decreasing the overhead of communication, computation, and storage. However, most existing blockchain sharding systems adopt a static sharding policy that cannot efficiently deal with the dynamic environment in the blockchain system, i.e., joining and leaving of nodes, and malicious attack. This paper presents SkyChain, a novel dynamic sharding-based blockchain framework to achieve a good balance between performance and security without compromising scalability under the dynamic environment. We first propose an adaptive ledger protocol to guarantee that the ledgers can merge or split efficiently based on the dynamic sharding policy. Then, to optimize the sharding policy under dynamic environment with high dimensional system states, a deep reinforcement learning-based sharding approach has been proposed, the goals of which include: 1) building a framework to evaluate the blockchain sharding systems from the aspects of performance and security; 2) adjusting the re-sharding interval, shard number and block size to maintain a long-term balance of the system’s performance and security. Experimental results show that SkyChain can effectively improve the performance and security of the sharding system without compromising scalability under the dynamic environment in the blockchain system.
Zicong Hong, Xiaoyu Qiu, Yufeng Zhan, Song Guo 0001, Wuhui Chen
ICPP4
2020 An Incentive Mechanism Design for Efficient Edge Learning by Deep Reinforcement Learning Approach
abstract
Emerging technologies and applications have generated large amounts of data at the network edge. Due to bandwidth, storage, and privacy concerns, it is often impractical to move the collected data to the cloud. With the rapid development of edge computing and distributed machine learning (ML), edge-based ML called federated learning has emerged to overcome the shortcomings of cloud-based ML. Existing works mainly focus on designing efficient learning algorithms, few works focus on designing the incentive mechanisms with heterogeneous edge nodes (EN) and uncertainty of network bandwidth. The incentive mechanisms affect various tradeoffs: (i) between computation and communication latency, and thus (ii) between the edge learning time and payment consumption. We fill this gap by designing an incentive mechanism that captures the tradeoff between latency and payment. Due to the network dynamics and privacy protection, we propose a deep reinforcement learning-based (DRL-based) solution that can automatically learn the best pricing strategy. To the best of our knowledge, this is the first work that applies the advances of DRL to design the incentive mechanism for edge learning. We evaluate the performance of the incentive mechanism using trace-driven experiments. The results demonstrate the superiority of our proposed approach as compared with the baselines.
Yufeng Zhan, Jiang Zhang 0003
INFOCOM1
2020 Experience-Driven Computational Resource Allocation of Federated Learning by Deep Reinforcement Learning
abstract
Federated learning is promising in enabling large-scale machine learning by massive mobile devices without exposing the raw data of users with strong privacy concerns. Existing work of federated learning struggles for accelerating the learning process, but ignores the energy efficiency that is critical for resource-constrained mobile devices. In this paper, we propose to improve the energy efficiency of federated learning by lowering CPU-cycle frequency of mobile devices who are faster in the training group. Since all the devices are synchronized by iterations, the federated learning speed is preserved as long as they complete the training before the slowest device in each iteration. Based on this idea, we formulate an optimization problem aiming to minimize the total system cost that is defined as a weighted sum of training time and energy consumption. Due to the hardness of nonlinear constraints and unawareness of network quality, we design an experience-driven algorithm based on the Deep Reinforcement Learning (DRL), which can converge to the near-optimal solution without knowledge of network quality. Experiments on a small-scale testbed and large-scale simulations are conducted to evaluate our proposed algorithm. The results show that it outperforms the start-of-the-art by 40% at most.
Yufeng Zhan, Peng Li 0017, Song Guo 0001
IPDPS1
2020 A Learning-Based Incentive Mechanism for Federated Learning
abstract
Internet of Things (IoT) generates large amounts of data at the network edge. Machine learning models are often built on these data, to enable the detection, classification, and prediction of the future events. Due to network bandwidth, storage, and especially privacy concerns, it is often impossible to send all the IoT data to the data center for centralized model training. To address these issues, federated learning has been proposed to let nodes use the local data to train models, which are then aggregated to synthesize a global model. Most of the existing work has focused on designing learning algorithms with provable convergence time, but other issues, such as incentive mechanism, are unexplored. Although incentive mechanisms have been extensively studied in network and computation resource allocation, yet they cannot be applied to federated learning directly due to the unique challenges of information unsharing and difficulties of contribution evaluation. In this article, we study the incentive mechanism for federated learning to motivate edge nodes to contribute model training. Specifically, a deep reinforcement learning-based (DRL) incentive mechanism has been designed to determine the optimal pricing strategy for the parameter server and the optimal training strategies for edge nodes. Finally, numerical experiments have been implemented to evaluate the efficiency of the proposed DRL-based incentive mechanism.
Yufeng Zhan, Peng Li 0017, Zhihao Qu, Deze Zeng, Song Guo 0001
IEEE Internet Things J.1
2020 An incentive mechanism design for mobile crowdsensing with demand uncertainties
Yufeng Zhan, Yuanqing Xia, Jiang Zhang 0003, Ting Li 0010, Yu Wang 0003
Inf. Sci.1
2020 A Deep Reinforcement Learning Based Offloading Game in Edge Computing
abstract
Edge computing is a new paradigm to provide strong computing capability at the edge of pervasive radio access networks close to users. A critical research challenge of edge computing is to design an efficient offloading strategy to decide which tasks can be offloaded to edge servers with limited resources. Although many research efforts attempt to address this challenge, they need centralized control, which is not practical because users are rational individuals with interests to maximize their benefits. In this article, we study to design a decentralized algorithm for computation offloading, so that users can independently choose their offloading decisions. Game theory has been applied in the algorithm design. Different from existing work, we address the challenge that users may refuse to expose their information about network bandwidth and preference. Therefore, it requires that our solution should make the offloading decision without such knowledge. We formulate the problem as a partially observable Markov decision process (POMDP), which is solved by a policy gradient deep reinforcement learning (DRL) based approach. Extensive simulation results show that our proposal significantly outperforms existing solutions.
Yufeng Zhan, Song Guo 0001, Peng Li 0017, Jiang Zhang 0003
IEEE Trans. Computers1
2020 Free Market of Multi-Leader Multi-Follower Mobile Crowdsensing: An Incentive Mechanism Design by Deep Reinforcement Learning
abstract
The explosive increase of mobile devices with built-in sensors such as GPS, accelerometer, gyroscope and camera has made the design of mobile crowdsensing (MCS) applications possible, which create a new interface between humans and their surroundings. Until now, various MCS applications have been designed, where the task initiators (TIs) recruit mobile users (MUs) to complete the required sensing tasks. In this paper, deep reinforcement learning (DRL) based techniques are investigated to address the problem of assigning satisfactory but profitable amount of incentives to multiple TIs and MUs as a MCS game. Specifically, we first formulate the problem as a multi-leader and multi-follower Stackelberg game, where TIs are the leaders and MUs are the followers. Then, the existence of the Stackelberg Equilibrium (SE) is proved. Considering the challenge to compute the SE, a DRL based Dynamic Incentive Mechanism (DDIM) is proposed. It enables the TIs to learn the optimal pricing strategies directly from game experiences without knowing the private information of MUs. Finally, numerical experiments are provided to illustrate the effectiveness of the proposed incentive mechanism compared with both state-of-the-art and baseline approaches.
Yufeng Zhan, Chi Harold Liu, Yinuo Zhao, Jiang Zhang 0003, Jian Tang 0008
IEEE Trans. Mob. Comput.1
2019 Future directions of networked control systems: A combination of cloud control and fog control approach
Yufeng Zhan, Yuanqing Xia, Athanasios V. Vasilakos
Comput. Networks1
2019 Energy-Efficient Distributed Mobile Crowd Sensing: A Deep Learning Approach
abstract
High-quality data collection is crucial for mobile crowd sensing (MCS) with various applications like smart cities and emergency rescues, where various unmanned mobile terminals (MTs), e.g., driverless cars and unmanned aerial vehicles (UAVs), are equipped with different sensors that aid to collect data. However, they are limited with fixed carrying capacity, and thus, MT's energy resource and sensing range are constrained. It is quite challenging to navigate a group of MTs to move around a target area to maximize their total amount of collected data with the limited energy reserve, while geographical fairness among those point-of-interests (PoIs) should also be maximized. It is even more challenging if fully distributed execution is enforced, where no central control is allowed at the backend. To this end, we propose to leverage emerging deep reinforcement learning (DRL) techniques for directing MT's sensing and movement and to present a novel and highly efficient control algorithm, called energy-efficient distributed MCS (Edics). The proposed neural network integrates convolutional neural network (CNN) for feature extraction and then makes decision under the guidance of multi-agent deep deterministic policy gradient (DDPG) method in a fully distributed manner. We also propose two enhancements into Edics with N-step return and prioritized experienced replay buffer. Finally, we evaluate Edics through extensive simulations and found the appropriate set of hyperparameters in terms of number of CNN hidden layers and neural units for all the fully connected layers. Compared with three commonly used baselines, results have shown its benefits.
Chi Harold Liu, Zheyu Chen 0001, Yufeng Zhan
IEEE J. Sel. Areas Commun.3
2018 Quality-aware incentive mechanism based on payoff maximization for mobile crowdsensing
Yufeng Zhan, Yuanqing Xia, Jinhui Zhang 0003
Ad Hoc Networks1
2018 Incentive mechanism in platform-centric mobile crowdsensing: A one-to-many bargaining approach
Yufeng Zhan, Yuanqing Xia, Jinhui Zhang 0003
Comput. Networks1
2018 Incentive Mechanism Design in Mobile Opportunistic Data Collection With Time Sensitivity
abstract
Mobile crowdsensing systems aim at providing various novel sensing applications by recruiting pervasive users with mobile devices, which are now equipped with enriched built-in sensors (e.g., GPS, microphone, camera, gyroscope, accelerometer, etc.). A key factor to enable such systems is substantial participation of large amount of mobile users. In this paper, we focus on the data collection in mobile opportunistic crowdsensing, where the data can be transferred between mobile users via opportunistic device-to-device communications. The goal is to deliver the sensed data from the collector to the corresponding requester, which can maximize the collector's rewards. Here, we assume that the data collection has time-sensitive characteristics, i.e., the reward is time-sensitive. We consider selfish mobile users with rational behaviors, and propose a credit-based incentive-aware mechanism to stimulate mobile users to participate in data collection for mobile opportunistic crowdsensing. Particularly, we propose an effective mechanism to define the expected rewards for the sensed data, and formulate the sensed data trading as a two-person cooperative game, whose solution is obtained through the Nash bargaining theory. Extensive simulations based on both synthetic and real-world mobility traces are conducted to validate the efficiency of our incentive-aware mechanisms.
Yufeng Zhan, Yuanqing Xia, Jinhui Zhang 0003, Yu Wang 0003
IEEE Internet Things J.1
2017 Incentive mechanism for computation offloading using edge computing: A Stackelberg game approach
Yang Liu 0038, Changqiao Xu, Yufeng Zhan, Zhixin Liu 0001, Jianfeng Guan, Hongke Zhang
Comput. Networks3
2016 Incentive Mechanism for Crowdsourced Mobile Video Offloading
abstract
In this work, we propose a time-sensitive incentive-aware mechanism for mobile video offloading by using the idea of crowdsourcing, where video packet holder cooperates with mobile users to deliver video packets to destination. The objective is to maximize video provider and mobile relay users' payoffs. We formulate the interaction among video packet provider and mobile relay users as a two-person cooperative game, where the video packets are treated as commodities. We apply the Nash bargain solution to obtain the optimal cooperation decision and payment. We carry out extensive simulation based on the real-world traces to validate the superiority of our proposed scheme.
Yufeng Zhan, Yang Liu 0038, Yuanqing Xia, Fan Li 0001, Hongyi Wu
MSN1
2016 GTS size adaptation algorithm for IEEE 802.15.4 wireless networks
Yufeng Zhan, Yuanqing Xia, Mashood Anwar
Ad Hoc Networks1
2016 TDMA-Based IEEE 802.15.4 for Low-Latency Deterministic Control Applications
abstract
In this paper, we propose a technique for making IEEE 802.15.4 standard suitable for low-latency deterministic networks for wireless control applications where the cyclic update time of sensor's information is not more than 10 ms. The IEEE 802.15.4 has shown good characteristics for deterministic networks in beacon-enabled mode, but with minimum cyclic update time for personal area network devices not less than 15.36 ms. Moreover, it is unsuitable for hard real-time systems when used in nonbeacon-enabled mode due to the random nature of channel access protocol. The proposed technique employs time-division multiple access (TDMA)-based protocol that works on slightly modified IEEE 802.15.4 in star topology. Each end device in the network transmits its data frame after certain time delay in response to periodic requests from the coordinator. This time delay is optimized for better channel bandwidth utilization and reliable data exchange. While this TDMA-based protocol eliminates the risk of frame collisions to a great extent, medium-access control sublayer modifications reduce the nondeterminism of the network and increase the bandwidth efficiency. A mathematical model is proposed to design and develop a practical communication network with commercial off-the-shelf radios that can predict worst case arrival time of data frames. Experiments were conducted to evaluate the model for the proposed protocol.
Mashood Anwar, Yuanqing Xia, Yufeng Zhan
IEEE Trans. Ind. Informatics3