EDBT 2026 Demo / reviewers in the wild / expert
Yangguang Cui
dblp:275/4394
· DBLP profile ↗
31ranked-venue papers
9as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 8 since 2021Computer networks · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two Heads Are Better Than One: Generalized Cross-Domain Federated Learning via Dual-PrototypeabstractCross-domain federated learning aims to collaboratively train a generalized model across clients with heterogeneous domain distributions without sharing data. Existing methods typically leverage prototypes to align intermediate representations among local models and enhance collaborative knowledge sharing, constructed either by directly aggregating class-center features across clients or by performing clustering to improve diversity. However, their performance is limited by the suboptimal ability to balance the learning of generalized and domain-specific features. To address this issue, this paper presents a novel dual-prototype guided FL framework named FedOrthrus, which decomposes the prototype into two components: i) the generalized prototype to capture cross-client domain-invariant features, and ii) the domain-specific prototype to extract the specific features of each domain. Specifically, the cloud server aggregates generalized prototypes to capture shared semantics across clients, thereby guiding each client to learn domain-invariant representations. Meanwhile, a clustering strategy is employed to adaptively construct domain-specific prototypes, ensuring that the representational capacity allocated to each domain is balanced according to its semantic complexity. Moreover, FedOrthrus employs a distribution-aware prototype construction scheme to dynamically assign the size of each part of prototypes, which enhances adaptability to different levels of domain heterogeneity. The experimental results on three datasets demonstrate that our FedOrthrus can achieve up to 14.56% and 3.96% accuracy improvement compared to traditional and state-of-the-art prototype-based FL methods. Our code is available at https://github.com/AAuZZ/FedOrthrus. Mingsheng Cao 0001, Tianci Chen, Ming Hu 0003, Zhuang Qi, Yangguang Cui, Junlong Zhou, Xiaofei Xie |
KDD (1) | 5 |
| 2026 | DHA-HFL: A Dynamic Hybrid Asynchronous Hierarchical Federated Learning Framework for NTN-Assisted IoT EnvironmentsabstractWith the growing demand for seamless global connectivity, the integration of Non-Terrestrial Networks (NTNs) and terrestrial networks presents new opportunities for largescale Federated Learning (FL). To reduce the communication overhead in NTN-assisted FL systems, NTN-assisted Hierarchical Federated Learning (HFL) frameworks have been proposed, which leverage aerial platforms as intermediate aggregation layers to effectively alleviate the communication burden on remote cloud servers. However, the high orbital altitude and mobility of Low-Earth Orbit (LEO) satellites introduce additional communication latency and frequent changes in network topology, making synchronous aggregation at the edge-cloud layer significantly degrade the training efficiency of HFL. Existing hybrid asynchronous HFL frameworks typically rely on static asynchronous triggering strategies that are not well-suited for the dynamic NTN environments and overlook the model drift caused by data heterogeneity among edge nodes. Furthermore, the data and heterogeneity among clients pose further challenges in HFL training performance. To address these issues, we propose a Dynamic Hybrid Asynchronous Hierarchical Federated Learning (DHA-HFL) framework for NTN scenarios. At the edge-cloud layer, we introduce a dynamic semi-asynchronous aggregation mechanism, consisting of an adaptive asynchronous aggregation trigger mechanism and an asynchronous dual-end aggregation strategy. The former adaptively adjusts the global aggregation timing to accommodate the dynamic NTN environments, while the latter enables efficient cloud-edge coordination to mitigate global model drift induced by data heterogeneity and inherent model staleness. At the client-edge layer, we employ a synchronous aggregation mechanism and propose a heterogeneityaware adaptive local iteration control strategy. By deriving the convergence bound of synchronous FL under non-IID data, we design an iterative algorithm to optimize the local iteration count of devices, minimizing training latency at the client-edge layer and accelerating global model convergence. Extensive experiments demonstrate the superior performance of DHA-HFL in terms of training latency, model accuracy, and convergence speed, providing an efficient distributed learning solution for NTN-assisted IoT scenarios. Siteng Liao, Tong Liu 0001, Yangguang Cui, Shaohua Wan 0001 |
IEEE Internet Things J. | 4 |
| 2026 | An Elastic Federated Learning Collaboration Framework for Computing-Constrained IoTabstractThrough exploiting decentralized data from multi-source Internet-of-Things (IoT) devices, federated learning (FL) can accomplish the training of deep neural network (DNN) models in a privacy-preserving manner to provide premium intelligent services. Due to portability considerations, most IoT devices are computing-constrained which cannot afford frequent DNN model training in FL. Existing approaches use model compression techniques to reduce computing cost of IoT devices, whereas accuracy degradation is inevitably incurred. To address this challenge, we propose an elastic federated learning collaboration framework, namely EFLCF, to accommodate limited computing resources of IoT devices. Specifically, we first design an FL-oriented elastic neural network model with multiple-width subnets, and couple it with an FL device-server collaboration framework to form EFLCF, thereby releasing computing cost pressure of IoT devices. We then develop a freezing-assisted wide-to-narrow training mechanism to realize efficient device-server distributed training and further reduce device computing cost. Finally, we design an entropy-based narrow-to-wide elastic inference mechanism to decrease computing cost of inference without compromising accuracy. Experiments demonstrate that compared to well-known benchmarks, our EFLCF can reduce up to 97.65% device computing cost and improve up to 48.3% accuracy in training, while reducing up to 42.5% computing cost in inference. Guobing Zou, Kun Cao 0001, Yangguang Cui, Tongquan Wei, Shiyan Hu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Multi-Agent Deep Reinforcement Learning-Based Distributed Task Assignment in Multi-UAV Cooperative Edge ComputingabstractLeveraging flexible deployment and wide-area coverage, Unmanned Aerial Vehicles (UAVs) assisted edge computing (EC) extends computation capability to the network edge and has become a key paradigm for improving the quality of service in the Internet of Things. However, constrained by onboard resources, UAVs struggle to independently process massive heterogeneous tasks in real time. Existing studies mainly focus on cooperation between UAVs and ground devices or the cloud, while the coupled cooperative relationships among UAVs and the performance scalability in large-scale UAV networks remain insufficiently characterized. To this end, we propose a distributed task assignment framework for a multi-UAV cooperative EC system, where each UAV explicitly accounts for cooperation with other UAVs and, leveraging controllable mobility, assigns a portion of local tasks to other UAVs or base stations to minimize the long-term system-wide total latency. To address the resulting joint trajectory control and task assignment optimization problem, we first formulate it as a Markov decision process. We then propose a Multi-head self-Attention critic-assisted MADDPG (MA$^{2}$DDPG) algorithm to train local online decision models for UAVs. Under the fully cooperative setting, we employ a centralized Critic to exploit global information and reduce parameter redundancy; concurrently, a multi-head self-attention mechanism is incorporated into the Critic to aggregate multi-UAV interaction information, delineate implicit cooperative dependencies, and alleviate the network input dimensionality curse arising from increasing UAV scale. Finally, extensive simulation experiments validate the effectiveness of the proposed method. Wendong Zuo, Siteng Liao, Yangguang Cui, Tong Liu 0001, Shaohua Wan 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2026 | Preference-Aware Fault-Tolerant Function Embedding in Energy-Harvesting Serverless Edge ComputingabstractServerless edge computing (SEC) that integrates serverless and edge computing paradigms has facilitated the deployment of intelligent Internet-of-things (IoT) applications. In SEC systems, energy efficiency and serverless pricing are essential to maintain operational sustainability. Nevertheless, most existing energy-saving techniques focus only on stable energy scenarios and are therefore inapplicable to energy-harvesting SEC systems powered by intermittent renewable sources. On the other hand, serverless pricing policies generally neglect the personalized perceptions of user quality-of-experience (QoE) preferences, thereby resulting in holistic user QoE degradation from a system perspective. Moreover, these approaches cannot guarantee functional correctness of serverless applications due to the appearance of computation and communication errors in practical SEC systems. To tackle these challenges, we investigate the preference-aware fault-tolerant function embedding problem for enhancing the holistic user QoE in energy-harvesting SEC systems. We first design a personalized QoE preference predictor to characterize trade-offs between service completion time and resultant service fees of individual users. Subsequently, we develop a reinforcement learning method to decide static function embedding decisions at the offline phase. Considering the intermittency of renewable sources, we further provide an energy-adaptive function replica freezing strategy at the online phase. Evaluations demonstrate that our approach boosts the holistic user QoE by 32.2% over state-of-the-art algorithms. Kun Cao 0001, Chaohong Tan, Yangguang Cui, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | One Arrow, Two Hawks: Sharpness-aware Minimization for Federated Learning via Global Model TrajectoryabstractFederated learning (FL) presents a promising strategy for distributed and privacy-preserving learning, yet struggles with performance issues in the presence of heterogeneous data distributions. Recently, a series of works based on sharpness-aware minimization (SAM) have emerged to improve local learning generality, proving to be effective in mitigating data heterogeneity effects. However, most SAM-based methods do not directly consider the global objective and require two backward pass per iteration, resulting in diminished effectiveness. To overcome these two bottlenecks, we leverage the global model trajectory to directly measure sharpness for the global objective, requiring only a single backward pass. We further propose a novel and general algorithm FedGMT to overcome data heterogeneity and the pitfalls of previous SAM-based methods. We analyze the convergence of FedGMT and conduct extensive experiments on visual and text datasets in a variety of scenarios, demonstrating that FedGMT achieves competitive accuracy with state-of-the-art FL methods while minimizing computation and communication overhead. Code is available at https://github.com/harrylee999/FL-SAM. Tong Liu 0001, Yangguang Cui, Xiaoqiang Li 0002 |
ICML | 3 |
| 2025 | Multi-Width Neural Network-Assisted Hierarchical Federated Learning in Heterogeneous Cloud-Edge-Device ComputingabstractFederated learning (FL), an emerging data-secure distributed training paradigm, unites massive isolated Internet of Things (IoT) device nodes to collaboratively train a global neural network (NN) model without the exposure of their local multimedia data. However, constrained by the synchronous NN model integration nature of FL, there is a training latency inconsistency among heterogeneous devices, which significantly deteriorates FL training efficiency. Meanwhile, frequent local NN training and transmission impose high energy consumption pressure on users. To tackle these issues, this paper proposes a premium multi-width NN-assisted hierarchical FL (HFL) framework in heterogeneous cloud-edge-device computing to achieve remarkable training speedup and energy conservation. Specifically, a heterogeneity-aware NN width coefficient determination algorithm, which flexibly assigns a subnet with a suitable width to each user device based on its computing ability, is first applied to shorten the HFL training latency. Subsequently, to integrate subnets with different width topologies, we design a width-aware adaptive NN model integration approach to effectively ensure the accuracy of the integrated global NN model. Finally, a latency-aware energy saving strategy is introduced to reduce energy consumption. Experimental results demonstrate that our proposed framework outperforms state-of-the-art benchmarks, and attains up to 42.42% enhancement in accuracy, 81.5% reduction in training latency, and 40.9% optimization in energy cost. Guobing Zou, Fei Xu 0009, Yangguang Cui, Tongquan Wei |
ACM Multimedia | 4 |
| 2025 | An inner-inter city graph neural network for predicting the course of COVID-19 casesabstractHighly infectious diseases like COVID-19 have affected millions worldwide, with transmission clearly linked to crowd mobility. However, research on the predictive power of city-level mobility remains limited. Most existing forecasting methods focus on state or national level, making it difficult for policymakers to quickly respond to emerging threats. To address this gap, we propose the Inner-Inter City Learner (IIL) for short- and mid-term COVID-19 case predictions at the city level, based on the correlation between human interaction and new cases. IIL consists of two key components: an inter-city transmission learner and an inner-city propagation learner. The first uses city mobility graphs and graph neural networks to learn how transmission in one city is influenced by others. The second, leveraging the highly contagious nature of the virus, captures key features of COVID-19 spread within cities. To overcome limited inner-city mobility data, we apply model-agnostic meta-learning to transfer common features across cities. We conduct various experiments and compare our methods with the state-of-art baselines. The results show the superiority of our method across various forecast horizons. Moussa Ndiaye, Tong Liu 0001, Yangguang Cui |
Intell. Data Anal. | 3 |
| 2025 | Latency-Aware Client Selection and Energy Management for Hierarchical Federated LearningabstractWith the prosperity of deep learning (DL) in Internet of Things (IoT) fields, federated learning (FL), viewed as a critical element of numerous DL-aided IoT intelligent applications, enables cooperative DL training across decentralized clients without revealing their personal data. However, the computational capacity heterogeneity and limited energy resources of IoT devices cause a huge negative influence on FL training in IoT intelligence applications. To address the above issues, this paper proposes an excellent distributed training mechanism for hierarchical FL to reduce training latency and energy cost for achieving the desirable accuracy. Specifically, taking into account computational capacity heterogeneity of clients, we first design a latency-regularization-aware client selection algorithm to appropriately select participating clients in training epochs and control their participation frequencies for boosting distributed training efficiency. Subsequently, after obtaining the selected client subset in each hierarchical FL training epoch, by leveraging the variable transmission delays of clients in distributed training, we propose a mixed integer linear programming-based transmission power management strategy for participating clients to alleviate their energy consumption burden. Extensive numerical results demonstrate that our proposed mechanism can attain 576.93% training speedup and achieve 10.94% accuracy enhancement compared with the baseline General FL, and yield up to 56.91% energy cost savings compared with the baseline HierFAVG. Jiamei Li, Kun Cao 0001, Yangguang Cui, Tong Liu 0001, Zhiquan Liu 0001 |
IEEE Internet Things J. | 4 |
| 2025 | Personalized Federated Learning for Green Industrial IoTabstractIn recent years, federated learning (FL) has gained increasing attention in industrial Internet-of-Things (IIoT) domains due to its privacy-preserving advantages. However, prior works commonly adopt a one-size-fits-all strategy for FL computation resource management and reward allocation, disregarding the time-varying participant states across different FL training rounds. Consequently, these methods fail to ensure the sustainability and active participation of IIoT devices in realistic FL deployments. To bridge this gap, we propose a personalized FL methodology for green IIoT systems powered by renewable energy sources. We first establish an incentive model along with its preference parameter-solving scheme to accurately characterize the incentive preferences of individual FL participants. Subsequently, a personalized participant scheduling approach is developed to accommodate dynamic resource usage patterns and diverse incentive preferences among FL participants. Our technique integrates empirical insights into conventional proximal policy optimization methods to accelerate policy learning within reinforcement learning frameworks. Experimental results on an FL prototype system show that our methodology improves the FL model accuracy by 25.92% compared with representative baseline algorithms. Kun Cao 0001, Yangguang Cui, Rui Xu 0013, Yuxia Sun, Zhiquan Liu 0001, Chaohong Tan |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Exploring High-Throughput and Low-Energy Federated Knowledge Tracking Using Pipeline and Energy-Efficient Techniques
Yangguang Cui, Jiamei Li, Lingzhi Shao |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Improving Generalization and Personalization in Long-Tailed Federated Learning via Classifier Retraining
Tong Liu 0001, Wenfeng Shen, Yangguang Cui, Weijia Lu |
Euro-Par (2) | 4 |
| 2024 | Bi-Level Reinforcement Learning-Based Task Offloading for Delay Optimization in Wireless Powered Mobile Edge ComputingabstractWireless powered mobile edge computing combines the benefits of wireless power transfer and mobile edge computing, enabling mobile devices (MDs) to overcome energy and computational constraints by obtaining energy from hybrid access point (HAP) and offloading computation tasks to HAP. However, given the devices' half-duplex nature, conflicts arise between time allocation and task offloading. Current research lacks a comprehensive examination of the joint impact of task offloading and time allocation on long-term delay. Furthermore, most existing centralized solutions may encounter issues such as dimensionality explosion and so on. This paper thoroughly investigates the task offloading problem in wireless powered mobile edge computing systems, focusing on minimizing long-term task completion delay under delay and energy constraints. To address this problem, we reformulate the task offloading process as a bi-level Markov decision process and convert it into an optimization problem of finding the optimal policy. Subsequently, we introduce a Bi-level Actor-Critic-based Task Offloading Approach (BiAC-TOA), where MDs and HAP leverage centrally trained agents for decisions on task offloading and time allocation in a decentralized manner. Finally, a series of experiments confirm the superiority of our proposed approach. Siteng Liao, Yangguang Cui, Tong Liu 0001, Yanmin Zhu 0006 |
MSN | 2 |
| 2024 | User-Distribution-Aware Federated Learning for Efficient Communication and Fast InferenceabstractDeep learning as a service (DLaaS) that promotes deep learning-based applications by selling computing services from IT companies to end-users has introduced potential privacy leaks from users and cloud servers. Federated learning (FL) provides an emerging distributed paradigm that enables numerous users to collaboratively train deep-learning models while protecting user privacy and data security. However, many FL-related existing works only focus on improving communication bottlenecks due to frequent model parameter transmission, but ignore the performance degradation incurred by imbalanced user distribution and high inference latency due to the high complexity of deep-learning models in the emerging IoT-edge-cloud FL. In this paper, we propose an efficient user-distribution-aware hierarchical FL for communication-efficient training and fast inference in the IoT-edge-cloud DLaaS architecture. Specifically, we propose a user-distribution-aware hierarchical FL architecture to cope with the performance degradation owing to the imbalanced user distribution. The proposed architecture also features a lightweight deep neural network that adopts the designed lightweight fire modules as components and has a side branch for communication-efficient training and fast inference. Extensive experiments demonstrate that the proposed schemes significantly boost the accuracy by up to 67.12%, save 47.98% communication costs, and accelerate inference by up to 87.24$\boldsymbol{\times}$compared to benchmarking methods. Yangguang Cui, Nuo Wang, Liying Li 0002, Chunwei Chang, Tongquan Wei |
IEEE Trans. Computers | 1 |
| 2024 | CPU-GPU Cooperative QoS Optimization of Personalized Digital Healthcare Using Machine Learning and Swarm IntelligenceabstractIn recent decades, the rapid advances in information technology have promoted a widespread deployment of medical cyber-physical systems (MCPS), especially in the area of digital healthcare. In digital healthcare, medical edge devices empowered by CPU-GPU (Graphics Processing Unit) cooperative multiprocessor system-on-chips (MPSoCs) have a great potential in processing and managing the massive amounts of health-related data. However, most of the existing works on CPU-GPU cooperative MPSoCs cannot maintain a high-precision workload estimation since they simply leverage the worst-case execution cycles to pessimistically predict the workload of digital healthcare applications. Besides, they neglect the personalized requirements of individual healthcare applications and the lifetime reliability demands of heterogeneous CPU-GPU cores. As a result, the normal functions of medical edge devices and the quality-of-services (QoS) of digital healthcare applications are likely to suffer from underlying failures and degradation. In this paper, we explore CPU-GPU cooperative QoS optimization of personalized digital healthcare applications running on reliability guaranteed edge devices with the help of machine learning and swarm intelligence techniques. We first develop two novel predictors: one is a machine learning based predictor for application workload estimation, and the other is a feature-driven predictor for application QoS estimation. We then incorporate the two predictors into a swarm intelligent application scheduling scheme upon the cooperative dual-population evolutionary algorithm (c-DPEA) to find optimal application mapping and partitioning settings. Experimental results show that our solution not only augments the average QoS of whole digital healthcare applications by 15.7%, but also balances the QoS of individual digital healthcare applications by 64.3%. Kun Cao 0001, Yangguang Cui, Liying Li 0002, Junlong Zhou, Shiyan Hu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Energy-Aware Incentive Mechanism for Hierarchical Federated Learning Using Water Filling TechniqueabstractFederated learning (FL) is an attractive industrial paradigm to accomplish distributed artificial intelligence (AI) training collaboratively in a data privacy-preserving manner. Most existing designs for FL systems assume that industrial user equipments (UEs) participate voluntarily in FL training. However, since both AI model training and transmission consume considerable energy, UEs are reluctant to participate without economic rewards. Hence, the lack of proper economic reward incentive mechanism results in low UE utility and frustrates UEs' enthusiasm for participating in training. To address the above challenge, in this article, we propose a two-phase energy-aware reward incentive mechanism for the edge-cloud-assisted hierarchical federated learning (HFL) system to optimize the overall UE utility, thereby, incentivizing UEs to participate more actively. Specifically, at the cloud server phase, we design an energy quantity-aware incentive mechanism for reasonably distributing rewards to its sub-edge-assisted FL systems. Subsequently, at the edge server phase, based on the quantitative analysis for the optimal reward allocation solution, we develop an energy-aware water filling-based reward incentive mechanism to adapt to individual needs of UEs and maximize the overall UE utility. Experiments verify that, compared to well-known benchmarks, our incentive mechanism can improve the overall UE utility by up to 55.94% and better incentivize UEs to participate in training. Yangguang Cui, Weiqin Tong, Tong Liu 0001, Kun Cao 0001, Junlong Zhou, Ming Xu 0010, Tongquan Wei |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | A Meta Reinforcement Learning-based Scheme for Adaptive Service Placement in Edge ComputingabstractService placement constitutes a prominent subject in edge computing, given its pivotal role in mitigating service latency. However, conventional heuristic placement strategies encounter difficulties in coping with the intricate dynamics of constantly evolving real-world settings, attributable to the non-stationary nature of service demand and time-varying network conditions. To tackle this challenge, deep reinforcement learning (DRL) techniques have received substantial attention. These approaches harness dynamic edge environments, encompassing cloud, edge, terminal devices, and wireless communication channels, to acquire placement policies through interactive learning. Nevertheless, a significant drawback of these learning-based approaches lie in their frequent necessity for retraining when confronted with novel edge environments, thereby resulting in inefficiencies in dynamic service placement across heterogeneous environments. To overcome this limitation, this paper proposes an innovative recurrent neural network (RNN)-based meta-reinforcement learning (meta-RL) technique. This technique capitalizes on prior historical experiences to expedite the learning of new policies in unfamiliar situations. Extensive simulation results are presented to illustrate the commendable performance of our proposed scheme. Jianfeng Rao, Tong Liu 0001, Yangguang Cui, Yanmin Zhu 0006 |
MSN | 3 |
| 2023 | Heterogeneity-Aware Federated Learning with Adaptive Local Epoch Size in Edge ComputingabstractFederated learning (FL) has been widely used in edge computing that enables artificial intelligence at the network edge as a distributed machine learning paradigm. In contrast to traditional cloud-based distributed training, the heterogeneity in edge computing may cause federated learning taking long training time. In this paper, we adapt control parameter (i.e., local epoch size) across devices to minimize wall-clock convergence time with joint consideration of resource heterogeneity and statistical heterogeneity. To analyze the influence of statistical heterogeneity, we derive a convergence upper bound for synchronous FL algorithm and establish the relationship between the number of training rounds and local epoch size under heterogeneous data distribution. Based on the convergence bound, we can solve the non-convex problem of minimizing FL training time with accuracy constraint and obtain near-optimal local epoch size. We develop a scheduling algorithm that estimates the statistical heterogeneity in initial training rounds and subsequently guides adaptive local training across devices. Practically, we evaluate our algorithm in a variety of heterogeneous scenarios. Extensive simulation results demonstrate that our algorithm performs high convergence speed over wall-clock time and spends less time reaching target accuracy compared with benchmark approaches. Wenying Yao, Tong Liu 0001, Yangguang Cui, Yanmin Zhu 0006 |
MSN | 3 |
| 2023 | FedEntropy: Information-entropy-aided training optimization of semi-supervised federated learning
Dongwei Qian, Yangguang Cui, Yufei Fu, Feng Liu 0039, Tongquan Wei |
J. Syst. Archit. | 2 |
| 2023 | MBSNN: A multi-branch scalable neural network for resource-constrained IoT devices
Liying Li 0002, Yangguang Cui, Nuo Wang, Fuke Shen, Tongquan Wei |
J. Syst. Archit. | 3 |
| 2023 | Filtering Out High Noise Data for Distributed Deep Neural NetworksabstractArtificial intelligence-based cyber-physical systems (CPS) applications have been spread across various fields such as smart cities, medical services, and industrial controls. When CPS devices are connected to a cloud server, big data streams generated by CPS devices impose enormous bandwidth pressure and exert excessive compute loads to the cloud server. Due to unpredictable environments and uncertainty in reality, these issues are mainly attributed to a large amount of high noise data captured and uploaded by CPS devices. To overcome these issues, this paper proposes a cyber-physical-cloud based framework for distributed deep neural networks (DDNNs) to prevent high noise data from being uploaded to the cloud. The proposed framework features a lightweight data filtering module enabled by depthwise separable convolutions to identify and filter out the high noise data that the cloud cannot recognize. Extensive experimental results demonstrate that the proposed data filtering module can achieve an accuracy of up to 83.72% in identifying high noise data and the proposed framework can effectively save bandwidth of up to 63.42% as compared to benchmarking methods. Note to Practitioners—This paper is motivated by the problems of enormous bandwidth pressure and excessive cloud compute loads in cyber-physical-cloud distributed computing paradigms. These problems are mainly caused by high noise data generated by CPS devices, because CPS devices often work in disturbing and unstable environments and there are uncontrollable uncertainties in reality. Especially for the emerging artificial intelligence-driven cyber-physical-cloud distributed paradigms, there is no existing research to solve the unnecessary transmission and cloud compute loads caused by high noise data. To tackle the challenge, this paper develops a novel cyber-physical-cloud distributed framework with data filtering capabilities to prevent high noise data from being uploaded. The proposed framework supports two popular loosely coupled and closely coupled distributed computing paradigms. Extensive experiments confirm that the proposed cyber-physical-cloud distributed framework can efficiently filter out high noise data and alleviate unnecessary transmission and needless cloud compute loads introduced by high noise data. Yangguang Cui, Liying Li 0002, Zhe Tao, Mingsong Chen 0001, Tongquan Wei |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2023 | Optimizing Training Efficiency and Cost of Hierarchical Federated Learning in Heterogeneous Mobile-Edge Cloud ComputingabstractFederated learning (FL), an emerging distributed machine learning (ML) technique, allows massive embedded devices and a server to work together for training a global ML model without collecting user data on a server. Most existing approaches adopt the traditional centralized FL paradigm with a single server: one is the cloud-centric FL paradigm and the other is the edge-centric FL paradigm. The cloud-centric FL paradigm is able to manage a large-scale FL system across massive user devices with high communication cost, whereas the edge-centric FL paradigm is capable of coordinating a small-scale FL system benefiting from the low communication delay over wireless networks. To fully exploit the advantages of both, in this article, we develop a distinctive hierarchical FL framework for the promising mobile-edge cloud computing (MECC) system, called HELCHFL, to achieve high-efficiency and low-cost hierarchical FL training. In particular, we formulate the corresponding theoretical foundation for our HELCHFL to ensure hierarchical training performance. Furthermore, to address the inherent communication and user heterogeneity issues of FL training, our HELCHFL develops a utility-driven and heterogeneity-aware heuristic user selection strategy to enhance training performance and reduce training delay. Subsequently, by analyzing and utilizing the slack time in FL training, our HELCHFL introduces a device operating frequency determination approach to reduce training energy cost. Experiments demonstrate that our HELCHFL can enhance the highest accuracy by up to 52.93%, gain the training speedup of up to 483.74%, and obtain up to 45.59% training energy savings compared to state-of-the-art baselines. Yangguang Cui, Kun Cao 0001, Junlong Zhou, Tongquan Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Reinforcement Learning-Based Device Scheduling for Renewable Energy-Powered Federated LearningabstractDue to its unique privacy protection advantages, emerging federated learning (FL) is regarded as a significant technique to enable Industry 4.0. However, the industrial deployment of FL encounters the primary obstacles of limited device energy and system communication resources. Nowadays, renewable energy-powered devices have been deployed in various industrial fields to tackle the challenges of unsustainable and limited energy of battery-powered devices. Inspired by this, this article proposes a novel FL protocol to groundbreakingly improve the performance of renewable energy-powered FL systems. Specifically, with the underlying theory of FL as the guide, the proposed protocol features a reinforcement learning-based device scheduling solution to adapt to intermittent renewable energy supply. Following this device scheduling solution, an integer linear programming-based bandwidth management scheme is introduced to optimize communication efficiency. Experimental results on two representative data distribution situations demonstrate that compared with the state-of-the-art schemes, our FL protocol can boost up to 36.63% and 50.99% accuracy, respectively. Yangguang Cui, Kun Cao 0001, Tongquan Wei |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | DBANet: A Dual Branch Attention-Based Deep Neural Network for Biological Iris RecognitionabstractEmerging iris recognition techniques are highly dependent on high-resolution iris images. However, existing iris recognition methods cannot effectively extract local texture features in low-resolution application scenarios, resulting in low recognition accuracy below expected. In this paper, we propose DBANet, a novel dual branch attention-based deep neural network for biological iris recognition that can achieve high accuracy for both high and low-resolution images. Specifically, we first design a spatial feature module with a small stride to preserve lower-level spatial detail features. Then, since the high-level feature can provide rich global context information, we propose a context feature module to generate high-level features. Finally, we develop a novel spatial attention module to fuse features generated by the above modules. We conduct the experiments on UBIRIS. v2, CASIA-V4-Distance, and MICHE-I datasets. Experimental results show that as compared to state-of-art methods, our proposed method can reduce equal error rates by up to 38.7%, 53.4%, and 71.9%, respectively. Yangguang Cui, Fuke Shen, Jianhua Shen, Tongquan Wei |
BIBM | 2 |
| 2022 | HELCFL: High-Efficiency and Low-Cost Federated Learning in Heterogeneous Mobile-Edge ComputingabstractFederated Learning (FL), an emerging distributed machine learning (ML), empowers a large number of embedded devices (e.g., phones and cameras) and a server to jointly train a global ML model without centralizing user private data on a server. However, when deploying FL in a mobile-edge computing (MEC) system, restricted communication resources of the MEC system, heterogeneity and constrained energy of user devices have a severe impact on FL training efficiency. To address these issues, in this article, we design a distinctive FL framework, called HELCFL, to achieve high-efficiency and low-cost FL training. Specifically, by analyzing the theoretical foundation of FL, our HELCFL first develops a utility-driven and greedy-decay user selection strategy to enhance FL performance and reduce training delay. Subsequently, by analyzing and utilizing the slack time in FL training, our HELCFL introduces a device operating frequency determination approach to reduce training energy costs. Experiments verify that our HELCFL can enhance the highest accuracy by up to 43.45 %, realize the training speedup of up to 275.03%, and save up to 58.25% training energy costs compared to state-of-the-art baselines. Yangguang Cui, Kun Cao 0001, Junlong Zhou, Tongquan Wei |
DATE | 1 |
| 2022 | ML-FORMER: Forecasting by Neighborhood and Long-Range Dependencies
Zengxiang Ke, Yangguang Cui, Liying Li 0002, Tongquan Wei |
ICANN (3) | 2 |
| 2022 | Edge Intelligent Joint Optimization for Lifetime and Latency in Large-Scale Cyber-Physical SystemsabstractIn recent years, the exploration on large-scale cyber–physical systems (CPSs) has become a fertile research field of significant impact. Large-scale CPS applications cover not only manufacturing and production areas but also daily living domains. Traditional solutions dedicated for large-scale CPSs mainly concentrate on the service latency or reliability optimization, but neglect the resultant negative impact on system lifetime. In this article, we conduct the first study on jointly optimizing the service latency and system lifetime subject to the constraints of reliability, energy consumption, and schedulability for large-scale CPSs. We propose an edge intelligent solution composed of offline and online phases. At the offline phase, the long short-term memory (LSTM) technique is leveraged to predict task offloading rates at individual user groups. Afterward, the multiobjective evolutionary algorithm with dual local search (DLS-MOEA) is exploited to determine optimal system static settings of computation offloading mapping and task replication number. At the online phase, an affinity-driven scheme incurring minimal system dynamic overheads is designed to deal with the inherent mobility of terminal users. We also build an algorithm validation platform upon which extensive simulation experiments are carried out. Experimental results show that our offline and online schemes outperform the state-of-the-art benchmarking methods by 27.1% and 43.5%, respectively. Kun Cao 0001, Yangguang Cui, Zhiquan Liu 0001, Wuzheng Tan, Jian Weng 0001 |
IEEE Internet Things J. | 2 |
| 2022 | Utility-driven renewable energy sharing systems for community microgrid
Liying Li 0002, Yangguang Cui, Fuke Shen, Meikang Qiu, Tongquan Wei |
J. Syst. Archit. | 3 |
| 2022 | Client Scheduling and Resource Management for Efficient Training in Heterogeneous IoT-Edge Federated LearningabstractFederated learning (FL) offers a promising paradigm that empowers numerous Internet of Things (IoT) devices to implement distributed learning on the premise of ensuring user privacy and data security. However, since FL adopts a synchronous distributed training mode, the heterogeneity of participating IoT devices and limited communication resources make FL encounter serious issues of low training efficiency in actual deployment. In this article, we propose an excellent FL policy for the heterogeneous IoT-edge FL system to improve distributed training efficiency. Specifically, first, by borrowing the idea of clustering, we explore an iterative self-organizing data analysis techniques algorithm (ISODATA)-based heterogeneous-aware client scheduling strategy to alleviate the issue of low training efficiency incurred by the heterogeneity of clients. Subsequently, to tackle the challenge of limited communication resources in FL, we first analyze the characteristics of the optimal resource block allocation solution theoretically and then introduce a mixed-integer linear programming (MILP)-based strategy to judiciously allocate resource blocks for scheduled clients. Comprehensive experimental results demonstrate that, compared with benchmarking strategies, our proposed FL policy can achieve up to 55.22% accuracy improvement in a relaxed time scenario, and attain up to$3.62\times $acceleration for reaching the specific expected accuracy. Yangguang Cui, Kun Cao 0001, Guitao Cao, Meikang Qiu, Tongquan Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Exploring Placement of Heterogeneous Edge Servers for Response Time Minimization in Mobile Edge-Cloud ComputingabstractIn the past few years, the study on placing edge servers for response time optimization in mobile edge-cloud computing systems has become increasingly popular. Most of the existing schemes neglect two important aspects: one is the heterogeneity of edge/cloud servers and the other is the response time fairness of base stations, which may significantly degrade the system quality of services to mobile users. In this article, we conduct the study of deploying heterogeneous edge servers to optimize the expected response time of both the whole and individual base stations. We propose an approach consisting of offline and online stages. At the offline stage, the optimal placement strategy of heterogeneous edge servers is produced by using an integer linear programming technique. At the online stage, a mobility-aware game-theory-based method is developed to deal with the dynamic characteristic of user movement. Experimental results reveal that compared to benchmarking methods, our approach not only reduces system-expected response time by 47.37%, but also improves response time fairness of base stations by 71.60%. Kun Cao 0001, Liying Li 0002, Yangguang Cui, Tongquan Wei, Shiyan Hu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Augmented Cross-Entropy-Based Joint Temperature Optimization of Real-Time 3-D MPSoC Systemsabstract3-D multiprocessor system-on-chip (MPSoC) systems can offer higher integration density, lower interaction cost, better bandwidth, and greater performance. However, vertically stacked silicon layers and limited heat dissipation paths result in high peak temperature and large temperature variation, which incur reliability reduction, lifetime decay, and performance degradation. In this article, we propose an offline augmented cross-entropy (CE)-based task scheduling strategy to jointly optimize peak temperature and temperature variation under the constraint of timeliness. Specifically, based on the conventional CE method, a heuristic iterative sampling method is designed to explore task-to-core assignment for balanced heat distribution between the top-layer and the bottom-layer cores. Subsequently, thermal characteristics of 3-D MPSoC systems are used to judiciously swap tasks between the two layers to improve the conventional CE-based task assignment and accelerate the iterative process. The peak temperature of individual cores is further reduced via sequencing, splitting, and slacking task execution. The experimental results demonstrate that compared to the existing state-of-the-art methods, the proposed scheme can reduce peak temperature by up to 8.02 °C and temperature variation by up to 24.78% without violating the timeliness of tasks. Yangguang Cui, Kun Cao 0001, Liying Li 0002, Junlong Zhou, Tongquan Wei, Shiyan Hu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |