Kun Cao 0001

dblp:65/8519-1 · DBLP profile ↗
← Back
37ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0003-4872-4908ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 7 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 An Elastic Federated Learning Collaboration Framework for Computing-Constrained IoT
abstract
Through exploiting decentralized data from multi-source Internet-of-Things (IoT) devices, federated learning (FL) can accomplish the training of deep neural network (DNN) models in a privacy-preserving manner to provide premium intelligent services. Due to portability considerations, most IoT devices are computing-constrained which cannot afford frequent DNN model training in FL. Existing approaches use model compression techniques to reduce computing cost of IoT devices, whereas accuracy degradation is inevitably incurred. To address this challenge, we propose an elastic federated learning collaboration framework, namely EFLCF, to accommodate limited computing resources of IoT devices. Specifically, we first design an FL-oriented elastic neural network model with multiple-width subnets, and couple it with an FL device-server collaboration framework to form EFLCF, thereby releasing computing cost pressure of IoT devices. We then develop a freezing-assisted wide-to-narrow training mechanism to realize efficient device-server distributed training and further reduce device computing cost. Finally, we design an entropy-based narrow-to-wide elastic inference mechanism to decrease computing cost of inference without compromising accuracy. Experiments demonstrate that compared to well-known benchmarks, our EFLCF can reduce up to 97.65% device computing cost and improve up to 48.3% accuracy in training, while reducing up to 42.5% computing cost in inference.
Guobing Zou, Kun Cao 0001, Yangguang Cui, Tongquan Wei, Shiyan Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 Preference-Aware Fault-Tolerant Function Embedding in Energy-Harvesting Serverless Edge Computing
abstract
Serverless edge computing (SEC) that integrates serverless and edge computing paradigms has facilitated the deployment of intelligent Internet-of-things (IoT) applications. In SEC systems, energy efficiency and serverless pricing are essential to maintain operational sustainability. Nevertheless, most existing energy-saving techniques focus only on stable energy scenarios and are therefore inapplicable to energy-harvesting SEC systems powered by intermittent renewable sources. On the other hand, serverless pricing policies generally neglect the personalized perceptions of user quality-of-experience (QoE) preferences, thereby resulting in holistic user QoE degradation from a system perspective. Moreover, these approaches cannot guarantee functional correctness of serverless applications due to the appearance of computation and communication errors in practical SEC systems. To tackle these challenges, we investigate the preference-aware fault-tolerant function embedding problem for enhancing the holistic user QoE in energy-harvesting SEC systems. We first design a personalized QoE preference predictor to characterize trade-offs between service completion time and resultant service fees of individual users. Subsequently, we develop a reinforcement learning method to decide static function embedding decisions at the offline phase. Considering the intermittency of renewable sources, we further provide an energy-adaptive function replica freezing strategy at the online phase. Evaluations demonstrate that our approach boosts the holistic user QoE by 32.2% over state-of-the-art algorithms.
Kun Cao 0001, Chaohong Tan, Yangguang Cui, Keqin Li 0001
IEEE Trans. Serv. Comput.1
2025 D3QN-based secure scheduling of microservice workflows in cloud environments
Saiqin Long, Chongxi Rao, Qingyong Deng, Kun Cao 0001
Comput. Networks5
2025 Latency-Aware Client Selection and Energy Management for Hierarchical Federated Learning
abstract
With the prosperity of deep learning (DL) in Internet of Things (IoT) fields, federated learning (FL), viewed as a critical element of numerous DL-aided IoT intelligent applications, enables cooperative DL training across decentralized clients without revealing their personal data. However, the computational capacity heterogeneity and limited energy resources of IoT devices cause a huge negative influence on FL training in IoT intelligence applications. To address the above issues, this paper proposes an excellent distributed training mechanism for hierarchical FL to reduce training latency and energy cost for achieving the desirable accuracy. Specifically, taking into account computational capacity heterogeneity of clients, we first design a latency-regularization-aware client selection algorithm to appropriately select participating clients in training epochs and control their participation frequencies for boosting distributed training efficiency. Subsequently, after obtaining the selected client subset in each hierarchical FL training epoch, by leveraging the variable transmission delays of clients in distributed training, we propose a mixed integer linear programming-based transmission power management strategy for participating clients to alleviate their energy consumption burden. Extensive numerical results demonstrate that our proposed mechanism can attain 576.93% training speedup and achieve 10.94% accuracy enhancement compared with the baseline General FL, and yield up to 56.91% energy cost savings compared with the baseline HierFAVG.
Jiamei Li, Kun Cao 0001, Yangguang Cui, Tong Liu 0001, Zhiquan Liu 0001
IEEE Internet Things J.3
2025 Reliability-Aware Personalized Deployment of Approximate Computation IoT Applications in Serverless Mobile Edge Computing
abstract
Over the past few years, the integration of mobile edge computing (MEC) and serverless computing, known as serverless MEC (SMEC), has garnered considerable attention. Despite abundant existing works on SMEC exploration, there remains an unaddressed gap in guaranteeing dependable application outputs due to ignoring the threat of both soft and bit errors on SMEC infrastructures. Furthermore, existing works fall short of accommodating the personalized requirements and approximate computation of Internet of Things (IoT) applications, thereby resulting in holistic quality-of-service (QoS) degradation of SMEC systems typically provisioned by limited edge resources. In this article, we investigate the reliability-aware personalized deployment of approximate computation IoT applications for QoS maximization in SMEC environments. To this end, we propose a hybrid methodology composed of offline and online optimization phases. At the offline phase, a decomposition-based function placement method is devised to accomplish function-to-server mapping by integrating convex optimization, cross-entropy method, and incremental control techniques. At the online phase, a lightweight reinforcement learning scheme based on proximal policy optimization (PPO) is developed to handle the inherent dynamicity of IoT applications. We also build a simulation platform upon the real-world base station distribution in Shanghai Telecom and the practical cluster trace in the Alibaba open program. Evaluations demonstrate that our hybrid approach boosts the holistic QoS by 63.9% compared with the state-of-the-art peer algorithms.
Kun Cao 0001, Mingsong Chen 0001, Stamatis Karnouskos, Shiyan Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Personalized Federated Learning With State-Adaptive IoT Device Scheduling in Mobile-Edge Computing
abstract
Federated learning (FL) is envisioned as a pioneering framework for the distributed training of artificial intelligence models in mobile edge computing (MEC) environments. Traditional MEC-empowered FL approaches commonly neglect the inherent competition for computation resources between uncertain user-own tasks and FL training activities on individual Internet-of-Things (IoT) devices. Meanwhile, these approaches fail to address the personalized reward perception that dominates the active participation of IoT devices in FL training. As a result, both the resource utilization on IoT devices and the overall performance of FL models are significantly degraded in practical MEC systems. To address these challenges, this paper investigates the personalized FL with state-adaptive IoT device scheduling in MEC scenarios. We first develop a collaborative device-state estimation method to effectively capture the uncertainty in future states of FL candidates. Subsequently, we design a user-personality inspired degree-of-satisfaction (DoS) prediction scheme to quantify the impact of computation resource competition on the satisfaction levels of personalized FL participants. Building on these efforts, we propose a state-adaptive IoT device scheduling technique to optimize the accuracy of FL models at the offline stage. An FL runtime management policy is also designed to deal with the timing failures of unsuccessful return of local training results at the online stage. Evaluations show that our approach enhances the accuracy of WideResNet FL model by up to 35.96% on CIFAR-10 and EuroSAT datasets. Our source code is available at https://github.com/superguymj/ACE.
Jun Mai, Kun Cao 0001, Tongquan Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Personalized Federated Learning for Green Industrial IoT
abstract
In recent years, federated learning (FL) has gained increasing attention in industrial Internet-of-Things (IIoT) domains due to its privacy-preserving advantages. However, prior works commonly adopt a one-size-fits-all strategy for FL computation resource management and reward allocation, disregarding the time-varying participant states across different FL training rounds. Consequently, these methods fail to ensure the sustainability and active participation of IIoT devices in realistic FL deployments. To bridge this gap, we propose a personalized FL methodology for green IIoT systems powered by renewable energy sources. We first establish an incentive model along with its preference parameter-solving scheme to accurately characterize the incentive preferences of individual FL participants. Subsequently, a personalized participant scheduling approach is developed to accommodate dynamic resource usage patterns and diverse incentive preferences among FL participants. Our technique integrates empirical insights into conventional proximal policy optimization methods to accelerate policy learning within reinforcement learning frameworks. Experimental results on an FL prototype system show that our methodology improves the FL model accuracy by 25.92% compared with representative baseline algorithms.
Kun Cao 0001, Yangguang Cui, Rui Xu 0013, Yuxia Sun, Zhiquan Liu 0001, Chaohong Tan
IEEE Trans. Ind. Informatics1
2024 CPU-GPU Cooperative QoS Optimization of Personalized Digital Healthcare Using Machine Learning and Swarm Intelligence
abstract
In recent decades, the rapid advances in information technology have promoted a widespread deployment of medical cyber-physical systems (MCPS), especially in the area of digital healthcare. In digital healthcare, medical edge devices empowered by CPU-GPU (Graphics Processing Unit) cooperative multiprocessor system-on-chips (MPSoCs) have a great potential in processing and managing the massive amounts of health-related data. However, most of the existing works on CPU-GPU cooperative MPSoCs cannot maintain a high-precision workload estimation since they simply leverage the worst-case execution cycles to pessimistically predict the workload of digital healthcare applications. Besides, they neglect the personalized requirements of individual healthcare applications and the lifetime reliability demands of heterogeneous CPU-GPU cores. As a result, the normal functions of medical edge devices and the quality-of-services (QoS) of digital healthcare applications are likely to suffer from underlying failures and degradation. In this paper, we explore CPU-GPU cooperative QoS optimization of personalized digital healthcare applications running on reliability guaranteed edge devices with the help of machine learning and swarm intelligence techniques. We first develop two novel predictors: one is a machine learning based predictor for application workload estimation, and the other is a feature-driven predictor for application QoS estimation. We then incorporate the two predictors into a swarm intelligent application scheduling scheme upon the cooperative dual-population evolutionary algorithm (c-DPEA) to find optimal application mapping and partitioning settings. Experimental results show that our solution not only augments the average QoS of whole digital healthcare applications by 15.7%, but also balances the QoS of individual digital healthcare applications by 64.3%.
Kun Cao 0001, Yangguang Cui, Liying Li 0002, Junlong Zhou, Shiyan Hu 0001
IEEE Trans. Comput. Biol. Bioinform.1
2024 Energy-Aware Incentive Mechanism for Hierarchical Federated Learning Using Water Filling Technique
abstract
Federated learning (FL) is an attractive industrial paradigm to accomplish distributed artificial intelligence (AI) training collaboratively in a data privacy-preserving manner. Most existing designs for FL systems assume that industrial user equipments (UEs) participate voluntarily in FL training. However, since both AI model training and transmission consume considerable energy, UEs are reluctant to participate without economic rewards. Hence, the lack of proper economic reward incentive mechanism results in low UE utility and frustrates UEs' enthusiasm for participating in training. To address the above challenge, in this article, we propose a two-phase energy-aware reward incentive mechanism for the edge-cloud-assisted hierarchical federated learning (HFL) system to optimize the overall UE utility, thereby, incentivizing UEs to participate more actively. Specifically, at the cloud server phase, we design an energy quantity-aware incentive mechanism for reasonably distributing rewards to its sub-edge-assisted FL systems. Subsequently, at the edge server phase, based on the quantitative analysis for the optimal reward allocation solution, we develop an energy-aware water filling-based reward incentive mechanism to adapt to individual needs of UEs and maximize the overall UE utility. Experiments verify that, compared to well-known benchmarks, our incentive mechanism can improve the overall UE utility by up to 55.94% and better incentivize UEs to participate in training.
Yangguang Cui, Weiqin Tong, Tong Liu 0001, Kun Cao 0001, Junlong Zhou, Ming Xu 0010, Tongquan Wei
IEEE Trans. Ind. Informatics4
2024 Reliability-Driven End-End-Edge Collaboration for Energy Minimization in Large-Scale Cyber-Physical Systems
abstract
In recent years, cyber-physical systems (CPS) have been widely deployed in industrial manufacturing fields and our daily living domains. End–end–edge collaboration, coupling mobile edge computing and device-to-device communication, is a promising computation paradigm to meet the stringent real-time demands of large-scale CPS applications. However, energy and reliability concerns should be carefully addressed in end–end–edge collaboration-empowered large-scale CPS due to the limited energy supply and inherent openness characteristic of end devices. In this article, we explore the reliability-driven energy optimization of end–end–edge collaborated large-scale CPS applications. We develop a reliability-driven end–end–edge collaboration approach to deal with the energy minimization problem. Our approach first designs a clustering method to quantify differentiated energy demands by analyzing the energy dissipation composition of heterogeneous applications. Afterward, our approach leverages incremental control and swarm intelligence-based techniques to obtain energy-efficient reliability-guaranteed task offloading solutions for differentiated application clusters. Experimental results reveal that our approach achieves 51.48% energy savings compared with peer algorithms.
Kun Cao 0001, Jian Weng 0001, Keqin Li 0001
IEEE Trans. Reliab.1
2024 REPFS: Reliability-Ensured Personalized Function Scheduling in Sustainable Serverless Edge Computing
abstract
In recent years, serverless edge computing has been widely employed in the deployments of Internet-of-things (IoT) applications. Despite considerable research efforts in this field, existing works fail to jointly consider essential factors such as energy, reliability, personalized user requirements, and stochastic application executions. This oversight results in an inefficient utilization of computation and communication resources within serverless edge computing networks, subsequently diminishing the profit of service providers and degrading the quality-of-experience (QoE) of end users. In this paper, we explore the problem of reliability-ensured personalized function scheduling (REPFS) to jointly optimize the profit of service providers and the holistic QoE of end users in sustainable serverless edge computing. A personality-driven user QoE prediction method is first designed to accurately estimate the QoE of individual end users with differentiated personality types. Afterward, a deterministic function scheduling policy is developed on the problem-specific augmented non-dominated sorting genetic algorithm II (PSA-NSGA-II). Given the inherent uncertainty of application executions, a stochastic function scheduling strategy that can be easily parallelized for modern multicore scheduler platforms is also devised to accelerate solution generation for stochastic applications. Experimental results show that our deterministic function scheduling policy achieves 15% performance enhancement compared with representative multiobjective evolutionary algorithms. Furthermore, our stochastic function scheduling strategy promotes the service profit by 78% and the holistic user QoE by 118% on average compared with the developed deterministic scheduling policy.
Kun Cao 0001, Jian Weng 0001
IEEE Trans. Sustain. Comput.1
2023 Optimizing Training Efficiency and Cost of Hierarchical Federated Learning in Heterogeneous Mobile-Edge Cloud Computing
abstract
Federated learning (FL), an emerging distributed machine learning (ML) technique, allows massive embedded devices and a server to work together for training a global ML model without collecting user data on a server. Most existing approaches adopt the traditional centralized FL paradigm with a single server: one is the cloud-centric FL paradigm and the other is the edge-centric FL paradigm. The cloud-centric FL paradigm is able to manage a large-scale FL system across massive user devices with high communication cost, whereas the edge-centric FL paradigm is capable of coordinating a small-scale FL system benefiting from the low communication delay over wireless networks. To fully exploit the advantages of both, in this article, we develop a distinctive hierarchical FL framework for the promising mobile-edge cloud computing (MECC) system, called HELCHFL, to achieve high-efficiency and low-cost hierarchical FL training. In particular, we formulate the corresponding theoretical foundation for our HELCHFL to ensure hierarchical training performance. Furthermore, to address the inherent communication and user heterogeneity issues of FL training, our HELCHFL develops a utility-driven and heterogeneity-aware heuristic user selection strategy to enhance training performance and reduce training delay. Subsequently, by analyzing and utilizing the slack time in FL training, our HELCHFL introduces a device operating frequency determination approach to reduce training energy cost. Experiments demonstrate that our HELCHFL can enhance the highest accuracy by up to 52.93%, gain the training speedup of up to 483.74%, and obtain up to 45.59% training energy savings compared to state-of-the-art baselines.
Yangguang Cui, Kun Cao 0001, Junlong Zhou, Tongquan Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Reinforcement Learning-Based Device Scheduling for Renewable Energy-Powered Federated Learning
abstract
Due to its unique privacy protection advantages, emerging federated learning (FL) is regarded as a significant technique to enable Industry 4.0. However, the industrial deployment of FL encounters the primary obstacles of limited device energy and system communication resources. Nowadays, renewable energy-powered devices have been deployed in various industrial fields to tackle the challenges of unsustainable and limited energy of battery-powered devices. Inspired by this, this article proposes a novel FL protocol to groundbreakingly improve the performance of renewable energy-powered FL systems. Specifically, with the underlying theory of FL as the guide, the proposed protocol features a reinforcement learning-based device scheduling solution to adapt to intermittent renewable energy supply. Following this device scheduling solution, an integer linear programming-based bandwidth management scheme is introduced to optimize communication efficiency. Experimental results on two representative data distribution situations demonstrate that compared with the state-of-the-art schemes, our FL protocol can boost up to 36.63% and 50.99% accuracy, respectively.
Yangguang Cui, Kun Cao 0001, Tongquan Wei
IEEE Trans. Ind. Informatics2
2023 Game Theoretical Task Offloading for Profit Maximization in Mobile Edge Computing
abstract
In this paper, a novel task offloading architecture called Flex-MEC is proposed, which achieves efficient task allocation and scheduling (TAS) between MEC servers. By adding metadata before task data, we redesign the offloading process in Flex-MEC, the TAS planning can be conducted without finishing the task data receiving. Once planning is done the task data can be directly forwarded to the allocated server and executed. This reduces latency compared to the traditional way of transmitting, planning, forwarding and executing sequentially. For TAS planning, a multi-server multi-task allocation and scheduling (MMAS) problem is formulated to maximize the MEC system profit. The MMAS problem is proven as an NP-complete problem, thus is challenging to solve. Then, a distributed scheme and a centralized scheme are proposed to solve the MMAS problem with low complexity. In the distributed scheme, the MMAS problem is converted into a non-cooperative game and the existence of Nash Equilibrium (NE) is proven and a low complexity response update algorithm is proposed to converge to NE. And the centralized scheme is based on a greedy idea and runs on a MEC controller in a centralized way. Verified by experiments, these two schemes can achieve better performance than compared schemes.
Haojun Teng, Zhetao Li, Kun Cao 0001, Saiqin Long, Song Guo 0001, Anfeng Liu
IEEE Trans. Mob. Comput.3
2022 HELCFL: High-Efficiency and Low-Cost Federated Learning in Heterogeneous Mobile-Edge Computing
abstract
Federated Learning (FL), an emerging distributed machine learning (ML), empowers a large number of embedded devices (e.g., phones and cameras) and a server to jointly train a global ML model without centralizing user private data on a server. However, when deploying FL in a mobile-edge computing (MEC) system, restricted communication resources of the MEC system, heterogeneity and constrained energy of user devices have a severe impact on FL training efficiency. To address these issues, in this article, we design a distinctive FL framework, called HELCFL, to achieve high-efficiency and low-cost FL training. Specifically, by analyzing the theoretical foundation of FL, our HELCFL first develops a utility-driven and greedy-decay user selection strategy to enhance FL performance and reduce training delay. Subsequently, by analyzing and utilizing the slack time in FL training, our HELCFL introduces a device operating frequency determination approach to reduce training energy costs. Experiments verify that our HELCFL can enhance the highest accuracy by up to 43.45 %, realize the training speedup of up to 275.03%, and save up to 58.25% training energy costs compared to state-of-the-art baselines.
Yangguang Cui, Kun Cao 0001, Junlong Zhou, Tongquan Wei
DATE2
2022 Edge Intelligent Joint Optimization for Lifetime and Latency in Large-Scale Cyber-Physical Systems
abstract
In recent years, the exploration on large-scale cyber–physical systems (CPSs) has become a fertile research field of significant impact. Large-scale CPS applications cover not only manufacturing and production areas but also daily living domains. Traditional solutions dedicated for large-scale CPSs mainly concentrate on the service latency or reliability optimization, but neglect the resultant negative impact on system lifetime. In this article, we conduct the first study on jointly optimizing the service latency and system lifetime subject to the constraints of reliability, energy consumption, and schedulability for large-scale CPSs. We propose an edge intelligent solution composed of offline and online phases. At the offline phase, the long short-term memory (LSTM) technique is leveraged to predict task offloading rates at individual user groups. Afterward, the multiobjective evolutionary algorithm with dual local search (DLS-MOEA) is exploited to determine optimal system static settings of computation offloading mapping and task replication number. At the online phase, an affinity-driven scheme incurring minimal system dynamic overheads is designed to deal with the inherent mobility of terminal users. We also build an algorithm validation platform upon which extensive simulation experiments are carried out. Experimental results show that our offline and online schemes outperform the state-of-the-art benchmarking methods by 27.1% and 43.5%, respectively.
Kun Cao 0001, Yangguang Cui, Zhiquan Liu 0001, Wuzheng Tan, Jian Weng 0001
IEEE Internet Things J.1
2022 Client Scheduling and Resource Management for Efficient Training in Heterogeneous IoT-Edge Federated Learning
abstract
Federated learning (FL) offers a promising paradigm that empowers numerous Internet of Things (IoT) devices to implement distributed learning on the premise of ensuring user privacy and data security. However, since FL adopts a synchronous distributed training mode, the heterogeneity of participating IoT devices and limited communication resources make FL encounter serious issues of low training efficiency in actual deployment. In this article, we propose an excellent FL policy for the heterogeneous IoT-edge FL system to improve distributed training efficiency. Specifically, first, by borrowing the idea of clustering, we explore an iterative self-organizing data analysis techniques algorithm (ISODATA)-based heterogeneous-aware client scheduling strategy to alleviate the issue of low training efficiency incurred by the heterogeneity of clients. Subsequently, to tackle the challenge of limited communication resources in FL, we first analyze the characteristics of the optimal resource block allocation solution theoretically and then introduce a mixed-integer linear programming (MILP)-based strategy to judiciously allocate resource blocks for scheduled clients. Comprehensive experimental results demonstrate that, compared with benchmarking strategies, our proposed FL policy can achieve up to 55.22% accuracy improvement in a relaxed time scenario, and attain up to$3.62\times $acceleration for reaching the specific expected accuracy.
Yangguang Cui, Kun Cao 0001, Guitao Cao, Meikang Qiu, Tongquan Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Throughput-Conscious Energy Allocation and Reliability-Aware Task Assignment for Renewable Powered In-Situ Server Systems
abstract
In-situ(InS) server systems are typically deployed in special environments to handleInSworkloads which are generated from environmentally sensitive areas or remote places lacking modern power supply infrastructure. This special operating environment ofInSservers urges such systems to be powered by renewable energy. In addition, theInSsystems are vulnerable to soft errors due to the harsh environments they deploy. This article tackles the problem of allocating harvested energy to renewable powered servers and assigning theInSworkloads to these servers for optimizing throughput of both the overall system and individual servers under energy and reliability constraints. We perform the energy allocation based on system state. In particular, for systems in low energy state, we propose a game theoretic approach that models the energy allocation as a cooperative game among multiple servers and derives a Nash bargaining solution. To meet the reliability constraint, we analyze the reliability optimality of assigning tasks to multiple servers and design a reliability-aware task assignment heuristic based on the analysis. Experimental results show that with a small time overhead, the proposed energy allocation approach achieves a high throughput from perspectives of both the overall system and individual servers, and the proposed task assignment approach ensures an increased system reliability.
Junlong Zhou, Kun Cao 0001, Xiumin Zhou, Mingsong Chen 0001, Tongquan Wei, Shiyan Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Power-Efficient Layer Mapping for CNNs on Integrated CPU and GPU Platforms: A Case Study
abstract
Heterogeneous MPSoCs consisting of integrated CPUs and GPUs are suitable platforms for embedded applications running on handheld devices such as smart phones. As the handheld devices are mostly powered by battery, the integrated CPU and GPU MPSoC is usually designed with an emphasis on low-power rather than performance. In this paper, we are interested in exploring a power-efficient layer mapping of convolution neural networks (CNNs) deployed on integrated CPU and GPU platforms. Specifically, we investigate the impact of layer mapping of YoloV3-Tiny (i.e., a widely-used CNN in both industry and academia) on system power consumption through numerous experiments on NVIDIA board Jetson TX2. The experimental results indicate that 1) almost all of the convolution layers are not suitable for mapping to CPU, 2) the pooling layer can be mapped to CPU for reducing power consumption, but the mapping may lead to a decrease in inference speed when the layer's output tensor size is large, 3) the detection layer can be mapped to CPU as long as its floating-point operation scale is not too large, and 4) the channel and upsampling layers are both suitable for mapping to CPU. These observations obtained in this study can be further utilized to guide the design of power-efficient layer mapping strategies for integrated CPU and GPU platforms.
Tian Wang 0001, Kun Cao 0001, Junlong Zhou, Gongxuan Zhang, Xiji Wang
ASP-DAC2
2021 BTMPP: Balancing Trust Management and Privacy Preservation for Emergency Message Dissemination in Vehicular Networks
abstract
As a potential application field of the sixth-generation (6G) communication technology and a promising part of massive Internet of Things (IoT), vehicular networks have attracted considerable attention from both academia and industry in recent years, where the cooperative safety applications are a significant branch. It is widely acknowledged that 6G is able to provide high-throughput and low-latency wireless communication capability for vehicular networks, support massive interconnectivity in vehicular networks with diverse service requirements, and significantly improve the performance of vehicular networks. Both trust management and privacy preservation play significant roles in vehicular networks, and there exists a tradeoff between them. For providing a satisfactory solution to balance the trust management and privacy preservation in vehicular networks, we put forward a novel scheme named as BTMPP (which is able to provide both the precise trust management and strong conditional privacy preservation simultaneously) in this article by leveraging the famous bloom filter (BF)-based private set intersection (PSI) technology. Furthermore, the theoretical analysis for the correctness, strong conditional privacy preservation capability, strong robustness, precise trust management, and the other features is detailed, and a series of simulations are conducted. The results reveal that the proposed scheme significantly outperforms the existing schemes in several aspects.
Zhiquan Liu 0001, Feiran Huang, Jian Weng 0001, Kun Cao 0001, Yinbin Miao, Yongdong Wu
IEEE Internet Things J.4
2021 Exploring reliable edge-cloud computing for service latency optimization in sustainable cyber-physical systems
abstract
Abstract In recent years, the advance in information technology has promoted a wide span of emerging cyber‐physical systems (CPS) applications such as autonomous automobile systems, healthcare monitoring, and process control systems. For these CPS applications, service latency management is extraordinarily important for the sake of providing high quality‐of‐experience to terminal users. Edge‐cloud computing, integrating both edge computing and cloud computing, is regarded as a promising computation paradigm to achieve low service latency for terminal users in CPS. However, existing latency‐aware edge‐cloud computing methods dedicated for CPS fail to jointly consider energy budgets and reliability requirements, which may greatly degrade the sustainability of CPS applications. In this article, we explore the problem of minimizing service latency of edge‐cloud computing coupled CPS under the constraints of energy budgets and reliability requirements. We propose a two‐stage approach composed of static and dynamic service latency optimization. At static stage, Monte‐Carlo simulation with integer‐linear‐programming technique is adopted to find the optimal computation offloading mapping and task backup number. At dynamic stage, a backup‐adaptive dynamic mechanism is developed to avoid redundant data transmissions and executions for achieving additional energy savings and service latency enhancement. Experimental results show that our solution is able to reduce system service latency by up to 18.3% compared with representative baseline solutions.
Kun Cao 0001, Tongquan Wei, Mingsong Chen 0001, Keqin Li 0001, Jian Weng 0001, Wuzheng Tan
Softw. Pract. Exp.1
2021 A Survey on Edge and Edge-Cloud Computing Assisted Cyber-Physical Systems
abstract
In recent years, the investigations on cyber-physical systems (CPS) have become increasingly popular in both academia and industry. A primary obstruction against the booming deployment of CPS applications lies in how to process and manage large amounts of generated data for decision making. To tackle this predicament, researchers advocate the idea of coupling edge computing, or edge-cloud computing into the design of CPS. However, this coupling process raises a diversity of challenges to the quality-of-services (QoS) of CPS applications. In this article, we present a survey on edge computing or edge-cloud computing assisted CPS designs from the QoS optimization perspective. We first discuss critical challenges in service latency, energy consumption, security, privacy, and reliability during the integration of CPS with edge computing or edge-cloud computing. Afterwards, we give an overview on the state-of-the-art works tackling different challenges for QoS optimization, and present a systematic classification during outlining literature for highlighting their similarities and differences. We finally summarize the experiences learned from surveyed works and envision future research directions on edge computing or edge-cloud computing assisted CPS optimization.
Kun Cao 0001, Shiyan Hu 0001, Yang Shi 0001, Armando W. Colombo, Stamatis Karnouskos, Xin Li 0001
IEEE Trans. Ind. Informatics1
2021 Exploring Placement of Heterogeneous Edge Servers for Response Time Minimization in Mobile Edge-Cloud Computing
abstract
In the past few years, the study on placing edge servers for response time optimization in mobile edge-cloud computing systems has become increasingly popular. Most of the existing schemes neglect two important aspects: one is the heterogeneity of edge/cloud servers and the other is the response time fairness of base stations, which may significantly degrade the system quality of services to mobile users. In this article, we conduct the study of deploying heterogeneous edge servers to optimize the expected response time of both the whole and individual base stations. We propose an approach consisting of offline and online stages. At the offline stage, the optimal placement strategy of heterogeneous edge servers is produced by using an integer linear programming technique. At the online stage, a mobility-aware game-theory-based method is developed to deal with the dynamic characteristic of user movement. Experimental results reveal that compared to benchmarking methods, our approach not only reduces system-expected response time by 47.37%, but also improves response time fairness of base stations by 71.60%.
Kun Cao 0001, Liying Li 0002, Yangguang Cui, Tongquan Wei, Shiyan Hu 0001
IEEE Trans. Ind. Informatics1
2020 Exploring Renewable-Adaptive Computation Offloading for Hierarchical QoS Optimization in Fog Computing
abstract
Fog computing is an emerging architectural paradigm for the implementation of the Internet of Things, where computation moves from cloud servers to network edges. Fog computing systems are with three characteristics: 1) low latency; 2) strong presence of real-time applications; and 3) reusability of end devices. Most existing designs of fog computing systems concentrate on reducing application processing latency, but neglect real-time requirements of applications and reusability of end devices, which may drastically degrade both functionality and quality-of-service (QoS) of applications. In this article, we investigate QoS optimization of real-time applications in fog computing systems equipped with reusable end devices and powered by hybrid energy of renewable generations and grid electricity. We propose a renewable-adaptive computation offloading approach. At the end device layer, local energy allocation schemes are designed at the application-level and component-level, where techniques of the cooperative game and mixed-integer linear programming (MILP) are leveraged, respectively. At the fog layer, the local energy allocation method is augmented to a local-remote scheduling solution by judiciously judging whether or not the computation offloading of an application needs to be triggered. The experimental results demonstrate that compared to benchmarking algorithms, our approach improves the overall and individual application QoS by up to 101.93% and 59.30%, respectively.
Kun Cao 0001, Junlong Zhou, Guo Xu, Tongquan Wei, Shiyan Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Augmented Cross-Entropy-Based Joint Temperature Optimization of Real-Time 3-D MPSoC Systems
abstract
3-D multiprocessor system-on-chip (MPSoC) systems can offer higher integration density, lower interaction cost, better bandwidth, and greater performance. However, vertically stacked silicon layers and limited heat dissipation paths result in high peak temperature and large temperature variation, which incur reliability reduction, lifetime decay, and performance degradation. In this article, we propose an offline augmented cross-entropy (CE)-based task scheduling strategy to jointly optimize peak temperature and temperature variation under the constraint of timeliness. Specifically, based on the conventional CE method, a heuristic iterative sampling method is designed to explore task-to-core assignment for balanced heat distribution between the top-layer and the bottom-layer cores. Subsequently, thermal characteristics of 3-D MPSoC systems are used to judiciously swap tasks between the two layers to improve the conventional CE-based task assignment and accelerate the iterative process. The peak temperature of individual cores is further reduced via sequencing, splitting, and slacking task execution. The experimental results demonstrate that compared to the existing state-of-the-art methods, the proposed scheme can reduce peak temperature by up to 8.02 °C and temperature variation by up to 24.78% without violating the timeliness of tasks.
Yangguang Cui, Kun Cao 0001, Liying Li 0002, Junlong Zhou, Tongquan Wei, Shiyan Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Queueing Theoretic Approach for Performance-Aware Modeling of Sustainable SDN Control Planes
abstract
Software Defined Networking (SDN) provides flexibility and programmability for network management by using a layered structure composed of data plane, control plane, and application plane. A key enabling technique for the sustainability of SDN-based network infrastructure is the modeling of power consumed by SDN control planes. However, power modeling of control planes is not extensively investigated yet, and no generic methods have been developed for performance and power comparison of sustainable SDN control planes. In this paper, we propose analytical performance and power models for different network controllers by using queuing theory, and design a generic framework for performance and power evaluation of different sustainable SDN control planes. Extensive simulation results show that the proposed solution can precisely model the power and performance of the concerned SDN control planes such that different control planes can be benchmarked under a general framework, which enables the identification of suitable control planes for various SDN network applications.
Xinli Huang, Fanshuo Li, Kun Cao 0001, Peijin Cong, Tongquan Wei, Shiyan Hu 0001
IEEE Trans. Sustain. Comput.3
2019 Lifetime-aware real-time task scheduling on fault-tolerant mixed-criticality embedded systems
Kun Cao 0001, Guo Xu, Junlong Zhou, Mingsong Chen 0001, Tongquan Wei, Keqin Li 0001
Future Gener. Comput. Syst.1
2019 A survey of optimization techniques for thermal-aware 3D processors
Kun Cao 0001, Junlong Zhou, Tongquan Wei, Mingsong Chen 0001, Shiyan Hu 0001, Keqin Li 0001
J. Syst. Archit.1
2019 QoS-Adaptive Approximate Real-Time Computation for Mobility-Aware IoT Lifetime Optimization
abstract
In recent years, the Internet of Things (IoT) has promoted many battery-powered emerging applications, such as smart home, environmental monitoring, and human healthcare monitoring, where energy management is of particular importance. Meanwhile, there is an accelerated tendency toward mobility of IoT devices, either being transported by humans or being mobile by itself. Existing energy management mechanisms for battery-powered IoT fail to consider the two significant characteristics of IoT: 1) the approximate real-time computation and 2) the mobility of IoT devices, resulting in unnecessary energy waste and network lifetime decay. In this paper, we explore mobility-aware network lifetime maximization for battery-powered IoT applications that perform approximate real-time computation under the quality-of-service (QoS) constraint. The proposed scheme is composed of offline and online stages. At offline stage, an optimal mobility-aware task schedule that maximizes network lifetime is derived by using mixed-integer linear programming technique. Redundant executions due to mobility-incurred overlapping of a single task on different IoT devices are avoided for energy savings. At online stage, a performance-guaranteed and time-efficient QoS-adaptive heuristic based on cross-entropy method is developed to adapt task execution to the fluctuating QoS requirements. Extensive simulations based on synthetic applications and real-life benchmarks have been implemented to validate the effectiveness of our proposed scheme. Experimental results demonstrate that the proposed technique can achieve up to 169.52% network lifetime improvement compared to benchmarking solutions.
Kun Cao 0001, Guo Xu, Junlong Zhou, Tongquan Wei, Mingsong Chen 0001, Shiyan Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Affinity-Driven Modeling and Scheduling for Makespan Optimization in Heterogeneous Multiprocessor Systems
abstract
With the advent of heterogeneous multiprocessor architectures, efficient scheduling for high performance has been of significant importance. However, joint considerations of reliability, temperature, and stochastic characteristics of precedence-constrained tasks for performance optimization make task scheduling particularly challenging. In this paper, we tackle this challenge by using an affinity (i.e., probability)-driven task allocation and scheduling approach that decouples schedule lengths and thermal profiles of processors. Specifically, we separately model the affinity of a task for processors with respect to schedule lengths and the affinity of a task for processors with regard to chip thermal profiles considering task reliability and stochastic characteristics of task execution time and intertask communication time. Subsequently, we combine the two types of affinities, and design a scheduling heuristic that assigns a task to the processor with the highest joint affinity. Extensive simulations based on randomly generated stochastic and real-world applications are performed to validate the effectiveness of the proposed approach. Experiment results show that the proposed scheme can reduce the system makespan by up to 30.1% without violating the temperature and reliability constraints compared to benchmarking methods.
Kun Cao 0001, Junlong Zhou, Peijin Cong, Liying Li 0002, Tongquan Wei, Mingsong Chen 0001, Shiyan Hu 0001, Xiaobo Sharon Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Game Theoretic Feedback Control for Reliability Enhancement of EtherCAT-Based Networked Systems
abstract
EtherCAT has become one of the leading real-time Ethernet solutions for networked industrial systems, where a reliable communication infrastructure is needed due to highly error-prone environments. However, existing work on EtherCAT mainly focuses on clock synchronization and timeliness improvement. The reliability of EtherCAT-based networked systems has largely been ignored. In this paper, we present a proportional integral derivative (PID)-based feedback control scheme that aims at enhancing reliability of networked systems under timing and system resource constraints. Instead of retransmitting data upon error detection, we use forward error control technique based on inequality of arithmetic and geometric means to achieve the required system reliability at a low deadline miss rate of messages. We further optimize the forward error control technique and design a fast and fair error resilient mechanism by using a cooperative game. In addition to reliability enhancement, our PID-based error control scheme can also improve the stability of a system in terms of deadline miss rate in the presence of burst errors. Simulation results show that the proposed scheme can achieve reliability enhancement of up to 91% compared to benchmarking methods.
Liying Li 0002, Peijin Cong, Kun Cao 0001, Junlong Zhou, Tongquan Wei, Mingsong Chen 0001, Shiyan Hu 0001, Xiaobo Sharon Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Feedback control of real-time EtherCAT networks for reliability enhancement in CPS
abstract
EtherCAT has become one of the leading real-time Ethernet solutions for networked industrial systems where a reliable communication infrastructure is needed due to highly error-prone environments. However, existing work on EtherCAT mainly focuses on clock synchronization and timeliness improvement. The reliability of EtherCAT-based networked systems has largely been ignored. In this paper, we present a PID-based feedback control scheme that aims at enhancing reliability of networked systems under timing and system resource constraints. Instead of automatic repeat request method (ARQ), a forward error control technique is introduced to achieve the required system reliability at a lower deadline miss rate of messages. The PID-based feedback control scheme can also improve the stability of a system in terms of deadline miss rate in the presence of bursty errors. Simulation results show that the proposed scheme can achieve reliability enhancement of up to 79% compared to benchmarking methods.
Liying Li 0002, Peijin Cong, Kun Cao 0001, Junlong Zhou, Tongquan Wei, Mingsong Chen 0001, Xiaobo Sharon Hu
DATE3
2018 Thermal-aware correlated two-level scheduling of real-time tasks with reduced processor energy on heterogeneous MPSoCs
Junlong Zhou, Jianming Yan, Kun Cao 0001, Yanchao Tan, Tongquan Wei, Mingsong Chen 0001, Gongxuan Zhang, Xiaodao Chen, Shiyan Hu 0001
J. Syst. Archit.3
2018 Cost-Constrained QoS Optimization for Approximate Computation Real-Time Tasks in Heterogeneous MPSoCs
abstract
Internet of Things devices, such as video-based detectors or road side units are being deployed in emerging applications like sustainable and intelligent transportation systems. Oftentimes, stringent operation and energy cost constraints are exerted on this type of applications, necessitating a hybrid supply of renewable and grid energy. The key issue of a cost-constrained hybrid of renewable and grid power is its uncertainty in energy availability. The characteristic of approximate computation that accepts an approximate result when energy is limited and executes more computations yielding better results if more energy is available, can be exploited to intelligently handle the uncertainty. In this paper, we first propose an energy-adaptive task allocation scheme that optimally assigns real-time approximate-computation tasks to individual processors and subsequently enables a matching of the cost-constrained hybrid supply of energy with the energy demand of the resultant task schedule. We then present a quality of service (QoS)-driven task scheduling scheme that determines the optional execution cycles of tasks on individual processors for optimization of system QoS. A dynamic task scheduling scheme is also designed to adapt at runtime the task execution to the varying amount of the available energy. Simulation results show that our schemes can reduce system energy consumption by up to 29% and improve system QoS by up to 108% as compared to benchmarking algorithms.
Tongquan Wei, Junlong Zhou, Kun Cao 0001, Peijin Cong, Mingsong Chen 0001, Gongxuan Zhang, Xiaobo Sharon Hu, Jianming Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Developing User Perceived Value Based Pricing Models for Cloud Markets
abstract
With the rapid deployment of cloud computing infrastructures, understanding the economics of cloud computing has become a pressing issue for cloud service providers. However, existing pricing models rarely consider the dynamic interactions between user requests and the cloud service provider. Thus, the law of supply and demand in marketing is not fully explored in these pricing models. In this paper, we propose a dynamic pricing model based on the concept of user perceived value that accurately captures the real supply and demand relationship in the cloud service market. Subsequently, a profit maximization scheme is designed based on the dynamic pricing model that optimizes profit of the cloud service provider without violating service-level agreement. Finally, a dynamic closed loop control scheme is developed to adjust the cloud service price and multiserver configurations according to the dynamics of the cloud computing environment such as fluctuating electricity and rental fees. Extensive simulations using the data extracted from real-world applications validate the effectiveness of the proposed user perceived value-based pricing model and the dynamic profit maximization scheme. Our algorithm can achieve up to 31.32 percent profit improvement compared to a state-of-the-art approach.
Peijin Cong, Liying Li 0002, Junlong Zhou, Kun Cao 0001, Tongquan Wei, Mingsong Chen 0001, Shiyan Hu 0001
IEEE Trans. Parallel Distributed Syst.4
2017 Reliability and temperature constrained task scheduling for makespan minimization on heterogeneous multi-core platforms
Junlong Zhou, Kun Cao 0001, Peijin Cong, Tongquan Wei, Mingsong Chen 0001, Gongxuan Zhang, Jianming Yan, Yue Ma 0001
J. Syst. Softw.2
2016 Game Theoretic Energy Allocation for Renewable Powered In-Situ Server Systems
abstract
In-situ server systems are deployed in very special operating environment to handle in-situ workloads that are normally generated from environmentally sensitive areas or remote places that lack established utility infrastructure. This very special operating environment of in-situ servers urges such systems to be 100 percent powered by renewable energy. However, existing energy management schemes assume a hybrid supply of grid and renewable energy, hence are not well suited for 100 percent renewable powered in-situ server systems. In this paper, we tackle the problem of allocating harvested energy to 100 percent renewable powered server systems for optimizing both the overall system throughput and throughput of individual servers. From a game theoretic perspective, we model the energy allocation problem as a cooperative game among multiple servers and derive a Nash bargaining solution. Based on the Nash bargaining solution, we then propose a heuristic algorithm that determines the energy allocation strategies according to system energy states. Experimental results show that our proposed game theoretic approach achieves a high throughput from perspectives of both the overall system and individual servers.
Junlong Zhou, Kun Cao 0001, Tongquan Wei, Mingsong Chen 0001
ICPADS3