Wentai Wu

dblp:181/3887 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0001-5851-327XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 6 · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Experiential Fairness: Bridging the Gap Between User Experience and Resource-Centric Fairness in Online LLM Services
abstract
Conventional fairness in multi-tenant Large Language Model (LLM) inference services is typically defined by system-centric metrics such as equitable resource allocation. We argue that this is unilateral and it creates a gap between measured system performance and actual user-perceived quality. We challenge this notion by introducing and formalizing Experiential Fairness, a user-centric paradigm that shifts the objective from equality of opportunity (resource access) to equity of outcome (user experience). With this motivation we propose ExFairS, a lightweight scheduling framework that perceives each user's satisfaction as a composite measure of Service Level Objective (SLO) compliance and resource consumption, and dynamically re-orders the serving queue guided by a credit-based priority mechanism. Extensive experiments on an 8-GPU NVIDIA V100 node show that ExFairS reduces the SLO violation rate by up to 100% and improves system throughput by 14-21.9%, outperforming state-of-the-art schedulers and delivering a demonstrably higher degree of Experiential Fairness.
Jiahua Huang, Wentai Wu, Yongheng Liu, Guozhi Liu, Yang Wang 0006, Weiwei Lin 0001
AAAI2
2026 HaS: Accelerating RAG Through Homology-Aware Speculative Retrieval
abstract
Retrieval-Augmented Generation (RAG) expands the knowledge boundary of large language models (LLMs) at inference by retrieving external documents as context. However, retrieval becomes increasingly time-consuming as the knowledge databases grow in size. Existing acceleration strategies either compromise accuracy through approximate retrieval, or achieve marginal gains by reusing results of strictly identical queries. We propose HaS, a homology-aware speculative retrieval framework that performs low-latency speculative retrieval over restricted scopes to obtain candidate documents, followed by validating whether they contain the required knowledge. The validation, grounded in the homology relation between queries, is formulated as a homologous query re-identification task: once a previously observed query is identified as a homologous re-encounter of the incoming query, the draft is deemed acceptable, allowing the system to bypass slow full-database retrieval. Benefiting from the prevalence of homologous queries under real-world popularity patterns, HaS achieves substantial efficiency gains. Extensive experiments demonstrate that HaS reduces retrieval latency by 23.74% and 36.99% across datasets with only a 1-2% marginal accuracy drop. As a plug-and-play solution, HaS also significantly accelerates complex multi-hop queries in modern agentic RAG pipelines. Source code is available at: https://github.com/ErrEqualsNil/HaS.
Wentai Wu, Yongheng Liu
ICDE3
2026 SegRNN: Segment Recurrent Neural Network for Long-Term Time-Series Forecasting
abstract
With the proliferation of Internet of Things (IoT) applications, advanced time series forecasting techniques have become increasingly critical for managing and responding to complex temporal dynamics. However, traditional RNN-based methods have faced challenges in the Long-term Time Series Forecasting (LTSF) domain when dealing with excessively long look-back windows and forecast horizons. Consequently, the dominance in this domain has shifted towards Transformer, MLP, and CNN approaches. The substantial number of recurrent iterations are the fundamental reasons behind the limitations of RNNs in LTSF. To address these issues, we propose two novel strategies to reduce the number of iterations in RNNs for LTSF tasks: Segment-wise Iterations and Parallel Multi-step Forecasting (PMF). RNNs that combine these strategies, called SegRNN, significantly reduce the required recurrent iterations for LTSF, resulting in notable improvements in forecast accuracy and inference speed. Extensive experiments demonstrate that SegRNN not only outperforms state-of-the-art Transformer-based models but also reduces runtime and memory usage by more than 78%, making it highly suitable for resource-constrained IoT scenarios. These achievements provide strong evidence that RNNs continue to excel in LTSF tasks and encourage further exploration of this domain with more RNN-based approaches. The code is available at: https://github.com/lss-1138/SegRNN.
Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Feiyu Zhao, Ruichao Mo, Haotong Zhang 0003
IEEE Internet Things J.3
2026 Federated Learning for Edge Computing Enabled Artificial Intelligence of Things: A comprehensive survey
abstract
Among contemporary AI computing paradigms, Federated Learning (FL) stands out as an innovative method and has shown great potential in conjunction with edge computing. The two techniques combined serve as a building block forthe development of the Artificial Intelligence of Things (AIoT). This paper sheds light on the synergistic integration of FL with edge computing to propel AIoT’s capabilities in decentralized environments. By executing computing tasks closer to the data, FL at the edge not only alleviates latency and bandwidth limitations inherent in cloud-centric architectures, but also presents a robust solution to privacy concerns—a crucial obstacle in traditional centralized training setups. This paper delves into how FL tackles these privacy issues, providing an intricate explanation of its operational principles, applications, and the resultant benefits for AIoT systems. Through this scrutiny, we highlight FL’s potential in bolstering the efficiency and privacy of AIoT deployments while also delineating future research directions and the expected impact across various domains. This study aims to comprehensively comprehend FL for Edge Computing-enabled AIoT and foster developments in intelligent technologies and applications in an interconnected world.
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Wentai Wu, Chen Wang 0011, Ahmed M. Abdelmoniem
Knowl. Based Syst.4
2026 SparseTSF: Lightweight and Robust Time Series Forecasting via Sparse Modeling
abstract
This paper introduces SparseTSF, a novel and extremely lightweight method for Long-term Time Series Forecasting (LTSF), designed to address the challenges of modeling complex temporal dependencies over extended horizons with minimal computational resources. At the heart of SparseTSF lies the Cross-Period Sparse Forecasting technique, which simplifies the forecasting task by downsampling the original sequences to focus on cross-period trend prediction. This technique not only significantly reduces model complexity and the number of parameters but also serves as an implicit regularization mechanism that enhances the model's robustness, achieving an optimal balance between performance and efficiency. Based on this technique, SparseTSF uses fewer than 1,000 parameters to achieve competitive performance compared to state-of-the-art methods, with evident advantages under longer look-back windows (e.g., 720) that allow the model to better exploit inherent periodicity and trend information. Furthermore, SparseTSF showcases remarkable generalization capabilities, making it well-suited for scenarios with limited computational resources, small samples, or low-quality data.
Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Haojun Chen, C. L. Philip Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 ComHA: Cloud-Edge-Device Cooperative Model Building Based on Hierarchical Automated Machine Learning
abstract
Cloud-edge-device (CED) cooperative computing is an emerging paradigm that extends the reach of cloud services, providing higher flexibility and scalability to modern AI-driven computing services. However, traditional “one-size-fits-all” AI model construction at the edge struggles to accommodate the strong heterogeneity of target devices. This leads to an increasing demand for specialized models tailored for local resources in AI applications. To this end, we introduce ComHA, a cooperative model building framework based on hierarchical Automated Machine Learning (AutoML) to bridge the gap between hyperparameter optimization on the cloud and local model customization at the edge. In this framework, the cloud performs high-level AutoML to reduce the search space of learning algorithms, model architectures, and relevant hyperparameters to a specific set based on target device specifications. Subsequently, edge devices execute low-level AutoML to identify and train the optimal model, customized for their local data and resources. This approach aims to strike a balance between the benefit and cost of customized model building. Through extensive experiments conducted on a real-world testbed with public datasets, our results demonstrate that ComHA outperforms traditional methods in producing tailored models of high accuracy and low inference latency in various environments.
Weiwei Lin 0001, Wangbo Shen, Wentai Wu, Keqin Li 0001
IEEE Trans. Computers3
2026 Client Selection in Federated Learning With Differential Privacy-Based Data Stream
abstract
Federated learning (FL) is a distributed machine learning (ML) paradigm designed for numerous networked devices. To face the massive data generated by devices and privacy concerns in model construction, the paradigm can execute ML tasks with differential privacy (DP) over private data streams. In each FL training iteration, a few clients are selected to participate and consume privacy budgets that determine the level of privacy protection. The client selection strategy plays a pivotal role in the final model performance. At present, reconciling model performance with privacy protection remains an open issue in online settings. Specifically, the DP model designed for data streams reduces the benefits of frequent participation by clients with high-quality data, as the DP model constrains the available privacy budget of continuous iterations. To address this issue, we propose a novel client selection framework for FL with DP-based data streams. At a macro level, we leverage fuzzy control to dynamically adjust the participation rate of clients, ensuring sufficient privacy budgets are allocated in each training iteration. At a fine-grained level, we design a dynamic scoring function based on the characteristics of the clients and a budget-aware client selection method to select clients further. Additionally, considering the diverse scenarios in data collection, we propose a relaxed semi-online setting and integrate reinforcement learning (RL) to enhance the framework performance. Extensive experiments demonstrate the remarkable advantages of our framework in accuracy, convergence rate, robustness, etc.
Wentai Wu, Young-June Choi, Hiroo Sekiya, Zhetao Li
IEEE Trans. Netw.3
2025 Towards imbalanced regression over distributionally biased data: A fast static approach
Wentai Wu, Ligang He, Weiwei Lin 0001, Jinyi Long, Zhiquan Liu 0001, C. L. Philip Chen
Inf. Softw. Technol.1
2025 Cacomp: A Cloud-Assisted Collaborative Deep Learning Compiler Framework for DNN Tasks on Edge
abstract
With the development of edge computing, DNN services have been widely deployed on edge devices. The deployment efficiency of deep learning models relies on the optimization of inference and scheduling policy. However, traditional optimization methods on edge devices still suffer from prohibitively long tuning time due to devices’ low computational power. Meanwhile, the widely used scheduling algorithm, the dominant resource fairness algorithm(DRF algorithm), struggles to maximize the efficiency of model execution on edge devices and inevitably increases average waiting time as it is not applicable in the real-time distributed computing environment. In this paper, we propose Cacomp, a distributed cloud-assisted deep learning compiler framework that features accelerating the optimization on edge devices with assistance from the cloud and a novel inference task scheduling algorithm. Our framework utilizes the tuning records from the cloud devices and proposes a two-step distillation strategy to obtain the best tuning record set for the edge device. For the scheduling process, we propose an RD-DRF algorithm to allocate inference tasks to edge devices based on dominant resource matching in real time. Extensive results show that our framework can achieve up to 2.19× improvement in the optimization time compared with other methods on edge devices. Our proposed scheduling algorithm significantly shortens the average waiting time of inference tasks by 30% and improves resource utilization by 20% on edge devices.
Weiwei Lin 0001, Jinhui Lin, Haotong Zhang 0003, Wentai Wu, Weizheng Wu, Zhetao Li, Keqin Li 0001
IEEE Trans. Computers4
2025 Pattern-Sensitive Local Differential Privacy for Finite-Range Time-Series Data in Mobile Crowdsensing
abstract
Time-series data is crucial for the development of mobile crowdsensing (MCS). Participant’s privacy is one of the major concerns because MCS data often contain sensitive individual information. Existing privacy-preserving mechanisms for time-series data do not preserve salient patterns of the time series and take into account that the perturbed data may fall outside the valid data interval, leading to data distortion. To overcome these deficiencies, we first perform dynamic feature extraction and incorporate an adaptive sampling scheme that is sensitive to the distinction of short-term patterns and stable patterns. Then a Bounded Laplace (BLP) mechanism is adopted with a theoretical guarantee on the data perturbation range so as to address the issue of data going beyond the valid range. We establish theoretically that the proposed Adaptive Sampling and Randomized perturbation mechanism based on dynamic Temporal patterns (ASRT) satisfies the metric-based$w$-event$\epsilon$-LDP for privacy protection. Empirical results of extensive experiments on realworld datasets demonstrate that our proposed method is superior to existing protection mechanisms and the efficacy of our ASRT in enhancing data utility without introducing outliers.
Zhetao Li, Xiyu Zeng, Wentai Wu, Haolin Liu 0001
IEEE Trans. Mob. Comput.5
2025 AdaptiveFL: Communication-Adaptive Federated Learning Under Dynamic Bandwidth
abstract
Federated learning (FL) is a distributed machine learning paradigm that enables heterogeneous devices to train a model collaboratively. Recognizing communication as a bottleneck in FL, existing communication-efficient solutions, e.g., HeteroFL and LotteryFL, etc., utilize gradient sparsification to reduce communication costs. However, existing solutions fail to address the dynamic bandwidth issue in which the bandwidth of each client is constantly changing throughout the training process. In this article, we propose AdaptiveFL, a communication-adaptive FL framework, considering the dynamic constraints of bandwidth. The design of AdaptiveFL follows two key steps: 1) in each round, each device selects a best-fit sub-model for communication per currently available bandwidth; and 2) to guarantee the performance of each sub-model sent under dynamic bandwidth constraints, AdaptiveFL employs a local training method that enables each device to train a "tailorable" local model, which can be tailored to any sparsity with competitive accuracy. We compare AdaptiveFL with several communication-efficient SOTA methods and demonstrate that AdaptiveFL outperforms other baselines by a large margin.
Guozhi Liu, Weiwei Lin 0001, Tiansheng Huang, Fang Shi, Wentai Wu, Li Shen 0008
IEEE Trans. Neural Networks Learn. Syst.5
2024 SparseTSF: Modeling Long-term Time Series Forecasting with *1k* Parameters
abstract
This paper introduces SparseTSF, a novel, extremely lightweight model for Long-term Time Series Forecasting (LTSF), designed to address the challenges of modeling complex temporal dependencies over extended horizons with minimal computational resources. At the heart of SparseTSF lies the Cross-Period Sparse Forecasting technique, which simplifies the forecasting task by decoupling the periodicity and trend in time series data. This technique involves downsampling the original sequences to focus on cross-period trend prediction, effectively extracting periodic features while minimizing the model’s complexity and parameter count. Based on this technique, the SparseTSF model uses fewer than 1k parameters to achieve competitive or superior performance compared to state-of-the-art models. Furthermore, SparseTSF showcases remarkable generalization capabilities, making it well-suited for scenarios with limited computational resources, small samples, or low-quality data. The code is publicly available at this repository: https://github.com/lss-1138/SparseTSF.
Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Haojun Chen
ICML3
2024 CycleNet: Enhancing Time Series Forecasting through Modeling Periodic Patterns
abstract
The stable periodic patterns present in time series data serve as the foundation for conducting long-horizon forecasts. In this paper, we pioneer the exploration of explicitly modeling this periodicity to enhance the performance of models in long-term time series forecasting (LTSF) tasks. Specifically, we introduce the Residual Cycle Forecasting (RCF) technique, which utilizes learnable recurrent cycles to model the inherent periodic patterns within sequences, and then performs predictions on the residual components of the modeled cycles. Combining RCF with a Linear layer or a shallow MLP forms the simple yet powerful method proposed in this paper, called CycleNet. CycleNet achieves state-of-the-art prediction accuracy in multiple domains including electricity, weather, and energy, while offering significant efficiency advantages by reducing over 90% of the required parameter quantity. Furthermore, as a novel plug-and-play technique, the RCF can also significantly improve the prediction accuracy of existing models, including PatchTST and iTransformer. The source code is available at: https://github.com/ACAT-SCUT/CycleNet.
Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Ruichao Mo, Haocheng Zhong
NeurIPS4
2024 Generative data augmentation with differential privacy for non-IID problem in decentralized clinical machine learning
Tianyu He, Peiyi Han, Shaoming Duan, Wentai Wu, Chuanyi Liu, Jianrun Han
Future Gener. Comput. Syst.5
2024 Energy-aware virtual machine placement based on a holistic thermal model for cloud data centers
Jianpeng Lin, Weiwei Lin 0001, Wentai Wu, Keqin Li 0001
Future Gener. Comput. Syst.3
2024 Reliable Task Offloading in Sustainable Edge Computing with Imperfect Channel State Information
abstract
As a promising paradigm, edge computing enhances service provisioning by offloading tasks to powerful servers at the network edge. Meanwhile, Non-Orthogonal Multiple Access (NOMA) and renewable energy sources are increasingly adopted for spectral efficiency and carbon footprint reduction. However, these new techniques inevitably introduce reliability risks to the edge system generally because of i) imperfect Channel State Information (CSI), which can misguide offloading decisions and cause transmission outages, and ii) unstable renewable energy supply, which complicates device availability. To tackle these issues, we first establish a system model that measures service reliability based on probabilistic principles for the NOMA-based edge system. As a solution, a Reliable Offloading method with Multi-Agent deep reinforcement learning (ROMA) is proposed. In ROMA, we first reformulate the reliability-critical constraint into an long-term optimization problem via Lyapunov optimization. We discretize the hybrid action space and convert the resource allocation on edge servers into a 0-1 knapsack problem. The optimization problem is then formulated as a Partially Observable Markov Decision Process (POMDP) and addressed by multi-agent proximal policy optimization (PPO). Experimental evaluations demonstrate the superiority of ROMA over existing methods in reducing grid energy costs and enhancing system reliability, achieving Pareto-optimal performance under various settings.
Peng Peng 0005, Wentai Wu, Weiwei Lin 0001, Fan Zhang 0112, Yongheng Liu, Keqin Li 0001
IEEE Trans. Netw. Serv. Manag.2
2023 Evolving Deep Multiple Kernel Learning Networks Through Genetic Algorithms
abstract
Today's Industrial Internet of Things (IIoT) have achieved excellent manufacturing efficiency and automation results by leveraging machine learning (ML) and deep learning (DL). However, trustworthiness of ML/DL brings significant challenges to IIoT. This article proposes an evolving deep multiple kernel learning network through genetic algorithm (KNGA). Our KNGA method uses genetic algorithm (GA) to find the best deep multiple kernel learning structure, including the weights and the topology of the model. Compared with the current well-known models, KNGA has advantages in three aspects: 1) It can achieve good results without using many samples during model training; 2) the model can evolve in the process of training, including self-growth, and self-pruning; and 3) its trustworthiness and reliability can be guaranteed. Moreover, the whole model ensures excellent performance and requires manual adjustment of only a few parameters. Extensive experiments on the UCI, KEEL, Caltech256, and MNIST datasets demonstrate the effectiveness and trustworthiness of the proposed method.
Wangbo Shen, Weiwei Lin 0001, Yulei Wu, Fang Shi, Wentai Wu, Keqin Li 0001
IEEE Trans. Ind. Informatics5
2023 Thermal-aware virtual machine placement based on multi-objective optimization
Bo Liu 0045, Weiwei Lin 0001, Wentai Wu, Jianpeng Lin, Keqin Li 0001
J. Supercomput.4
2023 Publisher Correction to: Thermal‑aware virtual machine placement based on multi‑objective optimization
Bo Liu 0045, Weiwei Lin 0001, Wentai Wu, Jianpeng Lin, Keqin Li 0001
J. Supercomput.4
2023 FedProf: Selective Federated Learning Based on Distributional Representation Profiling
abstract
Federated Learning (FL) has shown great potential as a privacy-preserving solution to learning from decentralized data that are only accessible to end devices (i.e., clients). The data locality constraint offers strong privacy protection but also makes FL sensitive to the condition of local data. Apart from statistical heterogeneity, a large proportion of the clients, in many scenarios, are probably in possession of low-quality data that are biased, noisy or even irrelevant. As a result, they could significantly slow down the convergence of the global model we aim to build and also compromise its quality. In light of this, we first present a new view of local data by looking into the representation space and observing that they converge in distribution to Normal distributions before activation. We provide theoretical analysis to support our finding. Further, we proposeFedProf, a novel algorithm for optimizing FL over non-IID data of mixed quality. The key of our approach is a distributional representation profiling and matching scheme that uses the global model to dynamically profile data representations and allows for low-cost, lightweight representation matching. Using the scheme we sample clients adaptively in FL to mitigate the impact of low-quality data on the training process. We evaluated our solution with extensive experiments on different tasks and data conditions under various FL settings. The results demonstrate that the selective behavior of our algorithm leads to a significant reduction in the number of communication rounds and the amount of time (up to 2.4× speedup) for the global model to converge and also provides accuracy gain.
Wentai Wu, Ligang He, Weiwei Lin 0001, Carsten Maple
IEEE Trans. Parallel Distributed Syst.1
2022 VFedCS: Optimizing Client Selection for Volatile Federated Learning
abstract
Federated learning (FL) has shown great potential as a privacy-preserving solution to training a centralized model based on local data from available clients. However, we argue that, over the course of training, the available clients may exhibit some volatility in terms of the client population, client data, and training status. Considering these volatilities, we propose a new learning scenario termed volatile federated learning (volatile FL) featuring set volatility, statistical volatility, and training volatility. The volatile client set along with the dynamic of clients’ data and the unreliable nature of clients (e.g., unintentional shutdown and network instability) greatly increase the difficulty of client selection. In this article, we formulate and decompose the global problem into two subproblems based on alternating minimization. For an efficient settlement for the proposed selection problem, we quantify the impact of clients’ data and resource heterogeneity for volatile FL and introduce the cumulative effective participation data (CEPD) as an optimization objective. Based on this, we propose upper confidence bound-based greedy selection, dubbed UCB-GS, to address the client selection problem in volatile FL. Theoretically, we prove that the regret of UCB-GS is strictly bounded by a finite constant, justifying its theoretical feasibility. Furthermore, experimental results show that our method significantly reduces the number of training rounds (by up to 62%) while increasing the global model’s accuracy by 7.51%.
Fang Shi, Chunchao Hu, Weiwei Lin 0001, Lisheng Fan, Tiansheng Huang, Wentai Wu
IEEE Internet Things J.6
2022 Developing an Unsupervised Real-Time Anomaly Detection Scheme for Time Series With Multi-Seasonality
abstract
On-line detection of anomalies in time series is a key technique used in various event-sensitive scenarios such as robotic system monitoring, smart sensor networks and data center security. However, the increasing diversity of data sources and the variety of demands make this task more challenging than ever. First, the rapid increase in unlabeled data means supervised learning is becoming less suitable in many cases. Second, a large portion of time series data have complex seasonality features. Third, on-line anomaly detection needs to be fast and reliable. In light of this, we have developed a prediction-driven, unsupervised anomaly detection scheme, which adopts a backbone model combining the decomposition and the inference of time series data. Further, we propose a novel metric, Local Trend Inconsistency (LTI), and an efficient detection algorithm that computes LTI in a real-time manner and scores each data point robustly in terms of its probability of being anomalous. We have conducted extensive experimentation to evaluate our algorithm with several datasets from both public repositories and production environments. The experimental results show that our scheme outperforms existing representative anomaly detection algorithms in terms of the commonly used metric, Area Under Curve (AUC), while achieving the desired efficiency.
Wentai Wu, Ligang He, Weiwei Lin 0001, Yuhua Cui, Carsten Maple, Stephen A. Jarvis
IEEE Trans. Knowl. Data Eng.1
2022 An On-Line Virtual Machine Consolidation Strategy for Dual Improvement in Performance and Energy Conservation of Server Clusters in Cloud Data Centers
abstract
As data centers are consuming massive amount of energy, improving the energy efficiency of cloud computing has emerged as a focus of research. However, it is challenging to reduce energy consumption while maintaining system performance without increasing the risk of Service Level Agreement violations. Most of the existing consolidation approaches for virtual machines (VMs) consider system performance and Quality of Service (QoS) metrics as constraints, which usually results in large scheduling overhead and impossibility to achieve effective improvement in energy efficiency without sacrificing some system performance and cloud service quality. In this article, we first define the metrics of peak power efficiency and optimal utilization for heterogeneous physical machines (PMs). Then we propose Peak Efficiency Aware Scheduling (PEAS), a novel strategy of VM placement and reallocation for achieving dual improvement in performance and energy conservation from the perspective of server clusters. PEAS allocates and reallocates VMs in an on-line manner and always attempts to maintain PMs working in their peak power efficiency via VM consolidation. Extensive experiments on Cloudsim show that PEAS outperforms several energy-aware consolidation algorithms with regard to energy consumption, system performance as well as multiple QoS metrics.
Weiwei Lin 0001, Wentai Wu, Ligang He
IEEE Trans. Serv. Comput.2
2021 SAFA: A Semi-Asynchronous Protocol for Fast Federated Learning With Low Overhead
abstract
Federated learning (FL) has attracted increasing attention as a promising approach to driving a vast number of end devices with artificial intelligence. However, it is very challenging to guarantee the efficiency of FL considering the unreliable nature of end devices while the cost of device-server communication cannot be neglected. In this article, we propose SAFA, a semi-asynchronous FL protocol, to address the problems in federated learning such as low round efficiency and poor convergence rate in extreme conditions (e.g., clients dropping offline frequently). We introduce novel designs in the steps of model distribution, client selection and global aggregation to mitigate the impacts of stragglers, crashes and model staleness in order to boost efficiency and improve the quality of the global model. We have conducted extensive experiments with typical machine learning tasks. The results demonstrate that the proposed protocol is effective in terms of shortening federated round duration, reducing local resource wastage, and improving the accuracy of the global model at an acceptable communication cost.
Wentai Wu, Ligang He, Weiwei Lin 0001, Rui Mao 0001, Carsten Maple, Stephen A. Jarvis
IEEE Trans. Computers1
2021 A Power Consumption Model for Cloud Servers Based on Elman Neural Network
abstract
Leveraging power consumption models in software systems can achieve easy deployment of low-cost, high-availability power monitoring in cloud datacenters that are usually large-scale, heterogeneous and frequently scaling up. However, traditional regression-based power consumption models generally have two drawbacks. First, their mathematical forms are usually fixed and determined a priori. This may cause unacceptable increase of error or over-fitting as the power signatures of cloud servers are usually uncertain. Second, the characteristic of workload dispatched to cloud servers is constantly changing while regression-based models can hardly generalize to a wide range of servers and workload types. As a novel solution, we in this paper propose a server power consumption model based on Elman Neural Network (PCM-ENN), aiming to allow accurate and flexible power estimation. PCM-ENN is an end-to-end black box model capable of learning the temporal relation between samples in a time series of power consumption. We trained and evaluated PCM-ENN on two power sequence datasets collected from heterogeneous hardware and operating systems running quasi-production benchmarks like CloudSuite. Experimental result shows that PCM-ENN generated accurate estimates on server power consumption with only small errors, outperforming widely-used linear regression model and NARX model in terms of accuracy.
Wentai Wu, Weiwei Lin 0001, Ligang He, Guangxin Wu, Ching-Hsien Hsu
IEEE Trans. Cloud Comput.1
2021 An Efficiency-Boosting Client Selection Scheme for Federated Learning With Fairness Guarantee
abstract
The issue of potential privacy leakage during centralized AI's model training has drawn intensive concern from the public. A Parallel and Distributed Computing (or PDC) scheme, termed Federated Learning (FL), has emerged as a new paradigm to cope with the privacy issue by allowing clients to perform model training locally, without the necessity to upload their personal sensitive data. In FL, the number of clients could be sufficiently large, but the bandwidth available for model distribution and re-upload is quite limited, making it sensible to only involve part of the volunteers to participate in the training process. The client selection policy is critical to an FL process in terms of training efficiency, the final model's quality as well as fairness. In this article, we will model the fairness guaranteed client selection as a Lyapunov optimization problem and then a C2MAB-based method is proposed for estimation of the model exchange time between each client and the server, based on which we design a fairness guaranteed algorithm termed RBCS-F for problem-solving. The regret of RBCS-F is strictly bounded by a finite constant, justifying its theoretical feasibility. Barring the theoretical results, more empirical data can be derived from our real training experiments on public datasets.
Tiansheng Huang, Weiwei Lin 0001, Wentai Wu, Ligang He, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.3
2021 Accelerating Federated Learning Over Reliability-Agnostic Clients in Mobile Edge Computing Systems
abstract
Mobile Edge Computing (MEC), which incorporates the Cloud, edge nodes, and end devices, has shown great potential in bringing data processing closer to the data sources. Meanwhile, Federated learning (FL) has emerged as a promising privacy-preserving approach to facilitating AI applications. However, it remains a big challenge to optimize the efficiency and effectiveness of FL when it is integrated with the MEC architecture. Moreover, the unreliable nature (e.g., stragglers and intermittent drop-out) of end devices significantly slows down the FL process and affects the global model's quality in such circumstances. In this article, a multi-layer federated learning protocol called HybridFL is designed for the MEC architecture. HybridFL adopts two levels (the edge level and the cloud level) of model aggregation enacting different aggregation strategies. Moreover, in order to mitigate stragglers and end device drop-out, we introduce regional slack factors into the stage of client selection performed at the edge nodes using a probabilistic approach without identifying or probing the state of end devices (whose reliability is agnostic). We demonstrate the effectiveness of our method in modulating the proportion of clients selected and present the convergence analysis for our protocol. We have conducted extensive experiments with machine learning tasks in different scales of MEC system. The results show that HybridFL improves the FL training process significantly in terms of shortening the federated round length, speeding up the global model's convergence (by up to 12×) and reducing end device energy consumption (by up to 58 percent).
Wentai Wu, Ligang He, Weiwei Lin 0001, Rui Mao 0001
IEEE Trans. Parallel Distributed Syst.1
2020 A novel syntax-aware automatic graphics code generation with attention-based deep neural network
Xiong Wen Pang, Yanqiang Zhou, Weiwei Lin 0001, Wentai Wu, James Zijun Wang
J. Netw. Comput. Appl.5
2018 Experimental and quantitative analysis of server power model for cloud data centers
Weiwei Lin 0001, Wentai Wu, James Zijun Wang, Ching-Hsien Hsu
Future Gener. Comput. Syst.2
2018 Energy-efficient hadoop for big data analytics and computing: A systematic review and research insights
Wentai Wu, Weiwei Lin 0001, Ching-Hsien Hsu, Ligang He
Future Gener. Comput. Syst.1
2017 An intelligent power consumption model for virtual machines under CPU-intensive workload in cloud environment
Wentai Wu, Weiwei Lin 0001, Zhiping Peng
Soft Comput.1