VLDB 2026 Research / reviewers in the wild / expert
Ne Wang
dblp:236/3200
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-3821-785XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient KV Cache Migration for Geo-Distributed LLM Inference in Collaborative Edge Computing
Mingjin Zhang, Jiannong Cao 0001, Xiangchun Chen, Ne Wang |
ICDCS | 5 |
| 2026 | HALO: A scalable framework for hotness-aware coding and transformation-efficient placement
Junmei Chen, Ne Wang, Zongpeng Li, Zhiquan Liu 0001, Dan Xiang |
Future Gener. Comput. Syst. | 2 |
| 2025 | SP-MoE: Expediting Mixture-of-Experts Training with Optimized Pipelining PlanningabstractSparsely activated Mixture-of-Experts (MoE) has emerged as a key technique to expand the size of Transformer-based large language models (LLMs) while maintaining low computational costs. However, MoE layers require to route the input data to distributed devices, incurring significant communication latency. Existing studies have primarily focused on alleviating this problem by overlapping computation and communication tasks within a single MoE layer, which fails to achieve sufficient overlap and results in limited performance gains. In this work, we introduce an orthogonal partitioning dimension from existing task-parallel methods by leveraging the autoregressive nature of causal Transformer-based LLMs, i.e. partitioning tasks along the sequence dimension. This provides more flexible and efficient overlaps among tasks from both non-MoE and MoE layers. To this end, we propose an efficient MoE training approach, SP-MoE, with two innovative designs. 1) It incorporates non-MoE layers into the overlapping with not only the current MoE layer but also the preceding MoE layer, thereby facilitating more efficient training; 2) It identifies the optimal combination of pipeline degrees for non-MoE and MoE layers and devises the best scheduling plans for load-imbalanced non-MoE and uniform MoE layers to achieve the goal of minimizing the total training latency. Extensive experiments conducted on two GPU clusters demonstrate that SP-MoE can effectively identify the optimal combination of pipeline degrees and achieve 16.1% - 34.3% reduction in training latency compared to three state-of-the-art MoE systems. Ne Wang, Wenxiang Lin, Lin Zhang 0059, Shaohuai Shi, Ruiting Zhou, Bo Li 0001 |
INFOCOM | 1 |
| 2025 | Efficient and Intelligent Multijob Federated Learning in Wireless NetworksabstractFederated learning (FL) has emerged as an innovative paradigm designed to protect privacy by enabling collaborative machine learning (ML) model training across multiple data owners (also known as clients) without the need to access clients’ raw data. The majority of existing FL research concentrates on scenarios where a single job necessitates training. In practical applications, multiple FL jobs can simultaneously undergo training using a common pool of clients, a scenario known as multijob FL. However, the problem of FL training with multiple jobs remains open and presents significant challenges of the escalated heterogeneity of jobs and clients, complex tradeoffs between training latency and energy consumption, uncertainty of client quality, and potential linear switching cost associated with client selection. This work aims to jointly optimize training efficiency in terms of latency, energy consumption, and switching cost for multiple jobs in stochastic and dynamic environments. Specifically, we propose a novel multijob FL framework, namedEffI-FL, incorporating three innovative designs: 1) to reduce switching cost, we extend the client selection interval from every round to multiple rounds, called a block, within which client subset switching is prohibited; 2) we employ multiarmed bandit (MAB) methods to measure clients’ latency and energy cost under uncertainty. Additionally, we utilize the virtual queue technique to trace clients’ battery usage patterns. By integrating the above client-side knowledge, we propose an adaptive client selection policy aimed at balancing latency, energy consumption, and battery condition; and 3) given that multiple jobs may compete for the same client, we devise a greedy algorithm to assign each client to a single job. We rigorously prove that the regret of our client selection policy and the cost of our block-wise client subset switching algorithm are both sublinear. Finally, we implementEffI-FLusing PyTorch and conduct experiments demonstrating thatEffI-FLreduces the weighted sum of latency, energy consumption, and switching cost by up to 52.3% compared to four state-of-the-art FL frameworks. Jiajin Wang, Ne Wang, Ruiting Zhou, Bo Li 0001 |
IEEE Internet Things J. | 2 |
| 2024 | Boosting Correlated Failure Repair in SSD Data CentersabstractCurrent data centers rely on failure protection mechanisms to ensure data reliability. However, recent research indicates that failures within the same node or rack are common in data centers that use flash-based solid-state drives (SSDs) as the primary storage medium. Such correlated failures bring challenges for traditional protection mechanisms to achieve high reliability and repair performance. To this end, we propose a product erasure code (PECode) that encodes data blocks in multiple stripes cooperatively to generate intrastripe and interstripe parity blocks. Then, we design a multistripe cooperative repair algorithm (MSCRepair). MSCRepair first creates the failure distribution matrix (FDM) to represent the distribution of failure blocks in nodes and racks, and then conducts FDM-guided repair to minimize cross-rack traffic upon correlated failures. We prove that MSCRepair achieves the least cross-rack repair traffic at the cost of a longer repair time. We further propose a correlated failure repair scheduling algorithm for MSCRepair, which reduces the repair time by balancing the load and delivering data from links with higher bandwidths. We evaluate MSCRepair through both large-scale simulations and real experiments. In the mise-en-scene of its state-of-the-art alternatives, MSCRepair stands out by reducing up to 19.6%–49.9% of cross-rack traffic, while simultaneously reducing 16.2%–51.4% of recovery time of correlated failures. Junmei Chen, Zongpeng Li, Qifu Tyler Sun, Ne Wang, Lina Su |
IEEE Internet Things J. | 4 |
| 2024 | Advanced Elastic Reed-Solomon Codes for Erasure-Coded Key-Value StoresabstractErasure coding is a storage-efficient redundancy scheme for modern key–value (KV) stores, storing stripes of data and parity chunks in multiple nodes. To accommodate the highly skewed and time-varying nature of the workload, KV stores require erasure code that dynamically optimizes its parameters, known as redundancy converting. Stretched Reed–Solomon (SRS) and elastic Reed–Solomon (ERS) codes represent promising candidates for meeting such requirements. However, both SRS and ERS are limited to RS$(d,r)\to $RS$(d^{\prime },r^{\prime })$converting, where$d^{\prime }>d,r^{\prime }=r$, failing to fully meet actual needs. This work presents an advanced ERS code (AERS code), which builds upon flexible encoding matrices and placement strategies, serving different types of redundancy converting, and minimizing converting traffic. We further prove that the AERS code is an optimal redundancy converting solution that achieves the theoretical lower bound on data traffic during redundancy converting while guaranteeing node-level fault tolerance. We evaluate the AERS code through both mathematical analysis and experiments. In the mise-en-scène of its state-of-the-art alternatives, AERS stands out by reducing network traffic up to 50%–85.7% while accelerating redundancy converting. Junmei Chen, Zongpeng Li, Ruiting Zhou, Lina Su, Ne Wang |
IEEE Internet Things J. | 5 |
| 2024 | Low-Latency Hierarchical Federated Learning in Wireless Edge NetworksabstractHierarchical federated learning (HFL) has recently emerged as a more practical machine learning (ML) paradigm, which enables edge servers (ESs) in close proximity to conduct partial model aggregation. Despite its utility, local training and model aggregation incur considerable computation and communication time. client selection (CS) has proven effective for minimizing latency. However, CS faces the following challenges in hierarchical federated learning (HFL). First, the accessible clients, computation resources and network bandwidth are time-varying and unpredictable. Second, certain dynamics can only be observed after the decisions are made. Third, multiple ESs face different unknown clients, increasing the difficulty of selecting clients in an online manner. Finally, resource usage may be excessively violated during the training process. Existing HFL researches are insufficient to tackle these challenges. This work proposes a multi- ESs CS framework (MCS), which is based on multiarmed bandit (MAB) technique. MCS aims to reduce the cumulative computation and communication time, using two algorithms: 1) an online learning-based CS algorithm (OCA) makes the CS decisions for each ES, based on empirical learning results; and 2) a randomized rounding algorithm (RRA) converts fractional decisions obtained by OCA into binary solutions. Theoretically, MCS can enjoy the sublinear regret and violation compared to the optimal strategy. Practically, extensive experiments on real-world data sets demonstrate the empirical superiority of MCS over multiple state-of-the-art algorithms in minimizing cumulative latency. Lina Su, Ruiting Zhou, Ne Wang, Junmei Chen, Zongpeng Li |
IEEE Internet Things J. | 3 |
| 2024 | Adaptive Pricing and Online Scheduling for Distributed Machine Learning JobsabstractLarge-scale distributed machine learning (ML) systems involve extensive and costly computational resources. Pricing and scheduling, as two promising techniques for resource management, have garnered significant attention. However, existing job pricing and scheduling algorithms in cloud computing either charge fixed resource fees based on known job runtime or implement dynamic price setting with job preemption, unsuitable for distributed ML systems with high uncertainties and switching cost. First, whether the resources of a distributed ML job are placed together or not results in different job runtime. Second, various time-varying factors, including job arrival rates and competitors’ pricing, affect resource prices. Third, frequent price changes for the same resource can easily lead to system instability, ultimately jeopardizing user satisfaction. Addressing these uncertainties is challenging. This article introducesAPOS, an adaptive pricing and online scheduling framework, aiming at maximizing the operator’s overall revenue.APOSincorporates two innovations: 1) Intelligent Pricing: We represent each price using a feature vector that encapsulates relevant factors. Subsequently, based on the linear upper confidence bound (UCB) techniques, we establish relationships between price features and two revenue-associated elements: a) job arrival rates and b) resource consumption rates. To ensure system stability, we introduce batch pricing to reduce the frequency of resource price updates and 2) Online Scheduling: We strive to compute a nonpreemptive schedule that balances job utility with corresponding resource cost. We rigorously prove thatAPOSachieves truthfulness, individual rationality, system stability, and sublinear regret in polynomial time. Finally, extensive trace-driven simulations confirm thatAPOSoutperforms four state-of-the-art baselines, yielding a minimum of 23.3% improvement in total operator revenue. Lina Su, Junmei Chen, Ne Wang, Zongpeng Li |
IEEE Internet Things J. | 4 |
| 2023 | Online Scheduling of Distributed Machine Learning Jobs for Incentivizing Sharing in Multi-Tenant SystemsabstractTo save cost, companies usually train machine learning (ML) models on a shared multi-tenant system. In this cooperative environment, one of the fundamental challenges is how to distribute resources fairly among tenants such that each tenant is satisfied. A satisfactory allocation policy needs to meet the following properties. First, the performance of each tenant in the shared cluster is at least the same as that in its exclusive cluster partition. Second, no tenant can get more benefits by lying about its demands. Third, tenants cannot use the idle resources of others for free. Moreover, the resource allocation for ML workloads should avoid costly migration overhead. To this end, we propose a three-layer scheduling framework Astraea: i) a batch scheduling framework groups unprocessed jobs into multiple batches; ii) a round-by-round algorithm enables tenants to reserve their share of resources and schedule jobs in a non-preemptive manner; iii) one-round algorithm based on primal-dual approach and posted pricing framework, which encourages tenants to report truthful demands. Astraea is proven to achieve performance guarantee and some desirable properties of sharing, including sharing incentive, strategy-proofness and gain-as-you-contribute fairness. Extensive trace-driven simulations show Astraea advances in both fairness and cluster efficiency compared to three state-of-the-art baselines. Ne Wang, Ruiting Zhou, Zongpeng Li |
IEEE Trans. Computers | 1 |
| 2023 | DPS: Dynamic Pricing and Scheduling for Distributed Machine Learning Jobs in Edge-Cloud Networksabstract5G and Internet of Things stimulate smart applications of edge computing, such as autonomous driving and smart city. As edge computing power increases, more and more machine learning (ML) jobs will be trained in the edge-cloud network, adopting the parameter server (PS) architecture. Due to the distinct features of the edge (low-latency and the scarcity of resources), the cloud (high delay and rich computing capacity) and ML jobs (frequent communication between workers and PSs and unfixed runtime), existing cloud job pricing and scheduling algorithms are not applicable. Therefore, how to price, deploy and schedule ML jobs in the edge-cloud network becomes a challenging problem. To solve it, we propose an auction-based online framework DPS. DPS consists of three major parts: job admission control, price function design and scheduling orchestrator. DPS dynamically prices workers and PSs based on historical job information and real-time system status, and decides whether to accept the job according to the deployment cost. DPS then deploys and schedules accepted ML jobs to pursue the maximum social welfare. Through theoretical analysis, we prove that DPS can achieve a good competition ratio and truthfulness in polynomial time. Large-scale simulations and testbed experiments show that DPS can improve social welfare by at least$95\%$, compared with benchmark algorithms in today's cloud system. Ruiting Zhou, Ne Wang, Jinlong Pang |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | An Online Learning Approach for Client Selection in Federated Edge Learning under Budget ConstraintabstractFederated learning (FL) has emerged as a new paradigm that enables distributed mobile devices to learn a global model collaboratively. Since mobile devices (a.k.a, clients) exhibit diversity in model training quality, client selection (CS) becomes critical for efficient FL. CS faces the following challenges: First, the client’s availability, the training data volumes, and the network connection status are time-varying and cannot be easily predicted. Second, clients for training and the number of local iterations would seriously affect the model accuracy. Thus, selecting a subset of available clients and controlling local iterations should guarantee model quality. Third, renting clients for model training needs cost. It is necessary to dynamically administrate the use of the long-term budget without knowledge of future inputs. To this end, we propose a federated edge learning (FedL) framework, which can select appropriate clients and control the number of training iterations in real-time. FedL aims to reduce the completion time while reaching the desired model convergence and satisfying the long-term budget for renting clients. FedL consists of two algorithms: i) the online learning algorithm makes CS and iteration decisions according to historic learning results; ii) the online rounding algorithm translates fractional decisions derived by the online learning algorithm into integers to satisfy feasibility constraints. Rigorous mathematical proof reveals that dynamic regret and dynamic fit have sub-linear upper-bounds with time for a given budget. Extensive experiments based on realistic datasets suggest that FedL outperforms multiple state-of-the-art algorithms. In particular, FedL reduces at least 38% completion time compared with others. Lina Su, Ruiting Zhou, Ne Wang, Guang Fang, Zongpeng Li |
ICPP | 3 |
| 2022 | Multi-agent Multi-armed Bandit Learning for Content Caching in Edge NetworksabstractAs a new paradigm, edge caching is deemed an effective alternative by fetching contents at the network edge. However, designing an efficient caching mechanism is challenging. First, the content library is a dynamic set rather than a static set. Second, the content may be prevalent in different small base stations (SBSs), resulting in different rewards. Thus, the above reasons require each SBS could learn its caching decisions in a multi-SBSs network. Existing reinforcement learning algorithms either fail to consider the non-stationary environment or do not provide any performance guarantee. Thus, previous algorithms work well no longer. This work proposes a multi-agent multi-armed bandit caching framework, MAMAB-C, which navigates SBSs to cache contents in a distributed manner. Specifically, we formulate the multi-SBSs caching optimization problem as an online integer linear program (ILP) and convert it into a multi-agent multi-armed bandit (MAMAB) problem with resource constraints. MAMAB-C can realize the sub-linear metric property and significantly outperform multiple state-of-the-art algorithms. Lina Su, Ruiting Zhou, Ne Wang, Junmei Chen, Zongpeng Li |
ICWS | 3 |
| 2022 | Adaptive Clustered Federated Learning for Clients with Time-Varying InterestsabstractClustered Federated Learning (FL) addresses heterogeneous objectives from different client groups, by capturing the intrinsic relationship between data distributions of clients. This work aims to minimize the completion time of clustered FL training while guaranteeing convergence, given the following challenges. First, clients’ data distributions are not static since their interests are usually time-varying. Obsolete data may incur training failures, requiring detection of distribution changes at runtime. Second, even with the same distribution, client datasets may have different contributions to model accuracy. Besides, the training data typically arrive at clients dynamically, which brings uncertainties to assessing the quality of client data. Third, the execution environments of clients and networks are often unstable and stochastic, leading to uncertainties in calculating computation and communication time. Given the above challenges, we propose Acct with two innovations: i) change detection: we first model the time-varying interests of clients as piecewise stationary based on practical observations, then apply generalized likelihood ratio detectors to FL for detecting changes in client distributions; ii) client selection: we adopt the multi-armed bandit (MAB) technique to account for the uncertainties in measuring data quality, computation and communication time. Based on the upper confidence bound (UCB) method, we construct a novel “double UCB” policy to adaptively select clients with high data quality and low computation and communication overhead. We rigorously prove the convergence of Acct and sub-linear regret regarding the proposed client selection policy. Finally, we implement Acct using PyTorch and conduct experiments showing that Acct reduces the completion time by almost 18.2% compared with three state-of-the-art FL frameworks. Ne Wang, Ruiting Zhou, Lina Su, Guang Fang, Zongpeng Li |
IWQoS | 1 |
| 2022 | Dynamic service placement and request scheduling for edge networks
Lina Su, Ne Wang, Ruiting Zhou, Zongpeng Li |
Comput. Networks | 2 |
| 2022 | Preemptive Scheduling for Distributed Machine Learning Jobs in Edge-Cloud NetworksabstractRecent advances in 5G and edge computing enable rapid development and deployment of edge-cloud systems, which are ideal for delay-sensitive machine learning (ML) applications such as autonomous driving and smart city. Distributed ML jobs often need to train a large model with enormous datasets, which can only be handled by deploying a distributed set of workers in an edge-cloud system. One common approach is to employ a parameter server (PS) architecture, in which training is carried out at multiple workers, while PSs are used for aggregation and model updates. In this architecture, one of the fundamental challenges is how to dispatch ML jobs to workers and PSs such that the average job completion time (JCT) can be minimized. In this work, we propose a novel online preemptive scheduling framework to decide the location and the execution time window of concurrent workers and PSs upon each job arrival. Specifically, our proposed scheduling framework consists of: i) a job dispatching and scheduling algorithm that assigns each ML job to workers and decides the schedule to train each data chunk; ii) a PS assignment algorithm that determines the placement of PS. We prove theoretically that our proposed algorithm is$D_{max}(1+1/\epsilon)$-competitive with$(1 + \epsilon)$-speed augmentation, where$D_{max}$is the maximal number of data chunks in any job. Extensive testbed experiments and trace-driven simulations show that our algorithm can reduce the average JCT by up to 30% compared with state-of-the-art baselines. Ne Wang, Ruiting Zhou, Lei Jiao 0002, Renli Zhang, Bo Li 0001, Zongpeng Li |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | Research on Fast and Parallel Clustering Method for Trajectory DataabstractIn the era of big data, the development of satellite technology and Internet of Things has produced a large amount of trajectory data. We can effectively understand and predict the movement of the objects by analyzing their trajectory data. Now, most of density-based clustering algorithms have some disadvantages including the difficulty to determine input parameters, large I/O, and so on. DPC (Clustering by fast search and find of Density Peaks) is a new density-based clustering algorithm, which is simple and has only one input parameter, and also it is not affected by the data dimension, therefore, it can be effectively applied for trajectory clustering. However, in DPC, the local density is complex to calculate, and the cutoff distance is subjective to determine. In addition, DPC does not consider the existence of multiple cluster centers in the same cluster when clustering. To solve these problems, in this paper a fast clustering algorithm for trajectory data is put forward. In addition, Spark memory computing technology and data partitioning method are used to parallelize the algorithm, which greatly improves the clustering efficiency. Finally, experiments with three months' ship trajectory data from the Yangtze River have demonstrated that the clustering efficiency and effectiveness of our algorithm are significantly improved. Ne Wang, Shu Gao, Xiangwen Peng, Minrui Wang |
ICPADS | 1 |