VLDB 2026 Research / reviewers in the wild / expert
Jingpu Duan
dblp:167/9016
· DBLP profile ↗
45ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0001-7507-2197ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 37 · 4 first-author · 30 since 2021Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SwitchNN: In-Network CNN Inference for Edge-Assisted Smart Roadside Networks
Jianqiang Zhong, Jingpu Duan, Wenfei Wu, Deke Guo, Bingyang Liu, Xu Chen 0004 |
IWQoS | 2 |
| 2026 | HetTraffic: Multi-link traffic prediction and allocation for 6G heterogeneous networks
Yali Lv, Jingpu Duan |
Ad Hoc Networks | 3 |
| 2026 | A novel hybrid neural network for high-accuracy vehicle-to-infrastructure network traffic predictionabstractTo address the challenges in Vehicle-to-Infrastructure (V2I) network traffic prediction, this study proposes an innovative solution. We first establish a novel paradigm that integrates physical models to systematically convert publicly available vehicle trajectory data into V2I traffic data. On this basis, a gCNN–BiLSTM–MHA deep learning model is constructed, whose core advantage lies in its use of a lightweight GhostNet-based convolutional network (gCNN) to improve computational efficiency, while leveraging the synergistic effect of a bidirectional long short-term memory network (BiLSTM) and a multi-head attention mechanism (MHA) to effectively balance prediction efficiency and accuracy. The model’s superiority is comprehensively validated: compared to baseline models like LSTM, it demonstrates significant advantages across a series of key evaluation metrics — including running time, MBD, MAE, MAPE, RMSE, and R 2 — achieving an overall balanced performance. Furthermore, the model exhibits excellent performance on multiple benchmark datasets, confirming its strong robustness and high applicability for complex V2I network traffic prediction tasks. Xiaosheng Ni, Jingpu Duan |
Adv. Eng. Informatics | 2 |
| 2026 | Learning from easy to hard: Curriculum meta-learning for few-shot node classification
Qilong Yan, Weinan Guan, Yifei Xing 0001, Jingpu Duan, Jian Yin 0001 |
Inf. Sci. | 4 |
| 2026 | Resource Allocation in RIS-Assisted Integrated Sensing, Communication, and Computation NetworkabstractIntegrated sensing and communication (ISAC) is an emerging paradigm designed to support next-generation wireless services and applications. However, ISAC systems with limited computation capabilities are unable to handle computation-intensive and latency-sensitive tasks. This paper proposes a novel integrated sensing, communication, and computation (ISCC) network empowered by a reconfigurable intelligent surface (RIS) to mitigate the performance degradation caused by interference between radar sensing and uplink offloading. To effectively coordinate the cross-layer resource allocation among communication, sensing, and computation, we propose a resource scheduling problem. Specifically, we maximize the total computation rate while satisfying the sensing signal-to-noise ratio (SNR) requirement by jointly optimizing the energy allocation for local computing and offloading, the transmit and receive beamforming at the base station (BS), and the RIS reflective beamforming. To address this complex non-convex problem, we develop an efficient scheduling algorithm based on the block coordinate descent (BCD) framework. The iterative algorithm employs the fractional programming algorithm based on Lagrangian dual transform and quadratic transform, the generalized eigenvector methods, the convex relaxation techniques, and the successive convex approximation (SCA) algorithms to solve each subproblem separately. Experimental results demonstrate that the proposed scheme outperforms several baseline methods, confirming that RIS technology can effectively enhance system performance. In addition, we reveal the impact of various parameters on system performance. Yingsheng Peng, Jinbei Zhang, Jingpu Duan, Weichao Li 0001, Yong Liu 0005 |
IEEE Trans. Commun. | 3 |
| 2026 | Joint Bitrate and Resource Adaptation for Super-Resolution Video Streaming in Multi-Cluster Edge Networks: A New Online Learning ApproachabstractToday's video streaming service providers have exploited cloud-edge collaborative networks for video delivery across geo-distributed edge clusters and end users. The existing content delivery network (CDN) scheduling and adaptive bitrate algorithms may not fully utilize edge resources or lack a global control to optimize resource sharing. The emerging super-resolution (SR) approach can unleash the potential of leveraging computation resources to compensate for bandwidth consumption, by producing high-quality videos from low-resolution contents. Yet the uncertain SR resource sensitivity and its interplay with bitrate adaptation are under-explored. In this work, we proposeRosevin, the first resource scheduler that jointly decides the bitrates and fine-grained resource allocation to perform SR at the edge, which can learn to optimize the long-term QoE for distributed end users. To handle the time-varying and complex space of decisions as well as a non-smooth objective function,Rosevinrealizes a novel online combinatorial learning algorithm, which nicely integrates convex optimization theories and online learning techniques, addressing the switching cost issues. In addition to theoretically analyzing its performance, we implement an SR-assisted video streaming prototype ofRosevinand demonstrate its advantages over several video delivery benchmarks. Xiaoxi Zhang 0001, Longhao Zou, Jingpu Duan, Chuan Wu 0001, Yali Xue, Zuozhou Chen, Chaoqi Zhou, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Online Location Planning for AI-Defined Vehicles: Optimizing Joint Tasks of Order Serving and Spatio-Temporal Heterogeneous Model Fine-TuningabstractAdvances in artificial intelligence (AI) including foundation models (FMs), are increasingly transforming human society, with smart city driving the evolution of urban living. Meanwhile, vehicle crowdsensing (VCS) has emerged as a key enabler, leveraging vehicles' mobility and sensor-equipped capabilities. In particular, ride-hailing vehicles can effectively facilitate flexible data collection and contribute towards urban intelligence, despite resource limitations. Therefore, this work explores a promising scenario, where edge-assisted vehicles perform joint tasks of order serving and the emerging foundation model finetuning using various urban data. However, integrating the VCS AI task with the conventional order serving task is challenging, due to their inconsistent spatio-temporal characteristics: (i) The distributions of ride orders and data point-of-interests (PoIs) may not coincide in geography, both following a priori unknown patterns; (ii) they have distinct forms of temporal effects, i.e., prolonged waiting makes orders become instantly invalid while data with increased staleness gradually reduces its utility for model fine-tuning. To overcome these obstacles, we propose an online framework based on multi-agent reinforcement learning (MARL) with careful augmentation. A new quality-of-service (QoS) metric is designed to characterize and balance the utility of the two joint tasks, under the effects of varying data volumes and staleness. We also integrate graph neural networks (GNNs) with MARL to enhance state representations, capturing graph-structured, time-varying dependencies among vehicles and across locations. Extensive experiments on our testbed simulator, utilizing various real-world foundation model fine-tuning tasks and the New York City Taxi ride order dataset, demonstrate the advantage of our proposed method. Bokeng Zheng, Bo Rao, Tianxiang Zhu, Chee-Wei Tan 0001, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004, Xiaoxi Zhang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | User-Intent-Driven Semantic Communication via Adaptive Deep UnderstandingabstractSemantic communication focuses on transmitting task-relevant semantic information, aiming for intent-oriented communication. While existing systems improve efficiency by extracting key semantics, they still fail to deeply understand and generalize users’ real intentions. To overcome this, we propose a user-intention-driven semantic communication system that interprets diverse abstract intents. First, we integrate multimodal Large Models as semantic knowledge base to generate user-intention prior. Next, a mask-guided attention module is proposed to effectively highlight critical semantic regions. Further, a channel state awareness module ensures adaptive, robust transmission across varying channel conditions. Extensive experiments demonstrate that our system achieves deep intent understanding and outperforms DeepJSCC, e.g., under a Rayleigh channel at an SNR of 5 dB, it achieves improvements of 8%, 6%, and 19% in PSNR, SSIM, and LPIPS, respectively. Peigen Ye, Jingpu Duan, Hongyang Du 0001, Yulan Guo |
GLOBECOM | 2 |
| 2025 | TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive CorrectionabstractNon-independent and identically distributed (Non-IID) data across edge clients have long posed significant challenges to federated learning (FL) training. Prior works have proposed various methods to mitigate this statistical heterogeneity. While these methods can achieve good theoretical performance, they may lead to the over-correction problem, which degrades model performance and even causes failures in model convergence. In this paper, we provide the first investigation into the hidden over-correction phenomenon brought by the uniform model correction coefficients across clients adopted by the existing methods. To address this problem, we propose TACO, a novel algorithm that addresses the non-IID nature of clients’ data by implementing fine-grained, client-specific gradient correction and model aggregation, steering local models towards a more accurate global optimum. Moreover, we verify that leading FL algorithms generally have better model accuracy in terms of communication rounds rather than wall-clock time, resulting from their extra computation overhead imposed on clients. To enhance the training efficiency, TACO deploys a lightweight model correction and tailored aggregation approach that requires minimum computation overhead and no extra information beyond the synchronized model parameters. To validate TACO’s effectiveness, we present the first FL convergence analysis that reveals the root cause of over-correction. Extensive experiments across various datasets confirm TACO’s superior and stable performance in practice. Ziwei Zhan, Carlee Joe-Wong, Edith C. H. Ngai, Jingpu Duan, Deke Guo, Xu Chen 0004, Xiaoxi Zhang 0001 |
ICDCS | 5 |
| 2025 | PASTA: Training Acceleration for Vertical Federated Learning via Adaptive Pipeline ParallelismabstractVertical federated learning (VFL) enables collaborative model training among geo-distributed participants, each with different features of the same samples, but only one party possesses the labels. Communication delays between active and passive parties in VFL significantly hinder its training efficiency. Existing VFL methods adopt asynchronous schemes or multiple local updates per communication round, but they either introduce heavy computation overhead or fail to adapt to dynamic network conditions. This work proposes PASTA, a novel framework employing Adaptive Pipeline Parallelism with Staleness Control for VFL, designed to mitigate these delays and balance training efficiency and model performance. PASTA enables concurrent communication and computation, maximizing resource utilization and minimizing idle time by strategically using stale gradients. Each passive party can send one or more batches of embeddings per communication and conduct stale local training, so that computation times can overlap with communication latency. Since staleness impedes model accuracy despite its benefits in reducing time, a dynamic feedback-based mechanism is proposed to adjust the numbers of embeddings sent and local training iterations based on system heterogeneity. Extensive experiments across various datasets demonstrate that PASTA significantly enhances convergence speed by$1.8 \times$to$4.6 \times$compared to leading VFL systems, without compromising final accuracy. The source code is available at https://github.com/PointerA/PASTA. Ziwei Zhan, Jingpu Duan, Chuan Wu 0001, Jinhang Zuo, Xu Chen 0004, Xiaoxi Zhang 0001 |
IWQoS | 5 |
| 2025 | Resource allocation and pricing for SFC deployment in Space-Air-Ground-Integrated Networks: An innovative auction-based strategy
Yali Lv, Xiaoxi Zhang 0001, Yingsheng Peng, Jingpu Duan, Bo Yi 0002, Qing Li 0006 |
Comput. Networks | 4 |
| 2025 | Stateless and Proactive Routing for Dynamic Multicast With Deep Reinforcement LearningabstractStateful multicast protocols manage multicast group memberships by maintaining state information about active groups and their members. They have seen limited adoption in the modern internet due to lack of scalability, simplicity, and flexibility. Although stateless multicast protocols, like BIER, eliminate extensive state management, they still face complex tree computation and limited scalability for concurrent requests. In this paper, we propose Hawkeye, a stateless multicast mechanism with deep reinforcement learning (DRL) for real-time responses to dynamic multicast requests with near-optimal multicast TE performance. This mechanism is suited for Software-Defined Networking (SDN) environment where the controller has a global view of the network and supports flexible configuration of network resources for traffic engineering. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership, and thus are able to build multicast trees proactively for upcoming requests. We develop a novel source aggregation mechanism to facilitate the convergence of the DRL agent under high volume of multicast requests. Moreover, to improve the practicality and robustness of Hawkeye, we design incremental deployment and single failure handling mechanisms, which take advantages of source aggregation and fit well with multicast routing. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye responds effectively to dynamic multicast requests. Itoffers rapid routing decisions, e.g., making routing decisions in under 5ms on a tested topology, and reduces path latency variation by up to 89.5%, with less than a 10% increase in bandwidth consumption compared to the offline theoretical minimum. Qing Li 0006, Lie Lu, Dan Zhao 0003, Zeyu Luan, Yuan Yang 0001, Yong Jiang 0001, Jingpu Duan, Ruobin Zheng, Shaoteng Liu, Dingding Chen |
IEEE Trans. Netw. | 7 |
| 2025 | Efficient Data Center Network Monitoring and Troubleshooting With LMon: Leveraging ECMP Hashing Linearity and Lightweight ProbingabstractNetwork performance monitoring and troubleshooting are crucial yet challenging tasks in datacenter management. Despite the numerous solutions that have been proposed in recent years, their efforts are often hindered by high costs and unreliable failure localization, making it difficult to deploy them in real-world environments. In this paper, we presentLMon, a highly reliable and efficient system for monitoring and troubleshooting in datacenter networks. LMon utilizes the characteristic of ECMP hashing linearity to control probe packet routing, enabling the monitoring of targeted paths without any modification of underlying protocols and devices. Additionally, LMon leverages a lightweight probing technique to reduce monitoring overhead, as well as integrates the improved LASSO regression and hypothesis testing for higher accuracy and faster processing in link failure localization. We evaluate the performance of LMon in our testing environment. Compared to the monitoring system Pingmesh, LMon generates only one-third probes while maintaining 99% accuracy and 1% false negatives. Qinglin Xun, Weichao Li 0001, Jianer Zhou, Jingpu Duan, Yi Wang 0004, Xiaofeng Tao 0001, Jinbei Zhang |
IEEE Trans. Netw. | 4 |
| 2024 | Cost-Driven Auction Mechanism for SFC Allocation in Space-Air-Ground Integrated NetworkabstractService Function Chaining (SFC) is a fundamental technology for resource management in Space-Air-Ground Integrated Network (SAGIN). The heterogeneity and dynamic nature of network resources in SAGIN increase the complexity of SFC-based resource allocation. However, existing work rarely considers the issue of economically efficient resource allocation under cost constraints. To solve the issue, this study explicitly analyzes resource characteristics and establishes a cost-driven online auction mechanism for SFC allocation. First, we formulate a novel SFC allocation problem for SAGIN, aiming at maximizing social welfare while considering operational costs. We then adopt Fenchel duality to convert the primal problem into a dual problem and design a payment strategy that facilitates the dynamic updating of resource marginal prices. Our algorithm achieves optimal SFC allocation and pricing outcomes while guaranteeing bidding truthfulness, individual rationality, and polynomial-time complexity. Finally, we validate the online auction’s competitiveness through rigorous theoretical analysis and simulation studies driven by real-world traces. Yali Lyu, Xiaoxi Zhang 0001, Jingpu Duan, Xu Chen 0004 |
HPCC | 4 |
| 2024 | Rosevin: Employing Resource- and Rate-Adaptive Edge Super-Resolution for Video StreamingabstractToday’s video streaming service providers have exploited cloud-edge collaborative networks for geo-distributed video delivery. The existing content delivery network (CDN) scheduling and adaptive bitrate algorithms may not fully utilize edge resources or lack a global control to optimize resource sharing. The emerging super-resolution (SR) approach can unleash the potential of leveraging computation resources to compensate for bandwidth consumption, by producing high-quality videos from low-resolution contents. Yet the uncertain SR resource sensitivity and its interplay with bitrate adaptation are underexplored. In this work, we propose Rosevin, the first resource scheduler that jointly decides the bitrates and fine-grained resource allocation to perform SR at the edge, which can learn to optimize the long-term QoE for distributed end users. To handle the time-varying and complex space of decisions as well as a non-smooth objective function, Rosevin realizes a novel online combinatorial learning algorithm, which nicely integrates convex optimization theories and online learning techniques. In addition to theoretically analyzing its performance, we implement an SR-assisted video streaming prototype of Rosevin and demonstrate its advantages over several video delivery benchmarks. Xiaoxi Zhang 0001, Longhao Zou, Jingpu Duan, Chuan Wu 0001, Yali Xue, Zuozhou Chen, Xu Chen 0004 |
INFOCOM | 4 |
| 2024 | Can You Do Both? Balancing Order Serving and Crowdsensing for Ride-Hailing VehiclesabstractGiven the high mobility and sensor-carrying capability, vehicle crowdsensing (VCS) has become a significant part of urban crowdsensing tasks in the development of smart cities. Ride-hailing vehicles, which are widely distributed in cities, can be a powerful tool for carrying out VCS. However, dispatching the vehicles to jointly benefit VCS and order serving is challenging, as the goals of these two tasks may not be consistent or even conflict. The distribution of ride orders and the distribution of point-of-interests (PoIs) may not coincide in time and geography. In addition, these orders and data PoIs have distinct forms of timeliness: prolonged waiting makes orders invalid and data with a larger age-of-information (AoI) has lower utility. We propose an online framework by extending multi-agent reinforcement learning (MARL) with careful augmentation to optimize the profit of order-serving and the data utility of crowdsensing. A new quality-of-service (QoS) metric is designed to characterize the utility of the two joint tasks, and formal mathematical modeling drives our MARL design. In particular, we integrated graph neural networks (GNN) to enhance state representations and capture the graph-structured dependencies among vehicles. We developed a simulator and conducted extensive experiments utilizing the New York City Taxi dataset. Experimental results demonstrate the advantage of our method in QoS improvement. Bo Rao, Xiaoxi Zhang 0001, Tianxiang Zhu, Yufei You, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004 |
IWQoS | 6 |
| 2024 | MPVSched: Multipath Transmissions and Video Frame Scheduling for Content Delivery NetworksabstractWith the widespread adoption of video streaming applications, effective video delivery solutions are crucial for providing seamless user experiences. Recent studies have revealed that multipath transmissions are beneficial to video streaming applications, given their potential of better load balancing and fault tolerance, relative to single path settings. However, the necessity of cross-layer co-design of multipath routing and video frame scheduling is overlooked. This work identifies that preset or path-oblivious frame scheduling used in existing works cannot adapt to network dynamics and fail to enhance the quality of experiences (QoE) in multipath transmissions. Therefore, we propose MPVSched, a novel framework that unifies the design of multipath routing and application-layer frame scheduling, with a particular focus on improving the rebuffer rate for short video delivery. At the network layer, we propose to use network-assisted routing that selects the optimal paths for each video transmission, with per-hop per-frame latency prediction. We implement an end-to-end QUIC-based video streaming system by integrating our routing strategy and application-layer frame scheduler, which effectively improves streaming efficiency and prevents user-side freezes. Our testbed experiments with real-world short video request traces demonstrate that MPVSched can achieve reductions of up to 28.58% in rebuffer ratio, compared to representative baseline methods. Xiaoxi Zhang 0001, Jingpu Duan, Chuan Wu 0001, Jinhang Zuo, Xuan Zeng 0002, Yubing Qiu, Xu Chen 0004 |
NAS | 3 |
| 2024 | rpkt: A Generic, Safe, and Efficient Userspace Packet Processing Library in Rust
Yupeng Xiao, Yaxuan Chen, Jingpu Duan, Xiaoxi Zhang 0001, Weichao Li 0001, Xiaofeng Tao 0001 |
NPC (2) | 4 |
| 2024 | MATE: When multi-agent Deep Reinforcement Learning meets Traffic Engineering in multi-domain networks
Zeyu Luan, Qing Li 0006, Yong Jiang 0001, Jingpu Duan, Ruobin Zheng, Dingding Chen, Shaoteng Liu |
Comput. Networks | 4 |
| 2024 | Stochastic Long-Term Energy Optimization in Digital Twin-Assisted Heterogeneous Edge NetworksabstractMobile edge computing (MEC) and digital twin (DT) technologies have been recognized as key enabling factors for the next generation of industrial Internet of Things (IoT) applications. In existing works, DT-assisted edge network resource optimization solutions mostly focus on short-term performance optimization, and long-term resource optimization has not been well studied. Thus, this paper introduces a digital twin-assisted heterogeneous edge network (DTHEN), aiming to minimize long-term energy consumption by jointly optimizing transmit power and computing resource. To solve the stochastic optimization problem, we propose a long-term queue-aware energy minimization (LQEM) scheme for joint communication and computing resource management. The proposed scheme uses Lyapunov optimization to transform the original problem with long-term time constraints into a deterministic upper bound problem for each time slot, decouples it into three independent sub-problems, and solves each sub-problem separately. We then theoretically prove the asymptotic optimality of the LQEM scheme and the tradeoff between system energy consumption and task queue backlog. Finally, experimental results verify the performance analysis of the LQEM scheme, demonstrating its superiority over several benchmark schemes, and reveal the impact of various parameters on the system. Yingsheng Peng, Jingpu Duan, Jinbei Zhang, Weichao Li 0001, Yong Liu 0005, Fuli Jiang |
IEEE J. Sel. Areas Commun. | 2 |
| 2024 | Online Management for Edge-Cloud Collaborative Continuous Learning: A Two-Timescale ApproachabstractDeep learning (DL) powered real-time applications usually need continuous training using data streams generated over time and across different geographical locations. Enabling data offloading among computation nodes through model training is promising to mitigate the problem that devices generating large datasets may have low computation capability. However, offloading can compromise model convergence and incur communication costs, which must be balanced with the long-term cost spent on computation and model synchronization. Therefore, this paper proposes EdgeC3, a novel framework that can optimize the frequency of model aggregation and dynamic offloading for continuously generated data streams, navigating the trade-off between long-term accuracy and cost. We first provide a new error bound to capture the impacts of data dynamics that are varying over time and heterogeneous across devices, as well as quantifying varied data heterogeneity between local models and the global one. Based on the bound, we design a two-timescale online optimization framework. We periodically learn the synchronization frequency to adapt with uncertain future offloading and network changes. In the finer timescale, we manage online offloading by extending Lyapunov optimization techniques to handle an unconventional setting, where our long-term global constraint can have abruptly changed aggregation frequencies that are decided in the longer timescale. Finally, we theoretically prove the convergence of EdgeC3 by integrating the coupled effects of our two-timescale decisions, and we demonstrate its advantage through extensive experiments performing distributed DL training for different domains. Shaohui Lin, Xiaoxi Zhang 0001, Yupeng Li 0001, Carlee Joe-Wong, Jingpu Duan, Dongxiao Yu, Yu Wu 0010, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | DYNAMITE: Dynamic Interplay of Mini-Batch Size and Aggregation Frequency for Federated Learning With Static and Streaming DatasetsabstractFederated Learning (FL) is a distributed learning paradigm that can coordinate heterogeneous edge devices to perform model training without sharing private data. While prior works have focused on analyzing FL convergence with respect to hyperparameters like batch size and aggregation frequency, the joint effects of adjusting these parameters on model performance, training time, and resource consumption have been overlooked, especially when facing dynamic data streams and network characteristics. This paper introduces novel analytical models and optimization algorithms that leverage the interplay between batch size and aggregation frequency to navigate the trade-offs among convergence, cost, and completion time for dynamic FL training. We establish a new convergence bound for training error considering heterogeneous datasets across devices and derive closed-form solutions for co-optimized batch size and aggregation frequency that are consistent across all devices. Additionally, we design an efficient algorithm for assigning different batch configurations across devices, improving model accuracy and addressing the heterogeneity of both data and system characteristics. Further, we propose an adaptive control algorithm that dynamically estimates network states, efficiently samples appropriate data batches, and effectively adjusts batch sizes and aggregation frequency on the fly. Extensive experiments demonstrate the superiority of our offline optimal solutions and online adaptive algorithm. Xiaoxi Zhang 0001, Jingpu Duan, Carlee Joe-Wong, Zhi Zhou 0006, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | An Offline-Transfer-Online Framework for Cloud-Edge Collaborative Distributed Reinforcement LearningabstractRecent advances in deep reinforcement learning (DRL) have made it possible to train various powerful agents to perform complex tasks in real-time environments. With the next-generation communication technologies, making cloud-edge collaborative artificial intelligence service with evolved DRL agents can be a significant scenario. However, agents with different algorithms and architectures in the same DRL scenario may not be compatible, and training them is either time-consuming or resource-demanding. In this paper, we design a novel cloud-edge collaborative DRL training framework, named Offline-Transfer-Online, which is a new approach that can speed up the convergence of online DRL agents at the edge by interacting with offline agents in the cloud, with the minimum data interchanged and without relying on high-quality offline datasets. Therein, we propose a novel algorithm-independent knowledge distillation algorithm for online RL agents, by leveraging pre-trained models and the interface between agents and the environment to transfer distilled knowledge among multiple heterogeneous agents efficiently. Extensive experiments show that our algorithm can accelerate the convergence of various online agents in a double to decuple speed, with comparable reward achieved in different environments. Tianyu Zeng, Xiaoxi Zhang 0001, Jingpu Duan, Chao Yu 0004, Chuan Wu 0001, Xu Chen 0004 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | An Online Control Approach of Collaborative Federated Learning with Constrained ResourcesabstractNo abstract available. Shaohui Lin, Xiaoxi Zhang 0001, Yupeng Li 0001, Carlee Joe-Wong, Jingpu Duan, Xu Chen 0004 |
APNet | 5 |
| 2023 | TapFinger: Task Placement and Fine-Grained Resource Allocation for Edge Machine LearningabstractMachine learning (ML) tasks are one of the major workloads in today's edge computing networks. Existing edge-cloud schedulers allocate the requested amounts of resources to each task, falling short of best utilizing the limited edge resources flexibly for ML task performance optimization. This paper proposes TapFinger, a distributed scheduler that minimizes the total completion time of ML tasks in a multi-cluster edge network, through co-optimizing task placement and fine-grained multi-resource allocation. To learn the tasks' uncertain resource sensitivity and enable distributed online scheduling, we adopt multi-agent reinforcement learning (MARL), and propose several techniques to make it efficient for our ML-task resource allocation. First, TapFinger uses a heterogeneous graph attention network as the MARL backbone to abstract inter-related state features into more learnable environmental patterns. Second, the actor network is augmented through a tailored task selection phase, which decomposes the actions and encodes the optimization constraints. Third, to mitigate decision conflicts among agents, we novelly combine Bayes' theorem and masking schemes to facilitate our MARL model training. Extensive experiments using synthetic and test-bed ML task traces show that TapFinger can achieve up to 28.6% reduction in the average task completion time and improve resource efficiency as compared to state-of-the- art resource schedulers. Tianyu Zeng, Xiaoxi Zhang 0001, Jingpu Duan, Chuan Wu 0001 |
INFOCOM | 4 |
| 2023 | AdaCoOpt: Leverage the Interplay of Batch Size and Aggregation Frequency for Federated LearningabstractFederated Learning (FL) is a distributed learning paradigm that can coordinate heterogeneous edge devices to perform model training without sharing private raw data. Many prior works have analyzed the FL convergence with respect to important hyperparameters, including batch size and aggregation frequency. However, adjusting the batch size and the number of local updates can affect the model performance, training time, and the cost of consuming computation and communication resources, in different and perhaps complex forms. Their joint effects have been overlooked and should be exploited to achieve accurate models with controllable operational expenditure. This paper proposes novel analytical models and optimization algorithms that leverage the interplay of batch size and aggregation frequency to navigate the trade-offs among convergence, cost, and completion time for FL. We first obtain a new convergence bound of the training error under heterogeneous training datasets across devices. Based on this bound, we derive closed-form solutions of a co-optimized batch size and aggregation frequency, a single configuration for all the devices. We then design an efficient exact algorithm for assigning different batch configurations across devices that can further improve the model accuracy to address the heterogeneity of both data and system characteristics. Further, we propose an adaptive control algorithm to dynamically adjust the solutions with estimated network states. Extensive experiments demonstrate the superiority of our offline optimal solutions and online adaptive algorithm. Xiaoxi Zhang 0001, Jingpu Duan, Carlee Joe-Wong, Zhi Zhou 0006, Xu Chen 0004 |
IWQoS | 3 |
| 2023 | Enabling Reliable and Efficient Performance Monitoring and Troubleshooting in Datacenter NetworksabstractNetwork performance monitoring and troubleshooting is a crucial but challenging task in datacenter management. Despite the numerous solutions that have been proposed in recent years, their efforts are often hindered by high costs and unreliable fault localization, making it difficult to deploy them in real-world environments. In this paper, we present LMon, a highly reliable and efficient system for monitoring and troubleshooting in datacenter networks. LMon utilizes the characteristic of ECMP hashing linearity to control the packet routing without any modification of the underlying protocols. Additionally, LMon leverages a lightweight probing technique to reduce monitoring overhead. Furthermore, the system integrates improved LASSO regression and statistical hypothesis testing for higher accuracy and faster processing in link failure localization. The effectiveness of LMon is demonstrated through its implementation and evaluation in ns-3 simulation. The results validate the reliability and efficiency of the system, making it a promising option for ensuring long-term network maintenance in datacenters. Qinglin Xun, Weichao Li 0001, Haorui Guo, Qianyi Huang, Jianer Zhou, Jingpu Duan, Yi Wang 0004, Jinbei Zhang |
IWQoS | 6 |
| 2023 | A Budget-aware Incentive Mechanism for Vehicle-to-Grid via Reinforcement LearningabstractWith the increasing penetration of renewable energy and electric vehicles (EVs), the behavior of EVs' charging and discharging has shown great impact on the Micro Grid power load, motivating the development of Vehicle-to-Grid (V2G) technologies. However, the V2G market is still in its infancy, due to insufficient understanding of EV users' willingness and concerns. While many studies consider direct EV control, it's more realistic to indirectly affect users' behavior through monetary incentives. For better implementation flexibility, we advocate to display at charging piles strategically chosen incentives that are combined with electricity prices. Technically, this is the first model-free learning algorithm that can optimize incentives under unknown EV user reactions, increase the load control effectiveness and users' quality-of-service (QoS) simultaneously under a long-term incentive budget, and provide theoretical performance guarantees. We first construct a bi-level optimization framework to model the time-dependencies across our solutions. We then integrate primal-dual theories and upper-confidence bounds into reinforcement learning to balance power control and incentive consumption. A dynamic programming based algorithm is also proposed to maximize the aggregate user QoS. Finally, we prove bounded sub-optimality of our learning algorithm through theoretical analysis and conduct trace-driven simulations to demonstrate the advantages of our bi-level framework. Tianxiang Zhu, Xiaoxi Zhang 0001, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004 |
IWQoS | 3 |
| 2023 | RLink: Accelerate On-Device Deep Reinforcement Learning with Inference Knowledge at the EdgeabstractDeep reinforcement learning (DRL) has been a successful paradigm in machine learning that enables solving complex control problems at the human level. However, the sampling and training efficiency of state-of-the-art DRL frameworks can not satisfy the stringent latency and throughput requirements of today’s mobile environments. Existing distributed and offline reinforcement learning algorithms along with the libraries for training acceleration are inherently designed for DRL tasks performed in the cloud rather than on distributed mobile devices, on which the computing resources are highly constrained, heterogeneous, and possibly dynamically changing. With the rise of edge computing and intelligence services, this paper presents RLink, a novel distributed training library to accelerate on-device deep reinforcement learning with inference knowledge at the edge. We leverage knowledge distillation to realize lightweight interaction between our on-device training task and the remote models that can provide inference knowledge. In this way, RLink is designed to be event-driven and agnostic to heterogeneous deep reinforcement learning algorithms and libraries. To tackle the communication bottleneck, a novel asynchronous sampling algorithm is proposed to facilitate real-time training in RLink. Tuned for unstable-connected mobile devices, RLink is robust and efficient by using a semantic-aware communication pipeline for lossless data compression. Extensive experimental results show that, compared with state-of-the-art algorithms and libraries, RLink can accelerate deep reinforcement learning at the edge with up to decuple speedups in convergence and ideal computational performance. Tianyu Zeng, Xiaoxi Zhang 0001, Daipeng Feng, Jingpu Duan, Zhi Zhou 0006, Xu Chen 0004 |
MSN | 4 |
| 2023 | EdgeC3: Online Management for Edge-Cloud Collaborative Continuous LearningabstractDeep learning (DL) powered real-time applications usually need continuous training using data streams generated geographically. Enabling data offloading among computation nodes through model training is promising to mitigate the problem that devices generating large datasets may have low computation capability. However, offloading can compromise model convergence and incur communication costs, which must be balanced with the cost spent on computation and model synchronization. Therefore, this paper proposes EdgeC3, a novel framework that can optimize the frequency of model aggregation and dynamic offloading for continuously generated data streams, navigating the trade-off between long-term accuracy and cost. We first provide a new error bound to capture the impacts of data dynamics that are varying over time and heterogeneous across devices. Based on the bound, we design a two-timescale online optimization framework. We periodically learn the synchronization frequency to adapt with uncertain future offloading and network changes. In the finer timescale, we manage online offloading by extending Lyapunov optimization techniques to handle an unconventional setting, where our long-term global constraint can have abruptly changed aggregation frequencies that are decided in the longer timescale. Finally, we theoretically prove the convergence of EdgeC3 by integrating the coupled effects of our two-timescale decisions, and we demonstrate its advantage through extensive experiments. Shaohui Lin, Xiaoxi Zhang 0001, Yupeng Li 0001, Carlee Joe-Wong, Jingpu Duan, Xu Chen 0004 |
SECON | 5 |
| 2023 | Cable: A framework for accelerating 5G UPF based on eBPF
Jianer Zhou, Zengxie Ma, Weijian Tu, Xinyi Qiu, Jingpu Duan, Zhenyu Li 0001, Qing Li 0006, Xinyi Zhang 0004, Weichao Li 0001 |
Comput. Networks | 5 |
| 2023 | A Machine Learning-Based Framework for Dynamic Selection of Congestion Control AlgorithmsabstractMost congestion control algorithms (CCAs) are designed for specific network environments. As such, there is no known algorithm that achieves uniformly good performance in all scenarios for all flows. Rather than devising a one-size-fits-all algorithm (which is a likely impossible task), we propose a system to dynamically switch between the most suitable CCAs for specific flows in specific environments. This raises a number of challenges, which we address through the design and implementation of Antelope, a system that can dynamically reconfigure the stack to use the most suitable CCA for individual flows. We build a machine learning model to learn which algorithm works best for individual conditions and implement kernel-level support for dynamically switching between CCAs. The framework also takes application requirements of performance into consideration to fine-tune the selection based on application-layer needs. Moreover, to reduce the overhead introduced by machine learning on individual front-end servers, we (optionally) implement the CCA selection process in the cloud, which allows the share of models and the selection among front-end servers. We have implemented Antelope in Linux, and evaluated it in both emulated and production networks. The results demonstrate the effectiveness of Antelope via dynamic adjusting the CCAs for individual flows. Specifically, Antelope achieves an average 16% improvement in throughput compared with BBR, and an average 19% improvement in throughput and 10% reduction in delay compared with CUBIC. Jianer Zhou, Xinyi Qiu, Zhenyu Li 0001, Qing Li 0006, Gareth Tyson, Jingpu Duan, Yi Wang 0004, Qinghua Wu 0004 |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | Task Placement and Resource Allocation for Edge Machine Learning: A GNN-Based Multi-Agent Reinforcement Learning ParadigmabstractMachine learning (ML) tasks are one of the major workloads in today's edge computing networks. Existing edge-cloud schedulers allocate the requested amounts of resources to each task, falling short of best utilizing the limited edge resources for ML tasks. This paper proposesTapFinger, a distributed scheduler for edge clusters that minimizes the total completion time of ML tasks through co-optimizing task placement and fine-grained multi-resource allocation. To learn the tasks’ uncertain resource sensitivity and enable distributed scheduling, we adopt multi-agent reinforcement learning (MARL) and propose several techniques to make it efficient, including a heterogeneous graph attention network as the MARL backbone, a tailored task selection phase in the actor network, and the integration of Bayes’ theorem and masking schemes. We first implement asingle-task schedulingversion, which schedules at most one task each time. Then we generalize to themulti-task schedulingcase, in which a sequence of tasks is scheduled simultaneously. Our design can mitigate the expanded decision space and yield fast convergence to optimal scheduling solutions. Extensive experiments using synthetic and test-bed ML task traces show thatTapFingercan achieve up to 54.9% reduction in the average task completion time and improve resource efficiency as compared to state-of-the-art schedulers. Xiaoxi Zhang 0001, Tianyu Zeng, Jingpu Duan, Chuan Wu 0001, Di Wu 0001, Xu Chen 0004 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Mousika: Enable General In-Network Intelligence in Programmable Switches by Knowledge DistillationabstractGiven the power efficiency and Tbps throughput of packet processing, several works are proposed to offload the decision tree (DT) to programmable switches, i.e., in-network intelligence. Though the DT is suitable for the switches’ match-action paradigm, it has several limitations. E.g., its range match rules may not be supported well due to the hardware diversity; and its implementation also consumes lots of switch resources (e.g., stages and memory). Moreover, as learning algorithms (particularly deep learning) have shown their superior performance, some more complicated learning models are emerging for networking. However, their high computational complexity and large storage requirement are cause challenges in the deployment on switches. Therefore, we propose Mousika, an in-network intelligence framework that addresses these drawbacks successfully. First, we modify the DT to the Binary Decision Tree (BDT). Compared with the DT, our BDT supports faster training, generates fewer rules, and satisfies switch constraints better. Second, we introduce the teacher-student knowledge distillation in Mousika, which enables the general translation from other learning models to the BDT. Through the translation, we can not only utilize the super learning capabilities of complicated models, but also avoid the computation/memory constraints when deploying them on switches directly for line-speed processing. Guorui Xie, Qing Li 0006, Yutao Dong, Guanglin Duan, Yong Jiang 0001, Jingpu Duan |
INFOCOM | 6 |
| 2022 | Pricing-based resource allocation in three-tier edge computing for social welfare maximization
Yupeng Li 0001, Mengjia Xia, Jingpu Duan, Yang Chen 0001 |
Comput. Networks | 3 |
| 2022 | Weighted NSFIB Aggregation With Generalized Next Hop of Strict Partial OrderabstractThe size of the global routing table has been growing at an alarming rate. With the exhaustion of IPv4 addresses and the gradual deployment of IPv6 networks, the growth rate will continue to accelerate in the future. Although modern high performance routers provide enough line-card memory, Internet Service Providers (ISPs) cannot afford to upgrade their routers as fast as the growth of global routing tables. In this paper, we propose an algorithm to calculate the generalized next hops with strict partial order (GSPO next hops) of a network prefix and use them for the aggregation of the Nexthop-Selectable Forwarding Information Base (NSFIB). Since the existing NSFIB aggregation algorithm may introduce path stretch, we also propose a weighted NSFIB aggregation algorithm to effectively control path stretch under a given upper limit of the FIB size. Experiment results show that our algorithm can shrink the FIB size by at most 97% under IPv4 networks, and at most 95% under IPv6 networks. Under a given upper limit of the FIB size, our algorithm can reduce the path stretch by at least 22%. Qing Li 0006, Yichao Wu, Jingpu Duan, Jiahai Yang 0001, Yong Jiang 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | Antelope: A Framework for Dynamic Selection of Congestion Control AlgorithmsabstractMost congestion control mechanisms are designed for specific network environments. Hence, there is no known algorithm that achieves uniformly good performance in all scenarios for all flows. Rather than devising such a one-size-fits-all algorithm, we propose a system to dynamically switch between the most suitable congestion control mechanisms for specific flows in specific environments. This raises a number of challenges, which we address through the design and implementation of Antelope, a system that can dynamically reconfigure to use the most suitable congestion control mechanism for an individual flow. We build a machine learning approach to learn which algorithm works best for individual conditions and implement kernel-level support for dynamically adjusting congestion control algorithms. We have implemented Antelope in Linux, and evaluated it in both emulated and production networks. We show that in WAN, DCN, and cellular networks, Antelope achieves an average 16% improvement in throughput compared with BBR; compared with Cubic, Antelope achieves an average 19% improvement in throughput and 10% reduction in delay. Jianer Zhou, Xinyi Qiu, Zhenyu Li 0001, Gareth Tyson, Qing Li 0006, Jingpu Duan, Yi Wang 0004 |
ICNP | 6 |
| 2021 | FlexNF: Flexible Network Function Orchestration on the Programmable Data PlaneabstractRecently, Programmable Data Plane (PDP) has been leveraged to offload Network Functions (NFs). Due to its high processing capabilities, programmable data plane can improve the performance of NFs to more than one order of magnitude. However, the coarse-grained NF orchestration granularity on the PDP makes it hard to fulfill the dynamic service chain demands. In this paper, we propose the Flexible Network Function (FlexNF) Deployment on the programmable data plane. We first design an NF Selection Framework which leverages labels and the pipeline re-enter operation to support Selective Serving Mechanism for flexible NF orchestration. We then design a two-stage service path construction algorithm to provide on-path service based on SSM with load balancing taken into account. We implement 7 types of real network functions in the commodity P4 switch, based which we construct the comprehensive experiments. The results show that FlexNF can reduce the traffic routing delay by about 42.6% while increasing service chain acceptance rate by 5 times compared with current solutions. Qing Li 0006, Jingpu Duan, Yong Jiang 0001 |
IWQoS | 3 |
| 2020 | A proactive auto-scaling scheme with latency guarantees for multi-tenant NFV cloud
Guangwu Hu, Qing Li 0006, Shuo Ai, Jingpu Duan, Yu Wu 0010 |
Comput. Networks | 5 |
| 2019 | FlowShader: a Generalized Framework for GPU-accelerated VNF Flow ProcessingabstractGPU acceleration has been widely investigated for packet processing in virtual network functions (NFs), but not for L7 flow-processing NFs. In L7 NFs, reassembled TCP messages of the same flow should be processed in order in the same processing thread, and the uneven sizes among flows pose a major challenge for full realization of GPU's parallel computation power. To exploit GPUs for L7 NF processing, this paper presents FlowShader, a GPU acceleration framework to achieve both high generality and throughput even under skewed flow size distributions. We carefully design an efficient scheduling algorithm that fully exploits available GPU and CPU capacities; in particular, we dispatch large flows which seriously break up the size balance to CPU and the rest of flows to GPU. Furthermore, FlowShader allows similar NF logic (as CPU-based NFs) to run on individual threads in a GPU, which is more generalized and easy to take on as compared to redesigning an NF for operation parallelism on GPU. We implemented a number of L7 flow processing NFs based on FlowShader. Evaluations are conducted under both synthetic and real-world traffic traces and results show that the throughput achieved by FlowShader is up to 6x that of the CPU-only baseline and 3x of the GPU-only design. Xiaodong Yi 0001, Jingpu Duan, Wei Bai 0001, Chuan Wu 0001, Yongqiang Xiong, Dongsu Han |
ICNP | 3 |
| 2019 | NetStar: A Future/Promise Framework for Asynchronous Network FunctionsabstractNetwork functions (NFs) are more than simple packet processors that apply various transformations to the packet content. Modern NFs often resort to various external services to achieve their purposes, e.g., storing flow states in an external storage or looking up a DNS. Working with external services is usually implemented using callback-based asynchronous programming, which is complex and error-prone. This paper proposes NetStar, a new NF programming framework that brings the future/promise abstraction to the NF dataplane for flow processing. NetStar simplifies asynchronous NF programming via a carefully designed async-flow interface that exploits the future/promise paradigm by chaining multiple continuation functions for asynchronous operations handling. The programs implemented using the NetStar framework mimic simple synchronous programming but are able to achieve full flow processing asynchrony. We have used NetStar to implement a number of representative NFs. Our experience and evaluation results show that NetStar can effectively simplify asynchronous NF programming by substantially reducing the lines of code, while still approaching line-rate packet processing speeds. Jingpu Duan, Xiaodong Yi 0001, Chuan Wu 0001, Franck Le |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | NFVactor: A Resilient NFV System Using the Distributed Actor ModelabstractResilience functionality, including failure resilience and flow migration, is of pivotal importance in practical network function virtualization (NFV) systems. However, existing failure recovery procedures incur high packet processing delay due to heavyweight process checkpointing, while flow migration has poor performance due to centralized control. This paper proposes NFVactor, a novel NFV system that aims to provide lightweight failure resilience and high-performance flow migration. NFVactorenables these by using actor model to provide a per-flow execution environment, so that each flow can replicate and migrate itself with improved parallelism, while the efficiency of the actor model is guaranteed by a carefully designed runtime system. Moreover, NFVactorachieves transparent resilience: once a new network function (NF) is implemented for NFVactor, the NF automatically acquires resilience support. Our evaluation result shows that NFVactorachieves 10-Gbps packet processing, flow migration completion time that is 144 times faster than the existing system, and packet processing delay stabilized at around 20 μs during replication. Jingpu Duan, Xiaodong Yi 0001, Shixiong Zhao, Chuan Wu 0001, Heming Cui, Franck Le |
IEEE J. Sel. Areas Commun. | 1 |
| 2017 | GPUNFV: a GPU-Accelerated NFV SystemabstractThis paper presents GPUNFV, a high-performance NFV system providing flow-level micro services for stateful service chains with Graphics Processing Unit (GPU) acceleration. GPUNFV exploits the massively-parallel processing power of GPU to maximize the throughput of the NFV system. Combined with the customized flow handler, GPUNFV achieves a much better throughput than the existing NFV systems. With a carefully designed GPU-based virtualized network function framework, GPUNFV is able to efficiently support both stateful and stateless network functions. We have implemented a number of GPU-based network functions and a preliminary GPUNFV system to demonstrate the lexibility and potential of our design. Xiaodong Yi 0001, Jingpu Duan, Chuan Wu 0001 |
APNet | 2 |
| 2017 | Dynamic Scaling of Virtualized, Distributed Service Chains: A Case Study of IMSabstractThe emerging paradigm of network function virtualization advocates deploying virtualized network functions (VNFs) on standard virtualization platforms for significant cost reduction and management flexibility. There have been system designs for managing dynamic deployment and scaling of VNF service chains within one cloud datacenter. Many real-world network services involve geo-distributed service chains, with prominent examples of mobile core networks and IP multimedia subsystems (IMSs)). Virtualizing these service chains requires efficient coordination of dynamic VNF deployment across geo-distributed data centers, calling for a new management system. This paper designs a dynamic scaling system for geo-distributed VNF service chains, using the case of an IMS. IMSs are widely used subsystems for delivering multimedia services among mobile users in a 3G/4G network, whose virtualization has been broadly advocated in the industry for reducing cost, improving network usage efficiency and enabling dynamic network topology reconfiguration for performance optimization. Our scaling system design caters to key control-plane and data-plane service chains in an IMS, combining proactive and reactive approaches for timely, cost-effective scaling of the service chains. The design principles are applicable to scaling of other systems with multiple related service chains. We evaluate our system using real-world experiments on both an emulation platform and a geo-distributed public cloud. Jingpu Duan, Chuan Wu 0001, Franck Le, Alex X. Liu, Yanghua Peng |
IEEE J. Sel. Areas Commun. | 1 |
| 2015 | Responsive multipath TCP in SDN-based datacentersabstractA basic need in datacenter networks is to provide high throughput for large flows such as the massive shuffle traffic flows in a MapReduce application. Multipath TCP (MPTCP) has been investigated as an effective approach toward this goal, by spreading one TCP flow onto multiple paths. However, the current MPTCP implementation has two major limitations: (1) a fixed number of subflows are used without reacting to the actual traffic condition; (2) the routing of subflows of a multipath TCP connection relies heavily on the ECMP-based random hashing. The former may lead to a waste of both the server and network resources, while the latter can cause throughput degradation when multiple subflows collide on the same path. This paper proposes a responsive MPTCP system to resolve the two limitations simultaneously. Our system employs a centralized controller for intelligent subflow route calculation and a monitor running on each server for actively adjusting the number of subflows. Working in synergy, the two modules enable MPTCP flows to respond to the traffic conditions and pursue high throughput on the fly, at very low computation and messaging overhead. NS3-based experiments show that our system achieves satisfactory throughput with less resource overhead, or better throughput at similar amounts of overhead, as compared to common alternatives. Jingpu Duan, Zhi Wang 0001, Chuan Wu 0001 |
ICC | 1 |