Jieling Yu

dblp:343/8680 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0485-8162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism
Boran Sun, Guoyong Jiang, Yuechen Tao, Zhishu Che, Jieling Yu, Shan Chang, Huaxi Gu, Fangming Liu
ICDCS7
2026 Clearing MCP Navigation Fog with Economics-Aware Hierarchical Tool Routing
Jieling Yu, Zhenheng Tang, Ruiting Zhou, Baochun Li, Bo Li 0001
ICDCS1
2026 Heterogeneous Federated Learning Frameworks for Balancing Job Completion Time and Model Accuracy
Ruobei Wang, Ruiting Zhou, Jieling Yu, Bo Li 0001, Yuqing Li 0001
IEEE Trans. Netw.3
2026 Intelligent Frameworks for Minimizing Job Completion Time in Clustered Federated Learning
abstract
Federated Learning (FL) enables potentially a large number of clients to collaboratively train a global model with the coordination of a central cloud server without exposing client raw data. However, the FL model convergence performance, often measured by the job completion time, is hindered by two critical factors: non independent and identically distributed (non-IID) data across clients and the straggler effect. In this work, we propose a clustered FL framework,MCFL, to minimize the job completion time by mitigating the influence of non-IID data and the straggler effect while guaranteeing the FL model convergence performance.MCFLbuilds upon a two-stage operation: i) a clustering algorithm constructs clusters, each containing clients with similar computing and communications capabilities to combat the straggler effect within a cluster; ii) a deep reinforcement learning (DRL) algorithm based on soft actor-critic with discrete actions intelligently selects a subset of clients from each cluster to mitigate the impact of non-IID data, and derives the number of intra-cluster aggregation iterations for each cluster to reduce the straggler effect among clusters. We next proposeA-MCFL, a semi-asynchronous clustered framework based on context-aware Multi-Armed Bandit (MAB) client selection algorithm to further reduce job completion time. Extensive testbed experiments are conducted under various configurations to verify the efficacy ofMCFLandA-MCFL. The results show thatMCFLcan reduce the job completion time by up to 70% compared with three state-of-the-art FL frameworks. In addition,A-MCFLhas a further performance improvement overMCFL, achieving an average 64% reduction in job completion time.
Jieling Yu, Ruiting Zhou, Bo Li 0001
IEEE Trans. Netw.1
2024 Efficient Online DNN Inference with Continuous Learning in Edge Computing
abstract
Compressed edge DNN models usually experience decreasing model accuracy when performing inference due to data drift. To maintain the inference accuracy, retraining models with continuous learning is usually employed in the edge. However, online edge DNN inference with continuous learning faces new challenges. First, introducing retraining jobs leads to resource competition with the existing edge inference tasks, which will affect the inference latency. Second, retraining jobs and inference tasks exhibit significant differences in workload and latency requirements. These two jobs cannot adopt the same scheduling policy. To overcome the challenges, we propose an Online scheduling algorithm for INference with Continuous learning (OINC). OINC minimizes the weighted sum of the latency of inference tasks and the completion time of retraining jobs with limited edge resources, while ensuring the satisfaction of the inference task’s service level objective (SLO) and meeting the deadlines of retraining jobs. OINC first reserves a portion of resources to complete all current inference tasks and allocates the remaining resources to retraining jobs. Subsequently, based on the reserved resource ratio, OINC invokes two sub-algorithms to select edges and allocate resources for each inference task and retraining job respectively. Compared with six state-of-the-art algorithms, OINC can reduce the weighted sum by up to 23.7%, and increase the success rate by up to 35.6%.
Ruiting Zhou, Lei Jiao 0002, Ziyi Han, Jieling Yu
IWQoS5
2023 ASFL: Adaptive Semi-asynchronous Federated Learning for Balancing Model Accuracy and Total Latency in Mobile Edge Networks
abstract
Federated learning (FL) is a new paradigm for privacy-preserving learning. This is particularly appealing in the mobile edge network (MEN), in which devices collectively train a global model with their own set of data. It is, however, routinely difficult for FL algorithms to satisfy different training task preferences in terms of the total latency and model accuracy due to a number of factors including the straggler effect, data heterogeneity, communication bottleneck and device mobility. To this end, we propose an Adaptive Semi-asynchronous Federated Learning (ASFL) framework, which adaptively balances the total latency and model accuracy according to the task preferences in MEN. Specifically, ASFL conducts a two-stage operation: i) Device selection stage. Each global round selects a set of devices that can maximize the model accuracy to eliminate data heterogeneity and communication bottlenecks; ii) Training stage. We first define a latency-accuracy objective value to model the balance between the latency and accuracy. Then in each global round, we use a deep reinforcement learning (DRL) algorithm based on soft actor-critic with discrete actions to intelligently derive the number of picked devices (i.e., participants in the current global aggregation) and the lag tolerance at each global round to maximize the latency-accuracy objective value. Extensive experiments show that ASFL can improve the latency-accuracy objective value by up to 94% compared with three state-of-the-art FL frameworks.
Jieling Yu, Ruiting Zhou, Chen Chen 0067, Bo Li 0001, Fang Dong 0001
ICPP1
2023 A Reinforcement Learning Approach for Minimizing Job Completion Time in Clustered Federated Learning
abstract
Federated Learning (FL) enables potentially a large number of clients to collaboratively train a global model with the coordination of a central cloud server without exposing client raw data. However, the FL model convergence performance, often measured by the job completion time, is hindered by two critical factors: non independent and identically distributed (non-IID) data across clients and the straggler effect. In this work, we propose a clustered FL framework, MCFL, to minimize the job completion time by mitigating the influence of non-IID data and the straggler effect while guaranteeing the FL model convergence performance. MCFL builds upon a two-stage operation: i) a clustering algorithm constructs clusters, each containing clients with similar computing and communications capabilities to combat the straggler effect within a cluster; ii) a deep reinforcement learning (DRL) algorithm based on soft actor-critic with discrete actions intelligently selects a subset of clients from each cluster to mitigate the impact of non-IID data, and derives the number of intra-cluster aggregation iterations for each cluster to reduce the straggler effect among clusters. Extensive testbed experiments are conducted under various configurations to verify the efficacy of MCFL. The results show that MCFL can reduce the job completion time by up to 70% compared with three state-of-the-art FL frameworks.
Ruiting Zhou, Jieling Yu, Ruobei Wang, Bo Li 0001
INFOCOM2
2023 An Incentive Auction for Heterogeneous Client Selection in Federated Learning
abstract
Federated Learning (FL) is a new distributed machine learning (ML) approach which enables thousands of mobile devices to collaboratively train artificial intelligence (AI) models using local data without compromising user privacy. Although FL represents a promising computing paradigm, such training process can not be fully realized without an appropriate economic mechanism that incentivizes the participation of heterogeneous clients. This work targets social cost minimization, and studies the incentive mechanism design in FL through a procurement auction. Different from existing literature, we consider a practical scenario of FL where clients are selected and scheduled at different global iterations to guarantee the completion of the FL job, and capture the distinct feature of FL that the number of global iterations is determined by the local accuracy of all participants to balance between computation and communication. Our auction framework$A_{FL}$first decomposes the social cost minimization problem into a series of winner determination problems (WDPs) based on the number of global iterations. To solve each WDP,$A_{FL}$invokes a greedy algorithm to determine the winners, and a payment algorithm for computing remuneration to winners. Finally,$A_{FL}$returns the best solution among all WDPs. We carried out theoretical analysis to prove that$A_{FL}$is truthful, individual rational, computationally efficient, and achieves a near-optimal social cost. We further extend our model to consider multiple FL jobs with corresponding budgets and propose another efficient algorithm$A_{FL-M}$to solve the extended problem. We conduct large-scale simulations based on the real-world data and testbed experiments by adopting FL frameworks FAVOR and CoCoA. Simulation and experiment results show that both$A_{FL}$and$A_{FL-M}$can reduce the social cost by up to 55% compared with state-of-the-art algorithms.
Jinlong Pang, Jieling Yu, Ruiting Zhou, John C. S. Lui
IEEE Trans. Mob. Comput.2
2022 Heterogeneous Federated Learning for Balancing Job Completion Time and Model Accuracy
abstract
Federated Learning (FL) is a secure distributed learning paradigm, which enables potentially a large number of devices to collaboratively train a global model based on their local dataset. FL exhibits two distinctive features in job requirement and client participation, where FL jobs may have different training criteria, and clients possess diverse device capabilities and data characteristics. In order to capture such heterogeneities, this paper proposes a new FL framework, Hca, which aims to strike a balance between the job completion time and model accuracy. Specifically, Hca builds upon a number of innovations in the following three phases: i) pre-estimation: we first derive the optimal set of parameters used in training in terms of the number of training rounds, the number of iterations and the number of participating clients in each round; ii) client selection: we design a novel device selection algorithm, which selects the most effective clients for participation based on both client historical contributions and data effectiveness; iii) model aggregation: we improve the classic FedAvg algorithm by integrating the model loss reduction in consecutive rounds as a weighted factor into aggregation computation. To evaluate the performance and effectiveness of Hca, we conduct theoretical analysis and testbed experiments over an FL platform FAVOR. Extensive results show that Hca can improve the job completion time by up to 34% and the model accuracy by up to 9.1%, and can reduce the number of communication rounds required in FL by up to 75% compared with two state-of-the-art FL frameworks.
Ruiting Zhou, Ruobei Wang, Jieling Yu, Bo Li 0001, Yuqing Li 0001
ICPADS3