Yajie Li 0001

dblp:49/5792-1 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-8751-7947ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Ltfc: Loss-Tolerant Flow Control with RDMA Network for Machine Learning Clusters
abstract
Current AI training clusters widely use RoCEv2 to improve the communication efficiency of the interconnect networks across machines. RoCEv2 relies on Priority Flow Control (PFC) to ensure a lossless network. However, PFC brings certain side effects, such as head-of-line blocking, congestion spreading, and deadlock. Numerous studies have been proposed to eliminate the side effects. However, unlike traditional high-performance computing applications, distributed machine learning (DML) is not 100 % loss-intolerant. In light of this observation, this paper proposes a loss-tolerant flow control (LTFC). LTFC does not rely on PFC to ensure a lossless environment, but to control packet loss ratio within the tolerance threshold. Compared to the traditional trigger condition, LTFC reduces the likelihood that the PFC will be triggered. Additionally, we replace RoCEv2's default Go-back-N mechanism with a non-retransmission mechanism to eliminate retransmission latency. We demonstrate the bounded-loss tolerance feature of DML on our testbed and evaluate the performance of LTFC in large-scale simulations. Simulation results show that LTFC reduces the average flow completion time (FCT) by up to 26.9 % and tail FCT by up to 16.9 % compared to existing solutions.
Wei Wang 0116, Qiaojun Hu, Yiyang Li 0009, Yajie Li 0001, Yongli Zhao 0001, Xiaoyu Wang 0017, Jie Zhang 0006
ICC5
2025 Resource Allocation in Flexible-Bandwidth Fine-Grained Optical Transport Networks for Geo-Distributed Machine Learning
abstract
Geo-distributed machine learning (GDML) can facilitate collaborative learning among geographically-dispersed data centers to meet the demands of distributed and privacy-preserving training for large-scale distributed Internet of Things applications. Unfortunately, the efficiency of distributed training tasks heavily depends on synchronized communication between multiple distributed models over bandwidth-limited wide area networks (WANs). The fine-grained Optical Transport Network (fgOTN), thanks to its adjustable bandwidth connections, represents more flexible transmission and has the ability for accurate synchronization across GDML tasks in WANs. However, flexible bandwidth assignment and complex interdependencies among tasks pose significant challenges to resource allocation for GDML in fgOTN. Specifically, flexible bandwidth assignment exacerbates resource competition among task flows, leading to decreased learning efficiency. This paper provides novel resource allocation solutions for GDML in fgOTN. We first formulate this problem as a linear programming aimed at maximizing the completion ratio of GDML tasks. Subsequently, we propose an innovative resource allocation algorithm based on genetic algorithm (GARA) for GDML in fgOTN. GARA considers both task completion and bandwidth adjustment through population generation based on prior knowledge and adaptive mutation based on completion ratio. Simulation analysis demonstrates that GARA effectively prioritizes resource allocation for high-priority tasks to alleviate resource competition, achieving the highest task completion ratio while avoiding excessive network reconfiguration.
Yongli Zhao 0001, Xin Li 0041, Wenhong Liu, Yajie Li 0001, Massimo Tornatore, Jie Zhang 0006
IEEE Internet Things J.5
2025 Distributed Model Training Task Migration for Hotspot Management in Intelligent Computing Center Interconnection With Tidal Characteristics
abstract
Intelligent computing center (ICC) is a new type of data center constructed with intelligent computing power, such as graphic processing units (GPUs) and artificial intelligence acceleration cards. With billions of parameters, the emergence of large models (e.g., ChatGPT) presents a significant demand of computing power. It may be challenging for a single ICC to provide the required computing power during large model training. Thus, ICC interconnections (ICCI) will become a typical and effective solution to provide intensive computing power. Due to human activities, traditional computing tasks (e.g., transaction processing and online entertainment) exhibit a tidal effect of computing demand, which leads to the tidal variation of remaining computing resources. Moreover, distributed model training (DMT) tasks are likely to cover peaks and valleys of the tidal effect in computing power. In this case, it is easy for DMT tasks to cause an ICC to become a hotspot (i.e., computing load in an ICC exceeds a desired threshold), which significantly degrades the reliability and performance of the ICC. This paper proposes DeepHM, a deep reinforcement learning-based hotspot management strategy through task migration in ICCI networks. To comprehensively consider the bandwidth metrics of the ICCI network, we further propose a dynamic wavelength allocation strategy, i.e., DeepHM-DWA. Simulation results show that the DeepHM and DeepHM-DWA reduce the hotspot compute unit time blocks by 19% and 18% with fewer number of migrated workers while balancing the computing load among multiple ICCs. DeepHM and DeepHM-DWA reduce the average completion time ratio of the DMT tasks by 2% and 5%, respectively.
Yingbo Fan, Yajie Li 0001, Carlos Natalino, Jiaxing Guo, Wanping Wu, Rongrong Ruan, Wei Wang 0116, Yongli Zhao 0001, Jie Zhang 0006
IEEE Trans. Netw. Serv. Manag.2
2023 Bulk Transfers With GCN Scheduling In Digital Twin Networks
abstract
Digital Twin Network (DTN) constructs a many-to-many mapping network by communicating and collaborating with massive Digital Twins (DTs), which can better assist the management and operation of large-scale modern systems. However, the emergence of DTN mirrors a growing demand for bulk data transfer between geo-distributed DT nodes in the physical network infrastructure. Conventional end-to-end connections are facing a great challenge in the presence of the background traffic fluctuation. In this paper, we propose a graph convolutional network (GCN)-enabled scheduling method to schedule bulk data transfers across the DTN in a store-and-forward (SnF) manner. Instead of solving the complex SnF scheduling problem on the entire network, the proposed method decomposes the problem into multiple sub-problems on different pre-computed routes. Instead of learning the scheduling results directly, the GCN model is used to predict the reachability of DT nodes. These unreachable nodes will be excluded from the scheduling process. Our studies show that the proposed method obtains better network performance and lower complexity when compared with the conventional scheduling methods.
Xiao Lin 0013, Lanfang Zheng, Keqin Shi, Yajie Li 0001
ICC4
2023 Infrastructure-efficient Virtual-Machine Placement and Workload Assignment in Cooperative Edge-Cloud Computing Over Backhaul Networks
abstract
Edge computing provides computing capability at close-user proximity to reduce service latency for end users. To improve the efficiency of edge computing infrastructures, geographically-distributed edge datacenters can co-work with each other and with cloud datacenters, forming a new paradigm referred to as cooperative edge-cloud computing. In this context, applications typically run on a virtual machine (VM) that can be replicated at multiple sites, and thus user traffic can be served at all the sites where corresponding VMs reside. For the performance of many applications, latency is a critical parameter. In this work, taking applications’ latencies as the primary constraint, we model the problem of “VM placement and workload assignment” as a mixed integer linear program and develop heuristic algorithms accordingly. The goal is to minimize the consumption of information technology (IT) infrastructures for placing VMs in cooperative edge-cloud computing, while meeting the heterogeneous latency demands of different applications. Some preliminary results indicate that edge datacenter's resource efficiency can be optimized by proper cross-site VM placement and workload re-direction.
Wei Wang 0116, Massimo Tornatore, Yongli Zhao 0001, Haoran Chen 0007, Yajie Li 0001, Abhishek Gupta 0003, Jie Zhang 0006, Biswanath Mukherjee
IEEE Trans. Cloud Comput.5
2022 Double-Machine-Learning-Based Resource Scheduling Method for Offloading Transfers
abstract
Dynamically scheduling the bandwidth based on the traffic variation is important for a task offloading system. However, it faces two challenges. On one hand, the time-varying nature of the offloading traffic makes it difficult to be predicted accurately. On the other hand, differentiated mechanisms are applied to different offloading task types, which greatly complicates the behavior of the task offloading system. It is hence difficult to estimate the performance metrics accurately, especially when the metric values are extremely small. To tackle this, we present a double-machine-learning-based resource scheduling (DML-RS) method for task offloading traffic in this paper. The features of DML-RS are as follows: i) the wavelet transform and the sliding time window are incorporated with the LSTM traffic prediction model, which can capture the periodic and volatile natures of the offloading traffic and hence improve the prediction accuracy; ii) the logarithmic converting is applied to the ANN estimation models, which can improve the sensitivity of the ANN models to the small values and hence provides higher estimation accuracy. As a result, DML-RS can predict the traffic demand of the next network reconfiguration time point and optimize the resource allocation based on the performance estimations in advance. Results show that DML-RS offers near-optimal results compared with the existing method.
Xiao Lin 0013, Songlei Lin, Junyi Shao, Keqin Shi, Yajie Li 0001
APCC6
2020 Service Function Path Provisioning With Topology Aggregation in Multi-Domain Optical Networks
abstract
Traffic flows are often processed by a chain of Service Functions (SFs) (known as Service Function Chaining (SFC)) to satisfy service requirements. The deployed path for a SFC is called Service Function Path (SFP). SFs can be virtualized and migrated to datacenters, thanks to the evolution of Software Defined Network (SDN) and Network Function Virtualization (NFV). In such a scenario, provisioning of paths (i.e., SFPs) between virtualized network functions is an important problem. SFP provisioning becomes more complex in a multi-domain network topology. `Topology aggregation' helps to create a single-domain view of such a network by abstracting multi-domain networks. However, traditional `topology aggregation' methods are unable to abstract SF resources properly, which is required for SFP provisioning. In this paper, we propose an SFC-Oriented Topology Aggregation (SOTA) method to enable abstraction for SFs in multi-domain optical networks. This study explores the node and the link aggregation degree to evaluate information compression during the `Topology aggregation' process. Additionally, we also propose a new data structure named wheel matrix and related operations to store routing information in the aggregated topology. Based on SOTA, we propose two cross-domain SFP provisioning algorithms named Ordered Anchor Selection (OAS) and ${k}$ -paths OAS (K-OAS), and a benchmark named Global OAS (GOAS). Simulation results show that SOTA could aggregate large-scale multi-domain optical networks into a small network that contains only 6.9% of the nodes and 10.1% of the links. Both OAS and K-OAS can calculate SFPs efficiently and reduce blocking probability up to 52.10% compared to the benchmark.
Boyuan Yan, Yongli Zhao 0001, Xiaosong Yu, Yajie Li 0001, Sabidur Rahman, Yongqi He, Xiangjun Xin 0001, Jie Zhang 0006
IEEE/ACM Trans. Netw.4