Yulong Li 0001

dblp:71/2140-1 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0003-0003-9653ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ECC: Efficient Concurrency Control for Disaggregated Memory Systems
abstract
Memory disaggregation architecture presents unique challenges in ensuring data consistency due to the limited computational power available at memory servers. One line of pessimistic solutions utilizes lock tables to deal with this challenge. Since the lock release signal cannot be immediately synchronized to the compute server, such solutions can only achieve suboptimal memory utilization. To address the aforementioned limitation, another line of solutions operates optimistically, polling the memory server until the operation completes successfully. However, these solutions trigger a massive number of unnecessary retries, resulting in performance collapse in high contention scenarios.high memory utilization while avoiding the need for retries. Our key idea is that compute servers proactively predict the appropriate retry interval and schedule accordingly. By analyzing the properties of requests on memory disaggregation, we find that the retry interval mainly consists of network fluctuations and host congestion delay. ECC design builds upon insights gained from this analysis. We have integrated ECC into Clover—a state-of-the-art memory disaggregation system and evaluated it through simulations and testbed experiments. Our testbed results show that ECC improves throughput by 50.8% over Clover.
Yang Li 0245, Yaozhen Li, Wenxin Li 0001, Yulong Li 0001, Song Zhang 0008, Renjie Pei, Xiancheng Meng, Keqiu Li
IEEE Trans. Computers4
2026 Rethinking Selective In-Network Aggregation for Multi-Tenant Learning
abstract
In-network aggregation accelerates distributed training by offloading gradient aggregation to the programmable switches. However, in multi-tenant learning environments, contention for limited switch memory can cause memory overflows that significantly degrade aggregation throughput. To mitigate memory overflow, prior work has proposed selective in-network aggregation, which allocates switch memory to a subset of jobs based on memory availability. This approach classifies jobs into INA jobs (which perform in-network aggregation) and PS jobs (which perform in-server aggregation). Despite these efforts, network congestion occurs and slows INA jobs, resulting in inefficient memory utilization and reduced aggregation throughput. In this paper, we present FlexINA, which rethinks selective in-network aggregation to deliver high aggregation throughput. The first challenge is how to avoid excessive memory under-utilization by INA jobs during network congestion. FlexINA introduces INA-aware congestion control, prioritizing reducing the sending window of PS jobs during network congestion. The second challenge is how to allow PS jobs to utilize under-utilized aggregators without affecting the aggregation of INA jobs. FlexINA implements an adaptive head-tail aggregation, optimizing memory usage by combining static head mapping for INA jobs (to use allocated head memory) and dynamic tail mapping for PS jobs (to use under-utilized tail memory). We implement a FlexINA prototype and evaluate it on both a small-scale testbed and in large-scale simulation experiments. Our evaluation shows that FlexINA improves aggregation throughput by up to$1.9\times $and$1.4\times $compared to selective in-network aggregation NetPack (ATP) and NetPack (A2TP), respectively.
Yulong Li 0001, Wenxin Li 0001, Song Zhang 0008, Jiawen Shen, Keqiu Li
IEEE Trans. Netw.1
2025 Fork: A Dual Congestion Control Loop for Small and Large Flows in Datacenters
abstract
Many existing transport designs aim to deliver ultra-low latency and high bandwidth for applications in high-speed datacenter networks. However, almost all of them intertwine the control of small and large flows using the same control entity (e.g., sender or receiver) and congestion feedback signal (e.g., ECN or credit), thus bringing significant performance impairments. By contrast, we seek to decouple the rate control of small flows from that of large ones.
Wenxin Li 0001, Yulong Li 0001, Lide Suo, Xuan Gao 0001, Xin Xie 0001, Sheng Chen 0015, Ziqi Fan, Wenyu Qu, Guyue Liu
EuroSys3
2025 Jarm: Automated Remote Memory System for Java Applications
Jiawen Shen, Wenxin Li 0001, Linxuan Zhong, Yulong Li 0001
ICA3PP (1)4
2025 GIT: Accelerating Distributed DNN Training via Similar Gradient Filtering
Yinan Yao, Yulong Li 0001, Wenxin Li 0001, Keqiu Li, Dehan Wen
ICA3PP (2)3
2025 Lark: A Buffer-aware Building Block for Programmable Packet Scheduling in Datacenters
abstract
Programmable packet scheduling enables users to customize scheduling algorithms flexibly without designing new ASICs. Existing schemes prefre to approximate optimal Push-In First-Out (PIFO) using First-In First-Out (FIFO) queues in commodity programmable switches. Despite its availability, these schemes suffer performance degradation due to the unawareness of available switch buffer. To be specific, when the port buffer is drained, existing schemes discard all incoming packets, even though these packets have higher priorities than the enqueued packets. In this paper, we reveal that the problem's key culprit is the lack of coordination between buffer management and packet scheduling in the switch. To fill this gap, we present Lark, a buffer-aware building block for programmable scheduling schemes designed to solve the above problem. Its key idea is to proactively drop the low-priority packets when the allocated buffer is to be drained, thereby admitting the later-arriving high-priority packets. Lark contains two modules, a lightweight gradient-based online prediction module and a simple priority-based decision module. Lark relies the former module to identify whether the allocated buffer is to be drained and uses the later to determine whether to drop the incoming packet. We have integrated Lark into two representative schemes, SP-PIFO and AIFO. Our large-scale evaluations over three realistic workloads show that Lark can significantly optimize their key metrics without sacrificing throughnut.
Song Zhang 0008, Wenxin Li 0001, Yulong Li 0001, Lide Suo, Sheng Chen 0015, Yitao Hu, Laiping Zhao, Keqiu Li
INFOCOM3
2025 Flexible Job Scheduling With Spatial-Temporal Compatibility for In-Network Aggregation
abstract
In-Network Aggregation (INA) solutions represent the forefront in advancing All-Reduce, utilizing limited switch memory for efficient gradient aggregation. However, existing INA solutions primarily focus on enhancing aggregation efficiency, often overlooking the efficient utilization of memory. Isolation solutions typically pre-allocate resources for each job, leading to memory wastage due to the uncontrolled use of resources. In contrast, the sharing solutions encounter significant memory contention, resulting in performance degradation within a multi-tenant environment. In this paper, we propose DynaINA, a flexible job scheduler to support multi-tenant training. The core idea of DynaINA is to provide spatial and temporal compatibility between jobs. For spatial compatibility, DynaINA utilizes multiple dynamic memory pools to provide job isolation. For temporal compatibility, DynaINA employs contention-aware job scheduling to facilitate memory sharing. Furthermore, DynaINA prioritizes communication-intensive jobs, leveraging the benefits of INA to enhance overall performance in training clusters. Extensive experiments with popular vision and language models demonstrate that DynaINA reduces training time by up to 65.16% and improves switch memory utilization by up to 85.02% compared to state-of-the-art solutions in a 100Gbps network.
Yulong Li 0001, Wenxin Li 0001, Yinan Yao, Song Zhang 0008, Linxuan Zhong, Keqiu Li
IEEE Trans. Computers1
2024 Efficient Disaggregated Memory Eviction with Glitter
abstract
Memory disaggregation, a promising technique allowing applications to use remote memory, is increasingly appealing in datacenters due to its high resource utilization. Operationally, the application’s host server constantly evicts unused data to remote to make room for memory allocation of new pages. Inefficient evictions allow memory usage to hit its limit, resulting in application blocking, which brings severe throughput degradation. However, most existing works neglect the importance of eviction. They offload the eviction to a background thread and set a fixed trigger timing, rendering a belated eviction. Worse still, they overlook the impact of network congestion on eviction efficiency, making their strategy flawed in large-scale scenarios. In this paper, we present Glitter, an adaptive, multi-level awareness eviction solution that accelerates applications by minimizing the overhead of application blocking from host and network aspects. For host, Glitter presents an adaptive eviction threshold adjustment to optimize the eviction timing, reducing the occurrence of application blocking. For network, Glitter adopts an eviction flow scheduling to address the hazards posed by flow contention at switches, decreasing the duration of each application blocking. Through comprehensive experiments, Glitter gives an average 1.4 throughput boost to Fastswap, a state-of-the-art disaggregated×memory system.
Linxuan Zhong, Wenxin Li 0001, Yulong Li 0001, Jiawen Shen, Song Zhang 0008, Wenyu Qu, Yitao Hu
HPCC3
2024 Host-driven In-Network Aggregation on RDMA
abstract
Large-scale datacenter networks are increasingly using in-network aggregation (INA) and remote direct memory access (RDMA) techniques to accelerate deep neural network (DNN) training. However, existing research trends suggest that these two techniques are on an inevitable collision course. To fill this gap, we present FreeINA, a host-driven in-network aggregation aimed at providing RDMA reliable connection (RC) for multi-tenant learning settings. FreeINA relies on dual transmission paths to support RC compatibility, with one path for INA and another one for aggregation on end-host parameter server. With dynamic control of these two paths, FreeINA can leave the traditional in-server aggregation unaffected while ensuring INA’s reliability without modifying RDMA network interfaces (RNICs). To support multi-tenant learning, FreeINA employs all-reduce-level memory allocation, which can capture the well-known "on and off" DNN training pattern and thus improve switch memory efficiency. We have implemented a FreeINA prototype using P4-programmable switch and commercial RNICs, and evaluated it extensively using 100Gbps testbed. The results show that compared to the state-of-the-art solution—ATP, FreeINA improves single-job training speedup ratio by 1.20×, while improving the aggregation throughput by 2.65× in multi-job scenario.
Yulong Li 0001, Wenxin Li 0001, Yinan Yao, Keqiu Li
INFOCOM1
2024 Mild: A Zero-Wait Multi-Round Proactive Transport
abstract
With the rapid growth of datacenter network link speed, multi-round matching based proactive solutions (e.g., dcPIM) has become increasingly attractive. Such solutions enable receivers to obtain as much global information as possible through multi-round matching, thereby facilitating them to make near-optimal decisions on bandwidth allocation. However, the matching phase before transmitting data introduces significant latency overhead. In this paper, we present Mild, a zero-wait solution that runs a second sender-driven control loop in parallel, leveraging in-network telemetry (INT) to detect and fill the spare bandwidth during the matching phase. Furthermore, we introduce a selective dropping mechanism to ensure that the packets from the second loop do not impact the data transmission of the primary loop. Additionally, we use the well-protected primary loop to perform loss recovery for the dropped packets efficiently. We integrate Mild into a representative proposal dcPIM and evaluate its performance through 100Gbps large-scale simulations. Compared to the state-of-the-art solution, Mild reduces the tail flow completion time (FCT) of short flows by up to 55% while achieving up to 57%/45% lower average FCT of medium/large flows.
Renjie Pei, Wenxin Li 0001, Yulong Li 0001, Song Zhang 0008, Yaozhen Li, Wenyu Qu
ISCC3
2024 Anole: Scheduling Flows for Fast Datacenter Networks With Packet Re-Prioritization
abstract
Many existing datacenter transports perform one-shot packet priority tagging at end-hosts and leave them fixed during the packet's transmission. In this paper, we experimentally show that: (1) such fixed packet priority is not sufficient for FCT (flow completion time) minimization, and (2) adjusting packet transmission priority in the network requires effective coordination among switches. Building on these insights, we present Anole, a new datacenter transport that advocates packet re-prioritization in near-bottleneck switches to minimize FCT. To this end, Anole integrates three simple-yet-effective techniques. First, it employs an in-network telemetry (INT) based approach to dynamically detect the bottleneck for each flow. Second, it adopts an on-off rate control mechanism for each sender to pause heavily congested flows but send lightly- and non-congested ones. Last, it leverages an altruistic scheduling policy at each switch to let the flows whose next hops are bottleneck switches give way to others. We implement an Anole prototype based on DPDK and show, through both testbed experiments and simulations, that Anole delivers significant performance advantages. For example, compared to EPN, Homa, and Aeolus, it shortens the average FCT of all (small) flows by up to 61.6% (89.1%).
Song Zhang 0008, Lide Suo, Wenxin Li 0001, Yulong Li 0001, Keqiu Li
IEEE Trans. Cloud Comput.5
2024 BRT: Buffer Management for RDMA/TCP Mix-Flows in Datacenter Networks
abstract
The coexistence of RDMA and TCP is prevalent in the datacenter. Despite the sound isolation at the end hosts, they share the same switches in the network. Their different networking behaviors (E.g., in hardware demand and transport protocols) lead to huge differentiated buffer demand for switches. However, existing buffer management schemes ignore these dissimilarities and simply treat such RDMA/TCP mix-flows as the typical multi-class traffic, resulting in inferior isolation and degrading networking performances. This paper presents BRT, a first systematic solution for the buffer management of RDMA/TCP mix-flows in the DCN. BRT’s key insight is to allocate buffer with the awareness of traffic’s networking characteristics while minimally impacting the other’s performance. Guided by this insight, it first employs a traffic characteristics-based window to detect whether queues are in the state of persistent long queues. Then, it adjusts the total allocated buffer for each traffic type based on the number of persistent long queues and the normalized dequeue rates to reduce the buffer occupancy of meaningless queuing. Last, it calculates the buffer threshold for RDMA/TCP queues separately and uses a simple yet effective approach to prioritize the absorption of small flows. Our large-scale packet-level evaluations show that BRT can effectively optimize the networking performances for RDMA/TCP mix-flows. For example, compared to current practice, BRT achieves up to 53.5%, 46.7%, and 48.5% lower average FCT for incast flows, RDMA small flows, and TCP small flows, respectively, without sacrificing the overall throughput.
Song Zhang 0008, Wenxin Li 0001, Lide Suo, Yulong Li 0001, Jien Kato, Keqiu Li
IEEE Trans. Netw. Serv. Manag.5
2024 Analyzing and Detecting Information Types of Developer Live Chat Threads
abstract
Online chatrooms serve as vital platforms for information exchange among software developers. With multiple developers engaged in rapid communication and diverse conversation topics, the resulting chat messages often manifest complexity and lack structure. To enhance the efficiency of extracting information from chat threads , automatic mining techniques are introduced for thread classification. However, previous approaches still grapple with unsatisfactory classification accuracy due to two primary challenges that they struggle to adequately capture long-distance dependencies within chat threads and address the issue of category imbalance in labeled datasets. To surmount these challenges, we present a topic classification approach for chat information types named EAEChat. Specifically, EAEChat comprises three core components: the text feature encoding component captures contextual text features using a multi-head self-attention mechanism-based text feature encoder, and a siamese network is employed to mitigate overfitting caused by limited data; the data augmentation component expands a small number of categories in the training dataset using a technique tailored to developer chat messages, effectively tackling the challenge of imbalanced category distribution; the non-text feature encoding component employs a feature fusion model to integrate deep text features with manually extracted non-text features. Evaluation across three real-world projects demonstrates that EAEChat, respectively, achieves an average precision, recall, and F1-score of 0.653, 0.651, and 0.644, and it marks a significant 7.60% improvement over the state-of-the-art approaches. These findings confirm the effectiveness of our method in proficiently classifying developer chat messages in online chatrooms.
Xiuwei Shang, Shikai Guo, Yulong Li 0001, Rong Chen 0003, Hui Li 0014, He Jiang 0001
ACM Trans. Softw. Eng. Methodol.5
2023 MiddleCache: Accelerating TCP based In-memory Key-value Stores using eBPF
abstract
In-memory key-value stores are widely used in modern web services to support large-scale user requests by caching popular data. Their performance is critical, and BMC, the state-of-the-art work, builds an in-kernel cache and processes requests before the stack using eBPF to reduce the overhead of the kernel network stack. However, BMC fails to support stateful protocol TCP because pre-stack processing creates TCP state bias between the client and server.TCP is widely used by in-memory key-value stores, is even the only choice for some applications (e.g., Redis), and also suffers from performance issues. In this work, we present MiddleCache, a TCP-enabled in-memory key-value store acceleration design. Our key observation is that the TCP state bias of the client and server can be inferred and eliminated with packet length. The design of MiddleCache has two key parts: (i) A compact TCP state maintenance mechanism that accumulates packet lengths and applies corrections to the packet header, which realize TCP support within the constrains of eBPF. (ii) Lock-free accumulation counters that support high-performance concurrent access by utilizing Receive Side Scaling (RSS). Our experiments show that, compared with Memcached, MiddleCache reduces 56% processing latency on cache hit and achieves a 3.8× throughput improvement on Facebook-like small-size requests workload.
Yiren Pang, Sheng Chen 0015, Wenxin Li 0001, Yulong Li 0001, Xin He 0043, Song Zhang 0008, Zewei Guan, Lide Suo
ICPADS5
2023 DupHunter: Detecting Duplicate Pull Requests in Fork-Based Development
abstract
The emergence of numerous fork-based development platforms facilitates the development of Open-Source Software (OSS) projects. Developers across the world can fork software projects and submit their Pull Requests (PRs) to the projects. However, as the number of forks increases, numerous duplicate PRs might be submitted. These duplicate PRs may cause extra code review workload and frustrate developers working on the projects. To detect duplicate PRs, many approaches have been proposed, which analyze the similarity of different elements in PRs. However, previous approaches still suffer from unsatisfied detection accuracy due to two challenges. That is, they ignore the syntactic structural information of text elements in PRs and lack the joint reasoning between different elements of two PRs. In this study, we propose an automated duplicate PRs detector namedDupHunter(Duplicate PRsHunter), which includes a graph embedding component and a duplicate PRs detection component to address the above challenges. The graph embedding component uses a feature graph to represent a PR. It encodes the syntactic structure and semantics of text elements (e.g., the title and the description), as well as the knowledge of non-text elements (e.g., the submission time), to address the syntactic structural information challenge. The duplicate PRs detection component tackles the joint reasoning challenge using a graph matching network, which enables the information exchange and matching across different elements of two feature graphs with an attention coefficient mechanism. Experiments on 26 open-source projects show that DupHunter achieves an averageF1-score@1value of 0.650, significantly outperforming the state-of-the-art approaches by 3.2% to 48.1%. DupHunter can accurately detect duplicate PRs, with an averagePrecision@1value of 0.922 and an averageRecall@1value of 0.502.
He Jiang 0001, Yulong Li 0001, Shikai Guo, Tao Zhang 0001, Hui Li 0014, Rong Chen 0003
IEEE Trans. Software Eng.2