EDBT 2026 Demo / reviewers in the wild / expert
Song Zhang 0008
dblp:77/3067-8
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-7448-0850ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Computer networks · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ECC: Efficient Concurrency Control for Disaggregated Memory SystemsabstractMemory disaggregation architecture presents unique challenges in ensuring data consistency due to the limited computational power available at memory servers. One line of pessimistic solutions utilizes lock tables to deal with this challenge. Since the lock release signal cannot be immediately synchronized to the compute server, such solutions can only achieve suboptimal memory utilization. To address the aforementioned limitation, another line of solutions operates optimistically, polling the memory server until the operation completes successfully. However, these solutions trigger a massive number of unnecessary retries, resulting in performance collapse in high contention scenarios.high memory utilization while avoiding the need for retries. Our key idea is that compute servers proactively predict the appropriate retry interval and schedule accordingly. By analyzing the properties of requests on memory disaggregation, we find that the retry interval mainly consists of network fluctuations and host congestion delay. ECC design builds upon insights gained from this analysis. We have integrated ECC into Clover—a state-of-the-art memory disaggregation system and evaluated it through simulations and testbed experiments. Our testbed results show that ECC improves throughput by 50.8% over Clover. Yang Li 0245, Yaozhen Li, Wenxin Li 0001, Yulong Li 0001, Song Zhang 0008, Renjie Pei, Xiancheng Meng, Keqiu Li |
IEEE Trans. Computers | 5 |
| 2026 | Rethinking Selective In-Network Aggregation for Multi-Tenant LearningabstractIn-network aggregation accelerates distributed training by offloading gradient aggregation to the programmable switches. However, in multi-tenant learning environments, contention for limited switch memory can cause memory overflows that significantly degrade aggregation throughput. To mitigate memory overflow, prior work has proposed selective in-network aggregation, which allocates switch memory to a subset of jobs based on memory availability. This approach classifies jobs into INA jobs (which perform in-network aggregation) and PS jobs (which perform in-server aggregation). Despite these efforts, network congestion occurs and slows INA jobs, resulting in inefficient memory utilization and reduced aggregation throughput. In this paper, we present FlexINA, which rethinks selective in-network aggregation to deliver high aggregation throughput. The first challenge is how to avoid excessive memory under-utilization by INA jobs during network congestion. FlexINA introduces INA-aware congestion control, prioritizing reducing the sending window of PS jobs during network congestion. The second challenge is how to allow PS jobs to utilize under-utilized aggregators without affecting the aggregation of INA jobs. FlexINA implements an adaptive head-tail aggregation, optimizing memory usage by combining static head mapping for INA jobs (to use allocated head memory) and dynamic tail mapping for PS jobs (to use under-utilized tail memory). We implement a FlexINA prototype and evaluate it on both a small-scale testbed and in large-scale simulation experiments. Our evaluation shows that FlexINA improves aggregation throughput by up to$1.9\times $and$1.4\times $compared to selective in-network aggregation NetPack (ATP) and NetPack (A2TP), respectively. Yulong Li 0001, Wenxin Li 0001, Song Zhang 0008, Jiawen Shen, Keqiu Li |
IEEE Trans. Netw. | 4 |
| 2025 | Lark: A Buffer-aware Building Block for Programmable Packet Scheduling in DatacentersabstractProgrammable packet scheduling enables users to customize scheduling algorithms flexibly without designing new ASICs. Existing schemes prefre to approximate optimal Push-In First-Out (PIFO) using First-In First-Out (FIFO) queues in commodity programmable switches. Despite its availability, these schemes suffer performance degradation due to the unawareness of available switch buffer. To be specific, when the port buffer is drained, existing schemes discard all incoming packets, even though these packets have higher priorities than the enqueued packets. In this paper, we reveal that the problem's key culprit is the lack of coordination between buffer management and packet scheduling in the switch. To fill this gap, we present Lark, a buffer-aware building block for programmable scheduling schemes designed to solve the above problem. Its key idea is to proactively drop the low-priority packets when the allocated buffer is to be drained, thereby admitting the later-arriving high-priority packets. Lark contains two modules, a lightweight gradient-based online prediction module and a simple priority-based decision module. Lark relies the former module to identify whether the allocated buffer is to be drained and uses the later to determine whether to drop the incoming packet. We have integrated Lark into two representative schemes, SP-PIFO and AIFO. Our large-scale evaluations over three realistic workloads show that Lark can significantly optimize their key metrics without sacrificing throughnut. Song Zhang 0008, Wenxin Li 0001, Yulong Li 0001, Lide Suo, Sheng Chen 0015, Yitao Hu, Laiping Zhao, Keqiu Li |
INFOCOM | 1 |
| 2025 | AirBFT: An Efficient and Robust Consensus Mechanism for Large-Scale Drone CollaborationabstractThe application scenarios of drone collaboration are rapidly expanding, such as the low-altitude economy and wildfire protection. Blockchain-based drone collaboration requires a consensus mechanism to ensure efficient and secure consistency among large-scale distributed nodes. However, the existing consensus mechanism has problems with poor fault tolerance of topology and rigid proposal concurrency. To this end, this paper proposes AirBFT, an efficient and robust consensus mechanism for large-scale drone collaboration. First, this paper designs a new four-layer network topology, using upper-member and lower-member communication, while ensuring the maximum 1/3 resilience and fanout of √N. Secondly, this paper proposes a dynamic pipelining algorithm to adjust the parallelism of proposals according to the real-time network status. Finally, this paper proposes a committee sampling technology based on the EigenTrust algorithm to reduce the impact of the malicious behavior of Byzantine nodes. Experiments based on the public consensus framework show that compared with Kauri and HotStuff, the proposed AirBFT reduces transaction confirmation delay by 58%, the throughput is increased by 1.9 times, and it can ensure efficient operation with a 1/3 Byzantine node ratio. Zhongju Yan, Chenyu Zhang 0008, Yiran Lv, Hao Xu 0025, Xiulong Liu 0001, Song Zhang 0008, Sheng Chen 0015, Xiaoyi Tao, Keqiu Li |
IEEE Internet Things J. | 7 |
| 2025 | Flexible Job Scheduling With Spatial-Temporal Compatibility for In-Network AggregationabstractIn-Network Aggregation (INA) solutions represent the forefront in advancing All-Reduce, utilizing limited switch memory for efficient gradient aggregation. However, existing INA solutions primarily focus on enhancing aggregation efficiency, often overlooking the efficient utilization of memory. Isolation solutions typically pre-allocate resources for each job, leading to memory wastage due to the uncontrolled use of resources. In contrast, the sharing solutions encounter significant memory contention, resulting in performance degradation within a multi-tenant environment. In this paper, we propose DynaINA, a flexible job scheduler to support multi-tenant training. The core idea of DynaINA is to provide spatial and temporal compatibility between jobs. For spatial compatibility, DynaINA utilizes multiple dynamic memory pools to provide job isolation. For temporal compatibility, DynaINA employs contention-aware job scheduling to facilitate memory sharing. Furthermore, DynaINA prioritizes communication-intensive jobs, leveraging the benefits of INA to enhance overall performance in training clusters. Extensive experiments with popular vision and language models demonstrate that DynaINA reduces training time by up to 65.16% and improves switch memory utilization by up to 85.02% compared to state-of-the-art solutions in a 100Gbps network. Yulong Li 0001, Wenxin Li 0001, Yinan Yao, Song Zhang 0008, Linxuan Zhong, Keqiu Li |
IEEE Trans. Computers | 5 |
| 2024 | Efficient Disaggregated Memory Eviction with GlitterabstractMemory disaggregation, a promising technique allowing applications to use remote memory, is increasingly appealing in datacenters due to its high resource utilization. Operationally, the application’s host server constantly evicts unused data to remote to make room for memory allocation of new pages. Inefficient evictions allow memory usage to hit its limit, resulting in application blocking, which brings severe throughput degradation. However, most existing works neglect the importance of eviction. They offload the eviction to a background thread and set a fixed trigger timing, rendering a belated eviction. Worse still, they overlook the impact of network congestion on eviction efficiency, making their strategy flawed in large-scale scenarios. In this paper, we present Glitter, an adaptive, multi-level awareness eviction solution that accelerates applications by minimizing the overhead of application blocking from host and network aspects. For host, Glitter presents an adaptive eviction threshold adjustment to optimize the eviction timing, reducing the occurrence of application blocking. For network, Glitter adopts an eviction flow scheduling to address the hazards posed by flow contention at switches, decreasing the duration of each application blocking. Through comprehensive experiments, Glitter gives an average 1.4 throughput boost to Fastswap, a state-of-the-art disaggregated×memory system. Linxuan Zhong, Wenxin Li 0001, Yulong Li 0001, Jiawen Shen, Song Zhang 0008, Wenyu Qu, Yitao Hu |
HPCC | 5 |
| 2024 | Mild: A Zero-Wait Multi-Round Proactive TransportabstractWith the rapid growth of datacenter network link speed, multi-round matching based proactive solutions (e.g., dcPIM) has become increasingly attractive. Such solutions enable receivers to obtain as much global information as possible through multi-round matching, thereby facilitating them to make near-optimal decisions on bandwidth allocation. However, the matching phase before transmitting data introduces significant latency overhead. In this paper, we present Mild, a zero-wait solution that runs a second sender-driven control loop in parallel, leveraging in-network telemetry (INT) to detect and fill the spare bandwidth during the matching phase. Furthermore, we introduce a selective dropping mechanism to ensure that the packets from the second loop do not impact the data transmission of the primary loop. Additionally, we use the well-protected primary loop to perform loss recovery for the dropped packets efficiently. We integrate Mild into a representative proposal dcPIM and evaluate its performance through 100Gbps large-scale simulations. Compared to the state-of-the-art solution, Mild reduces the tail flow completion time (FCT) of short flows by up to 55% while achieving up to 57%/45% lower average FCT of medium/large flows. Renjie Pei, Wenxin Li 0001, Yulong Li 0001, Song Zhang 0008, Yaozhen Li, Wenyu Qu |
ISCC | 4 |
| 2024 | Flow Scheduling with Imprecise Knowledge
Wenxin Li 0001, Xin He 0043, Keqiu Li, Kai Chen 0005, Zhao Ge, Zewei Guan, Heng Qi, Song Zhang 0008, Guyue Liu |
NSDI | 9 |
| 2024 | Anole: Scheduling Flows for Fast Datacenter Networks With Packet Re-PrioritizationabstractMany existing datacenter transports perform one-shot packet priority tagging at end-hosts and leave them fixed during the packet's transmission. In this paper, we experimentally show that: (1) such fixed packet priority is not sufficient for FCT (flow completion time) minimization, and (2) adjusting packet transmission priority in the network requires effective coordination among switches. Building on these insights, we present Anole, a new datacenter transport that advocates packet re-prioritization in near-bottleneck switches to minimize FCT. To this end, Anole integrates three simple-yet-effective techniques. First, it employs an in-network telemetry (INT) based approach to dynamically detect the bottleneck for each flow. Second, it adopts an on-off rate control mechanism for each sender to pause heavily congested flows but send lightly- and non-congested ones. Last, it leverages an altruistic scheduling policy at each switch to let the flows whose next hops are bottleneck switches give way to others. We implement an Anole prototype based on DPDK and show, through both testbed experiments and simulations, that Anole delivers significant performance advantages. For example, compared to EPN, Homa, and Aeolus, it shortens the average FCT of all (small) flows by up to 61.6% (89.1%). Song Zhang 0008, Lide Suo, Wenxin Li 0001, Yulong Li 0001, Keqiu Li |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | BRT: Buffer Management for RDMA/TCP Mix-Flows in Datacenter NetworksabstractThe coexistence of RDMA and TCP is prevalent in the datacenter. Despite the sound isolation at the end hosts, they share the same switches in the network. Their different networking behaviors (E.g., in hardware demand and transport protocols) lead to huge differentiated buffer demand for switches. However, existing buffer management schemes ignore these dissimilarities and simply treat such RDMA/TCP mix-flows as the typical multi-class traffic, resulting in inferior isolation and degrading networking performances. This paper presents BRT, a first systematic solution for the buffer management of RDMA/TCP mix-flows in the DCN. BRT’s key insight is to allocate buffer with the awareness of traffic’s networking characteristics while minimally impacting the other’s performance. Guided by this insight, it first employs a traffic characteristics-based window to detect whether queues are in the state of persistent long queues. Then, it adjusts the total allocated buffer for each traffic type based on the number of persistent long queues and the normalized dequeue rates to reduce the buffer occupancy of meaningless queuing. Last, it calculates the buffer threshold for RDMA/TCP queues separately and uses a simple yet effective approach to prioritize the absorption of small flows. Our large-scale packet-level evaluations show that BRT can effectively optimize the networking performances for RDMA/TCP mix-flows. For example, compared to current practice, BRT achieves up to 53.5%, 46.7%, and 48.5% lower average FCT for incast flows, RDMA small flows, and TCP small flows, respectively, without sacrificing the overall throughput. Song Zhang 0008, Wenxin Li 0001, Lide Suo, Yulong Li 0001, Jien Kato, Keqiu Li |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | dBFC: Destination-based Backpressure Flow Control for IncastabstractIncast happens in data center networks (DCNs) when many senders simultaneously send flows to one receiver. Incast occurs frequently and compromises flow performance severely. However, widely deployed end-to-end congestion control (CC) protocols are inefficient at handling incast, as they rely on delayed congestion signals and can not isolate incast flows. Recently, researchers have favored per-hop flow control protocols since they could achieve fast reaction and incast isolation. Nevertheless, these impressive advantages are impractical because of hardware resource limitations.In this paper, we present a new per-hop flow control protocol dBFC (Destination-based Backpressure Flow Control), which achieves these advantages within limited hardware resources. Our key insight is that the number of destinations reached by queued packets in a port is minuscule. dBFC is compatible with existing CC protocols. We use large-scale NS3 simulations to evaluate dBFC. In our evaluation, compared with deployed CC protocols and the state-of-the-art per-hop flow control protocol, dBFC reduces the maximum buffer occupancy by 5 − 23× and thus provides up to 21× lower flow completion time of small flows in limited hardware resources. Zewei Guan, Wenxin Li 0001, Xin He 0043, Song Zhang 0008, Keqiu Li |
ICPADS | 4 |
| 2023 | MiddleCache: Accelerating TCP based In-memory Key-value Stores using eBPFabstractIn-memory key-value stores are widely used in modern web services to support large-scale user requests by caching popular data. Their performance is critical, and BMC, the state-of-the-art work, builds an in-kernel cache and processes requests before the stack using eBPF to reduce the overhead of the kernel network stack. However, BMC fails to support stateful protocol TCP because pre-stack processing creates TCP state bias between the client and server.TCP is widely used by in-memory key-value stores, is even the only choice for some applications (e.g., Redis), and also suffers from performance issues. In this work, we present MiddleCache, a TCP-enabled in-memory key-value store acceleration design. Our key observation is that the TCP state bias of the client and server can be inferred and eliminated with packet length. The design of MiddleCache has two key parts: (i) A compact TCP state maintenance mechanism that accumulates packet lengths and applies corrections to the packet header, which realize TCP support within the constrains of eBPF. (ii) Lock-free accumulation counters that support high-performance concurrent access by utilizing Receive Side Scaling (RSS). Our experiments show that, compared with Memcached, MiddleCache reduces 56% processing latency on cache hit and achieves a 3.8× throughput improvement on Facebook-like small-size requests workload. Yiren Pang, Sheng Chen 0015, Wenxin Li 0001, Yulong Li 0001, Xin He 0043, Song Zhang 0008, Zewei Guan, Lide Suo |
ICPADS | 7 |
| 2022 | Efficient Control of Unscheduled Packets for Credit-based Proactive TransportabstractProactive transport has been attractive in modern high-speed and shallow buffered datacenter networks. At its heart, the link capacity is proactively allocated as credit, and then following scheduled packets are triggered by active senders according to credits, providing (near) zero loss rate and extremely low latency. Despite being promising, when waiting for credits in the first RTT, called the “pre-credit phase a substantial amount of spare bandwidth is being underutilized. To bridge this gap, current practices send a Bandwidth-Delay-Product (BDP) worth of unscheduled packets with line rate in the pre-credit phase, however, degrading throughput or tail latency. One key insight is that they consider the spare bandwidth in the pre-credit phase as a fixed value but actually changes across time and space. In this paper, we present Schef, a novel switch-based bandwidth calculation mechanism to address the pre-credit phase challenge without introducing network congestion or throughput degradation. The main idea is to calculate the spare bandwidth and apply efficient control of unscheduled packets at switches. Schef is compatible with existing proactive transports. We integrate it into a representative proposal NDP and evaluate its performance through large-scale simulations. Compared with the state-of-art, Schef reduces the average and 99th flow completion time (FCT) of short flows by up to 23% and 31%, and achieves 13% higher goodput simultaneously. Xin He 0043, Wenxin Li 0001, Song Zhang 0008, Keqiu Li |
ICPADS | 3 |
| 2022 | Efficient Online Scheduling for Coflow-Aware Machine Learning ClustersabstractDistributed machine learning (DML) is an increasingly important workload. In a DML job, each communication phase can comprise acoflow, and there are dependencies among its coflows. Thus, efficient coflow scheduling becomes critical for DML jobs. However, the majority of existing solutions focus on scheduling single-stage coflows with no dependencies. While there are a few studies schedule dependent coflows of multi-stage jobs, they suffer from either practical or theoretical issues. Motivated by this situation, we study how to schedule dependent coflows of multiple DML jobs to minimize the total JCT in a shared cluster. We present a formal mathematical formulation for this problem and prove its NP-hardness. To solve this problem without job size information, we present an online coflow-aware optimization framework calledParrot. The core idea inParrotis to infer the job with the shortest remaining processing time (SRPT) each time and dynamically control the inferred job's bandwidth based on how confident it is an SRPT job while being mindful of not starving any other job. Specifically, in the design ofParrot, we present a least per-coflow attained service (LPCAS) policy to infer the SRPT job. We further propose a dynamic job weight assignment mechanism and a linear program (LP) based weighted bandwidth scaling strategy for sharing bandwidth among DML jobs. We have proved thatParrotalgorithm has a non-trivial competitive ratio. The results from large-scale trace-driven simulations further demonstrate that ourParrotcan reduce the total JCT by up to 58.4 percent, compared to the state-of-the-art Aalo solution. Wenxin Li 0001, Sheng Chen 0015, Keqiu Li, Heng Qi, Renhai Xu, Song Zhang 0008 |
IEEE Trans. Cloud Comput. | 6 |