EDBT 2026 Demo / reviewers in the wild / expert
Haoyang Li 0006
dblp:118/0004-6
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0009-0008-8728-5294ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCC: Re-architecting Delay-based Congestion Control in Datacenter Networks
Wanchun Jiang, Haoyang Li 0006, Danfeng Shan, Fengyuan Ren, Jiawei Huang 0001, Jianxin Wang 0001 |
NSDI | 2 |
| 2025 | Introspective Congestion Control for Consistent High PerformanceabstractThe congestion control (CC) algorithm is expected to achieve consistent high performance under different network environments. Traditionally, classic CCs are designed with the methodology of inferring path conditions to guide the rate adjustment. However, this methodology suffers from wrong path condition inferences in certain cases, which mislead the rate adjustment and lead to performance degradation. To avoid wrong path condition inferences, we develop the projection-based introspective method and design the introspective congestion control (ICC) algorithm in this paper. Specifically, the rate adjustment rules are designed to possess a specialized profile such that the projection of the profile can be distinguished under unchanged path conditions. In this way, the projection, which can be distinguished from the time series of delay signals in the frequency domain, facilitates ICC to extract more information for path condition inferences. Consequently, with the introspection on the projection, ICC can avoid being misled by wrong path condition inferences and thus achieve consistent high performance under different conditions. The advantages of ICC are confirmed through extensive experiments conducted on various locally emulated scenarios, global testbeds over the Internet, and the Alipay platform. Wanchun Jiang, Haoyang Li 0006, Jia Wu 0011, Fengyuan Ren, Jianxin Wang 0001 |
EuroSys | 2 |
| 2025 | Scalable Bolt: Taming the Burst Queue via Scalable Token ManagementabstractEfficient congestion control (CC) is critical to maintaining high performance in modern datacenter networks. Recently, Bolt, proposed in NSDI 23, builds a sub-RTT control loop that enables ultra-low-latency reactions to congestion. Combined with an effective ramp-up mechanism driven by the proactive tokens, Bolt achieves superior performance compared to other existing CC schemes, especially in the face of workloads dominated by short flows. However, our study reveals that Bolt suffers from a severe burst bottleneck queue in heterogeneous topologies. This issue is particularly concerning because such heterogeneity is common in widely deployed Fat-tree and Leaf-spine topologies, and is expected to increase as data centers continue to scale out. To solve this problem, Scalable Bolt (S-Bolt) is proposed. By scaling the generation and consumption of tokens, S-Bolt eliminates the burst bottleneck queue while retaining the advantages of fast responding to congestion and spare bandwidth. Moreover, S-Bolt is friendly to the implementation on switches. Evaluating results show that S-Bolt works well under the current workloads and heterogeneous topologies. Specifically, S-Bolt reduces the burst queue by 70.6% and cuts down the tail flow completion time by 12.5%, compared with Bolt. Haoyang Li 0006, Peile Chen, Ang Jiang, Wanchun Jiang |
ICPADS | 1 |
| 2025 | Cut the Response Time of Key-Value Stores by the SDN-Based SchedulerabstractAs the foundational components of large-scale applications, distributed key-value stores must respond to user requests quickly. However, a user request typically comprises multiple key-value access operations, which are processed in parallel across different servers, and the response time is determined by the slowest operation. To reduce the mean response time of requests, existing approaches schedule the sequence of key-value access operations across different servers so that all operations of a request complete at approximately the same time. Nevertheless, all of these approaches operate in a distributive manner, and their theoretical performance boundaries are unknown. To address these issues, we designed SDN-KVS (Key-Value Scheduler based on Software Defined Network), which migrates the waiting queue of key-value access operations from overloaded servers to the SDN controller. In this way, SDN-KVS centrally schedules the requests from different clients to overloaded servers without extra latency overhead. The scheduling result is proven to be$(1+2 \eta)$-approximation, i.e., the mean response time of requests is smaller than ($1+2 \eta$) times of the optimal value, where$\eta$is the parameter to make a trade-off between mean and tail response time. Simulation results confirm the excellent performance of the SDN-KVS algorithm. Specifically, SDN-KVS outperforms existing algorithms up to 37.6% and 79.8% in terms of mean and tail response time, respectively. Wanchun Jiang, Haoyang Li 0006, Chengke Wen, Rongfei Zeng, Jiawei Huang 0001, Jianxin Wang 0001 |
IWQoS | 4 |
| 2025 | Teaching to Fish Rather Than Giving a Fish: The Concentrator Method of Teaching Classic Congestion Control With Learning-Based moduleabstractNowadays, Congestion Control (CC) algorithms are expected to satisfy the diverse demands of applications running over diverse networks. To achieve this goal, the combinations, which are expected to inherit both the advantages of classic CC in terms of convergence, overhead, and explainability, and the advantages of learning-based CC on adapting to diverse networks and demands, become a hot topic. In this paper, we reveal the existing combination works are eithergiving a fishorteaching to fish. Based on the insight of their essential issues, we develop the Concentrator method ofteaching to fish. According to this method, we propose Seagull as a step further. Specifically, Seagull captures the network characteristics and application demands in a coarse-grained manner via an online learning module. Moreover, the online learning module guides the customization of the rate adjustment rules of the classic CC module for fine-grained system evolution. Replacing the assumption on networks by the captured characteristics, the classic CC module of Seagull can fulfill the specified application demands. Real-world experimental results show Seagull respectively outperforms Orca, PCC-Vivace, and CUBIC by$49.3\%,\ 30.4\%$, and 24.9% in terms of throughput over the Internet, and improves the video quality of experience (QoE) by$12.9\sim 33.5\%$compared to CUBIC over cellular links. Haoyang Li 0006, Wanchun Jiang, Jie Wang 0067, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Analysis and Improvement of PowerTCPabstractNowadays, Congestion Control (CC) algorithms based on In-Network Telemetry (INT) are popular in datacenter networks, because INT is supported by many commercial devices and provides comprehensive information about network congestion. In this work, we analyze the recent INT-based CC algorithm PowerTCP and reveal its fairness and large delay issues under the conditions of a large number of flows. Inspired by the analytical results, we propose Trident, which enhances PowerTCP by accelerating the speed of converging to fairness and maintaining a low queuing delay with a small equilibrium point. Simulations based on the open source codes confirm that Trident outperforms existing INT-based CC algorithms HPCC, PowerTCP and Poseidon by 14.3%, 14.8%, and 63.5% in terms of FCT of short flows. Moreover, Trident outperforms existing CC algorithms such as Timely, DCQCN and DCTCP, benefiting from the advantages inherited from PowerTCP. Wanchun Jiang, Haoyang Li 0006, Jiawei Huang 0001, Jianxin Wang 0001 |
IWQoS | 3 |
| 2024 | Improvement of Copa: Behaviors and Friendliness of Delay-Based Congestion Control AlgorithmabstractDelay-based congestion control has drawn a lot of attention in both academics and industry recently. Specifically, the Copa algorithm proposed in NSDI can achieve consistent high performance under various network environments and has already been deployed on Facebook. In this paper, we theoretically analyze Copa and reveal its large queuing delay and poor fairness issue under certain conditions. The root cause is that Copa fails to achieve its expected behaviors, i.e., clear the bottleneck buffer occupancy periodically. Moreover, we also reveal that the pathological competitive mode of Copa fails to guarantee friendliness. To address these issues, we propose Copa+, which enhances Copa with a parameter adaptation mechanism and an optimized competitive mode. Designed based on our theoretical analysis, Copa+ can adaptively clear the bottleneck buffer occupancy and become friendly to Cubic in the competitive mode. As a result, Copa+ inherits the advantages of Copa but achieves lower queuing delay and better fairness under different environments, as confirmed by real-world experiments and simulations. Specifically, Copa+ has the highest average throughput over different Internet links among different cloud nodes, compared to Cubic, BBR, PCC Vivace, Remy, and Indigo. Meanwhile, Copa+ has an 8.1% increase in throughput and similar low queuing delay compared to Copa. Moreover, Copa+ achieves 14.6% lower queuing delay and 2.4% higher throughput compared to Sprout over emulated cellular links. Wanchun Jiang, Haoyang Li 0006, Jia Wu 0002, Zheyuan Liu 0008, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Consistent Low Latency Scheduler for Distributed Key-Value StoresabstractNowadays, the distributed key-value stores have become the basic building block for large-scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for a single end-user request. Accordingly, the completion time of an end-user request is determined by the last completed key-value access operation. Scheduling the order of serving key-value access operations can effectively reduce the completion times of end requests, thereby improving the user experience. However, existing scheduling algorithms hardly achieve consistent low latency due to the following challenges: the large overhead of cooperating clients and servers, the time-varying load and performance of servers, the traffic distribution can be either heavy-tailed or light-tailed and both the mean and the tail completion time are expected to be low. In this paper, we formalize the problem of scheduling key-value access operations and show it is NP-hard. Furthermore, we heuristically design the distributed adaptive scheduler (DAS), which distributively combines the largest remaining processing time last and the shortest remaining process time first algorithms. Theoretical analysis shows that DAS is adaptive to the time-varying traffic and server performance and can achieve consistent low mean and tail latency regardless of traffic distributions. Extensive simulations show that DAS reduces the mean request completion time by$17 \! \sim \! 50\%$with heavy-tailed traffic and$2 \! \sim 26 \! \%$with light-tailed traffic, while keeping the smallest tail completion time, compared to the default first come first served algorithm. Moreover, DAS outperforms the existing Rein-SBF algorithm under various scenarios. Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jiawei Huang 0001, Jianxin Wang 0001, Tong Zhang 0018 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Copa+: Analysis and Improvement of the Delay-based Congestion Control Algorithm CopaabstractCopa is a delay-based congestion control algorithm proposed in NSDI recently. It can achieve consistent high performance under various network environments and has already been deployed in Facebook. In this paper, we theoretically analyze Copa and reveal its large queuing delay and poor fairness issue under certain conditions. The root cause is that Copa fails to clear the bottleneck buffer occupancy periodically as expected. Accordingly, Copa may get a wrong base RTT estimation and enter its competitive mode by mistake, leading to large delay and unfairness. To address these issues, we propose Copa+, which enhances Copa with a parameter adaptation mechanism and an optimized competitive mode entrance criterion. Designed based on our theoretical analysis, Copa+ can adaptively clear the bottleneck buffer occupancy for correct estimation of base RTT. Consequently, Copa+ inherits the advantages of Copa but achieves lower queuing delay and better fairness under different environments, as confirmed by the real-world experiments and simulations. Specifically, Copa+ has the highest throughput similar to Copa but 11.9% lower queuing delay over different Internet links among different cloud nodes, and achieves 39.4% lower queuing delay and 8.9% higher throughput compared to Sprout over emulated cellular links. Wanchun Jiang, Haoyang Li 0006, Zheyuan Liu 0008, Jia Wu 0002, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
INFOCOM | 2 |
| 2021 | Cutting the Request Completion Time in Key-value Stores with Distributed Adaptive SchedulerabstractNowadays, the distributed key-value stores have become the basic building block for large scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for the data required by a single end-user request. Hence, the completion time of the end request is determined by the last completed key-value access operation. Accordingly, scheduling the order of key-value access operations of different end requests can effectively reduce their completion time, improving the user experience. However, existing algorithms are either hard to employ in distributed key-value stores due to the relatively large cooperation overhead for centralized information or unable to adapt to the time-varying load and server performance under different traffic patterns. In this paper, we first formalize the scheduling problem for small mean request completion time. As a step further, because of the NP-hardness of this problem, we heuristically design the distributed adaptive scheduler (DAS) for distributed key-value stores. DAS reduces the average request completion time by a distributed combination of the largest remaining processing time last and shortest remaining process time first algorithms. Moreover, DAS is adaptive to the time-varying server load and performance. Extensive simulations show that DAS reduces the mean request completion time by more than 15 ~ 50% compared to the default first come first served algorithm and outperforms the existing Rein-SBF algorithm under various scenarios. Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jianxin Wang 0001, Tong Zhang 0018 |
ICDCS | 2 |
| 2021 | Analysis and improvement of the latency-based congestion control algorithm DX
Wanchun Jiang, Haoyang Li 0006, Lijuan Peng, Jia Wu 0011, Chang Ruan, Jianxin Wang 0001 |
Future Gener. Comput. Syst. | 2 |