EDBT 2026 Demo / reviewers in the wild / expert
Wei Bai 0001
dblp:45/661-1
· DBLP profile ↗
50ranked-venue papers
10as first author
20since 2021 · last 2025
0000-0002-8898-8070ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 40 · 10 first-author · 13 since 2021Systems, architecture and hardware · 8 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unlocking Superior Performance in Reconfigurable Data Center Networks with Credit-Based TransportabstractThe large-scale, end-to-end implementation of microsecond-switched reconfigurable data center networks (RDCNs), coupled with innovative routing and topology designs that provide continuous routes abstracting away frequent topology changes, demonstrates promise as a viable alternative to Clos networks in the post-Moore's Law era. However, the gap remains in transport performance, with current transport solutions falling short of unlocking their full performance. In this paper, we introduce Flare, a novel credit-based transport protocol that ensures reliable traffic delivery, low latency, and leverages the rapidly reconfiguring circuits of the RDCN to opportunistically route traffic over short paths, maximizing throughput. In simulations, Flare enables RDCNs to outperform Clos networks, achieving up to 1.15× higher throughput even under adversarial traffic. Additionally, it delivers up to 2× and 1.5× higher throughput than NDP and ExpressPass, and up to 10×, 15×, and 3.5× shorter flow completion time (FCT) than ExpressPass, TDTCP, and Bolt. Our testbed implementation further demonstrates the feasibility of Flare's mechanisms with programmable switches and DPDK. Federico De Marchi 0002, Jialong Li 0006, Ying Zhang 0022, Wei Bai 0001, Yiting Xia |
SIGCOMM | 4 |
| 2024 | Towards Domain-Specific Network Transport for Distributed DNN Training
Hao Wang 0116, Han Tian, Jingrong Chen 0004, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai 0001, Junchen Jiang, Yong Wang 0046, Kai Chen 0005 |
NSDI | 7 |
| 2024 | Reverie: Low Pass Filter-Based Switch Buffer Sharing for Datacenters with RDMA and TCP Traffic
Vamsi Addanki, Wei Bai 0001, Stefan Schmid 0001, Maria Apostolaki |
NSDI | 2 |
| 2024 | Harmonic: Hardware-assisted RDMA Performance Isolation for Public Clouds
Jiaqi Lou, Xinhao Kong, Jinghan Huang 0001, Wei Bai 0001, Nam Sung Kim, Danyang Zhuo |
NSDI | 4 |
| 2024 | Uniform-Cost Multi-Path Routing for Reconfigurable Data Center NetworksabstractReconfigurable data center networks (RDCNs) are arising as a promising data center network (DCN) design in the post-Moore's law era. However, the constantly reconfigured network topology in RDCNs invalidates the assumption of using hop count as the cost metric for routing, e.g., the status quo Equal-Cost Multi-Path routing (ECMP) in traditional DCNs. Unfortunately, existing routing solutions in RDCNs stick to the old assumption and deliver suboptimal performance either high in latency or low in bandwidth efficiency. In this paper, we redefine the cost metric for RDCN routing with uniform cost to unify the effects of topology disruption and hop count on latency and bandwidth efficiency. We propose Uniform-Cost Multi-Path routing (UCMP), an ECMP equivalent for RDCNs, where minimizing uniform cost leads flows of various sizes to the right balance between latency and bandwidth efficiency. Our simulation shows that UCMP achieves 53% to 98% lower flow completion time (FCT) and 1.55× bandwidth efficiency compared to the state-of-the-art RDCN routing strategy, and our testbed implementation demonstrates sustainable switch resource usage of UCMP as RDCNs scale. Jialong Li 0006, Haotian Gong, Federico De Marchi 0002, Aoyu Gong, Yiming Lei 0002, Wei Bai 0001, Yiting Xia |
SIGCOMM | 6 |
| 2024 | Distributed Network Telemetry With Resource Efficiency and Full AccuracyabstractNetwork telemetry is essential for administrators to monitor massive data traffic in a network-wide manner. Existing telemetry solutions often face the dilemma between resource efficiency (i.e., low CPU, memory, and bandwidth overhead) and full accuracy (i.e., error-free and holistic measurement). We break this dilemma via a network-wide architectural design, which simultaneously achieves resource efficiency and full accuracy in flow-level telemetry for large-scale data centers. carefully coordinates the collaboration among different types of entities in the whole network to execute telemetry operations, such that the resource constraints of each entity are satisfied without compromising full accuracy. It further addresses consistency in network-wide epoch synchronization and accountability in error-free packet loss inference. We prototype in DPDK and P4. Testbed experiments on commodity servers and Tofino switches demonstrate the effectiveness of over state-of-the-art solutions. Haifeng Sun 0004, Qun Huang 0001, Patrick P. C. Lee, Wei Bai 0001, Yungang Bao |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | FlexPass: A Case for Flexible Credit-based Transport for Datacenter NetworksabstractProactive transports explicitly allocate bandwidth to each sender with credits which schedule packet transmission. While promising, existing proactive solutions share a stringent deployment requirement; they assume the perfect control of every link and packet in the network. However, the assumption breaks in practice because new transports are usually deployed gradually over time and legacy traffic always coexists. In this paper, we present FlexPass, a credit-based transport that takes deployment flexibility as a first-class citizen. FlexPass uses a novel combination of network and end-host designs to solve the problem of co-existence and gradual deployment. FlexPass leverages a proactive control loop to send credit-scheduled packets and a complementary reactive control loop to send unscheduled packets to utilize the spare bandwidth. Finally, FlexPass prevents queue buildups of both scheduled and unscheduled packets, and recovers lost packets efficiently. Our evaluation on the testbed shows that FlexPass maintains co-existence with legacy transports (DCTCP), while preserving the high-performance properties of the proactive transport. In large-scale simulations, we show that FlexPass delivers the best incremental benefits during the gradual deployment. We find traffic upgraded to FlexPass benefits from the bounded queue and reduced flow completion time by up to 44% compared to the legacy traffic, while minimizing the side-effect on the legacy flows. Hwijoon Lim, Jaehong Kim 0002, Inho Cho, Keon Jang, Wei Bai 0001, Dongsu Han |
EuroSys | 5 |
| 2023 | Towards a Manageable Intra-Host NetworkabstractIntra-host networks, including heterogeneous devices and interconnect fabrics, have become increasingly complex and crucial. However, intra-host networks today do not provide sufficient manageability. This prevents data center operators from running a reliable and efficient end-to-end network, especially for multi-tenant clouds. In this paper, we analyze the main manageability deficiencies of intra-host networks and argue that a systematic solution should be implemented to bridge this function gap. We propose two key building blocks for a manageable intra-host network: a fine-grained monitoring system and a holistic resource manager. We discuss the research questions associated with realizing these two building blocks. Xinhao Kong, Jiaqi Lou, Wei Bai 0001, Nam Sung Kim, Danyang Zhuo |
HotOS | 3 |
| 2023 | Empowering Azure Storage with RDMA
Wei Bai 0001, Shanim Sainul Abdeen, Ankit Agrawal 0013, Krishan Kumar Attre, Paramvir Bahl, Ameya Bhagat, Gowri Bhaskara, Tanya Brokhman, Ahmad Cheema, Rebecca Chow, Jeff Cohen, Mahmoud Elhaddad, Vivek Ette, Igal Figlin, Daniel Firestone, Mathew George, Ilya German, Lakhmeet Ghai, Eric Green, Albert G. Greenberg, Randy Haagens, Matthew Hendel, Ridwan Howlader, Neetha John, Julia Johnstone, Tom Jolly, Greg Kramer, David Kruse, Erica Lan, Avi Levy, Marina Lipshteyn, Guohan Lu, Yuemin Lu, Xiakun Lu, Vadim Makhervaks, Ulad Malashanka, David A. Maltz, Ilias Marinos, Rohan Mehta, Sharda Murthi, Anup Namdhari, Aaron Ogus, Jitendra Padhye, Madhav Pandya, Douglas Phillips, Adrian Power, Suraj Puri, Shachar Raindel, Jordan Rhee, Anthony Russo, Maneesh Sah, Ali Sheriff, Chris Sparacino, Ashutosh Srivastava, Weixiang Sun, Nick Swanson, Fuhou Tian, Lukasz Tomczyk, Vamsi Vadlamuri, Alec Wolman, Joyce Yom, Yanzhao Zhang, Brian Zill |
NSDI | 1 |
| 2023 | Understanding RDMA Microarchitecture Resources for Performance Isolation
Xinhao Kong, Jingrong Chen 0002, Wei Bai 0001, Yechen Xu, Mahmoud Elhaddad, Shachar Raindel, Jitendra Padhye, Alvin R. Lebeck, Danyang Zhuo |
NSDI | 3 |
| 2023 | Understanding the Micro-Behaviors of Hardware Offloaded Network Stacks with LuminaabstractHardware offloaded network stacks are widely adopted in modern datacenters to meet the demand for high throughput, ultra-low latency and low CPU overhead. To fully leverage their exceptional performance, users need to have a deep understanding of their behaviors. Despite many efforts on testing software network stacks, hardware network stacks impose unique challenges to testing tools due to their kernel bypass nature and high performance. Zhuolong Yu, Wei Bai 0001, Shachar Raindel, Vladimir Braverman, Xin Jin 0008 |
SIGCOMM | 3 |
| 2023 | Enabling ECN for Datacenter Networks With RTT VariationsabstractECN has been widely employed in production datacenters to deliver high throughput low latency communications. Despite being successful, prior ECN-based transports have an important drawback: they adopt a fixed RTT value in calculating instantaneous ECN marking threshold while overlooking the RTT variations in practice. In this paper, we reveal that the current practice of using a fixed high-percentile RTT for ECN threshold calculation can lead to persistent queue buildups, significantly increasing packet latency. On the other hand, directly adopting lower percentile RTTs results in throughput degradation. To handle the problem, we introduce$\sf{ECN}^{\unicode{x266F}}$, a simple yet effective solution to enable ECN for RTT variations. At its heart,$\sf{ECN}^{\unicode{x266F}}$inherits the current instantaneous ECN marking (based on a high-percentile RTT) to achieve high throughput and burst tolerance, while further marking packets (conservatively) upon detecting long-term queue buildups to eliminate unnecessary queueing delay without degrading throughput. We implement$\sf{ECN}^{\unicode{x266F}}$on a Barefoot Tofino switch and evaluate it through extensive testbed experiments and large-scale simulations. Our evaluation confirms that$\sf{ECN}^{\unicode{x266F}}$can effectively reduce latency without hurting throughput. For example, compared to the current practice,$\sf{ECN}^{\unicode{x266F}}$achieves up to$23.4\%$($31.2\%$) lower average (99th percentile) flow completion time (FCT) for short flows while delivering similar FCT for large flows under production workloads. Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005 |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Aeolus: A Building Block for Proactive Transport in Datacenter NetworksabstractAs datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth isproactivelyallocated as “credits” to senders who then can send “scheduled packets” at a right rate to ensure high link utilization, low latency, and zero packet loss. Consequently, proactive solutions such as ExpressPass, NDP, Homa, etc., have been proposed recently. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT “pre-credit” phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on “pre-credit” packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. Aeolus is compatible with all existing proactive solutions and readily implementable with commodity switches. We have integrated Aeolus into ExpressPass, NDP and Homa, and shown, via both implementation and simulations, that the Aeolus-enhanced solutions deliver significant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by$20\times $, while achieving similar performance as NDP without switch modifications. Shuihai Hu, Gaoxiong Zeng, Wei Bai 0001, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run oncross-datacenter networkthat consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network poses significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., buffer depths, RTTs). In this paper, we find that existing DCN or WAN transport reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself only, can simultaneously capture the location and degree of congestion, mainly due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low buffer occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31%, 76% and 2% reduction of small flow average completion times, and up to 34%, 39%, 9% and 58% reduction of large flow average completion times compared to TCP Cubic, DCTCP, BBR and TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2021 | Towards timeout-less transport in commodity datacenter networksabstractDespite recent advances in datacenter networks, timeouts caused by congestion packet losses still remain a major cause of high tail latency. Priority-based Flow Control (PFC) was introduced to make the network lossless, but its Head-of-Line blocking nature causes various performance and management problems. In this paper, we ask if it is possible to design a network that achieves (near) zero timeout only using commodity hardware in datacenters. Hwijoon Lim, Wei Bai 0001, Yibo Zhu 0001, Youngmok Jung, Dongsu Han |
EuroSys | 2 |
| 2021 | 1Pipe: scalable total order communication in data center networksabstractThis paper proposes 1Pipe, a novel communication abstraction that enables different receivers to process messages from senders in a consistent total order. More precisely, 1Pipe provides both unicast and scattering (i.e., a group of messages to different destinations) in a causally and totally ordered manner. 1Pipe provides a best effort service that delivers each message at most once, as well as a reliable service that guarantees delivery and provides restricted atomic delivery for each scattering. 1Pipe can simplify and accelerate many distributed applications, e.g., transactional key-value stores, log replication, and distributed data structures. Bojie Li, Gefei Zuo, Wei Bai 0001 |
SIGCOMM | 3 |
| 2021 | Providing Bandwidth Guarantees, Work Conservation and Low Latency Simultaneously in the CloudabstractToday's cloud is shared among multiple tenants running different applications, and a desirable multi-tenant datacenter network infrastructure should provide bandwidth guarantees for throughput-intensive applications, low latency for latency-sensitive short messages, as well as work conservation to fully utilize the network bandwidth. Despite significant efforts in recent years, none of them can achieve these three properties simultaneously. In this paper, we identify the key deficiency of prior solutions and use this insight to motivate our design of Trinity-a simple, practical yet effective solution that achieves bandwidth guarantees, work conservation and low latency simultaneously in the cloud. We implement Trinity using existing commodity hardwares and demonstrate its superior performance over prior solutions using testbed experiments. Shuihai Hu, Wei Bai 0001, Kai Chen 0005, Chen Tian 0001, Ying Zhang 0022 |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | One More Config is Enough: Saving (DC)TCP for High-Speed Extremely Shallow-Buffered DatacentersabstractThe link speed in production datacenters is growing fast, from 1 Gbps to 40 Gbps or even 100 Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4 MB at 1 Gbps to 16 MB at 100 Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC,1a simple yet effective solution that requires only one more ECN configuration (i.e., shared buffer ECN/RED) at commodity switches. BCC operates upon real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; When it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing a small amount of throughput. BCC is readily deployable with existing commodity switches. We validate BCC's efficacy in a 100G testbed and evaluate its performance using extensive simulations. Our results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while only degrading average FCT for large flows by up to 3%. Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong |
IEEE/ACM Trans. Netw. | 1 |
| 2021 | Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and SchedulingabstractIn this article, we investigate the performance bottleneck of existing deep learning (DL) systems and propose DLBooster to improve the running efficiency of deploying DL applications on GPU clusters. At its core, DLBooster leverages two-level optimizations to boost the end-to-end DL workflow. On the one hand, DLBooster selectively offloads some key decoding workloads to FPGAs to provide high-performance online data preprocessing services to the computing engine. On the other hand, DLBooster reorganizes the computational workloads of training neural networks with the backpropagation algorithm and schedules them according to their dependencies to improve the utilization of GPUs at runtime. Based on our experiments, we demonstrate that compared with baselines, DLBooster can improve the image processing throughput by 1.4× - 2.5× and reduce the processing latency by 1/3 in several real-world DL applications and datasets. Moreover, DLBooster consumes less than 1 CPU core to manage FPGA devices at runtime, which is at least 90 percent less than the baselines in some cases. DLBooster shows its potential to accelerate DL workflows in the cloud. Dan Li 0001, Binyao Jiang, Jinkun Geng, Wei Bai 0001, Yongqiang Xiong |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | RepNet: Cutting Latency with Flow Replication in Data Center NetworksabstractData center networks need to provide low latency, especially at the tail, as demanded by many interactive applications. To improve tail latency, existing approaches require modifications to switch hardware and/or end-host operating systems, making them difficult to be deployed. We present the design, implementation, and evaluation of RepNet, an application layer transport that can be deployed today. RepNet exploits the fact that only a few paths among many are congested at any moment in the network, and applies simple flow replication to mice flows to opportunistically use the less congested path. RepNet has two designs for flow replication: (1) RepSYN, which only replicates SYN packets and uses the first connection that finishes TCP handshaking for data transmission, and (2) RepFlow which replicates the entire mice flow. We implement RepNet on node.js, one of the most commonly used platforms for networked interactive applications. node's single threaded event-loop and non-blocking I/O make flow replication highly efficient. Performance evaluation on a real network testbed and in Mininet reveals that RepNet is able to reduce the tail latency of mice flows, as well as application completion times, by more than 50 percent. Shuhao Liu 0001, Hong Xu 0001, Libin Liu 0001, Wei Bai 0001, Kai Chen 0005, Zhiping Cai |
IEEE Trans. Serv. Comput. | 4 |
| 2020 | One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered DatacentersabstractThe link speed in production datacenters is growing fast, from 1Gbps to 40Gbps or even 100Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4MB at 1Gbps to 16MB at 100Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet loss or substantial throughput degradation.To this end, we present BCC1, a simple yet effective solution that requires just one more ECN config (i.e., shared buffer ECN/RED) over prior solutions. BCC operates based on real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; Once it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing little throughput. BCC is readily deployable with existing commodity switches. We validate BCC's hardware feasibility in a small 100G testbed and evaluate its performance using large-scale simulations. Our results show that BCC maintains low packet loss rate while slightly degrading throughput when the available buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while degrading average FCT for large flows by up to 3%. Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong |
INFOCOM | 1 |
| 2020 | Aeolus: A Building Block for Proactive Transport in DatacentersabstractAs datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth is proactively allocated as "credits" to senders who then can send "scheduled packets" at a right rate to ensure high link utilization, low latency, and zero packet loss. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT "pre-credit" phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on "pre-credit" packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. We have integrated Aeolus into ExpressPass[14], NDP[18] and Homa[29], and shown, through both implementation and simulations, that the Aeolus-enhanced solutions deliver signiicant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by 20x, while achieving similar performance as NDP without switch modifications. Shuihai Hu, Wei Bai 0001, Gaoxiong Zeng, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004 |
SIGCOMM | 2 |
| 2020 | OmniMon: Re-architecting Network Telemetry with Resource Efficiency and Full AccuracyabstractNetwork telemetry is essential for administrators to monitor massive data traffic in a network-wide manner. Existing telemetry solutions often face the dilemma between resource efficiency (i.e., low CPU, memory, and bandwidth overhead) and full accuracy (i.e., error-free and holistic measurement). We break this dilemma via a network-wide architectural design OmniMon, which simultaneously achieves resource efficiency and full accuracy in flow-level telemetry for large-scale data centers. OmniMon carefully coordinates the collaboration among different types of entities in the whole network to execute telemetry operations, such that the resource constraints of each entity are satisfied without compromising full accuracy. It further addresses consistency in network-wide epoch synchronization and accountability in error-free packet loss inference. We prototype OmniMon in DPDK and P4. Testbed experiments on commodity servers and Tofino switches demonstrate the effectiveness of OmniMon over state-of-the-art telemetry designs. Qun Huang 0001, Haifeng Sun 0004, Patrick P. C. Lee, Wei Bai 0001, Yungang Bao |
SIGCOMM | 4 |
| 2019 | Rethinking Transport Layer Design for Distributed Machine LearningabstractMotivated by the increasing scale of data, we see a growing need of high performance distributed machine learning systems. Many research works are being proposed to improve distributed machine learning performance. Jiacheng Xia, Gaoxiong Zeng, Junxue Zhang 0001, Weiyan Wang, Wei Bai 0001, Junchen Jiang, Kai Chen 0005 |
APNet | 5 |
| 2019 | Enabling ECN for datacenter networks with RTT variationsabstractECN has been widely employed in production datacenters to deliver high throughput low latency communications. Despite being successful, prior ECN-based transports have an important drawback: they adopt a fixed RTT value in calculating instantaneous ECN marking threshold while overlooking the RTT variations in practice. Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005 |
CoNEXT | 2 |
| 2019 | FlowShader: a Generalized Framework for GPU-accelerated VNF Flow ProcessingabstractGPU acceleration has been widely investigated for packet processing in virtual network functions (NFs), but not for L7 flow-processing NFs. In L7 NFs, reassembled TCP messages of the same flow should be processed in order in the same processing thread, and the uneven sizes among flows pose a major challenge for full realization of GPU's parallel computation power. To exploit GPUs for L7 NF processing, this paper presents FlowShader, a GPU acceleration framework to achieve both high generality and throughput even under skewed flow size distributions. We carefully design an efficient scheduling algorithm that fully exploits available GPU and CPU capacities; in particular, we dispatch large flows which seriously break up the size balance to CPU and the rest of flows to GPU. Furthermore, FlowShader allows similar NF logic (as CPU-based NFs) to run on individual threads in a GPU, which is more generalized and easy to take on as compared to redesigning an NF for operation parallelism on GPU. We implemented a number of L7 flow processing NFs based on FlowShader. Evaluations are conducted under both synthetic and real-world traffic traces and results show that the throughput achieved by FlowShader is up to 6x that of the CPU-only baseline and 3x of the GPU-only design. Xiaodong Yi 0001, Jingpu Duan, Wei Bai 0001, Chuan Wu 0001, Yongqiang Xiong, Dongsu Han |
ICNP | 4 |
| 2019 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run on cross-datacenter network that consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network imposes significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., butter depths, RTTs). In this paper, we find that existing DCN or WAN transports reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself, can simultaneously capture the location and degree of congestion. This is due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low butter occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31% and 76% reduction of small flow average completion times compared to TCP Cubic, DCTCP and BBR; and up to 58% reduction of large flow average completion times compared to TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
ICNP | 2 |
| 2019 | DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing PipelinesabstractIn recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference. Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong |
ICPP | 9 |
| 2019 | Socksdirect: datacenter sockets can be fast and compatibleabstractCommunication intensive applications in hosts with multi-core CPU and high speed networking hardware often put considerable stress on the native socket system in an OS. Existing socket replacements often leave significant performance on the table, as well have limitations on compatibility and isolation. Bojie Li, Tianyi Cui, Wei Bai 0001 |
SIGCOMM | 4 |
| 2019 | Accelerating Rule-matching Systems with Learned Rankers
Zhao Lucis Li, Chieh-Jan Mike Liang, Wei Bai 0001, Yongqiang Xiong, Guangzhong Sun |
USENIX ATC | 3 |
| 2018 | Augmenting Proactive Congestion Control with AeolusabstractRecently, proactive congestion control solutions have drawn great attention in the community. By explicitly scheduling data transmissions based on the availability of network bandwidth, proactive solutions offer a lossless, near-zero queueing network for serving network transfers. Despite the advantages, proactive solutions require an extra RTT to allocate the ideal sending rate for new arrival flows. To resolve this, current solutions let new flows blindly transmit unscheduled packets in the first RTT, and assign these packets with high priority in the network. The unscheduled packets, however, can cause serious network congestion, resulting in large queue buildups and excessive packet losses. Shuihai Hu, Wei Bai 0001, Baochen Qiao, Kai Chen 0005, Kun Tan 0002 |
APNet | 2 |
| 2017 | Congestion Control for High-speed Extremely Shallow-buffered Datacenter NetworksabstractThe link speed in datacenters is growing fast, from 1Gbps to 100Gbps. However, the buffer size of commodity switches increases slowly, thus significantly outpaced by the link speed. In such extremely shallow-buffered datacenter networks, prior TCP/ECN solutions suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC, a simple yet effective solution with only one more configuration (shared buffer ECN/RED) at commodity switches. BCC operates based on real-time shared buffer utilization. When the buffer is abundant, BCC delivers both high throughput and low packet loss rate. When it becomes scarce, BCC triggers shared buffer ECN/RED to prevent packet losses at the cost of sacrificing a small amount of throughput. Our preliminary results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. Compared to current practice, BCC achieves up to 94.4% lower 99th percentile completion time for small flows while only degrading large flows by up to 2.8%. Wei Bai 0001, Kai Chen 0005, Shuihai Hu, Kun Tan 0002, Yongqiang Xiong |
APNet | 1 |
| 2017 | Combining ECN and RTT for Datacenter TransportabstractDatacenter transports should provide low average and tail flow completion times (FCT) to achieve desired application performance. While most prior datacenter transports take either ECN or RTT as congestion signal, this paper makes a case that both signals are indispensable: ECN, as a per-hop signal, is more effective to prevent packet loss; while RTT, as an end-to-end signal, controls end-to-end queueing delay better. As persistent low flow completion times imply low queueing delay and near zero packet loss, we introduce EAR, a new datacenter transport that hears and reacts to both ECN and RTT. Our preliminary results show that: 1) compared to delay-based DCTCP, EAR achieves up to 91% lower packet losses and 93% fewer timeouts; 2) compared to ECN-based DCTCP, EAR reduces RTT by up to 32% for cross-rack traffic in a 4-level fattree. As a result, EAR delivers persistent low average and tail completion times under various scenarios in large scale simulations. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
APNet | 2 |
| 2017 | Rate-aware flow scheduling for commodity data center networksabstractFlow completion times (FCTs) are critical for many cloud applications. To minimize the average FCT, recent transport designs, such as pFabric, PASE, and PIAS, approximate the Shortest Remaining Time First (SRTF) scheduling. A common, implicit assumption of these solutions is that the remaining time is only determined by the remaining flow size. However, this assumption does not hold in many real-world scenarios where applications generate data at diverse rates that are smaller than the network capacity. In this paper, we look into this issue from system perspective and find that the operating system (OS) kernel can be exploited to better estimate the remaining time of a flow. In particular, we use the rate of copying data from user space to kernel space to measure the data generation rate. We design RAX, a rate aware flow scheduling method, that calculates the remaining time of a flow more accurately, based on not only the flow size but also the data generation rate. We have implemented a RAX prototype in Linux kernel and evaluated it through testbed experiments and ns-2 simulations. Our testbed results show that RAX reduces FCT by up to 14.9%/41.8% and 7.8%/22.9% over DCTCP and PIAS for all/medium flows respectively. Ziyang Li 0003, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yiming Zhang 0003, Dongsheng Li 0001, Hong-Fang Yu |
INFOCOM | 2 |
| 2017 | Resilient Datacenter Load Balancing in the WildabstractProduction datacenters operate under various uncertainties such as traffic dynamics, topology asymmetry, and failures. Therefore, datacenter load balancing schemes must be resilient to these uncertainties; i.e., they should accurately sense path conditions and timely react to mitigate the fallouts. Despite significant efforts, prior solutions have important drawbacks. On the one hand, solutions such as Presto and DRB are oblivious to path conditions and blindly reroute at fixed granularity. On the other hand, solutions such as CONGA and CLOVE can sense congestion, but they can only reroute when flowlets emerge; thus, they cannot always react timely to uncertainties. To make things worse, these solutions fail to detect/handle failures such as blackholes and random packet drops, which greatly degrades their performance. Hong Zhang 0025, Junxue Zhang 0001, Wei Bai 0001, Kai Chen 0005, Mosharaf Chowdhury |
SIGCOMM | 3 |
| 2017 | PIAS: Practical Information-Agnostic Flow Scheduling for Commodity Data CentersabstractMany existing data center network (DCN) flow scheduling schemes, that minimize flow completion times (FCT) assume prior knowledge of flows and custom switch functions, making them superior in performance but hard to implement in practice. By contrast, we seek to minimize FCT with no prior knowledge and existing commodity switch hardware. To this end, we present PIAS, a DCN flow scheduling mechanism that aims to minimize FCT by mimicking shortest job first (SJF) on the premise that flow size is not knowna priori. At its heart, PIAS leverages multiple priority queues available in existing commodity switches to implement a multiple level feedback queue, in which a PIAS flow is gradually demoted from higher-priority queues to lower-priority queues based on the number of bytes it has sent. As a result, short flows are likely to be finished in the first few high-priority queues and thus be prioritized over long flows in general, which enables PIAS to emulate SJF without knowing flow sizes beforehand. We have implemented a PIAS prototype and evaluated PIAS through both testbed experiments and ns-2 simulations. We show that PIAS is readily deployable with commodity switches and backward compatible with legacy TCP/IP stacks. Our evaluation results show that PIAS significantly outperforms existing information-agnostic schemes, for example, it reduces FCT by up to 50% compared to DCTCP[11]and L2DCT[32]; and it only has a 1.1% performance gap to an ideal information-aware scheme, pFabric[13], for short flows under a production DCN workload. Wei Bai 0001, Li Chen 0008, Kai Chen 0005, Dongsu Han, Chen Tian 0001, Hao Wang 0022 |
IEEE/ACM Trans. Netw. | 1 |
| 2017 | Guaranteeing Deadlines for Inter-Data Center TransfersabstractInter-data center wide area networks (inter-DC WANs) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. In this paper, we address the problem by introducing a deadline-based network abstraction (DNA) for inter-DC WANs. DNA allows users to explicitly specify the amount of data to be delivered and the deadline by which it has to be completed. The malleability of DNA provides flexibility in resource allocation. Based on this, we develop a system calledAmoebathat implements DNA. Our simulations and test bed experiments show thatAmoeba, by harnessing DNA’s malleability, accommodates 15% more user requests with deadlines, while achieving 60% higher WAN utilization than prior solutions. Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005 |
IEEE/ACM Trans. Netw. | 3 |
| 2016 | Enabling ECN over Generic Packet SchedulingabstractExplicit Congestion Notification (ECN) is crucial for production datacenters, but current queue-length based ECN/RED implementation does not work with generic packet schedulers, leading to either degraded network performance or violated scheduling policies. In this paper, we first dive into this issue and reveal that the invalidity of ECN/RED lies in the difficulty of measuring changing queue capacities under various schedulers and traffic dynamics. Then we present Time-based Congestion Notification (TCN), a simple yet effective ECN solution, by combining two successful ideas: the sojourn time from CoDel and the instantaneous marking from DCTCP. Using packet sojourn-time, as opposed to queue-length, as the congestion signal, TCN eliminates the need of measuring dynamic queue capacities, making it suitable for arbitrary schedulers with traffic dynamics. By performing stateless instantaneous ECN marking rather than complex stateful dropping, TCN is designed to be inexpensive to implement on commodity switching chips. Through extensive testbed experiments and large-scale simulations, we show TCN can strictly preserve scheduling policies while providing desirable network performance. For example, TCN significantly reduces the average and 99th percentile completion times for small flows by up to 82.8% and 95.3% compared to current practice in a testbed experiment with production workload. Wei Bai 0001, Kai Chen 0005, Li Chen 0008, Changhoon Kim |
CoNEXT | 1 |
| 2016 | Providing bandwidth guarantees, work conservation and low latency simultaneously in the cloudabstractToday's cloud is shared among multiple tenants running different applications, and a desirable multi-tenant datacenter network infrastructure should provide bandwidth guarantees for throughput-intensive applications, low latency for latency-sensitive short messages, as well as work conservation to fully utilize the network bandwidth. Despite significant efforts in recent years, none of them can achieve these three properties simultaneously. In this paper, we identify the key deficiency of prior solutions and use this insight to motivate our design of Trinity - a simple, practical yet effective solution that achieves bandwidth guarantees, work conservation and low latency simultaneously in the cloud. We implement Trinity using existing commodity hardwares and demonstrate its superior performance over prior solutions using testbed experiments. Shuihai Hu, Wei Bai 0001, Kai Chen 0005, Chen Tian 0001, Ying Zhang 0022 |
INFOCOM | 2 |
| 2016 | Enabling ECN in Multi-Service Multi-Queue Data Centers
Wei Bai 0001, Li Chen 0008, Kai Chen 0005 |
NSDI | 1 |
| 2016 | Scheduling Mix-flows in Commodity Datacenters with KarunaabstractCloud applications generate a mix of flows with and without deadlines. Scheduling such mix-flows is a key challenge; our experiments show that trivially combining existing schemes for deadline/non-deadline flows is problematic. For example, prioritizing deadline flows hurts flow completion time (FCT) for non-deadline flows, with minor improvement for deadline miss rate. Li Chen 0008, Kai Chen 0005, Wei Bai 0001, Mohammad Alizadeh |
SIGCOMM | 3 |
| 2016 | Explicit Path Control in Commodity Data Centers: Design and ApplicationsabstractMany data center network DCN applications require explicit routing path control over the underlying topologies. In this paper, we present XPath, a simple, practical and readily-deployable way to implement explicit path control, using existing commodity switches. At its core, XPath explicitly identifies an end-to-end path with a path ID and leverages a two-step compression algorithm to pre-install all the desired paths into IP TCAM tables of commodity switches. Our evaluation and implementation show that XPath scales to large DCNs and is readily-deployable. Furthermore, on our testbed, we integrate XPath into four applications to showcase its utility. Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 4 |
| 2016 | Towards Comprehensive Traffic Forecasting in Cloud Computing: Design and ApplicationabstractIn this paper, we present our effort towards comprehensive traffic forecasting for big data applications using external, light-weighted file system monitoring. Our idea is motivated by the key observations that rich traffic demand information already exists in the log and meta-data files of many big data applications, and that such information can be readily extracted through run-time file system monitoring. As the first step, we use Hadoop as a concrete example to explore our methodology and develop a system called HadoopWatch to predict traffic demands of Hadoop applications. We further implement HadoopWatch in a small-scale testbed with 10 physical servers and 30 virtual machines. Our experiments over a series of MapReduce applications demonstrate that HadoopWatch can forecast the traffic demand with almost 100% accuracy and time advance. Furthermore, it makes no modification on the Hadoop framework, and introduces little overhead to the application performance. Finally, to showcase the utility of accurate traffic prediction made by HadoopWatch, we design and implement a simple HadoopWatch-enabled network optimization module into the HadoopWatch controller, and with realistic Hadoop job benchmarks we find that even a simple algorithm can leverage the forecasting results provided by HadoopWatch to significantly improve the Hadoop job completion time by up to 14.72%. Kai Chen 0005, Wei Bai 0001, Yangming Zhao, Hao Wang 0022, Yanhui Geng, Zhiqiang Ma 0002, Lin Gu 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2015 | Guaranteeing deadlines for inter-datacenter transfersabstractInter-datacenter wide area networks (inter-DC WAN) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005 |
EuroSys | 3 |
| 2015 | Rapier: Integrating routing and scheduling for coflow-aware data center networksabstractIn the data flow models of today's data center applications such as MapReduce, Spark and Dryad, multiple flows can comprise a coflow group semantically. Only completing all flows in a coflow is meaningful to an application. To optimize application performance, routing and scheduling must be jointly considered at the level of a coflow rather than individual flows. However, prior solutions have significant limitation: they only consider scheduling, which is insufficient. To this end, we present Rapier, a coflow-aware network optimization framework that seamlessly integrates routing and scheduling for better application performance. Using a small-scale testbed implementation and large-scale simulations, we demonstrate that Rapier significantly reduces the average coflow completion time (CCT) by up to 79.30% compared to the state-of-the-art scheduling-only solution, and it is readily implementable with existing commodity switches. Yangming Zhao, Kai Chen 0005, Wei Bai 0001, Minlan Yu, Chen Tian 0001, Yanhui Geng, Yiming Zhang 0003, Dan Li 0001, Sheng Wang 0006 |
INFOCOM | 3 |
| 2015 | Information-Agnostic Flow Scheduling for Commodity Data Centers
Wei Bai 0001, Kai Chen 0005, Hao Wang 0022, Li Chen 0008, Dongsu Han, Chen Tian 0001 |
NSDI | 1 |
| 2015 | Explicit Path Control in Commodity Data Centers: Design and Applications
Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo |
NSDI | 4 |
| 2014 | PIAS: Practical Information-Agnostic Flow Scheduling for Data Center NetworksabstractMany existing data center network (DCN) flow scheduling schemes minimize flow completion times (FCT) based on prior knowledge of flows and custom switch designs, making them hard to use in practice. This paper introduces, Pias, a practical flow scheduling approach that minimizes FCT with no prior knowledge using commodity switches. At its heart, Pias leverages multiple priority queues available in commodity switches to implement a Multiple Level Feedback Queue (MLFQ), in which a PIAS flow gradually demotes from higher-priority queues to lower-priority queues based on the bytes it has sent. In this way, short flows are prioritized over long flows, which enables Pias to emulate Shortest Job First (SJF) scheduling without knowing the flow sizes beforehand. Our preliminary evaluation shows that Pias significantly outperforms all existing information-agnostic solutions. It improves average FCT for short flows by up to 50% and 40% over DCTCP [3] and L2DCT [16]. Compared to an ideal information-aware DCN transport, p-Fabric [5], it only shows 4.9% performance degradation for short flows in a production datacenter workload. Wei Bai 0001, Li Chen 0008, Kai Chen 0005, Dongsu Han, Chen Tian 0001, Weicheng Sun |
HotNets | 1 |
| 2014 | PAC: Taming TCP Incast Congestion Using Proactive ACK ControlabstractTCP in cast congestion which can introduce hundreds of milliseconds delay and up to 90% throughput degradation, severely affecting application performance, has been a practical issue in high-bandwidth low-latency data enter networks. Despite continuous efforts, prior solutions have significant drawbacks. They either only support quite a limited number of senders (e.g., 40-60), which is not sufficient, or require non-trivial system modifications, which is impractical and not incrementally deployable. We present PAC, a simple yet very effective design to tame TCP in cast congestion via Proactive ACK Control at the receiver. The key design principle behind PAC is that we treat ACK not only as the acknowledgement of received packets but also as the trigger for new packets. Leveraging data center network characteristics, PAC enforces a novel ACK control to release ACKs in such a way that the ACK-triggered in-flight data can fully utilize the bottleneck link without causing in cast collapse even when faced with over a thousand senders. We implement PAC on both Windows and Linux platforms, and extensively evaluate PAC using small-scale test bed experiments and large-scale ns-2 simulations. Our results show that PAC significantly outperforms the previous representative designs such as ICTCP and DCTCP by supporting 40X (i.e., 40?1600) more senders, further, it does not introduce spurious timeout and retransmission even when the measured 99th percentile RTT is only 3.6ms. Our implementation experiences show that PAC is readily deployable in production data enters, while requiring minimal system modification compared to prior designs. Wei Bai 0001, Kai Chen 0005, Wuwei Lan, Yangming Zhao |
ICNP | 1 |
| 2014 | HadoopWatch: A first step towards comprehensive traffic forecasting in cloud computingabstractThis paper presents our effort towards comprehensive traffic forecasting for big data applications using external, light-weighted file system monitoring. Our idea is motivated by the key observations that rich traffic demand information already exists in the log and meta-data files of many big data applications, and that such information can be readily extracted through run-time file system monitoring. As the first step, we use Hadoop1 as a concrete example to explore our methodology and develop a system called HadoopWatch to predict traffic demand of Hadoop applications. We further implement HadoopWatch in our real small-scale testbed with 10 physical servers and 30 virtual machines. Our experiments over a series of MapReduce applications demonstrate that HadoopWatch can forecast the traffic demand with almost 100% accuracy and time advance. Furthermore, it makes no modification of the Hadoop framework, and introduces little overhead to the application performance. Kai Chen 0005, Wei Bai 0001, Zhiqiang Ma 0002, Lin Gu 0001 |
INFOCOM | 4 |