Kun Tan 0002

dblp:42/5784-2 · DBLP profile ↗
← Back
28ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0009-0413-363XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 9 since 2021Systems, architecture and hardware · 2Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Advancing RDMA Scalability With High Performance
abstract
Due to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections. We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end in a trusted network. Leveraging the asymmetric communication pattern in data center applications, StaRlets the communication end with low NIC memory usage to save states for the other end with high NIC memory usage, thus making the RNIC on the bottleneck side stateless. We implemented StaR on an FPGA board with a 10Gbps network port and NS-3, evaluating its performance on a testbed with 9 machines, each equipped with StaR NICs, and verified its scalability stability by conducting a larger-scale simulation with 200 fully connected nodes using a 100Gbps link. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively.
Xijin Yin, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002
IEEE Trans. Netw.8
2025 PRED: Performance-oriented Random Early Detection for Consistently Stable Performance in Datacenters
Xinle Du, Tong Li 0014, Guangmeng Zhou, Zhuotao Liu, Hanlin Huang, Mowei Wang, Kun Tan 0002, Ke Xu 0002
NSDI8
2025 Revisiting Random Early Detection Tuning for High-Performance Datacenter Networks
abstract
Random Early Detection (RED) has been integrated into datacenter switches as a fundamental Active Queue Management (AQM) for decades. The accurate configuration of RED parameters is crucial to achieving high throughput and low latency. However, due to the highly dynamic nature of workloads in datacenter networks, maintaining consistently high performance with statically configured RED thresholds poses a challenge. Prior work applies reinforcement learning to predict proper thresholds, but their real-world deployment has been hindered by poor tail performance caused by instability. In this paper, we propose$\textsf {PRED}$, a novel system that enables automatic and stable RED parameter adjustment in response to traffic dynamics. Specifically, the system employs a Multiplicative-Increase Multiplicative-Decrease (MIMD) strategy to dynamically adapt to flow concurrency while utilizing an Additive-Increase Additive-Decrease (AIAD) mechanism to adapt to flow distribution. We perform extensive evaluations on our physical testbed and large-scale simulations. The results demonstrate that$\textsf {PRED}$can keep up with the real-time network dynamics generated by realistic workloads. For instance, compared with the static-threshold-based methods,$\textsf {PRED}$keeps 66% shorter switch queue length and obtains up to 80% lower Flow Completion Time (FCT). Compared with the state-of-the-art learning-based method,$\textsf {PRED}$reduces the tail FCT by 34%.
Tong Li 0014, Xinle Du, Guangmeng Zhou, Hanlin Huang, Zhuotao Liu, Mowei Wang, Kun Tan 0002, Ke Xu 0002
IEEE Trans. Netw.8
2023 FastWake: Revisiting Host Network Stack for Interrupt-mode RDMA
abstract
Polling and interrupt has long been a trade-off in RDMA systems. Polling has lower latency but each CPU core can only run one thread. Interrupt enables time sharing among multiple threads but has higher latency. Many applications such as databases have hundreds of threads, which is much larger than the number of cores. So, they have to use interrupt mode to share cores among threads, and the resulting RDMA latency is much higher than the hardware limits. In this paper, we analyze the root cause of high costs in RDMA interrupt delivery, and present FastWake, a practical redesign of interrupt-mode RDMA host network stack using commodity RDMA hardware, Linux OS, and unmodified applications. Our first approach to fast thread wake-up completely removes interrupts. We design a per-core dispatcher thread to poll all the completion queues of the application threads on the same core, and utilize a kernel fast path to context switch to the thread with an incoming completion event. The approach above would keep CPUs running at 100% utilization, so we design an interrupt-based approach for scenarios with power constraints. Observing that waking up a thread on the same core as the interrupt is much faster than threads on other cores, we dynamically adjust RDMA event queue mappings to improve interrupt core affinity. In addition, we revisit the kernel path of thread wake-up, and remove the overheads in virtual file system (VFS), locking, and process scheduling. Experiments show that FastWake can reduce RDMA latency by 80% on x86 and 77% on ARM at the cost of < 30% higher power utilization than traditional interrupts, and the latency is only 0.3 ∼ 0.4 μ s higher than the limits of underlying hardware. When power saving is desired, our interrupt-based approach can still reduce interrupt-mode RDMA latency by 59% on x86 and 52% on ARM.
Bojie Li, Zihao Xiang, Xiaoliang Wang 0001, Han Ruan, Jingbin Zhou, Kun Tan 0002
APNet6
2022 Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN Training
abstract
Bulk synchronous parallel (BSP) is the de-facto paradigm for distributed DNN training in today’s production clusters. However, due to the global synchronization nature, its performance can be significantly influenced by network bottlenecks caused by either static topology heterogeneity or dynamic bandwidth contentions. Existing solutions, either system-level optimizations strengthening BSP (e.g., Ring or Hierarchical All-reduce) or algorithmic optimizations replacing BSP (e.g., ASP or SSP, which relax the global barriers), do not completely solve the problem, as they may still suffer from communication inefficiency or risk convergence inaccuracy.In this paper, we present a novel divide-and-shuffle synchronization (DS-Sync) to realize communication efficiency without sacrificing convergence accuracy for distributed DNN training. At its heart, by taking into account the network bottlenecks, DS-Sync improves communication efficiency by dividing workers into non-overlap groups to synchronize independently in a bottleneck-free manner. Meanwhile, it maintains convergence accuracy by iteratively shuffling workers among different groups to ensure a global consensus. We theoretically prove that DS-Sync converges properly in non-convex and smooth conditions like DNN. We further implement DS-Sync and integrate it with PyTorch, and our testbed experiments show that DS-Sync can achieve up to 94% improvements on the end-to-end training time with existing solutions while maintaining the same accuracy.
Weiyan Wang, Cengguang Zhang, Liu Yang 0008, Kai Chen 0005, Kun Tan 0002
INFOCOM5
2022 Aeolus: A Building Block for Proactive Transport in Datacenter Networks
abstract
As datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth isproactivelyallocated as “credits” to senders who then can send “scheduled packets” at a right rate to ensure high link utilization, low latency, and zero packet loss. Consequently, proactive solutions such as ExpressPass, NDP, Homa, etc., have been proposed recently. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT “pre-credit” phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on “pre-credit” packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. Aeolus is compatible with all existing proactive solutions and readily implementable with commodity switches. We have integrated Aeolus into ExpressPass, NDP and Homa, and shown, via both implementation and simulations, that the Aeolus-enhanced solutions deliver significant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by$20\times $, while achieving similar performance as NDP without switch modifications.
Shuihai Hu, Gaoxiong Zeng, Wei Bai 0001, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004
IEEE/ACM Trans. Netw.7
2021 StaR: Breaking the Scalability Limit for RDMA
abstract
Due to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections.We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end. Leveraging the asymmetric communication pattern in data center applications, StaR lets the communication end with low concurrency save states for the other end with high concurrency, thus making the RNIC on the bottleneck side to be stateless. We have implemented StaR on an FPGA board with 10Gbps network port and evaluated its performance on a testbed with 9 machines all equipped with StaR NICs. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively.
Xizheng Wang, Guo Chen 0001, Xijin Yin, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002
ICNP7
2021 S2Net: Preserving Privacy in Smart Home Routers
abstract
At present, wireless home routers are becoming increasingly smart. While these smart routers provide rich functionalities to users, they also raise security concerns. Although the existing end-to-end encryption techniques can be applied to protect personal data, such rich functionalities become unavailable due to the encrypted payloads. On the other hand, if the smart home routers are allowed to process and store the personal data of users, once compromised, the users' sensitive data will be exposed. As a consequence, users face a difficult trade-off between the benefits of the rich functionalities and potential privacy risks. To deal with this dilemma, we propose a novel system named Secure and Smart Network (S2Net) for home routers. For S2Net, we propose a secure OS that can distinguish and manage multiple sessions belonging to different users. The secure OS and all the router applications are placed in the secure world using the ARM TrustZone technology. In S2Net, we also confine the router applications in sandboxes provided by the proposed secure OS to prevent data leakage. As a result, S2Net can provide rich functionalities for users while preserving strong privacy for home routers. In addition, we develop a crypto-worker model that provides an abstraction layer of cryptographic tasks performed by a heterogeneous multi-core system. The other important role of crypto-worker is to parallelize the computations in order to resolve the high computation cost of cryptographic functions. We report the system design of S2Net and the details of our implementation. Experimental results with benchmarks and real applications demonstrate that our implementation is capable of achieving high performance in terms of throughput while mitigating the overhead of S2Net design.
SeungSeob Lee, Kun Tan 0002, Yunxin Liu 0001, Yong Cui 0001
IEEE Trans. Dependable Secur. Comput.3
2021 One More Config is Enough: Saving (DC)TCP for High-Speed Extremely Shallow-Buffered Datacenters
abstract
The link speed in production datacenters is growing fast, from 1 Gbps to 40 Gbps or even 100 Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4 MB at 1 Gbps to 16 MB at 100 Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC,1a simple yet effective solution that requires only one more ECN configuration (i.e., shared buffer ECN/RED) at commodity switches. BCC operates upon real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; When it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing a small amount of throughput. BCC is readily deployable with existing commodity switches. We validate BCC's efficacy in a 100G testbed and evaluate its performance using extensive simulations. Our results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while only degrading average FCT for large flows by up to 3%.
Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong
IEEE/ACM Trans. Netw.4
2021 Revisiting Acknowledgment Mechanism for Transport Control: Modeling, Analysis, and Implementation
abstract
The shared nature of the wireless medium induces contention between data transport and backward signaling, such as acknowledgment. The current way of TCP acknowledgment induces control overhead which is counter-productive for TCP performance especially in wireless local area network (WLAN) scenarios. In this paper, we present a new acknowledgment called TACK (“Tame ACK”), as well as its TCP implementation TCP-TACK. TACK seeks to minimize ACK frequency, which is exactly what is required by transport. TCP-TACK works on top of commodity WLAN, delivering high wireless transport goodput with minimal control overhead in the form of ACKs, without any hardware modification. Evaluation results show that TCP-TACK achieves significant advantages over legacy TCP in WLAN scenarios due to less contention between data packets and ACKs. Specifically, TCP-TACK reduces over 90% of ACKs and also obtains an improvement of up to 28% on goodput. A TACK-based protocol is a good replacement of the legacy TCP to compensate for scenarios where the acknowledgment overhead is non-negligible.
Tong Li 0014, Kai Zheng 0003, Ke Xu 0002, Rahul Arvind Jadhav, Keith Winstein, Kun Tan 0002
IEEE/ACM Trans. Netw.7
2020 One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered Datacenters
abstract
The link speed in production datacenters is growing fast, from 1Gbps to 40Gbps or even 100Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4MB at 1Gbps to 16MB at 100Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet loss or substantial throughput degradation.To this end, we present BCC1, a simple yet effective solution that requires just one more ECN config (i.e., shared buffer ECN/RED) over prior solutions. BCC operates based on real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; Once it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing little throughput. BCC is readily deployable with existing commodity switches. We validate BCC's hardware feasibility in a small 100G testbed and evaluate its performance using large-scale simulations. Our results show that BCC maintains low packet loss rate while slightly degrading throughput when the available buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while degrading average FCT for large flows by up to 3%.
Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong
INFOCOM4
2020 MasQ: RDMA for Virtual Private Cloud
abstract
RDMA communication in virtual private cloud (VPC) networks is still a challenging job due to the difficulty in fulfilling all virtualization requirements without sacrificing RDMA communication performance. To address this problem, this paper proposes a software-defined solution, namely, MasQ, which is short for "queue masquerade". The core insight of MasQ is that all RDMA communications should associate with at least one queue pair (QP). Thus, the requirements of virtualization, such as network isolation and the application of security rules, can be easily fulfilled if QP's behavior is properly defined. In particular, MasQ exploits the virtio-based paravirtualization technique to realize the control path. Moreover, to avoid performance overhead, MasQ leaves all data path operations, such as sending and receiving, to the hardware. We have implemented MasQ in the OpenFabrics Enterprise Distribution (OFED) framework and proved its scalability and performance efficiency by evaluating it against typical applications. The results demonstrate that MasQ achieves almost the same performance as bare-metal RDMA for data communication.
Binzhang Fu, Kun Tan 0002, Bei Hua, Zhi-Li Zhang, Kai Zheng 0003
SIGCOMM4
2020 Aeolus: A Building Block for Proactive Transport in Datacenters
abstract
As datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth is proactively allocated as "credits" to senders who then can send "scheduled packets" at a right rate to ensure high link utilization, low latency, and zero packet loss. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT "pre-credit" phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on "pre-credit" packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. We have integrated Aeolus into ExpressPass[14], NDP[18] and Homa[29], and shown, through both implementation and simulations, that the Aeolus-enhanced solutions deliver signiicant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by 20x, while achieving similar performance as NDP without switch modifications.
Shuihai Hu, Wei Bai 0001, Gaoxiong Zeng, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004
SIGCOMM7
2020 TACK: Improving Wireless Transport Performance by Taming Acknowledgments
abstract
The shared nature of the wireless medium induces contention between data transport and backward signaling, such as acknowledgement. The current way of TCP acknowledgment induces control overhead which is counter-productive for TCP performance especially in wireless local area network (WLAN) scenarios.
Tong Li 0014, Kai Zheng 0003, Ke Xu 0002, Rahul Arvind Jadhav, Keith Winstein, Kun Tan 0002
SIGCOMM7
2019 Towards Stateless RNIC for Data Center Networks
abstract
Because of small NIC on-chip memory, the massive connection states maintained on Remote Direct Memory Access (RDMA) NIC (RNIC) significantly limit its scalability. When the number of concurrent connections grows, RNICs have to frequently fetch connection states from host memory, leading to dramatic performance degradation. In this paper, we propose StaR, which fundamentally solves this scalability issue by making RNIC stateless. Leveraging the asymmetric communication pattern in data center applications, the StaR RNIC stores zero connection-related states by moving all the connection states to the other end. Through careful design, StaR RNICs can maintain unchanged RDMA semantics and avoid security issues even when processing traffic statelessly. Preliminary simulation results show that StaR can improve the aggregate throughput by more than 160x (stress test) and 4x (application) compared to original RNICs.
Pulin Pan, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002
APNet7
2019 MP-RDMA: Enabling RDMA With Multi-Path Transport in Datacenters
abstract
RDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%.
Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda
IEEE/ACM Trans. Netw.4
2019 Wireless Network Instabilities in the Wild: Measurement, Applications (Non)Resilience, and OS Remedy
abstract
While the bandwidth and latency improvement of both WiFi and cellular data networks in the past decades are plenty evident, the extent of signal strength fluctuation and network disruptions (unexpected switching or disconnections) experienced by mobile users in today's network deployment remains less clear. This paper makes three contributions. First, we conduct the first extensive measurement of network disruptions and significant signal strength fluctuations (together denoted as network instabilities) experienced by 2000 smartphones in the wild. Our results show that network disruptions and signal strength fluctuations remains prevalent as we moved into the 4G era. Second, we study how well popular mobile apps today handle such network instabilities. Our results show that even some of the most popular mobile apps do not implement any disruption-tolerant mechanisms. Third, we present Janus, an intelligent interface management framework that exploits the multiple interfaces on a handset to transparently handle network disruptions and satisfy apps' performance requirement. We have implemented a prototype of Janus and our evaluation using a set of popular apps shows that Janus can: 1) transparently and efficiently handle network disruptions; 2) reduce video stalls by 2.9 times and increase 31% of the time of good voice quality; 3) reduce traffic size by 26.4% and energy consumption by 16.3% compared to naive solutions.
Yong Cui 0001, Zeqi Lai, Y. Charlie Hu, Kun Tan 0002, Minglong Dai, Kai Zheng 0003, Yi Li 0015
IEEE/ACM Trans. Netw.6
2018 Augmenting Proactive Congestion Control with Aeolus
abstract
Recently, proactive congestion control solutions have drawn great attention in the community. By explicitly scheduling data transmissions based on the availability of network bandwidth, proactive solutions offer a lossless, near-zero queueing network for serving network transfers. Despite the advantages, proactive solutions require an extra RTT to allocate the ideal sending rate for new arrival flows. To resolve this, current solutions let new flows blindly transmit unscheduled packets in the first RTT, and assign these packets with high priority in the network. The unscheduled packets, however, can cause serious network congestion, resulting in large queue buildups and excessive packet losses.
Shuihai Hu, Wei Bai 0001, Baochen Qiao, Kai Chen 0005, Kun Tan 0002
APNet5
2018 Building Generic Scalable Middlebox Services Over Encrypted Protocols
abstract
The trends of the increasing middleboxes make the middle network more and more complex. Today, many middleboxes work on application layer and offer significant network services by the plain-text traffic, such as firewalling, intrusion detecting and application layer gateways. At the same time, more and more network applications are encrypting their data transmission to protect security and privacy. It is becoming a critical task and hot topic to continue providing application-layer middlebox services in the encrypted Internet, however, the state of the art is far from being able to be deployed in the real network. In this paper, we propose a practical architecture, named PlainBox, to enable session key sharing between the communication client and the middleboxes in the network path. It employs Attribute-Based Encryption (ABE) in the key sharing protocol to support multiple chaining middleboxes efficiently and securely. We develop a prototype system and apply it to popular security protocols such as TLS and SSH. We have tested our prototype system in a lab testbed as well as real-world websites. Our result shows PlainBox introduces very little overhead and the performance is practically deployable.
Cong Liu 0029, Yong Cui 0001, Kun Tan 0002, Quan Fan, Kui Ren 0001
INFOCOM3
2018 Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda
NSDI4
2018 FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
IEEE/ACM Trans. Netw.5
2017 Congestion Control for High-speed Extremely Shallow-buffered Datacenter Networks
abstract
The link speed in datacenters is growing fast, from 1Gbps to 100Gbps. However, the buffer size of commodity switches increases slowly, thus significantly outpaced by the link speed. In such extremely shallow-buffered datacenter networks, prior TCP/ECN solutions suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC, a simple yet effective solution with only one more configuration (shared buffer ECN/RED) at commodity switches. BCC operates based on real-time shared buffer utilization. When the buffer is abundant, BCC delivers both high throughput and low packet loss rate. When it becomes scarce, BCC triggers shared buffer ECN/RED to prevent packet losses at the cost of sacrificing a small amount of throughput. Our preliminary results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. Compared to current practice, BCC achieves up to 94.4% lower 99th percentile completion time for small flows while only degrading large flows by up to 2.8%.
Wei Bai 0001, Kai Chen 0005, Shuihai Hu, Kun Tan 0002, Yongqiang Xiong
APNet4
2017 Wireless network instabilities in the wild: Prevalence, App (non)resilience, and OS remedy
abstract
While the bandwidth and latency improvement of both WiFi and cellular data networks in the past decade are plenty evident, the extent of signal strength fluctuation and network disruptions (unexpected switching or disconnections) experienced by mobile users in today's network deployment remains less clear. This paper makes three contributions. First, we conduct the first extensive measurement of network disruptions and signal strength fluctuations (together denoted as instabilities) experienced by 2000 smartphones in the wild. Our results show that network disruptions and signal strength fluctuations remain prevalent as we moved into the 4G era. Second, we study how well popular mobile apps today handle such network instabilities. Our results show that even some of the most popular mobile apps do not implement any disruption-tolerant mechanisms. Third, we present JANUS, an intelligent interface management framework that exploits the multiple interfaces on a handset to transparently handle network disruptions and improve apps' QoE. We have implemented JANUS on Android and our evaluation using a set of popular apps shows that Janus can (1) transparently and efficiently handle network disruptions, (2) reduce video stalls by 2.9 times and increase 31% of the time of good voice quality compared to naive solutions.
Zeqi Lai, Yong Cui 0001, Y. Charlie Hu, Kun Tan 0002, Minglong Dai, Kai Zheng 0003
ICNP6
2017 One more queue is enough: Minimizing flow completion time with explicit priority notification
abstract
Ideally, minimizing the flow completion time (FCT) requires millions of priorities supported by the underlying network so that each flow has its unique priority. However, in production datacenters, the available switch priority queues for flow scheduling are very limited (merely 2 or 3). This practical constraint seriously degrades the performance of previous approaches. In this paper, we introduce Explicit Priority Notification (EPN), a novel scheduling mechanism which emulates fine-grained priorities (i.e., desired priorities or DP) using only two switch priority queues. EPN can support various flow scheduling disciplines with or without flow size information. We have implemented EPN on commodity switches and evaluated its performance with both testbed experiments and extensive simulations. Our results show that, with flow size information, EPN achieves comparable FCT as pFabric that requires clean-slate switch hardware. And EPN also outperforms TCP by up to 60.5% if it bins the traffic into two priority queues according to flow size. In information-agnostic setting, EPN outperforms PIAS with two priority queues by up to 37.7%. To the best of our knowledge, EPN is the first system that provides millions of priorities for flow scheduling with commodity switches.
Yuanwei Lu, Guo Chen 0001, Larry Luo, Kun Tan 0002, Yongqiang Xiong, Xiaoliang Wang 0001, Enhong Chen
INFOCOM4
2017 Latency-based WiFi congestion control in the air for dense WiFi networks
abstract
WiFi has become the primary method to access the Internet. However, the WiFi-hop latency, particularly in dense-WiFi environments, is far from satisfactory [1], to support delay-sensitive applications such as Web browsing and VoIP. The WiFi latency mainly comes from two kinds of queues: the host queue and the distributed queue, which is caused by CSMA/CA mechanism when multiple nodes contend for the channel. While the host queue can be easily bypassed using priority scheduling at end-host, the distributed queue is not. Previously, IEEE 802.11e tries to provide priorities in this distributed queue by adjusting the MAC layer parameters, but it does not scale when there are increasing number of delay-sensitive flows. In this paper, we propose and design QAir, a practical solution to reduce WiFi latency of delay-sensitive flows in dense WiFi networks. QAir takes a different approach to transfer this distributed queue to host queue. Consequently, the delay-sensitive flows can bypass the entire queue and their latency can be greatly reduced. QAir works in a distributed manner with no centralized scheduler. We have implemented QAir on commodity WiFi devices. Experimental results show that, compared to the 802.11 DCF baseline, QAir can reduce the average WiFi-hop latency of delay-sensitive flows by 50-75%.
Changhua Pei, Youjian Zhao, Yunxin Liu 0001, Kun Tan 0002, Jiansong Zhang 0001, Yuan Meng 0002, Dan Pei
IWQoS4
2017 Protego: Cloud-Scale Multitenant IPsec Gateway
Jeongseok Son, Yongqiang Xiong, Kun Tan 0002, Ze Gan, Sue B. Moon
USENIX ATC3
2016 ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable Hardware
abstract
Highly flexible software network functions (NFs) are crucial components to enable multi-tenancy in the clouds. However, software packet processing on a commodity server has limited capacity and induces high latency. While software NFs could scale out using more servers, doing so adds significant cost. This paper focuses on accelerating NFs with programmable hardware, i.e., FPGA, which is now a mature technology and inexpensive for datacenters. However, FPGA is predominately programmed using low-level hardware description languages (HDLs), which are hard to code and difficult to debug. More importantly, HDLs are almost inaccessible for most software programmers. This paper presents ClickNP, a FPGA-accelerated platform for highly flexible and high-performance NFs with commodity servers. ClickNP is highly flexible as it is completely programmable using high-level C-like languages, and exposes a modular programming abstraction that resembles Click Modular Router. ClickNP is also high performance. Our prototype NFs show that they can process traffic at up to 200 million packets per second with ultra-low latency ($< 2\mu$s). Compared to existing software counterparts, with FPGA, ClickNP improves throughput by 10x, while reducing latency by 10x. To the best of our knowledge, ClickNP is the first FPGA-accelerated platform for NFs, written completely in high-level language and achieving 40 Gbps line rate at any packet size.
Bojie Li, Kun Tan 0002, Layong Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, Peng Cheng 0005
SIGCOMM2
2016 Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
USENIX ATC5