Xin Zhe Khooi

dblp:155/4219 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
13since 2021 · last 2026
0009-0005-8048-8556ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Handling Network Faults in Distributed AI Training: Failover is Now an Option
abstract
Distributed AI training often suffers from network faults. Network faults, especially at the last hop between a switch and a host, result in loss of connectivity, resulting in training job stalls and eventual failure. This is typically managed through a fail-stop mechanism, followed by a restart, incurring significant inefficiencies. We present ReCCL, the first network fault-tolerant collective communication library (CCL) that allows training progress to be preserved by seamlessly failing over to alternate paths when a network fault occurs. During failover, ReCCL keeps communication states synchronized while using dynamic channel load balancing and intra-host GPU routing to improve communication performance. Our evaluations demonstrate that ReCCL can perform failover seamlessly with minimal performance losses. Additionally, our simulations also demonstrate that failover can be effectively used to achieve significant savings in GPU hours for large-scale distributed AI training workloads.
Xin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui, Meng Wang 0018, Yuze Jin, Pengfei Huo, Lulu Chen, Liaoyuan Feng, Qinlong Wang, Yongcan Wang, Jinshuai Sun, Yingkai Zhao, Haiquan Chen 0002, Yi Li 0098, Jianxi Ye, Mun Choon Chan
EuroSys1
2026 How to Hardware Accelerate Your 5G CU
Xin Zhe Khooi, Satis Kumar Permal, Cha Hwan Song, Nishant Budhdev, Raj Joshi, Mun Choon Chan
INFOCOM1
2025 How to Update Your 5G vRAN
abstract
The ongoing virtualization of Radio Access Networks (vRANs) promises increased velocity for updating RAN software with new features and bug fixes. To support this, we developed SwapRAN, a live update system for both vRAN software components: the Distributed Unit (DU) and the Centralized Unit (CU). Unlike previous systems, SwapRAN operates in-place without requiring additional hardware like staging servers or a programmable switch. For DU updates, SwapRAN presents two techniques: (1) using OS thread priorities to safely initialize the new DU while overlapping with the old DU which is active, and (2) using the network interface card's embedded switch to redirect fronthaul traffic to the new DU. For CU updates, SwapRAN is the first working live update system, which we achieve by (1) decoupling the stateful CU-DU connection and transparently rerouting DU messages to the new CU, and (2) repurposing existing midhaul control plane messages to move users to the new CU. We evaluate SwapRAN on real 5G testbeds and demonstrate its practical deployability via integration with Kubernetes. Our evaluations show that SwapRAN completes DU or CU updates with just 1–2 seconds of user downtime.
Xin Zhe Khooi, Anuj Kalia, Mun Choon Chan
MobiCom1
2025 Demo: Towards Seamless 5G vRAN Software Updates
abstract
We present SwapRAN, a live update system that brings Continuous Integration/Continuous Deployment (CI/CD) to virtualized RANs (vRANs), which includes both the Distributed Unit (DU) and the Centralized Unit (CU). In contrast to prior solutions, SwapRAN performs in-place software updates without relying on additional infrastructure such as staging servers or programmable switches. For DU updates, SwapRAN introduces two techniques: (1) leveraging OS thread priorities to safely bring up the new DU while the old DU remains active, and (2) redirecting fronthaul traffic to the new DU using the embedded switch found on modern network interface cards. For CU updates, SwapRAN is the first system to enable live updates, made possible by (1) decoupling the stateful connection between the CU and DU and transparently rerouting DU messages to the new CU, and (2) repurposing existing midhaul control plane messages to transfer users to the new CU. We demonstrate SwapRAN on our O-RAN testbed equipped with a commercial O-RU, showing that it can perform DU or CU updates with significantly reduced downtime, as low as 1–2 seconds, compared to existing update strategies in Kubernetes.
Xin Zhe Khooi, Anuj Kalia, Mun Choon Chan
MobiCom1
2024 JUNCTION: A Scalable Multi-Access Solution Using Programmable Switches
abstract
Multi-access networks are increasingly important for reliable end-to-end connectivity and enhanced throughput performance. A scalable multi-access solution is required to roll out multi-access networks at scale. However, existing CPU-based solutions can no longer scale sustainably, as network traffic has outgrown the CPU performance growth. Consequently, hardware accelerators offer a compelling alternative. This paper introduces JUNCTION, a scalable multi-access solution designed using programmable switches. JUNCTION features a multipath protocol tailored to the hardware constraints and optimized for efficient memory utilization, enabling it to handle a large number of multipath sessions. We validate JUNCTION on a 5G-WiFi multi-access testbed. Our analysis demonstrates that it can scale an order of magnitude better than existing solutions.
Xin Zhe Khooi, Cha Hwan Song, Satis Kumar Permal, Nishant Budhdev, Levente Csikor, Raj Joshi, Mun Choon Chan
SECON1
2023 Masking Corruption Packet Losses in Datacenter Networks with Link-local Retransmission
abstract
Packet loss due to link corruption is a major problem in large warehouse-scale datacenters. The current state-of-the-art approach of disabling corrupting links is not adequate because, in practice, all the corrupting links cannot be disabled due to capacity constraints. In this paper, we show that, it is feasible to implement link-local retransmission at sub-RTT timescales to completely mask corruption packet losses from the transport endpoints. Our system, LinkGuardian, employs a range of techniques to (i) keep the packet buffer requirement low, (ii) recover from tail packet losses without employing timeouts, and (iii) preserve packet ordering. We implement LinkGuardian on the Intel Tofino switch and show that for a 100G link with a loss rate of 10−3, LinkGuardian can reduce the loss rate by up to 6 orders of magnitude while incurring only 8% reduction in effective link speed. By eliminating tail packet losses, LinkGuardian improves the 99.9th percentile flow completion time (FCT) for TCP and RDMA by 51x and 66x respectively. Finally, we also show that in the context of datacenter networks, simple out-of-order retransmission is often sufficient to significantly mitigate the impact of corruption packet loss for short TCP flows.
Raj Joshi, Cha Hwan Song, Xin Zhe Khooi, Nishant Budhdev, Ayush Mishra, Mun Choon Chan, Ben Leong
SIGCOMM3
2023 Poster: Towards Accelerating the 5G Centralized Unit with Programmable Switches
abstract
5G networks are envisioned to support various emerging use cases, such as telemedicine, remote construction, autonomous driving, industrial automation, drone control, and immersive entertainment. These applications demand low latency, high reliability, and in some cases require ultra-high-bandwidths. Specifically, these applications require 5G networks to provide 1ms end-to-end latency with 99.99% reliability [12] for ultra-reliable low-latency communications (URLLC). Various studies [11, 19, 26] have shown that the radio access network (RAN) remains the bottleneck in realizing low-latency communications.
Xin Zhe Khooi, Archit Bhatnagar, Satis Kumar Permal, Nishant Budhdev, Cha Hwan Song, Mun Choon Chan
SIGCOMM1
2023 Network Load Balancing with In-network Reordering Support for RDMA
abstract
Remote Direct Memory Access (RDMA) is widely used in high-performance computing (HPC) and data center networks. In this paper, we first show that RDMA does not work well with existing load balancing algorithms because of its traffic flow characteristics and assumption of in-order packet delivery. We then propose ConWeave, a load balancing framework designed for RDMA. The key idea of ConWeave is that with the right design, it is possible to perform fine granularity rerouting and mask the effect of out-of-order packet arrivals transparently in the network datapath using a programmable switch. We have implemented ConWeave on a Tofino2 switch. Evaluations show that ConWeave can achieve up to 42.3% and 66.8% improvement for average and 99-percentile FCT, respectively compared to the state-of-the-art load balancing algorithms.
Cha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi, Jialin Li 0001, Mun Choon Chan
SIGCOMM2
2023 NeoBFT: Accelerating Byzantine Fault Tolerance Using Authenticated In-Network Ordering
abstract
Mission critical systems deployed in data centers today are facing more sophisticated failures. Byzantine fault-tolerant (BFT) protocols are capable of masking these types of failures, but are rarely deployed due to their performance cost and complexity. In this work, we propose a new approach to designing high performance BFT protocols in data centers. By re-examining the ordering responsibility between the network and the BFT protocol, we advocate a new abstraction offered by the data center network infrastructure. Concretely, we design a new authenticated ordered multicast primitive (aom) that provides transferable authentication and non-equivocation guarantees. Feasibility of the design is demonstrated by two hardware implementations of aom- one using HMAC and the other using public key cryptography for authentication - on new-generation programmable switches. We then co-design a new BFT protocol, NeoBFT, that leverages the guarantees of aom to eliminate cross-replica coordination and authentication in the common case. Evaluation results show that NeoBFT outperforms state-of-the-art protocols on both latency and throughput metrics by a wide margin, demonstrating the benefit of our new network ordering abstraction for BFT systems.
Guangda Sun, Mingliang Jiang, Xin Zhe Khooi, Jialin Li 0001
SIGCOMM3
2023 DySO: Enhancing application offload efficiency on programmable switches
abstract
Application offloads on modern high-speed programmable switches have been proposed in a variety of systems (e.g., key–value store systems and network middleboxes) so as to efficiently scale up the traditional server-oriented deployments. However, they largely achieve sub-optimal offloading efficiency due to the lack of (1) capability to perform control actions at sufficient rates, and (2) adaptability to workload changes. In this paper, we scrutinize the common stumbling blocks of existing frameworks with performance evaluations on real workloads. We present DySO (Dynamic State Offloading), a framework which enables expeditious on-demand control actions and self-tuning of management rules. DySO’s key insight is to perform control actions via a data-path instead of the switch control channel which is the bottleneck to read/write states into data plane. Our software simulations show up to 100% performance improvement compared to existing systems for various real world traces. On top of that, we implement and evaluate DySO on a commodity programmable switch, showing two orders of magnitude faster responsiveness to sudden workload changes compared to the existing systems.
Cha Hwan Song, Xin Zhe Khooi, Dinil Mon Divakaran, Mun Choon Chan
Comput. Networks2
2021 Towards a Framework for One-sided RDMA Multicast
abstract
We present the design and prototyping of a framework to support multicast for remote direct memory accesses (RDMA), specifically the one-sided WRITE operation. We use P4 programmable hardware to augment fixed-function RDMA transport hardware found on commodity NICs to enable one-sided RDMA multicast with zero-CPU overhead. Finally, we outline the potential challenges and future directions in realizing the framework for large-scale data center deployments.
Xin Zhe Khooi, Cha Hwan Song, Mun Choon Chan
ANCS1
2021 In-Network Applications: Beyond Single Switch Pipelines
abstract
The emergence of commodity programmable switches have spawned a series of innovations in the network data plane. By making the traditionally stateless network architectures to be stateful, we can realize a diverse set of applications, e.g., networking monitoring, load-balancing, firewalls, entirely in the data plane. On the other hand, many existing in-network applications assume that the underlying switch is single-pipelined, however, in reality, commodity programmable switches are designed with multiple pipelines in mind. While this approach enables high scalability, it has introduced a serious disadvantage: maintaining states across the pipelines is non-trivial. For instance, without involving the control plane it is infeasible to keep track of a request and its response in different pipelines, thereby rendering many in-network proposals impractical.In this paper, we highlight this fundamental limitation that holds back the practical widespread adoption of stateful applications in today’s multi-pipeline switches. By scrutinizing recent in-network approaches, we identify that majority of them cannot operate as they are proposed on multi-pipeline switches. After raising awareness of this inevitable consequence, we discuss a set of possible workarounds for in-network applications to overcome this issue on multi-pipeline switches.
Xin Zhe Khooi, Levente Csikor, Jialin Li 0001, Dinil Mon Divakaran
NetSoft1
2021 Revisiting Heavy-Hitter Detection on Commodity Programmable Switches
abstract
Existing in-network heavy-hitter detection algorithms suffer from several shortcomings. On the one hand, most of the algorithms perform monitoring in intervals and reset the data structures in between; consequently, a notable amount of heavy hitters (HH) spanning across the intervals go undetected. On the other hand, the algorithms consume substantial hardware resources, potentially hindering other data plane functionalities to be integrated on the same device.In this work, we revisit the state-of-the-art in-network approaches in this regard and identify that they fall short in over-coming the aforementioned issues. In particular, we investigate whether it is possible to design a heavy-hitter detection algorithm that provides high accuracy without consuming substantial re-sources, thereby making it feasible to integrate with concurrent applications. To this end, we propose dSketch, a time-decaying algorithm for in-network heavy-hitter detection. Trace-driven simulations and evaluations on the Intel Tofino-based commodity switches show that dSketch significantly improves the detection rate of HHs by 5–10% while being resource- and operation-efficient in contrast to state-of-the-art approaches. Moreover, we show that dSketch can be integrated with standard switch functionalities such as switch. p4 with additional resources spared, offering itself as a compelling solution for switch data plane designers.
Xin Zhe Khooi, Levente Csikor, Jialin Li 0001, Min Suk Kang, Dinil Mon Divakaran
NetSoft1
2020 DIDA: Distributed In-Network Defense Architecture Against Amplified Reflection DDoS Attacks
abstract
With each new DDoS attack potentially becoming a higher intensity attack than the previous ones, current ISP measures of over-provisioning or employing a scrubbing service are becoming ineffective and inefficient. We argue that we need an in-network solution (i.e., entirely in the data plane), to detect DDoS attacks, identify the corresponding traffic and mitigate promptly. In this paper, we propose the first distributed in-network defense architecture, DIDA, to cope with the sophisticated amplified reflection DDoS (AR-DDoS) attacks. We leverage programmable stateful data planes and efficient data structures and show that it is possible to keep track of per-user connections in an automated and distributed manner without overwhelming the network controller. Building on top of this data, DIDA can easily detect if unsolicited attack packets are sent towards a victim within an ISP network. Once an attack is detected, the routers at the network edge automatically block the malicious sources. We prototype DIDA in P4. Our preliminary experiments show that DIDA can detect and mitigate 99.8% of amplification attacks containing 7, 000 different sources while requiring less than 1% of the memory of current programmable switches.
Xin Zhe Khooi, Levente Csikor, Dinil Mon Divakaran, Min Suk Kang
NetSoft1
2014 Human Visualisation of Cryptographic Code Using Progressive Multi-Scale Resolution
abstract
High-entropy codewords frequently occur in the context of cryptographic protocols, and typically range in length from 128 to 1024 bits. Human vision is not well equipped to compare or recognise such codewords, due to the high information (length) and entropy content. In this paper, we propose a human visualisation mechanism to enable representation of long high-entropy codewords via perceptually significant visual images. The main contribution is a mechanism capable of representation at more detailed scales of resolution in progressive steps, so as to allow human visual inspection which is both secure and ergonomic. The generation of these visual representations is either dependent on key-specific or context-sensitive system inputs. The featured representation allows for machine-to-human authentication and authorization.
Alwyn Goh, Geong Sen Poh, Voon-Yee Vee, Kok Boon Chong, Xin Zhe Khooi, Chanan Zhuo Ern Loh, Zhi Yuan Eng
SIN5