Anuj Kalia

dblp:149/9171 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
10since 2021 · last 2025
0009-0004-1586-655XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 4 first-author · 9 since 2021Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2025 How to Update Your 5G vRAN
abstract
The ongoing virtualization of Radio Access Networks (vRANs) promises increased velocity for updating RAN software with new features and bug fixes. To support this, we developed SwapRAN, a live update system for both vRAN software components: the Distributed Unit (DU) and the Centralized Unit (CU). Unlike previous systems, SwapRAN operates in-place without requiring additional hardware like staging servers or a programmable switch. For DU updates, SwapRAN presents two techniques: (1) using OS thread priorities to safely initialize the new DU while overlapping with the old DU which is active, and (2) using the network interface card's embedded switch to redirect fronthaul traffic to the new DU. For CU updates, SwapRAN is the first working live update system, which we achieve by (1) decoupling the stateful CU-DU connection and transparently rerouting DU messages to the new CU, and (2) repurposing existing midhaul control plane messages to move users to the new CU. We evaluate SwapRAN on real 5G testbeds and demonstrate its practical deployability via integration with Kubernetes. Our evaluations show that SwapRAN completes DU or CU updates with just 1–2 seconds of user downtime.
Xin Zhe Khooi, Anuj Kalia, Mun Choon Chan
MobiCom2
2025 Demo: Towards Seamless 5G vRAN Software Updates
abstract
We present SwapRAN, a live update system that brings Continuous Integration/Continuous Deployment (CI/CD) to virtualized RANs (vRANs), which includes both the Distributed Unit (DU) and the Centralized Unit (CU). In contrast to prior solutions, SwapRAN performs in-place software updates without relying on additional infrastructure such as staging servers or programmable switches. For DU updates, SwapRAN introduces two techniques: (1) leveraging OS thread priorities to safely bring up the new DU while the old DU remains active, and (2) redirecting fronthaul traffic to the new DU using the embedded switch found on modern network interface cards. For CU updates, SwapRAN is the first system to enable live updates, made possible by (1) decoupling the stateful connection between the CU and DU and transparently rerouting DU messages to the new CU, and (2) repurposing existing midhaul control plane messages to transfer users to the new CU. We demonstrate SwapRAN on our O-RAN testbed equipped with a commercial O-RU, showing that it can perform DU or CU updates with significantly reduced downtime, as low as 1–2 seconds, compared to existing update strategies in Kubernetes.
Xin Zhe Khooi, Anuj Kalia, Mun Choon Chan
MobiCom2
2025 Towards Energy Efficient 5G vRAN Servers
Anuj Kalia, Nikita Lazarev, Leyang Xue, Xenofon Foukas, Bozidar Radunovic, Francis Y. Yan
NSDI1
2024 Savannah: Efficient mmWave Baseband Processing with Minimal and Heterogeneous Resources
abstract
5G new radio (NR) employs frequency range 2 (FR2) in the millimeter-wave (mmWave) bands, which employs a much shorter slot duration compared to FR1 (sub-7 GHz) systems and, therefore, poses significant challenges for softwarized baseband processing in virtualized radio access networks (vRANs). Existing systems supporting software baseband processing focus on enabling (massive) multiple-input and multiple-output (MIMO) using multi-core edge server(s). These solutions may fail to meet the more stringent processing deadline in FR2 or require more intensive computational resources. In this paper, we present Savannah, an efficient mmWave baseband processing framework using minimal and heterogeneous computing resources including CPU and eASIC. Savannah addresses the challenges associated with baseband processing in FR2 by applying techniques for vectorizing matrix operations and memory access patterns, supporting heterogeneous computation via offloading LDPC decoding to an eASIC, and enabling single-core operation. We show that Savannah, using a single CPU core and the ACC100 accelerator, can support a 2×2 MIMO link with 100 MHz bandwidth, yielding a data rate of up to 487 Mbps.
Zhenzhou Qi, Chung-Hsuan Tung, Anuj Kalia, Tingjun Chen
MobiCom3
2024 Savannah: A Real-time Programmable mmWave Baseband Processing Framework
abstract
5G new radio (NR) frequency range 2 (FR2) in the millimeter-wave (mmWave) band has a much shorter baseband processing deadline compared to that in the sub-7 GHz FR1 band. This tight deadline requires an efficient real-time system for baseband processing using minimal computational resources. We demonstrate Savannah, a software framework for efficient mmWave baseband processing using minimal and heterogeneous computing resources, including CPU and eASIC. Savannah vectorizes matrix operations and memory access patterns in multi-input multi-output (MIMO) arithmetic, offloads low-density parity-check (LDPC) coding to an eASIC, and enables single-core operation. We demonstrate that Savannah, using a single CPU core and an eASIC, can support a 2×2 MIMO link with 100 MHz bandwidth under full uplink traffic load, yielding a data rate of up to 487 Mbps.
Zhenzhou Qi, Chung-Hsuan Tung, Anuj Kalia, Tingjun Chen
MobiCom3
2024 Honeycomb: Ordered Key-Value Store Acceleration on an FPGA-Based SmartNIC
abstract
In-memory ordered key-value stores are an important building block in modern distributed applications. We present Honeycomb, a hybrid software-hardware system for accelerating read-dominated workloads on ordered key-value stores that provides linearizability for all operations including scans. Honeycomb stores a B-Tree in host memory. It executesput,updateanddeleteon a CPU. At the same time, it offloadsscanandgetonto an FPGA-based SmartNIC. This approach enables large stores and simplifies the FPGA implementation but raises the challenge of data access and synchronization across the slow PCIe bus. We describe how Honeycomb overcomes this challenge with careful data structure design, caching, request parallelism with out-of-order execution, wait-free read operations, and fast synchronization between the CPU and the FPGA. For read-heavy YCSB workloads, Honeycomb increases the throughput of a state-of-the-art ordered key-value store by at least$1.8\times$. For scan-heavy workloads inspired by cloud storage, Honeycomb increases the throughput by more than$2\times$. The cost-performance, which is more important for large-scale deployments, is improved by at least$1.5\times$on these workloads.
Aleksandar Dragojevic, Shane T. Fleming, Antonios Katsarakis, Dario Korolija, Igor Zablotchi, Ho-Cheung Ng, Anuj Kalia, Miguel Castro 0001
IEEE Trans. Computers8
2023 Accelerating Open RAN Research Through an Enterprise-scale 5G Testbed
abstract
Open RAN is an emerging paradigm in mobile networks where the Radio Access Network (RAN) functions are disaggregated and virtualized on commodity servers. Despite the importance of Open RAN research, existing platforms often lack the fidelity and stability required to address a wide range of research problems. In response to this limitation, we have developed an enterprise-scale Open RAN testbed aimed at conducting state-of-the-art research in key areas that have received limited attention due to the lack of suitable platforms. In this poster, we provide an overview of the testbed we have created and examples of the research it has enabled, with the hope of catalyzing future open RAN research and innovation.
Paramvir Bahl, Matthew Balkwill, Xenofon Foukas, Anuj Kalia, Daehyeok Kim, Manikanta Kotaru, Zhihua Lai, Sanjeev Mehrotra, Bozidar Radunovic, Stefan Saroiu, Connor Settle, Alec Wolman, Francis Y. Yan, Yongguang Zhang
MobiCom4
2023 Enabling Resilience in Virtualized RANs with Atlas
abstract
Virtualized radio access networks (vRANs), which allow running RAN processing on commodity servers instead of proprietary hardware, are gaining adoption in cellular networks. Two properties of the vRAN's "Distributed Unit (DU)" that implements the lower RAN layers---its real-time deadlines and its black-box nature---make it challenging to provide resilience features such as upgrades and failover without long service disruptions. These properties preclude the use of existing resilience techniques like virtual machine migration or state replication that are used for typical workloads. This paper presents Atlas, the first system that provides resilience for the DU. The central insight in Atlas is to repurpose existing cellular mechanisms for wireless resilience, namely handovers and cell reselection, to provide software resilience for the DU. For planned resilience events like upgrades, we design a novel technique that simultaneously serves cells from both the old and new DUs via the same radio, and uses handovers between these cells to migrate user devices. For unplanned failures, we identify deficiencies in existing RAN protocols that disrupt cell reselection after DU failure, and show how we can eliminate these disruptions using a middlebox between the DU and higher layers. Our evaluation with a state-of-the-art 5G vRAN testbed shows that Atlas achieves minimal disruption to cellular connectivity during resilience events, while incurring low overhead.
Jiarong Xing, Junzhi Gong, Xenofon Foukas, Anuj Kalia, Daehyeok Kim, Manikanta Kotaru
MobiCom4
2023 Scalable Distributed Massive MIMO Baseband Processing
Junzhi Gong, Anuj Kalia, Minlan Yu
NSDI2
2023 Resilient Baseband Processing in Virtualized RANs with Slingshot
abstract
In cellular networks, there is a growing adoption of virtualized radio access networks (vRANs), where operators are replacing the traditional specialized hardware for RAN processing with software running on commodity servers. Today's vRAN deployments lack resilience, since there is no support for vRAN failover or upgrades without long service interruptions. Enabling these features for vRANs is challenging because of their strict real-time latency requirements and black-box nature. Slingshot is a new system that transparently provides resilience for the vRAN's most performance-critical layer: the physical layer (PHY). We design new techniques for realtime workload migration with fast RAN protocol middle-boxes, and realtime RAN failure detection. A key insight in our design is to view the transient disruptions from resilience events to RAN computation state and I/O similarly to regular wireless signal impairments, and leverage the inherent resilience of cellular networks to these events. Experiments with a state-of-the-art 5G vRAN testbed show that Slingshot handles PHY failover with no disruption to video conferencing, and under 110 ms disruption to a TCP connection, and it also enables zero-downtime upgrades.
Nikita Lazarev, Anuj Kalia, Daehyeok Kim, Ilias Marinos, Francis Y. Yan, Christina Delimitrou, Zhiru Zhang, Aditya Akella
SIGCOMM3
2020 Challenges and solutions for fast remote persistent memory access
abstract
Non-volatile main memory DIMMs (NVMMs), such as Intel's Optane DC Persistent Memory modules, provide data durability with orders of magnitude higher performance than prior durable technologies. This paper explores the unique challenges that arise when building high-performance networked systems for NVMM. Compared to DRAM, we find that NVMMs have distinctive fundamental properties that pose unique challenges for networked access to NVMM, both from the NIC and the CPU. We show that much of the challenges in efficient access to remote NVMM arises from the fact that CPU caches are not optimized for NVMM. To address these challenges, we propose a menu of solutions for current hardware and evaluate their benefits.
Anuj Kalia, David G. Andersen, Michael Kaminsky
SoCC1
2020 Agora: Real-time massive MIMO baseband processing in software
abstract
Massive multiple-input multiple-output (MIMO) is a key technology in 5G New Radio (NR) to improve spectral efficiency. A major challenge in its realization is the huge amount of real-time computation required. All existing massive MIMO baseband processing solutions use dedicated and specialized hardware like FPGAs, which can efficiently process baseband data but are expensive, inflexible and difficult to program. In this paper, we show that a software-only system called Agora can handle the high computational demand of real-time massive MIMO baseband processing on a single many-core server. To achieve this goal, we identify the rich dimensions of parallelism in massive MIMO baseband processing, and exploit them across multiple CPU cores. We optimize Agora to best use CPU hardware and software features, including SIMD extensions to accelerate computation, cache optimizations to accelerate data movement, and kernel-bypass packet I/O. We evaluate Agora with up to 64 antennas and show that it meets the data rate and latency requirements of 5G NR.
Rahman Doost-Mohammady, Anuj Kalia, Lin Zhong 0001
CoNEXT3
2020 Lightweight Preemptible Functions
Sol Boucher, Anuj Kalia, David G. Andersen, Michael Kaminsky
USENIX ATC2
2019 Datacenter RPCs can be General and Fast
Anuj Kalia, Michael Kaminsky, David G. Andersen
NSDI1
2018 Putting the "Micro" Back in Microservice
Sol Boucher, Anuj Kalia, David G. Andersen, Michael Kaminsky
USENIX ATC2
2016 FaSST: Fast, Scalable and Simple Distributed Transactions with Two-Sided (RDMA) Datagram RPCs
Anuj Kalia, Michael Kaminsky, David G. Andersen
OSDI1
2016 Design Guidelines for High Performance RDMA Systems
Anuj Kalia, Michael Kaminsky, David G. Andersen
USENIX ATC1
2016 Full-Stack Architecting to Achieve a Billion-Requests-Per-Second Throughput on a Single Key-Value Store Server Platform
abstract
Distributed in-memory key-value stores (KVSs), such as memcached, have become a critical data serving layer in modern Internet-oriented data center infrastructure. Their performance and efficiency directly affect the QoS of web services and the efficiency of data centers. Traditionally, these systems have had significant overheads from inefficient network processing, OS kernel involvement, and concurrency control. Two recent research thrusts have focused on improving key-value performance. Hardware-centric research has started to explore specialized platforms including FPGAs for KVSs; results demonstrated an order of magnitude increase in throughput and energy efficiency over stock memcached. Software-centric research revisited the KVS application to address fundamental software bottlenecks and to exploit the full potential of modern commodity hardware; these efforts also showed orders of magnitude improvement over stock memcached. We aim at architecting high-performance and efficient KVS platforms, and start with a rigorous architectural characterization across system stacks over a collection of representative KVS implementations. Our detailed full-system characterization not only identifies the critical hardware/software ingredients for high-performance KVS systems but also leads to guided optimizations atop a recent design to achieve a record-setting throughput of 120 million requests per second (MRPS) (167MRPS with client-side batching) on a single commodity server. Our system delivers the best performance and energy efficiency (RPS/watt) demonstrated to date over existing KVSs including the best-published FPGA-based and GPU-based claims. We craft a set of design principles for future platform architectures, and via detailed simulations demonstrate the capability of achieving a billion RPS with a single server constructed following our principles.
Sheng Li 0007, Hyeontaek Lim, Victor W. Lee, Jung Ho Ahn, Anuj Kalia, Michael Kaminsky, David G. Andersen, Seongil O, Sukhan Lee 0002, Pradeep Dubey
ACM Trans. Comput. Syst.5
2015 Architecting to achieve a billion requests per second throughput on a single key-value store server platform
abstract
Distributed in-memory key-value stores (KVSs), such as memcached, have become a critical data serving layer in modern Internet-oriented datacenter infrastructure. Their performance and efficiency directly affect the QoS of web services and the efficiency of datacenters. Traditionally, these systems have had significant overheads from inefficient network processing, OS kernel involvement, and concurrency control. Two recent research thrusts have focused upon improving key-value performance. Hardware-centric research has started to explore specialized platforms including FPGAs for KVSs; results demonstrated an order of magnitude increase in throughput and energy efficiency over stock memcached. Software-centric research revisited the KVS application to address fundamental software bottlenecks and to exploit the full potential of modern commodity hardware; these efforts too showed orders of magnitude improvement over stock memcached.
Sheng Li 0007, Hyeontaek Lim, Victor W. Lee, Jung Ho Ahn, Anuj Kalia, Michael Kaminsky, David G. Andersen, Seongil O, Sukhan Lee 0002, Pradeep Dubey
ISCA5
2015 Raising the Bar for Using GPUs in Software Packet Processing
Anuj Kalia, Dong Zhou 0006, Michael Kaminsky, David G. Andersen
NSDI1
2014 Using RDMA efficiently for key-value services
abstract
This paper describes the design and implementation of HERD, a key-value system designed to make the best use of an RDMA network. Unlike prior RDMA-based key-value systems, HERD focuses its design on reducing network round trips while using efficient RDMA primitives; the result is substantially lower latency, and throughput that saturates modern, commodity RDMA hardware.
Anuj Kalia, Michael Kaminsky, David G. Andersen
SIGCOMM1