Raj Joshi

dblp:137/0874 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-1146-9931ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 3 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 How to Hardware Accelerate Your 5G CU
Xin Zhe Khooi, Satis Kumar Permal, Cha Hwan Song, Nishant Budhdev, Raj Joshi, Mun Choon Chan
INFOCOM5
2026 SyncWise: Error-Aware Time Synchronization for Reconfigurable Data Center Networks
Yiming Lei 0002, Jialong Li 0006, Zhengqing Liu, Raj Joshi, Yiting Xia
NSDI4
2026 OpenOptics: Enabling Open Research and Implementation of Optical Data Center Networks
Yiming Lei 0002, Federico De Marchi 0002, Jialong Li 0006, Raj Joshi, Shu-Ting Wang, Balakrishnan Chandrasekaran 0002, Yiting Xia
NSDI4
2026 Managing Congestion Control Heterogeneity on the Internet with Approximate Performance Isolation
Ayush Mishra, Archit Bhatnagar, Ben Leong, Raj Joshi
NSDI6
2026 Capybara: Dynamic Load Balancing with Microsecond-Scale TCP Migration
abstract
Layer-4 load balancers are a popular solution to high tail latencies but perform poorly under unpredictable skewed workloads because they statically assign connections to servers. We present Capybara, a new load balancer architecture that enables dynamic rebalancing of established connections. Capybara divides load balancing responsibility into a fast L4 load balancer, a host-switch co-designed connection migration protocol, and a transport interface for application-level connection state migration. Capybara leverages two trends - programmable switches and kernel-bypass - to efficiently implement connection migration without disruption, while maintaining transparency to clients. Under realistic workloads, Capybara achieves up to 149× lower tail latency and more than 2× higher throughput for scale-out services compared to state-of-the-art load balancing approaches.
Inho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi, Joshua Fried, Omar S. Navarro Leija, Dan R. K. Ports, Irene Zhang, Jialin Li 0001
SIGCOMM4
2026 Unlocking Diversity of Fast-Switched Optical Data Center Networks With Unified Routing
abstract
Optical data center networks (DCNs) are emerging as a promising solution for cloud infrastructure in the post-Moore’s Law era, particularly with the advent of “fast-switched” optical architectures capable of circuit reconfiguration at microsecond or even nanosecond scales. However, frequent reconfiguration of optical circuits introduces a unique challenge: in-flight packets risk loss during these transitions, hindering the deployment of many mature optical hardware designs due to the lack of suitable routing solutions. In this paper, we presentUnifiedRouting forOptical networks (URO), a general routing framework designed to support fast-switched optical DCNs across various hardware architectures. URO combines theoretical modeling of this novel routing problem with practical implementation on programmable switches, enabling precise, time-based packet transmission. Our prototype on Intel Tofino2 switches achieves a minimum circuit duration of$\mathrm {2~\mu \text {s} }$, ensuring end-to-end, loss-free application performance. Large-scale simulations using production DCN traffic validate URO’s generality across different hardware configurations, demonstrating its effectiveness and efficient system resource utilization.
Jialong Li 0006, Federico De Marchi 0002, Yiming Lei 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia
IEEE Trans. Netw.4
2025 Your network doesn't end at the NIC: A case for unifying the inter-host and intra-host networks in (AI) datacenters
abstract
Modern ML workloads increasingly rely on direct communication between host devices—such as GPUs, NVMe SSDs, and DRAM—spanning intra-host and inter-host networks. However, today's intra-host network lacks hardware-level primitives for routing across heterogeneous interconnects, hindering efficient use of alternative paths and leading to sub-optimal performance under failures or congestion. Furthermore, the inter-host network treats the NIC as the endpoint, with intra-host interconnects like PCIe running oblivious to inter-host network protocols. This prevents leveraging multiple paths for communication between host devices across different servers. To address these limitations, we propose expanding the datacenter network layer to encompass the intra-host network, making intra-host devices first-class network endpoints. Our scheme envisions hardware-level routing and forwarding across multiple intra-host interconnects and makes intra-host devices visible to the inter-host network. This unified approach provides a principled foundation for robust, efficient peer-to-peer communication between storage and compute hardware devices in AI datacenters.
Raj Joshi, Saksham Agarwal, ChonLam Lao, Minlan Yu
HotNets1
2024 JUNCTION: A Scalable Multi-Access Solution Using Programmable Switches
abstract
Multi-access networks are increasingly important for reliable end-to-end connectivity and enhanced throughput performance. A scalable multi-access solution is required to roll out multi-access networks at scale. However, existing CPU-based solutions can no longer scale sustainably, as network traffic has outgrown the CPU performance growth. Consequently, hardware accelerators offer a compelling alternative. This paper introduces JUNCTION, a scalable multi-access solution designed using programmable switches. JUNCTION features a multipath protocol tailored to the hardware constraints and optimized for efficient memory utilization, enabling it to handle a large number of multipath sessions. We validate JUNCTION on a 5G-WiFi multi-access testbed. Our analysis demonstrates that it can scale an order of magnitude better than existing solutions.
Xin Zhe Khooi, Cha Hwan Song, Satis Kumar Permal, Nishant Budhdev, Levente Csikor, Raj Joshi, Mun Choon Chan
SECON6
2024 Keeping an Eye on Congestion Control in the Wild with Nebby
abstract
The Internet congestion control landscape is rapidly evolving. Since the introduction of BBR and the deployment of QUIC, it has become increasingly commonplace for companies to modify and implement their own congestion control algorithms (CCAs). To respond effectively to these developments, it is crucial to understand the state of CCA deployments in the wild. Unfortunately, existing CCA identification tools are not future-proof and do not work well with modern CCAs and encrypted protocols like QUIC. In this paper, we articulate the challenges in designing a future-proof CCA identification tool and propose a measurement methodology that directly addresses these challenges. The resulting measurement tool, called Nebby, can identify all the CCAs currently available in the Linux kernel and BBRv2 with an average accuracy of 96.7%. We found that among the Alexa Top 20k websites, the share of BBR has shrunk since 2019 and that only 8% of them responded to QUIC requests. Among these QUIC servers, CUBIC and BBR seem equally popular. We show that Nebby is extensible by extending it for Copa and an undocumented family of CCAs that is deployed by 6% of the measured websites, including major corporations like Hulu and Apple.
Ayush Mishra, Lakshay Rastogi, Raj Joshi, Ben Leong
SIGCOMM3
2023 Masking Corruption Packet Losses in Datacenter Networks with Link-local Retransmission
abstract
Packet loss due to link corruption is a major problem in large warehouse-scale datacenters. The current state-of-the-art approach of disabling corrupting links is not adequate because, in practice, all the corrupting links cannot be disabled due to capacity constraints. In this paper, we show that, it is feasible to implement link-local retransmission at sub-RTT timescales to completely mask corruption packet losses from the transport endpoints. Our system, LinkGuardian, employs a range of techniques to (i) keep the packet buffer requirement low, (ii) recover from tail packet losses without employing timeouts, and (iii) preserve packet ordering. We implement LinkGuardian on the Intel Tofino switch and show that for a 100G link with a loss rate of 10−3, LinkGuardian can reduce the loss rate by up to 6 orders of magnitude while incurring only 8% reduction in effective link speed. By eliminating tail packet losses, LinkGuardian improves the 99.9th percentile flow completion time (FCT) for TCP and RDMA by 51x and 66x respectively. Finally, we also show that in the context of datacenter networks, simple out-of-order retransmission is often sufficient to significantly mitigate the impact of corruption packet loss for short TCP flows.
Raj Joshi, Cha Hwan Song, Xin Zhe Khooi, Nishant Budhdev, Ayush Mishra, Mun Choon Chan, Ben Leong
SIGCOMM1
2023 Network Load Balancing with In-network Reordering Support for RDMA
abstract
Remote Direct Memory Access (RDMA) is widely used in high-performance computing (HPC) and data center networks. In this paper, we first show that RDMA does not work well with existing load balancing algorithms because of its traffic flow characteristics and assumption of in-order packet delivery. We then propose ConWeave, a load balancing framework designed for RDMA. The key idea of ConWeave is that with the right design, it is possible to perform fine granularity rerouting and mask the effect of out-of-order packet arrivals transparently in the network datapath using a programmable switch. We have implemented ConWeave on a Tofino2 switch. Evaluations show that ConWeave can achieve up to 42.3% and 66.8% improvement for average and 99-percentile FCT, respectively compared to the state-of-the-art load balancing algorithms.
Cha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi, Jialin Li 0001, Mun Choon Chan
SIGCOMM3
2022 LinkGuardian: Mitigating the impact of packet corruption loss with link-local retransmission
abstract
Packet corruption loss is a serious problem in datacenter networks. A large-scale study by Microsoft reported that the number of packets lost due to corruption is comparable to those lost due to congestion. Previous attempts to mitigate the impact of packet corruption loss seek to avoid the faulty links by routing around them, at the cost of reduced link capacities and disruption to the rest of the network.
Raj Joshi, Nishant Budhdev, Ayush Mishra, Mun Choon Chan, Ben Leong
APNet1
2022 Hop-On Hop-Off Routing: A Fast Tour across the Optical Data Center Network for Latency-Sensitive Flows
abstract
Optical data center networks show promise to serve as the next-generation cloud infrastructure especially with their cost and power benefits. The need to set up dedicated optical circuits between endpoints before they can exchange data, however, delays latency-sensitive (“mice”) flows. We find the state-of-the-art solution to reducing flow latency produces sub-optimal paths. To address this issue, we leverage programmable switches to realize Hop-On Hop-Off (HOHO) routing, where mice flows are forwarded along the minimal-latency paths. We prove the optimality and robustness of our algorithm and sketch an implementation on programmable switches. In our packet-level simulations, HOHO routing reduces the flow-completion times for mice flows by up to 35% and the average path length by 15% compared to the state-of-the-art solution.
Jialong Li 0006, Yiming Lei 0002, Federico De Marchi 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia
APNet4
2021 Conjecture: Existence of Nash Equilibria in Modern Internet Congestion Control
abstract
The Internet’s congestion control landscape is currently in the midst of an unprecedented paradigm shift. A recent measurement study found that BBR, a congestion control algorithm introduced by Google in 2016, has seen rapid adoption and is deployed at more than 20% of the Alexa Top 20,000 websites. Encouraging early deployment results from Google, Dropbox and Spotify suggest that BBR could potentially replace traditional loss-based congestion control algorithms like CUBIC. In this paper, we study the interactions between CUBIC and BBR and show that the underlying interactions can be modeled as a normal form game. Our game-theoretic analysis and testbed measurements suggest that while BBR seems to achieve somewhat better performance than CUBIC on the Internet today, this advantage will decrease as the proportion of BBR flows increases. The distribution of congestion control algorithms on the Internet would likely reach a Nash Equilibrium, where no flow has the incentive to switch from CUBIC to BBR, or vice versa. We also found that the distribution of CUBIC and BBR flows in this Nash Equilibrium will be dependent mainly on the size of the bottleneck buffer, and marginally on the RTT distribution of the flows. Our results suggest that the future Internet will likely be more heterogeneous and that buffer sizing will continue to have a significant impact on Internet congestion control.
Ayush Mishra, Jingzhi Zhang, Melodies Sim, Sean Ng, Raj Joshi, Ben Leong
APNet5
2021 FSA: fronthaul slicing architecture for 5G using dataplane programmable switches
abstract
5G networks are gaining pace in development and deployment in recent years. One of 5G's key objective is to support a variety of use cases with different Service Level Objectives (SLOs). Slicing is a key part of 5G that allows operators to provide a tailored set of resources to different use cases in order to meet their SLOs. Existing works focus on slicing in the frontend or the C-RAN. However, slicing is missing in the fronthaul network that connects the frontend to the C-RAN. This leads to over-provisioning in the fronthaul and the C-RAN, and also limits the scalability of the network.
Nishant Budhdev, Raj Joshi, Pravein G. Kannan, Mun Choon Chan, Tulika Mitra
MobiCom2
2021 Debugging Transient Faults in Data Centers using Synchronized Network-wide Packet Histories
Pravein G. Kannan, Nishant Budhdev, Raj Joshi, Mun Choon Chan
NSDI3
2020 Slicing 5G fronthaul networks using programmable switches
abstract
Slicing is a critical technology in 5G, as it allows operators to slice a physical network into multiple virtual networks, each dedicated to a different use case/Mobile Virtual Network Operator (MVNO) [2]. Network slicing enables network operators to deploy a tailored set of resources for specific use cases or MVNO. For example, high performance reliable hardware is required only for ultra-reliable low-latency (uRLLC) use cases such as autonomous vehicle networks. Such tailoring of services reduces costs for network operators. Further, 5G systems can now be deployed more quickly due to virtualization provided by slicing, thereby enabling faster time-to-market. To this end, there exists a large body of work that introduces slicing in different parts of the cellular network (see Fig. 1). PRAN [12] and FlexRAN [13] provide slicing in the Radio Access Network (RAN) while Orion [14] provides slicing for the frontend (wireless spectrum). The fronthaul connects the frontend base station to the RAN and carries digitized radio signals between the two parts of the cellular network. However, to the best of our knowledge, there exists no work on slicing in the fronthaul. This severely limits the benefits of slicing in the RAN and the frontend (see §1.1).
Nishant Budhdev, Raj Joshi, Pravein G. Kannan, Mun Choon Chan, Tulika Mitra
CoNEXT2
2019 SQR: In-network Packet Loss Recovery from Link Failures for Highly Reliable Datacenter Networks
abstract
In datacenter networks, flows need to complete as quickly as possible because the flow completion time (FCT) directly impacts user experience, and thus revenue. Link failures can have a significant impact on short latency-sensitive flows because they increase their FCTs by several fold. Existing link failure management techniques cannot keep the FCTs low under link failures because they cannot completely eliminate packet loss during such failures. We observe that to completely mask the effect of packet loss and the resulting long recovery latency, the network has to be responsible for packet loss recovery instead of relying on end-to-end recovery. To this end, we propose Shared Queue Ring (SQR), an on-switch mechanism that completely eliminates packet loss during link failures by diverting the affected flows seamlessly to alternative paths. We implemented SQR on a Barefoot Tofino switch using the P4 programming language. Our evaluation on a hardware testbed shows that SQR can completely mask link failures and reduce tail FCT by up to 4 orders of magnitude for latency-sensitive workloads.
Ting Qu 0003, Raj Joshi, Mun Choon Chan, Ben Leong, Deke Guo, Zhong Liu 0002
ICNP2
2018 EleTrack: Ultra-Low-Power Retrofitted Monitoring for Elevators
Mobashir Mohammad, Raj Joshi, Mun Choon Chan
EWSN2
2016 HaptiColor: Interpolating Color Information as Haptic Feedback to Assist the Colorblind
abstract
Most existing colorblind aids help their users to distinguish and recognize colors but not compare them. We present HaptiColor, an assistive wristband that encodes discrete color information into spatiotemporal vibrations to support colorblind users to recognize and compare colors. We ran three experiments: the first found the optimal number and placement of motors around the wrist-worn prototype, and the second tested the optimal way to represent discrete points between the vibration motors. Results suggested that using three vibration motors and pulses of varying duration to encode proximity information in spatiotemporal patterns is the optimal solution. Finally, we evaluated the HaptiColor prototype and encodings with six colorblind participants. Our results show that the participants were able to easily understand the encodings and perform color comparison tasks accurately (94.4% to 100%).
Marta Gonzalez Carcedo, Soon Hau Chua, Simon T. Perrault, Pawel W. Wozniak, Raj Joshi, Mohammad Obaid, Morten Fjeld, Shengdong Zhao 0001
CHI5