VLDB 2026 Research / reviewers in the wild / expert
Rachee Singh
dblp:180/5453
· DBLP profile ↗
24ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-8118-3026ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 6 first-author · 15 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconfigurable Torus Fabrics for Multi-tenant MLabstractWe develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML datacenters with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of accelerator failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72x improvement in finetuning throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator with a healthy one in 1.2 seconds. Abhishek Vijaya Kumar, Ding Ding 0005, Arjun Devraj, Darius Bunandar, Rachee Singh |
ASPLOS (2) | 5 |
| 2026 | HEDGE: Traffic Engineering with Probabilistic Link Capacities
Arjun Devraj, Bill Owens, Umesh Krishnaswamy, Rachee Singh |
NSDI | 5 |
| 2026 | Learning to Tune Optical WANs: A Field Deployment of Noise Models in Optical Networks
Bhaskar Kataria, Howard Hua, Andrea D'Amico, Bill Owens, Rachee Singh |
NSDI | 5 |
| 2026 | Opus: Photonic Rail-Optimized Fabric in ML DatacentersabstractRail-optimized network fabrics have become the de facto data-center scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-to-all connectivity in rails imposes substantial power and cost. We propose a rethinking of the rail abstraction by retaining its communication semantics, but realizing it using optical circuit switches. The key challenge is that optical switches support one-to-one connectivity at a time, limiting the fan-out of traffic in ML workloads using hybrid parallelisms. We overcome this through parallelism-driven rail reconfiguration, which exploits the non-overlapping communication phases of different parallelism dimensions. This time-multiplexes a single set of physical ports across circuit configurations tailored to each phase within a training iteration. We design and implement Opus, a control plane that orchestrates this in-job reconfiguration of photonic rails at parallelism phase boundaries, and evaluate it on a physical OCS testbed, the Perlmutter supercomputer, and in simulation at up to 2,048 GPUs. Our results show that photonic rails can achieve over 23× network power reduction and 4× cost savings while incurring only modest training overhead at production-relevant OCS reconfiguration latencies. Ding Ding 0005, Barry Lyu, Bhaskar Kataria, Rachee Singh |
SIGCOMM | 4 |
| 2026 | λλ: A Programming Language for Silicon PhotonicsabstractWe present λλ1, a programming language for silicon photonics. λλ uses a linear type system to encode the physical constraints of optics, rejecting unrealizable programs at compile time. The compiler lowers well-typed programs to a graph-based intermediate representation, then solves a constrained embedding problem to map these graphs onto arbitrary silicon photonic switch targets while minimizing signal loss. We validate λλ on a commercial photonic switch, demonstrating correct operation for circuit switching, time-varying rotor switching and analog in-network computation. Across various hardware targets and programs, the λλ compiler scales to silicon photonic switches with over 100,000 programmable elements and handles switch programs with 128 input-output pairs. Finally, we develop a synthesizer to automatically generate λλ programs from high-level specifications, allowing users to program photonic hardware without reasoning about optical primitives. Vaibhav Mehta, Arjun Devraj, Bill Owens, Justin Hsu, Rachee Singh |
SIGCOMM | 5 |
| 2025 | Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU DomainsabstractInference on large-language models (LLMs) is constrained by GPU memory capacity. A sudden increase in the number of inference requests to a cloud-hosted LLM can deplete GPU memory, leading to contention between multiple prompts for limited resources. Modern LLM serving engines deal with the challenge of limited GPU memory using admission control, which causes them to be unresponsive during request bursts. We propose that preemptive scheduling of prompts in time slices is essential for ensuring responsive LLM inference, especially under conditions of high load and limited GPU memory. However, preempting prompt inference incurs a high paging overhead, which reduces inference throughput. We present Aqua, a GPU memory management framework that significantly reduces the overhead of paging inference state; achieving both responsive and high throughput inference even under bursty request patterns. We evaluate Aqua by hosting several state-of-the-art large generative ML models of different modalities on servers with 8 Nvidia H100 80G GPUs. Aqua improves the responsiveness of LLM inference by 20X compared to the state-of-the-art. It improves LLM inference throughput over a single long prompt by 4X. Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh |
ASPLOS (2) | 3 |
| 2025 | Photonic Rails in ML DatacentersabstractRail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-to-all connectivity in rails imposes massive power, cost, and complexity overheads. We propose a rethinking of the rail abstraction by retaining its communication semantics, but realizing it using optical circuit switches. The key challenge is that optical switches support only one-to-one connectivity at a time, limiting the fan-out of traffic in ML workloads using hybrid parallelisms. We introduce parallelism-driven rail reconfiguration as a solution that leverages the sequential ordering between traffic from different parallelisms. We design a control plane, Opus, to enable time-multiplexed emulation of electrical rail switches using optical switches. More broadly, our work discusses a new research agenda: datacenter fabrics that co-evolve with the model parallelism dimensions within each job, as opposed to the prevailing mindset of reconfiguring networks before a job begins. Ding Ding 0005, Chuhan Ouyang, Rachee Singh |
HotNets | 3 |
| 2025 | Software Managed Networks via CoarseningabstractWe propose moving from Software Defined Networks (SDN) to Software Managed Networks (SMN) where all information for managing the life cycle of a network (from deployment to operations to upgrades), across all layers (from Layer 1 through 7) is stored in a central repository. Crucially, a SMN also has a generalized control plane that, unlike SDN, controls all aspects of the cloud including traffic management (e.g., capacity planning) and reliability (e.g., incident routing) at both short (minutes) and large (years) time scales. Just as SDN allows better routing, a SMN improves visibility and enables cross-layer optimizations for faster response to failures and better network planning and operations. Implemented naively, SMN for planetary sc6ale networks requires orders of magnitude larger and more heterogeneous data (e.g., alerts, logs) than SDN. We address this using coarsening — mapping complex data to a more compact abstract representation that has approximately the same effect, and is more scalable, maintainable, and learnable. We show examples including Coarse Bandwidth Logs for capacity planning and Coarse Dependency Graphs for incident routing. Coarse Dependency Graphs improve an incident routing metric from 45% to 78% while for a distributed approach like Scouts the same metric was 22%. We end by discussing how to realize SMN, and suggest cross-layer optimizations and coarsenings for other operational and planning problems in networks. Pradeep Dogga, Rachee Singh, Suman Nath, Ravi Netravali, Jens Palsberg, George Varghese |
HotNets | 2 |
| 2025 | The Reality of Chasing Shannon's Limit in Optical Wide-Area NetworksabstractWide-area network operators are pushing optical channels toward their theoretical capacity limits, i.e., Shannon's limit, by dynamically adjusting data rates as signal quality fluctuates. This rate-adaptive approach promises to maximize the utilization of expensive fiber infrastructure, yet it introduces a key challenge: each adjustment in the data rate risks temporary wavelength failures that compromise network reliability. Through a six-month empirical study of a production ISP network in the United States, we quantify the fundamental trade-off between capacity efficiency and operational reliability in rate-adaptive optical networks. Our analysis highlights two contributors to this trade-off. First, wider spectral channels, while offering higher peak capacity, can experience up to 4x more frequent capacity fluctuations than narrower channels despite identical signal quality. Second, the precise signal-to-noise ratio (SNR) at which wavelengths operate creates distinct reliability profiles, with certain SNR values located in stability ''sweet spots'' while others could trigger frequent and disruptive rate adaptations. Our findings highlight how operators can strategically approach Shannon's limit by jointly considering the allocation of optical spectral widths, SNR distributions of wavelengths, and transponder capabilities. Arjun Devraj, Bill Owens, Rachee Singh |
IMC | 3 |
| 2025 | FlashMoE: Fast Distributed MoE in a Single KernelabstractThe computational sparsity of Mixture-of-Experts (MoE) models enables sub-linear growth in compute cost as model size increases, thus offering a scalable path to training massive neural networks. However, existing implementations suffer from low GPU utilization, significant latency overhead, and a fundamental inability to leverage task locality, primarily due to CPU-managed scheduling, host-initiated communication, and frequent kernel launches. To overcome these limitations, we develop FlashMoE, a fully GPU-resident MoE operator that fuses expert computation and inter-GPU communication into a single persistent GPU kernel. FlashMoE enables fine-grained pipelining of dispatch, compute, and combine phases, eliminating launch overheads and reducing idle gaps. Unlike existing work, FlashMoE obviates bulk-synchronous collectives for one-sided, device-initiated, inter-GPU (R)DMA transfers, thus unlocking payload efficiency, where we eliminate bloated or redundant network payloads in sparsely activated layers. When evaluated on an 8-H100 GPU node with MoE models having up to 128 experts and 16K token sequences, FlashMoE achieves up to 9× higher GPU utilization, 6× lower latency, 5.7× higher throughput, and 4× better overlap efficiency compared to state-of-the-art baselines—despite using FP32 while baselines use FP16. FlashMoE shows that principled GPU kernel-hardware co-design is key to unlocking the performance ceiling of large-scale distributed ML. We provide code at https://github.com/osayamenja/FlashMoE. Osayamen Jonathan Aimuyo, Byungsoo Oh, Rachee Singh |
NeurIPS | 3 |
| 2025 | Efficient Multi-WAN Transport for 5G with OTTER
Mary Hogan, Gerry Wan, Yiming Qiu 0001, Sharad Agarwal, Ryan Beckett, Rachee Singh, Paramvir Bahl |
NSDI | 6 |
| 2024 | A case for server-scale photonic connectivityabstractThe commoditization of machine learning is fuelling the demand for compute required to both train large models and infer from them. At the same time, scaling the performance of individual microprocessors to satisfy the demand for compute has become increasingly difficult since the end of Moore's law and Dennard scaling. As a result, compute resources in modern servers are distributed across multiple accelerators on the server board. In this work, we make the case for using optics to interconnect accelerators within a server. A key benefit of on-board chip-to-chip optical connectivity is its ability to dynamically allocate bandwidth between accelerators, where necessary, rather than the common practice of statically dividing bandwidth among links within the topology of a multi-accelerator server, as seen in popular direct-connect architectures. This property prevents bandwidth under-utilization in state-of-the-art rack-scale multi-accelerator deployments. Moreover, server-scale optical connectivity can reduce the blast radius of individual accelerator failures in rack-scale ML deployments. Our early experiments with the prototype of a newly commercialized server-scale photonic interconnect show how the capability of the hardware can enable our vision. Abhishek Vijaya Kumar, Arjun Devraj, Darius Bunandar, Rachee Singh |
HotNets | 4 |
| 2024 | CHISEL: An optical slice of the wide-area network
Abhishek Vijaya Kumar, Bill Owens, Nikolaj S. Bjørner, Binbin Guan, Yawei Yin, Paramvir Bahl, Rachee Singh |
NSDI | 7 |
| 2023 | OneWAN is better than two: Unifying a split WAN architecture
Umesh Krishnaswamy, Rachee Singh, Paul Mattes, Paul-Andre C. Bissonnette, Nikolaj S. Bjørner, Zahira Nasrin, Sonal Kothari, Prabhakar Reddy, John Abeln, Srikanth Kandula, Himanshu Raj, Luis Irún-Briz, Jamie Gaudette, Erica Lan |
NSDI | 2 |
| 2023 | Teal: Learning-Accelerated Optimization of WAN Traffic EngineeringabstractThe rapid expansion of global cloud wide-area networks (WANs) has posed a challenge for commercial optimization engines to efficiently solve network traffic engineering (TE) problems at scale. Existing acceleration strategies decompose TE optimization into concurrent subproblems but realize limited parallelism due to an inherent tradeoff between run time and allocation performance. Zhiying Xu, Francis Y. Yan, Rachee Singh, Justin T. Chiu, Alexander M. Rush, Minlan Yu |
SIGCOMM | 3 |
| 2023 | Glowing in the Dark: Uncovering IPv6 Address Discovery and Scanning Strategies in the Wild
Hammas Bin Tanveer, Rachee Singh, Paul Pearce, Rishab Nithyanand |
USENIX Security Symposium | 2 |
| 2022 | Decentralized cloud wide-area network traffic engineering with BLASTSHIELD
Umesh Krishnaswamy, Rachee Singh, Nikolaj S. Bjørner, Himanshu Raj |
NSDI | 2 |
| 2021 | Cost-effective Cloud Edge Traffic Engineering with Cascara
Rachee Singh, Sharad Agarwal, Matt Calder, Paramvir Bahl |
NSDI | 1 |
| 2021 | Cost-effective capacity provisioning in wide area networks with ShooflyabstractIn this work we propose Shoofly, a network design tool that minimizes hardware costs of provisioning long-haul capacity by optically bypassing network hops where conversion of signals from optical to electrical domain is unnecessary and uneconomical. Shoofly leverages optical signal quality and traffic demand telemetry from a large commercial cloud provider to identify optical bypasses in the cloud WAN that reduce the hardware cost of long-haul capacity by 40%. A key challenge is that optical bypasses cause signals to travel longer distances on fiber before re-generation, potentially reducing link capacities and resilience to optical link failures. Despite these challenges, Shoofly provisions bypass-enabled topologies that meet 8X the present-day demands using existing network hardware. Even under aggressive stochastic and deterministic link failure scenarios, these topologies save 32% of the cost of long-haul capacity. Rachee Singh, Nikolaj S. Bjørner, Sharon Shoham, Yawei Yin, John Arnold, Jamie Gaudette |
SIGCOMM | 1 |
| 2018 | Characterizing the Deployment and Performance of Multi-CDNs
Rachee Singh, Arun Dunna, Phillipa Gill |
Internet Measurement Conference | 1 |
| 2018 | RADWAN: rate adaptive wide area networkabstractFiber optic cables connecting data centers are an expensive but important resource for large organizations. Their importance has driven a conservative deployment approach, with redundancy and reliability baked in at multiple layers. In this work, we take a more aggressive approach and argue for adapting the capacity of fiber optic links based on their signal-to-noise ratio (SNR). We investigate this idea by analyzing the SNR of over 8,000 links in an optical backbone for a period of three years. We show that the capacity of 64% of 100 Gbps IP links can be augmented by at least 75 Gbps, leading to an overall capacity gain of over 134 Tbps. Moreover, adapting link capacity to a lower rate can prevent up to 25% of link failures. Our analysis shows that using the same links, we get higher capacity, better availability, and 32% lower cost per gigabit per second. To accomplish this, we propose RADWAN, a traffic engineering system that allows optical links to adapt their rate based on the observed SNR to achieve higher throughput and availability while minimizing the churn during capacity reconfigurations. We evaluate RADWAN using a testbed consisting of 1,540 km fiber with 16 amplifiers and attenuators. We then simulate the throughput gains of RADWAN at scale and compare them to the gains of state-of-the-art traffic engineering systems. Our data-driven simulations show that RADWAN improves the overall network throughput by 40% while also improving the average link availability. Rachee Singh, Manya Ghobadi, Klaus-Tycho Förster, Mark Filer, Phillipa Gill |
SIGCOMM | 1 |
| 2017 | Run, Walk, Crawl: Towards Dynamic Link CapacitiesabstractFiber optic cables are the workhorses of today's Internet services. Operators spend millions of dollars to purchase, lease and maintain their optical backbone, making the efficiency of fiber essential to their business. In this work, we make a case for adapting the capacity of optical links based on their signal-to-noise ratio (SNR). We show two immediate benefits of this by analyzing the SNR of over 2000 links in an optical backbone over a period of 2.5 years. First, the capacity of 80% of IP links can be augmented by 75% or more, leading to an overall capacity gain of 145 Tbps in a large optical backbone in North America. Second, at least 25% of link failures are caused by SNR degradation, not complete loss-of-light, highlighting the opportunity to replace link failures by link flaps wherein the capacity is adjusted according to the new SNR. Given these benefits, we identify the disconnect between current optical and networking infrastructure which hinders the deployment of dynamic capacity links in wide area networks (WANs). To bridge this gap, we propose a graph abstraction that enables existing traffic engineering algorithms to benefit from dynamic link capacities. We evaluate the feasibility of dynamic link capacities using a small testbed and simulate the throughput gains from deploying our approach. Rachee Singh, Manya Ghobadi, Klaus-Tycho Förster, Mark Filer, Phillipa Gill |
HotNets | 1 |
| 2017 | Characterizing the Nature and Dynamics of Tor Exit Blocking
Rachee Singh, Rishab Nithyanand, Sadia Afroz 0001, Paul Pearce, Michael Carl Tschantz, Phillipa Gill, Vern Paxson |
USENIX Security Symposium | 1 |
| 2016 | PathCache: A Path Prediction ToolkitabstractPath prediction on the Internet has been a topic of research in the networking community for close to a decade. Applications of path prediction solutions have ranged from optimizing selection of peers in peer- to-peer networks to improving and debugging CDN predictions. Recently, revelations of traffic correlation and surveillance on the Internet have raised the topic of path prediction in the context of network security. Specifically, predicting network paths can allow us to identify and avoid given organizations on network paths (e.g., to avoid traffic correlation attacks in Tor) or to infer the impact of hijacks and interceptions when direct measurements are not available. Rachee Singh, Phillipa Gill |
SIGCOMM | 1 |