VLDB 2026 Research / reviewers in the wild / expert
Swapnil Gandhi
dblp:266/6109
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3689-9591ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse Checkpointing for Fast and Reliable MoE Training
Swapnil Gandhi, Christoforos E. Kozyrakis |
NSDI | 1 |
| 2026 | SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
Athinagoras Skiadopoulos, Mark Zhao, Swapnil Gandhi, Thomas Norrie, Shrijeet Mukherjee, Christoforos E. Kozyrakis |
NSDI | 3 |
| 2024 | ReCycle: Resilient Training of Large DNNs using Pipeline AdaptationabstractTraining large Deep Neural Network (DNN) models requires thousands of GPUs over the course of several days or weeks. At this scale, failures are frequent and can have a big impact on training throughput. Utilizing spare GPU servers to mitigate performance loss becomes increasingly costly as model sizes grow. ReCycle is a system designed for efficient DNN training in the presence of failures, without relying on spare servers. It exploits the inherent functional redundancy in distributed training systems - where servers across data-parallel groups store the same model parameters - and pipeline schedule bubbles within each data-parallel group. When servers fails, ReCycle dynamically re-routes microbatches to data-parallel peers, allowing for uninterrupted training despite multiple failures. However, this re-routing can create imbalances across pipeline stages, leading to reduced training throughput. To address this, ReCycle introduces two key optimizations that ensure re-routed microbatches are processed within the original pipeline schedule's bubbles. First, it decouples the backward pass into two phases: one for computing gradients for the input and another for calculating gradients for the parameters. Second, it avoids synchronization across pipeline stages by staggering the optimizer step. Together, these optimizations enable adaptive pipeline schedules that minimize or even eliminate training throughput degradation during failures. We describe a prototype for ReCycle and show that it achieves high training throughput under multiple failures, outperforming recent proposals for fault-tolerant training such as Oobleck and Bamboo by up to 1.46× and 1.64×, respectively. Swapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christoforos E. Kozyrakis |
SOSP | 1 |
| 2024 | Improving DNN Inference Throughput Using Practical, Per-Input Compute AdaptationabstractMachine learning inference platforms continue to face high request rates and strict latency constraints. Existing solutions largely focus on compressing models to substantially lower compute costs (and time) with mild accuracy degradations. This paper explores an alternate (but complementary) technique that trades off accuracy and resource costs on a perinput granularity: early exit models, which selectively allow certain inputs to exit a model from an intermediate layer. Though intuitive, early exits face fundamental deployment challenges, largely owing to the effects that exiting inputs have on batch size (and resource utilization) throughout model execution. We present E3, the first system that makes early exit models practical for realistic inference deployments. Our key insight is to split and replicate blocks of layers in models in a manner that maintains a constant batch size throughout execution, all the while accounting for resource requirements and communication overheads. Evaluations with NLP and vision models show that E3 can deliver up to 1.74× improvement in goodput (for a fixed cost) or 1.78× reduction in cost (for a fixed goodput). Additionally, E3's goodput wins generalize to autoregressive LLMs (2.8--3.8×) and compressed models (1.67×). Anand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan 0003, Swapnil Gandhi, Ravi Netravali |
SOSP | 5 |
| 2022 | Maintaining Social Distancing in Pandemic Using Smartphones With Acoustic WavesabstractThe awareness of “social distancing” is something we have heard because of the coronavirus (COVID-19) pandemic. Second wave of COVID-19 is started, and maintaining social distance and managing our psychological and physical well-being are the new challenges today. We are constantly reminded to maintain a safe distance from other people around us to stop the spread of COVID-19. But what if there was any solution that could make us aware if someone is in your social distance radius and alert us to maintain enough distance? Our article describes a solution using smartphones that delivers near-field peer-to-peer communication by using the present hardware to convert the messages in sound waves. Vaibhav Rupapara, Manideep Narra, Naresh Kumar Gunda, Swapnil Gandhi, Kaushika Reddy Thipparthy |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2021 | P3: Distributed Deep Graph Learning at Scale
Swapnil Gandhi, Anand Padmanabha Iyer |
OSDI | 1 |
| 2020 | An Interval-centric Model for Distributed Computing over Temporal GraphsabstractAlgorithms for temporal property graphs may be time-dependent (TD), navigating the structure and time concurrently, or time-independent (TI), operating separately on different snapshots. Currently, there is no unified and scalable programming abstraction to design TI and TD algorithms over large temporal graphs. We propose an interval-centric computing model (ICM) for distributed and iterative processing of temporal graphs, where a vertex's time-interval is a unit of data-parallel computation. It introduces a unique time-warp operator for temporal partitioning and grouping of messages that hides the complexity of designing temporal algorithms, while avoiding redundancy in user logic calls and messages sent. GRAPHITE is our implementation of ICM over Apache Giraph, and we use it to design 12 TI and TD algorithms from literature. We rigorously evaluate its performance for diverse real-world temporal graphs - as large as 131M vertices and 5.5B edges, and as long as 219 snapshots. Our comparison with 4 baseline platforms on a 10-node commodity cluster shows that ICM shares compute and messaging across intervals to out-perform them by up to 25×, and matches them even in worst-case scenarios. GRAPHITE also exhibits weak-scaling with near-perfect efficiency. Swapnil Gandhi, Yogesh L. Simmhan |
ICDE | 1 |