VLDB 2026 Research / reviewers in the wild / expert
Nipun Kwatra
dblp:53/6666
· DBLP profile ↗
16ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-0354-6204ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QoServe: Breaking the Silos of LLM Inference Serving
Kanishk Goel, Jayashree Mohan, Nipun Kwatra, Ravi Shreyas Anupindi, Ramachandran Ramjee |
ASPLOS (2) | 3 |
| 2025 | SmartKC++: Improving Performance of Smartphone-Based Corneal TopographersabstractKeratoconus, an ocular condition marked by progressive corneal thinning and outward bulging, presents diagnostic challenges due to the high cost and lack of portability in conventional corneal topographers. These limitations restrict accessibility for many, necessitating affordable and mobile alternatives. Innovations like SmartKC [9] offer a low-cost and portable alternative, however, there still remains some gaps in performance when compared to commercial topographers. In this paper, we introduce SmartKC++, a series of innovative methodological improvements to the image processing pipeline of SmartKC, aimed at significantly enhancing its diagnostic precision and reliability. Our comprehensive evaluation on a dataset comprising 303 eye images reveals that SmartKC++ boosts the accuracy of automated keratoconus diagnosis by 7.69% relative to SmartKC. Vaibhav Ganatra, Siddhartha Gairola, Pallavi Joshi, Anand Balasubramaniam, Kaushik Murali, Arivunithi Varadharajan, Bellamkonda Mallikarjuna, Nipun Kwatra |
WACV | 8 |
| 2024 | Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresabstractDeep Learning training jobs process large amounts of training data using many GPU devices, often running for weeks or months. When hardware or software failures happen, these jobs need to restart, losing the memory state for the Deep Neural Network (DNN) model trained so far, unless checkpointing mechanisms are used to save training state periodically. However, for large models, periodic checkpointing incurs significant steady state overhead, and during recovery, a large number of GPUs need to redo work since the last checkpoint. This is especially problematic when failures are frequent for large DNN (such as Large Language Model) training jobs using many GPUs. In this paper, we present a novel approach of just-in-time checkpointing when failures happen, which enables recovery from failures with just a single minibatch iteration of work replayed by all GPUs. This reduces the cost of error recovery from several minutes to a few seconds per GPU, with nearly zero steady state overhead. This also avoids the guesswork of choosing a checkpointing frequency since failure rates usually have high variance. We discuss how just-in-time checkpointing can be enabled in training code, as well as design of key mechanisms for transparent just-in-time checkpointing without user code change. We analyze the wasted GPU work of just-in-time checkpointing and show that it is less than periodic checkpointing for large numbers of GPUs. We present results from our implementation in modern AI cluster infrastructure. Tanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev, Bhargav S. Gulavani, Nipun Kwatra, Ramachandran Ramjee, Muthian Sivathanu |
EuroSys | 6 |
| 2024 | Accuracy is Not All You NeedabstractWhen Large Language Models (LLMs) are compressed using techniques such as quantization, the predominant way to demonstrate the validity of such techniques is by measuring the model's accuracy on various benchmarks. If the accuracies of the baseline model and the compressed model are close, it is assumed that there was negligible degradation in quality. However, even when the accuracy of baseline and compressed model are similar, we observe the phenomenon of flips, wherein answers change from correct to incorrect and vice versa in proportion. We conduct a detailed study of metrics across multiple compression techniques, models and datasets, demonstrating that the behavior of compressed models as visible to end-users is often significantly different from the baseline model, even when accuracy is similar. We further evaluate compressed models qualitatively and quantitatively using MT-Bench and show that compressed models exhibiting high flips are worse than baseline models in this free-form generative task. Thus, we argue that accuracy and perplexity are necessary but not sufficient for evaluating compressed models, since these metrics hide large underlying changes that have not been observed by previous work. Hence, compression techniques should also be evaluated using distance metrics. We propose two such distance metrics, KL-Divergence and flips, and show that they are well correlated. Abhinav Dutta, Sanjeev Krishnan, Nipun Kwatra, Ramachandran Ramjee |
NeurIPS | 3 |
| 2024 | Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, Ramachandran Ramjee |
OSDI | 5 |
| 2023 | Towards Intermediated Workflows for Hybrid TelemedicineabstractThe growing platformization of health has spurred new avenues for healthcare access and reinvigorated telemedicine as a viable pathway to care. Telemedicine adoption during the COVID-19 pandemic has surfaced barriers to patient-centered care that call for attention. Our work extends current Human-Computer Interaction (HCI) research on telemedicine and the challenges to remote care, and investigates the scope for enhancing remote care seeking and provision through telemedicine workflows involving intermediation. Our study, focused on the urban Indian context, involved providing doctors with videos of remote clinical examinations to aid in telemedicine. We present a qualitative evaluation of this modified telemedicine experience, highlighting how workflows involving intermediation could bridge existing gaps in telemedicine, and how their acceptance among doctors could shift interaction dynamics between doctors and patients. We conclude by discussing the implications of such telemedicine workflows on patient-centered care and the future of care work. Karthik S. Bhat, Neha Kumar 0001, Karthik Shamanna, Nipun Kwatra |
CHI | 4 |
| 2023 | Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate ScheduleabstractSeveral papers argue that wide minima generalize better than narrow minima. In this paper, through detailed experiments that not only corroborate the generalization properties of wide minima, we also provide empirical evidence for a new hypothesis that the density of wide minima is likely lower than the density of narrow minima. Further, motivated by this hypothesis, we design a novel explore-exploit learning rate schedule. On a variety of image and natural language datasets, compared to their original hand-tuned learning rate baselines, we show that our explore-exploit schedule can result in either up to 0.84% higher absolute accuracy using the original training budget or up to 57% reduced training time while achieving the original reported accuracy. Nikhil Iyer, V. Thejas, Nipun Kwatra, Ramachandran Ramjee, Muthian Sivathanu |
J. Mach. Learn. Res. | 3 |
| 2022 | Varuna: scalable, low-cost training of massive deep learning modelsabstractSystems for training massive deep learning models (billions of parameters) today assume and require specialized "hyperclusters": hundreds or thousands of GPUs wired with specialized high-bandwidth interconnects such as NV-Link and Infiniband. Besides being expensive, such dependence on hyperclusters and custom high-speed inter-connects limits the size of such clusters, creating (a) scalability limits on job parallelism; (b) resource fragmentation across hyperclusters. Sanjith Athlur, Nitika Saran, Muthian Sivathanu, Ramachandran Ramjee, Nipun Kwatra |
EuroSys | 5 |
| 2020 | Balancing efficiency and fairness in heterogeneous GPU clusters for deep learningabstractWe present Gandivafair, a distributed, fair share scheduler that balances conflicting goals of efficiency and fairness in GPU clusters for deep learning training (DLT). Gandivafair provides performance isolation between users, enabling multiple users to share a single cluster, thus, maximizing cluster efficiency. Gandivafair is the first scheduler that allocates cluster-wide GPU time fairly among active users. Shubham Chaudhary 0004, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Srinidhi Viswanatha |
EuroSys | 4 |
| 2020 | Unsupervised Clustering using Pseudo-semi-supervised Learning
Divam Gupta, Ramachandran Ramjee, Nipun Kwatra, Muthian Sivathanu |
ICLR | 3 |
| 2018 | Gandiva: Introspective Cluster Scheduling for Deep Learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Quanlu Zhang, Fan Yang 0024, Lidong Zhou |
OSDI | 5 |
| 2011 | Asynchronous Evolution for Fully-Implicit and Semi-Implicit Time IntegrationabstractAbstract We propose a series of techniques for hybridizing implicit and semi‐implicit time integration methods in a manner that retains much of the speed of the implicit method without sacrificing all of the higher quality vibrations one obtains with methods that handle elastic forces explicitly. We propose our scheme in the context of asynchronous methods, where different parts of the mesh are evolved at different time steps. Whereas traditional asynchronous methods evolve each element independently, we partition all of our elements into two groups: one group evolved at the frame rate using a fully implicit scheme, and another group which takes a number of substeps per frame using a scheme that is implicit on damping forces and explicit on the elastic forces. This allows for a straightforward coupling between the implicit and semi‐implicit methods at frame boundaries for added stability. As has been stressed by various authors, asynchronous schemes take some of the pressure off of mesh generation, allowing time evolution to remain efficient even in the face of sliver elements. Finally, we propose a force distributing projection method which allows one to redistribute the forces felt on boundaries between implicit and semi‐implicit regions of the mesh in a manner that yields improved visual quality. Craig A. Schroeder, Nipun Kwatra, Zheng Wen 0002, Ronald Fedkiw |
Comput. Graph. Forum | 2 |
| 2010 | Fluid Simulation with Articulated BodiesabstractWe present an algorithm for creating realistic animations of characters that are swimming through fluids. Our approach combines dynamic simulation with data-driven kinematic motions (motion capture data) to produce realistic animation in a fluid. The interaction of the articulated body with the fluid is performed by incorporating joint constraints with rigid animation and by extending a solid/fluid coupling method to handle articulated chains. Our solver takes as input the current state of the simulation and calculates the angular and linear accelerations of the connected bodies needed to match a particular motion sequence for the articulated body. These accelerations are used to estimate the forces and torques that are then applied to each joint. Based on this approach, we demonstrate simulated swimming results for a variety of different strokes, including crawl, backstroke, breaststroke, and butterfly. The ability to have articulated bodies interact with fluids also allows us to generate simulations of simple water creatures that are driven by simple controllers. Nipun Kwatra, Christopher Wojtan, Mark T. Carlson, Irfan A. Essa, Peter J. Mucha, Greg Turk |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2008 | Two-Way Coupled SPH and Particle Level Set Fluid SimulationabstractGrid-based methods have difficulty resolving features on or below the scale of the underlying grid. Although adaptive methods (e.g. RLE, octrees) can alleviate this to some degree, separate techniques are still required for simulating small-scale phenomena such as spray and foam, especially since these more diffuse materials typically behave quite differently than their denser counterparts. In this paper, we propose a two-way coupled simulation framework that uses the particle level set method to efficiently model dense liquid volumes and a smoothed particle hydrodynamics (SPH) method to simulate diffuse regions such as sprays. Our novel SPH method allows us to simulate both dense and diffuse water volumes, fully incorporates the particles that are automatically generated by the particle level set method in under-resolved regions, and allows for two way mixing between dense SPH volumes and grid-based liquid representations. Frank Losasso, Jerry O. Talton, Nipun Kwatra, Ronald Fedkiw |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2007 | Texturing FluidsabstractWe present a novel technique for synthesizing textures over dynamically changing fluid surfaces. We use both image textures as well as bump maps as example inputs. Image textures can enhance the rendering of the fluid by either imparting realistic appearance to it or by stylizing it, whereas bump maps enable the generation of complex micro-structures on the surface of the fluid that may be very difficult to synthesize using simulation. To generate temporally coherent textures over a fluid sequence, we transport texture information, i.e. color and local orientation, between free surfaces of the fluid from one time step to the next. This is accomplished by extending the texture information from the first fluid surface to the 3D fluid domain, advecting this information within the fluid domain along the fluid velocity field for one time step, and interpolating it back onto the second surface -- this operation, in part, uses a novel vector advection technique for transporting orientation vectors. We then refine the transported texture by performing texture synthesis over the second surface using our "surface texture optimization" algorithm, which keeps the synthesized texture visually similar to the input texture and temporally coherent with the transported one. We demonstrate our novel algorithm for texture synthesis on dynamically evolving fluid surfaces in several challenging scenarios. Vivek Kwatra, David Adalsteinsson, Theodore Kim, Nipun Kwatra, Mark T. Carlson, Ming C. Lin |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2005 | Texture optimization for example-based synthesisabstractWe present a novel technique for texture synthesis using optimization. We define a Markov Random Field (MRF)-based similarity metric for measuring the quality of synthesized texture with respect to a given input sample. This allows us to formulate the synthesis problem as minimization of an energy function, which is optimized using an Expectation Maximization (EM)-like algorithm. In contrast to most example-based techniques that do region-growing, ours is a joint optimization approach that progressively refines the entire texture. Additionally, our approach is ideally suited to allow for controllable synthesis of textures. Specifically, we demonstrate controllability by animating image textures using flow fields. We allow for general two-dimensional flow fields that may dynamically change over time. Applications of this technique include dynamic texturing of fluid animations and texture-based flow visualization. Vivek Kwatra, Irfan A. Essa, Aaron F. Bobick, Nipun Kwatra |
ACM Trans. Graph. | 4 |