EDBT 2026 Demo / reviewers in the wild / expert
Gargi Alavani Prabhu
dblp:214/0400 · also Gargi Alavani, Gargi Kabirdas Alavani
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-2758-4694ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PowerQuant: Architecture-Agnostic GPU Power Estimation via Quantile RegressionabstractAccurate prediction of NVIDIA GPU power consumption remains challenging due to rapid architectural evolution. Existing machine-learning–based power models are tightly coupled to specific GPU architectures and degrade sharply on unseen platforms, requiring retraining and extensive power measurements, which hinder scalability. This paper presents a quantile-regression–based GPU power prediction framework that enables architecture-agnostic power estimation using static analysis-based compile-time CUDA kernel features. The key insight is that architectural changes primarily induce systematic shifts in power scale, while the relative ordering of kernel power demands remains preserved. By learning power quantiles that capture this ordering and mapping them to new GPUs through one-time calibration, the proposed approach mitigates cross-architecture distribution shift. Extensive evaluation across multiple NVIDIA GPU generations shows that, on unseen architectures, the proposed method improves prediction accuracy by up to 30–50% over existing regression models, while maintaining comparable accuracy in in-distribution settings. The resulting low-overhead, generalizable power estimates make the approach practical for power-aware scheduling, energy budgeting, and sustainability-oriented resource management in large HPC systems. Aditya Challa, Tanish Desai, Gargi Alavani Prabhu, Snehanshu Saha, Santonu Sarkar |
HPDC | 3 |
| 2025 | Predicting Executability and Performance of CNN Kernels on Tenstorrent Hardware Using Machine LearningabstractMaximizing the performance of deep learning models on AI accelerators like Tenstorrent Wormhole requires precise control over hardware resources such as compute cores and onchip memory. Operations such as convolution expose a range of tunable parameters, such as parallelization strategies, buffer sizes, compute datatypes, and memory hierarchies (SRAM vs. DRAM) that involve trade-offs between performance, memory usage, and numerical accuracy. Tenstorrent's TT-NN API reimplements PyTorch's conv2d while exposing all these lowlevel controls, allowing fine-grained control, but presenting a steep learning curve for users familiar with PyTorch's high-level abstractions. We propose a predictive software layer that facilitates the execution of conv2d on Tenstorrent hardware by modeling the relationship between input tensors, configuration parameters, and execution outcomes. We train two machine learning models: one predicts execution success with 99.82% accuracy, and the other estimates run-time performance with an$\mathrm{R}^{2}$score of 0.9984. These models enable automated hardware-aware tuning (TuneNet) of the TT-NN configuration space, reducing trial and error, avoiding out-of-memory errors, and improving performance. TuneNetselected configurations reduce conv2d latency by 13.29% compared to TT-NN defaults on Tenstorrent Wormhole N150s, while incurring minimal computational overhead (9s). This work introduces the first hardware-aware autotuning pipeline for Tenstorrent accelerators, significantly reducing tuning overhead while improving performance. Param Gandhi, Sharvil Potdar, Nayan Gogari, Gargi Alavani Prabhu, Sankar Manoj, Santonu Sarkar |
HiPC | 4 |
| 2025 | Adaptive GPU Power Capping: Balancing Energy Efficiency, Thermal Control and PerformanceabstractAs GPUs become increasingly popular in commodity hardware as well as High Performance Computing(HPC) systems, the need for sustainable computing is more critical. This work addresses the challenge of identifying the optimal operating power for GPUs to minimize energy consumption and operational temperature while incurring only minimal performance overhead. We propose a machine learning-based solution that leverages tree-based models to predict the optimal GPU power cap using key system parameters, including GPU utilization, Memory utilization, Temperature, and Frequency. Our experimental results demonstrate that our model can achieve a maximum energy saving of 12. 87% and a temperature reduction of 11. 38%, with only a 2.69% increase in execution time. These findings highlight the potential of our approach to enhance energy efficiency and thermal management in GPU-based systems, paving the way for more sustainable computing practices. Tanish Desai, Jainam Shah, Gargi Alavani Prabhu, Snehanshu Saha, Santonu Sarkar |
HPDC | 3 |
| 2024 | Estimating Power Consumption of GPU Application Using Machine Learning ToolabstractAs Graphic Processing Units (GPU)s play an increasingly important role in High-Performance Computing (HPC) and data-intensive Machine Learning (ML) tasks, accurate power prediction is essential. Traditional methods, relying on architecture-specific models like DVFS and hardware counters, limit cross-architecture applicability of these models. We propose a static analysis framework that predicts an application's power usage across different NVIDIA GPU architectures without execution. Extensive experiments with state-of-the-art ML approaches show promising results, demonstrating generalizability in predicting power consumption for a newer architecture without the need for complete retraining.11This research is partially supported by the New Faculty Seed Grant of BITS Pilani under Grant No.NFSG/GOA/2023/G0916. Gargi Alavani Prabhu, Tanish Desai, Sharvil Potdar, Nayan Gogari, Snehanshu Saha, Santonu Sarkar |
ICTAI | 1 |
| 2023 | Inspect-GPU: A Software to Evaluate Performance Characteristics of CUDA Kernels Using Microbenchmarks and Regression Models
Gargi Alavani Prabhu, Santonu Sarkar |
ICSOFT | 1 |
| 2022 | Performance modeling of graphics processing unit application using static and dynamic analysisabstractSummary Graphics processing units (GPUs) have become an integral part of high‐performance computing to achieve an exascale performance. Understanding and estimating GPU performance is crucial for developers to design performance‐driven as well as energy‐efficient applications for a given architecture. This work presents a model developed using a static analysis of CUDA code to predict the execution time of NVIDIA GPU kernels without the need for running it. Here a PTX code is statically analyzed to extract instruction features, control flow, and data dependence. We propose a scheduling algorithm that satisfies resource reservation constraints to schedule these instructions in threads across streaming multiprocessors (SMs). We use dynamic analysis to build a set of memory access penalty models and use these models in conjunction with the scheduling information to estimate the execution time of the code. We present the experimental results which support that this approach works across architectures of NVIDIA GPUs. We first tested our model on two Kepler machines, where the mean percentage error (MPE)/mean absolute percentage error (MAPE) was 8.88%/28.3% for Tesla K20 and 5.66%/29.4% for Quadro K4200. We further tested the model on Maxwell and Pascal architectures and recorded the MPEs/MAPEs to be 10.64%/47.8% and %/28.5%, respectively. Gargi Alavani Prabhu, Santonu Sarkar |
Concurr. Comput. Pract. Exp. | 1 |