EDBT 2026 Demo / reviewers in the wild / expert
Akhil Guliani
dblp:199/8910
· DBLP profile ↗
5ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0002-1437-6506ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Energy-efficient computing · 40% GPUs and heterogeneous computing · 20% Memory systems · 15% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
thermal management |
0.6 | 2 | 2018 | Machine Learning-Based Temperature Prediction for Runtime Thermal Management Across System Components · IEEE Trans. Parallel Distributed Syst. 2018 Adaptive Thermal Management for 3D ICs with Stacked DRAM Caches · DAC 2017 |
GPUs and heterogeneous computing
GPU performance analysis |
0.6 | 1 | 2022 | Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich Systems · SC 2022 |
Performance modeling and evaluation
performance variability |
0.6 | 1 | 2022 | Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich Systems · SC 2022 |
Energy-efficient computing
power management |
0.5 | 2 | 2019 | Per-Application Power Delivery · EuroSys 2019 Adaptive Thermal Management for 3D ICs with Stacked DRAM Caches · DAC 2017 |
Cloud and datacenter computing › datacenter architecture
datacenter server |
0.4 | 1 | 2019 | Per-Application Power Delivery · EuroSys 2019 |
Energy-efficient computing › datacenter power management
power provisioning |
0.4 | 1 | 2019 | Per-Application Power Delivery · EuroSys 2019 |
Memory systems › cache
DRAM cache |
0.3 | 1 | 2017 | Adaptive Thermal Management for 3D ICs with Stacked DRAM Caches · DAC 2017 |
Memory systems › DRAM
refresh management |
0.3 | 1 | 2017 | Adaptive Thermal Management for 3D ICs with Stacked DRAM Caches · DAC 2017 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.1 | 1 | 2017 | Adaptive Thermal Management for 3D ICs with Stacked DRAM Caches · DAC 2017 |
Methods — techniques the papers use, named apart from their topics
power management analysis · 0.6performance characterization · 0.6neural network · 0.3linear regression · 0.3gaussian process · 0.3feature selection · 0.3LASSO · 0.3runtime frequency modulation · 0.3adaptive thermal management · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich SystemsabstractScientists are increasingly exploring and utilizing the massive parallelism of general-purpose accelerators such as GPUs for scientific breakthroughs. As a result, datacenters, hyperscalers, national computing centers, and supercomputers have procured hardware to support this evolving application paradigm. These systems contain hundreds to tens of thousands of accelerators, enabling peta- and exa-scale levels of compute for scientific workloads. Recent work demonstrated that power management (PM) can impact application performance in CPU-based HPC systems, even when machines have the same architecture and SKU (stock keeping unit). This variation occurs due to manufacturing variability and the chip's PM. However, while modern HPC systems widely employ accelerators such as GPUs, it is unclear how much this variability affects applications. Accordingly, we seek to characterize the extent of variation due to GPU PM in modern HPC and supercomputing systems. We study a variety of applications that stress different GPU components on five large-scale computing centers with modern GPUs: Oak Ridge's Summit, Sandia's Vortex, TACC's Frontera and Longhorn, and Livermore's Corona. These clusters use a variety of cooling methods and GPU vendors. In total, we collect over 18,800 hours of data across more than 90% of the GPUs in these clusters. Regardless of the application, cluster, GPU vendor, and cooling method, our results show significant variation: 8% (max 22%) average performance variation even though the GPU architecture and vendor SKU are identical within each cluster, with outliers up to 1.5× slower than the median GPU. These results highlight the difficulty in efficiently using existing GPU clusters for modern HPC and scientific workloads, and the need to embrace variability in future accelerator-based systems. Prasoon Sinha, Akhil Guliani, Rutwik Jain, Brandon Tran, Matthew D. Sinclair, Shivaram Venkataraman |
SC | 2 |
| 2019 | Per-Application Power DeliveryabstractDatacenter servers are often under-provisioned for peak power consumption due to the substantial cost of providing power. When there is insufficient power for the workload, servers can lower voltage and frequency levels to reduce consumption, but at the cost of performance. Current processors provide power limiting mechanisms, but they generally apply uniformly to all CPUs on a chip. For servers running heterogeneous jobs, though, it is necessary to differentiate the power provided to different jobs. This prevents interference when a job may be throttled by another job hitting a power limit. While some recent CPUs support per-CPU power management, there are no clear policies on how to distribute power between applications. Current hardware power limiters, such as Intel's RAPL throttle the fastest core first, which harms high-priority applications. Akhil Guliani, Michael M. Swift |
EuroSys | 1 |
| 2018 | Machine Learning-Based Temperature Prediction for Runtime Thermal Management Across System ComponentsabstractElevated temperatures limit the peak performance of systems because of frequent interventions by thermal throttling. Non-uniform thermal states across system nodes also cause performance variation within seemingly equivalent nodes leading to significant degradation of overall performance. In this paper we present a framework for creating a lightweight thermal prediction system suitable for run-time management decisions. We pursue two avenues to explore optimized lightweight thermal predictors. First, we use feature selection algorithms to improve the performance of previously designed machine learning methods. Second, we develop alternative methods using neural network and linear regression-based methods to perform a comprehensive comparative study of prediction methods. We show that our optimized models achieve improved performance with better prediction accuracy and lower overhead as compared with the Gaussian process model proposed previously. Specifically we present a reduced version of the Gaussian process model, a neural network-based model, and a linear regression-based model. Using the optimization methods, we are able to reduce the average prediction errors in the Gaussian process from 4.2°C to 2.9°C. We also show that the newly developed models using neural network and Lasso linear regression have average prediction errors of 2.9°C and 3.8°C respectively. The prediction overheads are 0.22, 0.097, and 0.026 ms per prediction for reduced Gaussian process, neural network, and Lasso linear regression models, respectively, compared with 0.57 ms per prediction for the previous Gaussian process model. We have implemented our proposed thermal prediction models on a two-node system configuration to help identify the optimal task placement. The task placement identified by the models reduces the average system temperature by up to 11.9°C without any performance degradation. Furthermore, these models respectively achieve 75, 82.5, and 74.17 percent success rates in correctly pointing to those task placements with better thermal response, compared with 72.5 percent success for the original model in achieving the same objective. Finally, we extended our analysis to a 16-node system and we were able to train models and execute them in real time to guide task migration and achieve on average 17 percent reduction in the overall system cooling power. Akhil Guliani, Seda Ogrenci Memik, Gokhan Memik, Kazutomo Yoshii, Rajesh Sankaran, Pete Beckman |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Adaptive Thermal Management for 3D ICs with Stacked DRAM CachesabstractWe describe an adaptive thermal management system for 3D-ICs with stacked DRAM cache memories. We present a detailed analysis of the impact of 3D-IC hotspot aggregation on the refresh behavior of the stacked DRAM-based L3 cache. We also present the consequence of the refresh variation on the overall system performance and cache energy consumption. Our analysis demonstrates that memory intensive applications are influenced more strongly by the DRAM refresh variation. We show that there is an optimal operating point where, with a reduced clock frequency, processor cores would actually recover any performance loss induced by DRAM refresh and at the same time the cache energy consumption could be optimized. We propose a low overhead run-time method that can identify the best CPU frequency modulation factor to cool the system to minimize accelerated refresh rates in the DRAM caches. Our system can provide a customizable trade-off between performance of the processor and energy savings of the memory. Akhil Guliani, Seda Ogrenci Memik |
DAC | 3 |
| 2017 | Dark Shadows: User-Level Guest/Host Linux Process ShadowingabstractThe concept of a shadow process simplifies the design and implementation of virtualization services such as system call forwarding and device file-level device virtualization. A shadow process on the host mirrors a process in the guest at the level of the virtual and physical address space, terminating in the host physical addresses. Previous shadow process mechanisms have required changes to the guest and host kernels. We describe a shadow process technique that is implemented at user-level in both the guest and the host. In our technique, we refer to the host shadow process as a dark shadow as it arranges its own elements to avoid conflicting with the guest process's elements. We demonstrate the utility of dark shadows by using our implementation to create system call forwarding and device file-level device virtualization prototypes that are compact and simple. Peter A. Dinda, Akhil Guliani |
IC2E | 2 |