VLDB 2026 Research / reviewers in the wild / expert
Leonardo Piga
dblp:123/0042
· DBLP profile ↗
9ranked-venue papers
4as first author
2since 2021 · last 2026
0009-0006-2436-5313ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power Sloshing in Compound Servers for Large-Scale AI Inference Workloads
Albert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia, Sultan Mahmud Sajal, Gefei Zuo, Krishna T. Malladi, Devon Akers, Kalyan Subramanian, Shobhit O. Kanaujia, Alexandros Daglis |
ISCA | 3 |
| 2024 | Expanding Datacenter Capacity with DVFS Boosting: A safe and scalable deployment experienceabstractCOVID-19 pandemic created unexpected demand for our physical infrastructure. We increased our computing supply by growing our infrastructure footprint as well as expanded existing capacity by using various techniques among those DVFS boosting. This paper describes our experience in deploying DVFS boosting to expand capacity. Leonardo Piga, Iyswarya Narayanan, Aditya Sundarrajan, Matt Skach, Qingyuan Deng, Biswadip Maity, Manoj Chakkaravarthy, Alison Huang, Abhishek Dhanotia, Parth Malani |
ASPLOS (1) | 1 |
| 2019 | Tangram: Integrated Control of Heterogeneous ComputersabstractResource control in heterogeneous computers built with subsystems from different vendors is challenging. There is a tension between the need to quickly generate local decisions in each subsystem and the desire to coordinate the different subsystems for global optimization. In practice, global coordination among subsystems is considered hard, and current commercial systems use centralized controllers. The result is high response time and high design cost due to lack of modularity. Raghavendra Pradyumna Pothukuchi, Joseph L. Greathouse, Karthik Rao, Christopher Erb, Leonardo Piga, Petros G. Voulgaris, Josep Torrellas |
MICRO | 5 |
| 2017 | Dynamic GPGPU Power Management Using Adaptive Model Predictive ControlabstractModern processors can greatly increase energy efficiency through techniques such as dynamic voltage and frequency scaling. Traditional predictive schemes are limited in their effectiveness by their inability to plan for the performance and energy characteristics of upcoming phases. To date, there has been little research exploring more proactive techniques that account for expected future behavior when making decisions. This paper proposes using Model Predictive Control (MPC) to attempt to maximize the energy efficiency of GPU kernels without compromising performance. We develop performance and power prediction models for a recent CPU-GPU heterogeneous processor. Our system then dynamically adjusts hardware states based on recent execution history, the pattern of upcoming kernels, and the predicted behavior of those kernels. We also dynamically trade off the performance overhead and the effectiveness of MPC in finding the best configuration by adapting the horizon length at runtime. Our MPC technique limits performance loss by proactively spending energy on the kernel iterations that will gain the most performance from that energy. This energy can then be recovered in future iterations that are less performance sensitive. Our scheme also avoids wasting energy on low-throughput phases when it foresees future high-throughput kernels that could better use that energy. Compared to state-of-the-practice schemes, our approach achieves 24.8% energy savings with a performance loss (including MPC overheads) of 1.8%. Compared to state-of-the-art history-based schemes, our approach achieves 6.6% chip-wide energy savings while simultaneously improving performance by 9.6%. Abhinandan Majumdar, Leonardo Piga, Indrani Paul, Joseph L. Greathouse, Wei Huang 0004, David H. Albonesi |
HPCA | 2 |
| 2016 | A Case for Criticality Models in Exascale SystemsabstractPerformance variation is a significant problem for large scale HPC systems and will increase on future exascale systems. In this work, we show that performance variation impacts the performance and energy efficiency of contemporary large-scale computing systems in highly temporally inconsistent ways. We thus present a case for criticality models, a learning based mechanism that allows a system to generate holistic models of performance variation as it occurs during application runtime. Criticality models are designed to provide a mechanism by which applications can detect performance variation at runtime and take action to mitigate its effects. We present a promising preliminary analysis of criticality models on a small scale cluster. Our results demonstrate that models based on logistic regression scan accurately model criticality at this scale. Brian Kocoloski, Leonardo Piga, Wei Huang 0004, Indrani Paul, Jack Lange |
CLUSTER | 2 |
| 2016 | Performance Boosting Opportunities under Communication Imbalance in Power-Constrained HPC ClustersabstractThis paper provides a detailed message-passing interface (MPI) communication characterization across representative HPC applications. It further evaluates performance and power efficiency improvement opportunities. Specifically, it shows that the traditional approach of active polling while waiting for MPI messages is extremely power inefficient, especially under a constrained cluster-level power budget, where processors can only operate at some percentage of their labeled thermal design power (TDP) due to data center infrastructure limits. To mitigate the communication imbalance among different nodes, one can choose to power gate waiting processes and shift remaining power budget to processes that are in the critical execution paths, a technique we call Gate&Shift. With considerations of overheads from power gating and control-loop, Gate&Shift leads to performance improvement without additional power overhead. Gate&Shift is a reactive scheme that does not require prediction mechanisms. With the aid of real MPI traces and hardware measured power data from an HPC cluster, we show that (1) 1 ms control period for power-shifting is sufficient to achieve most potential performance gains, and (2) for a cluster with processors running at 65% of their labeled TDP, Gate&Shift can achieve 7%, 8.5% and 9% performance improvement for AMR Boxlib, Fill Boundary and Big FFT, respectively. Leonardo Piga, Indrani Paul, Wei Huang 0004 |
ICPP | 1 |
| 2014 | Adaptive global power optimization for Web servers
Leonardo Piga, Reinaldo A. Bergamaschi, Maurício Breternitz, Sandro Rigo |
J. Supercomput. | 1 |
| 2013 | Assessing computer performance with stocsabstractSeveral aspects of a computer system cause performance measurements to include random errors. Moreover, these systems are typically composed of a non-trivial combination of individual components that may cause one system to perform better or worse than another depending on the workload. Hence, properly measuring and comparing computer systems performance are non-trivial tasks. Leonardo Piga, Gabriel F. T. Gomes, Rafael Auler, Bruno Rosa 0001, Sandro Rigo, Edson Borin |
ICPE | 1 |
| 2012 | Cloud Workload Analysis with SWATabstractThis note describes the Synthetic Workload Application Toolkit (SWAT) and presents the results from a set of experiments on some key cloud workloads. SWAT is a software platform that automates the creation, deployment, provisioning, execution, and (most importantly) data gathering of synthetic compute workloads on clusters of arbitrary size. SWAT collects and aggregates data from application execution logs, operating system call interfaces, and micro architecture-specific program counters. The data collected by SWAT are used to characterize the effects of network traffic, file I/O, and computation on program performance. The output is analyzed to provide insight into the design and deployment of cloud workloads and systems. Each workload is characterized according to its scalability with the number of server nodes and Hadoop server jobs, sensitivity to network characteristics (bandwidth, latency, statistics on packet size), and computation vs. I/O intensity as these values adjusted via workload-specific parameters. (In the future, we will use SWAT's benchmark synthesizer capability.) We also characterize micro-architectural characteristics that give insight on the micro architecture of processors better suited for this class of workloads. We contrast our results with prior work on Cloud Suite [5], validating some conclusions and providing further insight into others. This illustrates SWAT's data collection capabilities and usefulness to obtain insight on cloud applications and systems. Maurício Breternitz, Keith Lowery, Anton Charnoff, Patryk Kaminski, Leonardo Piga |
SBAC-PAD | 5 |