EDBT 2026 Demo / reviewers in the wild / expert
Tapasya Patki
dblp:20/1263
· DBLP profile ↗
19ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-2543-9688ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Priority-Aware GPU Co-Scheduling for High Performance Computing
Naman Kulshreshtha, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Rong Ge 0002 |
CCGrid | 2 |
| 2026 | Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's RabbitsabstractModern HPC systems are placing increasing demands on job schedulers due to their scale and novel hardware. El Capitan’s Rabbit nodes exemplify this challenge: unlike traditional systems where storage is remote and shared, Rabbit nodes wire local NVMe SSDs directly to compute nodes via PCIe, forcing schedulers to actively track storage topology, capacity, and cross-job persistence, concerns they were never designed to handle. We introduce Flux Fiction, a fully plugin-based HPC system emulator built on top of Flux that replays historical job traces to evaluate scheduling policies in Flux. We validate Flux Fiction against the LLNL Tuolumne cluster using two workloads across four queueing policies, achieving a P99-bounded slowdown error below 1 in 7 of 8 experiments and a maximum utilization error of 1.2%. We then use Flux Fiction to explore Rabbit storage scheduling, demonstrating its ability to explore novel scheduling scenarios. Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer |
HPDC | 8 |
| 2025 | xAMM: "Attention" to Details Improves Cross-Platform Prediction AccuracyabstractAs computing becomes the major enabler in more and more fields, computing platforms also have become more heterogeneous than ever before to support different needs. Inevitably, high performance computing (HPC) centers and cloud vendors offer a diverse array of computing platforms to the user, often to a point where it overwhelms users as well as system managers. Therefore, a cross-platform performance prediction model, which leverages observations from one platform to predict performance on another, can be extremely valuable. However, building such a model for numerous platforms requires an enormous amount of effort to collect training data, which is often prohibitively expensive. To overcome this challenge, we propose$\times \text{AMM}^{1}$11Pronounced as “Exam”, an end-to-end Machine Learning (ML) pipeline that uses the attention mechanism, a transformative concept in generative AI, for two purposes: learning smart embeddings from raw application performance samples and constructing Abstract Machine Models (AMMs)-compact representations of machine properties. By integrating performance sample embeddings with AMMs where available, xAMM improves the accuracy of the state-of-the-art XGBoost model by 49.64 % for CPU$\rightarrow$CPU and 99.07 % for CPU$\rightarrow$GPU prediction compared to building the model using raw data, a common approach in the existing literature. Aakash Dhakal, Tanzima Z. Islam, Arunavo Dey, Daniel Nichols, Abhinav Bhatele, Tapasya Patki, Thomas Scogland, Jae-Seung Yeom |
CCGrid | 6 |
| 2025 | Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPCabstractEl Capitan, currently the world's largest supercomputer at 1.742 Ex-aflop/s, introduces challenges in scheduling due to its scale and innovative rabbit nodes, which traditional schedulers cannot efficiently handle. Flux, a resource and job management system, handles dynamic resource allocation tailored for exascale systems through its graph-based scheduler, Fluxion. This work introduces the Flux Emulator, a tool designed to test scheduling policies in Fluxion without impacting production systems. The emulator plugs into the real components of Flux and Fluxion to mimic job execution, emulate resource usage, and collect information on how the job behaves. Preliminary tests show negligible overhead introduced by the emulator and demonstrate its effectiveness in evaluating scheduli ng policies, like conservative backfilling, in a fraction of the time required with a real system. Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Dewi Yokelson, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer |
HPDC | 9 |
| 2025 | ModelX : A Novel Transfer Learning Approach Across Heterogeneous DatasetsabstractLeveraging an existing performance model to predict the runtime of a new application on a new system can save days and weeks of data collection time. However, knowledge transfer between High Performance Computing (HPC) systems can be challenging due to data heterogeneity caused by differences in data collection methods, architectural or application-specific individuality. This results in (1) sets of performance features that have significantly different names, orders, or the number of performance features that do not match between two datasets (heterogeneous domains), or (2) distribution shifts between datasets although their feature names match (homogeneous domains). While existing transfer learning techniques can handle mild distribution shifts, they fail to transfer knowledge when the source and target features do not match. This work introduces a novel transfer learning methodology-Cross Prediction Model (ModelX), which overcomes the large distribution discrepancy between homogeneous domains and enables transfer learning between heterogeneous domains. Extensive evaluations show that ModelX outperforms traditional transfer learning methods for all experiments using 11 HPC and 4 Machine Learning (ML) datasets. To the best of our knowledge, this is the first methodology to enable knowledge transfer between two heterogeneous domains with no matching features. Finally, we demonstrate an application of ModelX to an HPC job scheduling scenario using real-world job traces where it helps to reduce the job turnaround time of a set of jobs by 71%. Arunavo Dey, Neil Antony, Aakash Dhakal, Kowshik Thopalli, Jayaraman J. Thiagarajan, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Jae-Seung Yeom, Tanzima Z. Islam |
HPDC | 6 |
| 2025 | A Global Perspective on Supercomputer Power Provisioning: Case Studies from United States and EuropeabstractElectrical provisioning in high performance computing is transitioning from simple nameplate Thermal Design Power (TDP) models to more nuanced approaches based on expected electrical load.This paper captures current power Tapasya Patki, Barry Rountree, Torsten Wilde, Andrea Bartolini, Stephanie Brink, Esa Heiskanen, Sachin Idgunji, Matthias Maiterth, James H. Rogers, Ermal Rrapaj, Ralf Schneider, Woong Shin, Kathleen Shoga, Christian Simmendinger, Nicholas J. Wright, Zhengji Zhao |
ICS | 1 |
| 2025 | Enabling Lightweight Performance Analysis of Complex Scientific Workflows with PerfFlowAspect
Aliza Lisan, Tapasya Patki, Stephanie Brink, Konstantinos Parasyris, Brian Gunnarson, Giorgis Georgakoudis, Hank Childs |
SSDBM | 2 |
| 2024 | Relative Performance Prediction Using Few-Shot LearningabstractHigh-performance computing system architectures are evolving rapidly, making exhaustive data collection for each architecture to build predictive performance models increasingly impractical. Concurrently, the arrival of new applications daily necessitates efficient performance prediction methods. Traditional data collection can take days or weeks, making it more efficient for scientists to leverage existing models to predict an application's performance on new architectures or use data from one application to predict another on the same architecture. The growing heterogeneity in applications and resources further complicates the exact matches needed for effective knowledge transfer. This work systematically studies various Machine Learning (ML) models to predict the relative performance of new applications on new platforms using existing data. Our findings demonstrate that few-shot learning using a few samples significantly enhances cross-platform knowledge transfer, multi-source models outperform single-source models, and Large Language Models (LLMs)-generated samples can effectively improve knowledge transfer efficacy. Arunavo Dey, Aakash Dhakal, Tanzima Z. Islam, Jae-Seung Yeom, Tapasya Patki, Daniel Nichols, Alexander Movsesyan, Abhinav Bhatele |
COMPSAC | 5 |
| 2024 | HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning ApproachabstractThe growing necessity for enhanced processing capabilities in edge devices with limited resources has led us to develop effective methods for improving high-performance computing (HPC) applications. In this paper, we introduce LASP (Lightweight Autotuning of Scientific Application Parameters), a novel strategy designed to address the parameter search space challenge in edge devices. Our strategy employs a multi-armed bandit (MAB) technique focused on online exploration and exploitation. Notably, LASP takes a dynamic approach, adapting seamlessly to changing environments. We tested LASP with four HPC applications: Lulesh, Kripke, Clomp, and Hypre. Its lightweight nature makes it particularly well-suited for resource-constrained edge devices. By employing the MAB framework to efficiently navigate the search space, we achieved significant performance improvements while adhering to the stringent computational limits of edge devices. Our experimental results demonstrate the effectiveness of LASP in optimizing parameter search on edge devices. Abrar Hossain, Abdel-Hameed A. Badawy, Mohammad A. Islam 0001, Tapasya Patki, Kishwar Ahmed |
HiPC | 4 |
| 2024 | Predicting Cross-Architecture Performance of Parallel ProgramsabstractA variety of hardware architectures, both CPUs and GPUs, are used today to build supercomputers and parallel clusters. Often times, users can choose which hardware platform they want to run on. Modern scientific workflows have multiple computational tasks, and each task may be better suited for a different architecture in terms of performance. Deciding where to run an application or workflow task is not straightforward because of the complexity of applications, and hardware architectures, which makes performance predictions challenging. Hence, modeling the performance of scientific applications across a variety of architectures is important for achieving the best performance. In this paper, we present a machine learning based methodology to model the relative performance of applications across multiple architectures using hardware performance counters. Our machine learning model can predict the relative performance of an application with a mean absolute error of 0.11, and can be used effectively to make performance-aware and multi-architecture scheduling decisions, reducing makespan by up to 20%. Daniel Nichols, Alexander Movsesyan, Jae-Seung Yeom, Abhik Sarkar, Daniel Milroy, Tapasya Patki, Abhinav Bhatele |
IPDPS | 6 |
| 2023 | Evaluating the Potential of Coscheduling on High-Performance Computing Systems
Jason Hall, Arjun Lathi, David K. Lowenthal, Tapasya Patki |
JSSPP | 4 |
| 2021 | Monitoring Large Scale Supercomputers: A Case Study with the Lassen SupercomputerabstractScalable management of user workloads on large-scale supercomputers remains a challenge due to the tradeoff between capturing adequate detail for analysis from various data sources and minimizing overhead. Co-designed frameworks, such as IBM’s Cluster System Management (CSM), provide a unified approach and novel insights for large-scale cluster management. This paper presents a longitudinal study and detailed analysis of a first-of-its-kind dataset collected by CSM from one of the world’s fastest supercomputers – comprised of over 1.4 million jobs on heterogenous nodes over multiple years. Furthermore, by focusing on a case study for power management, we identify the strengths and limitations of current CSM power measurement techniques in production. We present a deep dive into a large-scale scientific workflow, where finer-grained monitoring reveals power fluctuations at megawatt-levels resulting from the dynamic nature of the application, which are not captured by CSM. We make our unique datasets available to the HPC community, and discuss potential mitigation strategies by analyzing both coarse-grained and fine-grained data. Tapasya Patki, Adam Bertsch, Ian Karlin, Dong H. Ahn, Brian Van Essen, Barry Rountree, Bronis R. de Supinski, Nathan Besaw |
CLUSTER | 1 |
| 2020 | Flux: Overcoming scheduling challenges for exascale workflows
Dong H. Ahn, Ned Bass, Albert Chu, Jim Garlick, Mark Grondona, Stephen Herbein, Helgi I. Ingólfsson, Joe Koning, Tapasya Patki, Thomas Scogland, Becky Springmeyer, Michela Taufer |
Future Gener. Comput. Syst. | 9 |
| 2019 | Performance optimality or reproducibility: that is the questionabstractThe era of extremely heterogeneous supercomputing brings with itself the devil of increased performance variation and reduced reproducibility. There is a lack of understanding in the HPC community on how the simultaneous consideration of network traffic, power limits, concurrency tuning, and interference from other jobs impacts application performance. Tapasya Patki, Jayaraman J. Thiagarajan, Alexis Ayala, Tanzima Z. Islam |
SC | 1 |
| 2018 | Analyzing Resource Trade-offs in Hardware Overprovisioned SupercomputersabstractHardware overprovisioned systems have recently been proposed as a viable alternative for a power-efficient design of next-generation supercomputers. A key challenge for such systems is to determine the degree of overprovisioning, which refers to the number of extra nodes that need to be installed under a given power constraint. In this paper, we first show that the degree of overprovisioning depends on dynamic parameters, such as the job mix as well as the global power constraint, and that static decisions can result in limited system throughput. We then study an exhaustive combination of adaptive resource management strategies that span three job scheduling algorithms, four power capping techniques, and three node boot-up mechanisms to understand the trade-off space involved. We then draw conclusions about how these strategies can adaptively control the degree of overprovisioning and analyze their impact on job throughput and power utilization. Ryuichi Sakamoto, Tapasya Patki, Masaaki Kondo, Koji Inoue, Masatsugu Ueda, Daniel A. Ellsworth, Barry Rountree, Martin Schulz 0001 |
IPDPS | 2 |
| 2017 | Production Hardware Overprovisioning: Real-World Performance Optimization Using an Extensible Power-Aware Resource Management FrameworkabstractLimited power budgets will be one of the biggest challenges for deploying future exascale supercomputers. One of the promising ways to deal with this challenge is hardware over provisioning, that is, installing more hardware resources than can be fully powered under a given power limit coupled with software mechanisms to steer the limited power to where it is needed most. Prior research has demonstrated the viability of this approach, but could only rely on small-scale simulations of the software stack. While such research is useful to understand the boundaries of performance benefits that can be achieved, it does not cover any deployment or operational concerns of using overprovisioning on production systems. This paper is the first to present an extensible power-aware resource management framework for production-sized overprovisioned systems based on the widely established SLURM resource manager. Our framework provides flexible plugin interfaces and APIs for power management that can be easily extended to implement site-specific strategies and for comparison of different power management techniques. We demonstrate our framework on a 965-node HA8000 production system at Kyushu University. Our results indicate that it is indeed possible to safely overprovision hardware in production. We also find that the power consumption of idle nodes, which depends on the degree of overprovisioning, can become a bottleneck. Using real-world data, we then draw conclusions about the impact of the total number of nodes provided in an overprovisioned environment. Ryuichi Sakamoto, Masaaki Kondo, Koji Inoue, Masatsugu Ueda, Tapasya Patki, Daniel A. Ellsworth, Barry Rountree, Martin Schulz 0001 |
IPDPS | 6 |
| 2015 | Practical Resource Management in Power-Constrained, High Performance ComputingabstractPower management is one of the key research challenges on the path to exascale. Supercomputers today are designed to be worst-case power provisioned, leading to two main problems --- limited application performance and under-utilization of procured power. Tapasya Patki, David K. Lowenthal, Anjana Sasidharan, Matthias Maiterth, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski |
HPDC | 1 |
| 2015 | Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputingabstractA key challenge in next-generation supercomputing is to effectively schedule limited power resources. Modern processors suffer from increasingly large power variations due to the chip manufacturing process. These variations lead to power inhomogeneity in current systems and manifest into performance inhomogeneity in power constrained environments, drastically limiting supercomputing performance. We present a first-of-its-kind study on manufacturing variability on four production HPC systems spanning four microarchitectures, analyze its impact on HPC applications, and propose a novel variation-aware power budgeting scheme to maximize effective application performance. Our low-cost and scalable budgeting algorithm strives to achieve performance homogeneity under a power constraint by deriving application-specific, module-level power allocations. Experimental results using a 1,920 socket system show up to 5.4X speedup, with an average speedup of 1.8X across all benchmarks when compared to a variation-unaware power allocation scheme. Yuichi Inadomi, Tapasya Patki, Koji Inoue, Mutsumi Aoyagi, Barry Rountree, Martin Schulz 0001, David K. Lowenthal, Yasutaka Wada, Keiichiro Fukazawa, Masatsugu Ueda, Masaaki Kondo, Ikuo Miyoshi |
SC | 2 |
| 2013 | Exploring hardware overprovisioning in power-constrained, high performance computingabstractMost recent research in power-aware supercomputing has focused on making individual nodes more efficient and measuring the results in terms of flops per watt. While this work is vital in order to reach exascale computing at 20 megawatts, there has been a dearth of work that explores efficiency at the whole system level. Traditional approaches in supercomputer design use worst-case power provisioning: the total power allocated to the system is determined by the maximum power draw possible per node. In a world where power is plentiful and nodes are scarce, this solution is optimal. However, as power becomes the limiting factor in supercomputer design, worst-case provisioning becomes a drag on performance. Tapasya Patki, David K. Lowenthal, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski |
ICS | 1 |