EDBT 2026 Demo / reviewers in the wild / expert
Aniruddha Marathe
dblp:51/10101
· DBLP profile ↗
20ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-0546-4472ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Priority-Aware GPU Co-Scheduling for High Performance Computing
Naman Kulshreshtha, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Rong Ge 0002 |
CCGrid | 3 |
| 2026 | Performance-Aligned LLMs for Generating Fast HPC CodeabstractOptimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. We demonstrate that our fine tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code. Daniel Nichols, Pranav Polasam, Harshitha Menon, Aniruddha Marathe, Todd Gamblin, Abhinav Bhatele |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | ModelX : A Novel Transfer Learning Approach Across Heterogeneous DatasetsabstractLeveraging an existing performance model to predict the runtime of a new application on a new system can save days and weeks of data collection time. However, knowledge transfer between High Performance Computing (HPC) systems can be challenging due to data heterogeneity caused by differences in data collection methods, architectural or application-specific individuality. This results in (1) sets of performance features that have significantly different names, orders, or the number of performance features that do not match between two datasets (heterogeneous domains), or (2) distribution shifts between datasets although their feature names match (homogeneous domains). While existing transfer learning techniques can handle mild distribution shifts, they fail to transfer knowledge when the source and target features do not match. This work introduces a novel transfer learning methodology-Cross Prediction Model (ModelX), which overcomes the large distribution discrepancy between homogeneous domains and enables transfer learning between heterogeneous domains. Extensive evaluations show that ModelX outperforms traditional transfer learning methods for all experiments using 11 HPC and 4 Machine Learning (ML) datasets. To the best of our knowledge, this is the first methodology to enable knowledge transfer between two heterogeneous domains with no matching features. Finally, we demonstrate an application of ModelX to an HPC job scheduling scenario using real-world job traces where it helps to reduce the job turnaround time of a set of jobs by 71%. Arunavo Dey, Neil Antony, Aakash Dhakal, Kowshik Thopalli, Jayaraman J. Thiagarajan, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Jae-Seung Yeom, Tanzima Z. Islam |
HPDC | 7 |
| 2024 | PERFGEN: A Synthesis and Evaluation Framework for Performance Data using Generative AIabstractCollecting data in High-Performance Computing (HPC) is a laborious task, demanding that application scientists execute the application multiple times with different configurations. Due to the essential nature of performance modeling and root cause analysis as initial phases of performance enhancement, the data collection phase prolongs the optimization process. Motivated by this observation, we investigate the feasibility of leveraging the recent advancement in the field of generative Artificial Intelligence (AI) to synthesize performance samples. However, generating synthetic performance data introduces an additional hurdle: the absence of ground truths to assess the quality of the synthetic data. This work takes a step toward bridging this gap where we propose a framework-PERFGEN-for generating performance data and evaluating its quality using a novel metric called Dissimilarity. Our experiments with three performance and five machine learning datasets (including three classification and two regression datasets), confirm that our proposed Dissimilarity correlates with model accuracy better than three of the state-of-the-art metrics-SD quality, Kullback-Leibler Divergence (KL), and TabSyndex, demonstrating that the Dissimilarity metric strongly correlates with the quality of generated scientific data. We evaluate the quality by measuring how well the generated data enables a downstream Machine Learning (ML) task to generalize. Since performance data is a special case of scientific data-typically stored in tabular format and consisting of numerical, categorical, and ordinal features-our methodologies and metrics apply to scientific data from other domains as well. Banooqa H. Banday, Tanzima Z. Islam, Aniruddha Marathe |
COMPSAC | 3 |
| 2024 | On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and ProgrammabilityabstractMatrix multiplication is a core computational part of deep learning and scientific workloads. The emergence of Matrix Cores in high-end AMD GPUs, a building block of Exascale computers, opens new opportunities for optimizing the performance and power efficiency of compute-intensive applications. This work provides a timely, comprehensive characterization of the novel Matrix Cores in AMD GPUs. We develop low-level micro-benchmarks for leveraging Matrix Cores at different levels of parallelism, achieving up to 350, 88, and 69 TFLOPS for mixed, float, and double precision on one GPU. Using results obtained from the micro-benchmarks, we provide a performance model of Matrix Cores that can guide application developers in performance tuning. We also provide the first quantitative study and modeling of the power efficiency of Matrix Cores at different floating-point data types. Finally, we evaluate the high- level programmability of Matrix Cores through the rocBLAS library in a wide range of matrix sizes from 16 to 64K. Our results indicate that application developers can transparently leverage Matrix Cores to deliver more than 92% peak computing throughput by properly selecting data types and interfaces. Gabin Schieffer, Daniel Araújo de Medeiros, Jennifer Faj, Aniruddha Marathe, Ivy Bo Peng |
ISPASS | 4 |
| 2022 | LIBNVCD: An Extendable and User-friendly Multi-GPU Performance Measurement ToolabstractCost and power efficiency considerations have driven High Performance Computing (HPC) system design inno-vations in accelerator-based heterogeneous computing. Complex interactions between applications and heterogeneous hardware make it difficult for users to extract maximum performance out of these systems. While there is a plethora of performance measurement and analysis tools for CPU s, the same is not the case for GPUs. Existing tools either provide too high-level information or are overly complicated to setup, impeding performance profiling. While NVIDIA's CUPTI profiling library enables basic kernel-level measurements on NVIDIA's GPUs, it does not provide root-causes of performance slowdown. This paper presents a low-overhead, flexible, and user-friendly tool, LIBNV CD, built on top of CUPTI to simplify performance measurement and analysis of NVIDIA GPUs. LIBNVCD simplifies obtaining fine-grained measurements, requiring only three function calls in source, while masking changes and complexities of CUPTI. By automatically discovering performance event groups, LIBNV CD reduces data collection overhead significantly as many events (not all) can be measured at once. This user-friendly multi-GPU performance measurement tool incurs a mean overhead of less than 1% as compared to CUPTI, and has been released publicly. Holland Schutte, Chase Phelps, Aniruddha Marathe, Tanzima Z. Islam |
COMPSAC | 3 |
| 2022 | Resource Utilization Aware Job Scheduling to Mitigate Performance VariabilityabstractResource contention on high performance computing (HPC) platforms can lead to significant variation in application performance. When several jobs experience such large variations in run times, it can lead to less efficient use of system resources. It can also lead to users over-estimating their job's expected run time, which degrades the efficiency of the system scheduler. Mitigating performance variation on HPC platforms benefits end users and also enables more efficient use of system resources. In this paper, we present a pipeline for collecting and analyzing system and application performance data for jobs submitted over long periods of time. We use a set of machine learning (ML) models trained on this data to classify performance variation using current system counters. Additionally, we present a new resource-aware job scheduling algorithm that utilizes the ML pipeline and current system state to mitigate job variation. We evaluate our pipeline, ML models, and scheduler using various proxy applications and an actual implementation of the scheduler on an Infiniband-based fat-tree cluster. Daniel Nichols, Aniruddha Marathe, Kathleen Shoga, Todd Gamblin, Abhinav Bhatele |
IPDPS | 2 |
| 2021 | Introducing Application Awareness Into a Unified Power Management StackabstractEffective power management in a data center is critical to ensure that power delivery constraints are met while maximizing the performance of users' workloads. Power limiting is needed in order to respond to greater-than-expected power demand. HPC sites have generally tackled this by adopting one of two approaches: (1) a system-level power management approach that is aware of the facility or site-level power requirements, but is agnostic to the application demands; OR (2) a job-level power management solution that is aware of the application design patterns and requirements, but is agnostic to the site-level power constraints. Simultaneously incorporating solutions from both domains often leads to conflicts in power management mechanisms. This, in turn, affects system stability and leads to irreproducibility of performance. To avoid this irreproducibility, HPC sites have to choose between one of the two approaches, thereby leading to missed opportunities for efficiency gains.This paper demonstrates the need for the HPC community to collaborate towards seamless integration of system-aware and application-aware power management approaches. This is achieved by proposing a new dynamic policy that inherits the benefits of both approaches from tight integration of a resource manager and a performance-aware job runtime environment. An empirical comparison of this integrated management approach against state-of-the-art solutions exposes the benefits of investing in end-to-end solutions to optimize for system-wide performance or efficiency objectives. With our proposed system-application integrated policy, we observed up to 7% reduction in system time dedicated to jobs and up to 11% savings in compute energy, compared to a baseline that is agnostic to system power and application design constraints. Daniel C. Wilson, Siddhartha Jana, Aniruddha Marathe, Stephanie Brink, Christopher Cantalupo, Diana R. Guttman, Brad Geltz, Lowren H. Lawson, Asma Al-Rawi, Ali Mohammad, Fuat Keceli, Federico Ardanaz, Jonathan Eastep, Ayse K. Coskun |
IPDPS | 3 |
| 2020 | Toward an End-to-End Auto-tuning Framework in HPC PowerStackabstractEfficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software stack to manage power and energy of high-performance computing systems and standardizes the interfaces between different components of the stack. This survey paper presents the findings of a working group focused on the end-to-end tuning of the PowerStack. First, we provide a background on the PowerStack layer-specific tuning efforts in terms of their high-level objectives, the constraints and optimization goals, layer-specific telemetry, and control parameters, and we list the existing software solutions that address those challenges. Second, we propose the PowerStack end-to-end auto-tuning framework, identify the opportunities in co-tuning different layers in the PowerStack, and present specific use cases and solutions. Third, we discuss the research opportunities and challenges for collective auto-tuning of two or more management layers (or domains) in the PowerStack. This paper takes the first steps in identifying and aggregating the important R&D challenges in streamlining the optimization efforts across the layers of the PowerStack. Xingfu Wu, Aniruddha Marathe, Siddhartha Jana, Ondrej Vysocky, Jophin John, Andrea Bartolini, Lubomir Riha, Michael Gerndt, Valerie Taylor 0001, Sridutt Bhalachandra |
CLUSTER | 2 |
| 2020 | Dynamic power management for value-oriented schedulers in power-constrained HPC system
Nirmal Kumbhare, Ali Akoglu, Aniruddha Marathe, Salim Hariri, Ghaleb Abdulla |
Parallel Comput. | 3 |
| 2020 | A Value-Oriented Job Scheduling Approach for Power-Constrained and Oversubscribed HPC SystemsabstractIn this article, we investigate limitations in the traditional value-based algorithms for a power-constrained HPC system and evaluate their impact on HPC productivity. We expose the trade-off between allocating system-wide power budget uniformly and greedily under different system-wide power constraints in an oversubscribed system. We experimentally demonstrate that, under the tightest power constraint, the mean productivity of the greedy allocation is 38 percent higher than the uniform allocation whereas, under the intermediate power constraint, the uniform allocation has a mean productivity of 6 percent higher than the greedy allocation. We then propose a new algorithm that adapts its behavior to deliver the combined benefits of the two allocation strategies. We design a methodology with online retraining capability to create application-specific power-execution time models for a class of HPC applications. These models are used in predicting the execution time of an application on the available resources at the time of making scheduling decisions in the power-aware algorithms. We evaluate the proposed algorithm using emulation and simulation environments, and show that our adaptive strategy results in improving HPC resource utilization while delivering a mean productivity that is almost the same as the best performing algorithm across various system-wide power constraints. Nirmal Kumbhare, Aniruddha Marathe, Ali Akoglu, Howard Jay Siegel, Ghaleb Abdulla, Salim Hariri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Adaptive Power Reallocation for Value-Oriented Schedulers in Power-Constrained HPCabstractIn the exascale era, HPC systems are expected to operate under different system-wide power-constraints. For such power-constrained systems, improving per-job flops-per-watt may not be sufficient to improve the total HPC productivity as more number of scientific applications with different compute intensities are migrating to the HPC systems. To measure HPC productivity for such applications, we utilize a monotonically decreasing time-dependent value function, called job-value, with each application. A job-value function represents the value of completing a job for an organization. We begin by exploring the trade-off between two commonly used static power allocation strategies (uniform and greedy) in a power-constrained oversubscribed system. We simulate a large-scale system and demonstrate that, at the tightest power constraint, the greedy allocation can lead to 30% higher productivity compared to the uniform allocation whereas, the uniform allocation can gain up to 6% higher productivity at the relaxed power constraint. We then propose a new dynamic power allocation strategy that utilizes power-performance models derived from offline data. We use these models for reallocating power from running jobs to newly arrived jobs to increase overall system utilization and productivity. In our simulation study, we show that compared to static allocation, the dynamic power allocation policy improves node utilization and job completion rates by 20% and 9%, respectively, at the tightest power constraint. Our dynamic approach consistently earns up to 8% higher productivity compared to the best performing static strategy under different power constraints. Nirmal Kumbhare, Aniruddha Marathe, Ali Akoglu, Salim Hariri, Ghaleb Abdulla |
PDCAT | 2 |
| 2018 | PShifter: feedback-based dynamic power shifting within HPC jobs for performanceabstractThe US Department of Energy (DOE) has set a power target of 20-30MW on the first exascale machines. To achieve one exaFLOPS under this power constraint, it is necessary to manage power intelligently while maximizing performance. Most production-level parallel applications suffer from computational load imbalance across distributed processes due to non-uniform work decomposition. Other factors like manufacturing variation and thermal variation in the machine room may amplify this imbalance. As a result of this imbalance, some processes of a job reach the blocking calls, collectives or barriers earlier and wait for others to reach the same point. This waiting results in a wastage of energy and CPU cycles which degrades application efficiency and performance. Neha Gholkar, Frank Mueller 0001, Barry Rountree, Aniruddha Marathe |
HPDC | 4 |
| 2018 | Bootstrapping Parameter Space Exploration for Fast TuningabstractThe task of tuning parameters for optimizing performance or other metrics of interest such as energy, variability, etc. can be resource and time consuming. Presence of a large parameter space makes a comprehensive exploration infeasible. In this paper, we propose a novel bootstrap scheme, called GEIST, for parameter space exploration to find performance-optimizing configurations quickly. Our scheme represents the parameter space as a graph whose connectivity guides information propagation from known configurations. Guided by the predictions of a semi-supervised learning method over the parameter graph, GEIST is able to adaptively sample and find desirable configurations using limited results from experiments. We show the effectiveness of GEIST for selecting application input options, compiler flags, and runtime/system settings for several parallel codes including LULESH, Kripke, Hypre, and OpenAtom. Jayaraman J. Thiagarajan, Rushil Anirudh, Alfredo Giménez, Rahul Sridhar, Aniruddha Marathe, Tao Wang 0077, Murali Emani, Abhinav Bhatele, Todd Gamblin |
ICS | 6 |
| 2017 | ScrubJay: deriving knowledge from the disarray of HPC performance dataabstractModern HPC centers comprise clusters, storage, networks, power and cooling infrastructure, and more. Analyzing the efficiency of these complex facilities is a daunting task. Increasingly, facilities deploy sensors and monitoring tools, but with millions of instrumented components, analyzing collected data manually is intractable. Data from an HPC center comprises different formats, granularities, and semantics, and handwritten scripts no longer suffice to transform the data into a digestible form. Alfredo Giménez, Todd Gamblin, Abhinav Bhatele, Chad Wood, Kathleen Shoga, Aniruddha Marathe, Peer-Timo Bremer, Bernd Hamann, Martin Schulz 0001 |
SC | 6 |
| 2017 | Performance modeling under resource constraints using deep transfer learningabstractTuning application parameters for optimal performance is a challenging combinatorial problem. Hence, techniques for modeling the functional relationships between various input features in the parameter space and application performance are important. We show that simple statistical inference techniques are inadequate to capture these relationships. Even with more complex ensembles of models, the minimum coverage of the parameter space required via experimental observations is still quite large. We propose a deep learning based approach that can combine information from exhaustive observations collected at a smaller scale with limited observations collected at a larger target scale. The proposed approach is able to accurately predict performance in the regimes of interest to performance analysts while outperforming many traditional techniques. In particular, our approach can identify the best performing configurations even when trained using as few as 1% of observations at the target scale. Aniruddha Marathe, Rushil Anirudh, Abhinav Bhatele, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Jae-Seung Yeom, Barry Rountree, Todd Gamblin |
SC | 1 |
| 2016 | Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2abstractThe use of clouds to execute high-performance computing (HPC) applications has greatly increased recently. Clouds provide several potential advantages over traditional supercomputers and in-house clusters. The most popular cloud is currently Amazon EC2, which provides fixed-cost and variable-cost, auction-based options. The auction market trades lower cost for potential interruptions that necessitate checkpointing; if the market price exceeds the bid price, a node is taken away from the user without warning. We explore techniques to maximize performance per dollar given a time constraint within which an application must complete. Specifically, we design and implement multiple techniques to reduce expected cost by exploiting redundancy in the EC2 auction market. We then design an adaptive algorithm that selects a scheduling algorithm and determines the bid price. We show that our adaptive algorithm executes programs up to seven times cheaper than using the on-demand market and up to 44 percent cheaper than the best non-redundant, auction-market algorithm. We extend our adaptive algorithm to incorporate application scalability characteristics for further cost savings. We show that the adaptive algorithm informed with scalability characteristics of applications achieves up to 56 percent cost savings compared to the expected cost for the base adaptive algorithm run at a fixed, user-defined scale. Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Finding the limits of power-constrained application performanceabstractAs we approach exascale systems, power is turning from an optimization goal to a critical operating constraint. With power bounds imposed by both stakeholders and the limitations of existing infrastructure, we need to develop new techniques that work with limited power to extract maximum performance. In this paper, we explore this area and provide an approach to find the theoretical upper bound of computational performance on a per-application basis in hybrid MPI + OpenMP applications. Peter E. Bailey, Aniruddha Marathe, David K. Lowenthal, Barry Rountree, Martin Schulz 0001 |
SC | 2 |
| 2014 | Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2abstractThe use of clouds to execute high-performance computing (HPC) applications has greatly increased recently. Clouds provide several potential advantages over traditional supercomputers and in-house clusters. The most popular cloud is currently Amazon EC2, which provides a fixed-cost option (called on-demand) and a variable-cost, auction-based option (called the spot market). The spot market trades lower cost for potential interruptions that necessitate checkpointing; if the market price exceeds the bid price, a node is taken away from the user without warning. Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001 |
HPDC | 1 |
| 2013 | A comparative study of high-performance computing on the cloud
Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001, Xin Yuan 0001 |
HPDC | 1 |