Ahmad Maroof Karimi

dblp:294/1449 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-7270-8847ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Lustre Unveiled: Evolution, Design, Advancements, and Current Trends
abstract
The Lustre filesystem serves as a vital element in high-performance parallel storage, meeting the rising demands of scientific, research, and enterprise environments. Widely deployed across HPC environments, ranging from small-scale applications in AI/ML, to domains like oil and gas, drug discovery, and meteorology, and manufacturing, Lustre addresses the universal challenge of efficiently accessing vast and ever-increasing volumes of data. Lustre is the filesystem of choice on six out of the top 10 fastest supercomputers in the world today, over 65% of the top 100, and also for over 60% of the top 500. Despite its widespread popularity, there is a lack of a complete and up-to-date reference, covering Lustre’s evolution, design, and various advancements made over the years. In this journal, we aim to fill this gap by providing a comprehensive journey of Lustre, including its history with significant contributions to HPC, detailed architecture and design elements, exploration of advancements added through its evolution, and future directions. Additionally, we present a comparison of Lustre with other prominent storage technologies of the era. To illustrate the current state of Lustre, we analyze several filesystem trends, including utilization, performance, and usage patterns on Orion, the Lustre filesystem on the first exascale supercomputer Frontier. We hope that this journal serves as a comprehensive educational reference for the current and future generations interested in HPC filesystem storage aspects.
Anjus George, Andreas Dilger, Michael J. Brim, Rick Mohr, Amir Shehata, Jong Choi 0001, Ahmad Maroof Karimi, Jesse Hanley, James Simmons, Dominic Manno, Verónica G. Vergara Larrea, Sarp Oral, Christopher Zimmer 0001
ACM Trans. Storage7
2024 Power Profile Monitoring and Tracking Evolution of System-Wide HPC Workloads
abstract
The power & energy demands of HPC machines have grown significantly. Modern exascale HPC systems require tens of megawatts of combined power for computing resources and cooling facilities at full capacity. The current energy trend is not sustainable for future HPC systems, and there is a need to work toward the energy efficiency aspect of HPC performance. Energy awareness of the HPC applications at the job level is essential for running an efficient HPC system. This work aims to develop a pipeline to provide a production-level system-wide overview of the HPC workloads' power profile while handling evolving workloads exhibiting new power trends. We developed an open-set classification model for HPC jobs based on the properties of power profiles to continuously provide a system-wide holistic view of recently completed jobs. The pipeline helps continuously monitor the job-level power usage pattern of HPC and enables us to capture the new trends in applications' power behavior. We employed a comprehensive set of techniques to generate job-level data, custom-designed feature extraction methods to extract critical features from jobs' power profiles, clustering techniques powered by generative modeling, and open-set classification for identifying job profiles into known classes or an unknown set. With extensive evaluations, we demonstrate the effectiveness of each component in our pipeline. We provide an analysis of the resulting clusters that characterize the power profile landscape of the Summit supercomputer from more than 60K jobs executed in a year. The open-set classification classifies the known data sets into known classes with high accuracy and identifies unknown data noints with over 85% accuracy.
Ahmad Maroof Karimi, Naw Safrin Sattar, Woong Shin, Feiyi Wang
ICDCS1
2023 Analyzing File Access Patterns on Large-Scale HPC Systems: Opportunities for File Prefetching
abstract
This paper explores the potential opportunities for implementing file prefetching techniques on large-scale high-performance computing (HPC) systems. Specifically, we investigate the file access patterns of various applications across multiple scientific domains using two years' worth of Darshan I/O traces obtained from the Summit supercomputer. We identify recurring trends and patterns which indicate that prefetching can be effectively leveraged to improve data access performance on HPC systems. This study serves as a valuable reference for system architects and developers in the HPC community, providing insights into the opportunities and challenges associated with enabling file prefetching on large-scale HPC systems.
Ahmad Maroof Karimi, Arnab Kumar Paul, Jong Choi 0001, Lipeng Wan 0001, Feiyi Wang
MASCOTS1
2022 Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production Load
abstract
Scientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack.
Jean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Surendra Byna, Philip H. Carns, Sarp Oral, Feiyi Wang, Jesse Hanley
HPDC2
2022 Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems
abstract
Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.
Arnab Kumar Paul, Jong Choi 0001, Ahmad Maroof Karimi, Feiyi Wang
HPDC3
2022 I/O performance analysis of machine learning workloads on leadership scale supercomputer
Ahmad Maroof Karimi, Arnab Kumar Paul, Feiyi Wang
Perform. Evaluation1
2021 Spatiotemporal Graph Neural Network for Performance Prediction of Photovoltaic Power Systems
abstract
In recent years, a large number of photovoltaic (PV) systems have been added to the electrical grid as well as installed as off-grid systems. The trend suggests that the deployment of PV systems will continue to rise in the future. Thus, accurate forecasting of PV performance is critical for the reliability of PV systems. Due to the complex non-linear variability in power output of the PV systems, forecasting PV power is a non-trivial task. This variability affects the stability and planning of a power system network, and accurate forecasting of the performance of the PV system can reduce the uncertainty caused during PV operation. In this work, we leverage spatial and temporal coherence among the power plants for PV power forecasting. Our approach is motivated by the observation that power plants in a region undergo similar environmental exposure. Thus, one power plant’s performance can help improve the forecast of other power plants' power values in the region. We utilize the relationship between PV plants to build a spatiotemporal graph neural network (st-GNN) and train machine learning models to forecast the PV power. The computational experiments on large-scale data from a network of 316 systems show that spatiotemporal forecasting of PV power performs significantly better than a model that only applies temporal convolution to isolated systems or nodes. Furthermore, the longer the future forecast time, the difference between the spatiotemporal forecasting and the isolated system forecast when only temporal convolution is applied increases further.
Ahmad Maroof Karimi, Yinghui Wu 0001, Mehmet Koyutürk, Roger H. French
AAAI1
2021 The Challenge of Disproportionate Importance of Temporal Features in Predicting HPC Power Consumption
abstract
In this work, we demonstrate the challenges in predicting HPC cluster power consumption in the face of significant temporal skew in power consumption behavioral patterns. Predicting large power swings that extend several megawatts has significant operational value for HPC centers, however, prediction is challenging due to the relative rarity of such events and also due to the abrupt or disjoint deviation from the average power consumption levels. To study the impact of this challenge, we have trained a recurrent neural network (RNN) as a reasonably sophisticated model to predict power consumption of the one-year worth of node power consumption data from the Summit supercomputer located in the Oak Ridge Leadership Computing Facility. By studying the prediction results, we have found that although simple usage of RNN models can provide good results on average power consumption levels, it would fail at predicting the power swings that have more operational value. With such results, we discuss potential next steps in addressing such issues aiming towards a robust usage of power prediction techniques in HPC operations.
Ahmad Maroof Karimi, Woong Shin, Hairong Qi 0001, Feiyi Wang
CLUSTER2
2021 Characterizing Machine Learning I/O Workloads on Leadership Scale HPC Systems
abstract
High performance computing (HPC) is no longer solely limited to traditional workloads such as simulation and modeling. With the increase in the popularity of machine learning (ML) and deep learning (DL) technologies, we are observing that an increasing number of HPC users are incorporating ML methods into their workflow and scientific discovery processes, across a wide spectrum of science domains such as biology, earth science, and physics. This gives rise to a diverse set of I/O patterns than the traditional checkpoint/restart-based HPC I/O behavior. The details of the I/O characteristics of such ML I/O workloads have not been studied extensively for large-scale leadership HPC systems. This paper aims to fill that gap by providing an in-depth analysis to gain an understanding of the I/O behavior of ML I/O workloads using darshan - an I/O characterization tool designed for lightweight tracing and profiling. We study the darshan logs of more than 23, 000 HPC ML I/O jobs over a time period of one year running on Summit - the second-fastest supercomputer in the world. This paper provides a systematic I/O characterization of ML I/O jobs running on a leadership scale supercomputer to understand how the I/O behavior differs across science domains and the scale of workloads, and analyze the usage of parallel file system and burst buffer by ML I/O workloads.
Arnab Kumar Paul, Ahmad Maroof Karimi, Feiyi Wang
MASCOTS2
2021 Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputer
abstract
As we approach the exascale computing era, the focused understanding of power consumption and its overall constraint on HPC architectures and applications are becoming increasingly paramount. Summit, located at the Oak Ridge Leadership Computing Facility (OLCF), is one of the fastest and largest pre-exascale platforms in operation today. This paper provides a first-order examination and analysis of power consumption at the component-level, node-level, and system-level, from all 4,626 Summit compute nodes, each with over 100 metrics at 1Hz frequency over the entire year of 2020. We also investigate the power characteristics and energy efficiency of over 840k Summit jobs and 250k GPU failure logs for further operational insights. To the best of our knowledge, this is the first systematic analysis of power data of HPC system at this scale.
Woong Shin, Vladyslav Oles, Ahmad Maroof Karimi, J. Austin Ellis, Feiyi Wang
SC3