EDBT 2026 Demo / reviewers in the wild / expert
Tanzima Z. Islam
dblp:10/7555 · also Tanzima Zerin Islam
· DBLP profile ↗
27ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0003-2877-5871ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | xAMM: "Attention" to Details Improves Cross-Platform Prediction AccuracyabstractAs computing becomes the major enabler in more and more fields, computing platforms also have become more heterogeneous than ever before to support different needs. Inevitably, high performance computing (HPC) centers and cloud vendors offer a diverse array of computing platforms to the user, often to a point where it overwhelms users as well as system managers. Therefore, a cross-platform performance prediction model, which leverages observations from one platform to predict performance on another, can be extremely valuable. However, building such a model for numerous platforms requires an enormous amount of effort to collect training data, which is often prohibitively expensive. To overcome this challenge, we propose$\times \text{AMM}^{1}$11Pronounced as “Exam”, an end-to-end Machine Learning (ML) pipeline that uses the attention mechanism, a transformative concept in generative AI, for two purposes: learning smart embeddings from raw application performance samples and constructing Abstract Machine Models (AMMs)-compact representations of machine properties. By integrating performance sample embeddings with AMMs where available, xAMM improves the accuracy of the state-of-the-art XGBoost model by 49.64 % for CPU$\rightarrow$CPU and 99.07 % for CPU$\rightarrow$GPU prediction compared to building the model using raw data, a common approach in the existing literature. Aakash Dhakal, Tanzima Z. Islam, Arunavo Dey, Daniel Nichols, Abhinav Bhatele, Tapasya Patki, Thomas Scogland, Jae-Seung Yeom |
CCGrid | 2 |
| 2025 | ModelX : A Novel Transfer Learning Approach Across Heterogeneous DatasetsabstractLeveraging an existing performance model to predict the runtime of a new application on a new system can save days and weeks of data collection time. However, knowledge transfer between High Performance Computing (HPC) systems can be challenging due to data heterogeneity caused by differences in data collection methods, architectural or application-specific individuality. This results in (1) sets of performance features that have significantly different names, orders, or the number of performance features that do not match between two datasets (heterogeneous domains), or (2) distribution shifts between datasets although their feature names match (homogeneous domains). While existing transfer learning techniques can handle mild distribution shifts, they fail to transfer knowledge when the source and target features do not match. This work introduces a novel transfer learning methodology-Cross Prediction Model (ModelX), which overcomes the large distribution discrepancy between homogeneous domains and enables transfer learning between heterogeneous domains. Extensive evaluations show that ModelX outperforms traditional transfer learning methods for all experiments using 11 HPC and 4 Machine Learning (ML) datasets. To the best of our knowledge, this is the first methodology to enable knowledge transfer between two heterogeneous domains with no matching features. Finally, we demonstrate an application of ModelX to an HPC job scheduling scenario using real-world job traces where it helps to reduce the job turnaround time of a set of jobs by 71%. Arunavo Dey, Neil Antony, Aakash Dhakal, Kowshik Thopalli, Jayaraman J. Thiagarajan, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Jae-Seung Yeom, Tanzima Z. Islam |
HPDC | 10 |
| 2025 | On The Role of Prompt Construction In Enhancing Efficacy and Efficiency of LLM-Based Tabular Data GenerationabstractLLM-based data generation for real-world tabular data can be challenged by the lack of sufficient semantic context in feature names used to describe columns. We hypothesize that enriching prompts with even minimal contextual information, such as a brief explanation of what each feature represents can improve both the quality and efficiency of data generation. To test this, we investigate three prompt construction methods: Expert-guided, LLM-guided, and Novel-Mapping, with the latter two being automated approaches. Using the GReaT framework, our experiments show that context-enriched prompts significantly enhance the quality of the generated data while improving training efficiency. Notably, the LLM-guided method performed on par with expert-guided approaches, demonstrating its effectiveness as a scalable alternative. Banooqa H. Banday, Kowshik Thopalli, Tanzima Z. Islam, Jayaraman J. Thiagarajan |
ICASSP | 3 |
| 2025 | Structure-Aware Representation Learning for Effective Performance PredictionabstractABSTRACT Application performance is a function of several unknowns stemming from the interactions between the application, runtime, OS, and underlying hardware, making it challenging to model performance using deep learning techniques, especially without a large labeled dataset. Collecting such labeled longitudinal datasets can take weeks. Intuitively, developers could save analysis time during code development by taking a comparative approach between multiple applications. However, the unknown dynamic interactions between applications and execution environments make it difficult for deep learning‐based models to predict the performance of new applications. In this paper, we address these problems by presenting a labeled dataset for the community and taking a comparative analysis approach to explore the source code differences between different correct implementations of the same problem. This paper assesses the feasibility of using purely static information, for example, Abstract Syntax Tree (AST), of applications to predict performance change based on code structure. We evaluate several deep learning‐based representation learning techniques for source code and propose an architecture for the tree‐based Long Short‐Term Memory (LSTM) models to discover latent representations for a source code's hierarchical structure. We demonstrate that our proposed architecture enables feed‐forward predictive models to predict change in performance using source code with up to 84% accuracy. Tarek Ramadan, Nathan Pinnow, Chase Phelps, Jayaraman J. Thiagarajan, Tanzima Z. Islam |
Concurr. Comput. Pract. Exp. | 5 |
| 2024 | PERFGEN: A Synthesis and Evaluation Framework for Performance Data using Generative AIabstractCollecting data in High-Performance Computing (HPC) is a laborious task, demanding that application scientists execute the application multiple times with different configurations. Due to the essential nature of performance modeling and root cause analysis as initial phases of performance enhancement, the data collection phase prolongs the optimization process. Motivated by this observation, we investigate the feasibility of leveraging the recent advancement in the field of generative Artificial Intelligence (AI) to synthesize performance samples. However, generating synthetic performance data introduces an additional hurdle: the absence of ground truths to assess the quality of the synthetic data. This work takes a step toward bridging this gap where we propose a framework-PERFGEN-for generating performance data and evaluating its quality using a novel metric called Dissimilarity. Our experiments with three performance and five machine learning datasets (including three classification and two regression datasets), confirm that our proposed Dissimilarity correlates with model accuracy better than three of the state-of-the-art metrics-SD quality, Kullback-Leibler Divergence (KL), and TabSyndex, demonstrating that the Dissimilarity metric strongly correlates with the quality of generated scientific data. We evaluate the quality by measuring how well the generated data enables a downstream Machine Learning (ML) task to generalize. Since performance data is a special case of scientific data-typically stored in tabular format and consisting of numerical, categorical, and ordinal features-our methodologies and metrics apply to scientific data from other domains as well. Banooqa H. Banday, Tanzima Z. Islam, Aniruddha Marathe |
COMPSAC | 2 |
| 2024 | Relative Performance Prediction Using Few-Shot LearningabstractHigh-performance computing system architectures are evolving rapidly, making exhaustive data collection for each architecture to build predictive performance models increasingly impractical. Concurrently, the arrival of new applications daily necessitates efficient performance prediction methods. Traditional data collection can take days or weeks, making it more efficient for scientists to leverage existing models to predict an application's performance on new architectures or use data from one application to predict another on the same architecture. The growing heterogeneity in applications and resources further complicates the exact matches needed for effective knowledge transfer. This work systematically studies various Machine Learning (ML) models to predict the relative performance of new applications on new platforms using existing data. Our findings demonstrate that few-shot learning using a few samples significantly enhances cross-platform knowledge transfer, multi-source models outperform single-source models, and Large Language Models (LLMs)-generated samples can effectively improve knowledge transfer efficacy. Arunavo Dey, Aakash Dhakal, Tanzima Z. Islam, Jae-Seung Yeom, Tapasya Patki, Daniel Nichols, Alexander Movsesyan, Abhinav Bhatele |
COMPSAC | 3 |
| 2024 | Reimagine Application Performance as a Graph: Novel Graph-Based Method for Performance Anomaly Classification in High-Performance ComputingabstractPerformance anomaly in High Performance Computing (HPC) can be defined as run-to-run variation of an application in repeated runs with the same set of configuration parameters. Such variations can occur for myriad reasons, including contention for shared resources such as the network and dynamic data distribution across application processes. Traditionally HPC researchers focus on real-time anomaly detection using different Machine Learning (ML) methods. These popular methods, such as auto-encoder, limit finding anomalous event patterns during training time. On top of that, in HPC, performance data are stored in tabular format. Though gradient-based methods have already proved their significant improvement over classification tasks, they explicitly use feature-feature relationships, ignoring the potential sample-sample relationship. To fill this gap, we build a performance anomaly classification technique leveraging the potential graph-based representation learning. We hypothesize that a meaningful and robust representation considering the sample-sample relationship for the given tabular datasets will improve the downstream anomaly classification technique. We conduct our experiment on 5 HPC datasets and 6 ML datasets. Our empirical study proves that graph-based anomaly classification outperforms the gradient-based approaches in 6 out of 11 experiments. We also explain how anomaly decisions are made inside the performance graph. Chase Phelps, Ankur Lahiry, Tanzima Z. Islam, Line C. Pouchard |
COMPSAC | 3 |
| 2024 | Characterize and Compare the Performance of Deep Learning Optimizers in Recurrent Neural Network ArchitecturesabstractThe emergence of newer hardware and limited accessibility to that hardware are driving the need to characterize application performance on commonly used architectures to gain valuable insights. While abundant performance events create opportunities for detailed performance characterization, they also make the process overwhelming. This paper takes a systematic approach to address the problem of performance characterization by taking advantage of such opportunities. Specifically, this paper 1) characterizes the hardware usage behaviors of commonly used optimizers in the context of Recurrent Neural Network (RNN) architectures, 2) connects observations to identify problems and opportunities for optimization in those optimizers, and 3) creates a large performance dataset. Comparing the hardware resource usage behaviors of Gradient Descent (GS), AdaGrad (AdaGrad), and ADAM (ADAM) optimizers for three commonly used recurrent neural network-based deep learning architectures running on AMD MI100 GPUs, this paper identifies non-uniform memory access as one of the most frequent issues and suggests strategies for optimizing such problems. Mohammad Zaeed, Tanzima Z. Islam, Vladimir Indic |
COMPSAC | 2 |
| 2023 | Signal Processing Based Method for Real-Time Anomaly Detection in High-Performance ComputingabstractPerformance anomalies can manifest as irregular execution times or abnormal execution events for many reasons, including network congestion and resource contention. Detecting such anomalies in real-time by analyzing the details of performance traces at scale is impractical due to the sheer volume of data High-Performance Computing (HPC) applications produce. In this paper, we propose formulating HPC performance anomaly detection as a signal-processing problem where anomalies can be treated as noise. We evaluate our proposed method in comparison with two other commonly used anomaly detection techniques of varying complexity based on their detection accuracy and scalability. Since real-time in-situ anomaly detection at a large scale requires lightweight methods that can handle a large volume of streaming data, we find that our proposed method provides the best trade-off. We then implement the proposed method in Chimbuko, the first online, distributed, and scalable workflow-level performance trace analysis framework. We compare our proposed signal-based anomaly detection algorithm with two other methods using a function of their accuracy, F1 score, and detection overhead. Our experiments demonstrate that our proposed approach achieves a 99% improvement for the benchmark datasets and a 93% improvement with Chimbuko traces. Arunavo Dey, Tanzima Z. Islam, Chase Phelps, Christopher Kelly |
COMPSAC | 2 |
| 2023 | Automatic Parallelization of Cellular Automata for Heterogeneous Platforms
Chase Phelps, Tanzima Z. Islam |
COMPSAC | 2 |
| 2023 | Building the I (Interoperability) of FAIR for Performance Reproducibility of Large-Scale Composable Workflows in RECUPabstractScientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment. Bogdan Nicolae, Tanzima Z. Islam, Robert B. Ross, Huub J. J. Van Dam, Kevin Assogba, Polina Shpilker, Mikhail Titov, Matteo Turilli, Tianle Wang 0001, Ozgur O. Kilic, Shantenu Jha, Line C. Pouchard |
e-Science | 2 |
| 2023 | Novel Representation Learning Technique Using Graphs for Performance AnalyticsabstractThe performance analytics domain in High Performance Computing (HPC) uses tabular data to solve regression problems, such as predicting the execution time. Existing Machine Learning (ML) techniques leverage the correlations among features given tabular datasets, not leveraging the relationships between samples directly. Moreover, since high-quality embed-dings from raw features improve the fidelity of the downstream predictive models, existing methods rely on extensive feature engineering and pre-processing steps, costing time and manual effort. To fill these two gaps, we propose a novel idea of transforming tabular performance data into graphs to leverage the advancement of Graph Neural Network-based (GNN) techniques in capturing complex relationships between features and samples. In contrast to other ML application domains, such as social networks, the graph is not given; instead, we need to build it. To address this gap, we propose graph-building methods where nodes represent samples, and the edges are automatically inferred iteratively based on the similarity between the features in the samples. We evaluate the effectiveness of the generated embeddings from GNNs based on how well they make even a simple feed-forward neural network perform for regression tasks compared to other state-of-the-art representation learning techniques. Our evaluation demonstrates that even with up to 25 % random missing values for each dataset, our method outperforms commonly used graph and Deep Neural Network (DNN)-based approaches and achieves up to 61.67% & 78.56 % improvement in MSE loss over the DNN baseline respectively for HPC dataset and Machine Learning Datasets. Tarek Ramadan, Ankur Lahiry, Tanzima Z. Islam |
ICMLA | 3 |
| 2022 | LIBNVCD: An Extendable and User-friendly Multi-GPU Performance Measurement ToolabstractCost and power efficiency considerations have driven High Performance Computing (HPC) system design inno-vations in accelerator-based heterogeneous computing. Complex interactions between applications and heterogeneous hardware make it difficult for users to extract maximum performance out of these systems. While there is a plethora of performance measurement and analysis tools for CPU s, the same is not the case for GPUs. Existing tools either provide too high-level information or are overly complicated to setup, impeding performance profiling. While NVIDIA's CUPTI profiling library enables basic kernel-level measurements on NVIDIA's GPUs, it does not provide root-causes of performance slowdown. This paper presents a low-overhead, flexible, and user-friendly tool, LIBNV CD, built on top of CUPTI to simplify performance measurement and analysis of NVIDIA GPUs. LIBNVCD simplifies obtaining fine-grained measurements, requiring only three function calls in source, while masking changes and complexities of CUPTI. By automatically discovering performance event groups, LIBNV CD reduces data collection overhead significantly as many events (not all) can be measured at once. This user-friendly multi-GPU performance measurement tool incurs a mean overhead of less than 1% as compared to CUPTI, and has been released publicly. Holland Schutte, Chase Phelps, Aniruddha Marathe, Tanzima Z. Islam |
COMPSAC | 4 |
| 2021 | College Life is Hard! - Shedding Light on Stress Prediction for Autistic College Students using Data-Driven AnalysisabstractAutistic college students face significant challenges in college settings and have a higher dropout rate than neurotypical college students. High physiological distress, depression, and anxiety are identified as critical challenges that contribute to this less than optimal college experience. In this paper, we leverage affordable mobile and wearable devices to collect large amounts of physiological and contextual data (biomarkers) and leverage a data-driven analysis approach for building stress prediction models. Such models can be used to provide real-time intervention for better stress management. We conducted a mixed-method study where we collected physiological and contextual data from 20 college students (10 neurotypical and 10 autistic). Our proposed data-driven analysis pipeline leverages an unsupervised representation learning technique with a semi-supervised label approximation method to predict the onset of stress based on biomarkers for autistic students, neurotypical students, and both populations with accuracies 69%, 72%, and 70%, respectively. Tanzima Z. Islam, Philip Wu Liang, Forest Sweeney, Cody Pranger, Jayaraman J. Thiagarajan, Moushumi Sharmin, Shameem Ahmed |
COMPSAC | 1 |
| 2021 | FILCIO: Application Agnostic I/O Aggregation to Scale Scientific WorkflowsabstractScientific workflows often consist of black-box stand-alone programs that each write to one or more files, resulting in significant I/O operations. When strung together, each stage of the pipeline is bottlenecked by I/O, without much ability to modify that behavior. Since computation capabilities of large-scale distributed systems grow much faster than their I/O bandwidth, processing such big data using these workflows does not scale as the amount of data grows. While most related work focuses on data aggregation by modifying source code, we propose a Function Interception Library for Collective I/O (FILCIO), to intercept I/O function calls in C’s stdio library on a POSIX file system. This aggregates outgoing data in memory before writing to disk. Our aim is to scale I/O intensive scientific workflows without needing to access and modify their source code. In this paper, we motivate the need for I/O aggregation, and discuss what applications are poised to benefit most from FILCIO. We use an I/O intensive program where nearly 100% of its overall runtime is spent writing to disk. We show that the use of FILCIO results in a speedup of 1.75x over the native approach, and extrapolate that the best improvement is when a program generates a large volume of small writes. FILCIO is available at https://github.com/jensenq/io_aggregation Quentin Jensen, Filip Jagodzinski, Tanzima Z. Islam |
COMPSAC | 3 |
| 2021 | Comparative Code Structure Analysis using Deep Learning for Performance PredictionabstractPerformance analysis has always been an afterthought during the application development process, focusing on application correctness first. The learning curve of the existing static and dynamic analysis tools are steep, which requires understanding low-level details to interpret the findings for actionable optimizations. Additionally, application performance is a function of a number of unknowns stemming from the application-, runtime-, and interactions between the OS and underlying hardware, making it difficult to model using any deep learning technique, especially without a large labeled dataset. In this paper, we address both of these problems by presenting a large corpus of a labeled dataset for the community and take a comparative analysis approach to mitigate all unknowns except their source code differences between different correct implementations of the same problem. We put the power of deep learning to the test for automatically extracting information from the hierarchical structure of abstract syntax trees to represent source code. This paper aims to assess the feasibility of using purely static information (e.g., abstract syntax tree or AST) of applications to predict performance change based on the change in code structure. This research will enable performance-aware application development since every version of the application will continue to contribute to the corpora, which will enhance the performance of the model. We evaluate several deep learning-based representation learning techniques for source code. Our results show that tree-based Long Short-Term Memory (LSTM) models can leverage source code's hierarchical structure to discover latent representations. Specifically, LSTM-based predictive models built using a single problem and a combination of multiple problems can correctly predict if a source code will perform better or worse up to 84% and 73% of the time, respectively. Tarek Ramadan, Tanzima Z. Islam, Chase Phelps, Nathan Pinnow, Jayaraman J. Thiagarajan |
ISPASS | 2 |
| 2020 | Towards Aggregation Based I/O Optimization for Scaling Bioinformatics ApplicationsabstractBioinformatics software often integrates multiple off-the-shelf programs into a single compute pipeline. Each stand-alone program generates output, that is frequently saved into a plain-text file, which is then processed as input by the program that is responsible for the next stage of the computation. Modern advances in genome sequencing and protein structure resolution methods have yielded large amounts of data, that when processed by bioinformatics compute pipelines, results in vast numbers of file reads and writes. Since computation capabilities of large-scale distributed systems grow much faster than their I/O bandwidth, processing such big data using these compute pipelines will not scale as the amount of data grows. For this work, we motivate and demonstrate a dynamic interception-based I/O analysis tool to assess the file read and write characteristics of a protein mutation generation pipeline. We discuss how our analysis tool can be further extended to apply compression and in-site analysis and has the potential to scale I/O-intensive bioinformatics applications on high performance computing (HPC) systems. Jack Stratton, Michael Albert 0006, Quentin Jensen, Max Ismailov, Filip Jagodzinski, Tanzima Z. Islam |
COMPSAC | 6 |
| 2019 | Performance optimality or reproducibility: that is the questionabstractThe era of extremely heterogeneous supercomputing brings with itself the devil of increased performance variation and reduced reproducibility. There is a lack of understanding in the HPC community on how the simultaneous consideration of network traffic, power limits, concurrency tuning, and interference from other jobs impacts application performance. Tapasya Patki, Jayaraman J. Thiagarajan, Alexis Ayala, Tanzima Z. Islam |
SC | 4 |
| 2018 | PADDLE: Performance Analysis Using a Data-Driven Learning EnvironmentabstractThe use of machine learning techniques to model execution time and power consumption, and, more generally, to characterize performance data is gaining traction in the HPC community. Although this signifies huge potential for automating complex inference tasks, a typical analytics pipeline requires selecting and extensively tuning multiple components ranging from feature learning to statistical inferencing to visualization. Further, the algorithmic solutions often do not generalize between problems, thereby making it cumbersome to design and validate machine learning techniques in practice. In order to address these challenges, we propose a unified machine learning framework, PADDLE, which is specifically designed for problems encountered during analysis of HPC data. The proposed framework uses an information-theoretic approach for hierarchical feature learning and can produce highly robust and interpretable models. We present user-centric workflows for using PADDLE and demonstrate its effectiveness in different scenarios: (a) identifying causes of network congestion; (b) determining the best performing linear solver for sparse matrices; and (c) comparing performance characteristics of parent and proxy application pairs. Jayaraman J. Thiagarajan, Rushil Anirudh, Bhavya Kailkhura, Tanzima Z. Islam, Abhinav Bhatele, Jae-Seung Yeom, Todd Gamblin |
IPDPS | 5 |
| 2017 | MetaKV: A Key-Value Store for Metadata Management of Distributed Burst BuffersabstractDistributed burst buffers are a promising storage architecture for handling I/O workloads for exascale computing. Their aggregate storage bandwidth grows linearly with system node count. However, although scientific applications can achieve scalable write bandwidth by having each process write to its node-local burst buffer, metadata challenges remain formidable, especially for files shared across many processes. This is due to the need to track and organize file segments across the distributed burst buffers in a global index. Because this global index can be accessed concurrently by thousands or more processes in a scientific application, the scalability of metadata management is a severe performance-limiting factor. In this paper, we propose MetaKV: a key-value store that provides fast and scalable metadata management for HPC metadata workloads on distributed burst buffers. MetaKV complements the functionality of an existing key-value store with specialized metadata services that efficiently handle bursty and concurrent metadata workloads: compressed storage management, supervised block clustering, and log-ring based collective message reduction. Our experiments demonstrate that MetaKV outperforms the state-of-the-art key-value stores by a significant margin. It improves put and get metadata operations by as much as 2.66× and 6.29×, respectively, and the benefits of MetaKV increase with increasing metadata workload demand. Teng Wang 0001, Adam Moody, Yue Zhu 0002, Kathryn Mohror, Kento Sato, Tanzima Z. Islam, Weikuan Yu |
IPDPS | 6 |
| 2016 | CMT-Bone - A Proxy Application for Compressible Multiphase Turbulent FlowsabstractCMT-bone is a proxy app of CMT-nek, which is a solver of the compressible Navier-Stokes equations for multiphase flows being developed at University of Florida. While the objective of CMT-nek is to perform high fidelity, predictive simulations of particle laden explosively dispersed turbulent flows, the goal of CMT-bone is to mimic the computational behavior of CMT-nek in terms of operation counts, memory access patterns for data and performance characteristics of hardware devices (memory, cache, floating point unit, etc.). CMT-bone, as a proxy app, has a tremendous potential to be an important benchmark to realize tradeoffs in HPC software, hardware, and algorithm design aspart of the co-design process. Tania Banerjee, Jason Hackl, Mrugesh Sringarpure, Tanzima Z. Islam, S. Balachandar 0001, Thomas L. Jackson, Sanjay Ranka |
HiPC | 4 |
| 2016 | I/O Aware Power ShiftingabstractPower limits on future high-performance computing (HPC) systems will constrain applications. However, HPC applications do not consume constant power over their lifetimes. Thus, applications assigned a fixed power bound may be forced to slow down during high-power computation phases, but may not consume their full power allocation during low-power I/O phases. This paper explores algorithms that leverage application semantics -- phase frequency, duration and power needs -- to shift unused power from applications in I/O phases to applications in computation phases, thus improving system-wide performance. We design novel techniques that include explicit staggering of applications to improve power shifting. Compared to executing without power shifting, our algorithms can improve average performance by up to 8% or improve performance of a single, high-priority application by up to 32%. Lee Savoie, David K. Lowenthal, Bronis R. de Supinski, Tanzima Z. Islam, Kathryn Mohror, Barry Rountree, Martin Schulz 0001 |
IPDPS | 4 |
| 2016 | A machine learning framework for performance coverage analysis of proxy applicationsabstractProxy applications are written to represent subsets of performance behaviors of larger, and more complex applications that often have distribution restrictions. They enable easy evaluation of these behaviors across systems, e.g., for procurement or co-design purposes. However, the intended correlation between the performance behaviors of proxy applications and their parent codes is often based solely on the developer's intuition. In this paper, we present novel machine learning techniques to methodically quantify the coverage of performance behaviors of parent codes by their proxy applications. We have developed a framework, VERITAS, to answer these questions in the context of on-node performance: (a) which hardware resources are covered by a proxy application and how well, and (b) which resources are important, but not covered. We present our techniques in the context of two benchmarks, STREAM and DGEMM, and two production applications, OpenMC and CMTnek, and their respective proxy applications. Tanzima Z. Islam, Jayaraman J. Thiagarajan, Abhinav Bhatele, Martin Schulz 0001, Todd Gamblin |
SC | 1 |
| 2014 | Batchsubmit: a high-volume batch submission system for earthquake engineering simulationabstractSUMMARY Network for Earthquake Engineering Simulation (NEES) is a network of 14 earthquake engineering labs distributed across the USA. As a part of the NEES effort NEESComm operates a comprehensive cyberinfrastructure that consists of the NEEShub and the NEES Project Warehouse. NEESComm provides consistent access to several high performance computing (HPC) venues. These venues include Extreme Science and Engineering Discovery Environment, the Open Science Grid, Purdue Supercomputers, and NEEShub servers. In this paper, we describe the system we developed, batchsubmit, which allows NEES researchers to make use of all these venues through the NEEShub science gateway. Copyright © 2014 John Wiley & Sons, Ltd. Anup Mohan, Thomas J. Hacker, Gregory P. Rodgers, Tanzima Z. Islam |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | Reliable and Efficient Distributed Checkpointing System for Grid Environments
Tanzima Z. Islam, Saurabh Bagchi, Rudolf Eigenmann |
J. Grid Comput. | 1 |
| 2012 | McrEngine: a scalable checkpointing system using data-aware aggregation and compressionabstractHigh performance computing (HPC) systems use checkpoint-restart to tolerate failures. Typically, applications store their states in checkpoints on a parallel file system (PFS). As applications scale up, checkpoint-restart incurs high overheads due to contention for PFS resources. The high overheads force large-scale applications to reduce checkpoint frequency, which means more compute time is lost in the event of failure. We alleviate this problem through a scalable checkpointrestart system, MCRENGINE. MCRENGINE aggregates checkpoints from multiple application processes with knowledge of the data semantics available through widely-used I/O libraries, e.g., HDF5 and netCDF, and compresses them. Our novel scheme improves compressibility of checkpoints up to 115% over simple concatenation and compression. Our evaluation with large-scale application checkpoints show that MCRENGINE reduces checkpointing overhead by up to 87% and restart overhead by up to 62% over a baseline with no aggregation or compression. Tanzima Z. Islam, Kathryn Mohror, Saurabh Bagchi, Adam Moody, Bronis R. de Supinski, Rudolf Eigenmann |
SC | 1 |
| 2009 | FALCON: a system for reliable checkpoint recovery in shared grid environmentsabstractIn Fine-Grained Cycle Sharing (FGCS) systems, machine owners voluntarily share their unused CPU cycles with guest jobs, as long as their performance degradation is tolerable. However, unpredictable evictions of guest jobs lead to fluctuating completion times. Checkpoint-recovery is an attractive mechanism for recovering from such "failures". Today's FGCS systems often use expensive, high-performance dedicated checkpoint servers. However, in geographically distributed clusters, this may incur high checkpoint transfer latencies. In this paper we present a system called Falcon that uses available disk resources of the FGCS machines as shared checkpoint repositories. However, an unavailable storage host may lead to loss of checkpoint data. Therefore, we model failures of storage hosts and develop a prediction algorithm for choosing reliable checkpoint repositories. We experiment with Falcon in the university-wide Condor testbed at Purdue and show improved and consistent performance for guest jobs in the presence of irregular resource availability. Tanzima Z. Islam, Saurabh Bagchi, Rudolf Eigenmann |
SC | 1 |