EDBT 2026 Demo / reviewers in the wild / expert
Omar Aaziz
dblp:170/2069
· DBLP profile ↗
11ranked-venue papers
6as first author
6since 2021 · last 2024
0009-0000-9651-0299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active LearningabstractWith the increasing scale and complexity of High-Performance Computing (HPC) systems, performance variations in applications caused by anomalies have become significant bottlenecks in system health and operational efficiency. As we move towards exascale systems, these variations become more prominent due to the increased sharing of resources. Such variations lead to lower energy efficiency and higher operational costs. To mitigate these problems, one must quickly and accurately diagnose the root cause of the anomalies at scale. One way to evaluate system health and identify the underlying causes is by manually examining certain performance metrics in telemetry data or using rule-based methods. Due to the daily size of telemetry data reaching terabytes and the fact that the numeric telemetry data contains thousands of metrics, manual analysis of telemetry to diagnose problems becomes challenging. Given these limitations, Machine Learning (ML)-based approaches have been gaining popularity as they have been shown to be effective and practical in diagnosing previously encountered performance anomalies. One primary challenge for supervised ML models is that they require a significant amount of labeled samples during training. However, obtaining many labels for anomalies is extremely difficult and costly, considering anomalies occur infrequently and real-world numeric system telemetry data is hard to label since it contains thousands of metrics. This paper proposes a novel active learning-based framework that diagnoses performance anomalies (i.e., identifying the type of an anomaly) in HPC systems at runtime using significantly fewer labeled samples compared to state-of-the-art ML-based approaches. We show that the proposed framework achieves the same F1-score compared to a supervised approach using much fewer labeled samples (i.e., 16x fewer samples for achieving a 0.78 F1-score, 11x fewer samples for achieving a 0.82 F1-score), even when there are previously unseen applications and application inputs in the test dataset. Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Prodigy: Towards Unsupervised Anomaly Detection in Production HPC SystemsabstractPerformance variations caused by anomalies in modern High Performance Computing (HPC) systems lead to decreased efficiency, impaired application performance, and increased operational costs. While machine learning (ML)-based frameworks for automated anomaly detection (often based on time series telemetry data) are gaining popularity in the literature, practical deployment challenges are often overlooked. Some ML-based frameworks require extensive customization, while others need a rich set of labeled samples, none of which are feasible for a production HPC system. Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun |
SC | 4 |
| 2022 | IncProf: Efficient Source-Oriented Phase Identification for Application Behavior UnderstandingabstractLong running applications often have varying behaviors, here called phases. While considerable work in computer architecture has been done in identifying application phases based on how the hardware is being exercised, comparatively less work has been focused on identifying application phases based on regions of source code being executed. In this paper we introduce a new methodology and an efficient tool framework, IncProf, for observing and capturing the time-varying source execution behavior of applications, and for then deducing application phases from the resulting data. Uses of this capability include simply better understanding the varying behavior of long running applications, and for efficiently tracking deployed application performance in the future by providing information to identify good instrumentation points. Omar Aaziz, Mohammad Al-Tahat, Strahinja Trecakov, Jonathan E. Cook 0001 |
CLUSTER | 1 |
| 2022 | ALBADross: Active Learning Based Anomaly Diagnosis for Production HPC SystemsabstractDiagnosing causes of performance variations in High-Performance Computing (HPC) systems is a daunting chal-lenge due to the systems' scale and complexity. Variations in application performance result in premature job termination, lower energy efficiency, or wasted computing resources. One potential solution is manual root-cause analysis based on system telemetry data. However, this approach has become an increasingly time-consuming procedure as the process relies on human expertise and the size of telemetry data is voluminous. Recent research employs supervised machine learning (ML) models to diagnose previously encountered performance anomalies in compute nodes automatically. However, these models generally necessitate vast amounts of labeled samples that represent anomalous and healthy states of an application during training. The demand for labeled samples is constraining because gathering labeled samples is difficult and costly, especially considering anomalies that occur infrequently. This paper proposes a novel active learning-based framework that diagnoses previously encountered performance anomalies in HPC systems using significantly fewer labeled samples compared to state-of-the-art ML-based frameworks. Our framework combines an active learning-based query strategy and a supervised classifier to minimize the number of labeled samples required to achieve a target performance score. We evaluate our framework on a production HPC system and a testbed HPC cluster using real and proxy applications. We show that our framework, ALBADross, achieves a 0.95 Fl-score using 28x fewer labeled samples compared to a supervised approach with equal Fl-score, even when there are previously unseen applications and application inputs in the test dataset. Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Ayse K. Coskun |
CLUSTER | 4 |
| 2022 | LDMS Darshan Connector: For Run Time Diagnosis of HPC Application I/O PerformanceabstractPeriodic capture of comprehensive, usable I/O performance data for scientific applications requires an easy-to-use technique to record information throughout the execution without causing substantial performance effects. In this paper, we introduce a unique framework that provides low latency monitoring of I/O event data during run time. We implement a system-level infrastructure that continuously collects I/O application data from an existing I/O characterization tool to enable insights into the I/O application behavior and the components affecting it through analyses and visualizations. In this effort, we evaluate our framework by analyzing sampled I/O data captured from two HPC benchmark applications to understand the I/O behavior during the execution life of the applications. The result shows the utility of capturing I/O application performance and behavior. Sara Walton, Omar Aaziz, Ana Veroneze Solórzano, Benjamin Schwaller |
CLUSTER | 2 |
| 2021 | E2EWatch: An End-to-End Anomaly Diagnosis Framework for Production HPC Systems
Burak Aksar, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun |
Euro-Par | 3 |
| 2019 | Proxy or Imposter? A Method and Case Study to Determine the AnswerabstractAs the HPC community moves toward exascale, understanding application behavior is more important due to the increase in size and complexity of systems. While applications also grow larger in size and complexity, the need for proxy applications is crucial because of their ease of use and fast execution. They have become an essential aid for system vendors to evaluate new advanced architectures and for application developers to more quickly resolve algorithm and optimization issues. Therefore, proxies must be representative in behavior and function of the applications they mimic. In this work, we present a methodology to understand if a proxy represents a parent application based on a comparison of computational kernels, appropriate dynamic execution characteristics, and hardware bottlenecks. Based on this method, we conclude that miniQMC is a fairly good proxy for QMCPACK, but could be improved based on our analysis. Omar Aaziz, Jeanine E. Cook, Courtenay T. Vaughan, David Richards |
CLUSTER | 1 |
| 2018 | A Methodology for Characterizing the Correspondence Between Real and Proxy ApplicationsabstractProxy applications are a simplified means for stake-holders to evaluate how both hardware and software stacks might perform on the class of real applications that they are meant to model. However, characterizing the relationship between them and their behavior is not an easy task. We present a data-driven methodology for characterizing the relationship between real and proxy applications based on collecting runtime data from both and then using data analytics to find their correspondence and divergence. We use new capabilities for application-level monitoring within LDMS (Lightweight Distributed Monitoring System) to capture hardware performance counter and MPI-related data. To demonstrate the utility of this methodology, we present experimental evidence from two system platforms, using four proxy applications from the current ECP Proxy Application Suite and their corresponding parent applications (in the ECP application portfolio). Results show that each proxy analyzed is representative of its parent with respect to computation and memory behavior. We also analyze communication patterns separately using mpiP data and show that communication for these four proxy/parent pairs is also similar. Omar Aaziz, Jeanine E. Cook, Jonathan E. Cook 0001, Tanner Juedeman, David Richards, Courtenay T. Vaughan |
CLUSTER | 1 |
| 2018 | Modeling Expected Application Runtime for Characterizing and Assessing Job PerformanceabstractIn this paper, we present a methodology for modeling the expected runtime of a job based on historical application data and data from the job itself. This estimation model is useful for both for HPC users and administrators as a metric to compare the actual job runtime to, thus establishing a measure of performance of the job. We used job data, system data, and hardware performance counters in a near-zero overhead manner to model and assess job performance, in particular whether or not the job runtime was in line with expectations from historical application performance. We show over three proxy applications and three real applications that our estimations are within 5% of actual performance. Omar Aaziz, Jonathan E. Cook 0001, Mohammed Tanash |
CLUSTER | 1 |
| 2017 | YAViT (Yet Another Viz Tool): Raising the Level of Abstraction in End-User HPC InteractionsabstractBecause data collection in HPC systems happens on the nodes and is easily related to the job running on the node, tools presenting the data and subsequent analyses to the user generally present them at the job level. Our position is that this is the wrong level of abstraction and thus limits the value of the analyses, often dissuading users from using any of the offered tools. In this paper we present the position that tools need to present analyses at the level users are interested in, which is their applications. Omar Aaziz, Ujjwal Panthi, Jonathan E. Cook 0001 |
CLUSTER | 1 |
| 2015 | Push Me Pull You: Integrating Opposing Data Transport Modes for Efficient HPC Application MonitoringabstractWhile HPC system monitoring is a necessary and accepted practice, applications are still basically opaque in the production environment. For better HPC platform management and utilization, especially as platforms push towards exascale size, HPC applications need to be more transparent in their execution in the production environment. PROMON is a framework for application monitoring in the production environment, but its design concentrated on the front end issues of offering easy to use application instrumentation. This paper presents the integration of PROMON with LDMS, a proven efficient HPC system monitoring framework. PROMON and LDMS offer a case study in integrating two disparate instrumentation and monitoring models, and the lessons are applicable to other HPC monitoring issues. Omar Aaziz, Jonathan E. Cook 0001, Hadi Sharifi |
CLUSTER | 1 |