VLDB 2026 Research / reviewers in the wild / expert
Jim M. Brandt
dblp:10/2130 · also James Brandt, James M. Brandt
· DBLP profile ↗
36ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-8605-5795ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 8 first-author · 10 since 2021Computer networks · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Autonomy Loop for Dynamic HPC Job Time Limit Adjustment
Thomas Jakobsche, Osman Seckin Simsek, Jim M. Brandt, Ann C. Gentile, Florina M. Ciorba |
Euro-Par (1) | 3 |
| 2025 | Bringing Differential Privacy to HPC: Privacy-Preserving Transformations of HPC TracesabstractMonitoring HPC systems yields valuable insights into user behavior, aiding resource management, collaborative research, and software design. However, privacy concerns raise the barrier for real-world HPC trace sharing between HPC facilities and researchers. Traditional anonymization methods fall short as user behavior remains identifiable. To address this, we propose a robust toolset for privacy protection of HPC traces using Differential Privacy (DP). Our toolset offers a set of DP algorithms, metrics, and visualizations to empower HPC operators to protect users' sensitive information under a privacy protection guarantee. We evaluated our toolset over real HPC systems traces for different parameters and data aggregations. Moreover, we show that machine learning models trained on privacy-preserved logs maintain accuracy compared to real data, which supports data publishing and sharing across different computing facilities. Ana Veroneze Solórzano, Rohan Basu Roy, Benjamin Schwaller, Sara Walton, Jim M. Brandt, Devesh Tiwari |
HPDC | 5 |
| 2025 | Job Grouping Based Intelligent Resource Prediction Framework
Beste Oztop, Benjamin Schwaller, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun |
JSSPP | 4 |
| 2024 | Toward Sustainable HPC: In-Production Deployment of Incentive-Based Power Efficiency Mechanism on the Fugaku SupercomputerabstractThis paper describes the deployment and operational experience of a novel incentive-based power-control strategy on the Fugaku supercomputer. Our incentive-based program, termed Fugaku Points, provides knobs to users to apply power control functions to improve the overall power efficiency of the supercomputer toward achieving HPC sustainability in terms of its environmental implications. We also discuss new operational opportunities, challenges, and future directions. Ana Veroneze Solórzano, Kento Sato, Keiji Yamamoto, Fumiyoshi Shoji, Jim M. Brandt, Benjamin Schwaller, Sara Walton, Jennifer Green, Devesh Tiwari |
SC | 5 |
| 2024 | Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active LearningabstractWith the increasing scale and complexity of High-Performance Computing (HPC) systems, performance variations in applications caused by anomalies have become significant bottlenecks in system health and operational efficiency. As we move towards exascale systems, these variations become more prominent due to the increased sharing of resources. Such variations lead to lower energy efficiency and higher operational costs. To mitigate these problems, one must quickly and accurately diagnose the root cause of the anomalies at scale. One way to evaluate system health and identify the underlying causes is by manually examining certain performance metrics in telemetry data or using rule-based methods. Due to the daily size of telemetry data reaching terabytes and the fact that the numeric telemetry data contains thousands of metrics, manual analysis of telemetry to diagnose problems becomes challenging. Given these limitations, Machine Learning (ML)-based approaches have been gaining popularity as they have been shown to be effective and practical in diagnosing previously encountered performance anomalies. One primary challenge for supervised ML models is that they require a significant amount of labeled samples during training. However, obtaining many labels for anomalies is extremely difficult and costly, considering anomalies occur infrequently and real-world numeric system telemetry data is hard to label since it contains thousands of metrics. This paper proposes a novel active learning-based framework that diagnoses performance anomalies (i.e., identifying the type of an anomaly) in HPC systems at runtime using significantly fewer labeled samples compared to state-of-the-art ML-based approaches. We show that the proposed framework achieves the same F1-score compared to a supervised approach using much fewer labeled samples (i.e., 16x fewer samples for achieving a 0.78 F1-score, 11x fewer samples for achieving a 0.82 F1-score), even when there are previously unseen applications and application inputs in the test dataset. Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | Prodigy: Towards Unsupervised Anomaly Detection in Production HPC SystemsabstractPerformance variations caused by anomalies in modern High Performance Computing (HPC) systems lead to decreased efficiency, impaired application performance, and increased operational costs. While machine learning (ML)-based frameworks for automated anomaly detection (often based on time series telemetry data) are gaining popularity in the literature, practical deployment challenges are often overlooked. Some ML-based frameworks require extensive customization, while others need a rich set of labeled samples, none of which are feasible for a production HPC system. Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun |
SC | 6 |
| 2022 | ALBADross: Active Learning Based Anomaly Diagnosis for Production HPC SystemsabstractDiagnosing causes of performance variations in High-Performance Computing (HPC) systems is a daunting chal-lenge due to the systems' scale and complexity. Variations in application performance result in premature job termination, lower energy efficiency, or wasted computing resources. One potential solution is manual root-cause analysis based on system telemetry data. However, this approach has become an increasingly time-consuming procedure as the process relies on human expertise and the size of telemetry data is voluminous. Recent research employs supervised machine learning (ML) models to diagnose previously encountered performance anomalies in compute nodes automatically. However, these models generally necessitate vast amounts of labeled samples that represent anomalous and healthy states of an application during training. The demand for labeled samples is constraining because gathering labeled samples is difficult and costly, especially considering anomalies that occur infrequently. This paper proposes a novel active learning-based framework that diagnoses previously encountered performance anomalies in HPC systems using significantly fewer labeled samples compared to state-of-the-art ML-based frameworks. Our framework combines an active learning-based query strategy and a supervised classifier to minimize the number of labeled samples required to achieve a target performance score. We evaluate our framework on a production HPC system and a testbed HPC cluster using real and proxy applications. We show that our framework, ALBADross, achieves a 0.95 Fl-score using 28x fewer labeled samples compared to a supervised approach with equal Fl-score, even when there are previously unseen applications and application inputs in the test dataset. Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Ayse K. Coskun |
CLUSTER | 6 |
| 2022 | Metrics for Packing Efficiency and Fairness of HPC Cluster Batch Job SchedulingabstractDevelopment of job scheduling algorithms, which directly influence High-Performance Computing (HPC) clusters performance, is hindered because popular scheduling quality metrics, such as Bounded Slowdown, poorly correlate with global scheduling objectives that include job packing efficiency and fairness. This report proposes Area Weighted Response Time, a metric that offers an unbiased representation of job packing efficiency, and presents a class of new metrics, Priority Weighted Specific Response Time, that assess both packing efficiency and fairness of schedules. The provided examples of simulation of scheduling of real workload traces and analysis of the resulting schedules with the help of these metrics and conventional metrics, demonstrate that although Bounded Slowdown can be readily improved by modifying the standard First Come First Served backfilling algorithm and by using existing techniques of estimating job runtime, these improvements are accompanied by significant degradation of job packing efficiency and fairness. In contrast, improving job packing efficiency and fairness over the standard backfilling algorithm, which is designed to target those objectives, is difficult. It requires further algorithm development and more accurate runtime estimation techniques that reduce frequency of underpredictions. Alexander V. Goponenko, Kenneth Lamar, Christina L. Peterson, Benjamin A. Allan, Jim M. Brandt, Damian Dechev |
SBAC-PAD | 5 |
| 2021 | Backfilling HPC Jobs with a Multimodal-Aware PredictorabstractJob scheduling aims to minimize the turnaround time on the submitted jobs while catering to the resource constraints of High Performance Computing (HPC) systems. The challenge with scheduling is that it must honor job requirements and priorities while actual job run times are unknown. Although approaches have been proposed that use classification techniques or machine learning to predict job run times for scheduling purposes, these approaches do not provide a technique for reducing underprediction, which has a negative impact on scheduling quality. A common cause of underprediction is that the distribution of the duration for a job class is multimodal, causing the average job duration to fall below the expected duration of longer jobs. In this work, we propose the Top Percent predictor, which uses a hierarchical classification scheme to provide better accuracy for job run time predictions than the user-requested time. Our predictor addresses multimodal job distributions by making a prediction that is higher than a specified percentage of the observed job run times. We integrate the Top Percent predictor into scheduling algorithms and evaluate the performance using schedule quality metrics found in literature. To accommodate the user policies of HPC systems, we propose priority metrics that account for job flow time, job resource requirements, and job priority. The experiments demonstrate that the Top Percent predictor outperforms the related approaches when evaluated using our proposed priority metrics. Kenneth Lamar, Alexander V. Goponenko, Christina L. Peterson, Benjamin A. Allan, Jim M. Brandt, Damian Dechev |
CLUSTER | 5 |
| 2021 | E2EWatch: An End-to-End Anomaly Diagnosis Framework for Production HPC Systems
Burak Aksar, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun |
Euro-Par | 5 |
| 2021 | Delay sensitivity-driven congestion mitigation for HPC systemsabstractModern high-performance computing (HPC) systems concurrently execute multiple distributed applications that contend for the high-speed network leading to congestion. Consequently, application runtime variability and suboptimal system utilization are observed in production systems. To address these problems, we propose Netscope, a congestion mitigation framework based on a novel delay sensitivity metric. Delay sensitivity of an application is used to quantify the impact of congestion on its runtime. Netscope uses delay sensitivity estimates to drive a congestion mitigation mechanism to selectively throttle applications that are less susceptible to congestion. We evaluate Netscope on two Cray Aries systems, including a production supercomputer, on common scientific applications. Our evaluation shows that Netscope has a low training cost and accurately estimates the impact of congestion on application runtime with a correlation between 0.7 and 0.9. Moreover, Netscope reduces application tail runtime increase by up to 16.3x while improving the median system utility by 12%. Archit Patke, Saurabh Jha, Haoran Qiu, Jim M. Brandt, Ann C. Gentile, Joe Greenseid, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICS | 4 |
| 2021 | Systematically inferring I/O performance variability by examining repetitive job behaviorabstractMonitoring and analyzing I/O behaviors is critical to the efficient utilization of parallel storage systems. Unfortunately, with increasing I/O requirements and resource contention, I/O performance variability is becoming a significant concern. This paper investigates I/O behavior and performance variability on a large-scale high-performance computing (HPC) system using a novel methodology that identifies similarity across jobs from the same application leveraging an I/O characterization tool and then, detects potential I/O performance variability across jobs of the same application. We demonstrate and discuss how our unique methodology can be used to perform temporal and feature analyses to detect interesting I/O performance variability patterns in production HPC systems, and their implications for operating/managing large-scale systems. Emily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt, Devesh Tiwari |
SC | 4 |
| 2020 | Towards workload-adaptive scheduling for HPC clustersabstractThe performance of HPC clusters depends on efficient scheduling of jobs. However, modern schedulers generally lack real-time information about resource utilization and require users to provide information, which is seldom accurate, on job requirements. The problem is exacerbated as HPC systems become increasingly more complicated and heterogeneous, which gives rise to new resource constraints (GPU, parallel file system, network bandwidth, burst buffers, etc.) In this work, we integrated data from LDMS, the Lightweight Distributed Metric Service, with Slurm, a popular job scheduler. To demonstrate the capabilities of such integration, we enabled scheduling based on the Lustre file system throughput. We demonstrated benefits of measurement of real-time utilization, prediction of applications requirements from historical data, and finer control of resources, in a preliminary evaluation of scheduling on a cluster of virtual machines. We also identified the possibility of further improving the scheduling efficiency through workload-adaptive scheduling, by adjusting the scheduling based on characteristics of the pending job. We validated the feasibility of this strategy by simulating job executions in our custom-made HPC cluster simulator. Alexander V. Goponenko, Ramin Izadpanah, Jim M. Brandt, Damian Dechev |
CLUSTER | 3 |
| 2020 | HPC System Data Pipeline to Enable Meaningful Insights through Analysis-Driven VisualizationsabstractThe increasing complexity of High Performance Computing (HPC) systems has created a growing need for facilitating insight into system performance and utilization for administrators and users. The strides made in HPC system monitoring data collection have produced terabyte/day sized time-series data sets rich with critical information, but it is onerous to extract and construe meaningful information from these metrics. We have designed and developed an architecture that enables flexible, as-needed, run-time analysis and presentation capabilities for HPC monitoring data. Our architecture enables quick and efficient data filtration and analysis. Complex runtime or historical analyses can be expressed as Python-based computations. Results of analyses and a variety of HPC oriented summaries are displayed in a Grafana front-end interface. To demonstrate our architecture, we have deployed it in production for a 1500-node HPC system and have developed analyses and visualizations requested by system administrators, and later employed by users, to track key metrics about the cluster at a job, user, and system level. Our architecture is generic, applicable to any *-nix based system, and it is extensible to supporting multi-cluster HPC centers. We structure it with easily replaced modules that allow unique customization across clusters and centers. In this paper, we describe the data collection and storage infrastructure, the application created to query and analyze data from a custom database, and the visual displays created to provide clear insights into HPC system behavior. Benjamin Schwaller, Nick Tucker, Tom Tucker, Benjamin A. Allan, Jim M. Brandt |
CLUSTER | 5 |
| 2020 | Measuring Congestion in High-Performance Datacenter Interconnects
Saurabh Jha, Archit Patke, Jim M. Brandt, Ann C. Gentile, Benjamin Lim, Michael T. Showerman, Gregory H. Bauer, Larry Kaplan, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer |
NSDI | 3 |
| 2019 | HPAS: An HPC Performance Anomaly Suite for Reproducing Performance VariationsabstractModern high performance computing (HPC) systems, including supercomputers, routinely suffer from substantial performance variations. The same application with the same input can have more than 100% performance variation, and such variations cause reduced efficiency and wasted resources. There have been recent studies on performance variability and on designing automated methods for diagnosing "anomalies" that cause performance variability. These studies either observe data collected from HPC systems, or they rely on synthetic reproduction of performance variability scenarios. However, there is no standardized way of creating performance variability inducing synthetic anomalies; so, researchers rely on designing ad-hoc methods for reproducing performance variability. Emre Ates, Yijia Zhang 0002, Burak Aksar, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun |
ICPP | 4 |
| 2019 | Online Diagnosis of Performance Variation in HPC Systems Using Machine LearningabstractAs the size and complexity of high performance computing (HPC) systems grow in line with advancements in hardware and software technology, HPC systems increasingly suffer from performance variations due to shared resource contention as well as software- and hardware-related problems. Such performance variations can lead to failures and inefficiencies, which impact the cost and resilience of HPC systems. To minimize the impact of performance variations, one must quickly and accurately detect and diagnose the anomalies that cause the variations and take mitigating actions. However, it is difficult to identify anomalies based on the voluminous, high-dimensional, and noisy data collected by system monitoring infrastructures. This paper presents a novel machine learning based framework to automatically diagnose performance anomalies at runtime. Our framework leverages historical resource usage data to extract signatures of previously-observed anomalies. We first convert collected time series data into easy-to-compute statistical features. We then identify the features that are required to detect anomalies, and extract the signatures of these anomalies. At runtime, we use these signatures to diagnose anomalies with negligible overhead. We evaluate our framework using experiments on a real-world HPC supercomputer and demonstrate that our approach successfully identifies 98 percent of injected anomalies and consistently outperforms existing anomaly diagnosis techniques. Ozan Tuncer, Emre Ates, Yijia Zhang 0002, Ata Turk, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Large-Scale System Monitoring Experiences and RecommendationsabstractMonitoring of High Performance Computing (HPC) platforms is critical to successful operations, can provide insights into performance-impacting conditions, and can inform methodologies for improving science throughput. However, monitoring systems are not generally considered core capabilities in system requirements specifications nor in vendor development strategies. In this paper we present work performed at a number of large-scale HPC sites towards developing monitoring capabilities that fill current gaps in ease of problem identification and root cause discovery. We also present our collective views, based on the experiences presented, on needs and requirements for enabling development by vendors or users of effective sharable end-to-end monitoring capabilities. Ville Ahlgren, Stefan Andersson, Jim M. Brandt, Nicholas Cardo, Sudheer Chunduri, Jeremy Enos, Parks Fields, Ann C. Gentile, Richard A. Gerber, Michael Gienger, Joe Greenseid, Annette Greiner, Bilel Hadri, Dennis Hoppe, Urpo Kaila, Kaki Kelly, Mark Klein 0002, Alex Kristiansen, Stephen Leak, Mike Mason, Kevin T. Pedretti, Jean-Guillaume Piccinali, Jason Repik, Jim Rogers, Susanna Salminen, Michael T. Showerman, Cary Whitney, Jim Williams |
CLUSTER | 3 |
| 2018 | Characterizing Supercomputer Traffic Networks Through Link-Level AnalysisabstractWe present techniques for characterizing bandwidth and congestion characteristics of supercomputer High-Speed Networks (HSN). By utilizing a link-level perspective, we gain generality over analyses which are tied to specific topologies. We illustrate these techniques using five months of a Blue Waters production dataset consisting of network utilization and congestion counters. We find that: i) execution time of the communication-heavy applications is highly correlated to network stalls observed in the network topology and increase in application runtime can be as high as 1.7x with nominal increase in stalls, ii) heterogeneity in the available link bandwidth in the network can lead to backpressure and congestion even when the network is not underprovisioned, and (iii) links connected to I/O nodes are no more likely to observe congestion during operational hours than any other link in the system. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
CLUSTER | 2 |
| 2018 | Taxonomist: Application Detection Through Rich Monitoring Data
Emre Ates, Ozan Tuncer, Ata Turk, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun |
Euro-Par | 5 |
| 2018 | Integrating Low-latency Analysis into HPC System MonitoringabstractThe growth of High Performance Computer (HPC) systems increases the complexity with respect to understanding resource utilization, system management, and performance issues. While raw performance data is increasingly exposed at the component level, the usefulness of the data is dependent on the ability to do meaningful analysis on actionable timescales. However, current system monitoring infrastructures largely focus on data collection, with analysis performed off-system in post-processing mode. This increases the time required to provide analysis and feedback to a variety of consumers. Ramin Izadpanah, Nichamon Naksinehaboon, Jim M. Brandt, Ann C. Gentile, Damian Dechev |
ICPP | 3 |
| 2018 | An Efficient Latch-free Database Index Based on Multi-dimensional ListsabstractIn the interests of improving database performance, researchers have considered lock-free data structures for their attractive progress guarantees and scalability. This paper considers the performance of a recently developed lock-free structure, multi-dimensional list (MDList), used as a database index in SOS, a high-performance, object-oriented database. In our tests, we find that MDList outperforms the existing locking structures in multi-threaded workloads. This is the first known use of MDList as an index structure in databases. Kenneth Lamar, Ramin Izadpanah, Jim M. Brandt, Damian Dechev |
IPCCC | 3 |
| 2017 | Holistic Measurement-Driven System AssessmentabstractIn high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer |
CLUSTER | 2 |
| 2016 | Continuous whole-system monitoring toward rapid understanding of production HPC applications and systems
Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Ann C. Gentile, Sophia Lefantzi, Steve Monk, Jeff Ogden, Mahesh Rajan, Joel Stevenson |
Parallel Comput. | 3 |
| 2015 | Toward Rapid Understanding of Production HPC Applications and SystemsabstractA detailed understanding of HPC application's resource needs and their complex interactions with each other and HPC platform resources is critical to achieving scalability and performance. Such understanding has been difficult to achieve because typical application profiling tools do not capture the behaviors of codes under the potentially wide spectrum of actual production conditions and because typical monitoring tools do not capture system resource usage information with high enough fidelity to gain sufficient insight into application performance and demands. In this paper we present both system and application profiling results based on data obtained through synchronized system wide monitoring on a production HPC cluster at Sandia National Laboratories (SNL). We demonstrate analytic and visualization techniques that we are using to characterize application and system resource usage under production conditions for better understanding of application resource needs. Our goals are to improve application performance (through understanding application-to-resource mapping and system throughput) and to ensure that future system capabilities match their intended workloads. Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Ann C. Gentile, Sophia Lefantzi, Steve Monk, Jeff Ogden, Mahesh Rajan, Joel Stevenson |
CLUSTER | 3 |
| 2015 | New Systems, New Behaviors, New Patterns: Monitoring Insights from System StandupabstractDisentangling significant and important log messages from those that are routine and unimportant can be a difficult task. Further, on a new system, understanding correlations between significant and possibly new types of messages and conditions that cause them can require significant effort and time. The initial standup of a machine can provide opportunities for investigating the parameter space of events and operations and thus for gaining insight into the events of interest. In particular, failure inducement and investigation of corner case conditions can provide knowledge of system behavior for significant issues that will enable easier diagnosis and mitigation of such issues for when they may actually occur during the platform lifetime. In this work, we describe the testing process and monitoring results from a testbed system in preparation for the ACES Trinity system. We describe how events in the initial standup including changes in configuration and software and corner case testing has provided insights that can inform future monitoring and operating conditions, both of our test systems and the eventual large-scale Trinity system. Jim M. Brandt, Ann C. Gentile, Cindy Martin, Jason Repik, Narate Taerat |
CLUSTER | 1 |
| 2015 | Extending LDMS to Enable Performance Monitoring in Multi-core ApplicationsabstractIdentifying design patterns that limit the performance of multi-core algorithms is a challenging task. There are many known methods by which threads synchronize their actions and each method may exhibit different behavior in different use cases. These use cases may vary in regards to the workload being executed, number of parallel tasks, dependencies between these tasks, and the behavior of the system scheduler. Restructuring algorithms to overcome performance limitations requires intimate knowledge on how these algorithms utilize the hardware. In our experience, we have found a lack of adequate tools to gain such knowledge. To address this, we have enhanced and implemented additional data sampler modules for OVIS's Lightweight Distributed Metric Service (LDMS) to enable scalable distributed collection of hardware performance counter data. These modules provide an interface by which LDMS can utilize the PAPI library, Linux perf tools, and RAPL to collect hardware performance data of interest. Using these samplers, we plan to monitor the intra-node behavior, including contention for node level shared resources, of multi-core applications for a diverse set of use cases. We are currently exploring how the values reported are affected by the level of concurrency, the synchronization methodologies, and progress guarantees. We hope to use this information to identify ways to restructure algorithms to increase their performance. Steven D. Feldman, Deli Zhang, Damian Dechev, Jim M. Brandt |
CLUSTER | 4 |
| 2014 | Demonstrating improved application performance using dynamic monitoring and task mappingabstractThis work demonstrates the integration of monitoring, analysis, and feedback to perform application-to-resource mapping that adapts to both static architecture features and dynamic resource state. In particular, we present a framework for mapping MPI tasks to compute resources based on run-time analysis of system-wide network data, architecture-specific routing algorithms, and application communication patterns. We address several challenges. Within each node, we collect local utilization data. We consolidate that information to form a global view of system performance, accounting for system-wide factors including competing applications. We provide an interface for applications to query the global information. Then we exploit the system information to change the mapping of tasks to nodes so that system bottlenecks are avoided. We demonstrate the benefit of this monitoring and feedback by remapping MPI tasks based on route-length, bandwidth, and credit-stalls metrics for a parallel sparse matrix-vector multiplication kernel. In the best case, remapping based on dynamic network information in a congested environment recovered 48.9% of the time lost to congestion, reducing matrix-vector multiplication time by 7.8%. Our experiments focus on the Cray XE/XK platform, but the integration concepts are generally applicable to any platform for which applicable metrics and route knowledge can be obtained. Jim M. Brandt, Karen D. Devine, Ann C. Gentile, Kevin T. Pedretti |
CLUSTER | 1 |
| 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and ApplicationsabstractUnderstanding how resources of High Performance Compute platforms are utilized by applications both individually and as a composite is key to application and platform performance. Typical system monitoring tools do not provide sufficient fidelity while application profiling tools do not capture the complex interplay between applications competing for shared resources. To gain new insights, monitoring tools must run continuously, system wide, at frequencies appropriate to the metrics of interest while having minimal impact on application performance. We introduce the Lightweight Distributed Metric Service for scalable, lightweight monitoring of large scale computing systems and applications. We describe issues and constraints guiding deployment in Sandia National Laboratories' capacity computing environment and on the National Center for Supercomputing Applications' Blue Waters platform including motivations, metrics of choice, and requirements relating to the scale and specialized nature of Blue Waters. We address monitoring overhead and impact on application performance and provide illustrative profiling results. Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Paul Cassella, Jeremy Enos, Joshi Fullop, Ann C. Gentile, Steve Monk, Nichamon Naksinehaboon, Jeff Ogden, Mahesh Rajan, Michael T. Showerman, Joel Stevenson, Narate Taerat, Thomas W. Tucker |
SC | 3 |
| 2012 | Filtering log data: Finding the needles in the HaystackabstractLog data is an incredible asset for troubleshooting in large-scale systems. Nevertheless, due to the ever-growing system scale, the volume of such data becomes overwhelming, bringing enormous burdens on both data storage and data analysis. To address this problem, we present a 2-dimensional online filtering mechanism to remove redundant and noisy data via feature selection and instance selection. The objective of this work is two-fold: (i) to significantly reduce data volume without losing important information, and (ii) to effectively promote data analysis. We evaluate this new filtering mechanism by means of real environmental data from the production supercomputers at Oak Ridge National Laboratory and Sandia National Laboratory. Our preliminary results demonstrate that our method can reduce more than 85% disk space, thereby significantly reducing analysis time. Moreover, it also facilitates better failure prediction and diagnosis by more than 20%, as compared to the conventional predictive approach relying on RAS (Reliability, Availability, and Serviceability) events alone. Li Yu 0006, Ziming Zheng, Zhiling Lan, Terry R. Jones, Jim M. Brandt, Ann C. Gentile |
DSN | 5 |
| 2010 | Using Cloud Constructs and Predictive Analysis to Enable Pre-Failure Process Migration in HPC SystemsabstractAccurate failure prediction in conjunction with efficient process migration facilities including some Cloud constructs can enable failure avoidance in large-scale high performance computing (HPC) platforms. In this work we demonstrate a prototype system that incorporates our probabilistic failure prediction system with virtualization mechanisms and techniques to provide a whole system approach to failure avoidance. This work utilizes a failure scenario based on a real-world HPC case study. Jim M. Brandt, Frank Chen 0001, Vincent De Sapio, Ann C. Gentile, Jackson R. Mayo, Philippe P. Pébay, Diana C. Roe, David C. Thompson 0001, Matthew Wong |
CCGRID | 1 |
| 2009 | Resource monitoring and management with OVIS to enable HPC in cloud computing environmentsabstractUsing the cloud computing paradigm, a host of companies promise to make huge compute resources available to users on a pay-as-you-go basis. These resources can be configured on the fly to provide the hardware and operating system of choice to the customer on a large scale. While the current target market for these resources in the commercial space is Web development/hosting, this model has the lure of savings of ownership, operation, and maintenance costs, and thus sounds like an attractive solution for people who currently invest millions to hundreds of millions of dollars annually on high performance computing (HPC) platforms in order to support large-scale scientific simulation codes. Given the current interconnect bandwidth and topologies utilized in these commercial offerings, however, the only current viable market in HPC would be small-memory-footprint embarrassingly parallel or loosely coupled applications, which inherently require little to no inter-processor communication. While providing the appropriate resources (bandwidth, latency, memory, etc.) for the HPC community would increase the potential to enable HPC in cloud environments, this would not address the need for scalability and reliability, crucial to HPC applications. Providing for these needs is particularly difficult in commercial cloud offerings where the number of virtual resources can far outstrip the number of physical resources, the resources are shared among many users, and the resources may be heterogeneous. Advanced resource monitoring, analysis, and configuration tools can help address these issues, since they bring the ability to dynamically provide and respond to information about the platform and application state and would enable more appropriate, efficient, and flexible use of the resources key to enabling HPC. Additionally such tools could be of benefit to non-HPC cloud providers, users, and applications by providing more efficient resource utilization in general. Jim M. Brandt, Ann C. Gentile, Jackson R. Mayo, Philippe P. Pébay, Diana C. Roe, David C. Thompson 0001, Matthew Wong |
IPDPS | 1 |
| 2008 | Using Probabilistic Characterization to Reduce Runtime Faults in HPC SystemsabstractThe current trend in high performance computing is to aggregate ever larger numbers of processing and interconnection elements in order to achieve desired levels of computational power, This, however, also comes with a decrease in the Mean Time To Interrupt because the elements comprising these systems are not becoming significantly more robust. There is substantial evidence that the Mean Time To Interrupt vs. number of processor elements involved is quite similar over a large number of platforms. In this paper we present a system that uses hardware level monitoring coupled with statistical analysis and modeling to select processing system elements based on where they lie in the statistical distribution of similar elements. These characterizations can be used by the scheduler/resource manager to deliver a close to optimal set of processing elements given the available pool and the reliability requirements of the application. Jim M. Brandt, Bert J. Debusschere, Ann C. Gentile, Jackson R. Mayo, Philippe P. Pébay, David C. Thompson 0001, Matthew Wong |
CCGRID | 1 |
| 2008 | Ovis-2: A robust distributed architecture for scalable RASabstractResource utilization in High Performance Compute clusters can be improved by increased awareness of system state information. Sophisticated run-time characterization of system state in increasingly large clusters requires a scalable fault-tolerant RAS framework. In this paper we describe the architecture of OVIS-2 and how it meets these requirements. We describe some of the sophisticated statistical analysis, 3-D visualization, and use cases for these. Using this framework and associated tools allows the engineer to explore the behaviors and complex interactions of low level system elements while simultaneously giving the system administrator their desired level of detail with respect to ongoing system and component health. Jim M. Brandt, Bert J. Debusschere, Ann C. Gentile, Jackson R. Mayo, Philippe P. Pébay, David C. Thompson 0001, M. H. Wong |
IPDPS | 1 |
| 2006 | OVIS: a tool for intelligent, real-time monitoring of computational clustersabstractTraditional cluster monitoring approaches consider nodes in singleton, using manufacturer-specified extreme limits as thresholds for failure "prediction". We have developed a tool, OVIS, for monitoring and analysis of large computational platforms which, instead, uses a statistical approach to characterize single device behaviors from those of a large number of statistically similar devices. Baseline capabilities of OVIS include the visual display of deterministic information about state variables (e.g., temperature, CPU utilization, fan speed) and their aggregate statistics. Visual consideration of the cluster as a comparative ensemble, rather than as singleton nodes, is an easy and useful method for tuning cluster configuration and determining effects of realtime changes. Additionally, OVIS incorporates a novel Bayesian inference scheme to dynamically infer models for the normal behavior of a system and to determine bounds on the probability of values evinced in the system. Individual node values that are unlikely given the current applicable model are flagged as aberrant. This can be a much earlier indicator of problems than waiting for the crossing of some threshold that is necessarily set high to preclude too many false alarms. We present OVIS and discuss its applications in cluster configuration and environmental tuning and to abnormality and problem discovery in our production clusters Jim M. Brandt, Ann C. Gentile, D. J. Hale, Philippe P. Pébay |
IPDPS | 1 |
| 2005 | Meaningful Automated Statistical Analysis of Large Computational ClustersabstractAs clusters utilizing commercial off-the-shelf technology have grown from tens to thousands of nodes and typical job sizes have likewise increased, much effort has been devoted to improving the scalability of message-passing fabrics, schedulers, and storage. Largely ignored, however, has been the issue of predicting node failure, which also has a large impact on scalability. In fact, more than ten years into cluster computing, we are still managing this issue on a node-by-node basis even though available diagnostic data has grown immensely. We have built a tool that uses the statistical similarity of the large number of nodes in a cluster to infer the health of each individual node. In the poster, we first present real data and statistical calculations as foundational material and justification for our claims of similarity. Next we present our methodology and its implications for early notification of deviation from normal behavior, problem diagnosis, automatic code restart via interaction with scheduler, and airflow distribution monitoring in the machine room. A framework addressing scalability is discussed briefly. Lastly, we present case studies showing how our methodology has been used to detect aberrant nodes whose deviations are still far below the detection level of traditional methods. A summary of the results of the case studies appears below Jim M. Brandt, Ann C. Gentile, Youssef Marzouk 0001, Philippe P. Pébay |
CLUSTER | 1 |