Michael Ott 0001

dblp:57/4484-1 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
5since 2021 · last 2023
0009-0007-0278-9497ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2023 Phase-aware System-Side Sampling for HPC
abstract
HPC compute centers always benefit from better insights into the application mix running on their systems. We present a low-overhead statistical sampling tool running in the background on the system side, which can capture application compute phases. Our tool leverages eBPF (extended Berkeley Packet Filter) from modern Linux kernels to extract phase information provided by developers via instrumentation or instruction pointers. We outline how this tool can be integrated into the monitoring framework DCDB, and we show resulting performance insights.
Julian Scheipl, Amir Raoofy, Michael Ott 0001, Josef Weidendorfer
CF3
2022 Always-on instrumentation for application introspection in HPC
abstract
Obtaining insights into the dynamic behavior of user code is crucial for supercomputing centers to support both better operation and co-design of future systems. To this end, always-on instrumentation is the key: enabling all running code to dynamically forward metadata such as compute phase changes to the system would provide important information for those goals. To keep the overhead low, the system must be able to deactivate instrumentation points with high trigger frequency on demand. In this poster, we present a simple always-on instrumentation method for C/C++ which can be easily used by developers, copying a single source file into their code base. Our evaluations show that the overhead in the deactivated state is low enough for the manual instrumentation to stay compiled in, all the time.
Amir Raoofy, Josef Weidendorfer, Michael Ott 0001
CF3
2022 Operational Data Analytics in practice: Experiences from design to deployment in production HPC environments
Alessio Netti, Michael Ott 0001, Carla Guillén, Daniele Tafani, Martin Schulz 0001
Parallel Comput.2
2021 A Conceptual Framework for HPC Operational Data Analytics
abstract
This paper provides a broad framework for understanding trends in Operational Data Analytics (ODA) for High-Performance Computing (HPC) facilities. The goal of ODA is to allow for the continuous monitoring, archiving, and analysis of near real-time performance data, providing immediately actionable information for multiple operational uses. In this work, we combine two models to provide a comprehensive HPC ODA framework: one is an evolutionary model of analytics capabilities that consists of four types, which are descriptive, diagnostic, predictive and prescriptive, while the other is a four-pillar model for energy-efficient HPC operations that covers facility, system hardware, system software, and applications. This new framework is then overlaid with a description of current development and production deployments of ODA within leading-edge HPC facilities. Finally, we perform a comprehensive survey of ODA works and classify them according to our framework, in order to demonstrate its effectiveness.
Alessio Netti, Woong Shin, Michael Ott 0001, Torsten Wilde, Natalie J. Bates
CLUSTER3
2021 Correlation-wise Smoothing: Lightweight Knowledge Extraction for HPC Monitoring Data
abstract
Modern High-Performance Computing (HPC) and data center operators rely more and more on data analytics techniques to improve the efficiency and reliability of their operations. They employ models that ingest time-series monitoring sensor data and transform it into actionable knowledge for system tuning: a process known as Operational Data Analytics (ODA). However, monitoring data has a high dimensionality, is hardware-dependent and difficult to interpret. This, coupled with the strict requirements of ODA, makes most traditional data mining methods impractical and in turn renders this type of data cumbersome to process. Most current ODA solutions use ad-hoc processing methods that are not generic, are sensible to the sensors' features and are not fit for visualization. In this paper we propose a novel method, called Correlation-wise Smoothing (CS), to extract descriptive signatures from time-series monitoring data in a generic and lightweight way. Our CS method exploits correlations between data dimensions to form groups and produces image-like signatures that can be easily manipulated, visualized and compared. We evaluate the CS method on HPC-ODA, a collection of datasets that we release with this work, and show that it leads to the same performance as most state-of-the-art methods while producing signatures that are up to ten times smaller and up to ten times faster, while gaining visualizability, portability across systems and clear scaling properties.
Alessio Netti, Daniele Tafani, Michael Ott 0001, Martin Schulz 0001
IPDPS3
2020 Global Experiences with HPC Operational Data Measurement, Collection and Analysis
abstract
As we move into the exascale era, supercomputers grow larger, denser, more heterogeneous, and ever more complex. Operating such machines reliably and efficiently requires deep insight into the operational parameters of the machine itself as well as its supporting infrastructure. To fulfill this need, early adopter sites have started the development and deployment of Operational Data Analytics (ODA) frameworks allowing the continuous monitoring, archiving, and analysis of near realtime performance data from the machine and infrastructure levels, providing immediately actionable information for multiple operational uses. To understand their ODA goals, requirements, and use cases, we have conducted a survey among eight early adopter sites from the US, Europe, and Japan that operate top 50 high-performance computing systems. We have assessed the technologies leveraged to build their ODA frameworks, identified use cases and other push and pull factors that drive the sites' ODA activities, and report on their operational lessons.
Michael Ott 0001, Woong Shin, Norman Bourassa, Torsten Wilde, Stefan Ceballos, Melissa Romanus, Natalie J. Bates
CLUSTER1
2020 DCDB Wintermute: Enabling Online and Holistic Operational Data Analytics on HPC Systems
abstract
As we approach the exascale era, the size and complexity of HPC systems continues to increase, raising concerns about their manageability and sustainability. For this reason, more and more HPC centers are experimenting with fine-grained monitoring coupled with Operational Data Analytics (ODA) to optimize efficiency and effectiveness of system operations. However, while monitoring is a common reality in HPC, there is no well-stated and comprehensive list of requirements, nor matching frameworks, to support holistic and online ODA. This leads to insular ad-hoc solutions, each addressing only specific aspects of the problem.
Alessio Netti, Micha Müller, Carla Guillén, Michael Ott 0001, Daniele Tafani, Gence Ozer, Martin Schulz 0001
HPDC4
2019 From facility to application sensor data: modular, continuous and holistic monitoring with DCDB
abstract
Today's HPC installations are highly-complex systems, and their complexity will only increase as we move to exascale and beyond. At each layer, from facilities to systems, from runtimes to applications, a wide range of tuning decisions must be made in order to achieve efficient operation. This, however, requires systematic and continuous monitoring of system and user data. While many insular solutions exist, a system for holistic and facility-wide monitoring is still lacking in the current HPC ecosystem.
Alessio Netti, Micha Müller, Axel Auweter, Carla Guillén, Michael Ott 0001, Daniele Tafani, Martin Schulz 0001
SC5
2017 A Novel Approach for Job Scheduling Optimizations Under Power Cap for ARM and Intel HPC Systems
abstract
The ever-increasing energy demands of modern High Performance Computing (HPC) platforms is undeniably one of the most critical aspects for the future design and evolution of such systems. The capability of managing their energy consumption not only allows for significant reduction in electricity costs but is also a step forward on the road towards the exascale. Powercapping is a widely studied technique that contributes to address this challenge by instantaneously setting and maintaining a predefined power threshold (power cap) that cannot be exceeded. However, the lack of a centralized mechanism responsible for efficiently allocating the available power among resources and jobs may ultimately yield to fragmentation, low system utilization and increased user waiting times. Additionally, power cap violations can lead to high risk scenarios and/or increase operational costs. This paper proposes to prevent such issues with the introduction of the Enhanced Power Adaptive Scheduling (E-PAS) algorithm. The E-PAS algorithm combines scheduling and resource management mechanisms, correlating estimated and real power consumption data in order to optimize the resource utilization of the platform under a predefined power cap. The algorithm has been implemented in the widely used open-source resource and job management system SLURM and is planned to be pushed in a future mainstream version. Its effectiveness has been evaluated through real-scale experiments respectively on an ARM- and an Intel-based cluster of comparable size. All experiments have been performed using synthetic workloads from a set of mini-applications.
Dineshkumar Rajagopal, Daniele Tafani, Yiannis Georgiou 0002, David Glesser, Michael Ott 0001
HiPC5
2010 Automatic performance analysis with periscope
abstract
Abstract Performance analysis is essential to fully exploit the potential of high‐performance computers. With the imminence of petascale systems which will consist of ten thousands or even hundred thousands of processor cores, this task will increase in complexity. Hence, tools are required that automatically detect the performance bottlenecks and thus ease the performance analysis of an application. On large‐scale systems, collecting information about performance‐relevant events of an application can easily produce a huge amount of data whose analysis is very challenging. Aggregating the performance data during runtime and conducting the search for performance properties online allows users to distill essential performance bottlenecks without overwhelming the user with an uncontrollable load of data. In this paper we present the recent developments on Periscope, a highly scalable tool for the automatic distributed online search for the performance properties of large‐scale applications on high‐end computers. It allows for both detection of the performance bottlenecks limiting the scalability on parallel systems as well as pinpointing the issues concerning the single‐node performance of an application. Copyright © 2009 John Wiley & Sons, Ltd.
Michael Gerndt, Michael Ott 0001
Concurr. Comput. Pract. Exp.2
2009 Load Balance in the Phylogenetic Likelihood Kernel
abstract
Recent advances in DNA sequencing techniques have led to an unprecedented accumulation and availability of molecular sequence data that needs to be analyzed. This data explosion in combination with the multi-core revolution also affects the computational kernels for phylogenetic inference (reconstruction of evolutionary trees from molecular sequence data) under the widely-used Maximum Likelihood (ML) model. At present, analyses of so called multi-gene or phylogenomic alignments, i.e., input data sets that comprise concatenated sequence data of several genes, are becoming increasingly popular. Usually such multi-gene analyses are partitioned, i.e., a separate set of likelihood model parameters is estimated for each gene/partition. While the phylogenetic likelihood function exhibits intrinsic fine-grained parallelism, the parallel computation of the likelihood function in such partitioned multigene analyses can lead to significant load-balance problems. Here, we describe these problems for the first time, discuss the implications on the design of "classic" ML-based as well as Bayesian search algorithms, and provide an initial solution that yields up to eight-fold improvements in speedup values on AMD Barcelona and Sun x4600 16-core systems for realistic application scenarios.
Alexandros Stamatakis, Michael Ott 0001
ICPP2
2007 Large-scale maximum likelihood-based phylogenetic analysis on the IBM BlueGene/L
abstract
Phylogenetic inference is a grand challenge in Bioinformatics due to immense computational requirements. The increasing popularity of multi-gene alignments in biological studies, which typically provide a stable topological signal due to a more favorable ratio of the number of base pairs to the number of sequences, coupled with rapid accumulation of sequence data in general, poses new challenges for high performance computing. In this paper, we demonstrate how state-of-the-art Maximum Likelihood (ML) programs can be efficiently scaled to the IBM BlueGene/L (BG/L) architecture, by porting RAxML, which is currently among the fastest and most accurate programs for phylogenetic inference under the ML criterion. We simultaneously exploit coarse-grained and fine-grained parallelism that is inherent in every ML-based biological analysis. Performance is assessed using datasets consisting of 212 sequences and 566,470 base pairs, and 2,182 sequences and 51,089 base pairs, respectively. To the best of our knowledge, these are the largest datasets analyzed under ML to date. The capability to analyze such datasets will help to address novel biological questions via phylogenetic analyses. Our experimental results indicate that the fine-grained parallelization scales well up to 1, 024 processors. Moreover, a larger number of processors can be efficiently exploited by a combination of coarse-grained and fine-grained parallelism. Finally, we demonstrate that our parallelization scales equally well on an AMD Opteron cluster with a less favorable network latency to processor speed ratio. We recorded super-linear speedups in several cases due to increased cache efficiency.
Michael Ott 0001, Jaroslaw Zola, Alexandros Stamatakis, Srinivas Aluru
SC1
2005 DRAxML at home: a distributed program for computation of large phylogenetic trees
Alexandros Stamatakis, Markus Lindermeier, Michael Ott 0001, Thomas Ludwig 0002, Harald Meier
Future Gener. Comput. Syst.3