EDBT 2026 Demo / reviewers in the wild / expert
Michael T. Showerman
dblp:50/7892 · also Mike Showerman
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-authorComputer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 73% High-performance computing · 10% Performance modeling and evaluation · 10% | |
| Computer networks
1 paper |
Datacenter networks · 50% Network measurement and analytics · 50% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Network measurement and analytics › network performance measurement
congestion measurement |
0.4 | 1 | 2020 | Measuring Congestion in High-Performance Datacenter Interconnects · NSDI 2020 |
Distributed systems › fault tolerance › failure diagnosis
failure localization |
0.4 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Distributed systems › fault tolerance
fault detection and diagnosis |
0.4 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Distributed systems
fault tolerance |
0.4 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Performance modeling and evaluation › profiling
application profiling |
0.2 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
High-performance computing
system monitoring |
0.2 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
Storage systems
distributed storage |
0.1 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Distributed systems › observability
large-scale monitoring |
0.1 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
Methods — techniques the papers use, named apart from their topics
machine learning · 0.4lightweight distributed monitoring · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Measuring Congestion in High-Performance Datacenter Interconnects
Saurabh Jha, Archit Patke, Jim M. Brandt, Ann C. Gentile, Benjamin Lim, Michael T. Showerman, Gregory H. Bauer, Larry Kaplan, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer |
NSDI | 6 |
| 2020 | Live forensics for HPC systems: a case study on distributed storage systemsabstractLarge-scale high-performance computing systems frequently experience a wide range of failure modes, such as reliability failures (e.g., hang or crash), and resource overload-related failures (e.g., congestion collapse), impacting systems and applications. Despite the adverse effects of these failures, current systems do not provide methodologies for proactively detecting, localizing, and diagnosing failures. We present Kaleidoscope, a near real-time failure detection and diagnosis framework, consisting of of hierarchical domain-guided machine learning models that identify the failing components, the corresponding failure mode, and point to the most likely cause indicative of the failure in near real-time (within one minute of failure occurrence). Kaleidoscope has been deployed on Blue Waters supercomputer and evaluated with more than two years of production telemetry data. Our evaluation shows that Kaleidoscope successfully localized 99.3% and pinpointed the root causes of 95.8% of 843 real-world production issues, with less than 0.01% runtime overhead. Saurabh Jha, Shengkun Cui, Subho S. Banerjee, Tianyin Xu, Jeremy Enos, Michael T. Showerman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 6 |
| 2018 | Large-Scale System Monitoring Experiences and RecommendationsabstractMonitoring of High Performance Computing (HPC) platforms is critical to successful operations, can provide insights into performance-impacting conditions, and can inform methodologies for improving science throughput. However, monitoring systems are not generally considered core capabilities in system requirements specifications nor in vendor development strategies. In this paper we present work performed at a number of large-scale HPC sites towards developing monitoring capabilities that fill current gaps in ease of problem identification and root cause discovery. We also present our collective views, based on the experiences presented, on needs and requirements for enabling development by vendors or users of effective sharable end-to-end monitoring capabilities. Ville Ahlgren, Stefan Andersson, Jim M. Brandt, Nicholas Cardo, Sudheer Chunduri, Jeremy Enos, Parks Fields, Ann C. Gentile, Richard A. Gerber, Michael Gienger, Joe Greenseid, Annette Greiner, Bilel Hadri, Dennis Hoppe, Urpo Kaila, Kaki Kelly, Mark Klein 0002, Alex Kristiansen, Stephen Leak, Mike Mason, Kevin T. Pedretti, Jean-Guillaume Piccinali, Jason Repik, Jim Rogers, Susanna Salminen, Michael T. Showerman, Cary Whitney, Jim Williams |
CLUSTER | 27 |
| 2017 | Holistic Measurement-Driven System AssessmentabstractIn high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer |
CLUSTER | 7 |
| 2015 | Real Time Visualization of Monitoring Data for Large Scale HPC SystemsabstractHigh Performance Computing (HPC) system users and administrators are often hampered in their ability understand application performance and system behavior due to a lack of sufficient information about how resources, such as memory, CPU, networks and filesystems are being used. While obtaining the related data is a necessary step, it is insufficient without tools that can turn the data into actionable information. Required capabilities of such tools are the ability to efficiently handle vast amounts of data in a timely fashion, the presentation of effective and understandable information representations for large node counts, and the correlation of that data with job and system events. This paper presents visualization approaches and tools that NCSA is developing, combined with the use of freely available web interfaces, to turn the eight billion platform related data points per day being collected from their 27,648 compute node Blue Waters platform into actionable information for both system administrators and users. Insights from the visualizations both at the system and the job levels are also presented. Michael T. Showerman |
CLUSTER | 1 |
| 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and ApplicationsabstractUnderstanding how resources of High Performance Compute platforms are utilized by applications both individually and as a composite is key to application and platform performance. Typical system monitoring tools do not provide sufficient fidelity while application profiling tools do not capture the complex interplay between applications competing for shared resources. To gain new insights, monitoring tools must run continuously, system wide, at frequencies appropriate to the metrics of interest while having minimal impact on application performance. We introduce the Lightweight Distributed Metric Service for scalable, lightweight monitoring of large scale computing systems and applications. We describe issues and constraints guiding deployment in Sandia National Laboratories' capacity computing environment and on the National Center for Supercomputing Applications' Blue Waters platform including motivations, metrics of choice, and requirements relating to the scale and specialized nature of Blue Waters. We address monitoring overhead and impact on application performance and provide illustrative profiling results. Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Paul Cassella, Jeremy Enos, Joshi Fullop, Ann C. Gentile, Steve Monk, Nichamon Naksinehaboon, Jeff Ogden, Mahesh Rajan, Michael T. Showerman, Joel Stevenson, Narate Taerat, Thomas W. Tucker |
SC | 12 |
| 2009 | GPU clusters for high-performance computingabstractLarge-scale GPU clusters are gaining popularity in the scientific computing community. However, their deployment and production use are associated with a number of new challenges. In this paper, we present our efforts to address some of the challenges with building and running GPU clusters in HPC environments. We touch upon such issues as balanced cluster architecture, resource sharing in a cluster environment, programming models, and applications for GPU clusters. Volodymyr V. Kindratenko, Jeremy Enos, Guochun Shi, Michael T. Showerman, Galen Wesley Arnold, John E. Stone, James C. Phillips, Wen-Mei W. Hwu |
CLUSTER | 4 |