VLDB 2026 Research / reviewers in the wild / expert
Jeremy Enos
dblp:82/1723
· DBLP profile ↗
9ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0003-3100-6424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 54% High-performance computing · 32% Performance modeling and evaluation · 8% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › fault tolerance › failure diagnosis
failure localization |
0.4 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Distributed systems › fault tolerance
fault detection and diagnosis |
0.4 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Distributed systems
fault tolerance |
0.4 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
Performance modeling and evaluation › profiling
application profiling |
0.2 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
High-performance computing
system monitoring |
0.2 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
High-performance computing
supercomputing |
0.2 | 2 | 2018 | Best practices and lessons from deploying and operating a sustained-petascale system: the blue waters experience · SC 2018 Software tools II - Running a Top-500 benchmark on a windows compute cluster server cluster · SC 2006 |
Storage systems
distributed storage |
0.1 | 1 | 2020 | Live forensics for HPC systems: a case study on distributed storage systems · SC 2020 |
High-performance computing
cluster computing |
0.1 | 1 | 2006 | Software tools II - Running a Top-500 benchmark on a windows compute cluster server cluster · SC 2006 |
High-performance computing › numerical linear algebra
linpack |
0.1 | 1 | 2006 | Software tools II - Running a Top-500 benchmark on a windows compute cluster server cluster · SC 2006 |
Distributed systems › observability
large-scale monitoring |
0.1 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2006 | Software tools II - Running a Top-500 benchmark on a windows compute cluster server cluster · SC 2006 |
Cloud and datacenter computing
job submission |
0.0 | 1 | 2006 | Software tools II - Running a Top-500 benchmark on a windows compute cluster server cluster · SC 2006 |
Methods — techniques the papers use, named apart from their topics
machine learning · 0.4lightweight distributed monitoring · 0.2cluster provisioning · 0.1benchmark compilation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Live forensics for HPC systems: a case study on distributed storage systemsabstractLarge-scale high-performance computing systems frequently experience a wide range of failure modes, such as reliability failures (e.g., hang or crash), and resource overload-related failures (e.g., congestion collapse), impacting systems and applications. Despite the adverse effects of these failures, current systems do not provide methodologies for proactively detecting, localizing, and diagnosing failures. We present Kaleidoscope, a near real-time failure detection and diagnosis framework, consisting of of hierarchical domain-guided machine learning models that identify the failing components, the corresponding failure mode, and point to the most likely cause indicative of the failure in near real-time (within one minute of failure occurrence). Kaleidoscope has been deployed on Blue Waters supercomputer and evaluated with more than two years of production telemetry data. Our evaluation shows that Kaleidoscope successfully localized 99.3% and pinpointed the root causes of 95.8% of 843 real-world production issues, with less than 0.01% runtime overhead. Saurabh Jha, Shengkun Cui, Subho S. Banerjee, Tianyin Xu, Jeremy Enos, Michael T. Showerman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 5 |
| 2019 | Best practices for management and operation of large HPC installationsabstractSummary To achieve their mission and goals, HPC centers continually strive to improve the effectiveness of their resources and services to best serve their constituencies. Collectively, the community has learned a great deal about how to manage and operate HPC centers, provide robust and effective services, and develop new communities as well as about other important aspects. Yet, cataloguing best practices to help inform and guide the broader HPC community is not often done. To improve the situation, the Blue Waters project has documented sets of best practices that have been adopted for the deployment and operation over the past five years of the Blue Waters leadership system, a large Cray XE6/XK7 supercomputer at NCSA. Those practices, described in this paper, cover aspects of managing and operating the system and its resources, supporting its users, and expanding the diversity of applications and communities. Although the technical practices are sometimes discussed relative to Cray systems and leadership‐scale systems, we believe that they would benefit the deployment and operation of other large HPC installations as well. Scott A. Lathrop, Celso L. Mendes, Jeremy Enos, Brett M. Bode, Gregory H. Bauer, Robert Sisneros, William T. Kramer |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | Large-Scale System Monitoring Experiences and RecommendationsabstractMonitoring of High Performance Computing (HPC) platforms is critical to successful operations, can provide insights into performance-impacting conditions, and can inform methodologies for improving science throughput. However, monitoring systems are not generally considered core capabilities in system requirements specifications nor in vendor development strategies. In this paper we present work performed at a number of large-scale HPC sites towards developing monitoring capabilities that fill current gaps in ease of problem identification and root cause discovery. We also present our collective views, based on the experiences presented, on needs and requirements for enabling development by vendors or users of effective sharable end-to-end monitoring capabilities. Ville Ahlgren, Stefan Andersson, Jim M. Brandt, Nicholas Cardo, Sudheer Chunduri, Jeremy Enos, Parks Fields, Ann C. Gentile, Richard A. Gerber, Michael Gienger, Joe Greenseid, Annette Greiner, Bilel Hadri, Dennis Hoppe, Urpo Kaila, Kaki Kelly, Mark Klein 0002, Alex Kristiansen, Stephen Leak, Mike Mason, Kevin T. Pedretti, Jean-Guillaume Piccinali, Jason Repik, Jim Rogers, Susanna Salminen, Michael T. Showerman, Cary Whitney, Jim Williams |
CLUSTER | 6 |
| 2018 | Best practices and lessons from deploying and operating a sustained-petascale system: the blue waters experience
Gregory H. Bauer, Brett M. Bode, Jeremy Enos, William T. Kramer, Scott A. Lathrop, Celso L. Mendes, Robert Sisneros |
SC | 3 |
| 2017 | Holistic Measurement-Driven System AssessmentabstractIn high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer |
CLUSTER | 6 |
| 2017 | BOSS-LDG: A Novel Computational Framework That Brings Together Blue Waters, Open Science Grid, Shifter and the LIGO Data Grid to Accelerate Gravitational Wave DiscoveryabstractWe present a novel computational framework that connects Blue Waters, the NSF-supported, leadership-class supercomputer operated by NCSA, to the Laser Interferometer Gravitational-Wave Observatory (LIGO) Data Grid via Open Science Grid technology. To enable this computational infrastructure, we configured, for the first time, a LIGO Data Grid Tier-1 Center that can submit heterogeneous LIGO workflows using Open Science Grid facilities. In order to enable a seamless connection between the LIGO Data Grid and Blue Waters via Open Science Grid, we utilize Shifter to containerize LIGO's workflow software. This work represents the first time Open Science Grid, Shifter, and Blue Waters are unified to tackle a scientific problem and, in particular, it is the first time a framework of this nature is used in the context of large scale gravitational wave data analysis. This new framework has been used in the last several weeks of LIGO's second discovery campaign to run the most computationally demanding gravitational wave search workflows on Blue Waters, and accelerate discovery in the emergent field of gravitational wave astrophysics. We discuss the implications of this novel framework for a wider ecosystem of Higher Performance Computing users. Eliu A. Huerta, Roland Haas, Edgar Fajardo Hernandez, Daniel S. Katz, Peter Couvares, Josh Willis 0002, Timothy Bouvet, Jeremy Enos, William T. Kramer, Hon Wai Leong, David Wheeler |
eScience | 9 |
| 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and ApplicationsabstractUnderstanding how resources of High Performance Compute platforms are utilized by applications both individually and as a composite is key to application and platform performance. Typical system monitoring tools do not provide sufficient fidelity while application profiling tools do not capture the complex interplay between applications competing for shared resources. To gain new insights, monitoring tools must run continuously, system wide, at frequencies appropriate to the metrics of interest while having minimal impact on application performance. We introduce the Lightweight Distributed Metric Service for scalable, lightweight monitoring of large scale computing systems and applications. We describe issues and constraints guiding deployment in Sandia National Laboratories' capacity computing environment and on the National Center for Supercomputing Applications' Blue Waters platform including motivations, metrics of choice, and requirements relating to the scale and specialized nature of Blue Waters. We address monitoring overhead and impact on application performance and provide illustrative profiling results. Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Paul Cassella, Jeremy Enos, Joshi Fullop, Ann C. Gentile, Steve Monk, Nichamon Naksinehaboon, Jeff Ogden, Mahesh Rajan, Michael T. Showerman, Joel Stevenson, Narate Taerat, Thomas W. Tucker |
SC | 5 |
| 2009 | GPU clusters for high-performance computingabstractLarge-scale GPU clusters are gaining popularity in the scientific computing community. However, their deployment and production use are associated with a number of new challenges. In this paper, we present our efforts to address some of the challenges with building and running GPU clusters in HPC environments. We touch upon such issues as balanced cluster architecture, resource sharing in a cluster environment, programming models, and applications for GPU clusters. Volodymyr V. Kindratenko, Jeremy Enos, Guochun Shi, Michael T. Showerman, Galen Wesley Arnold, John E. Stone, James C. Phillips, Wen-Mei W. Hwu |
CLUSTER | 2 |
| 2006 | Software tools II - Running a Top-500 benchmark on a windows compute cluster server clusterabstractIn June 2006 Microsoft in conjunction with NCSA completed a Top 500 benchmark on a 900 processor Dell PowerEdge 1855 cluster running Windows CCS Version one. The result was a 4.1 Tflop Rmax number; placing this cluster at number 130 in the July 2006 Top 500 List. This was a significant accomplishment for an offering focused on a design point of 64 nodes.Attend this session to hear about our experiences compiling and running the HPC Linpack benchmark in a large-scale Windows environment with CCS. We will cover the design of a large Windows cluster, the tools used to provision the cluster, the tools used to compile the benchmark, the job submission process for the parallel execution of the program, and the actual results. We will also talk about what we learned from the process and how that information will lead to improvements for future versions of CCS. Frank Chism, Jeremy Enos |
SC | 2 |