EDBT 2026 Demo / reviewers in the wild / expert
Benjamin A. Allan
dblp:79/414
· DBLP profile ↗
13ranked-venue papers
4as first author
2since 2021 · last 2022
0009-0008-6180-2556ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 44% Performance modeling and evaluation · 44% Distributed systems · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › profiling
application profiling |
0.2 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
High-performance computing
system monitoring |
0.2 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
Distributed systems › observability
large-scale monitoring |
0.1 | 1 | 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications · SC 2014 |
Methods — techniques the papers use, named apart from their topics
lightweight distributed monitoring · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Metrics for Packing Efficiency and Fairness of HPC Cluster Batch Job SchedulingabstractDevelopment of job scheduling algorithms, which directly influence High-Performance Computing (HPC) clusters performance, is hindered because popular scheduling quality metrics, such as Bounded Slowdown, poorly correlate with global scheduling objectives that include job packing efficiency and fairness. This report proposes Area Weighted Response Time, a metric that offers an unbiased representation of job packing efficiency, and presents a class of new metrics, Priority Weighted Specific Response Time, that assess both packing efficiency and fairness of schedules. The provided examples of simulation of scheduling of real workload traces and analysis of the resulting schedules with the help of these metrics and conventional metrics, demonstrate that although Bounded Slowdown can be readily improved by modifying the standard First Come First Served backfilling algorithm and by using existing techniques of estimating job runtime, these improvements are accompanied by significant degradation of job packing efficiency and fairness. In contrast, improving job packing efficiency and fairness over the standard backfilling algorithm, which is designed to target those objectives, is difficult. It requires further algorithm development and more accurate runtime estimation techniques that reduce frequency of underpredictions. Alexander V. Goponenko, Kenneth Lamar, Christina L. Peterson, Benjamin A. Allan, Jim M. Brandt, Damian Dechev |
SBAC-PAD | 4 |
| 2021 | Backfilling HPC Jobs with a Multimodal-Aware PredictorabstractJob scheduling aims to minimize the turnaround time on the submitted jobs while catering to the resource constraints of High Performance Computing (HPC) systems. The challenge with scheduling is that it must honor job requirements and priorities while actual job run times are unknown. Although approaches have been proposed that use classification techniques or machine learning to predict job run times for scheduling purposes, these approaches do not provide a technique for reducing underprediction, which has a negative impact on scheduling quality. A common cause of underprediction is that the distribution of the duration for a job class is multimodal, causing the average job duration to fall below the expected duration of longer jobs. In this work, we propose the Top Percent predictor, which uses a hierarchical classification scheme to provide better accuracy for job run time predictions than the user-requested time. Our predictor addresses multimodal job distributions by making a prediction that is higher than a specified percentage of the observed job run times. We integrate the Top Percent predictor into scheduling algorithms and evaluate the performance using schedule quality metrics found in literature. To accommodate the user policies of HPC systems, we propose priority metrics that account for job flow time, job resource requirements, and job priority. The experiments demonstrate that the Top Percent predictor outperforms the related approaches when evaluated using our proposed priority metrics. Kenneth Lamar, Alexander V. Goponenko, Christina L. Peterson, Benjamin A. Allan, Jim M. Brandt, Damian Dechev |
CLUSTER | 4 |
| 2020 | LDMS Monitoring of EDR InfiniBand NetworksabstractWe introduce a new HPC system high-speed network fabric production monitoring tool, the ibnet sampler plugin for LDMS version 4. Large-scale testing of this tool is our work in progress. When deployed appropriately, the ibnet sampler plugin can provide extensive counter data, at frequencies up to 1 Hz. This allows the LDMS monitoring system to be useful for tracking the impact of new network features on production systems. We present preliminary results concerning reliability, performance impact, and usability of the sampler. Benjamin A. Allan, Michael J. Aguilar, Benjamin Schwaller, Steven Langer |
CLUSTER | 1 |
| 2020 | HPC System Data Pipeline to Enable Meaningful Insights through Analysis-Driven VisualizationsabstractThe increasing complexity of High Performance Computing (HPC) systems has created a growing need for facilitating insight into system performance and utilization for administrators and users. The strides made in HPC system monitoring data collection have produced terabyte/day sized time-series data sets rich with critical information, but it is onerous to extract and construe meaningful information from these metrics. We have designed and developed an architecture that enables flexible, as-needed, run-time analysis and presentation capabilities for HPC monitoring data. Our architecture enables quick and efficient data filtration and analysis. Complex runtime or historical analyses can be expressed as Python-based computations. Results of analyses and a variety of HPC oriented summaries are displayed in a Grafana front-end interface. To demonstrate our architecture, we have deployed it in production for a 1500-node HPC system and have developed analyses and visualizations requested by system administrators, and later employed by users, to track key metrics about the cluster at a job, user, and system level. Our architecture is generic, applicable to any *-nix based system, and it is extensible to supporting multi-cluster HPC centers. We structure it with easily replaced modules that allow unique customization across clusters and centers. In this paper, we describe the data collection and storage infrastructure, the application created to query and analyze data from a custom database, and the visual displays created to provide clear insights into HPC system behavior. Benjamin Schwaller, Nick Tucker, Tom Tucker, Benjamin A. Allan, Jim M. Brandt |
CLUSTER | 4 |
| 2019 | Standardized Environment for Monitoring Heterogeneous ArchitecturesabstractIncreasingly diverse architectures and operating systems continue to emerge in the HPC industry. As such, HPC centers are becoming more heterogeneous which introduces a variety of challenges for system administrators. Monitoring a wide array of different platforms by itself is difficult, but the problem compounds in an environment where new platforms are frequently added. Creating a standard monitoring environment across these platforms that allows for simple administration with minimal setup becomes necessary in such situations. This paper presents the solutions introduced in the HPC Development department at Sandia National Laboratories to meet these challenges. This includes our adoption of a multi-stage data-collection pipeline across our clusters that is implemented from the ground up with our Golden Image. We also discuss our infrastructure to support a heterogeneous environment and activities in progress to improve our center. These advances simplify system standup and make monitoring integration easier and faster for new systems which is necessary for our center's domain. Connor Brown, Benjamin Schwaller, Nathan Gauntt, Benjamin A. Allan, Kevin Davis |
CLUSTER | 4 |
| 2017 | Measuring Minimum Switch Port Metric Retrieval Time and Impact for Multi-layer InfiniBand FabricsabstractIn this work, we seek to gain an understanding of the InfiniBand network processing limitations that might exist in gathering performance metric information from InfiniBand switches using our new LDMS ibfabric sampler. The limitations studied consist of delays in gathering InfiniBand metric information from a specific switch device due to the switch's processor response delays or RDMA contention for network bandwidth. Michael Aguilar, Benjamin A. Allan, Sergei Polevitzky |
CLUSTER | 2 |
| 2016 | Continuous whole-system monitoring toward rapid understanding of production HPC applications and systems
Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Ann C. Gentile, Sophia Lefantzi, Steve Monk, Jeff Ogden, Mahesh Rajan, Joel Stevenson |
Parallel Comput. | 2 |
| 2015 | Toward Rapid Understanding of Production HPC Applications and SystemsabstractA detailed understanding of HPC application's resource needs and their complex interactions with each other and HPC platform resources is critical to achieving scalability and performance. Such understanding has been difficult to achieve because typical application profiling tools do not capture the behaviors of codes under the potentially wide spectrum of actual production conditions and because typical monitoring tools do not capture system resource usage information with high enough fidelity to gain sufficient insight into application performance and demands. In this paper we present both system and application profiling results based on data obtained through synchronized system wide monitoring on a production HPC cluster at Sandia National Laboratories (SNL). We demonstrate analytic and visualization techniques that we are using to characterize application and system resource usage under production conditions for better understanding of application resource needs. Our goals are to improve application performance (through understanding application-to-resource mapping and system throughput) and to ensure that future system capabilities match their intended workloads. Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Ann C. Gentile, Sophia Lefantzi, Steve Monk, Jeff Ogden, Mahesh Rajan, Joel Stevenson |
CLUSTER | 2 |
| 2014 | The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and ApplicationsabstractUnderstanding how resources of High Performance Compute platforms are utilized by applications both individually and as a composite is key to application and platform performance. Typical system monitoring tools do not provide sufficient fidelity while application profiling tools do not capture the complex interplay between applications competing for shared resources. To gain new insights, monitoring tools must run continuously, system wide, at frequencies appropriate to the metrics of interest while having minimal impact on application performance. We introduce the Lightweight Distributed Metric Service for scalable, lightweight monitoring of large scale computing systems and applications. We describe issues and constraints guiding deployment in Sandia National Laboratories' capacity computing environment and on the National Center for Supercomputing Applications' Blue Waters platform including motivations, metrics of choice, and requirements relating to the scale and specialized nature of Blue Waters. We address monitoring overhead and impact on application performance and provide illustrative profiling results. Anthony M. Agelastos, Benjamin A. Allan, Jim M. Brandt, Paul Cassella, Jeremy Enos, Joshi Fullop, Ann C. Gentile, Steve Monk, Nichamon Naksinehaboon, Jeff Ogden, Mahesh Rajan, Michael T. Showerman, Joel Stevenson, Narate Taerat, Thomas W. Tucker |
SC | 2 |
| 2009 | On optimizing I/O through InfiniBand RDMA for commodity clustersabstractOur goal was to enable pNFS as a highperformance parallel file system by using network file system (NFS) storage objects and InfiniBand remote direct memory access (RDMA) transport in the Linux mainline. One obstacle to this approach was the performance bottleneck in NFS/RDMA streaming-writes from the compute nodes. We benchmarked, tuned, and improved the streaming-write efficiency of the Linux NFS client. However, deeper analyses of the benchmarks and the various I/O short-circuit schemes established upper bounds on the performance of the NFS client — even with an infinitely fast network — so that the performance was substantially less than the theoretical streaming bandwidth of the fast interconnections. The complex interactions between the Linux virtual file system, Linux virtual memory management, and the IB network subsystems apparently impose a limit on further improvement. Benjamin A. Allan, Helen Chen, Scott Cranford, Ron Minnich, Don W. Rudish, Lee Ward |
CLUSTER | 1 |
| 2006 | The CCA component model for high-performance scientific computingabstractAbstract The Common Component Architecture (CCA) is a component model for high‐performance computing, developed by a grass‐roots effort of computational scientists. Although the CCA is usable with CORBA‐like distributed‐object components, its main purpose is to set forth a component model for high‐performance, parallel computing. Traditional component models are not well suited for performance and massive parallelism. We outline the design pattern for the CCA component model, discuss our strategy for language interoperability, describe the development tools we provide, and walk through an illustrative example using these tools. Performance and scalability, which are distinguishing features of CCA components, affect choices throughout design and implementation. Copyright © 2005 John Wiley & Sons, Ltd. Robert C. Armstrong, Gary Kumfert, Lois C. McInnes, Steven G. Parker, Benjamin A. Allan, Matthew J. Sottile, Thomas Epperly, Tamara Dahlgren |
Concurr. Comput. Pract. Exp. | 5 |
| 2004 | ODEPACK++: Refactoring the LSODE Fortran Library for Use in the CCA High Performance Component Software ArchitectureabstractWe present a case study of the alternatives and design trade-offs encountered when adapting an established numerical library into a form compatible with modern component-software implementation practices. Our study will help scientific software users, authors, and maintainers develop their own roadmaps for shifting to component-oriented software. The primary library studied, LSODE, and the issues involved in the adaptation are typical of many commonly used numerical libraries. We examine the adaptation of a related library, CVODE, and compare the impact on applications of the two different designs. The LSODE-derived components solve models composed with CCA components developed independently at the Argonne and Oak Ridge National Laboratories. The resulting applications run in the Ccaffeine framework implementation of the common component architecture specification. We provide CCA-style interface specifications appropriate to linear equations, ordinary differential equations (ODE), and differential algebraic equations (DAE) solvers. Benjamin A. Allan, Sophia Lefantzi, Jaideep Ray |
HIPS | 1 |
| 2002 | The CCA core specification in a distributed memory SPMD frameworkabstractAbstract We present an overview of the Common Component Architecture (CCA) core specification and CCAFFEINE, a Sandia National Laboratories framework implementation compliant with the draft specification. CCAFFEINE stands for CCA Fast Framework Example In Need of Everything; that is, CCAFFEINE is fast, lightweight, and it aims to provide every framework service by using external, portable components instead of integrating all services into a single, heavy framework core. By fast, we mean that the CCAFFEINE glue does not get between components in a way that slows down their interactions. We present the CCAFFEINE solutions to several fundamental problems in the application of component software approaches to the construction of single program multiple data (SPMD) applications. We demonstrate the integration of components from three organizations, two within Sandia and one at Oak Ridge National Laboratory. We outline some requirements for key enabling facilities needed for a successful component approach to SPMD application building. Copyright © 2002 John Wiley & Sons, Ltd. Benjamin A. Allan, Robert C. Armstrong, Alicia P. Wolfe, Jaideep Ray, David E. Bernholdt, James Arthur Kohl |
Concurr. Comput. Pract. Exp. | 1 |