EDBT 2026 Demo / reviewers in the wild / expert
Adam Moody
dblp:189/1241 · also Adam T. Moody
· DBLP profile ↗
38ranked-venue papers
3as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Distributed systems · 37% Storage systems · 29% High-performance computing · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems › distributed file system
parallel file system |
1.0 | 6 | 2016 | An ephemeral burst-buffer file system for scientific applications · SC 2016 Detailed Modeling and Evaluation of a Scalable Multilevel Checkpointing System · IEEE Trans. Parallel Distributed Syst. 2014 A 1 PB/s file system to checkpoint three million MPI tasks · HPDC 2013 |
Distributed systems › fault tolerance
checkpointing |
0.8 | 5 | 2014 | Detailed Modeling and Evaluation of a Scalable Multilevel Checkpointing System · IEEE Trans. Parallel Distributed Syst. 2014 A 1 PB/s file system to checkpoint three million MPI tasks · HPDC 2013 Design and modeling of a non-blocking checkpointing system · SC 2012 |
Distributed systems
fault tolerance |
0.8 | 5 | 2014 | Detailed Modeling and Evaluation of a Scalable Multilevel Checkpointing System · IEEE Trans. Parallel Distributed Syst. 2014 A 1 PB/s file system to checkpoint three million MPI tasks · HPDC 2013 Design and modeling of a non-blocking checkpointing system · SC 2012 |
High-performance computing › supercomputing
supercomputer deployment |
0.3 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
Distributed systems › fault tolerance › checkpointing
multi-level checkpointing |
0.3 | 2 | 2014 | Detailed Modeling and Evaluation of a Scalable Multilevel Checkpointing System · IEEE Trans. Parallel Distributed Syst. 2014 Design, Modeling, and Evaluation of a Scalable Multi-level Checkpointing System · SC 2010 |
Storage systems
checkpoint storage |
0.3 | 2 | 2012 | Design and modeling of a non-blocking checkpointing system · SC 2012 Design, Modeling, and Evaluation of a Scalable Multi-level Checkpointing System · SC 2010 |
Emerging computing paradigms › neuromorphic computing
brain-inspired computing |
0.2 | 1 | 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016 |
Emerging computing paradigms
neuromorphic computing |
0.2 | 1 | 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016 |
Emerging computing paradigms
neuromorphic hardware |
0.2 | 1 | 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016 |
High-performance computing
scientific data management |
0.2 | 1 | 2016 | An ephemeral burst-buffer file system for scientific applications · SC 2016 |
Storage systems
data compression |
0.1 | 1 | 2012 | McrEngine: a scalable checkpointing system using data-aware aggregation and compression · SC 2012 |
Parallel and multicore computing › parallel programming models › message passing
MPI runtime |
0.1 | 1 | 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes · SC 2012 |
Distributed systems › fault tolerance › checkpointing
nonblocking checkpointing |
0.1 | 1 | 2012 | Design and modeling of a non-blocking checkpointing system · SC 2012 |
Storage systems
storage reliability |
0.1 | 1 | 2012 | McrEngine: a scalable checkpointing system using data-aware aggregation and compression · SC 2012 |
Distributed systems
topology discovery |
0.1 | 1 | 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes · SC 2012 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
Software maintenance and evolution › software ecosystems
dependency management |
0.1 | 1 | 2015 | The Spack package manager: bringing order to HPC software chaos · SC 2015 |
Parallel and multicore computing
MPI |
0.0 | 1 | 2013 | A 1 PB/s file system to checkpoint three million MPI tasks · HPDC 2013 |
Distributed systems › communication optimization
topology-aware communication |
0.0 | 1 | 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes · SC 2012 |
High-performance computing
collective communication |
0.0 | 1 | 2003 | Scalable NIC-based Reduction on Large-scale Clusters · SC 2003 |
Storage systems
data reduction |
0.0 | 1 | 2003 | Scalable NIC-based Reduction on Large-scale Clusters · SC 2003 |
Interconnection networks and networks-on-chip › network interface
network interface card |
0.0 | 1 | 2003 | Scalable NIC-based Reduction on Large-scale Clusters · SC 2003 |
Methods — techniques the papers use, named apart from their topics
probabilistic modeling · 0.3markov model · 0.3software ecosystem · 0.2server-side read clustering · 0.2scalable systems · 0.2pipelining · 0.2co-located i/o delegation · 0.2multi-level checkpointing · 0.1data-aware aggregation · 0.1compression · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | UnifyFS: A User-level Shared File System for Unified Access to Distributed Local StorageabstractWe introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration. Michael J. Brim, Adam Moody, Seung-Hwan Lim, Ross G. Miller, Swen Böhm, Cameron Stanavige, Kathryn Mohror, Sarp Oral |
IPDPS | 2 |
| 2022 | DFMan: A Graph-based Optimization of Dataflow Scheduling on High-Performance Computing SystemsabstractScientific research and development campaigns are materialized by workflows of applications executing on high-performance computing (HPC) systems. These applications con-sist of tasks that can have inter- or intra-application flows of data to achieve the research goals successfully. These dataflows create dependencies among the tasks and cause resource con-tention on shared storage systems, thus limiting the aggregated I/O bandwidth achieved by the workflow. However, these I/O performance issues are often solved by tedious and manual efforts that demand holistic knowledge about the data dependencies in the workflow and the information about the infrastructure being utilized. Taking this into consideration, we design DFMan, a graph-based dataflow management and optimization framework for maximizing I/O bandwidth by leveraging the powerful storage stack on HPC systems to manage data sharing optimally among the tasks in the workflows. In particular, we devise a graph-based optimization algorithm that can leverage an intuitive graph representation of dataflow- and system-related information, and automatically carry out co-scheduling of task and data placement. According to our experiments, DFMan optimizes a wide variety of scientific workflows such as Hurricane 3D on Cloud Model 1 (CM1), Montage Carina Nebula (NGC3372), and an emulated dataflow kernel of the Multiscale Machine-learned Modeling Infrastructure (MuMMI I/O) on the Lassen supercomputer, and improves their aggregated I/O bandwidth by up to 5.42 x, 2.12 x and 1.29 x, respectively, compared to the baseline bandwidth. Fahim Chowdhury, Francesco Di Natale, Adam Moody, Kathryn Mohror, Weikuan Yu |
IPDPS | 3 |
| 2020 | Understanding HPC Application I/O Behavior Using System Level StatisticsabstractThe processor performance of high performance computing (HPC) systems is increasing at a much higher rate than storage performance. This imbalance leads to I/O performance bottlenecks in massively parallel HPC applications. Therefore, there is a need for improvements in storage and file system designs to meet the ever-growing I/O needs of HPC applications. Storage and file system designers require a deep understanding of how HPC application I/O behavior affects current storage system installations in order to improve them. In this work, we contribute to this understanding using application-agnostic file system statistics gathered on compute nodes as well as metadata and object storage file system servers. We analyze file system statistics of more than 4 million jobs over a period of three years on two systems at Lawrence Livermore National Laboratory that include a 15 PiB Lustre file system for storage. The results of our study add to the state-of-the-art in I/O understanding by providing insight into how general HPC workloads affect the performance of large-scale storage systems. Some key observations in our study show that reads and writes are evenly distributed across the storage system; applications which perform I/O, spread that I/O across ~78% of the minutes of their runtime on average; less than 22% of HPC users who submit write-intensive jobs perform efficient writes to the file system; and I/O contention seriously impacts I/O performance. Arnab Kumar Paul, Olaf Faaland, Adam Moody, Elsa Gonsiorowski, Kathryn Mohror, Ali Raza Butt |
HiPC | 3 |
| 2019 | Efficient User-Level Storage Disaggregation for Deep LearningabstractOn large-scale high performance computing (HPC) systems, applications are provisioned with aggregated resources to meet their peak demands for brief periods. This results in resource underutilization because application requirements vary a lot during execution. This problem is particularly pronounced for deep learning applications that are running on leadership HPC systems with a large pool of burst buffers in the form of flash or non-volatile memory (NVM) devices. In this paper, we examine the I/O patterns of deep neural networks and reveal their critical need of loading many small samples randomly for successful training. We have designed a specialized Deep Learning File System (DLFS) that provides a thin set of APIs. Particularly, we design the metadata management of DLFS through an in-memory tree-based sample directory and its file services through the user-level SPDK protocol that can disaggregate the capabilities of NVM Express (NVMe) devices to parallel training tasks. Our experimental results show that DLFS can dramatically improve the throughput of training for deep neural networks on NVMe over Fabric, compared with the kernel-based Ext4 file system. Furthermore, DLFS achieves efficient user-level storage disaggregation with very little CPU utilization. Yue Zhu 0002, Weikuan Yu, Bing Jiao, Kathryn Mohror, Adam Moody, Fahim Chowdhury |
CLUSTER | 5 |
| 2019 | A Novel Gesomin Detection Method Based on Microwave SpectroscopyabstractGeosmin contamination in water is a leading cause of odor related complaints to water companies in UK, tainting water with an earthy smell that is detectable by humans in quantities as low as 4 nanograms per liter. Current Geosmin detection methods depend on lab-based equipment, requiring samples to be collected and transported before Geosmin can be tested. This research presents a novel method for the detection of Geosmin in water using Microwave spectroscopy capable of detecting differentiating between four levels of Geosmin contamination: 5 ng/L, 10 ng/L, 0.5 mg/L and 1 mg/L as well as control samples. Frequencies within the 5.4 GHz to 5.9, 6.4 GHz to 6.5 GHz and 7.2 GHz to 7.5 GHz ranges showed significant separation between the sample classes. Samuel Ryecroft, Andy Shaw, Paul Fergus, Patryk Kot, Khalid Hashim, Laura Conway, Adam Moody |
DeSE | 7 |
| 2019 | I/O Characterization and Performance Evaluation of BeeGFS for Deep LearningabstractParallel File Systems (PFSs) are frequently deployed on leadership High Performance Computing (HPC) systems to ensure efficient I/O, persistent storage and scalable performance. Emerging Deep Learning (DL) applications incur new I/O and storage requirements to HPC systems with batched input of small random files. This mandates PFSs to have commensurate features that can meet the needs of DL applications. BeeGFS is a recently emerging PFS that has grabbed the attention of the research and industry world because of its performance, scalability and ease of use. While emphasizing a systematic performance analysis of BeeGFS, in this paper, we present the architectural and system features of BeeGFS, and perform an experimental evaluation using cutting-edge I/O, Metadata and DL application benchmarks. Particularly, we have utilized AlexNet and ResNet-50 models for the classification of ImageNet dataset using the Livermore Big Artificial Neural Network Toolkit (LBANN), and ImageNet data reader pipeline atop TensorFlow and Horovod. Through extensive performance characterization of BeeGFS, our study provides a useful documentation on how to leverage BeeGFS for the emerging DL applications. Fahim Chowdhury, Yue Zhu 0002, Todd Heer, Saul Paredes, Adam Moody, Robin Goldstone, Kathryn Mohror, Weikuan Yu |
ICPP | 5 |
| 2019 | VeloC: Towards High Performance Adaptive Asynchronous Checkpointing at Large ScaleabstractGlobal checkpointing to external storage (e.g., a parallel file system) is a common I/O pattern of many HPC applications. However, given the limited I/O throughput of external storage, global checkpointing can often lead to I/O bottlenecks. To address this issue, a shift from synchronous checkpointing (i.e., blocking until writes have finished) to asynchronous checkpointing (i.e., writing to faster local storage and flushing to external storage in the background) is increasingly being adopted. However, with rising core count per node and heterogeneity of both local and external storage, it is non-trivial to design efficient asynchronous checkpointing mechanisms due to the complex interplay between high concurrency and I/O performance variability at both the node-local and global levels. This problem is not well understood but highly important for modern supercomputing infrastructures. This paper proposes a versatile asynchronous checkpointing solution that addresses this problem. To this end, we introduce a concurrency-optimized technique that combines performance modeling with lightweight monitoring to make informed decisions about what local storage devices to use in order to dynamically adapt to background flushes and reduce the checkpointing overhead. We illustrate this technique using the VeloC prototype. Extensive experiments on a pre-Exascale supercomputing system show significant benefits. Bogdan Nicolae, Adam Moody, Elsa Gonsiorowski, Kathryn Mohror, Franck Cappello |
IPDPS | 2 |
| 2018 | Requirements of an Underwater Sensor-Networking Platform for Environmental MonitoringabstractWireless sensor networks have thrived over recent years, with conventional wireless sensor networks becoming an ever-growing area of research. However, underwater wireless sensor networks are underdeveloped in comparison with wireless sensor networks. This paper looks at several important considerations for wireless sensor networks and underwater sensor networking platform targeted at environmental monitoring. Many requirements are transferable from conventional wireless sensor networks and this paper looks at how these challenges have been tackled. However, considerations such as energy efficiency hold a greater weight where sensor nodes are more difficult to retrieve, consideration must also be given to other factors such as the impact of communication technologies, communication distances, data rates and ease of deployment. Samuel Ryecroft, Andy Shaw, Paul Fergus, Patryk Kot, Magomed Muradov, Adam Moody, Laura Conroy |
DeSE | 6 |
| 2018 | PRIONN: Predicting Runtime and IO using Neural NetworksabstractFor job allocation decision, current batch schedulers have access to and use only information on the number of nodes and runtime because it is readily available at submission time from user job scripts. User-provided runtimes are typically inaccurate because users overestimate or lack understanding of job resource requirements. Beyond the number of nodes and runtime, other system resources, including IO and network, are not available but play a key role in system performance. There is the need for automatic, general, and scalable tools that provide accurate resource usage information to schedulers so that, by becoming resource-aware, they can better manage system resources. Michael R. Wyatt II, Stephen Herbein, Todd Gamblin, Adam Moody, Dong H. Ahn, Michela Taufer |
ICPP | 4 |
| 2018 | Effective Quantization Approaches for Recurrent Neural NetworksabstractDeep learning, Recurrent Neural Networks (RNN) in particular have shown superior accuracy in a large variety of tasks including machine translation, language understanding, and movie frames generation. However, these deep learning approaches are very expensive in terms of computation. In most cases, Graphic Processing Units (GPUs) are in used for large scale implementations. Meanwhile, energy efficient RNN approaches are proposed for deploying solutions on special purpose hardware including Field Programming Gate Arrays (FPGAs) and mobile platforms. In this paper, we propose an effective quantization approach for Recurrent Neural Networks (RNN) techniques including Long Short-Term Memory (LSTM), Gated Recurrent Units (GRU), and Convolutional Long Short Term Memory (ConvLSTM). We have implemented different quantization methods including Binary Connect {-1, 1}, Ternary Connect {-1, 0, 1}, and Quaternary Connect {-1, -0.5, 0.5, 1}. These proposed approaches are evaluated on different datasets for sentiment analysis on IMDB and video frame predictions on the moving MNIST dataset. The experimental results are compared against the full precision versions of the LSTM, GRU, and ConvLSTM. They show promising results for both sentiment analysis and video frame prediction. Md. Zahangir Alom, Adam Moody, Naoya Maruyama, Brian Van Essen, Tarek M. Taha |
IJCNN | 2 |
| 2018 | Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC SystemsabstractDeep neural networks have recently gained tremendous interest due to their capabilities in a wide variety of application areas such as computer vision and speech recognition. Thus it is important to exploit the unprecedented power of leadership High-Performance Computing (HPC) systems for greater potential of deep learning. While much attention has been paid to leverage the latest processors and accelerators, I/O support also needs to keep up with the growth of computing power for deep neural networks. In this research, we introduce an entropy-aware I/O framework called DeepIO for large-scale deep learning on HPC systems. Its overarching goal is to coordinate the use of memory, communication, and I/O resources for efficient training of datasets. DeepIO features an I/O pipeline that utilizes several novel optimizations: RDMA (Remote Direct Memory Access)-assisted in-situ shuffling, input pipelining, and entropy-aware opportunistic ordering. In addition, we design a portable storage interface to support efficient I/O on any underlying storage system. We have implemented DeepIO as a prototype for the popular TensorFlow framework and evaluated it on a variety of different storage systems. Our evaluation shows that DeepIO delivers significantly better performance than existing memory-based storage systems. Yue Zhu 0002, Fahim Chowdhury, Huansong Fu, Adam Moody, Kathryn Mohror, Kento Sato, Weikuan Yu |
MASCOTS | 4 |
| 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin |
SC | 31 |
| 2017 | A spike-based long short-term memory on a neurosynaptic processorabstractLow-power brain-inspired hardware systems have gained significant traction in recent years. They offer high energy efficiency and massive parallelism due to the distributed and asynchronous nature of neural computation through low-energy spikes. One such platform is the IBM TrueNorth Neurosynaptic System. Recently TrueNorth compatible representation learning algorithms have emerged, achieving close to state-of-the-art performance in various datasets. An exception is its application in temporal sequence processing models such as recurrent neural networks (RNNs), which is still at the proof of concept level. This is partly due to the hardware constraints in connectivity and syn-aptic weight resolution, and the inherent difficulty in capturing temporal dynamics of an RNN using spiking neurons. This work presents a design flow that overcomes the aforementioned difficulties and maps a special case of recurrent networks called Long Short-Term Memory (LSTM) onto a spike-based platform. The framework is built on top of various approximation techniques, weight and activation discretization, spiking neuron sub-circuits that implements the complex gating mechanisms and a store-and-release technique to enable neuron synchronization and faithful storage. While many of the techniques can be applied to map LSTM to any SNN simulator/emulator, here we demonstrate this approach on the TrueNorth chip adhering to its constraints. Two benchmark LSTM applications, parity check and Extended Reber Grammar, are evaluated and their accuracy, energy and speed tradeoffs are analyzed. Amar Shrestha, Khadeer Ahmed, Yanzhi Wang 0001, David P. Widemann, Adam Moody, Brian Van Essen, Qinru Qiu |
ICCAD | 5 |
| 2017 | Convolutional sparse coding on neurosynaptic cognitive systemabstractImage features can be learned and subsequently used for reconstruction and classification tasks in the fields of machine learning and computer vision. In this work, we propose image reconstruction using Convolutional Sparse Coding (CSC) on IBM's TrueNorth Neuromorphic computing system. CSC explicitly models local interactions through the convolution operations. Convolutional kernels define a dictionary and Sparse Feature Maps (SFMs) that are generated through a training process. The images are reconstructed with convolutional operations on SFMs and respective kernels. In this paper, we report on experimental results demonstrating promising sparse reconstructions on the IBM Neuromorphic TrueNorth hardware for two different benchmarks: MNIST and CIFAR-10. It is noted that this is the first ever important step towards convolutional sparse coding on neuromorphic hardware. Md. Zahangir Alom, Brian Van Essen, Adam Moody, David P. Widemann, Tarek M. Taha |
IJCNN | 3 |
| 2017 | Quadratic Unconstrained Binary Optimization (QUBO) on neuromorphic computing systemabstractThe problems of Artificial intelligence (AI) naturally maps to NP-hard optimization problems. This trend has significance to achieve human-level computation capability from machines. This computational ability can be achieved by developing evolutionary algorithms or mapping those evolutionary algorithms onto new generation computing systems: Quantum or Neuromorphic hardware. In this paper, we implemented the NP-hard optimization problem called Quadratic Unconstrained Binary Optimization (QUBO) problem for the solution of graph problems on the IBM's Neurosynaptic TrueNorth System. We have experimented on different types of graph problems with different levels of complexities and achieved encouraging results on IBM's Neuromorphic TrueNorth chip. Moreover, there are a set of potential applications have been discussed based on this proposed QUBO solution. Along with the QUBO on quantum annealing, it is the important first step towards the solutions of QUBO on Neuromorphic computing systems. Md. Zahangir Alom, Brian Van Essen, Adam Moody, David P. Widemann, Tarek M. Taha |
IJCNN | 3 |
| 2017 | MetaKV: A Key-Value Store for Metadata Management of Distributed Burst BuffersabstractDistributed burst buffers are a promising storage architecture for handling I/O workloads for exascale computing. Their aggregate storage bandwidth grows linearly with system node count. However, although scientific applications can achieve scalable write bandwidth by having each process write to its node-local burst buffer, metadata challenges remain formidable, especially for files shared across many processes. This is due to the need to track and organize file segments across the distributed burst buffers in a global index. Because this global index can be accessed concurrently by thousands or more processes in a scientific application, the scalability of metadata management is a severe performance-limiting factor. In this paper, we propose MetaKV: a key-value store that provides fast and scalable metadata management for HPC metadata workloads on distributed burst buffers. MetaKV complements the functionality of an existing key-value store with specialized metadata services that efficiently handle bursty and concurrent metadata workloads: compressed storage management, supervised block clustering, and log-ring based collective message reduction. Our experiments demonstrate that MetaKV outperforms the state-of-the-art key-value stores by a significant margin. It improves put and get metadata operations by as much as 2.66× and 6.29×, respectively, and the benefits of MetaKV increase with increasing metadata workload demand. Teng Wang 0001, Adam Moody, Yue Zhu 0002, Kathryn Mohror, Kento Sato, Tanzima Z. Islam, Weikuan Yu |
IPDPS | 2 |
| 2016 | Machine Learning Predictions of Runtime and IO Traffic on High-End ClustersabstractWe use supervised machine learning algorithms (i.e., Decision Trees, Random Forest, and K-nearest Neighbors) to predict performance characteristics such as runtime and IO traffic of batch jobs on high-end clusters, using only user job scripts as input. We show that decision trees outperform other algorithms and accurately predict the runtime of 73% of jobs within a error tolerance of 10 minutes, which is a 51% improvement over the user requested runtime. Ryan McKenna, Stephen Herbein, Adam Moody, Todd Gamblin, Michela Taufer |
CLUSTER | 3 |
| 2016 | Managing I/O Interference in a Shared Burst Buffer SystemabstractIn this work, we investigate the problem of inter-application interference in a shared Burst Buffer (BB) system. A BB is a new storage technology for HPC architectures that acts as an intermediate layer between performance-hungry HPC applications and the slow parallel file system. While the BB is meant to alleviate the problem of slow I/O in HPC systems, it is itself prone to performance degradation under interference. We observe that the magnitude of interference effects can reach a level that matters to the HPC system and the jobs that run on it. We investigate I/O scheduling techniques as a mechanism to mitigate BB I/O interference. With our results, we show that scheduling techniques tuned to BBs can control interference and significant performance benefits can be achieved. Sagar Thapaliya, Purushotham V. Bangalore, Jay F. Lofstead, Kathryn Mohror, Adam Moody |
ICPP | 5 |
| 2016 | System Noise Revisited: Enabling Application Scalability and Reproducibility with SMTabstractDespite significant advances in reducing system noise, the scalability and performance of scientific applications running on production commodity clusters today continue to suffer from the effects of noise. Unlike custom and expensive leadership systems, the Linux ecosystem provides a rich set of services that application developers utilize to increase productivity and to ease porting. The cost is the overhead that these services impose on a running application, negatively impacting its scalability and performance reproducibility. In this work, we propose and evaluate a simple yet effective way to isolate an application from system processes by leveraging Simultaneous Multi-Threading (SMT), a pervasive architectural feature on current systems. Our method requires no changes to the operating system or to the application. We quantify its effectiveness on a diverse set of scientific applications of interest to the U. S. Department of Energy showing performance improvements of up to 2.4 times at 16,384 tasks for a high-order finite elements shock hydrodynamics application. Finally, we provide guidance to system and application developers on how to best leverage SMT under different application characteristics and scales. Edgar A. León, Ian Karlin, Adam Moody |
IPDPS | 3 |
| 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applicationsabstractAbstract not provided Jun Sawada, Filipp Akopyan, Andrew S. Cassidy, Brian Taba, Michael DeBole, Pallab Datta, Rodrigo Alvarez-Icaza, Arnon Amir, John V. Arthur, Alexander Andreopoulos, Rathinakumar Appuswamy, Heinz Baier, Davis Barch, David J. Berg, Carmelo di Nolfo, Steven K. Esser, Myron Flickner, Thomas A. Horvath, Bryan L. Jackson, Jeffrey A. Kusnitz, Scott Lekuch, Michael Mastro, Timothy Melano, Paul Merolla, Steven E. Millman, Tapan K. Nayak, Norm Pass, Hartmut Penner, William P. Risk, Kai Schleupen, Ben Shaw 0001, Hayley Wu, Brian Giera, Adam Moody, T. Nathan Mundhenk, Brian Van Essen, Eric X. Wang, David P. Widemann, William E. Murphy, Jamie K. Infantolino, James A. Ross, Dale R. Shires, Manuel M. Vindiola, Raju Namburu, Dharmendra S. Modha |
SC | 34 |
| 2016 | An ephemeral burst-buffer file system for scientific applicationsabstractBurst buffers are becoming an indispensable hardware resource on large-scale supercomputers to buffer the bursty I/O from scientific applications. However, there is a lack of software support for burst buffers to be efficiently shared by applications within a batch-submitted job and recycled across different batch jobs. In addition, burst buffers need to cope with a variety of challenging I/O patterns from data-intensive scientific applications. In this study, we have designed an ephemeral Burst Buffer File System (BurstFS) that supports scalable and efficient aggregation of I/O bandwidth from burst buffers while having the same life cycle as a batch-submitted job. BurstFS features several techniques including scalable metadata indexing, co-located I/O delegation, and server-side read clustering and pipelining. Through extensive tuning and analysis, we have validated that BurstFS has accomplished our design objectives, with linear scalability in terms of aggregated I/O bandwidth for parallel writes and reads. Teng Wang 0001, Kathryn Mohror, Adam Moody, Kento Sato, Weikuan Yu |
SC | 3 |
| 2015 | Non-Blocking PMI Extensions for Fast MPI StartupabstractAn efficient implementation of the Process Management Interface (PMI) is crucial to enable fast start-up of MPI jobs. We propose three extensions to the PMI specification: 1) a blocking all gather collective (PMIX_Allgather), 2) a non-blocking all gather collective (PMIX_Iallgather), and 3) a non-blocking fence (PMIX_KVS_Ifence). We design and evaluate several PMI implementations to demonstrate how such extensions reduce MPI start-up cost. In particular, when sufficient work can be overlapped, these extensions allow for a constant initialization cost of MPI jobs at different core counts. At 16,384 cores, the designs lead to a speedup of 2.88 times over the state-of-the-art start-up schemes. Sourav Chakraborty 0003, Hari Subramoni, Adam Moody, Akshay Venkatesh, Jonathan L. Perkins, Dhabaleswar K. Panda 0001 |
CCGRID | 3 |
| 2015 | The Role of Container Technology in Reproducible Computer Systems ResearchabstractEvaluating experimental results in the field of computer systems is a challenging task, mainly due to the many changes in software and hardware that computational environments go through. In this position paper, we analyze salient features of container technology that, if leveraged correctly, can help reduce the complexity of reproducing experiments in systems research. We present a use case in the area of distributed storage systems to illustrate the extensions that we envision, mainly in terms of container management infrastructure. We also discuss the benefits and limitations of using containers as a way of reproducing research in other areas of experimental systems research. Ivo Jimenez, Carlos Maltzahn, Adam Moody, Kathryn Mohror, Jay F. Lofstead, Remzi H. Arpaci-Dusseau, Andrea C. Arpaci-Dusseau |
IC2E | 3 |
| 2015 | The Spack package manager: bringing order to HPC software chaosabstractLarge HPC centers spend considerable time supporting software for thousands of users, but the complexity of HPC software is quickly outpacing the capabilities of existing software management tools. Scientific applications require specific versions of compilers, MPI, and other dependency libraries, so using a single, standard software stack is infeasible. However, managing many configurations is difficult because the configuration space is combinatorial in size. Todd Gamblin, Matthew P. LeGendre, Michael R. Collette, Gregory L. Lee, Adam Moody, Bronis R. de Supinski, Scott Futral |
SC | 5 |
| 2014 | A User-Level InfiniBand-Based File System and Checkpoint Strategy for Burst BuffersabstractCheckpoint/Restart is an indispensable fault tolerance technique commonly used by high-performance computing applications that run continuously for hours or days at a time. However, even with state-of-the-art checkpoint/restart techniques, high failure rates at large scale will limit application efficiency. To alleviate the problem, we consider using burst buffers. Burst buffers are dedicated storage resources positioned between the compute nodes and the parallel file system, and this new tier within the storage hierarchy fills the performance gap between node-local storage and parallel file systems. With burst buffers, an application can quickly store checkpoints with increased reliability. In this work, we explore how burst buffers can improve efficiency compared to using only node-local storage. To fully exploit the bandwidth of burst buffers, we develop a user-level Infini Band-based file system (IBIO). We also develop performance models for coordinated and uncoordinated checkpoint/restart strategies, and we apply those models to investigate the best checkpoint strategy using burst buffers on future large-scale systems. Kento Sato, Kathryn Mohror, Adam Moody, Todd Gamblin, Bronis R. de Supinski, Naoya Maruyama, Satoshi Matsuoka |
CCGRID | 3 |
| 2014 | FMI: Fault Tolerant Messaging Interface for Fast and Transparent RecoveryabstractFuture supercomputers built with more components will enable larger, higher-fidelity simulations, but at the cost of higher failure rates. Traditional approaches to mitigating failures, such as checkpoint/restart (C/R) to a parallel file system incur large overheads. On future, extreme-scale systems, it is unlikely that traditional C/R will recover a failed application before the next failure occurs. To address this problem, we present the Fault Tolerant Messaging Interface (FMI), which enables extremely low-latency recovery. FMI accomplishes this using a survivable communication runtime coupled with fast, in-memory C/R, and dynamic node allocation. FMI provides message-passing semantics similar to MPI, but applications written using FMI can run through failures. The FMI runtime software handles fault tolerance, including check pointing application state, restarting failed processes, and allocating additional nodes when needed. Our tests show that FMI runs with similar failure-free performance as MPI, but FMI incurs only a 28% overhead with a very high mean time between failures of 1 minute. Kento Sato, Adam Moody, Kathryn Mohror, Todd Gamblin, Bronis R. de Supinski, Naoya Maruyama, Satoshi Matsuoka |
IPDPS | 2 |
| 2014 | Detailed Modeling and Evaluation of a Scalable Multilevel Checkpointing SystemabstractHigh-performance computing (HPC) systems are growing more powerful by utilizing more components. As the system mean time before failure correspondingly drops, applications must checkpoint frequently to make progress. However, at scale, the cost of checkpointing becomes prohibitive. A solution to this problem is multilevel checkpointing, which employs multiple types of checkpoints in a single run. Lightweight checkpoints can handle the most common failure modes, while more expensive checkpoints can handle severe failures. We designed a multilevel checkpointing library, the Scalable Checkpoint/Restart (SCR) library, that writes lightweight checkpoints to node-local storage in addition to the parallel file system. We present probabilistic Markov models of SCR's performance. We show that on future large-scale systems, SCR can lead to a gain in machine efficiency of up to 35 percent, and reduce the load on the parallel file system by a factor of two. Additionally, we predict that checkpoint scavenging, or only writing checkpoints to the parallel file system on application termination, can reduce the load on the parallel file system by 20 × on today's systems and still maintain high application efficiency. Kathryn Mohror, Adam Moody, Greg Bronevetsky, Bronis R. de Supinski |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | A 1 PB/s file system to checkpoint three million MPI tasks
Raghunath Rajachandrasekar, Adam Moody, Kathryn Mohror, Dhabaleswar K. Panda 0001 |
HPDC | 2 |
| 2013 | Efficient and Scalable Retrieval Techniques for Global File PropertiesabstractLarge-scale systems typically mount many different file systems with distinct performance characteristics and capacity. Applications must efficiently use this storage in order to realize their full performance potential. Users must take into account potential file replication throughout the storage hierarchy as well as contention in lower levels of the I/O system, and must consider communicating the results of file I/O between application processes to reduce file system accesses. Addressing these issues and optimizing file accesses requires detailed runtime knowledge of file system performance characteristics and the location(s) of files on them. In this paper, we propose Fast Global File Status (FGFS), a scalable mechanism to retrieve file information, such as its degree of distribution or replication and consistency. We use a novel node-local technique that turns expensive, non-scalable file system calls into simple string comparison operations. FGFS raises the namespace of a locally-defined file path to a global namespace with little or no file system calls to obtain global file properties efficiently. Our evaluation on a large multi-physics application shows that most FGFS file status queries on its executable and 848 shared library files complete in 272 milliseconds or faster at 32,768 MPI processes. Even the most expensive operation, which checks global file consistency, completes in under 7 seconds at this scale, an improvement of several orders of magnitude over the traditional checksum technique. Dong H. Ahn, Michael J. Brim, Bronis R. de Supinski, Todd Gamblin, Gregory L. Lee, Matthew P. LeGendre, Barton P. Miller, Adam Moody, Martin Schulz 0001 |
IPDPS | 8 |
| 2012 | Designing Non-blocking Allreduce with Collective Offload on InfiniBand Clusters: A Case Study with Conjugate Gradient SolversabstractScientists across a wide range of domains increasingly rely on computer simulation for their investigations. Such simulations often spend a majority of their run-times solving large systems of linear equations that require vast amounts of computational power and memory. It is hence critical to design solvers in a highly efficient and scalable manner. Hypre is a high performance, scalable software library that offers several optimized linear solver routines and pre-conditioners. In this paper, we study the characteristics of Hypre's Preconditioned Conjugate Gradient (PCG) solver algorithm. The PCG routine is known to spend a majority of its communication time in the MPI All reduce operation to compute a global summation during the inner product operation. The MPI All reduce is a blocking operation, whose latency is often a limiting factor to the overall efficiency of the PCG solver routine, and correspondingly the performance of simulations that rely on this solver. Hence, hiding the latency of the MPI All reduce operation is critical towards scaling the PCG solver routine and improving the performance of many simulations. The upcoming revision of MPI, MPI-3, will provide support for non-blocking collective communication to enable latency-hiding. The latest Infini Band adapter from Mellanox, ConnectX-2, enables offloading of generalized lists of communication operations to the network interface. Such an interface can be leveraged to design non-blocking collective operations. In this paper, we design fully functional, scalable algorithms for the MPI Iall reduce operation, based on the network offload technology. To the best of our knowledge, these network offload-based algorithms are the first to be presented for the MPI Iall reduce operation. Our designs scale beyond 512 processes and we achieve near perfect communication/computation overlap. We also re-design the PCG solver routine to leverage our proposed MPI Iall reduce operation to hide the latency of the global reduction operations. We observe up to 21% improvements in the run-times of the PCG routine, when compared to the default PCG implementation in Hypre. We also note that about 16% of the overall benefits are due to overlapping the All reduce operations. Krishna Chaitanya Kandalla, Ulrike Meier Yang, Jeff Keasler, Tzanio V. Kolev, Adam Moody, Hari Subramoni, Karen A. Tomko, Jérôme Vienne, Bronis R. de Supinski, Dhabaleswar K. Panda 0001 |
IPDPS | 5 |
| 2012 | McrEngine: a scalable checkpointing system using data-aware aggregation and compressionabstractHigh performance computing (HPC) systems use checkpoint-restart to tolerate failures. Typically, applications store their states in checkpoints on a parallel file system (PFS). As applications scale up, checkpoint-restart incurs high overheads due to contention for PFS resources. The high overheads force large-scale applications to reduce checkpoint frequency, which means more compute time is lost in the event of failure. We alleviate this problem through a scalable checkpointrestart system, MCRENGINE. MCRENGINE aggregates checkpoints from multiple application processes with knowledge of the data semantics available through widely-used I/O libraries, e.g., HDF5 and netCDF, and compresses them. Our novel scheme improves compressibility of checkpoints up to 115% over simple concatenation and compression. Our evaluation with large-scale application checkpoints show that MCRENGINE reduces checkpointing overhead by up to 87% and restart overhead by up to 62% over a baseline with no aggregation or compression. Tanzima Z. Islam, Kathryn Mohror, Saurabh Bagchi, Adam Moody, Bronis R. de Supinski, Rudolf Eigenmann |
SC | 4 |
| 2012 | Design and modeling of a non-blocking checkpointing systemabstractAs the capability and component count of systems increase, the MTBF decreases. Typically, applications tolerate failures with checkpoint/restart to a parallel file system (PFS). While simple, this approach can suffer from contention for PFS resources. Multi-level checkpointing is a promising solution. However, while multi-level checkpointing is successful on today's machines, it is not expected to be sufficient for exascale class machines, which are predicted to have orders of magnitude larger memory sizes and failure rates. Our solution combines the benefits of non-blocking and multi-level checkpointing. In this paper, we present the design of our system and model its performance. Our experiments show that our system can improve efficiency by 1.1 to 2.0x on future machines. Additionally, applications using our checkpointing system can achieve high efficiency even when using a PFS with lower bandwidth. Kento Sato, Naoya Maruyama, Kathryn Mohror, Adam Moody, Todd Gamblin, Bronis R. de Supinski, Satoshi Matsuoka |
SC | 4 |
| 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processesabstractOver the last decade, InfiniBand has become an increasingly popular interconnect for deploying modern supercomputing systems. However, there exists no detection service that can discover the underlying network topology in a scalable manner and expose this information to runtime libraries and users of the high performance computing systems in a convenient way. In this paper, we design a novel and scalable method to detect the InfiniBand network topology by using Neighbor-Joining techniques (NJ). To the best of our knowledge, this is the first instance where the neighbor joining algorithm has been applied to solve the problem of detecting InfiniBand network topology. We also design a network-topology-aware MPI library that takes advantage of the network topology service. The library places processes taking part in the MPI job in a network-topology-aware manner with the dual aim of increasing intra-node communication and reducing the long distance inter-node communication across the InfiniBand fabric. Hari Subramoni, Sreeram Potluri, Krishna Chaitanya Kandalla, William L. Barth, Jérôme Vienne, Jeff Keasler, Karen A. Tomko, Karl W. Schulz, Adam Moody, Dhabaleswar K. Panda 0001 |
SC | 9 |
| 2011 | Exascale Algorithms for Generalized MPI_Comm_split
Adam Moody, Dong H. Ahn, Bronis R. de Supinski |
EuroMPI | 1 |
| 2010 | Design, Modeling, and Evaluation of a Scalable Multi-level Checkpointing SystemabstractHigh-performance computing (HPC) systems are growing more powerful by utilizing more hardware components. As the system mean-time-before-failure correspondingly drops, applications must checkpoint more frequently to make progress. However, as the system memory sizes grow faster than the bandwidth to the parallel file system, the cost of checkpointing begins to dominate application run times. Multi-level checkpointing potentially solves this problem through multiple types of checkpoints with different costs and different levels of resiliency in a single run. This solution employs lightweight checkpoints to handle the most common failure modes and relies on more expensive checkpoints for less common, but more severe failures. This theoretically promising approach has not been fully evaluated in a large- scale, production system context. We have designed the Scalable Checkpoint/Restart (SCR) library, a multi-level checkpoint system that writes checkpoints to RAM, Flash, or disk on the compute nodes in addition to the parallel file system. We present the performance and reliability properties of SCR as well as a probabilistic Markov model that predicts its performance on current and future systems. We show that multi-level checkpointing improves efficiency on existing large-scale systems and that this benefit increases as the system size grows. In particular, we developed low-cost checkpoint schemes that are 100x-1000x faster than the parallel file system and effective against 85% of our system failures. This leads to a gain in machine efficiency of up to 35%, and it reduces the the load on the parallel file system by a factor of two on current and future systems. Adam Moody, Greg Bronevetsky, Kathryn Mohror, Bronis R. de Supinski |
SC | 1 |
| 2009 | Topology agnostic hot-spot avoidance with InfiniBandabstractAbstract InfiniBand has become a very popular interconnect due to its advanced features and open standard. Large‐scale InfiniBand clusters are becoming very popular, as reflected by the TOP 500 supercomputer rankings. However, even with popular topologies such as constant bi‐section bandwidth Fat Tree, hot‐spots may occur with InfiniBand due to inappropriate configuration of network paths, presence of other jobs in the network and un‐availability of adaptive routing. In this paper, we present a hot‐spot avoidance layer (HSAL) for InfiniBand, which provides hot‐spot avoidance using path bandwidth estimation and multi‐pathing using LMC mechanism, without taking the network topology into account. We propose an adaptive striping policy with batch‐based striping and sorting approach, for efficient utilization of disjoint network paths. Integration of HSAL with MPI, thede factoprogramming model of clusters, shows promising results with collective communication primitives and MPI applications. Copyright © 2008 John Wiley & Sons, Ltd. Abhinav Vishnu, Matthew J. Koop, Adam Moody, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Hot-Spot Avoidance With Multi-Pathing Over InfiniBand: An MPI PerspectiveabstractLarge scale InfiniBand clusters are becoming increasingly popular, as reflected by the TOP 500 supercomputer rankings. At the same time, fat tree has become a popular interconnection topology for these clusters, since it allows multiple paths to be available in between a pair of nodes. However, even with fat tree, hot-spots may occur in the network depending upon the route configuration between end nodes and communication pattern(s) in the application. To make matters worse, the deterministic routing nature of InfiniBand limits the application from effective use of multiple paths transparently and avoid the hot-spots in the network. Simulation based studies for switches and adapters to implement congestion control have been proposed in the literature. However, these studies have focussed on providing congestion control for the communication path, and not on utilizing multiple paths in the network for hot-spot avoidance. In this paper, we design an MPI functionality, which provides hot-spot avoidance for different communications, without a priori knowledge of the pattern. We leverage LMC (LID mask count) mechanism of InfiniBand to create multiple paths in the network and present the design issues (scheduling policies, selecting number of paths, scalability aspects) of our design. We implement our design and evaluate it with Pallas collective communication and MPI applications. On an InfiniBand cluster with 48 processes, MPI All-to-all personalized shows an improvement of 27%. Our evaluation with NAS parallel benchmarks on 64 processes shows significant improvement in execution time with this functionality. Abhinav Vishnu, Matthew J. Koop, Adam Moody, Amith R. Mamidala, Sundeep Narravula, Dhabaleswar K. Panda 0001 |
CCGRID | 3 |
| 2003 | Scalable NIC-based Reduction on Large-scale ClustersabstractMany parallel algorithms require efficient reduction collectives. In response, researchers have designed algorithms considering a range of parameters including data size, system size, and communication characteristics. Throughout this past work, however, processing was limited to the host CPU. Today, modern Network Interface Cards (NICs) sport programmable processors with substantial memory, and thus introduce a fresh variable into the equation. In this paper, we investigate this new option in the context of large-scale clusters. Through experiments on the 960-node, 1920-processor ASCI Linux Cluster (ALC) at Lawrence Livermore National Laboratory, we show that NIC-based reductions outperform host-based algorithms in terms of reduced latency and increased consistency. In particular, in the largest configuration tested - 1812 processors - our NIC-based algorithm summed single-element vectors of 32-bit integers and 64-bit floating-point numbers in 73 µs and 118 µs, respectively. These results represent respective improvements of 121% and 39% over the production-level MPI library. Adam Moody, Juan Fernández Peinador, Fabrizio Petrini, Dhabaleswar K. Panda 0001 |
SC | 1 |