EDBT 2026 Demo / reviewers in the wild / expert
Matthew Wolf
dblp:15/434
· DBLP profile ↗
66ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0002-8393-4436ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Revisiting Erasure Codes: A Configuration PerspectiveabstractErasure coding (EC) plays a crucial role in the fault tolerance of modern distributed storage systems (DSS). Inspired by recent research on storage configuration, we study the configuration sensitivity of EC in real DSS in this paper. We systematically inject faults to trigger EC recovery under various configurations, and measure the impact on recovery time and storage overhead quantitatively. Our results show that configurations may affect the EC recovery time significantly (e.g., up to 426%). More interestingly, theoretically superior codes may perform worse in DSS under certain configurations. Also, there is a system checking period before EC recovery that accounts for 41% to 58% of the overall system recovery time, which has been largely ignored in previous studies. Finally, in terms of storage overhead, EC may introduce 32.3% to 72.0% more write amplification (WA) than the theoretical expectation, and we derive a formula to help estimate WA more precisely. Our work suggests the importance of considering the context of real DSS for EC research, and we hope the methodology and findings can contribute to a firmer footing for EC optimization in practice. Runzhou Han, Tabassum Mahmud, Zeren Yang, Vladislav Esaulov, Lipeng Wan 0001, Yong Chen 0001, Jim Wayda, Matthew Wolf, Mai Zheng |
HotStorage | 9 |
| 2022 | Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path ForwardabstractThe ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows. Kshitij Mehta, Ashley Cliff, Frédéric Suter, Angelica M. Walker, Matthew Wolf, Daniel A. Jacobson, Scott Klasky |
e-Science | 5 |
| 2022 | F*** workflows: when parts of FAIR are missingabstractThe FAIR principles for scientific data (Findable, Accessible, Interoperable, Reusable) are also relevant to other digital objects such as research software and scientific workflows that operate on scientific data. The FAIR principles can be applied to the data being handled by a scientific workflow as well as the processes, software, and other infrastructure which are necessary to specify and execute a workflow. The FAIR principles were designed as guidelines, rather than rules, that would allow for differences in standards for different communities and for different degrees of compliance. There are many practical considerations which impact the level of FAIR-ness that can actually be achieved, including policies, traditions, and technologies. Because of these considerations, obstacles are often encountered during the workflow lifecycle that trace directly to shortcomings in the implementation of the FAIR principles. Here, we detail some cases, without naming names, in which data and workflows were Findable but otherwise lacking in areas commonly needed and expected by modern FAIR methods, tools, and users. We describe how some of these problems, all of which were overcome successfully, have motivated us to push on systems and approaches for fully FAIR workflows. Sean R. Wilkinson, Greg Eisenhauer, Anuj J. Kapadia, Kathryn Knight, Jeremy Logan, Patrick M. Widener, Matthew Wolf |
e-Science | 7 |
| 2022 | P-ckpt: Coordinated Prioritized CheckpointingabstractGood prediction accuracy and adequate lead time to failure are key to the success of failure-aware Check-point/Restart (C/R) models on current and future large-scale High-Performance Computing (HPC) systems. This paper develops a novel checkpointing technique, called p-ckpt, that aims to maintain the performance efficiency of failure-aware C/R models even when failures are predicted with a small lead time. The p-ckpt technique is developed for HPC systems with multi-level memory systems to prioritize checkpoints from vulnerable nodes (nodes with predicted failure) in the event of failure prediction. It applies coordination among the nodes within an application so that vulnerable nodes' checkpoint data is stored to the Parallel File System (PFS) first by assigning priorities based on the lead time to failure. Vulnerable nodes thus have low-latency access on the critical path to the PFS before any failure happens. Further, we create the hybrid p-ckpt model by integrating Live Migration (LM) because of its cost-effectiveness and to reduce checkpoint frequency. Our hybrid p-ckpt C/R model considers prediction lead time and checkpoint latency to the PFS to decide on a feasible proactive action such as p-ckpt and LM. Simulations of six real-world applications for the Summit supercomputer show a ≈53-65% reduction in overhead due to the hybrid p-ckpt model compared to a ≈31-61% reduction in a state-of-the-art solution. We assess our C/R models against multiple failure distributions and consider lead time variability and failure prediction accuracy. Based on this evaluation and assessment, we discuss the trade-offs of using these models and their impact on application overhead. Subhendu Behera, Lipeng Wan 0001, Frank Mueller 0001, Matthew Wolf, Scott Klasky |
IPDPS | 4 |
| 2022 | A codesign framework for online data analysis and reductionabstractAbstract Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation‐analysis‐reduction pipelines coupled in‐memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application‐ and system‐specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. This toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next‐generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers. Kshitij Mehta, Bryce Allen, Matthew Wolf, Jeremy Logan, Eric Suchyta, Swati Singhal, Jong Choi 0001, Keichi Takahashi, Kevin A. Huck, Igor Yakushin, Alan Sussman, Todd S. Munson, Ian T. Foster, Scott Klasky |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | MGARD+: Optimizing Multilevel Methods for Error-Bounded Scientific Data ReductionabstractNowadays, data reduction is becoming increasingly important in dealing with the large amounts of scientific data. Existing multilevel compression algorithms offer a promising way to manage scientific data at scale, but may suffer from relatively low performance and reduction quality. In this paper, we propose MGARD+, a multilevel data reduction and refactoring framework drawing on previous multilevel methods, to achieve high-performance data decomposition and high-quality error-bounded lossy compression. Our contributions are four-fold: 1) We propose to leverage a level-wise coefficient quantization method, which uses different error tolerances to quantize the multilevel coefficients. 2) We propose an adaptive decomposition method which treats the multilevel decomposition as a preconditioner and terminates the decomposition process at an appropriate level. 3) We leverage a set of algorithmic optimization strategies to significantly improve the performance of multilevel decomposition/recomposition. 4) We evaluate our proposed method using four real-world scientific datasets and compare with several state-of-the-art lossy compressors. Experiments demonstrate that our optimizations improve the decomposition/recomposition performance of the existing multilevel method by up to$70 \times$, and the proposed compression method can improve compression ratio by up to$2 \times$compared with other state-of-the-art error-bounded lossy compressors under the same level of data distortion. Xin Liang 0001, Ben Whitney, Jieyang Chen, Lipeng Wan 0001, Qing Liu 0002, Dingwen Tao, James Kress, David Pugmire, Matthew Wolf, Norbert Podhorszki, Scott Klasky |
IEEE Trans. Computers | 9 |
| 2021 | Creating a Tools Ecosystem for Cross-Discipline Environmental Data ReuseabstractReusing data is difficult even within well-defined science communities and only gets worse when combining data from multiple communities and disciplines. Through the lens of current work on constructing an environmental epidemiological data set from multiple disciplinary sources, we demonstrate the need for a new tool ecosystem to support heterogeneous Big Data science. Extending existing community standards for schemas and/or data formats through human auditing and wrangling of the data is not feasible at scale. This work therefore suggests new approaches for the multi-disciplinary communities to build a shared tool ecosystem for big data. We discuss both the larger context of data wrangling of epidemiological data sets for novel artificial intelligence algorithms and the specific lessons from working with these multi-disciplinary data sets. Adopting a more model-driven, automatable approach promises not only better efficiency but also removes key sources of human-generated errors and promotes reuse and reproducibility of science data. Jeremy Logan, Greeshma A. Agasthya, Heidi A. Hanson, Matthew Wolf, Heechan Lee, Shaheen Dewji, Hong-Jun Yoon, Anuj J. Kapadia |
IEEE BigData | 4 |
| 2021 | Reusability First: Toward FAIR WorkflowsabstractThe FAIR principles of open science (Findable, Accessible, Interoperable, and Reusable) have had transformative effects on modern large-scale computational science. In particular, they have encouraged more open access to and use of data, an important consideration as collaboration among teams of researchers accelerates and the use of workflows by those teams to solve problems increases. How best to apply the FAIR principles to workflows themselves, and software more generally, is not yet well understood. We argue that the software engineering concept of technical debt management provides a useful guide for application of those principles to workflows, and in particular that it implies reusability should be considered as ‘first among equals’. Moreover, our approach recognizes a continuum of reusability where we can make explicit and selectable the tradeoffs required in workflows for both their users and developers.To this end, we propose a new abstraction approach for reusable workflows, with demonstrations for both synthetic workloads and real-world computational biology workflows. Through application of novel systems and tools that are based on this abstraction, these experimental workflows are refactored to rightsize the granularity of workflow components to efficiently fill the gap between end-user simplicity and general customizability. Our work makes it easier to selectively reason about and automate the connections between trade-offs across user and developer concerns when exposing degrees of freedom for reuse. Additionally, by exposing fine-grained reusability abstractions we enable performance optimizations, as we demonstrate on both institutional-scale and leadership-class HPC resources. Matthew Wolf, Jeremy Logan, Kshitij Mehta, Daniel A. Jacobson, Mikaela Cashman, Angelica M. Walker, Greg Eisenhauer, Patrick M. Widener, Ashley Cliff |
CLUSTER | 1 |
| 2021 | Towards System for Knowledge Representation of Campaign ExperimentationabstractThe campaign is an experimentation construct for codesign activity wherein multiple researchers carry out computational experiments that individually contribute to a shared goal. The larger objective of our research is a system that exists in the experimental environment that constructs a knowledge representation of campaigns and products both produced and consumed such that the campaign can as efficient as possible and the products richly contextualized for reuse. Using campaign experiments running on the Summit machine at Oak Ridge National Labs, we demonstrate early results of support for discovery queries and for detecting when two sweeps are similar. Sachith Withana, Kshitij Mehta, Matthew Wolf, Beth Plale |
e-Science | 3 |
| 2021 | Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUsabstractRapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput-83% of theoretical peak-on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software. Jieyang Chen, Lipeng Wan 0001, Xin Liang 0001, Ben Whitney, Qing Liu 0002, David Pugmire, Nicholas Thompson, Jong Choi 0001, Matthew Wolf, Todd S. Munson, Ian T. Foster, Scott Klasky |
IPDPS | 9 |
| 2020 | Orchestrating Fault Prediction with Live Migration and CheckpointingabstractCheckpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices. Subhendu Behera, Lipeng Wan 0001, Frank Mueller 0001, Matthew Wolf, Scott Klasky |
HPDC | 4 |
| 2019 | Scalable Performance Awareness for In Situ Scientific ApplicationsabstractPart of the promise of exascale computing and the next generation of scientific simulation codes is the ability to bring together time and spatial scales that have traditionally been treated separately. This enables creating complex coupled simulations and in situ analysis pipelines, encompassing such things as "whole device" fusion models or the simulation of cities from sewers to rooftops. Unfortunately, the HPC analysis tools that have been built up over the preceding decades are ill suited to the debugging and performance analysis of such computational ensembles. In this paper, we present a new vision for performance measurement and understanding of HPC codes, MonitoringAnalytics (MONA). MONA is designed to be a flexible, high performance monitoring infrastructure that can perform monitoring analysis in place or in transit by embedding analytics and characterization directly into the data stream, without relying upon delivering all monitoring information to a central database for post-processing. It addresses the trade-offs between the prohibitively expensive capture of all performance characteristics and not capturing enough to detect the features of interest. We demonstrate several uses of MONA; capturing and indexing multi-executable performance profiles to enable later processing, extraction of performance primitives to enable the generation of customizable benchmarks and performance skeletons, and extracting communication and application behaviors to enable better control and placement for the current and future runs of the science ensemble. Relevant performance information based on a system for MONA built from ADIOS and SOSflow technologies is provided for DOE science applications and leadership machines. Matthew Wolf, Julien Dominski, Gabriele Merlo, Jong Choi 0001, Greg Eisenhauer, Stéphane Ethier, Kevin A. Huck, Scott Klasky, Jeremy Logan, Allen D. Malony, Chad Wood |
eScience | 1 |
| 2019 | A Vision for Managing Extreme-Scale Data HoardsabstractScientific data collections grow ever larger, both in terms of the size of individual data items and of the number and complexity of items. To use and manage them, it is important to directly address issues of robust and actionable provenance. We identify three key drivers as our focus: managing the size and complexity of metadata, lack of a priori information to match usage intents between publishers and consumers of data, and support for campaigns over collections of data driven by multi-disciplinary, collaborating teams. We introduce the Hoarde abstraction as an attempt to formalize a way of looking at collections of data to make them more tractable for later use. Hoarde leverages middleware and systems infrastructures for scientific and technical data management. Through the lens of a select group of challenging data usage scenarios, we discuss some of the aspects of implementation, usage, and forward portability of this new view on data management. Jeremy Logan, Kshitij Mehta, Gerd Heber, Scott Klasky, Tahsin M. Kurç, Norbert Podhorszki, Patrick M. Widener, Matthew Wolf |
ICDCS | 8 |
| 2019 | MPI jobs within MPI jobs: A practical way of enabling task-level fault-tolerance in HPC workflows
Justin M. Wozniak, Matthieu Dorier, Robert B. Ross, Tong Shu, Tahsin M. Kurç, Li Tang 0007, Norbert Podhorszki, Matthew Wolf |
Future Gener. Comput. Syst. | 8 |
| 2019 | Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems?abstractFor a production high-performance computing (HPC) system, where storage devices are shared between multiple applications and managed in a best effort manner, I/O contention is often a major problem. In this paper, we propose a balanced messaging-based re-routing in conjunction with throttling at the middleware level. This work tackles two key challenges that have not been fully resolved in the past: whether I/O variability can be reduced on a QoS-less HPC storage system, and how to design a runtime scheduling system that can scale up to a large amount of cores. The proposed scheme uses a two-level messaging system to re-route I/O requests to a less congested storage location so that write performance is improved, while limiting the impact on read by throttling re-routing. An analytical model is derived to guide the setup of optimal throttling factor. We thoroughly analyze the virtual messaging layer overhead and explore whether the in-transit buffering is effective in managing I/O variability. Contrary to the intuition, in-transit buffer cannot completely solve the problem. It can reduce the absolute variability but not the relative variability. The proposed scheme is verified against a synthetic benchmark as well as being used by production applications. Dan Huang 0001, Qing Liu 0002, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Jeremy Logan, George Ostrouchov, Xubin He, Matthew Wolf |
IEEE Trans. Computers | 9 |
| 2018 | Coupling Exascale Multiphysics Applications: Methods and Lessons LearnedabstractWith the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration. Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood |
eScience | 36 |
| 2018 | LADR: low-cost application-level detector for reducing silent output corruptionsabstractApplications running on future high performance computing (HPC) systems are more likely to experience transient faults due to technology scaling trends with respect to higher circuit density, smaller transistor size and near-threshold voltage (NTV) operations. A transient fault could corrupt application state without warning, possibly leading to incorrect application output. Such errors are called silent data corruptions (SDCs). Chao Chen 0024, Greg Eisenhauer, Matthew Wolf, Santosh Pande |
HPDC | 3 |
| 2018 | A View from ORNL: Scientific Data Research Opportunities in the Big Data AgeabstractOne of the core issues across computer and computational science today is adapting to, managing, and learning from the influx of "Big Data". In the commercial space, this problem has led to a huge investment in new technologies and capabilities that are well adapted to dealing with the sorts of human-generated logs, videos, texts, and other large-data artifacts that are processed and resulted in an explosion of useful platforms and languages (Hadoop, Spark, Pandas, etc.). However, translating this work from the enterprise space to the computational science and HPC community has proven somewhat difficult, in part because of some of the fundamental differences in type and scale of data and timescales surrounding its generation and use. We describe a forward-looking research and development plan which centers around the concept of making Input/Output (I/O) intelligent for users in the scientific community, whether they are accessing scalable storage or performing in situ workflow tasks. Much of our work is based on our experience with the Adaptable I/O System (ADIOS 1.X), and our next generation version of the software ADIOS 2.X [1]. Scott Klasky, Matthew Wolf, Mark Ainsworth, Chuck Atkins, Jong Choi 0001, Greg Eisenhauer, Berk Geveci, William F. Godoy, Mark Kim, James Kress, Tahsin M. Kurç, Qing Liu 0002, Jeremy Logan, Arthur B. Maccabe, Kshitij Mehta, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Eric Suchyta, Lipeng Wan 0001 |
ICDCS | 2 |
| 2018 | Understanding and Modeling Lossy Compression Schemes on HPC Scientific DataabstractScientific simulations generate large amounts of floating-point data, which are often not very compressible using the traditional reduction schemes, such as deduplication or lossless compression. The emergence of lossy floating-point compression holds promise to satisfy the data reduction demand from HPC applications; however, lossy compression has not been widely adopted in science production. We believe a fundamental reason is that there is a lack of understanding of the benefits, pitfalls, and performance of lossy compression on scientific data. In this paper, we conduct a comprehensive study on state-of-the-art lossy compression, including ZFP, SZ, and ISABELA, using real and representative HPC datasets. Our evaluation reveals the complex interplay between compressor design, data features and compression performance. The impact of reduced accuracy on data analytics is also examined through a case study of fusion blob detection, offering domain scientists with the insights of what to expect from fidelity loss. Furthermore, the trial and error approach to understanding compression performance involves substantial compute and storage overhead. To this end, we propose a sampling based estimation method that extrapolates the reduction ratio from data samples, to guide domain scientists to make more informed data reduction decisions. Tao Lu 0014, Qing Liu 0002, Xubin He, Huizhang Luo, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Matthew Wolf, Tong Liu 0030, Zhenbo Qiao |
IPDPS | 9 |
| 2017 | TGE: Machine Learning Based Task Graph Embedding for Large-Scale Topology MappingabstractTask mapping is an important problem in parallel and distributed computing. The goal in task mapping is to find an optimal layout of the processes of an application (or a task) onto a given network topology. We target this problem in the context of staging applications. A staging application consists of two or more parallel applications (also referred to as staging tasks) which run concurrently and exchange data over the course of computation. Task mapping becomes a more challenging problem in staging applications, because not only data is exchanged between the staging tasks, but also the processes of a staging task may exchange data with each other. We propose a novel method, called Task Graph Embedding (TGE), that harnesses the observable graph structures of parallel applications and network topologies. TGE employs a machine learning based algorithm to find the best representation of a graph, called an embedding, onto a space in which the task-to-processor mapping problem can be solved. We evaluate and demonstrate the effectiveness of TGE experimentally with the communication patterns extracted from runs of XGC, a large-scale fusion simulation code, on Titan. Jong Choi 0001, Jeremy Logan, Matthew Wolf, George Ostrouchov, Tahsin M. Kurç, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Melissa Romanus, Manish Parashar, Michael Churchill, Choong-Seock Chang |
CLUSTER | 3 |
| 2017 | Extending Skel to Support the Development and Optimization of Next Generation I/O SystemsabstractAs the memory and storage hierarchy get deeper and more complex, it is important to have new benchmarks and evaluation tools that allow us to explore the emerging middleware solutions to use this hierarchy. Skel is a tool aimed at automating and refining this process of studying HPC I/O performance. It works by generating application I/O kernel/benchmarks as determined by a domain-specific model. This paper provides some techniques for extending Skel to address new situations and to answer new research questions. For example, we document use cases as diverse as using Skel to troubleshoot I/O performance issues for remote users, refining an I/O system model, and facilitating the development and testing of a mechanism for runtime monitoring and performance analytics. We also discuss data oriented extensions to Skel to support the study of compression techniques for Exascale scientific data management. Jeremy Logan, Jong Choi 0001, Matthew Wolf, George Ostrouchov, Lipeng Wan 0001, Norbert Podhorszki, William F. Godoy, Scott Klasky, Erich Lohrmann, Greg Eisenhauer, Chad Wood, Kevin A. Huck |
CLUSTER | 3 |
| 2017 | Introducing Weirs: An Abstraction for Next Generation Streaming WorkflowsabstractIn HPC applications, it is widely understood that in situ systems will play a significant role in next generation systems. The rate that next-generation leadership machines will be able to generate data will exceed the bandwidths of the planned I/O systems, leading to a need for in situ processing of the resulting data to reduce it. There have been a number of techniques proposed for in situ workflow systems; we are concerned here specifically with the needs of publish/subscribe or streaming-based approaches, which allow the in situ system to be composed of completely separate executables running concurrently on the hardware. In this paper we introduce an abstract known as a Weir. A weir is a useful mechanism to control state in a point to point streaming workflow in the absence of a third party publish subscribe system. Erich Lohrmann, Greg Eisenhauer, Matthew Wolf |
CLUSTER | 3 |
| 2017 | Canopus: A Paradigm Shift Towards Elastic Extreme-Scale Data Analytics on HPC StorageabstractScientific simulations on high performance computing (HPC) platforms generate large quantities of data. To bridge the widening gap between compute and I/O, and enable data to be more efficiently stored and analyzed, simulation outputs need to be refactored, reduced, and appropriately mapped to storage tiers. However, a systematic solution to support these steps has been lacking on the current HPC software ecosystem. To that end, this paper develops Canopus, a progressive JPEGlike data management scheme for storing and analyzing big scientific data. It co-designs the data decimation, compression and data storage, taking the hardware characteristics of each storage tier into considerations. With reasonably low overhead, our approach refactors simulation data into a much smaller, reduced-accuracy base dataset, and a series of deltas that is used to augment the accuracy if needed. The base dataset and deltas are compressed and written to multiple storage tiers. Data saved on different tiers can then be selectively retrieved to restore the level of accuracy that satisfies data analytics. Thus, Canopus provides a paradigm shift towards elastic data analytics and enables end users to make trade-offs between analysis speed and accuracy on-the-fly. We evaluate the impact of Canopus on unstructured triangular meshes, a pervasive data model used by scientific modeling and simulations. In particular, we demonstrate the progressive data exploration of Canopus using the “blob detection” use case on the fusion simulation data. Tao Lu 0014, Eric Suchyta, David Pugmire, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Norbert Podhorszki, Mark Ainsworth, Matthew Wolf |
CLUSTER | 9 |
| 2017 | Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo |
Euro-Par | 27 |
| 2017 | Canopus: Enabling Extreme-Scale Data Analytics on Big HPC Storage via Progressive Refactoring
Tao Lu 0014, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Qing Liu 0002, David Pugmire, Matthew Wolf, Mark Ainsworth |
HotStorage | 8 |
| 2017 | RBAY: A Scalable and Extensible Information Plane for Federating Distributed Datacenter ResourcesabstractWhile many institutions, whether industrial, academic, or governmental, satisfy their computing needs through public cloud providers, many others still manage their own resources, often as geographically distributed datacenters. Spare capacity from these geographically distributed datacenters could be offered to others, provided there were a mechanism to discover, and then request these resources. Unfortunately, single datacenter administrators tend not to cooperate due to issues of scalability, diverse administrative policies, and site-specific monitoring infrastructure. This paper describes RBAY, an integrated information plane that enables secure and scalable sharing between geographically distributed datacenters. RBAY's key design features are twofold. First, RBAY employs a decentralized `hierarchical aggregation tree' structure to seamlessly aggregate spare resources from geographically distributed datacenters to a global information plane. Second, RBAY attaches to each participating server a `admin-customized' handler, which follows site-specific policy to expose, hide, add, remove resources to RBAY, and thus fulfill the task of `which resource to expose to whom, when, and how'. An experimental evaluation on eight real-world geo-distributed sites demonstrates RBAY's rapid response to composite queries, as well as its extensible, scalable, and lightweight nature. Liting Hu, Douglas M. Blough, Michael A. Kozuch, Matthew Wolf |
ICDCS | 5 |
| 2017 | Exacution: Enhancing Scientific Data Management for ExascaleabstractAs we continue toward exascale, scientific data volume is continuing to scale and becoming more burdensome to manage. In this paper, we lay out opportunities to enhance state of the art data management techniques. We emphasize well-principled data compression, and using it to achieve progressive refinement. This can both accelerate I/O and afford the user increased flexibility when she interacts with the data. The formulation naturally maps onto enabling partitioning of the progressively improving-quality representations of a data quantity into different media-type destinations, to keep the highest priority information as close as possible to the computation, and take advantage of deepening memory/storage hierarchies in ways not previously possible. Careful monitoring is requisite to our vision, not only to verify that compression has not eliminated salient features in the data, but also to better understand the performance of massively parallel scientific applications. Increased mathematical rigor would be ideal,to help bring compression on a better-understood theoretical footing, closer to the relevant scientific theory, more aware of constraints imposed by the science, and more tightly error-controlled. Throughout, we highlight pathfinding research we have begun exploring related these topics, and comment toward future work that will be needed. Scott Klasky, Eric Suchyta, Mark Ainsworth, Qing Liu 0002, Ben Whitney, Matthew Wolf, Jong Choi 0001, Ian T. Foster, Mark Kim, Jeremy Logan, Kshitij Mehta, Todd S. Munson, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Lipeng Wan 0001 |
ICDCS | 6 |
| 2017 | Comprehensive Measurement and Analysis of the User-Perceived I/O Performance in a Production Leadership-Class Storage SystemabstractWith the increase of the scale and intensity of the parallel I/O workloads generated by those scientific applications running on high performance computing facilities, understanding the I/O dynamics, especially the root cause of the I/O performance variability and degradation in HPC environment, have become extremely critical to the HPC community. In this paper, we run extensive I/O measuring tests on a production leadership-class storage system to capture the performance variabilities of large-scale parallel I/O. Analyzing these results and its statistic correlation revealed some valuable insights into the characteristics of the storage system and the root cause of I/O performance variability. Further, we leverage these findings and propose an I/O middleware design refactoring which can improve the performance of the parallel I/O by optimizing the data striping and placement. Our preliminary evaluation results demonstrate the proposed approach can reduce the average per-process write latency by at least 80% and the maximum per-process write latency by at least 20%. Lipeng Wan 0001, Matthew Wolf, Feiyi Wang, Jong Choi 0001, George Ostrouchov, Scott Klasky |
ICDCS | 2 |
| 2016 | Landrush: Rethinking In-Situ Analysis for GPGPU WorkflowsabstractIn-situ analysis on the output data of scientific simulations has been made necessary by ever-growing output data volumes and increasing costs of data movement as supercomputing is moving towards exascale. With hardware accelerators like GPUs becoming increasingly common in high end machines, new opportunities arise to co-locate scientific simulations and online analysis performed on the scientific data generated by the simulations. However, the asynchronous nature of GPGPU programming models and the limited context-switching capabilities on the GPU pose challenges to co-locating the scientific simulation and analysis on the same GPU. This paper dives deeper into these challenges to understand how best to co-locate analysis with scientific simulations on the GPUs in HPC clusters. Specifically, our 'Landrush' approach to GPU sharing proposes a solution that utilizes idle cycles on the GPU to provide an improved time-to-answer, that is, the total time to run the scientific simulation and analysis of the generated data. Landrush is demonstrated with experimental results obtained from leadership high-end applications on ORNL's Titan supercomputer, which show that (i) GPU-based scientific simulations have varying degrees of idle cycles to afford useful analysis task co-location, and (ii) the inability to context switch on the GPU at instruction granularity can be overcome by careful control of the analysis kernel launches and software-controlled early completion of analysis kernel executions. Results show that Landrush is superior in terms of time-to-answer compared to serially running simulations followed by analysis or by relying on the GPU driver and hardwired thread dispatcher to run analysis concurrently on a single GPU. Anshuman Goswami, Yuan Tian 0004, Karsten Schwan, Fang Zheng 0003, Jeffrey Young 0001, Matthew Wolf, Greg Eisenhauer, Scott Klasky |
CCGrid | 6 |
| 2016 | SuperGlue: Standardizing Glue Components for HPC WorkflowsabstractThis poster describes our work on SuperGlue, a set of generic, reusable components for composing scientific workflows. These are distributed data analysis and manipulation tools that can be chained together to form a variety of real-time workflows providing analytical results during the execution of the primary scientific code. Unlike existing components used in IAWs, SuperGlue components do not have a fixed data type. This one change enables using these components on completely different kinds of simulations that share nothing in their output format. Key to making this work is using a typed transport mechanism between different components. Many options exist for these transports and the particular mechanism selected is not critical. Jay F. Lofstead, Alexis Champsaur, Jai Dayal, Matthew Wolf, Greg Eisenhauer |
CLUSTER | 4 |
| 2016 | FlashStager: Improving the Performance of SSD-Based Data Staging Systems via Write RedirectionabstractWhen SSDs are used for in-situ execution of data-intensive scientific workflows, it is challenging to obtain consistently high I/O throughput because its I/O efficiency can be compromised for serving write and read requests simultaneously. This issue is so-called write-read interference. In this paper, we propose a novel scheme named FlashStager, which can isolate writes from reads using write redirection to improve data staging performance of SSDs by minimizing the write-read interference. Not only can it detect the interference, but also evaluate whether it is cost-effective to resolve it by executing the write redirection according to its correlation with write ratio and request size. Our experiments with both micro-benchmarks and real scientific applications show than FlashStager can improve I/O performance of staging by 40% on average. Xuechen Zhang 0001, Fang Zheng 0003, Karsten Schwan, Matthew Wolf |
CLUSTER | 4 |
| 2016 | GraphIn: An Online High Performance Incremental Graph Processing Framework
Dipanjan Sengupta, Narayanan Sundaram, Theodore L. Willke, Jeffrey Young 0001, Matthew Wolf, Karsten Schwan |
Euro-Par | 6 |
| 2016 | Performance analysis, design considerations, and applications of extreme-scale in situ infrastructuresabstractA key trend facing extreme-scale computational science is the widening gap between computational and I/O rates, and the challenge that follows is how to best gain insight from simulation data when it is increasingly impractical to save it to persistent storage for subsequent visual exploration and analysis. One approach to this challenge is centered around the idea of in situ processing, where visualization and analysis processing is performed while data is still resident in memory. This paper examines several key design and performance issues related to the idea of in situ processing at extreme scale on modern platforms: scalability, overhead, performance measurement and analysis, comparison and contrast with a traditional post hoc approach, and interfacing with simulation codes. We illustrate these principles in practice with studies, conducted on large-scale HPC platforms, that include a miniapplication and multiple science application codes, one of which demonstrates in situ methods in use at greater than 1M-way concurrency. Utkarsh Ayachit, Andrew C. Bauer, Earl P. N. Duque, Greg Eisenhauer, Nicola J. Ferrier, Junmin Gu, Kenneth E. Jansen, Burlen Loring, Zarija Lukic, Suresh Menon, Dmitriy Morozov, Patrick O'Leary, Reetesh Ranjan, Michel E. Rasquin, Christopher P. Stone, Venkatram Vishwanath, Gunther H. Weber, Brad Whitlock, Matthew Wolf, Kesheng Wu, E. Wes Bethel |
SC | 19 |
| 2015 | SODA: Science-Driven Orchestration of Data AnalyticsabstractAs scientific simulation applications evolve on the path towards exascale, a new model of scientific inquiry is required where concurrently with the running simulation, online analytics operate on the data it produces. By avoiding offline data storage except when absoluately necessary, it enables speeding up the scientific discovery process by providing rapid insights into the simulated science phenomena and affording more frequent, detailed data analytics than is possible with the traditional purely offline approach of using disk for intermediate data storage. However, a challenge for online analytics is to respond to behavior dynamics caused by changing simulation outputs and by unforeseen events on the underlying hardware/software platforms. This paper presents SODA, a set of run-time abstractions for online orchestration of data analytics, realized by embedding analytics tasks into workstations that monitor component behavior and enable responses to run-time changes in their resource demands and in the platform's resource availability. For high end simulations running on a leadership class machine, experimental evaluations show SODA can invoke efficient orchestration operations responding to a diverse set of run-time dynamics at different granularities to meet end-user and analysis specific requirements. Jai Dayal, Jay F. Lofstead, Greg Eisenhauer, Karsten Schwan, Matthew Wolf, Hasan Abbasi, Scott Klasky |
e-Science | 5 |
| 2014 | Flexpath: Type-Based Publish/Subscribe System for Large-Scale Science AnalyticsabstractAs high-end systems move toward exascale sizes, a new model of scientific inquiry being developed is one in which online data analytics run concurrently with the high end simulations producing data outputs. Goals are to gain rapid insights into the ongoing scientific processes, assess their scientific validity, and/or initiate corrective or supplementary actions by launching additional computations when needed. The Flex path system presented in this paper addresses the fundamental problem of how to structure and efficiently implement the communications between high end simulations and concurrently running online data analytics, the latter comprised of componentized dynamic services and service pipelines. Using a type-based publish/subscribe approach, Flexpath encourages diversity by permitting analytics services to differ in their computational and scaling characteristics and even in their internal execution models. Flex path uses direct and MxN connections between interacting services to reduce data movements, to allow for runtime connectivity changes to accommodate component arrivals/departures, and to support the multiple underlying communication protocols used for analytics workflows in which simulation outputs are processed by analytics services residing on the same nodes where they are generated, on the same machine, and/or on attached or remote analytics engines. This paper describes the design and implementation of Flex path, and evaluates it with two widely used scientific applications and their associated data analytics methods. Jai Dayal, Drew Bratcher, Greg Eisenhauer, Karsten Schwan, Matthew Wolf, Xuechen Zhang 0001, Hasan Abbasi, Scott Klasky, Norbert Podhorszki |
CCGRID | 5 |
| 2014 | Scibox: Online Sharing of Scientific Data via the CloudabstractCollaborative science demands global sharing of scientific data. But it cannot leverage universally accessible cloud-based infrastructures like Drop Box, as those offer limited interfaces and inadequate levels of access bandwidth. We present the Scibox cloud facility for online sharing scientific data. It uses standard cloud storage solutions, but offers a usage model in which high end codes can write/read data to/from the cloud via the APIs they already use for their I/O actions. With Scibox, data upload/download volumes are controlled via Data Reduction-functions stated by end users and applied at the data source, before data is moved, with further gains in efficiency obtained by combining DR-functions to move exactly what is needed by current data consumers. We evaluate Scibox with science applications and their representative data analytics - the GTS fusion and the combustion image processing - demonstrating the potential for ubiquitous data access with substantial reductions in network traffic. Jian Huang 0006, Xuechen Zhang 0001, Greg Eisenhauer, Karsten Schwan, Matthew Wolf, Stéphane Ethier, Scott Klasky |
IPDPS | 5 |
| 2014 | Hello ADIOS: the challenges and lessons of developing leadership class I/O frameworksabstractSUMMARY Applications running on leadership platforms are more and more bottlenecked by storage input/output (I/O). In an effort to combat the increasing disparity between I/O throughput and compute capability, we created Adaptable IO System (ADIOS) in 2005. Focusing on putting users first with a service oriented architecture, we combined cutting edge research into new I/O techniques with a design effort to create near optimal I/O methods. As a result, ADIOS provides the highest level of synchronous I/O performance for a number of mission critical applications at various Department of Energy Leadership Computing Facilities. Meanwhile ADIOS is leading the push for next generation techniques including staging and data processing pipelines. In this paper, we describe the startling observations we have made in the last half decade of I/O research and development, and elaborate the lessons we have learned along this journey. We also detail some of the challenges that remain as we look toward the coming Exascale era. Copyright © 2013 John Wiley & Sons, Ltd. Qing Liu 0002, Jeremy Logan, Yuan Tian 0004, Hasan Abbasi, Norbert Podhorszki, Jong Choi 0001, Scott Klasky, Roselyne Tchoua, Jay F. Lofstead, Ron A. Oldfield, Manish Parashar, Nagiza F. Samatova, Karsten Schwan, Arie Shoshani, Matthew Wolf, Kesheng Wu, Weikuan Yu |
Concurr. Comput. Pract. Exp. | 15 |
| 2013 | FlexQuery: An online query system for interactive remote visual data exploration at large scaleabstractThe remote visual exploration of live data generated by scientific simulations is useful for scientific discovery, performance monitoring, and online validation for the simulation results. Online visualization methods are challenged, however, by the continued growth in the volume of simulation output data that has to be transferred from its source - the simulation running on the high end machine - to where it is analyzed, visualized, and displayed. A specific challenge in this context is limits in the communication bandwidth between data source(s) and sinks. Previous work places queries `near' data sources, exploiting their data reduction capabilities, but such work does not address the common scenario in which scientists make multiple different queries on the data being produced. This paper considers the general case in which science users are interested in different (sub)sets of the data produced by a high end simulation. We offer the FlexQuery online data query system that can deploy and execute data queries `along' the I/O and analytics pipelines. FlexQuery carefully extends such analytics pipelines, using online performance monitoring and data location tracking, to realize data queries in ways that minimize additional data movement and offer low latency in data query execution. Using a real-world scientific application - the Maya astrophysics code and its analytics workflow - we demonstrate FlexQuery's ability to dynamically deploy queries for low-latency remote data visualization. Hongbo Zou, Karsten Schwan, Magdalena Slawiñska, Matthew Wolf, Greg Eisenhauer, Fang Zheng 0003, Jai Dayal, Jeremy Logan, Qing Liu 0002, Scott Klasky, Tanja Bode, Michael Clark, Matthew Kinsey |
CLUSTER | 4 |
| 2013 | ADIOS Visualization Schema: A First Step Towards Improving Interdisciplinary Collaboration in High Performance ComputingabstractScientific communities have benefitted from a significant increase of available computing and storage resources in the last few decades. For science projects that have access to leadership scale computing resources, the capacity to produce data has been growing exponentially. Teams working on such projects must now include, in addition to the traditional application scientists, experts in various disciplines including applied mathematicians for development of algorithms, visualization specialists for large data, and I/O specialists. Sharing of knowledge and data is becoming a requirement for scientific discovery, providing useful mechanisms to facilitate this sharing is a key challenge for e-Science. Our hypothesis is that in order to decrease the time to solution for application scientists we need to lower the barrier of entry into related computing fields. We aim at improving users' experience when interacting with a vast software ecosystem and/or huge amount of data, while maintaining focus on their primary research field. In this context we present our approach to bridge the gap between the application scientists and the visualization experts through a visualization schema as a first step and proof of concept for a new way to look at interdisciplinary collaboration among scientists dealing with big data. The key to our approach is recognizing that our users are scientists who mostly work as islands. They tend to work in very specialized environment but occasionally have to collaborate with other researchers in order to take full advantage of computing innovations and get insight from big data. We present an example of identifying the connecting elements between one of such relationships and offer a liaison schema to facilitate their collaboration. Roselyne Tchoua, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Jeremy Logan, Kenneth Moreland, Jingqing Mu, Manish Parashar, Norbert Podhorszki, David Pugmire, Matthew Wolf |
e-Science | 11 |
| 2013 | FlexIO: I/O Middleware for Location-Flexible Scientific Data AnalyticsabstractIncreasingly severe I/O bottlenecks on High-End Computing machines are prompting scientists to process simulation output data online while simulations are running and before storing data on disk. There are several options to place data analytics along the I/O path: on compute nodes, on separate nodes dedicated to analytics, or after data is stored on persistent storage. Since different placements have different impact on performance and cost, there is a consequent need for flexibility in the location of data analytics. The FlexIO middleware described in this paper makes it easy for scientists to obtain such flexibility, by offering simple abstractions and diverse data movement methods to couple simulation with analytics. Various placement policies can be built on top of FlexIO to exploit the trade-offs in performing analytics at different levels of the I/O hierarchy. Experimental results demonstrate that FlexIO can support a variety of simulation and analytics workloads at large scale through flexible placement options, efficient data movement, and dynamic deployment of data manipulation functionalities. Fang Zheng 0003, Hongbo Zou, Greg Eisenhauer, Karsten Schwan, Matthew Wolf, Jai Dayal, Jianting Cao, Hasan Abbasi, Scott Klasky, Norbert Podhorszki, Hongfeng Yu 0001 |
IPDPS | 5 |
| 2013 | GoldRush: resource efficient in situ scientific data analytics using fine-grained interference aware executionabstractSevere I/O bottlenecks on High End Computing platforms call for running data analytics in situ. Demonstrating that there exist considerable resources in compute nodes un-used by typical high end scientific simulations, we leverage this fact by creating an agile runtime, termed GoldRush, that can harvest those otherwise wasted, idle resources to efficiently run in situ data analytics. GoldRush uses fine-grained scheduling to "steal" idle resources, in ways that minimize interference between the simulation and in situ analytics. This involves recognizing the potential causes of on-node resource contention and then using scheduling methods that prevent them. Experiments with representative science applications at large scales show that resources harvested on compute nodes can be leveraged to perform useful analytics, significantly improving resource efficiency, reducing data movement costs incurred by alternate solutions, and posing negligible impact on scientific simulations. Fang Zheng 0003, Hongfeng Yu 0001, Can Hantas, Matthew Wolf, Greg Eisenhauer, Karsten Schwan, Hasan Abbasi, Scott Klasky |
SC | 4 |
| 2013 | Multi-course approaches to curriculum 2013's parallel and distributed computing (abstract only)abstractThe emerging CS2013 Curriculum recommendations call for greatly expanded emphasis on parallel and distributed computing (PDC), in response to recent industry changes. CS2013's PDC knowledge units relate to many undergraduate courses. Participants in this BOF will consider responses to CS2013 PDC recommendations that involve multiple undergraduate CS courses at an institution, as opposed to approaches that concentrate PDC topics primarily within a single course. This sharing and brainstorming session will bring together: people having experience with a multi-course or multi-level approach to teaching PDC; people contemplating a multi-course approach to introducing PDC material; and people wishing to provide and/or hear rationale for a multi-course strategy for teaching PDC. Richard A. Brown, Joel Adams 0001, David P. Bunde, Jens Mache, Elizabeth Shoop, Michael A. Smith 0002, Paul F. Steinberg, Matthew Wolf |
SIGCSE | 8 |
| 2012 | Mining hidden mixture context with ADIOS-P to improve predictive pre-fetcher accuracyabstractPredictive pre-fetcher, which predicts future data access events and loads the data before users requests, has been widely studied, especially in file systems or web contents servers, to reduce data load latency. Especially in scientific data visualization, pre-fetching can reduce the IO waiting time. In order to increase the accuracy, we apply a data mining technique to extract hidden information. More specifically, we apply a data mining technique for discovering the hidden contexts in data access patterns and make prediction based on the inferred context to boost the accuracy. In particular, we performed Probabilistic Latent Semantic Analysis (PLSA), a mixture model based algorithm popular in the text mining area, to mine hidden contexts from the collected user access patterns and, then, we run a predictor within the discovered context. We further improve PLSA by applying the Deterministic Annealing (DA) method to overcome the local optimum problem. In this paper we demonstrate how we can apply PLSA and DA optimization to mine hidden contexts from users data access patterns and improve predictive pre-fetcher performance. Jong Choi 0001, Hasan Abbasi, David Pugmire, Norbert Podhorszki, Scott Klasky, Cristian Capdevila, Manish Parashar, Matthew Wolf, Judy Qiu, Geoffrey C. Fox |
eScience | 8 |
| 2012 | Understanding I/O Performance Using I/O Skeletal Applications
Jeremy Logan, Scott Klasky, Hasan Abbasi, Qing Liu 0002, George Ostrouchov, Manish Parashar, Norbert Podhorszki, Yuan Tian 0004, Matthew Wolf |
Euro-Par | 9 |
| 2012 | A system-aware optimized data organization for efficient scientific analyticsabstractLarge-scale scientific applications on High End Computing systems produce a large volume of highly complex datasets. Such data imposes a grand challenge to conventional storage systems for the need of efficient I/O solutions during both the simulation runtime and data post-processing phases. With the mounting needs of scientific discovery, the read performance of large-scale simulations has becomes a critical issue for the HPC community. In this study, we propose a system-aware optimized data organization strategy that can organize data blocks of multidimensional scientific data efficiently based on simulation output and the underlying storage systems, thereby enabling efficient scientific analytics. Our experimental results demonstrate a performance speedup up to 72 times for the combustion simulation S3D, compared to the logically contiguous data layout. Yuan Tian 0004, Scott Klasky, Weikuan Yu, Hasan Abbasi, Bin Wang 0019, Norbert Podhorszki, Ray W. Grout, Matthew Wolf |
HPDC | 8 |
| 2012 | SMART-IO: SysteM-AwaRe Two-Level Data Organization for Efficient Scientific AnalyticsabstractCurrent I/O techniques have pushed the write performance close to the system peak, but they usually overlook the read side of problem. With the mounting needs of scientific discovery, it is important to provide good read performance for many common access patterns. Such demand requires an organization scheme that can effectively utilize the underlying storage system. However, the mismatch between conventional data layout on disk and common scientific access patterns leads to significant performance degradation when a subset of data is accessed. To this end, we design a system-aware Optimized Chunking model, which aims to find an optimized organization that can strike for a good balance between data transfer efficiency and processing overhead. To enable such model for scientific applications, we propose SMART-IO, a two-level data organization framework that can organize the blocks of multidimensional data efficiently. This scheme can adapt data layouts based on data characteristics and underlying storage systems, and enable efficient scientific analytics. Our experimental results demonstrate that SMART-IO can significantly improve the read performance for challenging access patterns, and speed up data analytics. For a mission critical combustion simulation code S3D, Smart-IO achieves up to 72 times speedup for planar reads of a 3-D variable compared to the logically contiguous data layout. Yuan Tian 0004, Scott Klasky, Weikuan Yu, Hasan Abbasi, Bin Wang 0019, Norbert Podhorszki, Ray W. Grout, Matthew Wolf |
MASCOTS | 8 |
| 2012 | VScope: Middleware for Troubleshooting Time-Sensitive Data Center Applications
Chengwei Wang, Infantdani Abel Rayan, Greg Eisenhauer, Karsten Schwan, Vanish Talwar, Matthew Wolf, Chad Huneycutt |
Middleware | 6 |
| 2012 | Sharing incremental approaches for adding parallelism to CS curricula (abstract only)abstractRecent industry changes, including multi-core processors, cloud computing, and GPU programming, increase the need to teach parallelism to CS undergraduates. But few CS programs can afford to add new courses or greatly alter syllabi, and the large parallelism body of knowledge relates to many courses. Participants in this BOF will share incremental approaches for adding parallelism to undergraduate CS curricula, where students study parallel computing in brief units. This networking event/ brainstorming session/ swap meet will bring together: " people with sharable parallelism expository readings, hands-on exercises, tech support ideas, etc.; "people wishing to include such materials in their courses; and" people curious about incremental approaches to teaching parallel computing. Richard A. Brown, Elizabeth Shoop, Joel Adams 0001, David P. Bunde, Jens Mache, Paul F. Steinberg, Matthew Wolf, Michael Wrinn |
SIGCSE | 7 |
| 2011 | Just in time: adding value to the IO pipelines of high performance applications with JITStagingabstractLarge scale applications are generating a tsunami of data, with understanding driven by finding information hidden within this data. The ever-increasing sizes of output, however, are making it difficult for science users to inspect the data generated by their applications, understand its important properties, and/or organize it for subsequent analysis and visualization. This paper presents JITStager, a software infrastructure with which end users can dynamically customize and thus, add value to the output pipelines of their HEC applications. JITStager is able to customize data at scale, by leveraging the computational power of both compute nodes and of additional `data staging' nodes allocated by end users. Using existing, componentized I/O interfaces to decouple the compile-time specification of the program and the run-time customization of the data pipeline, JITStager employs efficient runtime methods for binary code generation and data movement to create custom pipelines for applications' output processes that provide end users with improved insights into the data being produced, without burdening the application's computational performance and without impeding output performance. This paper describes the JITStager architecture, evaluates its performance, and demonstrates the advantages derived from its use with representative HPC applications. Hasan Abbasi, Greg Eisenhauer, Matthew Wolf, Karsten Schwan, Scott Klasky |
HPDC | 3 |
| 2011 | Six degrees of scientific data: reading patterns for extreme scale science IOabstractPetascale science simulations generate 10s of TBs of application data per day, much of it devoted to their checkpoint/restart fault tolerance mechanisms. Previous work demonstrated the importance of carefully managing such output to prevent application slowdown due to IO blocking, resource contention negatively impacting simulation performance and to fully exploit the IO bandwidth available to the petascale machine. This paper takes a further step in understanding and managing extreme-scale IO. Specifically, its evaluations seek to understand how to efficiently read data for subsequent data analysis, visualization, checkpoint restart after a failure, and other read-intensive operations. In their entirety, these actions support the 'end-to-end' needs of scientists enabling the scientific processes being undertaken. Contributions include the following. First, working with application scientists, we define 'read' benchmarks that capture the common read patterns used by analysis codes. Second, these read patterns are used to evaluate different IO techniques at scale to understand the effects of alternative data sizes and organizations in relation to the performance seen by end users. Third, defining the novel notion of a 'data district' to characterize how data is organized for reads, we experimentally compare the read performance seen with the ADIOS middleware's log-based BP format to that seen by the logically contiguous NetCDF or HDF5 formats commonly used by analysis tools. Measurements assess the performance seen across patterns and with different data sizes, organizations, and read process counts. Outcomes demonstrate that high end-to-end IO performance requires data organizations that offer flexibility in data layout and placement on parallel storage targets, including in ways that can make tradeoffs in the performance of data writes vs. reads. Jay F. Lofstead, Milo Polte, Garth A. Gibson, Scott Klasky, Karsten Schwan, Ron A. Oldfield, Matthew Wolf, Qing Liu 0002 |
HPDC | 7 |
| 2010 | PreDatA - preparatory data analytics on peta-scale machinesabstractPeta-scale scientific applications running on High End Computing (HEC) platforms can generate large volumes of data. For high performance storage and in order to be useful to science end users, such data must be organized in its layout, indexed, sorted, and otherwise manipulated for subsequent data presentation, visualization, and detailed analysis. In addition, scientists desire to gain insights into selected data characteristics `hidden' or `latent' in these massive datasets while data is being produced by simulations. PreDatA, short for Preparatory Data Analytics, is an approach to preparing and characterizing data while it is being produced by the large scale simulations running on peta-scale machines. By dedicating additional compute nodes on the machine as `staging' nodes and by staging simulations' output data through these nodes, PreDatA can exploit their computational power to perform select data manipulations with lower latency than attainable by first moving data into file systems and storage. Such intransit manipulations are supported by the PreDatA middleware through asynchronous data movement to reduce write latency, application-specific operations on streaming data that are able to discover latent data characteristics, and appropriate data reorganization and metadata annotation to speed up subsequent data access. PreDatA enhances the scalability and flexibility of the current I/O stack on HEC platforms and is useful for data pre-processing, runtime data analysis and inspection, as well as for data exchange between concurrently running simulations. Fang Zheng 0003, Hasan Abbasi, Ciprian Docan, Jay F. Lofstead, Qing Liu 0002, Scott Klasky, Manish Parashar, Norbert Podhorszki, Karsten Schwan, Matthew Wolf |
IPDPS | 10 |
| 2010 | Managing Variability in the IO Performance of Petascale Storage SystemsabstractSignificant challenges exist for achieving peak or even consistent levels of performance when using IO systems at scale. They stem from sharing IO system resources across the processes of single largescale applications and/or multiple simultaneous programs causing internal and external interference, which in turn, causes substantial reductions in IO performance. This paper presents interference effects measurements for two different file systems at multiple supercomputing sites. These measurements motivate developing a 'managed' IO approach using adaptive algorithms varying the IO system workload based on current levels and use areas. An implementation of these methods deployed for the shared, general scratch storage system on Oak Ridge National Laboratory machines achieves higher overall performance and less variability in both a typical usage environment and with artificially introduced levels of 'noise'. The latter serving to clearly delineate and illustrate potential problems arising from shared system usage and the advantages derived from actively managing it. Jay F. Lofstead, Fang Zheng 0003, Qing Liu 0002, Scott Klasky, Ron A. Oldfield, Todd Kordenbrock, Karsten Schwan, Matthew Wolf |
SC | 8 |
| 2009 | Extending I/O through high performance data servicesabstractThe complexity of HPC systems has increased the burden on the developer as applications scale to hundreds of thousands of processing cores. Moreover, additional efforts are required to achieve acceptable I/O performance, where it is important how I/O is performed, which resources are used, and where I/O functionality is deployed. Specifically, by scheduling I/O data movement and by effectively placing operators affecting data volumes or information about the data, tremendous gains can be achieved both in the performance of simulation output and in the usability of output data. Previous studies have shown the value of using asynchronous I/O, of employing a staging area, and of performing select operations on data before it is written to disk. Leveraging such insights, this paper develops and experiments with higher level I/O abstractions, termed ldquodata servicesrdquo, that manage output data from `source to sink': where/when it is captured, transported towards storage, and filtered or manipulated by service functions to improve its information content. Useful services include data reduction, data indexing, and those that manage how I/O is performed, i.e., the control aspects of data movement. Our data services implementation distinguishes control aspects - the control plane - from data movement - the data plane, so that both may be changed separably. This results in runtime flexibility not only in which services to employ, but also in where to deploy them and how they use I/O resources. The outcome is consistently high levels of I/O performance at large scale, without requiring application change. Hasan Abbasi, Jay F. Lofstead, Fang Zheng 0003, Karsten Schwan, Matthew Wolf, Scott Klasky |
CLUSTER | 5 |
| 2009 | DataStager: scalable data staging services for petascale applicationsabstractKnown challenges for petascale machines are that (1) the costs of I/O for high performance applications can be substantial, especially for output tasks like checkpointing, and (2) noise from I/O actions can inject undesirable delays into the runtimes of such codes on individual compute nodes. This paper introduces the flexible 'DataStager' framework for data staging and alternative services within that jointly address (1) and (2). Data staging services moving output data from compute nodes to staging or I/O nodes prior to storage are used to reduce I/O overheads on applications' total processing times, and explicit management of data staging offers reduced perturbation when extracting output data from a petascale machine's compute partition. Experimental evaluations of DataStager on the Cray XT machine at Oak Ridge National Laboratory establish both the necessity of intelligent data staging and the high performance of our approach, using the GTC fusion modeling code and benchmarks running on 1000+ processors. Hasan Abbasi, Matthew Wolf, Greg Eisenhauer, Scott Klasky, Karsten Schwan, Fang Zheng 0003 |
HPDC | 2 |
| 2008 | Exploiting Latent I/O Asynchrony in Petascale Science ApplicationsabstractCurrent and emerging large-scale HPC applications face daunting I/O challenges. In existing codes, problems arise both from large data volumes and from the need to perform complex online data manipulations, including data staging, reorganization, and transformation. We describe three related techniques for enabling, encouraging, and exploiting latent I/O asynchrony in HPC applications: data taps, IOgraphs, and Metabots. Mary Payne, Patrick M. Widener, Matthew Wolf, Hasan Abbasi, Scott McManus, Patrick G. Bridges, Karsten Schwan |
eScience | 3 |
| 2007 | LIVE data workspace: A flexible, dynamic and extensible platform for petascale applicationsabstractThe data needs of current and future PetaScale applications have increased over the last half decade to the extent that appropriate data management has become a crucial requirement. This concerns not only the storage of data produced by the new class of PetaScale applications, but also the data exchanges needed for coupling applications with concurrent analysis, online data visualization for validation, and others. To address such dynamic code coupling, we introduce the concept of an extensible, dynamic, and flexible data workspace, termed LIVE. In contrast to the data exchanges programmed with MPI, MPI-IO, or grid software, LIVE focuses on data exchanges carried out without a priori knowledge of potential data requirements. Examples include exchanges required by ad-hoc or dynamically determined methods for data validation, for general data analysis tasks, or for data visualization. Run on an execution environment comprised of integrated dynamic discovery and on-line management services, LIVE is used to create a dasiadata workspacepsila for a working molecular dynamics code base utilized by mechanical and materials engineers at Georgia Tech, for multi-scale materials modeling. Measurements of both this applicationpsilas data workspace and of the basic primitives in the LIVE framework demonstrate that the environmentpsilas substantial flexibility has minimal impact on overall performance, and in fact, that it improves performance in a number of usage scenarios. In particular, for a visualization pipeline example derived from our collaborators, we see a slight improvement over a solution based on MPI-IO, and a further improvement of up to 5% by utilizing LIVEpsilas ability to overlap communication with user-specified computation. Hasan Abbasi, Matthew Wolf, Karsten Schwan |
CLUSTER | 2 |
| 2007 | LIVE: : a light-weight data workspace for computational scienceabstractWe present the Lightweight Information Validation Environment, LIVE as asolution to the high complexity and data sizes of modern day computational science applications. LIVE is a data workspace that facilitates the creation of dynamic data processing overlays we call I/O graphs. We use LIVE as aplatform for dynamic extension of scientific applications using lightweight data extraction, runtime discovery and flexible data selection. Hasan Abbasi, Matthew Wolf, Karsten Schwan |
HPDC | 2 |
| 2006 | IQ-Services: network-aware middleware for interactive large-data applicationsabstractAbstract IQ‐Services are application‐specific, resource‐aware code modules executed by data transport middleware. They constitute a ‘thin’ layer between application components and the underlying computational and communication resources. This layer implements the data manipulations necessary to permit wide‐area collaborations to proceed smoothly in the presence of dynamic resource variations. IQ‐Services interact with the application and resource layers via dynamic performance attributes, and end‐to‐end implementations of such attributes also permit clients to interact with data providers. The joint middleware/resource and provider/consumer interactions implemented with performance attributes may be used to realize effective methods for managing the data flows in the large‐data, distributed Grid applications targeted by our research. Experimental results in this paper demonstrate substantial performance improvements. These are attained by coordinating network‐level with service‐level adaptations of the data being transported and by permitting end users to dynamically deploy and use application‐specific services for manipulating data in ways suitable for their current needs. Copyright © 2005 John Wiley & Sons, Ltd. Zhongtang Cai, Greg Eisenhauer, Qi He 0001, Vibhore Kumar, Karsten Schwan, Matthew Wolf |
Concurr. Comput. Pract. Exp. | 6 |
| 2005 | Service Augmentation for High End Interactive Data ServicesabstractAdvances in computational science, combined with the increasingly interdisciplinary and geographically distributed research teams, have led to a need to support multi-tiered, data- and meta-data-rich collaboration infrastructures. Our research addresses the interactive, remote tasks undertaken in such collaborations, which require a flexible software infrastructure able to dynamically deploy services where and when needed, and to provide data to clients in the forms in which they require it with suitable levels of end-to-end performance. The concept of service augmentation advanced in this paper seeks to continuously adjust the differences or degrees of incompatibility between the data received and the data displayed or stored by clients. Difference adjustments occur anywhere on the paths between data providers and clients, and compatibility computations leverage all of the resources that may be brought to bear, including CPUs and GPUs on servers and additional data manipulations on server, overlay, and client nodes. A formal structure and experimental evaluations of this concept are performed with the SmartPointer scientific visualization and annotation framework, for which we show that data-driven SLAs provide improved client flexibility and the ability to maintain application-specific notions of quality of service Matthew Wolf, Hasan Abbasi, Benjamin Collins, David Spain, Karsten Schwan |
CLUSTER | 1 |
| 2004 | XChange: coupling parallel applications in a dynamic environmentabstractModern computational science applications are becoming increasingly multidisciplinary, involving widely distributed research teams and their underlying computational platforms. A common problem for the grid applications used in these environments is the necessity to couple multiple, parallel subsystems, with examples ranging from data exchanges between cooperating, linked parallel programs, to concurrent data streaming to distributed storage engines. This work presents the XChange/sub mxn/ middleware infrastructure for coupling componentized distributed applications. XChange/sub mxn/ implements the basic functionality of well-known services like the CCA Forum's MxN project, by providing efficient data redistribution across parallel application components. Beyond such basic functionality, however, XChange/sub mxn/ also addresses two of the problems faced by wide area scientific collaborations, which are (1) the need to deal with dynamic application/component behaviors, such as dynamic arrivals and departures due to the availability of additional resources, and (2) the need to 'match' data formats across disparate application components and research teams. In response to these needs, XChange/sub mxn/ uses an anonymous publish/subscribe model for linking interacting components, and the data being exchanged is dynamically specialized and transformed to match end point requirements. The pub/sub paradigm makes it easy to deal with dynamic component arrivals and departures. Dynamic data transformation enables the 'inflight' correction of data or needs mismatches for cooperating components. This work describes the design and implementation of XChange/sub mxn/, and it evaluates its implementation compared to those of less flexible transports like MPI. It also highlights the utility ofXChange/sub mxn/'s 'inflight' data specialization, by applying it to the SmartPointer parallel data visualization environment developed at our institution. Interestingly, using XChange/sub mxn/ did not significantly affect performance but led to a reduction in the size of the code base. Hasan Abbasi, Matthew Wolf, Karsten Schwan, Greg Eisenhauer, A. Hilton |
CLUSTER | 2 |
| 2004 | StreamGen: A Workload Generation Tool for Distributed Information Flow ApplicationsabstractThis work presents the StreamGen load generator, which is targeted at distributed information flow applications. These include the event streaming services used in wide-area publish/subscribe systems or in operational information systems, the data streaming services used in remote visualization or collaboration, and the continuous data streams occurring in download services. Running across heterogeneous distributed platforms, these services are implemented by computational component that capture, manipulate, and produce information streams and are linked via overlay topologies. StreamGen can be used to produce the distributed computational and communication loads imposed by these applications. Dynamic application behaviors can be created with mathematical specifications or with behavior traces collected from application-level traces. An interesting set of traces presented in this paper is derived from long -term observations of the FTP download patterns observed at the Linux mirror site being run by the CERCS research center at the Georgia Institute of Technology. Two different flow-based applications are created and evaluated with StreamGen. The first emulates the data streaming behavior in a distributed scientific collaboration, where a scientific simulation (i.e., a molecular dynamics code) produces simulation data sent to and displayed for multiple, interactive remote users. The second emulates portions of the event-streaming behavior of an operational information system used by,a large U.S. corporation. Parametric studies with StreamGen's FTP traces applied to these applications are used to evaluate different load balancing strategies for the cluster machines manipulating these applications' data streams. Mohamed S. Mansour, Matthew Wolf, Karsten Schwan |
ICPP | 2 |
| 2004 | IQ-Services: Resource-Aware Middleware for Heterogeneous ApplicationsabstractSummary form only given. Heterogeneous computing platforms constitute a challenging execution environment for distributed applications. This article presents a 'systems' view of effective platform usage, by demonstrating the need for application software to be continuously 'aware' of the resources currently available on their underlying heterogeneous computing platforms. Our approach to the implementation of resource awareness is one that (1) provides a 'thin' middleware layer of resource aware services that permit applications to react to changes in resource availability and resources to be managed in accordance with application needs, and that (2) develops compiler- and application-level techniques for dynamic 'service morphing', the goal being to make it easy for application-level services to adjust to runtime changes in application needs or in platform resources. The specific results presented in this article are focused on large-data applications, for which the IQ-services "morphing" layer implements the data manipulations necessary to permit wide-area interactive or multimedia applications to proceed smoothly despite variations in underlying computing and network resources. Experimental results demonstrate substantial performance improvements attained by coordinating network-level with service-level adaptations of the data being transported and by permitting end users to dynamically deploy and use application-specific services for manipulating data in ways suitable for their current needs. Zhongtang Cai, Greg Eisenhauer, Christian Poellabauer, Karsten Schwan, Matthew Wolf |
IPDPS | 5 |
| 2004 | Dynamic Data Access to the GT/CERCS Linux Mirror SiteabstractSummary form only given. The purpose of this work is to better understand the dynamic data sharing behavior of certain classes of grid end users. Toward this end, we study a large-scale data repository and the end users' access behaviors to this repository. Interesting insights from the study include that (1) the use of parallel methods for data downloads, via download accelerators, is common, despite qualms expressed by the community about the impacts of such behavior on wide area data distribution networks, (2) high levels of burstiness exist for such data movements, as also observed for scientific or populist data repositories and Web sites (e.g., space imagery, sports events), and (3) a large number of remote data retrievals are by single clients for single files. We finally discuss the impact of our observations on grid applications in general. Mohamed S. Mansour, Matthew Wolf, Karsten Schwan |
IPDPS | 2 |
| 2003 | Resource-Aware Stream Management with the Customizable dproc Distributed Monitoring MechanismsabstractMonitoring the resources of distributed systems is essential to the successful deployment and execution of grid applications, particularly when such applications have well-defined QoS requirements. The dproc system-level monitoring mechanisms implemented for standard Linux kernels have several key components. First, utilizing the familiar /proc filesystem, dproc extends this interface with resource information collected from both local and remote hosts. Second, to predictably capture and distribute monitoring information, dproc uses a kernel-level group communication facility, termed KECho, which is based on events and event channels. Third and the focus of this paper is dproc's run-time customizability for resource monitoring, which includes the generation and deployment of monitoring functionality within remote operating system kernels. Using dproc, we show that: (a) data streams can be customized according to a client's resource availabilities (dynamic stream management); (b) by dynamically varying distributed monitoring (dynamic filtering of monitoring information), appropriate balance can be maintained between monitoring overheads and application quality; and (c) by performing monitoring at kernel-level, the information captured enables decision making that takes into account the multiple resources used by applications. Sandip Agarwala, Christian Poellabauer, Jiantao Kong, Karsten Schwan, Matthew Wolf |
HPDC | 5 |
| 2003 | System-Level Resource Monitoring in High-Performance Computing Environments
Sandip Agarwala, Christian Poellabauer, Jiantao Kong, Karsten Schwan, Matthew Wolf |
J. Grid Comput. | 5 |
| 2002 | SmartPointers: personalized scientific data portals in your handabstractThe SmartPointer system provides a paradigm for utilizing multiple light-weight client endpoints in a real-time scientific visualization infrastructure. Together, the client and server infrastructure form a new type of data portal for scientific computing. The clients can be used to personalize data for the needs of the individual scientist. This personalization of a shared dataset is designed to allow multiple scientists, each with their laptops or iPaqs to explore the dataset from different angles and with different personalized filters. As an example, iPaq clients can display 2D derived data functions which can be used to dynamically update and annotate the shared data space, which might be visualized separately on a large immersive display such as a CAVE. Measurements are presented for such a system, built upon the ECho middleware system developed at Georgia Tech. Matthew Wolf, Zhongtang Cai, Weiyun Huang, Karsten Schwan |
SC | 1 |