Justin M. Wozniak

dblp:03/4784 · DBLP profile ↗
← Back
44ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-2441-2048ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Benchmarking community drug response prediction models: datasets, models, tools, and metrics for cross-dataset generalization analysis
abstract
Deep learning and machine learning models have shown promise in drug response prediction (DRP), yet their ability to generalize across datasets remains an open question, raising concerns about their real-world applicability. Due to the lack of standardized benchmarking approaches, model evaluations and comparisons often rely on inconsistent datasets and evaluation criteria, making it difficult to assess true predictive capabilities. In this work, we introduce a benchmarking framework for evaluating cross-dataset prediction generalization in DRP models. Our framework incorporates five publicly available drug screening datasets, seven standardized DRP models, and a scalable workflow for systematic evaluation. To assess model generalization, we introduce a set of evaluation metrics that quantify both absolute performance (e.g. predictive accuracy across datasets) and relative performance (e.g. performance drop compared to within-dataset results), enabling a more comprehensive assessment of model transferability. Our results reveal substantial performance drops when models are tested on unseen datasets, underscoring the importance of rigorous generalization assessments. While several models demonstrate relatively strong cross-dataset generalization, no single model consistently outperforms across all datasets. Furthermore, we identify CTRPv2 as the most effective source dataset for training, yielding higher generalization scores across target datasets. By sharing this standardized evaluation framework with the community, our study aims to establish a rigorous foundation for model comparison, and accelerate the development of robust DRP models for real-world applications.
Alexander Partin, Priyanka Vasanthakumari, Oleksandr Narykov, Andreas Wilke, Natasha Koussa, Sara E. Jones, Yitan Zhu, Jamie C. Overbeek, Rajeev Jain, Gayara Demini Fernando, Cesar Sanchez-Villalobos, Cristina Garcia-Cardona, Jamaludin Mohd-Yusof, Nicholas Chia, Justin M. Wozniak, Souparno Ghosh, Ranadip Pal, Thomas S. Brettin, M. Ryan Weil, Rick L. Stevens
Briefings Bioinform.15
2026 A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva
Future Gener. Comput. Syst.25
2024 Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows
abstract
The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.
Bogdan Nicolae, Justin M. Wozniak, Tekin Bicer, Hai Nguyen 0005, Haochen Pan, Amal Gueroudji, Maxime Gonthier, Valérie Hayot-Sasson, Eliu A. Huerta, Kyle Chard, Ryan Chard, Matthieu Dorier, Nageswara S. V. Rao, Anees Al-Najjar, Alessandra Corsi, Ian T. Foster
e-Science2
2023 PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs
abstract
It is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead.
Mihael Hategan, André Merzky, Nicholson T. Collier, Ketan Maheshwari, Jonathan Ozik, Matteo Turilli, Andreas Wilke, Justin M. Wozniak, Kyle Chard, Ian T. Foster, Rafael Ferreira da Silva, Shantenu Jha, Daniel E. Laney
e-Science8
2023 An Automation Framework for Comparison of Cancer Response Models Across Configurations
abstract
Machine learning has made significant advancements in precision medicine, resulting in the development of various deep learning applications. For instance, in cancer drug response prediction, numerous deep learning models have been created. However, comparing these models across vast configurations of hyperparameters and data sets can be challenging. In this paper, we introduce a new scalable workflow suite that aims to answer questions that arise when comparing different models developed by different teams on similar or the same problems. We explain the problem in more detail and discuss our approach using near-exascale or exascale computers.
Justin M. Wozniak, Rajeev Jain, Andreas Wilke, Rylie Weaver, Alexander Partin, Thomas S. Brettin, Rick L. Stevens
e-Science1
2022 Tracking Dubious Data: Protecting Scientific Workflows from Invalidated Experiments
abstract
Provenance systems automate record keeping so that humans and/or machines can determine how a given result was obtained. In so doing, they enable a variety of reproducibility and reconstruction capabilities, while tracking the impact of older artifacts on newer ones. Large-scale scientific experiments are increasingly relying on workflows and other automation techniques to keep up with data-rates and perform on-line computation, notably training of machine learning models, and to provide rapid feedback to experimentalists. However, these workflows pose the challenges of: 1) adapting to errors in the experimental process both at the experiment site as well as in computation and 2) complex data provenance patterns that can result from the use machine learning and other methods that can arise from a feedback pattern in which initial experimental results drive the creation of new experimental parameters. The Braid Provenance Engine (Braid-DB) addresses this domain by integrating with workflow systems used in large-scale science and providing the additional capability to drive additional workflows or other automation in response to errors or other causes for elements of the workflow to be considered invalid. In this paper, we describe how Braid-DB responds to data marked as invalid, a common case in experimental science, and demonstrate its ability to retain artifacts unaffected by the invalid data.
Jim Pruyne, Justin M. Wozniak, Ian T. Foster
e-Science2
2022 A cross-study analysis of drug response prediction in cancer cell lines
abstract
To enable personalized cancer treatment, machine learning models have been developed to predict drug response as a function of tumor and drug features. However, most algorithm development efforts have relied on cross-validation within a single study to assess model accuracy. While an essential first step, cross-validation within a biological data set typically provides an overly optimistic estimate of the prediction performance on independent test sets. To provide a more rigorous assessment of model generalizability between different studies, we use machine learning to analyze five publicly available cell line-based data sets: National Cancer Institute 60, ancer Therapeutics Response Portal (CTRP), Genomics of Drug Sensitivity in Cancer, Cancer Cell Line Encyclopedia and Genentech Cell Line Screening Initiative (gCSI). Based on observed experimental variability across studies, we explore estimates of prediction upper bounds. We report performance results of a variety of machine learning models, with a multitasking deep neural network achieving the best cross-study generalizability. By multiple measures, models trained on CTRP yield the most accurate predictions on the remaining testing data, and gCSI is the most predictable among the cell line data sets included in this study. With these experiments and further simulations on partial data, two lessons emerge: (1) differences in viability assays can limit model generalizability across studies and (2) drug diversity, more than tumor diversity, is crucial for raising model generalizability in preclinical screening.
Fangfang Xia, Jonathan E. Allen, Prasanna Balaprakash, Thomas S. Brettin, Cristina Garcia-Cardona, Austin Clyde, Judith D. Cohn, James H. Doroshow, Xiaotian Duan, Veronika Dubinkina, Yvonne A. Evrard, Ya-Ju Fan, Jason Gans, Stewart He, Pinyi Lu, Sergei Maslov, Alexander Partin, Maulik Shukla, Eric A. Stahlberg, Justin M. Wozniak, Hyun Seung Yoo, George F. Zaki, Yitan Zhu, Rick L. Stevens
Briefings Bioinform.20
2021 In-situ workflow auto-tuning through combining component models
abstract
In-situ parallel workflows couple multiple component applications via streaming data transfer to avoid data exchange via shared file systems. Such workflows are challenging to configure for optimal performance due to the huge space of possible configurations. Here, we propose an in-situ workflow auto-tuning method, ALIC, which integrates machine learning techniques with knowledge of in-situ workflow structures to enable automated workflow configuration with a limited number of performance measurements. Experiments with real applications show that ALIC identify better configurations than existing methods given a computer time budget.
Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç
PPoPP3
2021 Bootstrapping in-situ workflow auto-tuning via combining performance models of component applications
abstract
In an in-situ workflow, multiple components such as simulation and analysis applications are coupled with streaming data transfers. The multiplicity of possible configurations necessitates an auto-tuner for workflow optimization. Existing auto-tuning approaches are computationally expensive because many configurations must be sampled by running the whole workflow repeatedly in order to train the auto-tuner surrogate model or otherwise explore the configuration space. To reduce these costs, we instead combine the performance models of component applications by exploiting the analytical workflow structure, selectively generating test configurations to measure and guide the training of a machine learning workflow surrogate model. Because the training can focus on well-performing configurations, the resulting surrogate model can achieve high prediction accuracy for good configurations despite training with fewer total configurations. Experiments with real applications demonstrate that our approach can identify significantly better configurations than other approaches for a fixed computer time budget.
Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç
SC3
2020 DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models
abstract
In the age of big data, deep learning has emerged as a powerful tool to extract insight and exploit its value, both in industry and scientific applications. One common pattern emerging in such applications is frequent checkpointing of the state of the learning model during training, needed in a variety of scenarios: analysis of intermediate states to explain features and correlations with training data, exploration strategies involving alternative models that share a common ancestor, knowledge transfer, resilience, etc. However, with increasing size of the learning models and popularity of distributed data-parallel training approaches, simple checkpointing techniques used so far face several limitations: low serialization performance, blocking I/O, stragglers due to the fact that only a single process is involved in checkpointing. This paper proposes a checkpointing technique specifically designed to address the aforementioned limitations, introducing efficient asynchronous techniques to hide the overhead of serialization and I/O, and distribute the load over all participating processes. Experiments with two deep learning applications (CANDLE and ResNet) on a pre-Exascale HPC platform (Theta) shows significant improvement over state-of-art, both in terms of checkpointing duration and runtime overhead.
Bogdan Nicolae, Justin M. Wozniak, George Bosilca, Matthieu Dorier, Franck Cappello
CCGRID3
2020 DeepClone: Lightweight State Replication of Deep Learning Models for Data Parallel Training
abstract
Training modern deep neural network (DNN) models involves complex workflows triggered by model exploration, sensitivity analysis, explainability, etc. A key primitive in this context is the ability to clone a model training instance, i.e. “fork” the training process in a potentially different direction, which enables comparisons of different evolution paths using variations of training data and model parameters. However, in a quest improve the training throughput, a mix of data parallel, model parallel, pipeline parallel and layer-wise parallel approaches are making the problem of cloning highly complex. In this paper, we explore the problem of efficient cloning under such circumstances. To this end, we leverage several properties of data-parallel training and layer-wise parallelism to design DeepClone, a cloning approach based on augmenting the execution graph to gain direct access to tensors, which are then sharded and reconstructed asynchronously in order to minimize runtime overhead, standby duration, readiness duration. Compared with state-of-art approaches, DeepClone shows orders of magnitude improvement for several classes of DNN models.
Bogdan Nicolae, Justin M. Wozniak, Matthieu Dorier, Franck Cappello
CLUSTER2
2019 Parsl: Pervasive Parallel Programming in Python
abstract
High-level programming languages such as Python are increasingly used to provide intuitive interfaces to libraries written in lower-level languages and for assembling applications from various components. This migration towards orchestration rather than implementation, coupled with the growing need for parallel computing (e.g., due to big data and the end of Moore's law), necessitates rethinking how parallelism is expressed in programs. Here, we present Parsl, a parallel scripting library that augments Python with simple, scalable, and flexible constructs for encoding parallelism. These constructs allow Parsl to construct a dynamic dependency graph of components that it can then execute efficiently on one or many processors. Parsl is designed for scalability, with an extensible set of executors tailored to different use cases, such as low-latency, high-throughput, or extreme-scale execution. We show, via experiments on the Blue Waters supercomputer, that Parsl executors can allow Python scripts to execute components with as little as 5 ms of overhead, scale to more than 250000 workers across more than 8000 nodes, and process upward of 1200 tasks per second. Other Parsl features simplify the construction and execution of composite programs by supporting elastic provisioning and scaling of infrastructure, fault-tolerant execution, and integrated wide-area data management. We show that these capabilities satisfy the needs of many-task, interactive, online, and machine learning applications in fields such as biology, cosmology, and materials science.
Yadu N. Babuji, Anna Woodard, Zhuozhao Li, Daniel S. Katz, Ben Clifford, Lukasz Lacinski, Ryan Chard, Justin M. Wozniak, Ian T. Foster, Michael Wilde, Kyle Chard
HPDC9
2019 Performance, Energy, and Scalability Analysis and Improvement of Parallel Cancer Deep Learning CANDLE Benchmarks
abstract
Training scientific deep learning models requires the significant compute power of high-performance computing systems. In this paper, we analyze the performance characteristics of the benchmarks from the exploratory research project CANDLE (Cancer Distributed Learning Environment) with a focus on the hyperparameters epochs, batch sizes, and learning rates. We present the parallel methodology that uses the distributed deep learning framework Horovod to parallelize the CANDLE benchmarks. We then use scaling strategies for both epochs and batch size with linear learning rate scaling to investigate how they impact the execution time and accuracy as well as the power, energy, and scalability of the parallel CANDLE benchmarks under conditions of strong scaling and weak scaling on the IBM Power9 heterogeneous system Summit at Oak Ridge National Laboratory and the Cray XC40 Theta at Argonne National Laboratory. This study provides insights into how to set the proper numbers of epochs, batch sizes, and compute resources for these benchmarks to preserve the high accuracy and to reduce the execution time of the benchmarks. We identify the data-loading performance bottleneck and then improve the performance and energy for better scalability. Results with the modified benchmarks on Summit indicate up to 78.25% in performance improvement and up to 78% in energy saving under strong scaling on up to 384 GPUs, and up to 79.5% in performance improvement and up to 77.11% in energy saving under weak scaling on up to 3,072 GPUs. On Theta, we achieve up to 45.22% performance improvement and up to 41.78% in energy saving under strong scaling on up to 384 nodes. Moreover, the modification dramatically reduces the broadcast overhead.
Xingfu Wu, Valerie Taylor 0001, Justin M. Wozniak, Rick L. Stevens, Thomas S. Brettin, Fangfang Xia
ICPP3
2019 MPI jobs within MPI jobs: A practical way of enabling task-level fault-tolerance in HPC workflows
Justin M. Wozniak, Matthieu Dorier, Robert B. Ross, Tong Shu, Tahsin M. Kurç, Li Tang 0007, Norbert Podhorszki, Matthew Wolf
Future Gener. Comput. Syst.1
2018 High-throughput cancer hypothesis testing with an integrated PhysiCell-EMEWS workflow
abstract
BACKGROUND: Cancer is a complex, multiscale dynamical system, with interactions between tumor cells and non-cancerous host systems. Therapies act on this combined cancer-host system, sometimes with unexpected results. Systematic investigation of mechanistic computational models can augment traditional laboratory and clinical studies, helping identify the factors driving a treatment's success or failure. However, given the uncertainties regarding the underlying biology, these multiscale computational models can take many potential forms, in addition to encompassing high-dimensional parameter spaces. Therefore, the exploration of these models is computationally challenging. We propose that integrating two existing technologies-one to aid the construction of multiscale agent-based models, the other developed to enhance model exploration and optimization-can provide a computational means for high-throughput hypothesis testing, and eventually, optimization. RESULTS: In this paper, we introduce a high throughput computing (HTC) framework that integrates a mechanistic 3-D multicellular simulator (PhysiCell) with an extreme-scale model exploration platform (EMEWS) to investigate high-dimensional parameter spaces. We show early results in applying PhysiCell-EMEWS to 3-D cancer immunotherapy and show insights on therapeutic failure. We describe a generalized PhysiCell-EMEWS workflow for high-throughput cancer hypothesis testing, where hundreds or thousands of mechanistic simulations are compared against data-driven error metrics to perform hypothesis optimization. CONCLUSIONS: While key notational and computational challenges remain, mechanistic agent-based models and high-throughput model exploration environments can be combined to systematically and rapidly explore key problems in cancer. These high-throughput computational experiments can improve our understanding of the underlying biology, drive future experiments, and ultimately inform clinical practice.
Jonathan Ozik, Nicholson T. Collier, Justin M. Wozniak, Charles M. Macal, Chase Cockrell, Samuel H. Friedman, Ahmadreza Ghaffarizadeh, Randy W. Heiland, Gary An, Paul Macklin
BMC Bioinform.3
2018 CANDLE/Supervisor: a workflow framework for machine learning applied to cancer research
abstract
BACKGROUND: Current multi-petaflop supercomputers are powerful systems, but present challenges when faced with problems requiring large machine learning workflows. Complex algorithms running at system scale, often with different patterns that require disparate software packages and complex data flows cause difficulties in assembling and managing large experiments on these machines. RESULTS: This paper presents a workflow system that makes progress on scaling machine learning ensembles, specifically in this first release, ensembles of deep neural networks that address problems in cancer research across the atomistic, molecular and population scales. The initial release of the application framework that we call CANDLE/Supervisor addresses the problem of hyper-parameter exploration of deep neural networks. CONCLUSIONS: Initial results demonstrating CANDLE on DOE systems at ORNL, ANL and NERSC (Titan, Theta and Cori, respectively) demonstrate both scaling and multi-platform execution.
Justin M. Wozniak, Rajeev Jain, Prasanna Balaprakash, Jonathan Ozik, Nicholson T. Collier, John Bauer, Fangfang Xia, Thomas S. Brettin, Rick L. Stevens, Jamaludin Mohd-Yusof, Cristina Garcia-Cardona, Brian Van Essen, Matt Baughman
BMC Bioinform.1
2018 Extreme-Scale Dynamic Exploration of a Distributed Agent-Based Model With the EMEWS Framework
abstract
Agent-based models (ABMs) integrate multiple scales of behavior and data to produce higher-order dynamic phenomena and are increasingly used in the study of important social complex systems in biomedicine, socio-economics and ecology/resource management. However, the development, validation and use of ABMs is hampered by the need to execute very large numbers of simulations in order to identify their behavioral properties, a challenge accentuated by the computational cost of running realistic, large-scale, potentially distributed ABM simulations. In this paper we describe the Extreme-scale Model Exploration with Swift (EMEWS) framework, which is capable of efficiently composing and executing large ensembles of simulations and other "black box" scientific applications while integrating model exploration (ME) algorithms developed with the use of widely available 3rd-party libraries written in popular languages such as R and Python. EMEWS combines novel stateful tasks with traditional run-to-completion many task computing (MTC) and solves many problems relevant to high-performance workflows, including scaling to very large numbers (millions) of tasks, maintaining state and locality information, and enabling effective multiple-language problem solving. We present the high-level programming model of the EMEWS framework and demonstrate how it is used to integrate an active learning ME algorithm to dynamically and efficiently characterize the parameter space of a large and complex, distributed Message Passing Interface (MPI) agent-based infectious disease model.
Jonathan Ozik, Nicholson T. Collier, Justin M. Wozniak, Charles M. Macal, Gary An
IEEE Trans. Comput. Soc. Syst.3
2017 Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo
Euro-Par28
2017 Report on the first workshop on negative and null results in eScience
abstract
New techniques and technologies, such as the use of large-scale computing, influence research approaches, methods, and scales and are rapidly changing the scientific landscape. Research projects in eScience ∗ thus start with many assumptions and many unknowns and are often complex. While the scientific process is sometimes viewed, at least in hindsight, as a linear progression from one good idea to the next, it is in fact fraught with false starts, wrong assumptions, and dead ends. The increasing reliance on computation adds to the scope of problems that occur. Researchers invest a significant amount of time and effort in their research. Funding agencies similarly make large investments to support such research, on the assumption that most of the research will be successful. When the research assumptions and hypotheses turn out to be false, causing results that are "negative" or "null", the natural bias is to judge that the research project "failed." The history of science, however, shows that negative results may be an opportunity to revolutionize a field of study. For example, Fleming noticed that his flu cultures were contaminated by mold, but that there was infection around that mold, leading to his discovery of Penicillin. Similarly, a project today may fail because of the misuse or failure of computational support. Such "failures" actually indicate that there is an opportunity for the cyberinfrastructure research community to improve computing resources and tools. The interaction of these modes of failure is multi-faceted. Negative results have been difficult to find in published papers in all scientific domains. We identify three reasons for this. First, negative results may not be identified as such but simply considered mistakes. Such cases may never be investigated further. Secondly, paper referees may demand a higher standard from such results, because they are more difficult to understand or challenge the conventional narrative. Third, researchers may self-select against publishing such results in light of the previous point. This paper contributes to the discussion about null or negative results in eScience. It also attempts to organize concepts about negative or null results in eScience in the form of a taxonomy. Falsifiability is the concept that a given statement can be refuted by a real-world measurement or observation 4. The empirical sciences are dominated by the construction of such statements and efforts to confirm or refute them. In eScience, such statements are rarely formally presented in a refutable manner. eScience projects typically merge goals from the physical science with computer science and engineering aspects. A failure in eScience may often be attributed to a computer engineering failure (software defects or unresolved performance shortcomings) or a collaboration misfit (the groups never came together). However, many important statements are never answered definitely, such as whether a given computational approach is effective for the physical science investigation. The formalization and confirmation/refutation of such statements have the potential to prevent efforts lost to engineering aspects. Post-mortem analysis of failed experiments provides "clues suggesting deeper lying forces," as Galison notes in How Experiments End, "Any historical reconstruction that ignores what seems in retrospect to be erroneous will be an inadequate account" 5. This means that the study of errors is not only relevant to students of history, as these "forces" can guide future investigations, suggest fundamental problems in experimental approaches, or even challenge prevailing theories. For example, relational database systems have been a well-accepted solution for information structuring, storage, and retrieval. This model is now being challenged by other database concepts, largely motivated by the need to cope with increasing data volumes. While the boundary between failure and success is not sharp in transitioning from relational to noSQL databases, the transition demonstrates a need to adapt and improve. This is often the normal path in research in computer science and cyberinfrastructure in particular, which could learn a lot from the various "failures" in eScience projects. Similarly, "short-term" examples of such "failures" are abound, such as the limit of being able to be a part of at most 16 Unix groups in NFS. It is likely that someone architecting an eScience collaboration system will face this limitation rather quickly. Are negative results as valuable as positive results in general? As Ayer points out, "What justifies scientific procedure ... is the success of the predictions to which it gives rise" 6. Following this line of thought, negative results are subordinate to the positive results that validate useful predictions. A negative result invalidates a previously held prediction, challenging or demolishing a theory or model. It is, however, incomplete. A negative result is an opportunity to pick up the pieces and fix the theory. Negative results are thus an important reminder of the limitations of science at any given point in time. "We forget about unpredictability when it is our turn to predict," Taleb says in The Black Swan 7, a book that attempts to analyze tumultuous events, including several scientific cases. Taleb makes the case that studying such cognitive upheavals is worthwhile in its own right. Professionals who act with the history of failed ideas and efforts in mind will be more resilient against similar changes in the future. Taleb describes the social and mental impact of experiencing (repeated) failure, indicating that without support, researchers can easily become demoralized and shy away from challenging, long-term problems. However, he notes that "Your finding nothing is very valuable... —hey, you know where not to look" 7. Venues such as the ERROR workshop are intended to encourage discussion of specific negative results. By co-locating with the eScience conference in Munich, the workshop attracted significant attention from this scientific community, with about 20 participants. The workshop accepted four papers out of six submitted after a peer review process. Each paper was reviewed by two to three members of the program committee, who evaluated the works based on originality, scientific rigor, significance, and presentation. The accepted papers were presented orally at the workshop, followed by a panel discussion on the topic "theory versus practice in eScience: gaps and gaping holes." The first presentation, by Gomes et al. 9, considered problems regarding interoperability between scientific workflows. Specifically, they discussed the problem of reusing workflows previously developed and implemented using one particular scientific workflow management framework with another one. To solve this problem, the authors developed an "intermediate" workflow language, with the idea that this intermediate language would preserve the workflow's semantic information across frameworks. However, they observed a loss of information about workflow semantics during the translation from a first workflow language to the "neutral" language and from the "neutral" language to the second workflow language. This happens because there are no ideal or standardized semantics for workflow languages, which is the key negative result in this research. A solution proposed to this problem is through the adoption of workflow patterns to describe richer workflow semantics. The second presentation, by Groen and Portgies Zwart 10, provided a high level overview of the authors' experience in constructing a distributed supercomputing system: CosmoGrid. The authors discussed how ambitious ideas can often be stymied by site-local resource allocation decisions. One of the negative results is the conclusion that harnessing multiple large machines is not feasible and therefore one should focus on harnessing a larger number of smaller machines. Additionally, the authors pointed out that a task as simple as getting software installed is significantly difficult at major computational sites, exposing the often overlooked reality of working with large scale computational infrastructures. The third presentation, by Cebrian et al. 11, presented an experience in designing two separate cache stores—for private and shared data— for multicore system architectures. The premise of the work was to improve efficiency by excluding private and shared read-only cache contents from coherency management. From the experiments and analysis of results obtained with this approach, the authors concluded that systems are less efficient with this kind of design, which is a negative result. This is because the overhead of classification mechanisms and increased concentration of access to shared data cause a bandwidth bottleneck to a particular portion of cache, resulting in higher latencies. The fourth presentation, by Jackson et al. 12, discussed an experience of performing an experiment in the context of a larger body of work. The experiment was on latency measurement between nodes within a single cluster, as well as across different clusters. Some of the main takeaways from the experiments as described by the authors were the technical and administrative obstacles faced when the experiment involves dependencies on several independently managed computational systems across administrative boundaries. Despite these obstacles, the authors were able to produce a significant body of latency data. One finding from this dataset was that an exhaustive study of latencies among systems was not necessarily, by itself, a good predictor of actual application performance. The topic of discussion for the panel session was "theory versus practice in eScience: gaps and gaping holes." This theme emerged from a recurring observation in the submitted papers, in which negative or null results are attributed to a mismatch between expectations, which are based on theory, and what is found in reality. The panelists were Daniel S. Katz, Simon Portegies Zwart, Kyle Chard, Juan M. Cebrián, and Gary Jackson. Each panelist spoke for 2–3 min, and then there was an open discussion between panelists and the audience. The rest of this subsection presents the highlights of the panelists talks and the discussion that followed. Jackson spoke about the importance of not losing research focus because of infrastructure complexities and problems—the proverbial "missing the forest for trees." Chard said that there are no good definitions of eScience, although we provide one taken from the eScience conference series website in Section 2. Furthermore, he raised questions as to how the scientific process, which had been mostly unchanged for hundreds of years and has long review cycles, has recently been changing with online data publication and open access journals. Katz spoke of the phenomenon of failures among research projects and endeavors by quoting from the opening of Tolstoy's Anna Karenina, "Happy families are all alike; every unhappy family is unhappy in its own way." He noted that successful research endeavors must have all their critical factors right to be successful and failing even one of them could jeopardize the complete project. This is popularly known as Anna Karenina Principle 13. Katz also emphasized the importance and value of scientific results in general and negative results in particular with the question/statement: "How do we decide if there is value in a result?" Portgies Zwart spoke about the importance of understanding the difference between core computer science and other sciences, as well as the scientists associated with each one. He argued that computer science is currently undergoing a crisis because it is hard to find interesting problems, because of competition with the industry. In particular, computer science is challenged by reproducibility. One solution, he suggested, is an establishment of a software museum to prevent loss of software. The open discussion that followed focused on diverse topics such as software preservation, publication and credit, training, and the definition of negative or null results. It began with participants expressing concerns about issues related to software in particular. Scenarios were discussed that introduce the "gaps between eScience theory and practice" connecting technologies, ideas, and people. One gap is that digital products in general and software in particular, including methods and knowledge (algorithms), can be lost over time, sometimes known as bit rot 14. Software hosting services such as GitHub can address this problem to a certain extent by preserving the files, but they still require much human effort to preserve the function delivered by the software as meaningful. A curation service for algorithms could be another solution, requiring additional effort. Can the software and algorithm hosting services be linked as concepts and implementation? Preservation and publication of negative results are a challenge. While there are no technical barriers, from a publishing culture point of view, there are few or no incentives for publishing negative results. In the presence of such incentives, people would develop the culture about explaining not only what they did but also why they did not do so in some other way. And, if the negative results are actually published, it is likely that the same approach will not be taken by other researchers and groups. Another identified gap is the lack of a comprehensive understanding of negative results because of the lack of a conceptual framework, for example, a taxonomy. The discussion also raised the gap introduced by a lack of a credit model for discovering, identifying, and reporting negative results. For example, can negative results and/or methods to obtain them be patented? For instance, who receives credit if a succession of graduate students working on a problem arrive at a negative result followed by positive result? Or what happens to the positive results obtained before further investigation leads to their negation and nullification? Another gap arises from the lack of proper eScience training of domain scientists—can domain scientists be trained to become eScientists? This also applies to principal investigators, many of whom were trained in an era when science was carried out differently than it is carried out today. Training imparting the knowledge of modern computational methods and capabilities could play a key role in filling such a gap. The difference between incomplete (such as obtained from samples of insufficient size) and negative results can also be unclear. This can often result in negative results that are subject to interpretation. For instance, in an MD simulation 14, it cannot be shown if sampling was sufficient. In the same vein, should incremental competitive results be considered negative? In general, there is no standard on how many simulated timesteps are needed to obtain the correct answer, although sometimes, one can validate against lab experiments. Another topic raised during the discussion compared research in academia versus science in the commercial sector. One prominent sentiment expressed in the discussion was that, in some areas, research done in the commercial domain is "ahead" of research in academia, particularly, where industry has larger-scale problems and data than academia. One possible reason for this could be that there are more negative and null results in academia compared to industry. However, academic research can be transferred to the commercial domain and vice versa. For instance, the patent system is in place to enable commercial contribution to public research. One question that arises here is should there be a distinction between commercial and academic research? How will this distinction manifest itself? It was noted that there is an asymmetry between positive and negative results: With negative results, it is more likely that some error was made. It may be harder to truly prove a negative result. For instance an "existence proof" is sufficient for a positive result while a "non-possible proof" is needed for negative results. The workshop led to concrete outcomes before, during, and as a follow-up of its realization. After the announcement of the workshop, the Mozilla Science Foundation hosted a guest post about the workshop by the organizers 15. The post discussed the importance of the theme of "negative" results and the goals for the workshop. One of the outcomes of the panel discussion was the call for a taxonomy of negative and null results in eScience. We respond to this call in this paper by proposing a taxonomy in Section 5. In Figure 1, we present a taxonomy of eScience results. The three kinds of results in eScience are positive, null, and negative. In the taxonomy, we focus on negative and null results. Negative results may be caused by one or more of the following reasons: technological, technical, human, and domain. For instance, an erroneous result obtained because of insufficient precision resulting from a limitation of a system library is an example of negative result caused by technical and technological limitations. Similarly, a simulation algorithm resulting from a flawed understanding of a natural phenomenon could be considered a negative result caused by domain and human factors. A mismatch between the problem/data size and the technology/methodology used is an example of a technological reason. Examples of technical causes include software bugs and cyberinfrastructure faults. Human causes include both incidental issues such as mistakes in measurements and systemic issues such as a false hypothesis or an insufficient sample size. Null results are obtained because of the lack of discriminating conditions to confirm or refute a hypothesis. Such situation may be caused by similar reasons as for negative results. An example of technological/technical reason is a statistical test has poor performance on the data because the implementation uses limited precision. Null results can also have a human cause, when insufficient samples are used in the experiments or when some bias in the data goes unnoticed. We are aware of two workshops with similar themes in related fields (Information and Communication Technologies). The first is NoISE (Workshop on Negative or Inconclusive Results in Semantic Web) 19. The second is NOPE (Workshop on Negative Outcomes, Post-mortems, and Experiences) 20. Both the workshops were organized for the first time in 2015. Similarly to the current special issue, there have been two special issues in the prominent journals focused on negative or null results: the Journal on Negative results in Empirical Software Engineering 21 and PLOS ONE Collection 22. The PLOS ONE Collection focuses on inconclusive results as a distinct type of negative results in addition to null results. Similarly to the workshops, both special issues were launched for the first time in 2015. In this section, we discuss some of the key implications that are drawn from the previous sections. These are the issues that are directly impacted by the occurrence of negative and null results in the research as conducted by the scientific community. We classify these implications into four categories: technological, technical, cultural, and domain specific. Publication, credit, and citation of the work that has yielded negative results are an important consideration from the research community point of view. Citations and credit are important measures of success for a research publication. Given the current trends of publishing positive results, it is a crucial decision for a researcher to invest efforts in publishing a negative result. Technical issues such as hardware faults and software bugs often go undetected until late in the research work. In these cases, the negative results are not necessarily of the same nature as the science domain unless the domain is computer science itself. It becomes difficult for a domain scientist to draw value from the publication and dissemination of such results. As a consequence, they often are ignored or fixed after the results were obtained and disseminated. For example, a bug in the third party library call up the toolchain of an application that limited the results of the actual science in scale or precision can be considered a negative result. In some cases, the problems are mismatched to the available computational infrastructure. Sometimes, the problems are too small for a given environment, leading to an inefficient use. In others cases, the problems are too large for the infrastructure, leading to generation of incomplete or no results. Publishing details about such cases can benefit the community by allowing it to better understand how to more optimally match problems and solutions. Finally, a cumulative effect from more than one cause is possible. One of the biggest concern about such issues is that they often go unnoticed by the larger community and hence the appropriate correction measures are not adapted. We are grateful to the program committee members, the panelists, and the authors who supported the realization of this first workshop. We also thank the reviewers of this special issue. The work by Katz was supported in part by the National Science Foundation while working at the Foundation. Any opinion, finding, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.
Ketan Maheshwari, Daniel S. Katz, Sílvia Delgado Olabarriaga, Justin M. Wozniak, Douglas Thain
Concurr. Comput. Pract. Exp.4
2017 Experimental evaluation of a flexible I/O architecture for accelerating workflow engines in ultrascale environments
Francisco Rodrigo Duro, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001, Justin M. Wozniak, Robert B. Ross
Parallel Comput.5
2016 Flexible Data-Aware Scheduling for Workflows over an In-memory Object Store
abstract
This paper explores novel techniques for improving the performance of many-task workflows based on the Swift scripting language. We propose novel programmer options for automated distributed data placement and task scheduling. These options trigger a data placement mechanism used for distributing intermediate workflow data over the servers of Hercules, a distributed key-value store that can be used to cache file system data. We demonstrate that these new mechanisms can significantly improve the aggregated throughput of many-task workflows with up to 86x, reduce the contention on the shared file system, exploit the data locality, and trade off locality and load balance.
Francisco Rodrigo Duro, Francisco Javier García Blas, Florin Isaila, Justin M. Wozniak, Jesús Carretero 0001, Robert B. Ross
CCGrid4
2015 Toward Interlanguage Parallel Scripting for Distributed-Memory Scientific Computing
abstract
Scripting languages such as Python and R have been widely adopted as tools for the productive development of scientific software because of the power and expressiveness of the languages and available libraries. However, deploying scripted applications on large-scale parallel computer systems such as the IBM Blue Gene/Q or Cray XE6 is a challenge because of issues including operating system limitations, interoperability challenges, parallel filesystem overheads due to the small file system accesses common in scripted approaches, and other issues. We present here a new approach to these problems in which the Swift scripting system is used to integrate high-level scripts written in Python, R, and Tcl, with native code developed in C, C++, and Fortran, by linking Swift to the library interfaces to the script interpreters. In this approach, Swift handles data management, movement, and marshaling among distributed-memory processes without direct user manipulation of low-level communication libraries such as MPI. We present a technique to efficiently launch scripted applications on large-scale supercomputers using a hierarchical programming model.
Justin M. Wozniak, Timothy G. Armstrong, Ketan Maheshwari, Daniel S. Katz, Michael Wilde, Ian T. Foster
CLUSTER1
2015 Porting Ordinary Applications to Blue Gene/Q Supercomputers
abstract
Efficiently porting ordinary applications to Blue Gene/Q supercomputers is a significant challenge. Codes are often originally developed without considering advanced architectures and related tool chains. Science needs frequently lead users to want to run large numbers of relatively small jobs (often called many-task computing, an ensemble, or a workflow), which can conflict with supercomputer configurations. In this paper, we discuss techniques developed to execute ordinary applications over leadership class supercomputers. We use the high-performance Swift parallel scripting framework and build two workflow execution techniques -- sub-jobs and main-wrap. The sub-jobs technique, built on top of the IBM Blue Gene/Q resource manager Cobalt's sub-block jobs, lets users submit multiple, independent, repeated smaller jobs within a single larger resource block. The main-wrap technique is a scheme that enables C/C++ programs to be defined as functions that are wrapped by a high-performance Swift wrapper and that are invoked as a Swift script. We discuss the needs, benefits, technicalities, and current limitations of these techniques. We further discuss the real-world science enabled by these techniques and the results obtained.
Ketan Maheshwari, Justin M. Wozniak, Timothy G. Armstrong, Daniel S. Katz, T. Andrew Binkowski, Xiaoliang Zhong, Olle Heinonen, Dmitry Karpeyev, Michael Wilde
e-Science2
2014 Compiler Optimization for Extreme-Scale Scripting
abstract
The data-driven task parallelism execution model can support parallel programming models that are well suited for large-scale distributed-memory parallel computing, for example, simulations and analysis pipelines running on clusters and clouds. We describe a novel compiler intermediate representation and optimizations for this execution model, including adaptions of standard techniques alongside novel techniques. These techniques are applied to Swift/T, a high-level scripting language for flexible data flow composition of functions, which may be serial or use lower-level parallel programming models such as MPI and OpenMP. This paper presents preliminary results, indicating that our compiler optimizations reduce communication overhead by 70% to 93% on distributed-memory systems.
Timothy G. Armstrong, Justin M. Wozniak, Michael Wilde, Ian T. Foster
CCGRID2
2014 Design and evaluation of the gemtc framework for GPU-enabled many-task computing
abstract
We present the design and first performance and usability evaluation of GeMTC, a novel execution model and runtime system that enables accelerators to be programmed with many concurrent and independent tasks of potentially short or variable duration. With GeMTC, a broad class of such "many-task" applications can leverage the increasing number of accelerated and hybrid high-end computing systems. GeMTC overcomes the obstacles to using GPUs in a many-task manner by scheduling and launching independent tasks on hardware designed for SIMD-style vector processing. We demonstrate the use of a high-level MTC programming model (the Swift parallel dataflow language) to run tasks on many accelerators and thus provide a high-productivity programming model for the growing number of supercomputers that are accelerator-enabled. While still in an experimental stage, GeMTC can already support tasks of fine (subsecond) granularity and execute concurrent heterogeneous tasks on 86,000 independent GPU warps spanning 2.7M GPU threads on the Blue Waters supercomputer.
Scott J. Krieder, Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Benjamin Grimmer, Ian T. Foster, Ioan Raicu
HPDC2
2014 Compiler Techniques for Massively Scalable Implicit Task Parallelism
abstract
Swift/T is a high-level language for writing concise, deterministic scripts that compose serial or parallel codes implemented in lower-level programming models into large-scale parallel applications. It executes using a data-driven task parallel execution model that is capable of orchestrating millions of concurrently executing asynchronous tasks on homogeneous or heterogeneous resources. Producing code that executes efficiently at this scale requires sophisticated compiler transformations: poorly optimized code inhibits scaling with excessive synchronization and communication. We present a comprehensive set of compiler techniques for data-driven task parallelism, including novel compiler optimizations and intermediate representations. We report application benchmark studies, including unbalanced tree search and simulated annealing, and demonstrate that our techniques greatly reduce communication overhead and enable extreme scalability, distributing up to 612 million dynamically load balanced tasks per second at scales of up to 262,144 cores without explicit parallelism, synchronization, or load balancing in application code.
Timothy G. Armstrong, Justin M. Wozniak, Michael Wilde, Ian T. Foster
SC2
2013 Evaluating Cloud Computing Techniques for Smart Power Grid Design Using Parallel Scripting
abstract
Applications used to evaluate next-generation electrical power grids(``smart grids'') are anticipated to be compute and data-intensive. In this work, we parallelize and improve performance of one such application which was run sequentially prior to the use of our cloud-based configuration. We examine multiple cloud computing offerings, both commercial and academic, to evaluate their potential for improving the turnaround time for application results. Since the target application does not fit well into existing computational paradigms for the cloud, we employ parallel scripting tool, as a first step toward a broader program of adapting portable, scalable computational tools for use as enablers of the future smart grids. We use multiple clouds as a way to reassure potential users that the risk of cloud-vendor lock-in can be managed. This paper discusses our methods and results. Our experience sheds light on some of the issues facing computational scientists and engineers tasked with adapting new paradigms and infrastructures for existing engineering design problems.
Ketan Maheshwari, Kenneth P. Birman, Justin M. Wozniak, Devin Van Zandt
CCGRID3
2013 Swift/T: Large-Scale Application Composition via Distributed-Memory Dataflow Processing
abstract
Many scientific applications are conceptually built up from independent component tasks as a parameter study, optimization, or other search. Large batches of these tasks may be executed on high-end computing systems, however, the coordination of the independent processes, their data, and their data dependencies is a significant scalability challenge. Many problems must be addressed, including load balancing, data distribution, notifications, concurrent programming, and linking to existing codes. In this work, we present Swift/T, a programming language and runtime that enables the rapid development of highly concurrent, task-parallel applications. Swift/Tis composed of several enabling technologies to address scalability challenges, offers a high-level optimizing compiler for user programming and debugging, and provides tools for binding user code in C/C++/Fortran into a logical script. In this work, we describe the Swift/T solution and present scaling results from the IBM Blue Gene/Pand Blue Gene/Q.
Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster
CCGRID1
2013 Enabling multi-task computation on Galaxy-based gateways using swift
abstract
The Galaxy science portal is a popular gateway to data analysis and computational tools for a broad range of life sciences communities. While Galaxy enables users to overcome the complexities of integrating diverse tools into unified workflows, it has only limited capabilities to execute those tools on the parallel and often distributed high-performance resources that the life sciences fields increasingly requires. We outline here an approach to meet this pressing requirement with the Swift parallel scripting language and its distributed runtime system. Swift's model of computation - implicitly parallel functional dataflow - is an elemental abstraction to which the core computing model of Galaxy maps very closely. We describe an integration between Galaxy and Swift that is transforming Galaxy into a much more powerful science gateway, retaining its user-friendly nature while extending its power to execute highly scalable workflows on diverse parallel environments.
Ketan Maheshwari, Alexis A. Rodriguez, David Kelly, Ravi K. Madduri, Justin M. Wozniak, Michael Wilde, Ian T. Foster
CLUSTER5
2013 MTC envelope: defining the capability of large scale computers in the context of parallel scripting applications
Zhao Zhang 0007, Daniel S. Katz, Michael Wilde, Justin M. Wozniak, Ian T. Foster
HPDC4
2013 Swift/T: scalable data flow programming for many-task applications
abstract
Swift/T, a novel programming language implementation for highly scalable data flow programs, is presented.
Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster
PPoPP1
2013 Dataflow coordination of data-parallel tasks via MPI 3.0
abstract
Scientific applications are often complex collections of many large-scale tasks. Mature tools exist for describing task-parallel workflows consisting of serial tasks, and a variety of tools exist for programming a single data-parallel operation. However, few tools cover the intersection of these two models. In this work, we extend the load balancing library ADLB to support parallel tasks. We demonstrate how applications can easily be composed of parallel tasks using Swift dataflow scripts, which are compiled to ADLB programs with performance comparable to hand-coded equivalents. By combining this framework with data-parallel analysis libraries, we are able to dynamically execute many instances of a parallel data analysis application in support of a parameter exploration workload.
Justin M. Wozniak, Tom Peterka, Timothy G. Armstrong, James Dinan, Ewing L. Lusk, Michael Wilde, Ian T. Foster
EuroMPI1
2013 Parallelizing the execution of sequential scripts
abstract
Scripting is often used in science to create applications via the composition of existing programs. Parallel scripting systems allow the creation of such applications, but each system introduces the need to adopt a somewhat specialized programming model. We present an alternative scripting approach, AMFS Shell, that lets programmers express parallel scripting applications via minor extensions to existing sequential scripting languages, such as Bash, and then execute them in-memory on large-scale computers. We define a small set of commands between the scripts and a parallel scripting runtime system, so that programmers can compose their scripts in a familiar scripting language. The underlying AMFS implements both collective (fast file movement) and functional (transformation based on content) file management. Tasks are handled by AMFS's built-in execution engine. AMFS Shell is expressive enough for a wide range of applications, and the framework can run such applications efficiently on large-scale computers.
Zhao Zhang 0007, Daniel S. Katz, Timothy G. Armstrong, Justin M. Wozniak, Ian T. Foster
SC4
2013 Turbine: A Distributed-memory Dataflow Engine for High Performance Many-task Applications
abstract
Efficiently utilizing the rapidly increasing concurrency of multi-petaflop computing systems is a significant programming challenge. One approach is to structure applications with an upper layer of many loosely coupled coarse-grained tasks, each comp
Justin M. Wozniak, Timothy G. Armstrong, Ketan Maheshwari, Ewing L. Lusk, Daniel S. Katz, Michael Wilde, Ian T. Foster
Fundam. Informaticae1
2013 JETS: Language and System Support for Many-Parallel-Task Workflows
Justin M. Wozniak, Michael Wilde, Daniel S. Katz
J. Grid Comput.1
2012 AESOP: Expressing Concurrency in High-Performance System Software
abstract
High-performance computing (HPC) and distributed systems rely on a diverse collection of system soft-ware to provide application services, including file systems, schedulers, and web services. Such system software services must manage highly concurrent requests, interact with a wide range of resources, and scale well in order to be successful. Unfortunately, no single programming model for distributed system software currently offers optimal performance and productivity for all these tasks. While numerous libraries, languages, and language extensions have been developed in recent years to simplify parallel computation, they do not address the challenges of distributed system software in which concurrency control involves a variety of hardware and network devices, not just computational resources. In this work we present AESOP, a new programming language and programming model designed to implement distributed system software with high development productivity and run-time efficiency. AESOP is a superset of the C language that describes blocks of code to be executed concurrently without dictating whether that concurrency will be provided by a threading, event, or other model. This decoupling enables system software to adjust to different architectures, device APIs, and workloads without any change to the core algorithm implementation. AESOP also provides additional language constructs to simplify common system software development tasks. We evaluate AESOP by implementing a basic file server and comparing its performance, memory efficiency, and developer productivity with several thread-based and event-based implementations. AESOP is shown to provide competitive performance to traditional distributed system software development models while at the same time reducing code complexity and enhancing developer productivity.
Dries Kimpe, Philip H. Carns, Kevin Harms, Justin M. Wozniak, Samuel Lang, Robert B. Ross
NAS4
2012 Design and analysis of data management in scalable parallel scripting
abstract
We seek to enable efficient large-scale parallel execution of applications in which a shared filesystem abstraction is used to couple many tasks. Such parallel scripting (many-task computing, MTC) applications suffer poor performance and utilization on large parallel computers because of the volume of filesystem I/O and a lack of appropriate optimizations in the shared filesystem. Thus, we design and implement a scalable MTC data management system that uses aggregated compute node local storage for more efficient data movement strategies. We co-design the data management system with the data-aware scheduler to enable dataflow pattern identification and automatic optimization. The framework reduces the time to solution of parallel stages of an astronomy data analysis application, Montage, by 83.2% on 512 cores; decreases the time to solution of a seismology application, CyberShake, by 7.9% on 2,048 cores; and delivers BLAST performance better than mpiBLAST at various scales up to 32,768 cores, while preserving the flexibility of the original BLAST application.
Zhao Zhang 0007, Daniel S. Katz, Justin M. Wozniak, Allan Espinosa, Ian T. Foster
SC3
2011 Swift: A language for distributed parallel scripting
Michael Wilde, Mihael Hategan, Justin M. Wozniak, Ben Clifford, Daniel S. Katz, Ian T. Foster
Parallel Comput.3
2008 Making the best of a bad situation: Prioritized storage management in GEMS
Justin M. Wozniak, Paul R. Brenner, Douglas Thain, Aaron Striegel, Jesús A. Izaguirre
Future Gener. Comput. Syst.1
2008 Biomolecular committor probability calculation enabled by processing in network storage
Paul R. Brenner, Justin M. Wozniak, Douglas Thain, Aaron Striegel, Jeffrey W. Peng, Jesús A. Izaguirre
Parallel Comput.2
2007 Overdrive Controllers for Distributed Scientific Computation
abstract
Large distributed computer systems have been successfully employed to solve modern scientific problems that were previously impracticable. Tools exist to bind together notably underutilized high speed computers and sizable storage resources, creating architectures that provide performance and capacity that scales with the quantity of resources invested. In various forms such as Internet computing, desktop grids, and even "big iron " grids, users and administrators find it difficult, however, to obtain aggregate systems that are more reliable than their underlying components and simple to utilize in concert. In this work, we will propose a model for controlling complex distributed systems and its application to the construction of scientific repositories.
Justin M. Wozniak
CCGRID1
2007 Biomolecular Path Sampling Enabled by Processing in Network Storage
abstract
Computationally complex and data intensive atomic scale biomolecular simulation is enabled via processing in network storage (PINS): a novel distributed system framework to overcome bandwidth, compute, storage, and security challenges inherent to the wide area computation and storage grid. High throughput data generation requirements for our scientific target are overcome through novel aggregate bandwidth capabilities. Biomolecular simulation methods are correlated with the client tools, hybrid database/file server (GEMS), computation engine (Condor), virtual file system adapter (Parrot), and local file servers (Chirp). PINS performance is reported for the path sampling of a solvated protein domain requiring over 1000 simulations with total output data generation on the order of 1TB.
Paul R. Brenner, Justin M. Wozniak, Douglas Thain, Aaron Striegel, Jeffrey W. Peng, Jesús A. Izaguirre
IPDPS2
2005 Generosity and gluttony in GEMS: grid enabled molecular simulations
abstract
Biomolecular simulations produce more output data than can be managed effectively by traditional computing systems. Researchers need distributed systems that allow the pooling of resources, the sharing of simulation data, and the reliable publication of both tentative and final results. To address this need, we have designed GEMS, a system that enables biomolecular researchers to store, search, and share large scale simulation data. The primary design problem is striking a balance between generosity and gluttony. On one hand, storage providers wish to be generous and share resources with their collaborators. On the other hand, an unchecked data producer can be gluttonous and easily replicate data unnecessarily until it fills all available space. To balance generosity and gluttony, GEMS allows both storage providers and data producers to state and enforce policies on the consumption of storage and the replication of data. By taking advantage of known properties of simulation data, the system is able to distinguish between high value final results that must be preserved and low value intermediate results that can be deleted and regenerated if necessary. We have built a prototype of GEMS on a cluster of workstations and demonstrate its ability to store new data, to replicate within policy limits, and to recover from failures.
Justin M. Wozniak, Paul R. Brenner, Douglas Thain, Aaron Striegel, Jesús A. Izaguirre
HPDC1
2005 Separating Abstractions from Resources in a Tactical Storage System
abstract
Sharing data and storage space in a distributed system remains a difficult task for ordinary users, who are constrained to the fixed abstractions and resources provided by administrators. To remedy this situation, we introduce the concept of a tactical storage system (TSS) that separates storage abstractions from storage resources, leaving users free to create, reconfigure, and destroy abstractions as their needs change. In this paper, we describe how a TSS can provide a variety of filesystem and database abstractions for unmodified applications without requiring special privileges or kernel changes. A TSS provides performance competitive with NFS for single clients and also scales well for multiple servers and multiple clients. A prototype TSS of 120 disks and 6 TB of storage has been deployed at the University of Notre Dame and used for applications in high energy physics and bioinformatics.
Douglas Thain, Sander Klous, Justin M. Wozniak, Paul R. Brenner, Aaron Striegel, Jesús A. Izaguirre
SC3