EDBT 2026 Demo / reviewers in the wild / expert
Rafael Ferreira da Silva
dblp:08/10040
· DBLP profile ↗
61ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0002-1720-0928ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 9 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 13 since 2021Software engineering, systems software and programming languages · 19 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging paradigms: Designing for HPC-Quantum convergence
Amir Shehata, Peter Groszkowski, Thomas J. Naughton, Muralikrishnan Gopalakrishnan Meena, Daniel Claudino, Rafael Ferreira da Silva, Thomas L. Beck |
Future Gener. Comput. Syst. | 7 |
| 2026 | A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva |
Future Gener. Comput. Syst. | 26 |
| 2025 | PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic WorkflowsabstractLarge Language Models (LLMs) and other foundation models are increasingly used as the core of AI agents. In agentic workflows, these agents plan tasks, interact with humans and peers, and influence scientific outcomes across federated and heterogeneous environments. However, agents can hallucinate or reason incorrectly, propagating errors when one agent’s output becomes another’s input. Thus, assuring that agents’ actions are transparent, traceable, reproducible, and reliable is critical to assess hallucination risks and mitigate their workflow impacts. While provenance techniques have long supported these principles, existing methods fail to capture and relate agent-centric metadata such as prompts, responses, and decisions with the broader workflow context and downstream outcomes. In this paper, we introduce PROV-AGENT, a provenance model that extends W3C PROV and leverages the Model Context Protocol (MCP) and data observability to integrate agent interactions into end-to-end workflow provenance. Our contributions include: (1) a provenance model tailored for agentic workflows, (2) a near real-time, open-source system for capturing agentic provenance, and (3) a cross-facility evaluation spanning edge, cloud, and HPC environments, demonstrating support for critical provenance queries and agent reliability analysis. Renan Souza 0001, Amal Gueroudji, Stephen DeWitt, Daniel Rosendo, Tirthankar Ghosal, Robert B. Ross, Prasanna Balaprakash, Rafael Ferreira da Silva |
eScience | 8 |
| 2025 | ControlA: Agentic Workflow Control Mechanisms for Reliable ScienceabstractAI-driven scientific discovery has emerged as a transformative fifth paradigm in research, with agentic AI playing an increasingly prominent role across scientific domains. Agentic AI can enable collaborative AI-human or even fully autonomous decision-making, but it also introduces significant reliability challenges due to the dynamic and evolutionary nature of the AI agents. Specifically, foundation model-powered agents are prone to generating hallucinated, misleading, or adversarial outputs that can propagate silently through workflows and corrupt downstream results. In this paper we present a conceptual framework for a unified approach that integrates agentic workflow-level instrumentation and agent-level safeguards to enhance the reliability of the wider system, particularly critical in science. Embedding these mechanisms into a provenance-augmented infrastructure enables early detection, containment, and recovery from erroneous behavior, ultimately enhancing reliability and reproducibility in AI-assisted scientific workflows. Amal Gueroudji, Tanwi Mallick, Renan Souza 0001, Rafael Ferreira da Silva, Robert B. Ross, Matthieu Dorier, Philip H. Carns, Kyle Chard, Ian T. Foster |
eScience | 4 |
| 2024 | Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AIabstractAs Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions. Renan Souza 0001, Silvina Caíno-Lores, Mark Coletti, Tyler J. Skluzacek, Alexandru Costan, Frédéric Suter, Marta Mattoso, Rafael Ferreira da Silva |
e-Science | 8 |
| 2024 | Integrating quantum computing resources into scientific HPC ecosystems
Thomas L. Beck, Alessandro Baroni 0003, Ryan S. Bennink, Gilles Buchs, Eduardo Antonio Coello Pérez, Markus Eisenbach 0002, Rafael Ferreira da Silva, Muralikrishnan Gopalakrishnan Meena, Kalyana C. Gottiparthi, Peter Groszkowski, Travis S. Humble, Ryan Landfield, Ketan Maheshwari, Sarp Oral, Michael A. Sandoval, Amir Shehata, In-Saeng Suh, Christopher Zimmer 0001 |
Future Gener. Comput. Syst. | 7 |
| 2024 | An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems
Jesse McDonald, John Dobbs, Yick Ching Wong, Rafael Ferreira da Silva, Henri Casanova |
Future Gener. Comput. Syst. | 4 |
| 2023 | WfCommons: Data Collection and Runtime Experiments using Multiple Workflow SystemsabstractScientific workflows have become ubiquitous across scientific fields, and their execution methods and systems continue to be the subject of research and development. Most experimental evaluations of these workflows rely on workflow instances, which can be either real-world or synthetic, to ensure relevance to current application domains or explore hypothetical/future scenarios. The WfCommons project addresses this need by providing data and tools to support such evaluations. In this paper, we present an overview of WfCommons and describe two recent developments. Firstly, we introduce a workflow execution "tracer" for Nextflow, which significantly enhances the set of real-world instances available in WfCommons. Secondly, we describe a workflow instance "translator" that enables the execution of any real-world or synthetic WfCommons workflow instance using Dask. Our contributions aim to provide researchers and practitioners with more comprehensive resources for evaluating scientific workflows. Henri Casanova, Kyle Berney, Serge Chastel, Rafael Ferreira da Silva |
COMPSAC | 4 |
| 2023 | PSI/J: A Portable Interface for Submitting, Monitoring, and Managing JobsabstractIt is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead. Mihael Hategan, André Merzky, Nicholson T. Collier, Ketan Maheshwari, Jonathan Ozik, Matteo Turilli, Andreas Wilke, Justin M. Wozniak, Kyle Chard, Ian T. Foster, Rafael Ferreira da Silva, Shantenu Jha, Daniel E. Laney |
e-Science | 11 |
| 2023 | Message from the IEEE eScience 2023 Conference Leadership eScience 2023abstractThe 19th IEEE Conference on eScience (eScience 2023), which took place from October 9th to 13th in Limassol (Cyprus), provided a platform where researchers, developers, and users of eScience applications and enabling IT technologies delved into the realms of interdisciplinary collaboration. The mission of the eScience conference was to drive innovation in data- and compute-intensive research, spanning a wide array of disciplines, including the physical and biological sciences, as well as the social sciences, arts, and humanities. IEEE eScience 2023 successfully dismantled traditional barriers, nurturing collaboration among interdisciplinary research communities, developers, and eScience application users. Throughout that five-day event, the conference focus was on advancing all aspects of eScience and its associated technologies, applications, algorithms, and tools, with a strong emphasis on practical solutions to real-world challenges. New Directions and Communities: This year, we have introduced four key topics: “Computational Science for sustainable development,” “FAIR,” “Research Infrastructures for eScience,” and “Continuum Computing: Convergence between Cloud Computing and the Internet of Things (IoT).” These tracks enriched our discussions and investigations. Keynote Speakers: In addition to these exciting tracks, we were delighted to present our distinguished keynote speakers: Dr. İlkay Altıntaş illuminated the convergence of machine learning, AI, and scientific research, showcasing innovative approaches and real-world applications. Professor Ian T. Foster took us on a journey into the realm of global science services, demonstrating how they can reshape collaborative research. Finally, Professor Paul Watson shared insights from the National Innovation Centre for Data, highlighting successful projects that leverage data science and AI to drive impact. Program Overview: As we embarked on this eScience journey, we invited participants to engage, collaborate, and explore the transformative potential of eScience and its associated technologies. Together, we addressed practical solutions, open challenges, and pushed the boundaries of interdisciplinary research. This year's conference featured three keynotes from different domains, presented 40 peer-reviewed papers, showcased 23 posters, featured 15 invited talks from renowned researchers, hosted 6 workshops, and provided 5 tutorials. The conference was held in person and was co-located with the 4th Global Research Platform Workshop (4GRP). Our sincere appreciation goes out to the authors who submitted exceptional papers, and we are immensely thankful for the dedicated efforts of our numerous volunteers. The Technical Paper Committee diligently reviewed 96 papers and 25 posters, with over 70 volunteers conducting 308 reviews. This collaborative endeavour has resulted in the exceptional collection of papers featured in this volume. We would also like to express our deep gratitude to the IEEE Computer Society and IEEE's Technical Committee on High-Performance Computing (TCHPC) for their unwavering sponsorship and support, which have been instrumental in making this conference possible. Finally, we extend our gratitude to you, our valued readers, for your interest in this volume. We are confident that the contents within will greatly contribute to the advancement of your research, whether it pertains to eScience or other domains. George Angelos Papadopoulos, Rafael Ferreira da Silva, Rosa Filgueira |
e-Science | 2 |
| 2023 | Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data ObservabilityabstractModern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer. Renan Souza 0001, Tyler J. Skluzacek, Sean R. Wilkinson, Maxim A. Ziatdinov, Rafael Ferreira da Silva |
e-Science | 5 |
| 2023 | Driving Next-Generation Workflows from the Data PlaneabstractWe observe the emergence of a new generation of scientific workflows that process data produced at a sustained rate by scientific instruments and large scale numerical simulations. This data is consumed by multiple analysis, visualization, or Machine Learning components not only to enable inference and justify the scientific program, but also to monitor and steer the evolution of these experiments. In such workflows, moving intermediate data efficiently is key to performance, more than efficiently scheduling computational tasks. However, most traditional workflow management systems focus on optimizing task scheduling and then deal with data management, assuming a “move little, compute for long” model, which makes them unfit to the efficient management of this new generation of workflows. Therefore, we advocate for a new way to manage scientific workflows. We propose to consider an efficiently and independently managed data plane that can store and stream data. Workflows compute components, in the application plane can then interact with the data plane, abstracted from complexities of data management. Then, the role of a workflow management system would become that of a control plane that allows users to connect services together to execute the workflow and manages connections between the application and data planes. In this position paper, we characterize several next-generation workflow motifs and describe how their interaction with the data plane is a challenge to traditional workflow management systems. Then, we express a set of requirements that a workflow management system should meet to efficiently manage next-generation workflows at different scales. Based on these requirements, we expose our vision of driving next-generation workflows from the data plane and list remaining open challenges. Frédéric Suter, Rafael Ferreira da Silva, Ana Gainaru, Scott Klasky |
e-Science | 2 |
| 2023 | Performance assessment of ensembles of in situ workflows under resource constraintsabstractSummary Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long‐running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble‐based simulations paired with in situ analyses using in‐memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co‐location scenarios, our proposed performance indicators demonstrate benefits of co‐locating simulation and coupled analyses within a compute node. Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Silvina Caíno-Lores, Michela Taufer, Ewa Deelman |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Automated generation of scientific workflow generators with WfChef
Tainã Coleman, Henri Casanova, Rafael Ferreira da Silva |
Future Gener. Comput. Syst. | 3 |
| 2022 | Pseudonymization at Scale: OLCF's Summit Usage Data Case StudyabstractThe analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available. Ketan Maheshwari, Sean R. Wilkinson, Alex May 0002, Tyler J. Skluzacek, Olga A. Kuchar, Rafael Ferreira da Silva |
IEEE Big Data | 6 |
| 2022 | SIM-SITU: A Framework for the Faithful Simulation of in situ ProcessingabstractThe amount of data generated by numerical simulations in various scientific domains led to a fundamental redesign of how the analysis and visualization of simulation outputs are performed. The throughput and capacity of storage subsystems have not evolved as fast as the computing power in extreme-scale supercomputers, making the classical post-hoc approach highly inefficient. In situ processing has then emerged as a solution in which simulation and data analysis/visualization are intertwined for better performance and greater interactivity. Determining the best allocation, i.e., how many resources to allocate to simulation and analysis respectively, mapping, i.e., where and at which frequency to run the analysis/visualization, and data transfer mode is a complex task whose performance assessment is crucial to the efficient execution of in situ processing. However, such a performance evaluation of different strategies usually relies either on directly running them on the targeted execution environments, which can rapidly become extremely time- and resource-consuming, or on resorting to simplified models of the components of an in situ application, which can lack of realism. In both cases, the validity of the performance evaluation is limited. In this paper, we present SIM-SITU, a framework for the faithful performance evaluation of in situ processing strategies. We designed SIM-SITU to reflect the typical features of in situ processing systems. Thanks to its modular design, Sim-Situ has the necessary flexibility to easily and faithfully evaluate the behavior and performance of various allocation, mapping, and data transfer strategies. We illustrate the capabilities of SIM-SITU on a Molecular Dynamics use case. We study the impact of different strategies on performance and show how users can leverage SIM-SITU to determine interesting tradeoffs when adding analysis/visualization components to their application. Valentin Honoré, Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Ewa Deelman, Frédéric Suter |
e-Science | 4 |
| 2022 | On the Feasibility of Simulation-Driven Portfolio Scheduling for Cyberinfrastructure Runtime Systems
Henri Casanova, Yick Ching Wong, Loïc Pottier, Rafael Ferreira da Silva |
JSSPP | 4 |
| 2022 | WfCommons: A framework for enabling scientific workflow research and development
Tainã Coleman, Henri Casanova, Loïc Pottier, Manav Kaushik, Ewa Deelman, Rafael Ferreira da Silva |
Future Gener. Comput. Syst. | 6 |
| 2021 | Modeling the Linux page cache for accurate simulation of data-intensive applicationsabstractThe emergence of Big Data in recent years has resulted in a growing need for efficient data processing solutions. While infrastructures with sufficient compute power are available, the I/O bottleneck remains. The Linux page cache is an efficient approach to reduce I/O overheads, but few experimental studies of its interactions with Big Data applications exist, partly due to limitations of real-world experiments. Simulation is a popular approach to address these issues, however, existing simulation frameworks do not simulate page caching fully, or even at all. As a result, simulation-based performance studies of data-intensive applications can lead to misleading results and inaccurate conclusions.In this paper, we propose an I/O simulation model that captures the key features of the Linux page cache. We have implemented this model as part of the WRENCH workflow simulation framework, which itself builds on the popular Sim-Grid distributed systems simulation framework. Our model and its implementation enable the simulation of both single-threaded and multithreaded applications, and of both writeback and writethrough caches for local or network-based filesystems. We evaluate the accuracy of our model in different conditions, including sequential and concurrent applications, as well as local and remote I/Os. We find that our page cache model reduces the simulation error by up to an order of magnitude when compared to state-of-the-art, cacheless simulations. Our model is publicly available in the WRENCH framework, making it usable in a wide range of simulation studies. Hoang-Dung Do, Valérie Hayot-Sasson, Rafael Ferreira da Silva, Christopher Steele, Henri Casanova, Tristan Glatard |
CLUSTER | 3 |
| 2021 | A Roadmap to Robust Science for High-throughput Applications: The Developers' PerspectiveabstractScientists using the high-throughput computing (HTC) paradigm for scientific discovery rely on complex software systems and heterogeneous architectures that must deliver robust science (i.e., ensuring performance scalability in space and time; trust in technology, people, and infrastructures; and reproducible or confirmable research). Developers must overcome a variety of obstacles to pursue workflow interoperability, identify tools and libraries for robust science, port codes across different architectures, and establish trust in non-deterministic results. This poster presents recommendations to build a roadmap to overcome these challenges and enable robust science for HTC applications and workflows. The findings were collected from an international community of software developers during a Virtual World Cafe in May 2021. Michela Taufer, Ewa Deelman, Rafael Ferreira da Silva, Trilce Estrada, Mary W. Hall, Miron Livny |
CLUSTER | 3 |
| 2021 | Serverless Containers - Rising Viable Approach to Scientific WorkflowsabstractThe increasing popularity of the serverless computing approach has led to the emergence of new cloud infrastructures working in Container-as-a-Service (CaaS) model like AWS Fargate, Google Cloud Run, or Azure Container Instances. New infrastructures facilitate an innovative approach to running cloud containers where developers are freed from managing underlying resources. In this paper, we focus on evaluating the capabilities of elastic containers and their usefulness for scientific computing in the scientific workflow paradigm using AWS Fargate and Google Cloud Run infrastructures. For the experimental evaluation of our approach, we extended the HyperFlow engine to support these CaaS platforms, together with adapting four scientific workflows composed of several dozen to hundreds of tasks organized into a dependency graph. Studied applications are used to create cost-performance benchmarks and flow execution plots, delay, elasticity, and scalability measurements. Results show that serverless containers can be successfully utilized for running scientific workflows. Moreover, the results allow for gaining insight into the specific advantages and limits of the studied platforms. Krzysztof Burkat, Maciej Pawlik, Bartosz Balis, Maciej Malawski, Karan Vahi, Mats Rynge, Rafael Ferreira da Silva, Ewa Deelman |
e-Science | 7 |
| 2021 | WfChef: Automated Generation of Accurate Scientific Workflow GeneratorsabstractScientific workflow applications have become mainstream and their automated and efficient execution on large-scale compute platforms is the object of extensive research and development. For these efforts to be successful, a solid experimental methodology is needed to evaluate workflow algorithms and systems. A foundation for this methodology is the availability of realistic workflow instances. Dozens of workflow instances for a few scientific applications are available in public repositories. While these are invaluable, they are limited: workflow instances are not available for all application scales of interest. To address this limitation, previous work has developed generators of synthetic, but representative, workflow instances of arbitrary scales. These generators are popular, but implementing them is a manual, labor-intensive process that requires expert application knowledge. As a result, these generators only target a handful of applications, even though hundreds of applications use workflows in production.In this work, we present WfChef, a framework that fully automates the process of constructing a synthetic workflow generator for any scientific application. Based on an input set of workflow instances, WfChef automatically produces a synthetic workflow generator. We define and evaluate several metrics for quantifying the realism of the generated workflows. Using these metrics, we compare the realism of the workflows generated by WfChef generators to that of the workflows generated by the previously available, hand-crafted generators. We find that the WfChef generators not only require zero development effort (because it is automatically produced), but also generate workflows that are more realistic than those generated by hand-crafted generators. Tainã Coleman, Henri Casanova, Rafael Ferreira da Silva |
e-Science | 3 |
| 2021 | A Roadmap to Robust Science for High-throughput Applications: The Scientists' PerspectiveabstractThis poster presents our first steps to define a roadmap to robust science for high-throughput applications used in scientific discovery. These applications combine multiple components into increasingly complex multi-modal workflows that are often executed in concert on heterogeneous systems. The increasing complexity hinders the ability of scientists to generate robust science (i.e., ensuring performance scalability in space and time; trust in technology, people, and infrastructures; and reproducible or confirmable research). Scientists must withstand and overcome adverse conditions such as heterogeneous and unreliable architectures at all scales (including extreme scale), rigorous testing under uncertainties, unexplainable algorithms in machine learning, and black-box methods. This poster presents findings and recommendations to build a roadmap to overcome these challenges and enable robust science. The data was collected from an international community of scientists during a virtual world café in February 2021. Michela Taufer, Ewa Deelman, Rafael Ferreira da Silva, Trilce Estrada, Mary W. Hall |
e-Science | 3 |
| 2021 | GLUME: A Strategy for Reducing Workflow Execution Times on Batch-Scheduled Platforms
Evan Hataishi, Pierre-François Dutot, Rafael Ferreira da Silva, Henri Casanova |
JSSPP | 3 |
| 2021 | End-to-end online performance data capture and analysis for scientific workflows
George Papadimitriou 0002, Cong Wang 0014, Karan Vahi, Rafael Ferreira da Silva, Anirban Mandal, Zhengchun Liu, Rajiv Mayani, Mats Rynge, Mariam Kiran, Vickie E. Lynch, Rajkumar Kettimuthu, Ewa Deelman, Jeffrey S. Vetter, Ian T. Foster |
Future Gener. Comput. Syst. | 4 |
| 2021 | Special issue on workflows in support of large-scale science
Rafael Ferreira da Silva, Sandra Gesing, Rizos Sakellariou, Ian J. Taylor |
Future Gener. Comput. Syst. | 1 |
| 2021 | Teaching parallel and distributed computing concepts in simulation with WRENCH
Henri Casanova, Ryan Tanaka, William Koch, Rafael Ferreira da Silva |
J. Parallel Distributed Comput. | 4 |
| 2021 | Artificial Intelligence for Modeling Complex Systems: Taming the Complexity of Expert Models to Improve Decision MakingabstractMajor societal and environmental challenges involve complex systems that have diverse multi-scale interacting processes. Consider, for example, how droughts and water reserves affect crop production and how agriculture and industrial needs affect water quality and availability. Preventive measures, such as delaying planting dates and adopting new agricultural practices in response to changing weather patterns, can reduce the damage caused by natural processes. Understanding how these natural and human processes affect one another allows forecasting the effects of undesirable situations and study interventions to take preventive measures. For many of these processes, there are expert models that incorporate state-of-the-art theories and knowledge to quantify a system's response to a diversity of conditions. A major challenge for efficient modeling is the diversity of modeling approaches across disciplines and the wide variety of data sources available only in formats that require complex conversions. Using expert models for particular problems requires integration of models with third-party data as well as integration of models across disciplines. Modelers face significant heterogeneity that requires resolving semantic, spatiotemporal, and execution mismatches, which are largely done by hand today and may take more than 2 years of effort. We are developing a modeling framework that uses artificial intelligence (AI) techniques to reduce modeling effort while ensuring utility for decision making. Our work to date makes several innovative contributions: (1) an intelligent user interface that guides analysts to frame their modeling problem and assists them by suggesting relevant choices and automating steps along the way; (2) semantic metadata for models, including their modeling variables and constraints, that ensures model relevance and proper use for a given decision-making problem; and (3) semantic representations of datasets in terms of modeling variables that enable automated data selection and data transformations. This framework is implemented in the MINT (Model INTegration) framework, and currently includes data and models to analyze the interactions between natural and human systems involving climate, water availability, agricultural production, and markets. Our work to date demonstrates the utility of AI techniques to accelerate modeling to support decision-making and uncovers several challenging directions for future work. Yolanda Gil, Daniel Garijo, Deborah Khider, Craig A. Knoblock, Varun Ratnakar, Maximiliano Osorio, Hernán Vargas, Minh Pham 0004, Jay Pujara, Basel Shbita, Yao-Yi Chiang, Dan Feldman, Yijun Lin 0001, Hayley Song, Vipin Kumar 0001, Ankush Khandelwal, Michael S. Steinbach, Kshitij Tayal, Shaoming Xu, Suzanne A. Pierce, Lissa Pearson, Daniel Hardesty-Lewis, Ewa Deelman, Rafael Ferreira da Silva, Rajiv Mayani, Armen R. Kemanian, Lorne Leonard, Scott D. Peckham, Maria Stoica 0001, Kelly M. Cobourn, Zeya Zhang, Christopher J. Duffy, Lele Shu |
ACM Trans. Interact. Intell. Syst. | 25 |
| 2020 | Modeling the Performance of Scientific Workflow Executions on HPC Platforms with Burst BuffersabstractScientific domains ranging from bioinformatics to astronomy and earth science rely on traditional high-performance computing (HPC) codes, often encapsulated in scientific workflows. In contrast to traditional HPC codes that employ a few programming and runtime approaches that are highly optimized for HPC platforms, scientific workflows are not necessarily optimized for these platforms. As an effort to reduce the gap between compute and I/O performance, HPC platforms have adopted intermediate storage layers known as burst buffers. A burst buffer (BB) is a fast storage layer positioned between the global parallel file system and the compute nodes. Two designs currently exist: (i) shared, where the BBs are located on dedicated nodes; and (ii) on-node, in which each compute node embeds a private BB. In this paper, using accurate simulations and realworld experiments, we study how to best use these new storage layers when executing scientific workflows. These applications are not necessarily optimized to run on HPC systems, and thus can exhibit I/O patterns that differ from that of HPC codes. Thus, we first characterize the I/O behaviors of a real-world workflow under different configuration scenarios on two leadership-class HPC systems (Cori at NERSC and Summit at ORNL). Then, we use these characterizations to calibrate a simulator for workflow executions on HPC systems featuring shared and private BBs. Last, we evaluate our approach against a large I/O-intensive workflow, and we provide insights on the performance levels and the potential limitations of these two BBs architectures. Loïc Pottier, Rafael Ferreira da Silva, Henri Casanova, Ewa Deelman |
CLUSTER | 2 |
| 2020 | Developing accurate and scalable simulators of production workflow management systems with WRENCH
Henri Casanova, Rafael Ferreira da Silva, Ryan Tanaka, Suraj Pandey, Gautam Jethwani, William Koch, Spencer Albrecht, James Oeth, Frédéric Suter |
Future Gener. Comput. Syst. | 2 |
| 2019 | Exploration of Workflow Management Systems Emerging Features from Users PerspectivesabstractThere has been a recent emergence of new workflow applications focused on data analytics and machine learning. This emergence has precipitated a change in the workflow management landscape, causing the development of new dataoriented workflow management systems (WMSs) in addition to the earlier standard of task-oriented WMSs. In this paper, we summarize three general workflow use-cases and explore the unique requirements of each use-case in order to understand how WMSs from both workflow management models meet the requirements of each workflow use-case from the user’s perspective. We analyze the applicability of the two models by carefully describing each model and by providing an examination of the different variations of WMSs that fall under the task driven model. To illustrate the strengths and weaknesses of each workflow management model, we summarize the key features of four production-ready WMSs: Pegasus, Makeflow, Apache Airflow, and Pachyderm. To deepen our analysis of the four WMSs examined in this paper,we implement three real-world use-cases to highlight the specifications and features of each WMS. We present our final assessment of each WMS after considering the following factors: usability, performance, ease of deployment, and relevance. The purpose of this work is to offer insights from the user’s perspective into the research challenges that WMSs currently face due to the evolving workflow landscape. Ryan Mitchell, Loïc Pottier, Steve Jacobs, Rafael Ferreira da Silva, Mats Rynge, Karan Vahi, Ewa Deelman |
IEEE BigData | 4 |
| 2019 | Empowering Agroecosystem Modeling with HTC Scientific Workflows: The Cycles Model Use CaseabstractScientific workflows have enabled large-scale scientific computations and data analysis, and lowered the entry barrier for performing computations in distributed heterogeneous platforms (e.g., HTC and HPC). In spite of impressive achievements to date, large-scale modeling, simulation, and data analytics in the long-tail still face several challenges such as efficient scheduling and execution of large-scale workflows (O(106)) with very short-running tasks (few seconds). While the current trend to support next-generation workflows on leadership class machines have gained much attention in the past years, at the other end of the spectrum scientific workflows from the long-tail science have become larger and require processing massive volumes of data. In this paper, we report on our experience in designing and implementing an HTC workflow for agroecosystem modeling. We leverage well-known (task clustering and co-scheduling) and emerging (hierarchical workflows and containers) workflow optimization techniques to make the workflow planning problem tractable, and maximize resource utilization and the degree of task parallelism. Experimental results, via the implementation of a use case, show that by strategically combining the above strategies and defining an appropriate set of optimization parameters, the overall workflow makespan can be improved by 3.5 orders of magnitude when compared to a regular (non-optimized) execution of the workflow. Rafael Ferreira da Silva, Rajiv Mayani, Armen R. Kemanian, Mats Rynge, Ewa Deelman |
IEEE BigData | 1 |
| 2019 | Bridging Concepts and Practice in eScience via Simulation-Driven EngineeringabstractThe CyberInfrastructure (CI) has been the object of intensive research and development in the last decade, resulting in a rich set of abstractions and interoperable software implementations that are used in production today for supporting ongoing and breakthrough scientific discoveries. A key challenge is the development of tools and application execution frameworks that are robust in current and emerging CI configurations, and that can anticipate the needs of upcoming CI applications. This paper presents WRENCH, a framework that enables simulation-driven engineering for evaluating and developing CI application execution frameworks. WRENCH provides a set of high-level simulation abstractions that serve as building blocks for developing custom simulators. These abstractions rely on the scalable and accurate simulation models that are provided by the SimGrid simulation framework. Consequently, WRENCH makes it possible to build, with minimum software development effort, simulators that that can accurately and scalably simulate a wide spectrum of large and complex CI scenarios. These simulators can then be used to evaluate and/or compare alternate platform, system, and algorithm designs, so as to drive the development of CI solutions for current and emerging applications. Rafael Ferreira da Silva, Henri Casanova, Ryan Tanaka, Frédéric Suter |
eScience | 1 |
| 2019 | Characterizing In Situ and In Transit Analytics of Molecular Dynamics Simulations for Next-Generation SupercomputersabstractMolecular Dynamics (MD) simulations executed on state-of-the-art supercomputers are producing data at rates faster than it can be written out to disk. In situ and in transit analysis of data generated by MD simulations reduce the original volume of information by several orders of magnitude, thereby alleviating the negative impact of I/O bottlenecks. This work focuses on characterizing the impact of in situ and in transit analytics on the overall MD workflow performance, and the capability for capturing rapid, rare events in the simulated molecular system. The MD simulation and analysis processes share data via remote direct memory access (RDMA) using DataSpaces. Our metrics of interest are time spent waiting in I/O by the MD simulation, lost frames of the MD simulation, and idle time of the analysis. We measure these metrics for a diverse set of molecular systems and characterize their trends for in situ and in transit configurations. We then model which frames are dropped and which ones are analyzed for a real use case. The insights gained from this study are generally applicable for in situ and in transit workflows that require optimization of parameters to minimize loss in workflow performance and analytic accuracy. Michela Taufer, Ewa Deelman, Michael R. Wyatt II, Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Harel Weinstein, Michel A. Cuendet, Trilce Estrada |
eScience | 7 |
| 2019 | Custom Execution Environments with Containers in Pegasus-Enabled Scientific WorkflowsabstractScience reproducibility is a cornerstone feature in scientific workflows. In most cases, this has been implemented as a way to exactly reproduce the computational steps taken to reach the final results. While these steps are often completely described, including the input parameters, datasets, and codes, the environment in which these steps are executed is only described at a higher level with endpoints and operating system name and versions. Though this may be sufficient for reproducibility in the short term, systems evolve and are replaced over time, breaking the underlying workflow reproducibility. A natural solution to this problem is containers, as they are well defined, have a lifetime independent of the underlying system, and can be user-controlled so that they can provide custom environments if needed. This paper highlights some unique challenges that may arise when using containers in distributed scientific workflows. Further, this paper explores how the Pegasus Workflow Management System implements container support to address such challenges. Karan Vahi, Michael Zink, Mats Rynge, George Papadimitriou 0002, Duncan A. Brown, Rajiv Mayani, Rafael Ferreira da Silva, Ewa Deelman, Anirban Mandal, Eric Lyons 0001 |
eScience | 7 |
| 2019 | Collaborative circuit designs using the CRAFT repository
Adam Brinckman, Ewa Deelman, Sandeep Gupta 0001, Jarek Nabrzyski, Soowang Park, Rafael Ferreira da Silva, Ian J. Taylor, Karan Vahi |
Future Gener. Comput. Syst. | 6 |
| 2019 | Measuring the impact of burst buffers on data-intensive scientific workflows
Rafael Ferreira da Silva, Scott Callaghan, Tu Mai Anh Do, George Papadimitriou 0002, Ewa Deelman |
Future Gener. Comput. Syst. | 1 |
| 2019 | Using simple PID-inspired controllers for online resilient resource management of distributed scientific workflows
Rafael Ferreira da Silva, Rosa Filgueira, Ewa Deelman, Erola Pairo-Castineira, Ian Michael Overton, Malcolm P. Atkinson 0001 |
Future Gener. Comput. Syst. | 1 |
| 2018 | A Job Sizing Strategy for High-Throughput Scientific WorkflowsabstractThe user of a computing facility must make a critical decision when submitting jobs for execution: how many resources (such as cores, memory, and disk) should be requested for each job? If the request is too small, the job may fail due to resource exhaustion; if the request is too large, the job may succeed, but resources will be wasted. This decision is especially important when running hundreds of thousands of jobs in a high throughput workflow, which may exhibit complex, long tailed distributions of resource consumption. In this paper, we present a strategy for solving the job sizing problem: (1) applications are monitored and measured in user-space as they run; (2) the resource usage is collected into an online archive; and (3) jobs are automatically sized according to historical data in order to maximize throughput or minimize waste. We evaluate the solution analytically, and present case studies of applying the technique to high throughput physics and bioinformatics workflows consisting of hundreds of thousands of jobs, demonstrating an increase in throughput of 10-400 percent compared to naive approaches. Benjamín Tovar, Rafael Ferreira da Silva, Gideon Juve, Ewa Deelman, William E. Allcock, Douglas Thain, Miron Livny |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Toward Prioritization of Data Flows for Scientific Workflows Using Virtual Software Defined ExchangesabstractRecent advances in cloud systems, on-demand circuits and software-defined networking have created new opportunities to enable complex, data-intensive scientific applications to run on dynamic networked cloud infrastructures. In this work, we present an end-to-end framework for autonomic adaptation for scientific workflows on networked cloud systems, which leverages novel network provisioning technologies. We present an application-independent controller framework called Mobius++ that includes dynamic network adaptation capabilities using Software-Defined Networking (SDN) mechanisms, which enables workflow management systems to address competing priorities of workflow operations, data movements in particular. We use a representative, data-intensive bioinformatics workflow as a driving use case to showcase the above capabilities. Experimental results show that the Mobius++ framework, in conjunction with a novel virtual Software Defined Exchange (SDX) platform, is able to dynamically prioritize bandwidths between different end-points, on-demand, and being driven by priority directives from a workflow management system. We show that data transfer jobs from two workflows with different priorities are accurately arbitrated as the relative priorities change. Anirban Mandal, Paul Ruth, Ilya Baldin, Rafael Ferreira da Silva, Ewa Deelman |
eScience | 4 |
| 2017 | Software architectures to integrate workflow engines in science gatewaysabstractScience gateways often rely on workflow engines to execute applications on distributed infrastructures. We investigate six software architectures commonly used to integrate workflow engines into science gateways. In tight integration, the workflow engine shares software components with the science gateway. In service invocation, the engine is isolated and invoked through a specific software interface. In task encapsulation, the engine is wrapped as a computing task executed on the infrastructure. In the pool model, the engine is bundled in an agent that connects to a central pool to fetch and execute workflows. In nested workflows, the engine is integrated as a child process of another engine. In workflow conversion, the engine is integrated through workflow language conversion. We describe and evaluate these architectures with metrics for assessment of integration complexity, robustness, extensibility, scalability and functionality. Tight integration and task encapsulation are the easiest to integrate and the most robust. Extensibility is equivalent in most architectures. The pool model is the most scalable one and meta-workflows are only available in nested workflows and workflow conversion. These results provide insights for science gateway architects and developers. Tristan Glatard, Marc-Etienne Rousseau, Sorina Camarasu-Pop, Reza Adalat, Natacha Beck, Samir Das, Rafael Ferreira da Silva, Najmeh Khalili-Mahani, Vladimir Korkhov, Pierre-Olivier Quirion, Pierre Rioux, Sílvia Delgado Olabarriaga, Lune Bellec, Alan C. Evans |
Future Gener. Comput. Syst. | 7 |
| 2017 | Reproducibility of execution environments in computational science using Semantics and Clouds
Idafen Santana-Pérez, Rafael Ferreira da Silva, Mats Rynge, Ewa Deelman, María S. Pérez 0001, Óscar Corcho |
Future Gener. Comput. Syst. | 2 |
| 2017 | A characterization of workflow management systems for extreme-scale applications
Rafael Ferreira da Silva, Rosa Filgueira, Ilia Pietri, Ming Jiang 0005, Rizos Sakellariou, Ewa Deelman |
Future Gener. Comput. Syst. | 1 |
| 2016 | Automating environmental computing applications with scientific workflowsabstractComputational environmental science applications have evolved and become more complex over the last decade. In order to cope with the needs of such applications, computational methods and technologies have emerged to support the execution of these applications on heterogeneous, distributed systems. Among them are workflow management systems such as Pegasus. Pegasus is being used by researchers to model seismic wave propagation, to discover new celestial objects, to study RNA critical to human brain development, and to investigate other important research questions. This paper provides an introduction to scientific workflows and describes Pegasus and its main features. The paper highlights how the environmental science community has used Pegasus to automate their scientific workflow executions on high performance and high throughput computing systems by presenting three use cases: two Earth science workflows, and a climate science workflow. Rafael Ferreira da Silva, Ewa Deelman, Rosa Filgueira, Karan Vahi, Mats Rynge, Rajiv Mayani, Benjamin Mayer |
eScience | 1 |
| 2016 | Science automation in practice: Performance data farming in workflowsabstractThis paper describes an approach to conduct large-scale parameter studies, where each data point in the study requires the execution of a whole scientific workflow. We show how a parameter studies system can be integrated with a workflow management system to seamlessly execute a large number of workflows, each with different input parameter values using large-scale computing infrastructure. The work is motivated by a need to collect performance-related data to conduct a sensitivity analysis in the context of relation between workflow input parameters and the performance of tasks in the workflow developed for the Spallation Neutron Source facility at the Oak Ridge National Laboratory. Dariusz Król 0002, Jacek Kitowski, Rafael Ferreira da Silva, Gideon Juve, Karan Vahi, Mats Rynge, Ewa Deelman |
ETFA | 3 |
| 2016 | Consecutive Job Submission Behavior at Mira SupercomputerabstractUnderstanding user behavior is crucial for the evaluation of scheduling and allocation performances in HPC environments. This paper aims to further understand the dynamic user reaction to different levels of system performance by performing a comprehensive analysis of user behavior in recorded data in the form of delays in the subsequent job submission behavior. Therefore, we characterize a workload trace covering one year of job submissions from the Mira supercomputer at ALCF (Argonne Leadership Computing Facility). We perform an in-depth analysis of correlations between job characteristics, system performance metrics, and the subsequent user behavior. Analysis results show that the user behavior is significantly influenced by long waiting times, and that complex jobs (number of nodes and CPU hours) lead to longer delays in subsequent job submissions. Stephan Schlagkamp, Rafael Ferreira da Silva, William E. Allcock, Ewa Deelman, Uwe Schwiegelshohn |
HPDC | 2 |
| 2016 | Dynamic and Fault-Tolerant Clustering for Scientific WorkflowsabstractTask clustering has proven to be an effective method to reduce execution overhead and to improve the computational granularity of scientific workflow tasks executing on distributed resources. However, a job composed of multiple tasks may have a higher risk of suffering from failures than a single task job. In this paper, we conduct a theoretical analysis of the impact of transient failures on the runtime performance of scientific workflow executions. We propose a general task failure modeling framework that uses a maximum likelihood estimation-based parameter estimation process to model workflow performance. We further propose three fault-tolerant clustering strategies to improve the runtime performance of workflow executions in faulty execution environments. Experimental results show that failures can have significant impact on executions where task clustering policies are not fault-tolerant, and that our solutions yield makespan improvements in such scenarios. In addition, we propose a dynamic task clustering strategy to optimize the workflow's makespan by dynamically adjusting the clustering granularity when failures arise. A trace-based simulation of five real workflows shows that our dynamic method is able to adapt to unexpected behaviors, and yields better makespans when compared to static methods. Weiwei Chen 0002, Rafael Ferreira da Silva, Ewa Deelman, Thomas Fahringer |
IEEE Trans. Cloud Comput. | 2 |
| 2015 | Practical Resource Monitoring for Robust High Throughput ComputingabstractRobust high throughput computing requires effective monitoring and enforcement of a variety of resources including CPU cores, memory, disk, and network traffic. Without effective monitoring and enforcement, it is easy to overload machines, causing failures and slowdowns, or underutilize machines, which results in wasted opportunities. This paper explores how to describe, measure, and enforce resources used by computational tasks. We focus on tasks running in distributed execution systems, in which a task requests the resources it needs, and the execution system ensures the availability of such resources. This presents two non-trivial problems: how to measure the resources consumed by a task, and how to monitor and report resource exhaustion in a robust and timely manner. For both of these tasks, operating systems have a variety of mechanisms with different degrees of availability, accuracy, overhead, and intrusiveness. We describe various forms of monitoring and the available mechanisms in contemporary operating systems. We then present two specific monitoring tools that choose different tradeoffs in overhead and accuracy, and evaluate them on a selection of benchmarks. Gideon Juve, Benjamín Tovar, Rafael Ferreira da Silva, Dariusz Król 0002, Douglas Thain, Ewa Deelman, William E. Allcock, Miron Livny |
CLUSTER | 3 |
| 2015 | Using imbalance metrics to optimize task clustering in scientific workflow executions
Weiwei Chen 0002, Rafael Ferreira da Silva, Ewa Deelman, Rizos Sakellariou |
Future Gener. Comput. Syst. | 2 |
| 2015 | Pegasus, a workflow management system for science automation
Ewa Deelman, Karan Vahi, Gideon Juve, Mats Rynge, Scott Callaghan, Philip Maechling, Rajiv Mayani, Weiwei Chen 0002, Rafael Ferreira da Silva, Miron Livny, R. Kent Wenger |
Future Gener. Comput. Syst. | 9 |
| 2014 | Community Resources for Enabling Research in Distributed Scientific WorkflowsabstractA significant amount of recent research in scientific workflows aims to develop new techniques, algorithms and systems that can overcome the challenges of efficient and robust execution of ever larger workflows on increasingly complex distributed infrastructures. Since the infrastructures, systems and applications are complex, and their behavior is difficult to reproduce using physical experiments, much of this research is based on simulation. However, there exists a shortage of realistic datasets and tools that can be used for such studies. In this paper we describe a collection of tools and data that have enabled research in new techniques, algorithms, and systems for scientific workflows. These resources include: 1) execution traces of real workflow applications from which workflow and system characteristics such as resource usage and failure profiles can be extracted, 2) a synthetic workflow generator that can produce realistic synthetic workflows based on profiles extracted from execution traces, and 3) a simulator framework that can simulate the execution of synthetic workflows on realistic distributed infrastructures. This paper describes how we have used these resources to investigate new techniques for efficient and robust workflow execution, as well as to provide improvements to the Pegasus Workflow Management System or other workflow tools. Our goal in describing these resources is to share them with other researchers in the workflow research community. All of the tools and data are freely available online for the community at http://www.workflowarchive.org. These data have already been leveraged for a number of studies. Rafael Ferreira da Silva, Weiwei Chen 0002, Gideon Juve, Karan Vahi, Ewa Deelman |
eScience | 1 |
| 2014 | Controlling fairness and task granularity in distributed, online, non-clairvoyant workflow executionsabstractSUMMARY Distributed computing infrastructures are commonly used for scientific computing, and science gateways provide complete middleware stacks to allow their transparent exploitation by end users. However, administrating such systems manually is time consuming and sub‐optimal because of the complexity of the execution conditions. Algorithms and frameworks aiming at automating system administration must deal with online and non‐clairvoyant conditions, where most parameters are unknown and evolve over time. We consider the problem of controlling task granularity and fairness among scientific workflows executed in these conditions. We present two self‐managing loops monitoring the fineness, coarseness, and fairness of workflow executions, comparing these metrics with thresholds extracted from knowledge acquired in previous executions and planning appropriate actions to maintain these metrics to appropriate ranges. Experiments on the European Grid Infrastructure show that our task granularity control can speed up executions up to a factor of 2 and that our fairness control reduces slowdown variability by 3–7 compared with first‐come, first‐served. We also study the interaction between granularity control and fairness control: our experiments demonstrate that controlling task granularity degrades fairness but that our fairness control algorithm can compensate this degradation. Copyright © 2014 John Wiley & Sons, Ltd. Rafael Ferreira da Silva, Tristan Glatard, Frédéric Desprez |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Introducing PRECIP: An API for Managing Repeatable Experiments in the CloudabstractCloud computing with its on-demand access to resources has emerged as a tool used by researchers from a wide range of domains to run computer-based experiments. In this paper we introduce a flexible experiment management API, written in Python that simplifies and formalizes the execution of scientific experiments on cloud infrastructures. We describe the features and functionality of PRECIP (Pegasus Repeatable Experiments for the Cloud in Python), and how PRECIP can be used to set up experiments on academic clouds such as OpenStack Eucalyptus, Nimbus, and commercial clouds such as Amazon EC2. Sepideh Azarnoosh, Mats Rynge, Gideon Juve, Ewa Deelman, Michal Niec, Maciej Malawski, Rafael Ferreira da Silva |
CloudCom (2) | 7 |
| 2013 | Balanced Task Clustering in Scientific WorkflowsabstractScientific workflows can be composed of many fine computational granularity tasks. The runtime of these tasks may be shorter than the duration of system overheads, for example, when using multiple resources of a cloud infrastructure. Task clustering is a runtime optimization technique that merges multiple short tasks into a single job such that the scheduling overhead is reduced and the overall runtime performance is improved. However, existing task clustering strategies only provide a coarse-grained approach that relies on an over-simplified workflow model. In our work, we examine the reasons that cause Runtime Imbalance and Dependency Imbalance in task clustering. Next, we propose quantitative metrics to evaluate the severity of the two imbalance problems respectively. Furthermore, we propose a series of task balancing methods to address these imbalance problems. Finally, we analyze their relationship with the performance of these task balancing methods. A trace-based simulation shows our methods can significantly improve the runtime performance of two widely used workflows compared to the actual implementation of task clustering. Weiwei Chen 0002, Rafael Ferreira da Silva, Ewa Deelman, Rizos Sakellariou |
e-Science | 2 |
| 2013 | Workflow Fairness Control on Online and Non-clairvoyant Distributed Computing Platforms
Rafael Ferreira da Silva, Tristan Glatard, Frédéric Desprez |
Euro-Par | 1 |
| 2013 | On-Line, Non-clairvoyant Optimization of Workflow Activity Granularity on Grids
Rafael Ferreira da Silva, Tristan Glatard, Frédéric Desprez |
Euro-Par | 1 |
| 2013 | Monte Carlo simulation on heterogeneous distributed systems: A computing framework with parallel merging and checkpointing strategies
Sorina Camarasu-Pop, Tristan Glatard, Rafael Ferreira da Silva, Pierre Gueth, David Sarrut, Hugues Benoit-Cattin |
Future Gener. Comput. Syst. | 3 |
| 2013 | Self-healing of workflow activity incidents on distributed computing infrastructures
Rafael Ferreira da Silva, Tristan Glatard, Frédéric Desprez |
Future Gener. Comput. Syst. | 1 |
| 2013 | A Virtual Imaging Platform for Multi-Modality Medical Image SimulationabstractThis paper presents the Virtual Imaging Platform (VIP), a platform accessible at http://vip.creatis.insa-lyon.fr to facilitate the sharing of object models and medical image simulators, and to provide access to distributed computing and storage resources. A complete overview is presented, describing the ontologies designed to share models in a common repository, the workflow template used to integrate simulators, and the tools and strategies used to exploit computing and storage resources. Simulation results obtained in four image modalities and with different models show that VIP is versatile and robust enough to support large simulations. The platform currently has 200 registered users who consumed 33 years of CPU time in 2011. Tristan Glatard, Carole Lartizien, Bernard Gibaud, Rafael Ferreira da Silva, Germain Forestier, Frederic Cervenansky, Martino Alessandrini, Hugues Benoit-Cattin, Olivier Bernard 0001, Sorina Camarasu-Pop, Nadia Cerezo, Patrick Clarysse, Alban Gaignard, Patrick Hugonnard, Hervé Liebgott, Simon Marache, Adrien Marion, Johan Montagnat, Joachim Tabary, Denis Friboulet |
IEEE Trans. Medical Imaging | 4 |
| 2012 | Self-Healing of Operational Workflow Incidents on Distributed Computing InfrastructuresabstractDistributed computing infrastructures are commonly used through scientific gateways, but operating these gateways requires important human intervention to handle operational incidents. This paper presents a self-healing process that quantifies incident degrees of workflow activities from metrics measuring long-tail effect, application efficiency, data transfer issues, and site-specific problems. These metrics are simple enough to be computed online and they make little assumptions on the application or resource characteristics. Incidents are classified in levels and associated to sets of healing actions that are selected based on association rules modeling correlations between incident levels. The healing process is parametrized on real application traces acquired in production on the European Grid Infrastructure. Implementation and experimental results obtained in the Virtual Imaging Platform show that the proposed method speeds up execution up to a factor of 4 and properly detects unrecoverable errors. Rafael Ferreira da Silva, Tristan Glatard, Frédéric Desprez |
CCGRID | 1 |
| 2011 | Multi-modality medical image simulation of biological models with the Virtual Imaging Platform (VIP)abstractThis paper describes a framework for the integration of medical image simulators in the Virtual Imaging Platform (VIP). Simulation is widely involved in medical imaging but its availability is hampered by the heterogeneity of software interfaces and the required amount of computing power. To address this, VIP defines a simulation workflow template which transforms object models from the IntermediAte Model Format (IAMF) into native simulator formats and parallelizes the simulation computation. Format conversions, geometrical scene definition and physical parameter generation are covered. The core simulator executables are directly embedded in the simulation workflow, enabling data parallelism exploitation without modifying the simulator. The template is instantiated on simulators of the four main medical imaging modalities, namely Positron Emission Tomography, Ultrasound imaging, Magnetic Resonance Imaging and Computed Tomography. Simulation examples and performance results on the European Grid Infrastructure are shown. Adrien Marion, Germain Forestier, Hugues Benoit-Cattin, Sorina Camarasu-Pop, Patrick Clarysse, Rafael Ferreira da Silva, Bernard Gibaud, Tristan Glatard, Patrick Hugonnard, Carole Lartizien, Hervé Liebgott, Svenja Specovius, Joachim Tabary, Sébastien Valette, Denis Friboulet |
CBMS | 6 |