EDBT 2026 Demo / reviewers in the wild / expert
Marta Mattoso
dblp:m/MartaLQueirosMattoso · also Marta L. Queiros Mattoso
· DBLP profile ↗
64ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0002-0870-3371ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 15Applied, interdisciplinary, general and emerging computing · 10 · 1 since 2021Artificial intelligence and machine learning · 9 · 1 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AIabstractAs Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions. Renan Souza 0001, Silvina Caíno-Lores, Mark Coletti, Tyler J. Skluzacek, Alexandru Costan, Frédéric Suter, Marta Mattoso, Rafael Ferreira da Silva |
e-Science | 7 |
| 2023 | ProvLight: Efficient Workflow Provenance Capture on the Edge-to-Cloud ContinuumabstractModern scientific workflows require hybrid infrastructures combining numerous decentralized resources on the IoT/Edge interconnected to Cloud/HPC systems (aka the Computing Continuum) to enable their optimized execution. Understanding and optimizing the performance of such complex Edge-to-Cloud workflows is challenging. Capturing the provenance of key performance indicators, with their related data and processes, may assist in understanding and optimizing workflow executions. However, the capture overhead can be prohibitive, particularly in resource-constrained devices, such as the ones on the IoT/Edge.To address this challenge, based on a performance analysis of existing systems, we propose ProvLight, a tool to enable efficient provenance capture on the IoT/Edge. We leverage simplified data models, data compression and grouping, and lightweight transmission protocols to reduce overheads. We further integrate ProvLight into the E2Clab framework to enable workflow provenance capture across the Edge-to-Cloud Continuum. This integration makes E2Clab a promising platform for the performance optimization of applications through reproducible experiments.We validate ProvLight at a large scale with synthetic workloads on 64 real-life IoT/Edge devices in the FIT IoT LAB testbed. Evaluations show that ProvLight outperforms state-of-the-art systems like ProvLake and DfAnalyzer in resource-constrained devices. ProvLight is 26—37x faster to capture and transmit provenance data; uses 5—7x less CPU; 2x less memory; transmits 2x less data; and consumes 2—2.5x less energy. ProvLight [1] and E2Clab [2] are available as open-source tools. Daniel Rosendo, Marta Mattoso, Alexandru Costan, Renan Souza 0001, Débora B. Pina, Patrick Valduriez, Gabriel Antoniu |
CLUSTER | 2 |
| 2022 | Workflow provenance in the lifecycle of scientific machine learningabstractAbstract Machine learning (ML) has already fundamentally changed several businesses. More recently, it has also been profoundly impacting the computational science and engineering domains, like geoscience, climate science, and health science. In these domains, users need to perform comprehensive data analyses combining scientific data and ML models to provide for critical requirements, such as reproducibility, model explainability, and experiment data understanding. However, scientific ML is multidisciplinary, heterogeneous, and affected by the physical constraints of the domain, making such analyses even more challenging. In this work, we leverage workflow provenance techniques to build a holistic view to support the lifecycle of scientific ML. We contribute with (i) characterization of the lifecycle and taxonomy for data analyses; (ii) design decisions to build this view, with a W3C PROV compliant data representation and a reference system architecture; and (iii) lessons learned after an evaluation in an Oil & Gas case using an HPC cluster with 393 nodes and 946 GPUs. The experiments show that the decisions enable queries that integrate domain semantics with ML models while keeping low overhead (<1%), high scalability, and an order of magnitude of query acceleration under certain workloads against without our representation. Renan Souza 0001, Leonardo Guerreiro Azevedo, Vítor N. Lourenço, Elton F. S. Soares, Raphael Thiago, Rafael Brandão 0001, Daniel Civitarese, Emilio Vital Brazil, Márcio Ferreira Moreno, Patrick Valduriez, Marta Mattoso, Renato Cerqueira, Marco Aurélio Stelmar Netto |
Concurr. Comput. Pract. Exp. | 11 |
| 2022 | A horizontal partitioning-based method for frequent pattern mining in transport timetableabstractAbstract Analysing transport timetables is an important task, as it brings the opportunity to discover which routes commonly lead to delays. Frequent pattern mining is a technique used to support such type of discovery. However, functional dependencies are intrinsic properties present in timetables, particularly related to attributes derived from the origin–destination matrix. Such functional dependencies compromise the search for patterns in timetables in both the number of association rules (ARs) generated and the computational cost. Several of these ARs refer to the same information. Redundancy removal techniques can reduce the number of ARs. However, these techniques are designed to be used after mining finishes, which increases the computational cost of finding useful ARs. This work presents timetable pattern mining (T‐mine), a novel method for frequent pattern mining that improves knowledge discovery in timetables. We evaluated T‐mine using Brazilian Flight Data and compared T‐mine with the direct application of frequent pattern mining approaches with and without functional dependencies. Our experiments indicate that T‐mine is about one order magnitude faster than other methods with functional dependencies. Claudio Teixeira, Luana Fragoso, Marta Mattoso, Diego Carvalho 0001, Eduardo Bezerra 0002, Jorge Soares 0001, Glauco Fiorott Amorim, Eduardo S. Ogasawara |
Expert Syst. J. Knowl. Eng. | 3 |
| 2022 | Latency and Energy-Awareness in Data Stream Processing for Edge Based IoT Systems
Egberto A. R. de Oliveira, Atslands Rego da Rocha, Marta Mattoso, Flávia Coimbra Delicato |
J. Grid Comput. | 3 |
| 2021 | A Real-time and Energy-aware Framework for Data Stream Processing in the Internet of Things
Egberto A. R. de Oliveira, Flávia Coimbra Delicato, Atslands Rego da Rocha, Marta Mattoso |
IoTBDS | 4 |
| 2020 | Capturing and Analyzing Provenance from Spark-based Scientific Workflows with SAMbA-RaP
Thaylon Guedes, Lucas Bertelli Martins, Maria Luiza Furtuozo Falci, Vítor Silva 0003, Kary A. C. S. Ocaña, Marta Mattoso, Marcos V. N. Bedo, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 6 |
| 2020 | Adding domain data to code profiling tools to debug workflow parallel executionabstractComputer simulations may be composed of several scientific programs chained in a coherent flow running in High Performance Computing and cloud environments. These runs may present different execution behavior associated to the parallel flow of data among programs. Gather insight into the parallel flow of data is important for several applications. The usual way of getting insight into code performance is by means of a code-profiler. Several parallel code-profiling tools already support performance analysis, such as Tuning and Analysis Utilities (TAU), or provide fine-grained performance statistics, e.g., System Activity Report (SAR). These tools are effective for code profiling, but are not connected to the concept of IO-intensive workflows. Analyzing the workflow execution with domain and performance data is important for users because they can identify anomalies, choose suitable machines to run their workflows, etc. This type of analysis may be performed by capturing execution data enriched with fine-grained domain data during the long-term run of a computer simulation. In this paper, we propose a monitoring data capture approach as a component that couples code-profiling tools to domain data from workflow executions. The goal is to profile and debug parallel executions of workflows through queries to a database that integrates performance, resource consumption, provenance, and domain data from simulation programs flow at runtime. We show how querying this database with domain-aware data at runtime allows to identify performance anomalies not detected by code-profiling tools. We evaluate our approach using the astronomy Montage workflow on a cluster environment and the SciPhy bioinformatics workflow on the Amazon cloud. In both cases computing time overhead imposed by our approach for gathering fine-grained domain, performance, and resource consumption data is negligible. Vítor Silva 0003, Leonardo Neves, Renan Souza 0001, Alvaro L. G. A. Coutinho, Daniel de Oliveira 0001, Marta Mattoso |
Future Gener. Comput. Syst. | 6 |
| 2020 | Data reduction in scientific workflows using provenance monitoring and user steering
Renan Souza 0001, Vítor Silva 0003, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 5 |
| 2019 | Efficient Runtime Capture of Multiworkflow Data Using ProvenanceabstractComputational Science and Engineering (CSE) projects are typically developed by multidisciplinary teams. Despite being part of the same project, each team manages its own workflows, using specific execution environments and data processing tools. Analyzing the data processed by all workflows globally is a core task in a CSE project. However, this analysis is hard because the data generated by these workflows are not integrated. In addition, since these workflows may take a long time to execute, data analysis needs to be done at runtime to reduce cost and time of the CSE project. A typical solution in scientific data analysis is to capture and relate the data in a provenance database while the workflows run, thus allowing for data analysis at runtime. However, the main problem is that such data capture competes with the running workflows, adding significant overhead to their execution. To mitigate this problem, we introduce in this paper a system called ProvLake, which adopts design principles for providing efficient distributed data capture from the workflows. While capturing the data, ProvLake logically integrates and ingests them into a provenance database ready for analyses at runtime. We validated ProvLake in a real use case in the O&G industry encompassing four workflows that process 5 TB datasets for a deep learning classifier. Compared with Komadu, the closest solution that meets our goals, our approach enables runtime multiworkflow data analysis with much smaller overhead, such as 0.1%. Renan Souza 0001, Marta Mattoso, Leonardo Guerreiro Azevedo, Raphael Thiago, Elton F. S. Soares, Marcelo Nery dos Santos, Marco Aurélio Stelmar Netto, Emilio Vital Brazil, Renato Cerqueira, Patrick Valduriez |
eScience | 2 |
| 2019 | Keeping track of user steering actions in dynamic workflows
Renan Souza 0001, Vítor Silva 0003, José J. Camata, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 6 |
| 2019 | Efficient Scheduling of Scientific Workflows Using Hot Metadata in a Multisite CloudabstractLarge-scale, data-intensive scientific applications are often expressed as scientific workflows (SWfs). In this paper, we consider the problem of efficient scheduling of a large SWf in a multisite cloud, i.e., a cloud with geo-distributed cloud data centers (sites). The reasons for using multiple cloud sites to run a SWf are that data is already distributed, the necessary resources exceed the limits at a single site, or the monetary cost is lower. In a multisite cloud, metadata management has a critical impact on the efficiency of SWf scheduling as it provides a global view of data location and enables task tracking during execution. Thus, it should be readily available to the system at any given time. While it has been shown that efficient metadata handling plays a key role in performance, little research has targeted this issue in multisite cloud. In this paper, we propose to identify and exploit hot metadata (frequently accessed metadata) for efficient SWf scheduling in a multisite cloud, using a distributed approach. We implemented our approach within a scientific workflow management system, which shows that our approach reduces the execution time of highly parallel jobs up to 64 percent and that of the whole SWfs up to 55 percent. Ji Liu 0003, Luis Pineda-Morales, Esther Pacitti, Alexandru Costan, Patrick Valduriez, Gabriel Antoniu, Marta Mattoso |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2018 | DfAnalyzer: Runtime Dataflow Analysis of Scientific Applications using ProvenanceabstractWe present DfAnalyzer, a tool that enables monitoring, debugging, steering, and analysis of dataflows while being generated by scientific applications. It works by capturing strategic domain data, registering provenance and execution data to enable queries at runtime. DfAnalyzer provides lightweight dataflow monitoring components to be invoked by high performance applications. It can be plugged in scientific code scripts, or Spark applications, in the same way users already plug visualization library components. During this demo, we will show how DfAnalyzer captures the dataflow, provenance, as well as how it provides runtime data analyses of applications. We will also encourage attendees to use DfAnalyzer for their own applications. Vítor Silva 0003, Daniel de Oliveira 0001, Marta Mattoso, Patrick Valduriez |
Proc. VLDB Endow. | 3 |
| 2017 | Deriving scientific workflows from algebraic experiment lines: A practical approach
Anderson Marinho, Daniel de Oliveira 0001, Eduardo S. Ogasawara, Vítor Silva 0003, Kary A. C. S. Ocaña, Leonardo Murta 0001, Vanessa Braganholo, Marta Mattoso |
Future Gener. Comput. Syst. | 8 |
| 2017 | Raw data queries during data-intensive parallel workflow execution
Vítor Silva 0003, José Leite, José J. Camata, Daniel de Oliveira 0001, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 7 |
| 2016 | Managing hot metadata for scientific workflows on multisite cloudsabstractLarge-scale scientific applications are often expressed as workflows that help defining data dependencies between their different components. Several such workflows have huge storage and computation requirements, and so they need to be processed in multiple (cloud-federated) datacenters. It has been shown that efficient metadata handling plays a key role in the performance of computing systems. However, most of this evidence concern only single-site, HPC systems to date. In this paper, we present a hybrid decentralized/distributed model for handling hot metadata (frequently accessed metadata) in multisite architectures. We couple our model with a scientific workflow management system (SWfMS) to validate and tune its applicability to different real-life scientific scenarios. We show that efficient management of hot metadata improves the performance of SWfMS, reducing the workflow execution time up to 50% for highly parallel jobs and avoiding unnecessary cold metadata operations. Luis Pineda-Morales, Ji Liu 0003, Alexandru Costan, Esther Pacitti, Gabriel Antoniu, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 7 |
| 2016 | Analyzing related raw data files through dataflowsabstractSummary Computer simulations may ingest and generate high numbers of raw data files. Most of these files follow a de facto standard format established by the application domain, for example, Flexible Image Transport System for astronomy. Although these formats are supported by a variety of programming languages, libraries, and programs, analyzing thousands or millions of files requires developing specific programs. Database management systems (DBMS) are not suited for this, because they require loading the raw data and structuring it, which becomes heavy at large scale. Systems like NoDB, RAW, and FastBit have been proposed to index and query raw data files without the overhead of using a database management system. However, these solutions are focused on analyzing one single large file instead of several related files. In this case, when related files are produced and required for analysis, the relationship among elements within file contents must be managed manually, with specific programs to access raw data. Thus, this data management may be time‐consuming and error‐prone. When computer simulations are managed by a scientific workflow management system (SWfMS), they can take advantage of provenance data to relate and analyze raw data files produced during workflow execution. However, SWfMS registers provenance at a coarse grain, with limited analysis on elements from raw data files. When the SWfMS is dataflow‐aware, it can register provenance data and the relationships among elements of raw data files altogether in a database, which is useful to access the contents of a large number of files. In this paper, we propose a dataflow approach for analyzing element data from several related raw data files. Our approach is complementary to the existing single raw data file analysis approaches. We use the Montage workflow from astronomy and a workflow from Oil and Gas domain as data‐intensive case studies. Our experimental results for the Montage workflow explore different types of raw data flows like showing all linear transformations involved in projection simulation programs, considering specific mosaic elements from input repositories. The cost for raw data extraction is approximately 3.7% of the total application execution time. Copyright © 2015 John Wiley & Sons, Ltd. Vítor Silva 0003, Daniel de Oliveira 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Multi-objective scheduling of Scientific Workflows in multisite clouds
Ji Liu 0003, Esther Pacitti, Patrick Valduriez, Daniel de Oliveira 0001, Marta Mattoso |
Future Gener. Comput. Syst. | 5 |
| 2015 | Data Analytics in Bioinformatics: Data Science in Practice for Genomics Analysis WorkflowsabstractWorkflow systems manage large-scale experiments and deliver a large volume of provenance data traces. The provenance repository of these systems contains information about the workflow execution, which allows for tracking and analyzing data transformations. However, provenance data may still be considered a black-box, when it comes to analyze the contents of resulting data files. Current solutions are focused on data transformation at coarse grain, they point to input and output files, but do not allow for exploring domain-specific data. Data analytics is essential for managing large-scale workflows executed in parallel, especially when tracking anomalous executions. In this paper, we present a data analytics approach, which is based on the use of provenance data enriched with domain-specific data coupled to a data mining tool. A real bioinformatics workflow was modeled and executed in parallel on top of Amazon clouds. It manipulates complex biological data, which is difficult to monitor like many other genomic workflows. We evaluate the benefits of using domain-specific data and provenance data for user steering while monitoring the execution with detailed filters, steering on specific conditions and performance evaluation. Results show that the provenance database coupled to workflow systems has an unexplored potential for raw data analytics, which may improve the user confidence and reduce overall execution time. Kary A. C. S. Ocaña, Vítor Silva 0003, Daniel de Oliveira 0001, Marta Mattoso |
e-Science | 4 |
| 2015 | Data-centric iteration in dynamic workflows
Jonas Dias, Gabriel Guerra, Fernando Rochinha, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 6 |
| 2015 | Dynamic steering of HPC scientific workflows: A survey
Marta Mattoso, Jonas Dias, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Flavio Costa, Felipe Horta, Vítor Silva 0003, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 1 |
| 2015 | A Survey of Data-Intensive Scientific Workflow Management
Ji Liu 0003, Esther Pacitti, Patrick Valduriez, Marta Mattoso |
J. Grid Comput. | 4 |
| 2013 | Algebraic dataflows for big data analysisabstractAnalyzing big data requires the support of dataflows with many activities to extract and explore relevant information from the data. Recent approaches such as Pig Latin propose a high-level language to model such dataflows. However, the dataflow execution is typically delegated to a MapRe-duce implementation such as Hadoop, which does not follow an algebraic approach, thus it cannot take advantage of the optimization opportunities of PigLatin algebra. In this paper, we propose an approach for big data analysis based on algebraic workflows, which yields optimization and parallel execution of activities and supports user steering using provenance queries. We illustrate how a big data processing dataflow can be modeled using the algebra. Through an experimental evaluation using real datasets and the execution of the dataflow with Chiron, an engine that supports our algebra, we show that our approach yields performance gains of up to 19.6% using algebraic optimizations in the dataflow and up to 39.1% of time saved on a user steering scenario. Jonas Dias, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 6 |
| 2013 | On the performance of the position() XPath functionabstractIn very large XML documents or collections, the query response times are not always satisfactory. To overcome this limitation, parallel processing can be applied. Data can be replicated in several processors and queries can be partitioned to run over different virtual data partitions on each processor, on an approach called virtual partitioning. PartiX-VP is a simple XML virtual partitioning approach that generates virtual data partitions by dividing the cardinality of the partitioning attribute by the number of allocated processors, resulting in intervals of equal size for each processor. In this approach, the XML query is rewritten and selection predicates are added to define the virtual partitions. These selection predicates use the position() XPath function that addresses a set of elements on a given position in the document. In this paper, we present an experimental evaluation of the position() XPath function in five XML native DBMS. We have identified differences in the processing time of the position() XPath function in large collections of XML documents. This may lead to load unbalancing in simple virtual partitioning approaches, thus this analysis opens space for improvements in virtual partitioning. Luiz Augusto Matos da Silva, Luiz Laerte N. da Silva Jr., Marta Mattoso, Vanessa Braganholo |
ACM Symposium on Document Engineering | 3 |
| 2013 | Chiron: a parallel engine for algebraic scientific workflowsabstractSUMMARY Large‐scale scientific experiments based on computer simulations are typically modeled as scientific workflows, which eases the chaining of different programs. These scientific workflows are defined, executed, and monitored by scientific workflow management systems (SWfMS). As these experiments manage large amounts of data, it becomes critical to execute them in high‐performance computing environments, such as clusters, grids, and clouds. However, few SWfMS provide parallel support. The ones that do so are usually labor‐intensive for workflow developers and have limited primitives to optimize workflow execution. To address these issues, we developed workflow algebra to specify and enable the optimization of parallel execution of scientific workflows. In this paper, we show how the workflow algebra is efficiently implemented in Chiron, an algebraic based parallel scientific workflow engine. Chiron has a unique native distributed provenance mechanism that enables runtime queries in a relational database. We developed two studies to evaluate the performance of our algebraic approach implemented in Chiron; the first study compares Chiron with different approaches, whereas the second one evaluates the scalability of Chiron. By analyzing the results, we conclude that Chiron is efficient in executing scientific workflows, with the benefits of declarative specification and runtime provenance support. Copyright © 2013 John Wiley & Sons, Ltd. Eduardo S. Ogasawara, Jonas Dias, Vítor Silva 0003, Fernando Seabra Chirigati, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 8 |
| 2013 | Designing a parallel cloud based comparative genomics workflow to improve phylogenetic analyses
Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
Future Gener. Comput. Syst. | 5 |
| 2013 | Performance evaluation of parallel strategies in public clouds: A study with phylogenomic workflows
Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Jonas Dias, João Carlos de A. R. Gonçalves, Fernanda Baião, Marta Mattoso |
Future Gener. Comput. Syst. | 7 |
| 2012 | Discovering drug targets for neglected diseases using a pharmacophylogenomic cloud workflowabstractIllnesses caused by parasitic protozoan are a research priority. A representative group of these illnesses is the commonly known as Neglected Tropical Diseases (NTD). NTD specially attack low socioeconomic population around the world and new anti-protozoan inhibitors are needed and several drug discovery projects focus on researching new drug targets. Pharmacophylogenomics is a novel bioinformatics field that aims at reducing the time and the financial cost of the drug discovery process. Pharmacophylogenomic analyses are applied mainly in the early stages of the research phase in drug discovery. Pharmacophylogenomic analysis executes several bioinformatics programs in a coherent flow to identify homologues sequences, construct phylogenetic trees and execute evolutionary and structural experiments. This way, it can be modeled as scientific workflows. Pharmacophylogenomic analysis workflows are complex, computing and data intensive and may execute during weeks. This way, it benefits from parallel execution. We propose SciPPGx, a scientific workflow that aims at providing thorough inferring support for pharmacophylogenomic hypotheses. SciPPGx is executed in parallel in a cloud using SciCumulus workflow engine. Experiments show that SciPPGx considerably reduces the total execution time up to 97.1% when compared to a sequential execution. We also present representative biological results taking advantage of the inference covering several related bioinformatics overviews. Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
eScience | 5 |
| 2012 | ProvManager: a provenance management system for scientific workflowsabstractSUMMARY Running scientific workflows in distributed and heterogeneous environments has been a motivating approach for provenance management, which is loosely coupled to the workflow execution engine. This kind of approach is interesting because it allows both storage and access to provenance data in a homogeneous way, even in an environment where different workflow management systems work together. However, current approaches overload scientists with many ad hoc tasks, such as script adaptations and implementations of extra functionalities to provide provenance independence. This paper proposes ProvManager, a provenance management approach that eases the gathering, storage, and analysis of provenance information in a distributed and heterogeneous environment scenario, without putting the burden of adaptations on the scientist. ProvManager leverages the provenance management at the experiment level by integrating different workflow executions from multiple workflow management systems. Copyright © 2011 John Wiley & Sons, Ltd. Anderson Marinho, Leonardo Murta 0001, Cláudia M. L. Werner, Vanessa Braganholo, Sérgio Manuel Serra da Cruz, Eduardo S. Ogasawara, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 7 |
| 2012 | An adaptive parallel execution strategy for cloud-based scientific workflowsabstractSUMMARY Many of the existing large‐scale scientific experiments modeled as scientific workflows are compute‐intensive. Some scientific workflow management systems already explore parallel techniques, such as parameter sweep and data fragmentation, to improve performance. In those systems, computing resources are used to accomplish many computational tasks in high performance environments, such as multiprocessor machines or clusters. Meanwhile, cloud computing provides scalable and elastic resources that can be instantiated on demand during the course of a scientific experiment, without requiring its users to acquire expensive infrastructure or to configure many pieces of software. In fact, because of these advantages some scientists have already adopted the cloud model in their scientific experiments. However, this model also raises many challenges. When scientists are executing scientific workflows that require parallelism, it is hard to decide a priori the amount of resources to use and how long they will be needed because the allocation of these resources is elastic and based on demand. In addition, scientists have to manage new aspects such as initialization of virtual machines and impact of data staging. SciCumulus is a middleware that manages the parallel execution of scientific workflows in cloud environments. In this paper, we introduce an adaptive approach for executing parallel scientific workflows in the cloud. This approach adapts itself according to the availability of resources during workflow execution. It checks the available computational power and dynamically tunes the workflow activity size to achieve better performance. Experimental evaluation showed the benefits of parallelizing scientific workflows using the adaptive approach of SciCumulus, which presented an increase of performance up to 47.1%. Copyright © 2011 John Wiley & Sons, Ltd. Daniel de Oliveira 0001, Eduardo S. Ogasawara, Kary A. C. S. Ocaña, Fernanda Baião, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 5 |
| 2012 | MTCProv: a practical provenance query framework for many-task scientific computing
Luiz M. R. Gadelha Jr., Michael Wilde, Marta Mattoso, Ian T. Foster |
Distributed Parallel Databases | 3 |
| 2012 | A Provenance-based Adaptive Scheduling Heuristic for Parallel Scientific Workflows in Clouds
Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Fernanda Baião, Marta Mattoso |
J. Grid Comput. | 4 |
| 2011 | A Performance Evaluation of X-Ray Crystallography Scientific Workflow Using SciCumulusabstractX-ray crystallography is an important field due to its role in drug discovery and its relevance in bioinformatics experiments of comparative genomics, phylogenomics, evolutionary analysis, ortholog detection, and three-dimensional structure determination. Managing these experiments is a challenging task due to the orchestration of legacy tools and the management of several variations of the same experiment. Workflows can model a coherent flow of activities that are managed by scientific workflow management systems (SWfMS). Due to the huge amount of variations of the workflow to be explored (parameters, input data) it is often necessary to execute X-ray crystallography experiments in High Performance Computing (HPC) environments. Cloud computing is well known for its scalable and elastic HPC model. In this paper, we present a performance evaluation for the X-ray crystallography workflow defined by the PC4 (Provenance Challenge series). The workflow was executed using the SciCumulus middleware at the Amazon EC2 cloud environment. SciCumulus is a layer for SWfMS that offers support for the parallel execution of scientific workflows in cloud environments with provenance mechanisms. Our results reinforce the benefits (total execution time × monetary cost) of parallelizing the X-ray crystallography workflow using SciCumulus. The results show a consistent way to execute X-ray crystallography workflows that need HPC using cloud computing. The evaluated workflow shares features of many scientific workflows and can be applied to other experiments. Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Jonas Dias, Fernanda Baião, Marta Mattoso |
IEEE CLOUD | 6 |
| 2011 | Optimizing Phylogenetic Analysis Using SciHmm Cloud-based Scientific WorkflowabstractPhylogenetic analysis and multiple sequence alignment (MSA) are closely related bioinformatics fields. Phylogenetic analysis makes extensive use of MSA in the construction of phylogenetic trees, which are used to infer the evolutionary relationships between homologous genes. These bioinformatics experiments are usually modeled as scientific workflows. There are many alternative workflows that use different MSA methods to conduct phylogenetic analysis and each one can produce MSA with different quality. Scientists have to explore which MSA method is the most suitable for their experiments. However, workflows for phylogenetic analysis are both computational and data intensive and they may run sequentially during weeks. Although there any many approaches that parallelize these workflows, exploring all MSA methods many become a burden and expensive task. If scientists know the most adequate MSA method a priori, it would spare time and money. To optimize the phylogenetic analysis workflow, we propose in this paper SciHmm, a bioinformatics scientific workflow based in profile hidden Markov models (pHMMs) that aims at determining the most suitable MSA method for a phylogenetic analysis prior than executing the phylogenetic workflow. SciHmm is also executed in parallel in a cloud environment using SciCumulus middleware. The results demonstrated that optimizing a phylogenetic analysis using SciHmm considerably reduce the total execution time of phylogenetic analysis (up to 80%). This optimization also demonstrates that the biological results presented more quality. In addition, the parallel execution of SciHmm demonstrates that this kind of bioinformatics workflow is suitable to be executed in the cloud. Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
eScience | 5 |
| 2011 | Towards a Cost Model for Scheduling Scientific Workflows Activities in Cloud EnvironmentsabstractCloud computing has emerged as a new paradigm that enables scientists to benefit from several distributed resources such as hardware and software. Clouds poses as an opportunity for scientists that need high performance computing infrastructure to execute their scientific experiments. Most of the experiments modeled as scientific workflows manage the execution of several activities and work with large amounts of data. In this way parallel techniques are often a key factor. Parallelizing a scientific workflow in the cloud environment is not trivial. One of the complex tasks is to define the number and types of virtual machines and to design the parallel execution strategy. Due to the number of options for configuring an environment it is a hard task to do it manually and it may produce negative impact on performance. This paper initially proposes a cost model based on concepts of quality of service (QoS) in clouds to help determining an adequate configuration of the environment according to restrictions imposed by scientists. Vitor Viana, Daniel de Oliveira 0001, Marta Mattoso |
SERVICES | 3 |
| 2011 | Many task computing for orthologous genes identification in protozoan genomes using HydraabstractSUMMARY One of the main advantages of using a scientific workflow management system (SWfMS) is to orchestrate data flows among scientific activities and register provenance of the whole workflow execution. Nevertheless, the execution control of distributed activities in high performance computing environments by SWfMS presents challenges such as steering control and provenance gathering. Such challenges may become a complex task to be accomplished in bioinformatics experiments, particularly in Many Task Computing scenarios. This paper presents a data parallelism solution for a bioinformatics experiment supported by Hydra, a middleware that bridges SWfMS and high performance computing to enable workflow parallelization with provenance gathering. Hydra Many Task Computing parallelization strategies can be registered and reused. Using Hydra, provenance may also be uniformly gathered. We have evaluated Hydra using an Orthologous Gene Identification workflow. Experimental results show that a systematic approach for distributing parallel activities is viable, sparing scientist time and diminishing operational errors, with the additional benefits of distributed provenance support. Copyright © 2011 John Wiley & Sons, Ltd. Fábio Coutinho, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Vanessa Braganholo, Alexandre A. B. Lima, Alberto M. R. Dávila, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 7 |
| 2011 | Provenance management in Swift
Luiz M. R. Gadelha Jr., Ben Clifford, Marta Mattoso, Michael Wilde, Ian T. Foster |
Future Gener. Comput. Syst. | 3 |
| 2011 | An Algebraic Approach for Data-Centric Scientific Workflows
Eduardo S. Ogasawara, Daniel de Oliveira 0001, Patrick Valduriez, Jonas Dias, Fábio Porto 0001, Marta Mattoso |
Proc. VLDB Endow. | 6 |
| 2010 | SciCumulus: A Lightweight Cloud Middleware to Explore Many Task Computing Paradigm in Scientific WorkflowsabstractMost of the large-scale scientific experiments modeled as scientific workflows produce a large amount of data and require workflow parallelism to reduce workflow execution time. Some of the existing Scientific Workflow Management Systems (SWfMS) explore parallelism techniques - such as parameter sweep and data fragmentation. In those systems, several computing resources are used to accomplish many computational tasks in homogeneous environments, such as multiprocessor machines or cluster systems. Cloud computing has become a popular high performance computing model in which (virtualized) resources are provided as services over the Web. Some scientists are starting to adopt the cloud model in scientific domains and are moving their scientific workflows (programs and data) from local environments to the cloud. Nevertheless, it is still difficult for the scientist to express a parallel computing paradigm for the workflow on the cloud. Capturing distributed provenance data at the cloud is also an issue. Existing approaches for executing scientific workflows using parallel processing are mainly focused on homogeneous environments whereas, in the cloud, the scientist has to manage new aspects such as initialization of virtualized instances, scheduling over different cloud environments, impact of data transferring and management of instance images. In this paper we propose SciCumulus, a cloud middleware that explores parameter sweep and data fragmentation parallelism in scientific workflow activities (with provenance support). It works between the SWfMS and the cloud. SciCumulus is designed considering cloud specificities. We have evaluated our approach by executing simulated experiments to analyze the overhead imposed by clouds on the workflow execution time. Daniel de Oliveira 0001, Eduardo S. Ogasawara, Fernanda Baião, Marta Mattoso |
IEEE CLOUD | 4 |
| 2010 | Data parallelism in bioinformatics workflows using HydraabstractLarge scale bioinformatics experiments are usually composed by a set of data flows generated by a chain of activities (programs or services) that may be modeled as scientific workflows. Current Scientific Workflow Management Systems (SWfMS) are used to orchestrate these workflows to control and monitor the whole execution. It is very common in bioinformatics experiments to process very large datasets. In this way, data parallelism is a common approach used to increase performance and reduce overall execution time. However, most of current SWfMS still lack on supporting parallel executions in high performance computing (HPC) environments. Additionally keeping track of provenance data in distributed environments is still an open, yet important problem. Recently, Hydra middleware was proposed to bridge the gap between the SWfMS and the HPC environment, by providing a transparent way for scientists to parallelize workflow executions while capturing distributed provenance. This paper analyzes data parallelism scenarios in bioinformatics domain and presents an extension to Hydra middleware through a specific cartridge that promotes data parallelism in bioinformatics workflows. Experimental results using workflows with BLAST show performance gains with the additional benefits of distributed provenance support. Fábio Coutinho, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Vanessa Braganholo, Alexandre A. B. Lima, Alberto M. R. Dávila, Marta Mattoso |
HPDC | 7 |
| 2010 | Adaptive Normalization: A novel data normalization approach for non-stationary time seriesabstractData normalization is a fundamental preprocessing step for mining and learning from data. However, finding an appropriated method to deal with time series normalization is not a simple task. This is because most of the traditional normalization methods make assumptions that do not hold for most time series. The first assumption is that all time series are stationary, i.e., their statistical properties, such as mean and standard deviation, do not change over time. The second assumption is that the volatility of the time series is considered uniform. None of the methods currently available in the literature address these issues. This paper proposes a new method for normalizing non-stationary heteroscedastic (with non-uniform volatility) time series. The method, named Adaptive Normalization (AN), was tested together with an Artificial Neural Network (ANN) in three forecast problems. The results were compared to other four traditional normalization methods, and showed AN improves ANN accuracy in both short- and long-term predictions. Eduardo S. Ogasawara, Leonardo C. Martinez, Daniel de Oliveira 0001, Geraldo Zimbrão, Gisele L. Pappa, Marta Mattoso |
IJCNN | 6 |
| 2010 | ARAXA: Storing and managing Active XML documents
Cláudio Ananias Ferraz, Vanessa Braganholo, Marta Mattoso |
J. Web Semant. | 3 |
| 2009 | Neural networks cartridges for data mining on time seriesabstractNeural networks is one of the techniques used for time series analysis. The performance of neural networks is affected by some parameters such as neural network structure and the quality of data preprocessing. These parameters need to be explored in order to obtain an optimal neural network. However, the manual establishment of different neural networks configurations for selecting the best ones may be error-prone and time-consuming. This paper proposes the creation of neural networks cartridges to systematically empower neural network performance by means of data mining activities, which obtain an optimal neural network structure. The experiments conducted in this paper use stock market and exchange rate series, and show that the usage of neural network cartridges can lead to configurations that double the performance of some ad-hoc neural network configuration. Eduardo S. Ogasawara, Leonardo Murta 0001, Geraldo Zimbrão, Marta Mattoso |
IJCNN | 4 |
| 2009 | Experiment Line: Software Reuse in Scientific Workflows
Eduardo S. Ogasawara, Carlos Eduardo Paulino Silva, Leonardo Murta 0001, Cláudia M. L. Werner, Marta Mattoso |
SSDBM | 5 |
| 2009 | Parallel OLAP query processing in database clusters with data replication
Alexandre A. B. Lima, Camille Furtado, Patrick Valduriez, Marta Mattoso |
Distributed Parallel Databases | 4 |
| 2008 | Provenance Services for Distributed WorkflowsabstractScientific experiments using workflows benefit from mechanisms to trace the generation of results. As workflows start to scale it is fundamental to have access to their underlying processes, parameters and data. Particularly in molecular dynamics (MD) simulations, a study of the interatomic interactions in proteins must use distributed high performance computing environments to produce timely results. Scientist's trust in experiments produced by gathering distributed partial results may be limited without provenance information. This paper presents a service architecture that captures and stores provenance data from distributed, autonomous, replicated and heterogeneous resources. Such provenance data can be used to trace the history of the distributed execution process. These services can be coupled to workflow management systems. The Kepler system was used as a basis to manage a grid workflow application. Experimental results regarding cluster and grid MD simulations were evaluated using the provenance services architecture. Sérgio Manuel Serra da Cruz, Patrícia M. Barros, Paulo Mascarello Bisch, Maria Luiza M. Campos, Marta Mattoso |
CCGRID | 5 |
| 2008 | A Lightweight Middleware Monitor for Distributed Scientific WorkflowsabstractMonitoring the execution of distributed tasks within the workflow execution is not easy and is frequently controlled manually. This work presents a lightweight middleware monitor to design and control the parallel execution of tasks from a distributed scientific workflow. This middleware can be connected into a workflow management system. This middleware implementation is evaluated with the Kepler workflow management system, by including new modules to control and monitor the distributed execution of the tasks. These middleware modules were added to a bio informatics workflow to monitor parallel BLAST executions. Results show potential to high performance process execution while preserving the original features of the workflow. Sérgio Manuel Serra da Cruz, Fabrício Nogueira da Silva, Luiz M. R. Gadelha Jr., Maria Cláudia Cavalcanti, Maria Luiza M. Campos, Marta Mattoso |
CCGRID | 6 |
| 2008 | Kairos: An Architecture for Securing Authorship and Temporal Information of Provenance Data in Grid-Enabled Workflow Management SystemsabstractSecure provenance techniques are essential in generating trustworthy provenance records, where one is interested in protecting their integrity, confidentiality, and availability. In this work, we suggest an architecture to provide protection of authorship and temporal information in grid-enabled provenance systems. It can be used in the resolution of conflicting intellectual property claims, and in the reliable chronological reconstitution of scientific experiments. We observe that some techniques from public key infrastructures can be readily applied for this purpose. We discuss the issues involved in the implementation of such architecture and describe some experiments realized with the proposed techniques. Luiz M. R. Gadelha Jr., Marta Mattoso |
eScience | 2 |
| 2008 | XCraft: boosting the performance of active XML materializationabstractAn active XML (AXML) document contains tags representing calls to Web services. Therefore, retrieving its contents consists in materializing its data elements by invoking the embedded service calls in a P2P network. In this process, the result of some service calls can be used as input of other calls. Also, usually several peers provide each requested Web service, and peers can collaborate to invoke these services. This often implies a huge search space of many equivalent materialization alternatives, each with different performance. In this paper, we model AXML documents from a workflow perspective and propose a dynamic cost-based optimization strategy to efficiently materialize them, considering the volatility of a typical P2P scenario. Our strategy enables the optimizer, called XCraft, to get more up-to-date information on the status of the peers, and to deliver partial results earlier. Based on a service-oriented algebra of plan operators, we exploit P2P collaboration to delegate both execution and optimization control. Our tests with an XCraft prototype show important performance gains w.r.t. a centralized approach, whilst the optimizer also achieved to drastically reduce the size of the search space. Gabriela Ruberg, Marta Mattoso |
EDBT | 2 |
| 2008 | RL-Based Scheduling Strategies in Actual Grid EnvironmentsabstractIn this work, we study the behaviour of different resource scheduling strategies when doing job orchestration in grid environments. We empirically demonstrate that scheduling strategies based on reinforcement learning are a good choice to improve the overall performance of grid applications and resource utilization. Bernardo Fortunato Costa, Inês de Castro Dutra, Marta Mattoso |
ISPA | 3 |
| 2008 | Parallel query processing for OLAP in gridsabstractAbstract OLAP query processing is critical for enterprise grids. Capitalizing on our experience with the ParGRES database cluster, we propose a middleware solution, GParGRES, which exploits database replication and inter‐ and intra‐query parallelism to efficiently support OLAP queries in a grid. GParGRES is designed as a wrapper that enables the use of ParGRES in PC clusters of a grid (in our case, Grid5000). Our approach has two levels of query splitting: grid‐level splitting, implemented by GParGRES, and node‐level splitting, implemented by ParGRES. GParGRES has been partially implemented as database grid services compatible with existing grid solutions such as the open grid service architecture and the Web services resource framework. We give preliminary experimental results obtained with two clusters of Grid5000 using queries of the TPC‐H Benchmark. The results show linear or almost linear speedup in query execution, as more nodes are added in all tested configurations. Copyright © 2008 John Wiley & Sons, Ltd. Nelson Kotowski, Alexandre A. B. Lima, Esther Pacitti, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 5 |
| 2007 | Preface to the Special Issue on Grid Data Management
Esther Pacitti, Marta Mattoso, Patrick Valduriez |
J. Grid Comput. | 2 |
| 2007 | Grid Data Management: Open Problems and New Issues
Esther Pacitti, Patrick Valduriez, Marta Mattoso |
J. Grid Comput. | 3 |
| 2006 | Odyssey-Search: A multi-agent system for component information search and retrieval
Regina Braga 0001, Cláudia M. L. Werner, Marta Mattoso |
J. Syst. Softw. | 3 |
| 2005 | Physical and Virtual Partitioning in OLAP Database ClustersabstractOn-line analytical processing (OLAP) applications require high performance database support to achieve good response time (crucial for decision making). Database clusters provide a cost-effective alternative to parallel database systems. For OLAP applications, that typically use heavy weight queries, intra-query parallelism yields better performance as it reduces the execution time of individual queries. Intra-query parallelism is based on processing the same query on different subsets of the query table. Combining physical and virtual partitioning to define table subsets provides flexibility in intra-query parallelism while optimizing disk space usage and data availability. Experiments with our partitioning technique using TPC-H benchmark queries on a 32-dual node cluster gave linear and super-linear speedup, thereby reducing significantly the time of typical OLAP heavy weight queries. Camille Furtado, Alexandre A. B. Lima, Esther Pacitti, Patrick Valduriez, Marta Mattoso |
SBAC-PAD | 5 |
| 2005 | Managing structural genomic workflows using Web services
Maria Cláudia Cavalcanti, Rafael Targino, Fernanda Baião, Shaila C. Rössle, Paulo Mascarello Bisch, Paulo F. Pires, Maria Luiza M. Campos, Marta Mattoso |
Data Knowl. Eng. | 8 |
| 2004 | OLAP Query Processing in a Database Cluster
Alexandre A. B. Lima, Marta Mattoso, Patrick Valduriez |
Euro-Par | 2 |
| 2004 | Automatic Composition of Web Services with Contingency PlansabstractThe semantic Web technology and the Web services description language extensibility may be combined to describe services in an unambiguous and machine interpretable way, automating Web services discovery, selection and invocation. In this paper, we present an algorithm and a prototype for the automatic composition of Web services that implement workflows described in a high level language. Our approach has many advantages comparing to the manual creation of a simple program composition, such as smaller implementation time and cost, reliability with the generation of contingency plans, greater capacity to evolve with the dynamic service discovery, and faster execution time with the use of heuristics. We use the OWLS ontology to semantically describe Web services metadata and indexes to help selecting them. The proposed algorithm considers that equivalent services may have different interfaces and also respects preferences of the users. Luiz A. G. da Costa, Paulo F. Pires, Marta Mattoso |
ICWS | 3 |
| 2004 | A Distribution Design Methodology for Object DBMS
Fernanda Baião, Marta Mattoso, Gerson Zaverucha |
Distributed Parallel Databases | 2 |
| 2003 | Applying Theory Revision to the Design of Distributed Databases
Fernanda Baião, Marta Mattoso, Jude W. Shavlik, Gerson Zaverucha |
ILP | 2 |
| 2002 | Estimating Costs of Path Expression Evaluation in Distributed Object Databases
Gabriela Ruberg, Fernanda Baião, Marta Mattoso |
DEXA | 3 |
| 2002 | An Architecture for Managing Distributed Scientific ResourcesabstractThere are many examples where cooperation among scientists takes place by exchanging scientific resources, such as data, programs and mathematical models. This is particularly true for environmental applications. Finding the right resource to apply in an environmental problem is a difficult task. Usually, this decision is based on previous experience. Scientists have to cooperate in order to solve such problems. To facilitate the exchange, reuse and dissemination of information we propose an architecture for managing distributed scientific resources. Our proposal combines a mediation-based heterogeneous distributed database system and an enhanced metadata support system for effective management of distributed scientific models and data. Maria Cláudia Cavalcanti, Marta Mattoso, Maria Luiza M. Campos, Eric Simon, François Llirbat |
SSDBM | 2 |
| 1999 | Mining a large database with a parallel database serverabstractData mining is a data-intensive computation activity. Parallel processing has often been used in data mining algorithms. However, when data do not fit in memory, some solutions do not apply and a database system may be required rather than flat files. Most of the implementations use the database system loosely coupled with the data mining techniques. Hence, the database system only issues queries to be processed on the client machine. In this work, we address the data consuming activities through parallel processing on a database server providing a tight integration with data mining techniques. Experimental results showing the potential benefits of this integration were obtained. Despite the difficulties in processing a complex application, we extracted rules and obtained high performance on all the data-intensive activities such as the construction of the decision tree, pruning and rule extraction. Mauro Sousa, Marta Mattoso, Nelson F. F. Ebecken |
Intell. Data Anal. | 2 |
| 1998 | Towards an Inductive Design of Distributed Object Oriented DatabasesabstractCooperative information systems (CIS) often consist of applications that access shared resources such as databases. Since centralized systems may have a great impact on the system performance, parallel and distribution techniques are needed for attaining scalability. Distributed databases are, then, crucial for the development of cooperative applications. However, in order to improve performance, it is very important to design information distribution properly, which is the goal of distribution design. Considering the various difficulties embedded in the design of distributed object oriented databases, this work presents an algorithm to assist distribution designers in their task. The analysis algorithm indicates the most adequate fragmentation technique (vertical, horizontal or mixed) for each class in the database schema, and we propose the use of a machine learning method-inductive logic programming-to uncover some implicit issues to be considered in the distribution design, thus revising the proposed analysis algorithm. Fernanda Baião, Marta Mattoso, Gerson Zaverucha |
CoopIS | 2 |