EDBT 2026 Demo / reviewers in the wild / expert
Eduardo S. Ogasawara
dblp:24/788
· DBLP profile ↗
42ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-0466-0626ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 3 since 2021Systems, architecture and hardware · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 8Software engineering, systems software and programming languages · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards a Spatiotemporal Fusion Approach to Precipitation NowcastingabstractWith the increasing availability of meteorological data from various sensors, numerical models and reanalysis products, the need for efficient data integration methods has become paramount for improving weather forecasts and hydrometeorological studies. In this work, we propose a data fusion approach for precipitation nowcasting by integrating data from meteorological and rain gauge stations in Rio de Janeiro metropolitan area with ERA5 reanalysis data and GFS numerical weather prediction. We employ the spatiotemporal deep learning architecture called STConvS2S, leveraging a structured dataset covering a 9 x 11 grid. The study spans from January 2011 to October 2024, and we evaluate the impact of integrating three surface station systems. Among the tested configurations, the fusion-based model achieves an F1-score of 0.1964 for forecasting heavy precipitation events (greater than 25 mm/h) at a one-hour lead time. Additionally, we present an experimental assessment of the contribution of each station network and propose a refined inference strategy for precipitation nowcasting, integrating the GFS numerical weather prediction (NWP) data with in-situ observations. Felipe Curcio, Augusto José Fonseca, Rafaela C. Nascimento, Raquel Franco, Eduardo S. Ogasawara, Victor Stepanenko, Fábio Porto 0001, Mariza Ferro, Eduardo Bezerra 0002 |
FUSION | 6 |
| 2025 | Fuzzy-Based Ensemble Method for Robust Concept Drift Detection in Multivariate Time SeriesabstractConcept drift detection (CDD) is the general problem of identifying significant changes in streaming data distribution over time. Effective drift detection is important in industrial processes such as oil and gas exploration to mitigate financial losses, ensure personnel safety, and reduce environmental risks. However, current CDD methods face challenges in large-scale, multivariate datasets, where single drift detectors (DD) often fail to capture variable interdependencies. While ensemble drift detectors (EDD) are usually adopted to mitigate the adoption of a single DD, EDD may suffer when detections do not converge. This misalignment can cause voting mechanisms to neglect critical intervals with high detection rates. To address this issue, we propose a fuzzy ensemble drift detector (FEDD) that integrates unsupervised threshold voting with fuzzy logic to provide time tolerance and reconcile minor temporal misalignments in drift detection. FEDD is evaluated using the 3W dataset, a realistic public benchmark with rare undesirable real events in oil wells. The results demonstrate that FEDD outperforms existing approaches by improving detection robustness and coverage, ensuring more reliable drift detection in high-dimensional, noisy environments. Lucas Giusti Tavares, Janio Lima, Matheus Melo, Chao Chen 0007, Jonathan M. Garibaldi, Gabriel dos Santos Scatena, Anna Helena Reali Costa, Edson S. Gomi, Rebecca Salles, Esther Pacitti, Ismael H. F. dos Santos, Isabela Guimarães Siqueira, Diego Carvalho 0001, Rafaelli de C. Coutinho, Fábio Porto 0001, Eduardo S. Ogasawara |
IJCNN | 16 |
| 2025 | Scalable and accurate online multivariate anomaly detection
Rebecca Salles, Benoit Lange, Reza Akbarinia, Florent Masseglia, Eduardo S. Ogasawara, Esther Pacitti |
Inf. Syst. | 5 |
| 2024 | Online Event Detection in Streaming Time Series: Novel Metrics and Practical InsightsabstractOnline event detection in streaming time series is a critical task with applications across various domains. For example, the right-on-time event detection for control systems is a key for correctly addressing the issues related to the events. However, events may not be identified right after their occurrence. Depending on the monitoring solution, a time difference may exist between the event’s occurrence and detection. This problem raises research questions regarding the study of such a temporal gap. The paper introduces novel metrics (detection probability and detection lag) to address these questions. It explores the impact of configurable batches on detection performance. The experimental evaluation of diverse datasets reveals nuanced insights into the interplay between batch parameters, detection accuracy, and computational performance. Janio Lima, Lucas Giusti Tavares, Esther Pacitti, João Eduardo Ferreira, Ismael H. F. dos Santos, Isabela Guimarães Siqueira, Diego Carvalho 0001, Fábio Porto 0001, Rafaelli de C. Coutinho, Eduardo S. Ogasawara |
IJCNN | 10 |
| 2024 | REMD: A Novel Hybrid Anomaly Detection Method Based on EMD and ARIMAabstractAnomalies are defined as behavioral deviations from expected patterns and pose challenges to identify them. Anomaly detection is a fundamental activity of time series analysis. It enables informed decision-making in many control and monitoring activities, such as healthcare, water quality, seismic reflection analysis, and oil exploration. Many anomaly detection methods exist, but choosing the appropriate methods is complex due to the intrinsic nature of the time series. There is a demand for robust and adaptable anomaly detection methods. This paper introduces Refined Empirical Mode Decomposition (REMD) as a hybrid approach addressing this need, integrating Empirical Mode Decomposition (EMD) and Autoregressive Integrated Moving Average (ARIMA) models. REMD's design aims to optimize the strengths of both methods and overcome their limitations. It is evaluated against state-of-the-art methods on diverse datasets. It demonstrates superior performance, with up to three times better F1 score. Jéssica Souza, Ellen Paixão Silva, Fernando Fraga, Laís Baroni, Ronaldo Fernandes Santos Alves, Kele T. Belloze, Joel A. F. dos Santos, Eduardo Bezerra 0002, Fábio Porto 0001, Eduardo S. Ogasawara |
IJCNN | 10 |
| 2023 | Towards accurate recommendations of merge conflicts resolution strategies
Paulo Elias, Heleno de S. Campos Junior, Eduardo S. Ogasawara, Leonardo Murta 0001 |
Inf. Softw. Technol. | 3 |
| 2022 | Forward and Backward Inertial Anomaly Detector: A Novel Time Series Event Detection MethodabstractTime series event detection is related to studying methods for detecting observations in a series with special meaning. These observations differ from the expected behavior of the data set. In data streaming scenarios, it is possible to observe an increase in the speed of data generation in time series. Therefore, adapting to time series changes becomes crucial. Thus, identifying events associated with these changes is essential for timely and correct decision-making. Although there are many methods to detect events, it is still possible to have difficulties detecting them correctly, particularly those associated with concept drift. In order to fill the gap in the literature, this work proposes a new method, named Forward and Backward Inertial Anomaly Detector (FBIAD), for detecting events in time series. It contributes by analyzing surrounding inertia around observations. FBIAD outperformed other methods both in accuracy and elapsed time. Janio Lima, Rebecca Salles, Fábio Porto 0001, Rafaelli de C. Coutinho, Pedro Alpis, Luciana E. G. Escobar, Esther Pacitti, Eduardo S. Ogasawara |
IJCNN | 8 |
| 2022 | A horizontal partitioning-based method for frequent pattern mining in transport timetableabstractAbstract Analysing transport timetables is an important task, as it brings the opportunity to discover which routes commonly lead to delays. Frequent pattern mining is a technique used to support such type of discovery. However, functional dependencies are intrinsic properties present in timetables, particularly related to attributes derived from the origin–destination matrix. Such functional dependencies compromise the search for patterns in timetables in both the number of association rules (ARs) generated and the computational cost. Several of these ARs refer to the same information. Redundancy removal techniques can reduce the number of ARs. However, these techniques are designed to be used after mining finishes, which increases the computational cost of finding useful ARs. This work presents timetable pattern mining (T‐mine), a novel method for frequent pattern mining that improves knowledge discovery in timetables. We evaluated T‐mine using Brazilian Flight Data and compared T‐mine with the direct application of frequent pattern mining approaches with and without functional dependencies. Our experiments indicate that T‐mine is about one order magnitude faster than other methods with functional dependencies. Claudio Teixeira, Luana Fragoso, Marta Mattoso, Diego Carvalho 0001, Eduardo Bezerra 0002, Jorge Soares 0001, Glauco Fiorott Amorim, Eduardo S. Ogasawara |
Expert Syst. J. Knowl. Eng. | 8 |
| 2022 | TSPred: A framework for nonstationary time series prediction
Rebecca Salles, Esther Pacitti, Eduardo Bezerra 0002, Fábio Porto 0001, Eduardo S. Ogasawara |
Neurocomputing | 5 |
| 2021 | DJEnsemble: a Cost-Based Selection and Allocation of a Disjoint Ensemble of Spatio-temporal ModelsabstractConsider a set of black-box models – each of them independently trained on a different dataset – answering the same predictive spatio-temporal query. Being built in isolation, each model traverses its own life-cycle until it is deployed to production, learning data patterns from different datasets and facing independent hyper-parameter tuning. In order to answer the query, the set of black-box predictors has to be ensembled and allocated to the spatio-temporal query region. However, computing an optimal ensemble is a complex task that involves selecting the appropriate models and defining an effective allocation strategy that maps the models to the query region. In this paper we present DJEnsemble, a cost-based strategy for the automatic selection and allocation of a disjoint ensemble of black-box predictors to answer predictive spatio-temporal queries. We conduct a set of extensive experiments that evaluate DJEnsemble and highlight its efficiency, selecting model ensembles that are almost as efficient as the optimal solution. When compared against the traditional ensemble approach, DJEnsemble achieves up to 4X improvement in execution time and almost 9X improvement in prediction accuracy. Rafael S. Pereira 0001, Yania Molina Souto, Anderson Chaves da Silva, Rocío Zorilla, Brian Tsan, Florin Rusu, Eduardo S. Ogasawara, Artur Ziviani, Fábio Porto 0001 |
SSDBM | 7 |
| 2021 | STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for weather forecasting
Rafaela C. Nascimento, Yania Molina Souto, Eduardo S. Ogasawara, Fábio Porto 0001, Eduardo Bezerra 0002 |
Neurocomputing | 3 |
| 2020 | Spatial-time motifs discoveryabstractDiscovering motifs in time series data has been widely explored. Various techniques have been developed to tackle this problem. However, when it comes to spatial-time series, a clear gap can be observed according to the literature review. This paper tackles such a gap by presenting an approach to discover and rank motifs in spatial-time series, denominated Combined Series Approach (CSA). CSA is based on partitioning the spatial-time series into blocks. Inside each block, subsequences of spatial-time series are combined in a way that hash-based motif discovery algorithm is applied. Motifs are validated according to both temporal and spatial constraints. Later, motifs are ranked according to their entropy, the number of occurrences, and the proximity of their occurrences. The approach was evaluated using both synthetic and seismic datasets. CSA outperforms traditional methods designed only for time series. CSA was also able to prioritize motifs that were meaningful both in the context of synthetic data and also according to seismic specialists. Heraldo Borges, Murillo Dutra, Amin Bazaz, Rafaelli de C. Coutinho, Fabio Perosi, Fábio Porto 0001, Florent Masseglia, Esther Pacitti, Eduardo S. Ogasawara |
Intell. Data Anal. | 9 |
| 2020 | An analysis of malaria in the Brazilian Legal Amazon using divergent association rules
Laís Baroni, Rebecca Salles, Samella Salles, Gustavo Paiva Guedes, Fábio Porto 0001, Eduardo Bezerra 0002, Christovam Barcellos, Marcel Pedroso, Eduardo S. Ogasawara |
J. Biomed. Informatics | 9 |
| 2019 | Nonstationary time series transformation methods: An experimental review
Rebecca Salles, Kele T. Belloze, Fábio Porto 0001, Pedro H. Gonzalez, Eduardo S. Ogasawara |
Knowl. Based Syst. | 5 |
| 2018 | Control and Security System for Classroom Access Based on Facial RecognitionabstractThis work presents a facial recognition system based on spectral analysis. With this system, it is possible to provide greater security of access to a classroom, avoiding irregularities, such as falsified signatures on the presence sheet, or use of adulterated identities. It applies an image recognition process in which it seeks to extract relevant information from an image, then encode and compare it with other facial data stored in an image database. This information of the images represents a set of characteristics that present the variations between the images of the faces collected by the system and those contained in the image database. The Facial Recognition System is composed of two processing modules: training and recognition. It was applied in a high school classroom to evaluate the usefulness and accuracy of these algorithms for people recognition. Cedric Monteiro, Eduardo S. Ogasawara, Laercio Gonçalves, João Roberto de Toledo Quadros |
CLEI | 2 |
| 2018 | Evaluating the Complementarity of Communication Tools for Learning Platforms
Leonardo Carvalho, Laura Assis, Leonardo Silva de Lima, Eduardo Bezerra 0002, Gustavo Paiva Guedes, Artur Ziviani, Fábio Porto 0001, Rafael Garcia Barbastefano, Eduardo S. Ogasawara |
CSEDU (2) | 9 |
| 2018 | Orthographic Educational Game for Portuguese Language Countries
Paula Chaves, Luan Paschoal, Tauan Velasco, Tiago Bento, Julliany S. Brandão, Carlos Schocair, João Roberto de Toledo Quadros, Talita Oliveira, Eduardo S. Ogasawara |
CSEDU (2) | 9 |
| 2018 | Discovering Tight Space-Time Sequences
Riccardo Campisano, Heraldo Borges, Fábio Porto 0001, Fabio Perosi, Esther Pacitti, Florent Masseglia, Eduardo S. Ogasawara |
DaWaK | 7 |
| 2018 | On Evaluating Data Preprocessing Methods for Machine Learning Models for Flight DelaysabstractFlight delays cause various inconveniences for airlines, airports, and passengers. According to data provided by the Brazilian National Civil Aviation Agency (ANAC), between 2009 and 2015, about 22% of domestic flights made in Brazil were delayed by more than 15 minutes. The prediction of these delays is fundamental to mitigate their occurrence and optimize the decision-making process of an air transport system. Particularly, airlines, airports, and users may be more interested in when delays are likely to occur than the accurate prediction of the absence of delays. This paper focuses on the unbalanced distribution of the classes of delay (presence and absence) by performing an experimental evaluation of several preprocessing methods for the development of machine-learning flight delay classification models. Those models were built from a dataset that integrates national flight operations with meteorological conditions of airports. Our results indicate the models that applied the balancing techniques performed much better in predicting the occurrence of delays, getting about 60% of hits. Leonardo Moreira, Christofer Dantas, Leonardo Oliveira, Jorge Soares 0001, Eduardo S. Ogasawara |
IJCNN | 5 |
| 2018 | Point pattern search in big dataabstractConsider a set of points P in space with at least some of the pairwise distances specified. Given this set P, consider the following three kinds of queries against a database D of points : (i) pure constellation query: find all sets S in D of size |P| that exactly match the pairwise distances within P up to an additive error ϵ; (ii) isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f for which the distances between pairs in S exactly match f times the distances between corresponding pairs of P up to an additive ϵ; (iii) non-isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f and for at least some pairs of points, a maximum stretch factor mi,j > 1 such that (f X mi,jXdist(pi, pj))+ϵ > dist(si,sj) > (f X dist(pi, pj)) - ϵ. Finding matches to such queries has applications to spatial data in astronomical, seismic, and any domain in which (approximate, scale-independent) geometrical matching is required. Answering the isotropic and non-isotropic queries is challenging because scale factors and stretch factors may take any of an infinite number of values. This paper proposes practically efficient sequential and distributed algorithms for pure, isotropic, and non-isotropic constellation queries. As far as we know, this is the first work to address isotropic and non-isotropic queries. Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Alberto Krone-Martins, Patrick Valduriez, Dennis E. Shasha |
SSDBM | 3 |
| 2017 | Pre-processing and Indexing Techniques for Constellation Queries in Big Data
Amir Khatibi, Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Patrick Valduriez, Dennis E. Shasha |
DaWaK | 4 |
| 2017 | A framework for benchmarking machine learning methods using linear models for univariate time series predictionabstractTime series prediction has been attracting interest of researchers due to its increasing importance in decision-making activities in many fields of knowledge. The demand for better accuracy in time series prediction furthered the arising of many machine learning time series prediction methods (MLM). Choosing a suitable method for a particular dataset is a challenge and demands established benchmark methods (BM) for performance assessment. Suppose a particular BM is selected, and an experimental comparison is made with a particular MLM. If the latter does not provide better prediction results for the same dataset, this indicates that some improvements are needed for the MLM. Regarding this matter, adopting a well-established, easy to interpret, and tuned BM is desirable. This paper presents a framework for systematic benchmarking some MLM against well-known Linear Methods (LM), namely Polynomial Regression and models in the ARIMA family, used as BM for univariate time series prediction. We implemented such a framework within the R-Package named TSPred. This implementation was evaluated using a wide number of datasets from past prediction competitions. The results show that fittest LM provided by TSPred are adequate BM for univariate time series predictions. Rebecca Salles, Laura Assis, Gustavo Paiva Guedes, Eduardo Bezerra 0002, Fábio Porto 0001, Eduardo S. Ogasawara |
IJCNN | 6 |
| 2017 | Deriving scientific workflows from algebraic experiment lines: A practical approach
Anderson Marinho, Daniel de Oliveira 0001, Eduardo S. Ogasawara, Vítor Silva 0003, Kary A. C. S. Ocaña, Leonardo Murta 0001, Vanessa Braganholo, Marta Mattoso |
Future Gener. Comput. Syst. | 3 |
| 2016 | Exploring machine learning methods for the Star/Galaxy Separation ProblemabstractFor recent or planned deep astronomical surveys, it is important to tell stars and galaxies apart, a task known as Star/Galaxy Separation Problem (SGSP). At faint magnitudes, the separation between pointy and extended sources is fuzzy, which makes SGSP a hard task. This problem is even harder for large surveys like Dark Energy Survey (DES) and, in a near future, the Large Synoptic Survey Telescope (LSST) due to their large data volume. Hence, the search for classification methods that are both accurate and efficient is highly relevant. In this work, we present a comparative analysis of several machine learning methods targeted at solving the SGSP at faint magnitudes. In order to train the classification models, the COSMOS survey was used. We use machine learning methods as distinct as artificial neural networks, k nearest-neighbor, Support Vector Machines, Random Forests and Naive Bayes. The exploratory process was modeled as data centric workflow. The workflow was implemented on top of Hadoop framework and was used to find the best parameter values for each classification method we considered, of which neural networks and random forest present superior performance. Eduardo Machado, Marcello Serqueira, Eduardo S. Ogasawara, Ricardo Ogando, Marcio A. G. Maia, Luiz Nicolaci da Costa, Riccardo Campisano, Gustavo Paiva Guedes, Eduardo Bezerra 0002 |
IJCNN | 3 |
| 2016 | Discovering top-k non-redundant clusterings in attributed graphs
Gustavo Paiva Guedes, Eduardo S. Ogasawara, Eduardo Bezerra 0002, Geraldo Xexéo |
Neurocomputing | 2 |
| 2015 | Dynamic steering of HPC scientific workflows: A survey
Marta Mattoso, Jonas Dias, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Flavio Costa, Felipe Horta, Vítor Silva 0003, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 4 |
| 2013 | Algebraic dataflows for big data analysisabstractAnalyzing big data requires the support of dataflows with many activities to extract and explore relevant information from the data. Recent approaches such as Pig Latin propose a high-level language to model such dataflows. However, the dataflow execution is typically delegated to a MapRe-duce implementation such as Hadoop, which does not follow an algebraic approach, thus it cannot take advantage of the optimization opportunities of PigLatin algebra. In this paper, we propose an approach for big data analysis based on algebraic workflows, which yields optimization and parallel execution of activities and supports user steering using provenance queries. We illustrate how a big data processing dataflow can be modeled using the algebra. Through an experimental evaluation using real datasets and the execution of the dataflow with Chiron, an engine that supports our algebra, we show that our approach yields performance gains of up to 19.6% using algebraic optimizations in the dataflow and up to 39.1% of time saved on a user steering scenario. Jonas Dias, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 2 |
| 2013 | Chiron: a parallel engine for algebraic scientific workflowsabstractSUMMARY Large‐scale scientific experiments based on computer simulations are typically modeled as scientific workflows, which eases the chaining of different programs. These scientific workflows are defined, executed, and monitored by scientific workflow management systems (SWfMS). As these experiments manage large amounts of data, it becomes critical to execute them in high‐performance computing environments, such as clusters, grids, and clouds. However, few SWfMS provide parallel support. The ones that do so are usually labor‐intensive for workflow developers and have limited primitives to optimize workflow execution. To address these issues, we developed workflow algebra to specify and enable the optimization of parallel execution of scientific workflows. In this paper, we show how the workflow algebra is efficiently implemented in Chiron, an algebraic based parallel scientific workflow engine. Chiron has a unique native distributed provenance mechanism that enables runtime queries in a relational database. We developed two studies to evaluate the performance of our algebraic approach implemented in Chiron; the first study compares Chiron with different approaches, whereas the second one evaluates the scalability of Chiron. By analyzing the results, we conclude that Chiron is efficient in executing scientific workflows, with the benefits of declarative specification and runtime provenance support. Copyright © 2013 John Wiley & Sons, Ltd. Eduardo S. Ogasawara, Jonas Dias, Vítor Silva 0003, Fernando Seabra Chirigati, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Designing a parallel cloud based comparative genomics workflow to improve phylogenetic analyses
Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
Future Gener. Comput. Syst. | 4 |
| 2013 | Performance evaluation of parallel strategies in public clouds: A study with phylogenomic workflows
Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Jonas Dias, João Carlos de A. R. Gonçalves, Fernanda Baião, Marta Mattoso |
Future Gener. Comput. Syst. | 3 |
| 2012 | Discovering drug targets for neglected diseases using a pharmacophylogenomic cloud workflowabstractIllnesses caused by parasitic protozoan are a research priority. A representative group of these illnesses is the commonly known as Neglected Tropical Diseases (NTD). NTD specially attack low socioeconomic population around the world and new anti-protozoan inhibitors are needed and several drug discovery projects focus on researching new drug targets. Pharmacophylogenomics is a novel bioinformatics field that aims at reducing the time and the financial cost of the drug discovery process. Pharmacophylogenomic analyses are applied mainly in the early stages of the research phase in drug discovery. Pharmacophylogenomic analysis executes several bioinformatics programs in a coherent flow to identify homologues sequences, construct phylogenetic trees and execute evolutionary and structural experiments. This way, it can be modeled as scientific workflows. Pharmacophylogenomic analysis workflows are complex, computing and data intensive and may execute during weeks. This way, it benefits from parallel execution. We propose SciPPGx, a scientific workflow that aims at providing thorough inferring support for pharmacophylogenomic hypotheses. SciPPGx is executed in parallel in a cloud using SciCumulus workflow engine. Experiments show that SciPPGx considerably reduces the total execution time up to 97.1% when compared to a sequential execution. We also present representative biological results taking advantage of the inference covering several related bioinformatics overviews. Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
eScience | 4 |
| 2012 | ProvManager: a provenance management system for scientific workflowsabstractSUMMARY Running scientific workflows in distributed and heterogeneous environments has been a motivating approach for provenance management, which is loosely coupled to the workflow execution engine. This kind of approach is interesting because it allows both storage and access to provenance data in a homogeneous way, even in an environment where different workflow management systems work together. However, current approaches overload scientists with many ad hoc tasks, such as script adaptations and implementations of extra functionalities to provide provenance independence. This paper proposes ProvManager, a provenance management approach that eases the gathering, storage, and analysis of provenance information in a distributed and heterogeneous environment scenario, without putting the burden of adaptations on the scientist. ProvManager leverages the provenance management at the experiment level by integrating different workflow executions from multiple workflow management systems. Copyright © 2011 John Wiley & Sons, Ltd. Anderson Marinho, Leonardo Murta 0001, Cláudia M. L. Werner, Vanessa Braganholo, Sérgio Manuel Serra da Cruz, Eduardo S. Ogasawara, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 6 |
| 2012 | An adaptive parallel execution strategy for cloud-based scientific workflowsabstractSUMMARY Many of the existing large‐scale scientific experiments modeled as scientific workflows are compute‐intensive. Some scientific workflow management systems already explore parallel techniques, such as parameter sweep and data fragmentation, to improve performance. In those systems, computing resources are used to accomplish many computational tasks in high performance environments, such as multiprocessor machines or clusters. Meanwhile, cloud computing provides scalable and elastic resources that can be instantiated on demand during the course of a scientific experiment, without requiring its users to acquire expensive infrastructure or to configure many pieces of software. In fact, because of these advantages some scientists have already adopted the cloud model in their scientific experiments. However, this model also raises many challenges. When scientists are executing scientific workflows that require parallelism, it is hard to decide a priori the amount of resources to use and how long they will be needed because the allocation of these resources is elastic and based on demand. In addition, scientists have to manage new aspects such as initialization of virtual machines and impact of data staging. SciCumulus is a middleware that manages the parallel execution of scientific workflows in cloud environments. In this paper, we introduce an adaptive approach for executing parallel scientific workflows in the cloud. This approach adapts itself according to the availability of resources during workflow execution. It checks the available computational power and dynamically tunes the workflow activity size to achieve better performance. Experimental evaluation showed the benefits of parallelizing scientific workflows using the adaptive approach of SciCumulus, which presented an increase of performance up to 47.1%. Copyright © 2011 John Wiley & Sons, Ltd. Daniel de Oliveira 0001, Eduardo S. Ogasawara, Kary A. C. S. Ocaña, Fernanda Baião, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | A Performance Evaluation of X-Ray Crystallography Scientific Workflow Using SciCumulusabstractX-ray crystallography is an important field due to its role in drug discovery and its relevance in bioinformatics experiments of comparative genomics, phylogenomics, evolutionary analysis, ortholog detection, and three-dimensional structure determination. Managing these experiments is a challenging task due to the orchestration of legacy tools and the management of several variations of the same experiment. Workflows can model a coherent flow of activities that are managed by scientific workflow management systems (SWfMS). Due to the huge amount of variations of the workflow to be explored (parameters, input data) it is often necessary to execute X-ray crystallography experiments in High Performance Computing (HPC) environments. Cloud computing is well known for its scalable and elastic HPC model. In this paper, we present a performance evaluation for the X-ray crystallography workflow defined by the PC4 (Provenance Challenge series). The workflow was executed using the SciCumulus middleware at the Amazon EC2 cloud environment. SciCumulus is a layer for SWfMS that offers support for the parallel execution of scientific workflows in cloud environments with provenance mechanisms. Our results reinforce the benefits (total execution time × monetary cost) of parallelizing the X-ray crystallography workflow using SciCumulus. The results show a consistent way to execute X-ray crystallography workflows that need HPC using cloud computing. The evaluated workflow shares features of many scientific workflows and can be applied to other experiments. Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Jonas Dias, Fernanda Baião, Marta Mattoso |
IEEE CLOUD | 3 |
| 2011 | Optimizing Phylogenetic Analysis Using SciHmm Cloud-based Scientific WorkflowabstractPhylogenetic analysis and multiple sequence alignment (MSA) are closely related bioinformatics fields. Phylogenetic analysis makes extensive use of MSA in the construction of phylogenetic trees, which are used to infer the evolutionary relationships between homologous genes. These bioinformatics experiments are usually modeled as scientific workflows. There are many alternative workflows that use different MSA methods to conduct phylogenetic analysis and each one can produce MSA with different quality. Scientists have to explore which MSA method is the most suitable for their experiments. However, workflows for phylogenetic analysis are both computational and data intensive and they may run sequentially during weeks. Although there any many approaches that parallelize these workflows, exploring all MSA methods many become a burden and expensive task. If scientists know the most adequate MSA method a priori, it would spare time and money. To optimize the phylogenetic analysis workflow, we propose in this paper SciHmm, a bioinformatics scientific workflow based in profile hidden Markov models (pHMMs) that aims at determining the most suitable MSA method for a phylogenetic analysis prior than executing the phylogenetic workflow. SciHmm is also executed in parallel in a cloud environment using SciCumulus middleware. The results demonstrated that optimizing a phylogenetic analysis using SciHmm considerably reduce the total execution time of phylogenetic analysis (up to 80%). This optimization also demonstrates that the biological results presented more quality. In addition, the parallel execution of SciHmm demonstrates that this kind of bioinformatics workflow is suitable to be executed in the cloud. Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
eScience | 4 |
| 2011 | Many task computing for orthologous genes identification in protozoan genomes using HydraabstractSUMMARY One of the main advantages of using a scientific workflow management system (SWfMS) is to orchestrate data flows among scientific activities and register provenance of the whole workflow execution. Nevertheless, the execution control of distributed activities in high performance computing environments by SWfMS presents challenges such as steering control and provenance gathering. Such challenges may become a complex task to be accomplished in bioinformatics experiments, particularly in Many Task Computing scenarios. This paper presents a data parallelism solution for a bioinformatics experiment supported by Hydra, a middleware that bridges SWfMS and high performance computing to enable workflow parallelization with provenance gathering. Hydra Many Task Computing parallelization strategies can be registered and reused. Using Hydra, provenance may also be uniformly gathered. We have evaluated Hydra using an Orthologous Gene Identification workflow. Experimental results show that a systematic approach for distributing parallel activities is viable, sparing scientist time and diminishing operational errors, with the additional benefits of distributed provenance support. Copyright © 2011 John Wiley & Sons, Ltd. Fábio Coutinho, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Vanessa Braganholo, Alexandre A. B. Lima, Alberto M. R. Dávila, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | An Algebraic Approach for Data-Centric Scientific Workflows
Eduardo S. Ogasawara, Daniel de Oliveira 0001, Patrick Valduriez, Jonas Dias, Fábio Porto 0001, Marta Mattoso |
Proc. VLDB Endow. | 1 |
| 2010 | SciCumulus: A Lightweight Cloud Middleware to Explore Many Task Computing Paradigm in Scientific WorkflowsabstractMost of the large-scale scientific experiments modeled as scientific workflows produce a large amount of data and require workflow parallelism to reduce workflow execution time. Some of the existing Scientific Workflow Management Systems (SWfMS) explore parallelism techniques - such as parameter sweep and data fragmentation. In those systems, several computing resources are used to accomplish many computational tasks in homogeneous environments, such as multiprocessor machines or cluster systems. Cloud computing has become a popular high performance computing model in which (virtualized) resources are provided as services over the Web. Some scientists are starting to adopt the cloud model in scientific domains and are moving their scientific workflows (programs and data) from local environments to the cloud. Nevertheless, it is still difficult for the scientist to express a parallel computing paradigm for the workflow on the cloud. Capturing distributed provenance data at the cloud is also an issue. Existing approaches for executing scientific workflows using parallel processing are mainly focused on homogeneous environments whereas, in the cloud, the scientist has to manage new aspects such as initialization of virtualized instances, scheduling over different cloud environments, impact of data transferring and management of instance images. In this paper we propose SciCumulus, a cloud middleware that explores parameter sweep and data fragmentation parallelism in scientific workflow activities (with provenance support). It works between the SWfMS and the cloud. SciCumulus is designed considering cloud specificities. We have evaluated our approach by executing simulated experiments to analyze the overhead imposed by clouds on the workflow execution time. Daniel de Oliveira 0001, Eduardo S. Ogasawara, Fernanda Baião, Marta Mattoso |
IEEE CLOUD | 2 |
| 2010 | Data parallelism in bioinformatics workflows using HydraabstractLarge scale bioinformatics experiments are usually composed by a set of data flows generated by a chain of activities (programs or services) that may be modeled as scientific workflows. Current Scientific Workflow Management Systems (SWfMS) are used to orchestrate these workflows to control and monitor the whole execution. It is very common in bioinformatics experiments to process very large datasets. In this way, data parallelism is a common approach used to increase performance and reduce overall execution time. However, most of current SWfMS still lack on supporting parallel executions in high performance computing (HPC) environments. Additionally keeping track of provenance data in distributed environments is still an open, yet important problem. Recently, Hydra middleware was proposed to bridge the gap between the SWfMS and the HPC environment, by providing a transparent way for scientists to parallelize workflow executions while capturing distributed provenance. This paper analyzes data parallelism scenarios in bioinformatics domain and presents an extension to Hydra middleware through a specific cartridge that promotes data parallelism in bioinformatics workflows. Experimental results using workflows with BLAST show performance gains with the additional benefits of distributed provenance support. Fábio Coutinho, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Vanessa Braganholo, Alexandre A. B. Lima, Alberto M. R. Dávila, Marta Mattoso |
HPDC | 2 |
| 2010 | Adaptive Normalization: A novel data normalization approach for non-stationary time seriesabstractData normalization is a fundamental preprocessing step for mining and learning from data. However, finding an appropriated method to deal with time series normalization is not a simple task. This is because most of the traditional normalization methods make assumptions that do not hold for most time series. The first assumption is that all time series are stationary, i.e., their statistical properties, such as mean and standard deviation, do not change over time. The second assumption is that the volatility of the time series is considered uniform. None of the methods currently available in the literature address these issues. This paper proposes a new method for normalizing non-stationary heteroscedastic (with non-uniform volatility) time series. The method, named Adaptive Normalization (AN), was tested together with an Artificial Neural Network (ANN) in three forecast problems. The results were compared to other four traditional normalization methods, and showed AN improves ANN accuracy in both short- and long-term predictions. Eduardo S. Ogasawara, Leonardo C. Martinez, Daniel de Oliveira 0001, Geraldo Zimbrão, Gisele L. Pappa, Marta Mattoso |
IJCNN | 1 |
| 2009 | Neural networks cartridges for data mining on time seriesabstractNeural networks is one of the techniques used for time series analysis. The performance of neural networks is affected by some parameters such as neural network structure and the quality of data preprocessing. These parameters need to be explored in order to obtain an optimal neural network. However, the manual establishment of different neural networks configurations for selecting the best ones may be error-prone and time-consuming. This paper proposes the creation of neural networks cartridges to systematically empower neural network performance by means of data mining activities, which obtain an optimal neural network structure. The experiments conducted in this paper use stock market and exchange rate series, and show that the usage of neural network cartridges can lead to configurations that double the performance of some ad-hoc neural network configuration. Eduardo S. Ogasawara, Leonardo Murta 0001, Geraldo Zimbrão, Marta Mattoso |
IJCNN | 1 |
| 2009 | Experiment Line: Software Reuse in Scientific Workflows
Eduardo S. Ogasawara, Carlos Eduardo Paulino Silva, Leonardo Murta 0001, Cláudia M. L. Werner, Marta Mattoso |
SSDBM | 1 |