Fábio Porto 0001

dblp:p/FabioPorto · also Fábio André Machado Porto · DBLP profile ↗
← Back
48ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-4597-4832ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 21 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 19 · 6 since 2021Systems, architecture and hardware · 11 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2025 Towards a Spatiotemporal Fusion Approach to Precipitation Nowcasting
abstract
With the increasing availability of meteorological data from various sensors, numerical models and reanalysis products, the need for efficient data integration methods has become paramount for improving weather forecasts and hydrometeorological studies. In this work, we propose a data fusion approach for precipitation nowcasting by integrating data from meteorological and rain gauge stations in Rio de Janeiro metropolitan area with ERA5 reanalysis data and GFS numerical weather prediction. We employ the spatiotemporal deep learning architecture called STConvS2S, leveraging a structured dataset covering a 9 x 11 grid. The study spans from January 2011 to October 2024, and we evaluate the impact of integrating three surface station systems. Among the tested configurations, the fusion-based model achieves an F1-score of 0.1964 for forecasting heavy precipitation events (greater than 25 mm/h) at a one-hour lead time. Additionally, we present an experimental assessment of the contribution of each station network and propose a refined inference strategy for precipitation nowcasting, integrating the GFS numerical weather prediction (NWP) data with in-situ observations.
Felipe Curcio, Augusto José Fonseca, Rafaela C. Nascimento, Raquel Franco, Eduardo S. Ogasawara, Victor Stepanenko, Fábio Porto 0001, Mariza Ferro, Eduardo Bezerra 0002
FUSION8
2025 Fuzzy-Based Ensemble Method for Robust Concept Drift Detection in Multivariate Time Series
abstract
Concept drift detection (CDD) is the general problem of identifying significant changes in streaming data distribution over time. Effective drift detection is important in industrial processes such as oil and gas exploration to mitigate financial losses, ensure personnel safety, and reduce environmental risks. However, current CDD methods face challenges in large-scale, multivariate datasets, where single drift detectors (DD) often fail to capture variable interdependencies. While ensemble drift detectors (EDD) are usually adopted to mitigate the adoption of a single DD, EDD may suffer when detections do not converge. This misalignment can cause voting mechanisms to neglect critical intervals with high detection rates. To address this issue, we propose a fuzzy ensemble drift detector (FEDD) that integrates unsupervised threshold voting with fuzzy logic to provide time tolerance and reconcile minor temporal misalignments in drift detection. FEDD is evaluated using the 3W dataset, a realistic public benchmark with rare undesirable real events in oil wells. The results demonstrate that FEDD outperforms existing approaches by improving detection robustness and coverage, ensuring more reliable drift detection in high-dimensional, noisy environments.
Lucas Giusti Tavares, Janio Lima, Matheus Melo, Chao Chen 0007, Jonathan M. Garibaldi, Gabriel dos Santos Scatena, Anna Helena Reali Costa, Edson S. Gomi, Rebecca Salles, Esther Pacitti, Ismael H. F. dos Santos, Isabela Guimarães Siqueira, Diego Carvalho 0001, Rafaelli de C. Coutinho, Fábio Porto 0001, Eduardo S. Ogasawara
IJCNN15
2024 Discovering Denial Constraints in Dynamic Datasets
abstract
Denial constraints (DCs) are data dependencies with high expressive power, offering great flexibility for modeling data quality rules. Specifying DCs manually is problematic, as the required domain expertise is expensive and scarce. Moreover, database updates can invalidate DCs thought to hold and simultaneously uncover new DCs. This fact leads to burdensome scenarios where experts must often revisit DC specifications. Several algorithms have been devised to discover DCs from data, among which only one considers DC discovery on data updates. However, that solution underperforms in many scenarios due to long runtime and excessive memory use. Also, it targets database inserts only, so no previous solution covers deletions. This paper proposes an efficient and flexible algorithm that covers the earlier limitations regarding performance and scope. The algorithm maintains small-footprint intermediate structures during database updates and a method that exploits the changes in this intermediate to update the DCs incrementally. The results of our extensive experimental evaluation show that our algorithm is orders of magnitude faster than the existing one, with much better scalability in the size of the data updates.
Eduardo H. M. Pena, Fábio Porto 0001, Felix Naumann
ICDE2
2024 Online Event Detection in Streaming Time Series: Novel Metrics and Practical Insights
abstract
Online event detection in streaming time series is a critical task with applications across various domains. For example, the right-on-time event detection for control systems is a key for correctly addressing the issues related to the events. However, events may not be identified right after their occurrence. Depending on the monitoring solution, a time difference may exist between the event’s occurrence and detection. This problem raises research questions regarding the study of such a temporal gap. The paper introduces novel metrics (detection probability and detection lag) to address these questions. It explores the impact of configurable batches on detection performance. The experimental evaluation of diverse datasets reveals nuanced insights into the interplay between batch parameters, detection accuracy, and computational performance.
Janio Lima, Lucas Giusti Tavares, Esther Pacitti, João Eduardo Ferreira, Ismael H. F. dos Santos, Isabela Guimarães Siqueira, Diego Carvalho 0001, Fábio Porto 0001, Rafaelli de C. Coutinho, Eduardo S. Ogasawara
IJCNN8
2024 REMD: A Novel Hybrid Anomaly Detection Method Based on EMD and ARIMA
abstract
Anomalies are defined as behavioral deviations from expected patterns and pose challenges to identify them. Anomaly detection is a fundamental activity of time series analysis. It enables informed decision-making in many control and monitoring activities, such as healthcare, water quality, seismic reflection analysis, and oil exploration. Many anomaly detection methods exist, but choosing the appropriate methods is complex due to the intrinsic nature of the time series. There is a demand for robust and adaptable anomaly detection methods. This paper introduces Refined Empirical Mode Decomposition (REMD) as a hybrid approach addressing this need, integrating Empirical Mode Decomposition (EMD) and Autoregressive Integrated Moving Average (ARIMA) models. REMD's design aims to optimize the strengths of both methods and overcome their limitations. It is evaluated against state-of-the-art methods on diverse datasets. It demonstrates superior performance, with up to three times better F1 score.
Jéssica Souza, Ellen Paixão Silva, Fernando Fraga, Laís Baroni, Ronaldo Fernandes Santos Alves, Kele T. Belloze, Joel A. F. dos Santos, Eduardo Bezerra 0002, Fábio Porto 0001, Eduardo S. Ogasawara
IJCNN9
2024 HIHISIV: a database of gene expression in HIV and SIV host immune response
abstract
In the battle of the host against lentiviral pathogenesis, the immune response is crucial. However, several questions remain unanswered about the interaction with different viruses and their influence on disease progression. The simian immunodeficiency virus (SIV) infecting nonhuman primates (NHP) is widely used as a model for the study of the human immunodeficiency virus (HIV) both because they are evolutionarily linked and because they share physiological and anatomical similarities that are largely explored to understand the disease progression. The HIHISIV database was developed to support researchers to integrate and evaluate the large number of transcriptional data associated with the presence/absence of the pathogen (SIV or HIV) and the host response (NHP and human). The datasets are composed of microarray and RNA-Seq gene expression data that were selected, curated, analyzed, enriched, and stored in a relational database. Six query templates comprise the main data analysis functions and the resulting information can be downloaded. The HIHISIV database, available at https://hihisiv.github.io , provides accurate resources for browsing and visualizing results and for more robust analyses of pre-existing data in transcriptome repositories.
Raquel Lopes Costa, Luiz M. R. Gadelha Jr., Mirela D'arc, Marcelo Ribeiro-Alves, David L. Robertson, Jean-Marc Schwartz, Marcelo A. Soares, Fábio Porto 0001
BMC Bioinform.8
2022 Forward and Backward Inertial Anomaly Detector: A Novel Time Series Event Detection Method
abstract
Time series event detection is related to studying methods for detecting observations in a series with special meaning. These observations differ from the expected behavior of the data set. In data streaming scenarios, it is possible to observe an increase in the speed of data generation in time series. Therefore, adapting to time series changes becomes crucial. Thus, identifying events associated with these changes is essential for timely and correct decision-making. Although there are many methods to detect events, it is still possible to have difficulties detecting them correctly, particularly those associated with concept drift. In order to fill the gap in the literature, this work proposes a new method, named Forward and Backward Inertial Anomaly Detector (FBIAD), for detecting events in time series. It contributes by analyzing surrounding inertia around observations. FBIAD outperformed other methods both in accuracy and elapsed time.
Janio Lima, Rebecca Salles, Fábio Porto 0001, Rafaelli de C. Coutinho, Pedro Alpis, Luciana E. G. Escobar, Esther Pacitti, Eduardo S. Ogasawara
IJCNN3
2022 TSPred: A framework for nonstationary time series prediction
Rebecca Salles, Esther Pacitti, Eduardo Bezerra 0002, Fábio Porto 0001, Eduardo S. Ogasawara
Neurocomputing4
2022 Fast Algorithms for Denial Constraint Discovery
abstract
Denial constraints (DCs) are an integrity constraint formalism widely used to detect inconsistencies in data. Several algorithms have been devised to discover DCs from data, as manually specifying them is burdensome and, worse yet, error-prone. The existing algorithms follow two basic steps: building an intermediate data structure from records, then enumerating the DCs from that intermediate. However, current algorithms are often inefficient in computing these intermediates. Also, it is still unclear which enumeration algorithm performs best since some of the available algorithms have not yet been compared to each other. In response, we present a set of new algorithms with improved design choices. We introduce a parallel pipeline for rapidly computing the intermediate using custom data representations, algorithms, and indexes. For DC enumeration, we propose an inverted index, pruning, and parallel search strategies. We present hybrid approaches that integrate our techniques with previous enumeration algorithms, improving their performance in many scenarios. Our experimental study shows that the proposed DC discovery algorithms are consistently much faster (up to an order of magnitude) than the current state-of-the-art.
Eduardo H. M. Pena, Fábio Porto 0001, Felix Naumann
Proc. VLDB Endow.2
2021 DJEnsemble: a Cost-Based Selection and Allocation of a Disjoint Ensemble of Spatio-temporal Models
abstract
Consider a set of black-box models – each of them independently trained on a different dataset – answering the same predictive spatio-temporal query. Being built in isolation, each model traverses its own life-cycle until it is deployed to production, learning data patterns from different datasets and facing independent hyper-parameter tuning. In order to answer the query, the set of black-box predictors has to be ensembled and allocated to the spatio-temporal query region. However, computing an optimal ensemble is a complex task that involves selecting the appropriate models and defining an effective allocation strategy that maps the models to the query region. In this paper we present DJEnsemble, a cost-based strategy for the automatic selection and allocation of a disjoint ensemble of black-box predictors to answer predictive spatio-temporal queries. We conduct a set of extensive experiments that evaluate DJEnsemble and highlight its efficiency, selecting model ensembles that are almost as efficient as the optimal solution. When compared against the traditional ensemble approach, DJEnsemble achieves up to 4X improvement in execution time and almost 9X improvement in prediction accuracy.
Rafael S. Pereira 0001, Yania Molina Souto, Anderson Chaves da Silva, Rocío Zorilla, Brian Tsan, Florin Rusu, Eduardo S. Ogasawara, Artur Ziviani, Fábio Porto 0001
SSDBM9
2021 Towards optimizing the execution of spark scientific workflows using machine learning-based parameter tuning
abstract
Summary In the last few years, Apache Spark has become a de facto the standard framework for big data systems on both industry and academy projects. Spark is used to execute compute‐ and data‐intensive workflows in distinct areas like biology and astronomy. Although Spark is an easy‐to‐install framework, it has more than one hundred parameters to be set, besides domain‐specific parameters of each workflow. In this way, to execute Spark‐based workflows efficiently, the user has to fine‐tune a myriad of Spark and workflow parameters (eg, partitioning strategy, the average size of a DNA sequence, etc.). This configuration task cannot be manually performed in a trial‐and‐error manner since it is tedious and error‐prone. This article proposes an approach that focuses on generating interpretable predictive machine learning models (ie, decision trees), and then extract useful rules (ie, patterns) from these models that can be applied to configure parameters of future executions of the workflow and Spark for nonexperts users. In the experiments presented in this article, the proposed parameter configuration approach led to better performance in processing Spark workflows. Finally, the approach introduced here reduced the number of parameters to be configured by identifying the most relevant domain‐specific ones related to the workflow performance in the predictive model.
Douglas E. M. de Oliveira, Fábio Porto 0001, Cristina Boeres, Daniel de Oliveira 0001
Concurr. Comput. Pract. Exp.2
2021 STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for weather forecasting
Rafaela C. Nascimento, Yania Molina Souto, Eduardo S. Ogasawara, Fábio Porto 0001, Eduardo Bezerra 0002
Neurocomputing4
2020 Parallel computation of PDFs on big spatial data using Spark
Ji Liu 0003, Noel Moreno Lemus, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez
Distributed Parallel Databases4
2020 BioinfoPortal: A scientific gateway for integrating bioinformatics applications on the Brazilian national high-performance computing network
Kary A. C. S. Ocaña, Marcelo Galheigo, Carla Osthoff, Luiz M. R. Gadelha Jr., Fábio Porto 0001, Antônio Tadeu A. Gomes, Daniel de Oliveira 0001, Ana Tereza Ribeiro de Vasconcelos
Future Gener. Comput. Syst.5
2020 Spatial-time motifs discovery
abstract
Discovering motifs in time series data has been widely explored. Various techniques have been developed to tackle this problem. However, when it comes to spatial-time series, a clear gap can be observed according to the literature review. This paper tackles such a gap by presenting an approach to discover and rank motifs in spatial-time series, denominated Combined Series Approach (CSA). CSA is based on partitioning the spatial-time series into blocks. Inside each block, subsequences of spatial-time series are combined in a way that hash-based motif discovery algorithm is applied. Motifs are validated according to both temporal and spatial constraints. Later, motifs are ranked according to their entropy, the number of occurrences, and the proximity of their occurrences. The approach was evaluated using both synthetic and seismic datasets. CSA outperforms traditional methods designed only for time series. CSA was also able to prioritize motifs that were meaningful both in the context of synthetic data and also according to seismic specialists.
Heraldo Borges, Murillo Dutra, Amin Bazaz, Rafaelli de C. Coutinho, Fabio Perosi, Fábio Porto 0001, Florent Masseglia, Esther Pacitti, Eduardo S. Ogasawara
Intell. Data Anal.6
2020 An analysis of malaria in the Brazilian Legal Amazon using divergent association rules
Laís Baroni, Rebecca Salles, Samella Salles, Gustavo Paiva Guedes, Fábio Porto 0001, Eduardo Bezerra 0002, Christovam Barcellos, Marcel Pedroso, Eduardo S. Ogasawara
J. Biomed. Informatics5
2019 Towards a Science Gateway for Bioinformatics: Experiences in the Brazilian System of High Performance Computing
abstract
Science gateways bring out the possibility of reproducible science as they are integrated into reusable techniques, data and workflow management systems, security mechanisms, and high performance computing (HPC). We introduce BioinfoPortal, a science gateway that integrates a suite of different bioinformatics applications using HPC and data management resources provided by the Brazilian National HPC System (SINAPAD). BioinfoPortal follows the Software as a Service (SaaS) model and the web server is freely available for academic use. The goal of this paper is to describe the science gateway and its usage, addressing challenges of designing a multiuser computational platform for parallel/distributed executions of large-scale bioinformatics applications using the Brazilian HPC resources. We also present a study of performance and scalability of some bioinformatics applications executed in the HPC environments and perform machine learning analyses for predicting features for the HPC allocation/usage that could better perform the bioinformatics applications via BioinfoPortal.
Kary A. C. S. Ocaña, Marcelo Galheigo, Carla Osthoff, Luiz M. R. Gadelha Jr., Antônio Tadeu A. Gomes, Daniel de Oliveira 0001, Fábio Porto 0001, Ana Tereza Ribeiro de Vasconcelos
CCGRID7
2019 Nonstationary time series transformation methods: An experimental review
Rebecca Salles, Kele T. Belloze, Fábio Porto 0001, Pedro H. Gonzalez, Eduardo S. Ogasawara
Knowl. Based Syst.3
2018 Evaluating the Complementarity of Communication Tools for Learning Platforms
Leonardo Carvalho, Laura Assis, Leonardo Silva de Lima, Eduardo Bezerra 0002, Gustavo Paiva Guedes, Artur Ziviani, Fábio Porto 0001, Rafael Garcia Barbastefano, Eduardo S. Ogasawara
CSEDU (2)7
2018 Discovering Tight Space-Time Sequences
Riccardo Campisano, Heraldo Borges, Fábio Porto 0001, Fabio Perosi, Esther Pacitti, Florent Masseglia, Eduardo S. Ogasawara
DaWaK3
2018 A Spatiotemporal Ensemble Approach to Rainfall Forecasting
abstract
This paper proposes a new ensemble method built upon a deep neural network architecture. We use a set of meteorological models for rain forecast as base predictors. Each meteorological model is provided to a channel of the network and, through a convolution operator, the prediction models are weighted and combined. As a result, the predicted value produced by the ensemble depends on both the spatial neighborhood and the temporal pattern. We conduct some computational experiments in order to compare our approach to other ensemble methods widely used for daily rainfall prediction. The results show that our architecture based on ConvLSTM networks is a strong candidate to solve the problem of combining predictions in a spatiotemporal context.
Yania Molina Souto, Fábio Porto 0001, Ana Maria de Carvalho Moura, Eduardo Bezerra 0002
IJCNN2
2018 Point pattern search in big data
abstract
Consider a set of points P in space with at least some of the pairwise distances specified. Given this set P, consider the following three kinds of queries against a database D of points : (i) pure constellation query: find all sets S in D of size |P| that exactly match the pairwise distances within P up to an additive error ϵ; (ii) isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f for which the distances between pairs in S exactly match f times the distances between corresponding pairs of P up to an additive ϵ; (iii) non-isotropic constellation queries: find all sets S in D of size |P| such that there exists some scale factor f and for at least some pairs of points, a maximum stretch factor mi,j > 1 such that (f X mi,jXdist(pi, pj))+ϵ > dist(si,sj) > (f X dist(pi, pj)) - ϵ. Finding matches to such queries has applications to spatial data in astronomical, seismic, and any domain in which (approximate, scale-independent) geometrical matching is required. Answering the isotropic and non-isotropic queries is challenging because scale factors and stretch factors may take any of an infinite number of values. This paper proposes practically efficient sequential and distributed algorithms for pure, isotropic, and non-isotropic constellation queries. As far as we know, this is the first work to address isotropic and non-isotropic queries.
Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Alberto Krone-Martins, Patrick Valduriez, Dennis E. Shasha
SSDBM1
2017 TARDIS: Optimal Execution of Scientific Workflows in Apache Spark
Daniel Gaspar, Fábio Porto 0001, Reza Akbarinia, Esther Pacitti
DaWaK2
2017 Pre-processing and Indexing Techniques for Constellation Queries in Big Data
Amir Khatibi, Fábio Porto 0001, João N. Rittmeyer, Eduardo S. Ogasawara, Patrick Valduriez, Dennis E. Shasha
DaWaK2
2017 A framework for benchmarking machine learning methods using linear models for univariate time series prediction
abstract
Time series prediction has been attracting interest of researchers due to its increasing importance in decision-making activities in many fields of knowledge. The demand for better accuracy in time series prediction furthered the arising of many machine learning time series prediction methods (MLM). Choosing a suitable method for a particular dataset is a challenge and demands established benchmark methods (BM) for performance assessment. Suppose a particular BM is selected, and an experimental comparison is made with a particular MLM. If the latter does not provide better prediction results for the same dataset, this indicates that some improvements are needed for the MLM. Regarding this matter, adopting a well-established, easy to interpret, and tuned BM is desirable. This paper presents a framework for systematic benchmarking some MLM against well-known Linear Methods (LM), namely Polynomial Regression and models in the ARIMA family, used as BM for univariate time series prediction. We implemented such a framework within the R-Package named TSPred. This implementation was evaluated using a wide number of datasets from past prediction competitions. The results show that fittest LM provided by TSPred are adequate BM for univariate time series predictions.
Rebecca Salles, Laura Assis, Gustavo Paiva Guedes, Eduardo Bezerra 0002, Fábio Porto 0001, Eduardo S. Ogasawara
IJCNN5
2016 A note on the complexity of the causal ordering problem
Bernardo Gonçalves, Fábio Porto 0001
Artif. Intell.2
2016 Database System Support of Simulation Data
abstract
Supported by increasingly efficient HPC infra-structure, numerical simulations are rapidly expanding to fields such as oil and gas, medicine and meteorology. As simulations become more precise and cover longer periods of time, they may produce files with terabytes of data that need to be efficiently analyzed. In this paper, we investigate techniques for managing such data using an array DBMS. We take advantage of multidimensional arrays that nicely models the dimensions and variables used in numerical simulations. However, a naive approach to map simulation data files may lead to sparse arrays, impacting query response time, in particular, when the simulation uses irregular meshes to model its physical domain. We propose efficient techniques to map coordinate values in numerical simulations to evenly distributed cells in array chunks with the use of equi-depth histograms and space-filling curves. We implemented our techniques in SciDB and, through experiments over real-world data, compared them with two other approaches: row-store and column-store DBMS. The results indicate that multidimensional arrays and column-stores are much faster than a traditional row-store system for queries over a larger amount of simulation data. They also help identifying the scenarios where array DBMSs are most efficient, and those where they are outperformed by column-stores.
Hermano Lustosa, Fábio Porto 0001, Patrick Valduriez
Proc. VLDB Endow.2
2015 Modeling and Implementing Scientific Hypothesis
abstract
Computational Simulations are important tools that enable scientists to study complex phenomena about which few data is available or that require dangerous human interventions. They involve complex and heterogeneous components, including: mathematical equations, hypothesis, computational models and data. In order to support in-silico scientific research this complex environment needs to be modeled and have its data and metadata managed enabling model evolution, prediction analysis and decision-making. This paper proposes a scientific hypothesis conceptual model that allows scientists to represent the phenomenon been investigated, the hypotheses formulated in the attempt to explain it, and provides the ability to store results of experiment simulations with their corresponding provenance metadata. The proposed model supports scientific life-cycle through: provenance management, exchange of hypothesis as data, experiment reproducibility, model steering and simulation result analyses. A cardiovascular numerical simulation illustrates the applicability of the model and an initial implementation using SciDB is discussed.
Fábio Porto 0001, Ramon G. Costa, Ana Maria de Carvalho Moura, Bernardo Gonçalves
J. Database Manag.1
2014 Υ-DB: Managing scientific hypotheses as uncertain data
abstract
In view of the paradigm shift that makes science ever more data-driven, we consider deterministic scientific hypotheses as uncertain data. This vision comprises a probabilistic database (p-DB) design methodology for the systematic construction and management of U-relational hypothesis DBs, viz., γ-DBs. It introduces hypothesis management as a promising new class of applications for p-DBs. We illustrate the potential of γ-DB as a tool for deep predictive analytics.
Bernardo Gonçalves, Fábio Porto 0001
Proc. VLDB Endow.2
2013 Algebraic dataflows for big data analysis
abstract
Analyzing big data requires the support of dataflows with many activities to extract and explore relevant information from the data. Recent approaches such as Pig Latin propose a high-level language to model such dataflows. However, the dataflow execution is typically delegated to a MapRe-duce implementation such as Hadoop, which does not follow an algebraic approach, thus it cannot take advantage of the optimization opportunities of PigLatin algebra. In this paper, we propose an approach for big data analysis based on algebraic workflows, which yields optimization and parallel execution of activities and supports user steering using provenance queries. We illustrate how a big data processing dataflow can be modeled using the algebra. Through an experimental evaluation using real datasets and the execution of the dataflow with Chiron, an engine that supports our algebra, we show that our approach yields performance gains of up to 19.6% using algebraic optimizations in the dataflow and up to 39.1% of time saved on a user steering scenario.
Jonas Dias, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso
IEEE BigData4
2013 Research lattices: towards a scientific hypothesis data model
abstract
As the problems of scientific interest raise in scale and complexity, scientists have to tacitly manage too many analytic elements. Hypotheses are worked out to drive research towards successful explanation and prediction, which characterizes science as a dynamic activity that is partially ordered towards progress. This paper motivates and introduces research lattices, carrying out a lattice-theoretic approach for hypothesis representation and management in large-scale science and engineering. The goal of this work is to equip scientists with tools to manipulate and query hypotheses while keeping track of research progress. We refer to SciDB's array data model and discuss how data and theories could be managed in a unified model management framework.
Bernardo Gonçalves, Fábio Porto 0001
SSDBM2
2013 Chiron: a parallel engine for algebraic scientific workflows
abstract
SUMMARY Large‐scale scientific experiments based on computer simulations are typically modeled as scientific workflows, which eases the chaining of different programs. These scientific workflows are defined, executed, and monitored by scientific workflow management systems (SWfMS). As these experiments manage large amounts of data, it becomes critical to execute them in high‐performance computing environments, such as clusters, grids, and clouds. However, few SWfMS provide parallel support. The ones that do so are usually labor‐intensive for workflow developers and have limited primitives to optimize workflow execution. To address these issues, we developed workflow algebra to specify and enable the optimization of parallel execution of scientific workflows. In this paper, we show how the workflow algebra is efficiently implemented in Chiron, an algebraic based parallel scientific workflow engine. Chiron has a unique native distributed provenance mechanism that enables runtime queries in a relational database. We developed two studies to evaluate the performance of our algebraic approach implemented in Chiron; the first study compares Chiron with different approaches, whereas the second one evaluates the scalability of Chiron. By analyzing the results, we conclude that Chiron is efficient in executing scientific workflows, with the benefits of declarative specification and runtime provenance support. Copyright © 2013 John Wiley & Sons, Ltd.
Eduardo S. Ogasawara, Jonas Dias, Vítor Silva 0003, Fernando Seabra Chirigati, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso
Concurr. Comput. Pract. Exp.6
2013 Data management for eScience in Brazil
abstract
In the 5th edition, submitted papers were subjected to a peer review process with up to three reviews per submission. The edition was organized into three tracks with research papers presentations; a keynote talk and a poster session. The scientific application track included two papers. The paper by Bustos et al. 1 used time series analysis in the forecast of fluids in oils reservoirs. The paper by Cugler et al. 2 is one whose version has been extended as an invited paper for this special issue and explores a new field of managing animal sound recordings using database technology. The adoption of techniques from high processing computing in eScience is the theme of track 2. The first two papers discussed techniques involved in natural phenomena modeling and simulation. The first paper 3 discusses the adoption of high performance computing (HPC) environment in the modeling of protein sequence. Next, the paper by Sabino et al. 4 presents the challenges of using clusters of graphics processing units to compute 3-D wave propagation simulations, using finite difference methods. Lastly, the paper by Chirigati et al. 5, also selected for submitting an extended version to this special issue, discusses techniques for deploying scientific workflows in HPC. Finally, the track eScience services included paper in roughly two lines. In the first line, Costa and colleagues 7 explore the availability of deploying eScience in the cloud as services. In the second line, the two other works are related to data modeling including metrics of quality in bio ontologies 8 and a database integration strategy for ecological datasources 9. The 5th BreSCI workshop also received Prof. Luiz Nicolaci da Costa, astronomer of the National Observatory in Brazil. Prof. da Costa is the head of the LIneA laboratory, responsible for the storage, analysis, and publishing of astronomy catalogs, produced by large surveys. In his talk, Prof. Nicolaci presented the Portal, a software infrastructure from which scientists can query and process scientific pipelines over catalogs. The workshop ended with a poster session. After the workshop, the steering committee of the BreSci, indicated the best papers, whose authors were invited to extend their contributions to a submission to this special issue. Four papers were invited and submitted extended versions, of which two were finally accepted for publication and take part of this special issue. The first paper, by Cluger et al. 2, ‘An architecture for semantic retrieval of animal sound recordings’ presents innovative research in biodiversity data management. The authors investigate techniques to support management of recordings of animal sounds captured in natura and mixed with contextual information, such as environmental conditions and social events. The approach proposes the retrieval of animal sound recordings on the basis of the analysis of contextual information associated to the recordings as metadata and the use of controlled vocabularies with ontological inference support. The authors present a first prototype that implements part of the ideas proposed in the paper. The second accepted paper is entitled ‘Chiron: A Parallel Engine for Algebraic Scientific Workflows’ by Ogasawara et al. 6 describes a parallel workflow engine designed to efficiently process data centric scientific workflows. The system assigns known algebraic operators to workflow activities, leveraging the latter processing semantics. On the basis of the analysis of algebraic operators semantics, the system can compute an optimized workflow. Moreover, during execution, activities and data compose a processing unit, called activation, which can be freely scheduled on a HPC cluster. Finally, Chiron implements two activation dispatching modes: static and dynamic. Chiron has been tested with complex scientific workflows from the Oil and Gaz industry and has shown very promising results, placing itself as a good candidate for composing the eScience workflow ecosystem. In this context, systems to be developed in support for huge volume of data, as the astronomy surveys, shall consider data partitioning as data storage strategies. This is in some aspect new to scientific computing infrastructure running tightly coupled methods, which has resorted to HPC platforms where nodes have very few permanent storage areas and storage devices are accessed via high throughput network connections. In a scientific workflow environment, for instance, in which workflow activities communicate through files, the huge data transfer may jeopardize gains obtained from activities parallelism. As astronomy data are stored in databases and, eventually, processed by scientific workflows, another issue appears in the lack of integration between scientific workflows management systems and distributed relational database systems. As a matter of fact the, two systems run in complete independence precluding any integrated optimization, such as placing activities in the same nodes as the data partition they shall process. Some parallel database systems have been extended to integrate parallelism à la MapReduce to process over partitioned data. Whereas such solutions are a nice response to queries over huge databases, they are incapable of running general program pipelines, as in scientific workflows. Given such panorama, the challenges for eScience in Brazil and worldwide are huge. The papers in this special issue deal with data management and scientific workflow processing and collaborate therein to bridge the gap between science and eScience. We expect the Brazilian eScience community to continue facing these challenges in future editions of the Brazilian science workshop. We would like to thank the authors for contributing papers on their research on Data Management for eScience for this special issue, and all the reviewers for providing constructive reviews and in helping to shape this special issue. Finally, we would like to thank Prof. Geoffrey Fox for providing us an opportunity to bring this special issue to the research community.
Fábio Porto 0001, Bruno Schulze
Concurr. Comput. Pract. Exp.1
2012 Dynamic Workload-Based Partitioning for Large-Scale Databases
Miguel Liroz-Gistau, Reza Akbarinia, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez
DEXA (2)4
2012 A metaphoric trajectory data warehouse for Olympic athlete follow-up
abstract
SUMMARY Sport science is a research discipline that aims to understand exercise and apply scientific methods in support of increasing an athlete's performance. In this paper, we present initial results on modeling, managing and analyzing an athlete's data gathered by sport scientists. An Olympic data warehouse is designed initially to support the monitoring of an athlete's biochemical data. A trajectory data model is extended to represent the athlete's measurements along his/her training states, referred to here as metaphoric trajectories. Furthermore, a data warehouse for metaphoric trajectories is designed and two analysis approaches — a relational and a multidimensional one — are evaluated. We compare both approaches and discuss their benefits to the athlete's follow‐up analyses applied by sport scientists. Copyright © 2011 John Wiley & Sons, Ltd.
Fábio Porto 0001, Ana Maria de Carvalho Moura, Frederico C. da Silva, Adriana Bassini, Daniele C. Palazzi, Maira Poltosi, Luis Eduardo Viveiros de Castro, L. C. Cameron
Concurr. Comput. Pract. Exp.1
2012 Middleware for Clouds and e-Science
abstract
This special issue focuses on middleware for Clouds and e-Science applications, and is based on extended, thoroughly revised papers from the 8th International Workshop on Middleware for Grid, Clouds and e-Science (MGC 2010) and the Workshop on Challenges in e-Science (CIS 2010).The authors were invited to provide extended versions of their original papers taking into account comments and suggestions raised during the peer review process and comments from the audience during the workshops.
Bruno Schulze, Rajkumar Buyya, Fábio Porto 0001
Concurr. Comput. Pract. Exp.3
2011 An Algebraic Approach for Data-Centric Scientific Workflows
Eduardo S. Ogasawara, Daniel de Oliveira 0001, Patrick Valduriez, Jonas Dias, Fábio Porto 0001, Marta Mattoso
Proc. VLDB Endow.5
2010 Query processing in a three-level ontology-based data integration system
abstract
In this paper, we present a three-level ontology-based framework for effectively designing GAV data integration systems. In our approach, the mediated schema is represented by a domain ontology, which provides a conceptual representation of the application. Each local source is described by an application ontology, whose vocabulary is restricted to be a subset of the vocabulary of domain ontology. The three-level architecture permits dividing the mapping definition in two stages: local mappings and mediated mappings. Due to this architecture the problem of query answering can also be broken into two steps. First, the query is decomposed, using the mediated mappings, into a set of elementary sub-queries expressed in terms of the application ontologies. Then, these sub-queries are rewritten, using the local mappings, in terms of their local sources schemas. This paper focus on a method for query processing that addresses the problem of efficient query answering. Our approach is illustrated by an example of a virtual store mediating access to online booksellers.
João Carlos Pinheiro, Vânia M. P. Vidal, José A. F. de Macêdo, Eveline R. Sacramento, Marco A. Casanova, Fábio Porto 0001
iiWAS6
2008 A conceptual view on trajectories
Stefano Spaccapietra, Christine Parent, Maria Luisa Damiani, José A. F. de Macêdo, Fábio Porto 0001, Christelle Vangenot
Data Knowl. Eng.5
2007 A Conceptual Data Model Language for the Molecular Biology Domain
abstract
In-silico experiments require molecular biological knowledge to be mapped into a data model. In a previous work, we have elicited a list of requirements to be attended by a conceptual language for the molecular biology domain. In this paper, we present the evolution of this research by introducing a such language that encompasses most of the identified biological domain requirements. The language offers logical constraints expressions that augments the semantics of data and relationships types, as well as monotonic inheritance. Those aspects address the uncertainty and variability of biological knowledge shortening the gap between scientists mental model and traditional conceptual models.
José A. F. de Macêdo, Fábio Porto 0001, Sérgio Lifschitz, Philippe Picouet
CBMS2
2007 QEF - Supporting Complex Query Applications
abstract
This paper describes QEF a query evaluation framework designed to support complex applications on the grid. QEF has been extended to support querying within a number of different applications, including supporting scientific visualization and implementing a web service semantic search engine. Application requests take a form of a workflow in which tasks are represented as algebraic operators and specific data types are enveloped into a common tuple structure. The implemented system is automatically deployed into schedule grid nodes and autonomously manages query evaluation according to grid environment conditions. The generality of our approach has been tested with a number of applications leading to a full grid web service implementation available at http://codims. epfl. ch.
Fábio Porto 0001, Othman Tajmouati, Vinícius F. V. da Silva, Bruno Schulze, Fausto V. M. Ayres
CCGRID1
2007 An Extensible and Personalized Approach to QoS-enabled Service Discovery
abstract
We present an extensible and customizable framework for the autonomous discovery of Semantic Web services based on their QoS properties. Using semantic technologies, users can specify the QoS matching model and customize the ranking of services flexibly according to their preferences. The formal modeling of the discovery process as a query execution plan facilitates the introduction of different discovery algorithms and the automatic generation of parallelized matchmaking evaluations. This enables adapting our approach to unpredictable arrival rates of user queries and scales up to high numbers of published service descriptions.
Le-Hung Vu, Fábio Porto 0001, Karl Aberer, Manfred Hauswirth
IDEAS2
2006 Symptoms Ontology for Mapping Diagnostic Knowledge Systems
abstract
This paper describes a methodology for increasing the scope and precision of diagnostic Knowledge Based (KB) Systems. It has been stated that medical KB systems are either highly specialised, lack accuracy or are just too simple. To resolve this problem of scope we propose the use of a phased approach to diagnosis. The first phase being the querying of a symptoms ontology, to direct diagnostic systems to the most appropriate domain or class reference given input symptoms. Additional symptoms can then be targeted, extracted and analysed with a domain specific set of KB systems. This process allows us to forecast key symptoms, patient characteristics and increase the value of available data in decision making. In addition this approach could allow a system to dynamically correct an inappropriate domain decision. Such an approach also has the potential to be used to build a bridge between existing specialised medical KB systems.
Robert Minchin, Fábio Porto 0001, Christelle Vangenot, Sven Hartmann
CBMS2
2006 An adaptive parallel query processing middleware for the Grid
abstract
Abstract Grid services provide an important abstract layer on top of heterogeneous components (hardware and software) that take part in a Grid environment. In this scenario, applications such as scientific visualization require access to data of non‐conventional data types, such as fluid path geometry, and the evaluation of special user programs and algebraic operators, such as spatial hash‐join, on these data. In order to support such applications we are developing a Configurable Data Integration Middleware System for the Grid (CoDIMS‐G). CoDIMS‐G provides a query execution environment adapted to the heterogeneity and variations found in a Grid environment by offering a node scheduling algorithm and an adaptive query execution strategy. The latter both adapts to performance variations in a scheduled node and deals efficiently with repetitive evaluation of a query execution plan fragment, as needed for computing a particle's, trajectory. Copyright © 2005 John Wiley & Sons, Ltd.
Vinícius F. V. da Silva, Márcio L. Dutra, Fábio Porto 0001, Bruno Schulze, Álvaro Cesar P. Barbosa, Jauvane Cavalcante de Oliveira
Concurr. Comput. Pract. Exp.3
2004 A workflow-based architecture for E-learning in the Grid
abstract
Effective E-learning distributed environments should promote high cooperation. Workflow techniques can certainly contribute to such effectiveness as the creation and delivery of learning contents are typically accomplished by groups of individuals executing specific and predefined sequences of activities. Reusable learning objects (RLO or LO) also play an important role in this context as learning content is broken into smaller parts to facilitate deployment and execution assignment. Additionally, some LO may require high levels of computation as in the case of simulations in hemodynamics in a fluid mechanics course. In order to cope with such an application, we envisioned a workflow-based learning management system (LMS) using Web services as the communication infrastructure allied to the computational power of the Grid.
Luiz Antônio M. Pereira, Rubens Nascimento Melo, Fábio Porto 0001, Bruno Schulze
CCGRID3
2004 ROSA: A Repository of Objects with Semantic Access for e-Learning
Fábio Porto 0001, Ana Maria de Carvalho Moura, Fábio José Coutinho da Silva
IDEAS1
2004 CoDIMS: an adaptable middleware system for scientific visualization in Grids
abstract
Abstract In this paper we propose a middleware infrastructure adapted for supporting scientific visualization applications over a Grid environment. We instantiate a middleware system from CoDIMS, which is an environment for the generation of configurable data integration middleware systems. CoDIMS adaptive architecture is based on the integration of special components managed by a control module that executes users workflows. We exemplify our proposal with a middleware system generated for computing particles' trajectories within the constraints imposed by a Grid environment. Copyright © 2004 John Wiley & Sons, Ltd.
Fábio Porto 0001, Gilson A. Giraldi, Jauvane Cavalcante de Oliveira, Rodrigo L. S. Silva, Bruno Schulze
Concurr. Pract. Exp.1
2001 Processing Queries with Expensive Functions and Large Objects in Distributed Mediator Systems
abstract
LeSelect is a mediator system which allows scientists to publish their resources (data and programs) so they can be transparently accessed. The scientists can typically issue queries which access distributed published data and involve the execution of expensive functions (corresponding to programs). Furthermore, the queries can involve large objects, such as images (e.g. archived meteorological satellite data). In this context, the costs of transmitting large objects and invoking expensive functions are the dominant factors of execution time. In this paper, we first propose three query execution techniques which minimize these costs by taking full advantage of the distributed architecture of mediator systems like LeSelect. Then we devise parallel processing strategies for queries including expensive functions. Based on experimentation, we show that it is hard to predict the optimal execution order when dealing with several functions. We propose a new hybrid parallel technique to solve this problem and give some experimental results.
Luc Bouganim, Françoise Fabret, Fábio Porto 0001, Patrick Valduriez
ICDE3