EDBT 2026 Demo / reviewers in the wild / expert
Daniel de Oliveira 0001
dblp:68/7553-1 · also Daniel C. M. de Oliveira, Daniel Cardoso Moraes de Oliveira, Daniel Oliveira 0001
· DBLP profile ↗
59ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0001-9346-7651ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14 · 6 since 2021Artificial intelligence and machine learning · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Language Model Fusion to Improve Generalization in Portuguese Hate Speech DetectionabstractAs the number of fine-tuned language models for specialized domains and tasks continues to grow, managing a diverse set of solutions presents increasing challenges regarding scalability and adaptability. In this context, model-fusion techniques offer a promising approach to enhancing generalization by leveraging knowledge from multiple trained models. This paper investigates the fusion of pre-trained BERT-based models using ten different fusion strategies: (i) Simple Merging, (ii) Select Simple Merging, (iii) Fisher Merging, (iv) Select Fisher Merging, (v) RegMean, (vi) Task Arithmetic, (vii) DARE-simple Merging, (viii) TIES-MERGING, (ix) Robust Fine-tuning, and (x) Select Epoch Merging. To assess the effectiveness of these techniques, we conducted experiments in the challenging task of detecting hate speech in Portuguese. Such a task benefits from combining knowledge with model fusion, given that individual models might overlook the context-dependent nature of offensive language and nuanced forms of hate speech. The results using six datasets indicate that TIES-MERGING, in particular, can outperform individual models by successfully integrating specialized knowledge into a single solution, a more efficient and robust model. Annie Amorim, Gabriel Assis, Daniel de Oliveira 0001, Aline Paes |
FUSION | 3 |
| 2025 | Scalable Reputation Management: A Multi-Task Prompting Approach Using Fine-Tuned PLMs for Sentiment and Topic ClassificationabstractSocial media platforms provide a direct, real-time channel for companies to engage with their audiences, making the management of corporate reputation on social media a pivotal factor for organizational success. Reputation is shaped by factors such as the relationship with the audience, crisis communication, public sentiment, and the topics most frequently associated with the company. Although reputation is often viewed as an intangible asset, it can be quantified and monitored through various metrics. Advances in AI and Pretrained Language Models (PLMs) have automated tasks such as sentiment analysis, sentiment strength assessment, and topic classification, creating new opportunities to manage reputation more efficiently. However, current PLM solutions typically address these tasks individually and require specialized training for each company, limiting scalability and flexibility. To address these challenges, we propose a novel system framework to assist public relations firms by leveraging PLMs for automatic classification. We evaluated this approach through experiments on four companies, comparing zero-shot and fine-tuned models in task-specific and multi-task configurations. Additionally, we explored the transfer of knowledge between client companies using fine-tuned models. Results indicate that multi-task, multi-company fine-tuned PLM models offer simpler system management with competitive performance compared to highly specialized models. Paulo R. S. da Costa Jr., Matheus Yasuo Ribeiro Utino, João Silva-Leite, Lyncoln S. de Oliveira, Rodrigo Salvador Monteiro, Rodrigo A. C. Dias, Daniel de Oliveira 0001, Paulo Mann, Marcos V. N. Bedo |
ICWSM | 7 |
| 2025 | Optimizing Resource Estimation for Scientific Workflows in HPC Environments: A Layered-Bucket Heuristic ApproachabstractABSTRACT As computational simulations become complex and the amount of processed data grows, executing scientific workflows in High‐Performance Computing (HPC) environments is increasingly essential. However, accurately estimating the required computational resources for such executions presents a significant challenge, requiring a thorough examination of the workflow structure and the characteristics of the computational environment. This manuscript introduces the GraspCC‐LB heuristic, based on the Greedy Randomized Adaptive Search Procedure (GRASP), for estimating the necessary resources for executing scientific workflows in HPC environments. Unlike existing methods, GraspCC‐LB incorporates the layered structure of workflows into its estimation process. The proposed approach was evaluated using real traces of workflows from the fields of bioinformatics and astronomy. The resource estimations produced by GraspCC‐LB were compared against the actual resource usage in a real‐world HPC environment to evaluate its effectiveness. The results demonstrate the effectiveness of GraspCC‐LB as a robust approach for resource optimization in the context of large‐scale scientific workflows that require HPC capabilities. Luis C. R. Alvarenga, Yuri Frota, Daniel de Oliveira 0001, Rafaelli de C. Coutinho |
Concurr. Comput. Pract. Exp. | 3 |
| 2025 | PATSA-BIL: Pipeline for automated texture and structure analysis of borehole image logs
André M. Souza, Matheus A. Cruz, Paola M. C. Braga, Rodrigo B. Piva, Rodrigo A. C. Dias, Paulo R. Siqueira, Willian A. Trevizan, Candida M. de Jesus, Camilla Bazzarella, Rodrigo Salvador Monteiro, Flavia Bernardini, Leandro A. F. Fernandes, Elaine P. M. Sousa, Daniel de Oliveira 0001, Marcos V. N. Bedo |
Expert Syst. Appl. | 14 |
| 2025 | Curio: A Dataflow-Based Framework for Collaborative Urban Visual AnalyticsabstractOver the past decade, several urban visual analytics systems and tools have been proposed to tackle a host of challenges faced by cities, in areas as diverse as transportation, weather, and real estate. Many of these tools have been designed through collaborations with urban experts, aiming to distill intricate urban analysis workflows into interactive visualizations and interfaces. However, the design, implementation, and practical use of these tools still rely on siloed approaches, resulting in bespoke systems that are difficult to reproduce and extend. At the design level, these tools undervalue rich data workflows from urban experts, typically treating them only as data providers and evaluators. At the implementation level, they lack interoperability with other technical frameworks. At the practical use level, they tend to be narrowly focused on specific fields, inadvertently creating barriers to cross-domain collaboration. To address these gaps, we present Curio, a framework for collaborative urban visual analytics. Curio uses a dataflow model with multiple abstraction levels (code, grammar, GUI elements) to facilitate collaboration across the design and implementation of visual analytics components. The framework allows experts to intertwine data preprocessing, management, and visualization stages while tracking the provenance of code and visualizations. In collaboration with urban experts, we evaluate Curio through a diverse set of usage scenarios targeting urban accessibility, urban microclimate, and sunlight access. These scenarios use different types of data and domain methodologies to illustrate Curio's flexibility in tackling pressing societal challenges. Curio is available at urbantk.org/curio. Gustavo Moreira, Carolina Veiga Ferreira de Souza, Lucas Alexandre, Nicola Colaninno, Daniel de Oliveira 0001, Nivan Ferreira, Marcos Lage, Fabio Miranda 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Aggregating embeddings from image and radiology reports for multimodal Chest-CT retrievalabstractThis paper proposes a multimodal retrieval system for Chest CT (called ChestFinder) that combines image and report embeddings’ into a filter-and-query strategy. ChestFinder is composed of three modules, namely (i) text transformation, (ii) feature extraction, and (iii) ranking aggregation. Text transformation is conducted by a fine-tuned Generative Pre-training Transformer model (GPT), and the image embeddings are extracted after (i) training a Residual Neural Network (ResNet-50), and (ii) reducing and scaling the encoded vectors. ChestFinder produces one list of similar images and another of related reports for each query input composed of a Chest CT with a radiology report. Then, the ChestFinder ranking aggregation module fuses those two lists to produce the final ordered set of retrieved objects, as in a top-k query. The aggregation is performed by a fine-tuned Threshold Algorithm (TA) whose weights are calculated by a Multi-Layer Perceptron (MLP) trained to label the reports. To examine the quality enhancement brought by this multimodal search, we constructed a dataset of Chest CTs from our University Hospital PACS/RIS systems by filtering distinct cases diagnosed with emphysema (one finding per case and with at least two radiologists agreeing on the diagnosis). A holdout experimental evaluation showed the ChestFinder search achieved higher Accuracy and Sensitivity than content-only top-k searches. Results also indicated quality gains drawn from the adjustments of ChestFinder modular components: (i) fine-tuned GPT achieved up to 0.89 F1-Score in data testing with a stable train/validation ratio for radiology reports, (ii) fine-tuned GPT significantly outperformed the zero-shot approach as well as a fine-tuned BERT, (iii) non-weighted ranking aggregation increased the search accuracy in up to 10%, and (iv) fine-tuned TA outperformed the baseline and non-weighted ranking aggregation in up to 52%. João Silva-Leite, Cristina A. P. Fontes, Alair S. Santos, Diogo G. Correa, Marcel Koenigkam-Santos, Paulo Mazzoncini de Azevedo Marques, Daniel de Oliveira 0001, Aline Paes, Marcos V. N. Bedo |
CBMS | 7 |
| 2024 | Enriching Hierarchical Navigable Small World Searches with Result Diversification
Mauro Weber, João Silva-Leite, Lúcio F. D. Santos, Daniel de Oliveira 0001, Marcos V. N. Bedo |
DEXA (1) | 4 |
| 2024 | MAESTRO: a lightweight ontology-based framework for composing and analyzing script-based scientific experiments
Luiz Gustavo Dias, Bruno Lopes 0001, Daniel de Oliveira 0001 |
Knowl. Inf. Syst. | 3 |
| 2024 | PW: A Visual Approach for Building, Managing, and Analyzing Weather Simulation Ensembles at RuntimeabstractWeather forecasting is essential for decision-making and is usually performed using numerical modeling. Numerical weather models, in turn, are complex tools that require specialized training and laborious setup and are challenging even for weather experts. Moreover, weather simulations are data-intensive computations and may take hours to days to complete. When the simulation is finished, the experts face challenges analyzing its outputs, a large mass of spatiotemporal and multivariate data. From the simulation setup to the analysis of results, working with weather simulations involves several manual and error-prone steps. The complexity of the problem increases exponentially when the experts must deal with ensembles of simulations, a frequent task in their daily duties. To tackle these challenges, we propose ProWis: an interactive and provenance-oriented system to help weather experts build, manage, and analyze simulation ensembles at runtime. Our system follows a human-in-the-loop approach to enable the exploration of multiple atmospheric variables and weather scenarios. ProWis was built in close collaboration with weather experts, and we demonstrate its effectiveness by presenting two case studies of rainfall events in Brazil. Carolina Veiga Ferreira de Souza, Suzanna Maria Bonnet, Daniel de Oliveira 0001, Márcio Cataldi, Fabio Miranda 0001, Marcos Lage |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | NewsCollab: Fostering Data-driven Journalism with CrowdsourcingabstractThe amount of data produced over the last decade has increased at a fast pace. Many domains of knowledge were completely transformed by the availability of massive datasets. Journalism is one of these domains. Commonly, news articles are produced by analyzing massive datasets containing raw data and previously published articles. This new type of journalism is called data-driven journalism. However, data-driven journalism requires that previously published articles are accessible, categorized, and organized. This can be a laborious task to be manually performed. Starting from the assumption that news articles should be able to be imported from multiple sources, and collaboratively analyzed and categorized by other journalists (to avoid bias), we propose the use of a new collaborative system named NewsCollab. NewsCollab allows a multitude of journalists to download news articles from different portals and provide feedback on these articles, giving to the journalistic community a better understanding of the published articles that can be used to produce new material. Ygor Rolim, Danielly Alves, Flávia Clemente, Daniel de Oliveira 0001 |
CSCWD | 4 |
| 2023 | Adding Result Diversification to kNN-Based Joins in a Map-Reduce Framework
Vinícius Souza, Luiz Olmes Carvalho, Daniel de Oliveira 0001, Marcos V. N. Bedo, Lúcio F. D. Santos |
DEXA (1) | 3 |
| 2023 | Optimizing computational costs of Spark for SARS-CoV-2 sequences comparisons on a commercial cloudabstractSummary Cloud computing is currently one of the prime choices in the computing infrastructure landscape. In addition to advantages such as the pay‐per‐use bill model and resource elasticity, there are technical benefits regarding heterogeneity and large‐scale configuration. Alongside the classical need for performance, for example, time, space, and energy, there is an interest in the financial cost that might come from budget constraints. Based on scalability considerations and the pricing model of traditional public clouds, a reasonable optimization strategy output could be the most suitable configuration of virtual machines to run a specific workload. From the perspective of runtime and monetary cost optimizations, we provide the adaptation of a Hadoop applications execution cost model extracted from the literature aiming at Spark applications modeled with the MapReduce paradigm. We evaluate our optimizer model executing an improved version of the Diff Sequences Spark application to perform SARS‐CoV‐2 coronavirus pairwise sequence comparisons using the AWS EC2's virtual machine instances. The experimental results with our model outperformed 80% of the random resource selection scenarios. By only employing spot worker nodes exposed to revocation scenarios rather than on‐demand workers, we obtained an average monetary cost reduction of 35.66% with a slight runtime increase of 3.36%. Alan L. Nunes, Alba Cristina Magalhaes Alves de Melo, Claude Tadonki, Cristina Boeres, Daniel de Oliveira 0001, Lúcia M. A. Drummond |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | AIS-based maritime anomaly traffic detection: A review
Cláudio Vasconcelos Ribeiro, Aline Paes, Daniel de Oliveira 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Pushing diversity into higher dimensions: The LID effect on diversified similarity searching
Daniel L. Jasbick, Lúcio F. D. Santos, Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina, Daniel de Oliveira 0001, Marcos V. N. Bedo |
Inf. Syst. | 5 |
| 2022 | A Provenance-based Execution Strategy for Variant GPU-accelerated Scientific Workflows in Clouds
Murilo B. Stockinger, Marcos A. Guerine, Ubiratam de Paula Junior, Filipe Santiago, Yuri Frota, Isabel Rosseti, Alexandre Plastino 0001, Daniel de Oliveira 0001 |
J. Grid Comput. | 8 |
| 2021 | An incremental reinforcement learning scheduling strategy for data-intensive scientific workflows in the cloudabstractSummary Most scientific experiments can be modeled as workflows. These workflows are usually computing‐ and data‐intensive, demanding the use of high‐performance computing environments such as clusters, grids, and clouds. This latter offers the advantage of the elasticity, which allows for changing the number of virtual machines (VMs) on demand. Workflows are typically managed using scientific workflow management systems (SWfMS). Many existing SWfMSs offer support for cloud‐based execution. Each SWfMS has its scheduler that follows a well‐defined cost function. However, such cost functions should consider the characteristics of a dynamic environment, such as live migrations or performance fluctuations, which are far from trivial to model. This article proposes a novel scheduling strategy, named ReASSIgN, based on reinforcement learning (RL). By relying on an RL technique, one may assume that there is an optimal (or suboptimal) solution for the scheduling problem, and aims at learning the best scheduling based on previous executions in the absence of a mathematical model of the environment. For this, an extension of a well‐known workflow simulator WorkflowSim is proposed to implement an RL strategy for scheduling workflows. Once the scheduling plan is generated via simulation, the workflow is executed in the cloud using SciCumulus SWfMS. We conducted a throughout evaluation of the proposed scheduling strategy using a real astronomy workflow named Montage. André Nascimento, Vítor Silva 0003, Aline Paes, Daniel de Oliveira 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Towards optimizing the execution of spark scientific workflows using machine learning-based parameter tuningabstractSummary In the last few years, Apache Spark has become a de facto the standard framework for big data systems on both industry and academy projects. Spark is used to execute compute‐ and data‐intensive workflows in distinct areas like biology and astronomy. Although Spark is an easy‐to‐install framework, it has more than one hundred parameters to be set, besides domain‐specific parameters of each workflow. In this way, to execute Spark‐based workflows efficiently, the user has to fine‐tune a myriad of Spark and workflow parameters (eg, partitioning strategy, the average size of a DNA sequence, etc.). This configuration task cannot be manually performed in a trial‐and‐error manner since it is tedious and error‐prone. This article proposes an approach that focuses on generating interpretable predictive machine learning models (ie, decision trees), and then extract useful rules (ie, patterns) from these models that can be applied to configure parameters of future executions of the workflow and Spark for nonexperts users. In the experiments presented in this article, the proposed parameter configuration approach led to better performance in processing Spark workflows. Finally, the approach introduced here reduced the number of parameters to be configured by identifying the most relevant domain‐specific ones related to the workflow performance in the predictive model. Douglas E. M. de Oliveira, Fábio Porto 0001, Cristina Boeres, Daniel de Oliveira 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Cache-aware scheduling of scientific workflows in a multisite cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
Future Gener. Comput. Syst. | 2 |
| 2020 | Distributed Caching of Scientific Workflows in Multisite Cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
DEXA (2) | 2 |
| 2020 | Some Branches May Bear Rotten Fruits: Diversity Browsing VP-Trees
Daniel L. Jasbick, Lúcio F. D. Santos, Daniel de Oliveira 0001, Marcos V. N. Bedo |
SISAP | 3 |
| 2020 | OLAP parallel query processing in clouds with C-ParGRESabstractSummary The advent of big data technologies has changed the way many companies manage their data. Several companies moved their data to the cloud using the concept of database‐as‐a‐service (DBaaS). Moving databases to the cloud presents several challenges related to flexible and scalable management of data. Although some of these companies migrated to NoSQL databases, most still rely on relational databases in the cloud to manage data, especially data that is critical to the decision making process. Online analytical processing (OLAP) queries take a long time to be processed, thus demanding high‐performance capabilities from their associated database systems to get results in a feasible time. In this article, we propose a middleware solution that can be deployed in any cloud provider, named C‐ParGRES, which explores database replication and interquery and intraquery parallelism to efficiently support OLAP queries in the cloud. C‐ParGRES is an extension of ParGRES, an open‐source database cluster middleware for high‐performance OLAP query processing in clusters. C‐ParGRES exploits cloud capabilities such as on‐demand resource provisioning and elasticity. In addition, C‐ParGRES can create multiple and independent virtual clusters for different database and users. We evaluate C‐ParGRES with two real‐world OLAP applications, both from the Brazilian Institute of Geography and Statistics. Results show that C‐ParGRES is a cost‐effective solution for OLAP query processing in the cloud. Marcello W. M. Ribeiro, Alexandre A. B. Lima, Daniel de Oliveira 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | Capturing and Analyzing Provenance from Spark-based Scientific Workflows with SAMbA-RaP
Thaylon Guedes, Lucas Bertelli Martins, Maria Luiza Furtuozo Falci, Vítor Silva 0003, Kary A. C. S. Ocaña, Marta Mattoso, Marcos V. N. Bedo, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 8 |
| 2020 | BioinfoPortal: A scientific gateway for integrating bioinformatics applications on the Brazilian national high-performance computing network
Kary A. C. S. Ocaña, Marcelo Galheigo, Carla Osthoff, Luiz M. R. Gadelha Jr., Fábio Porto 0001, Antônio Tadeu A. Gomes, Daniel de Oliveira 0001, Ana Tereza Ribeiro de Vasconcelos |
Future Gener. Comput. Syst. | 7 |
| 2020 | Adding domain data to code profiling tools to debug workflow parallel executionabstractComputer simulations may be composed of several scientific programs chained in a coherent flow running in High Performance Computing and cloud environments. These runs may present different execution behavior associated to the parallel flow of data among programs. Gather insight into the parallel flow of data is important for several applications. The usual way of getting insight into code performance is by means of a code-profiler. Several parallel code-profiling tools already support performance analysis, such as Tuning and Analysis Utilities (TAU), or provide fine-grained performance statistics, e.g., System Activity Report (SAR). These tools are effective for code profiling, but are not connected to the concept of IO-intensive workflows. Analyzing the workflow execution with domain and performance data is important for users because they can identify anomalies, choose suitable machines to run their workflows, etc. This type of analysis may be performed by capturing execution data enriched with fine-grained domain data during the long-term run of a computer simulation. In this paper, we propose a monitoring data capture approach as a component that couples code-profiling tools to domain data from workflow executions. The goal is to profile and debug parallel executions of workflows through queries to a database that integrates performance, resource consumption, provenance, and domain data from simulation programs flow at runtime. We show how querying this database with domain-aware data at runtime allows to identify performance anomalies not detected by code-profiling tools. We evaluate our approach using the astronomy Montage workflow on a cluster environment and the SciPhy bioinformatics workflow on the Amazon cloud. In both cases computing time overhead imposed by our approach for gathering fine-grained domain, performance, and resource consumption data is negligible. Vítor Silva 0003, Leonardo Neves, Renan Souza 0001, Alvaro L. G. A. Coutinho, Daniel de Oliveira 0001, Marta Mattoso |
Future Gener. Comput. Syst. | 5 |
| 2019 | A Two-Phase Learning Approach for the Segmentation of Dermatological WoundsabstractTissue segmentation in photographs of lower limb chronic ulcers is a non-intrusive approach that supports dermatological analyses. This paper presents 2PLA, a method that combines supervised and unsupervised learning strategies for enhancing the segmentation of dermatological wounds. Given an ulcer photo captured according to a fixed protocol, 2PLA first phase performs a pixelwise classification of points of interest, whereas pre-processing filters are employed for the smoothing of image noise. The cleaned image is further sent to the 2PLA divide-and-conquer second phase. It builds upon SLIC superpixel construction algorithm for dividing the lower limb into regions of interest with well-defined borders, and clusters the superpixels by taking advantage of the similarity-based DBSCAN algorithm. We set up the phases of our method by using a real annotated set of dermatological wounds, and empirical evaluations on representative samples up to 100,000 points showed a compact Multi-Layer Perceptron with Levenberg-Marquardt training algorithm (Cohen-Kappa = .971, Sensitivity = .98, and Specificity = .98) outperformed other classifiers as 2PLA first phase. Additionally, experimental trials on DBSCAN with five distance functions (L1, L2, L∞, Canberra, and BrayCurtis) indicated L1function provided fewer groups in comparison to the competitors, and the number of clusters was an exponential decay to the similarity ratio. Accordingly, we used the elbow criterion for finding the L1-based DBSCAN threshold as 2PLA second phase parameterization. We evaluated the fine-tuned setting of our method over a labeled set of ulcer images, and wounded tissues were segmented within a .05 Mean Absolute Error ratio. These results illustrate the impact of learning parameters on 2PLA as well as the method efficacy for wound segmentation. Wellington S. Silva, Daniel L. Jasbick, Rodrigo Erthal Wilson, Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina, Lúcio F. D. Santos, Ana Elisa Serafim Jorge, Daniel de Oliveira 0001, Marcos V. N. Bedo |
CBMS | 8 |
| 2019 | Towards a Science Gateway for Bioinformatics: Experiences in the Brazilian System of High Performance ComputingabstractScience gateways bring out the possibility of reproducible science as they are integrated into reusable techniques, data and workflow management systems, security mechanisms, and high performance computing (HPC). We introduce BioinfoPortal, a science gateway that integrates a suite of different bioinformatics applications using HPC and data management resources provided by the Brazilian National HPC System (SINAPAD). BioinfoPortal follows the Software as a Service (SaaS) model and the web server is freely available for academic use. The goal of this paper is to describe the science gateway and its usage, addressing challenges of designing a multiuser computational platform for parallel/distributed executions of large-scale bioinformatics applications using the Brazilian HPC resources. We also present a study of performance and scalability of some bioinformatics applications executed in the HPC environments and perform machine learning analyses for predicting features for the HPC allocation/usage that could better perform the bioinformatics applications via BioinfoPortal. Kary A. C. S. Ocaña, Marcelo Galheigo, Carla Osthoff, Luiz M. R. Gadelha Jr., Antônio Tadeu A. Gomes, Daniel de Oliveira 0001, Fábio Porto 0001, Ana Tereza Ribeiro de Vasconcelos |
CCGRID | 6 |
| 2019 | Adaptive Caching for Data-Intensive Scientific Workflows in the Cloud
Gaëtan Heidsieck, Daniel de Oliveira 0001, Esther Pacitti, Christophe Pradal, François Tardieu, Patrick Valduriez |
DEXA (2) | 2 |
| 2019 | A k-Skyband Approach for Feature Selection
Marcos V. N. Bedo, Paolo Ciaccia, Davide Martinenghi, Daniel de Oliveira 0001 |
SISAP | 4 |
| 2019 | A provenance-based heuristic for preserving results confidentiality in cloud-based scientific workflows
Marcos A. Guerine, Murilo B. Stockinger, Isabel Rosseti, Luidi Simonetti, Kary A. C. S. Ocaña, Alexandre Plastino 0001, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 7 |
| 2018 | Exploring Diversified Similarity with KundahaabstractExploring large medical image sets by means of traditional similarity query criteria (e.g., neighborhood) can be fruitless if retrieved images are too similar among themselves. This demonstration introduces Kundaha, an exploration tool that assists experts in retrieving and navigating on results from a diversified similarity perspective of user-posed queries. Its implementation includes a wide set of metrics, descriptors, and indexes for enhancing query execution. Users can combine such features with diversified similarity criteria for the organized exploration of result sets and also employ relevance feedback cycles for finding new query-based viewpoints. Lúcio F. D. Santos, Gustavo Blanco, Daniel de Oliveira 0001, Agma J. M. Traina, Caetano Traina Jr., Marcos V. N. Bedo |
CIKM | 3 |
| 2018 | Towards Safer (Smart) Cities: Discovering Urban Crime Patterns Using Logic-based Relational Machine LearningabstractSmart cities initiatives have the potential to improve the life of citizens in a huge number of dimensions. One of them is the development of techniques and services capable of contributing to the enhancement of security public policies. Finding criminal patterns from historical data would arguably help in predicting and even preventing thefts and burglaries that continuously increase in urban centers worldwide. However, accessing such history and finding patterns across the interrelated crime occurrences data are challenging tasks, particularly to underdevelopment countries. In this paper, we address these problems by combining three techniques: we collect crime data from existing crowd-sourcing systems, we automatically induce patterns with relational machine learning, and we manage the entire process using scientific workflows. The framework developed under these lines is named CRiMINaL (Crime patteRn MachINe Learning). Experimental results conducted from a popular Brazilian source of data and a traditional relational learning system shows that CRiMINaL is a promising tool to induce interpretable models that can assist police departments on crime prevention. Vítor N. Lourenço, Paulo Mann, Artur Guimaraes, Aline Paes, Daniel de Oliveira 0001 |
IJCNN | 5 |
| 2018 | DfAnalyzer: Runtime Dataflow Analysis of Scientific Applications using ProvenanceabstractWe present DfAnalyzer, a tool that enables monitoring, debugging, steering, and analysis of dataflows while being generated by scientific applications. It works by capturing strategic domain data, registering provenance and execution data to enable queries at runtime. DfAnalyzer provides lightweight dataflow monitoring components to be invoked by high performance applications. It can be plugged in scientific code scripts, or Spark applications, in the same way users already plug visualization library components. During this demo, we will show how DfAnalyzer captures the dataflow, provenance, as well as how it provides runtime data analyses of applications. We will also encourage attendees to use DfAnalyzer for their own applications. Vítor Silva 0003, Daniel de Oliveira 0001, Marta Mattoso, Patrick Valduriez |
Proc. VLDB Endow. | 2 |
| 2017 | Deriving scientific workflows from algebraic experiment lines: A practical approach
Anderson Marinho, Daniel de Oliveira 0001, Eduardo S. Ogasawara, Vítor Silva 0003, Kary A. C. S. Ocaña, Leonardo Murta 0001, Vanessa Braganholo, Marta Mattoso |
Future Gener. Comput. Syst. | 2 |
| 2017 | Raw data queries during data-intensive parallel workflow execution
Vítor Silva 0003, José Leite, José J. Camata, Daniel de Oliveira 0001, Alvaro L. G. A. Coutinho, Patrick Valduriez, Marta Mattoso |
Future Gener. Comput. Syst. | 4 |
| 2017 | A hybrid evolutionary algorithm for task scheduling and data assignment of data-intensive scientific workflows on clouds
Luan Teylo, Ubiratam de Paula Junior, Yuri Frota, Daniel de Oliveira 0001, Lúcia M. A. Drummond |
Future Gener. Comput. Syst. | 4 |
| 2017 | Managing Provenance of Implicit Data Flows in Scientific ExperimentsabstractScientific experiments modeled as scientific workflows may create, change, or access data products not explicitly referenced in the workflow specification, leading to implicit data flows. The lack of knowledge about implicit data flows makes the experiments hard to understand and reproduce. In this article, we present ProvMonitor, an approach that identifies the creation, change, or access to data products even within implicit data flows. ProvMonitor links this information with the workflow activity that generated it, allowing for scientists to compare data products within and throughout trials of the same workflow, identifying side effects on data evolution caused by implicit data flows. We evaluated ProvMonitor and observed that it could answer queries for scenarios that demand specific knowledge related to implicit provenance. Vitor C. Neves, Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Vanessa Braganholo, Leonardo Murta 0001 |
ACM Trans. Internet Techn. | 2 |
| 2016 | Analyzing related raw data files through dataflowsabstractSummary Computer simulations may ingest and generate high numbers of raw data files. Most of these files follow a de facto standard format established by the application domain, for example, Flexible Image Transport System for astronomy. Although these formats are supported by a variety of programming languages, libraries, and programs, analyzing thousands or millions of files requires developing specific programs. Database management systems (DBMS) are not suited for this, because they require loading the raw data and structuring it, which becomes heavy at large scale. Systems like NoDB, RAW, and FastBit have been proposed to index and query raw data files without the overhead of using a database management system. However, these solutions are focused on analyzing one single large file instead of several related files. In this case, when related files are produced and required for analysis, the relationship among elements within file contents must be managed manually, with specific programs to access raw data. Thus, this data management may be time‐consuming and error‐prone. When computer simulations are managed by a scientific workflow management system (SWfMS), they can take advantage of provenance data to relate and analyze raw data files produced during workflow execution. However, SWfMS registers provenance at a coarse grain, with limited analysis on elements from raw data files. When the SWfMS is dataflow‐aware, it can register provenance data and the relationships among elements of raw data files altogether in a database, which is useful to access the contents of a large number of files. In this paper, we propose a dataflow approach for analyzing element data from several related raw data files. Our approach is complementary to the existing single raw data file analysis approaches. We use the Montage workflow from astronomy and a workflow from Oil and Gas domain as data‐intensive case studies. Our experimental results for the Montage workflow explore different types of raw data flows like showing all linear transformations involved in projection simulation programs, considering specific mosaic elements from input repositories. The cost for raw data extraction is approximately 3.7% of the total application execution time. Copyright © 2015 John Wiley & Sons, Ltd. Vítor Silva 0003, Daniel de Oliveira 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Multi-objective scheduling of Scientific Workflows in multisite clouds
Ji Liu 0003, Esther Pacitti, Patrick Valduriez, Daniel de Oliveira 0001, Marta Mattoso |
Future Gener. Comput. Syst. | 4 |
| 2016 | A Dynamic Cloud Dimensioning Approach for Parallel Scientific Workflows: a Case Study in the Comparative Genomics Domain
Rafaelli de C. Coutinho, Yuri Frota, Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Lúcia M. A. Drummond |
J. Grid Comput. | 4 |
| 2015 | Data Analytics in Bioinformatics: Data Science in Practice for Genomics Analysis WorkflowsabstractWorkflow systems manage large-scale experiments and deliver a large volume of provenance data traces. The provenance repository of these systems contains information about the workflow execution, which allows for tracking and analyzing data transformations. However, provenance data may still be considered a black-box, when it comes to analyze the contents of resulting data files. Current solutions are focused on data transformation at coarse grain, they point to input and output files, but do not allow for exploring domain-specific data. Data analytics is essential for managing large-scale workflows executed in parallel, especially when tracking anomalous executions. In this paper, we present a data analytics approach, which is based on the use of provenance data enriched with domain-specific data coupled to a data mining tool. A real bioinformatics workflow was modeled and executed in parallel on top of Amazon clouds. It manipulates complex biological data, which is difficult to monitor like many other genomic workflows. We evaluate the benefits of using domain-specific data and provenance data for user steering while monitoring the execution with detailed filters, steering on specific conditions and performance evaluation. Results show that the provenance database coupled to workflow systems has an unexplored potential for raw data analytics, which may improve the user confidence and reduce overall execution time. Kary A. C. S. Ocaña, Vítor Silva 0003, Daniel de Oliveira 0001, Marta Mattoso |
e-Science | 3 |
| 2015 | Optimizing virtual machine allocation for parallel scientific workflows in federated clouds
Rafaelli de C. Coutinho, Lúcia M. A. Drummond, Yuri Frota, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 4 |
| 2015 | Dynamic steering of HPC scientific workflows: A survey
Marta Mattoso, Jonas Dias, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Flavio Costa, Felipe Horta, Vítor Silva 0003, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 8 |
| 2014 | Evaluating Grasp-based cloud dimensioning for comparative genomics: A practical approachabstractCloud computing establishes a new computing model where a wide range of computing resources are provided to several types of users. Especially for bioinformatics experiments modeled as scientific workflows, clouds provide several types of resources as virtual machines (VM), storage, databases and computing power that can be combined for empowering the scientific workflow execution. These workflows usually require high performance environments and parallelism techniques since their activities are data and computing intensive and can execute for a long time. There are then some Scientific Workflow Management Systems (SWfMS) that already manage the parallel execution of scientific workflows in clouds. Most of them instantiate a virtual cluster for the execution. However, they rely on the user to estimate the amount of VMs to be instantiated to create this virtual cluster. Estimating the amount of VMs to instantiate is then a crucial task to avoid negative impacts on the workflow performance with under or over estimations. This dimensioning also is not a trivial task in clouds due to the large number of VM types to choose in a cloud provider. Previously proposed approach named GraspCC already provides a near optimal estimation of the amount of VM for general applications, not scientific workflows. In this paper, we coupled the GraspCC to SciCumulus (Cloud-based Parallel Engine for Scientific Workflows) engine to estimate the necessary amount of VMs for bioinformatics workflows. We have evaluated GraspCC by comparing the estimative with real executions of a set of large-scale comparative genomics workflows. It showed the suitability of GraspCC to estimate the amount of VMs in real bioinformatics cloud workflows. Rafaelli de C. Coutinho, Lúcia M. A. Drummond, Yuri Frota, Daniel de Oliveira 0001, Kary A. C. S. Ocaña |
CLUSTER | 4 |
| 2013 | Algebraic dataflows for big data analysisabstractAnalyzing big data requires the support of dataflows with many activities to extract and explore relevant information from the data. Recent approaches such as Pig Latin propose a high-level language to model such dataflows. However, the dataflow execution is typically delegated to a MapRe-duce implementation such as Hadoop, which does not follow an algebraic approach, thus it cannot take advantage of the optimization opportunities of PigLatin algebra. In this paper, we propose an approach for big data analysis based on algebraic workflows, which yields optimization and parallel execution of activities and supports user steering using provenance queries. We illustrate how a big data processing dataflow can be modeled using the algebra. Through an experimental evaluation using real datasets and the execution of the dataflow with Chiron, an engine that supports our algebra, we show that our approach yields performance gains of up to 19.6% using algebraic optimizations in the dataflow and up to 39.1% of time saved on a user steering scenario. Jonas Dias, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
IEEE BigData | 3 |
| 2013 | An Artificial Emotional Agent-Based Architecture for Games Simulation
Rainier Sales, Esteban Walter Gonzalez Clua, Daniel de Oliveira 0001, Aline Paes |
ICEC | 3 |
| 2013 | Chiron: a parallel engine for algebraic scientific workflowsabstractSUMMARY Large‐scale scientific experiments based on computer simulations are typically modeled as scientific workflows, which eases the chaining of different programs. These scientific workflows are defined, executed, and monitored by scientific workflow management systems (SWfMS). As these experiments manage large amounts of data, it becomes critical to execute them in high‐performance computing environments, such as clusters, grids, and clouds. However, few SWfMS provide parallel support. The ones that do so are usually labor‐intensive for workflow developers and have limited primitives to optimize workflow execution. To address these issues, we developed workflow algebra to specify and enable the optimization of parallel execution of scientific workflows. In this paper, we show how the workflow algebra is efficiently implemented in Chiron, an algebraic based parallel scientific workflow engine. Chiron has a unique native distributed provenance mechanism that enables runtime queries in a relational database. We developed two studies to evaluate the performance of our algebraic approach implemented in Chiron; the first study compares Chiron with different approaches, whereas the second one evaluates the scalability of Chiron. By analyzing the results, we conclude that Chiron is efficient in executing scientific workflows, with the benefits of declarative specification and runtime provenance support. Copyright © 2013 John Wiley & Sons, Ltd. Eduardo S. Ogasawara, Jonas Dias, Vítor Silva 0003, Fernando Seabra Chirigati, Daniel de Oliveira 0001, Fábio Porto 0001, Patrick Valduriez, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 5 |
| 2013 | Designing a parallel cloud based comparative genomics workflow to improve phylogenetic analyses
Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
Future Gener. Comput. Syst. | 2 |
| 2013 | Performance evaluation of parallel strategies in public clouds: A study with phylogenomic workflows
Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Jonas Dias, João Carlos de A. R. Gonçalves, Fernanda Baião, Marta Mattoso |
Future Gener. Comput. Syst. | 1 |
| 2012 | Discovering drug targets for neglected diseases using a pharmacophylogenomic cloud workflowabstractIllnesses caused by parasitic protozoan are a research priority. A representative group of these illnesses is the commonly known as Neglected Tropical Diseases (NTD). NTD specially attack low socioeconomic population around the world and new anti-protozoan inhibitors are needed and several drug discovery projects focus on researching new drug targets. Pharmacophylogenomics is a novel bioinformatics field that aims at reducing the time and the financial cost of the drug discovery process. Pharmacophylogenomic analyses are applied mainly in the early stages of the research phase in drug discovery. Pharmacophylogenomic analysis executes several bioinformatics programs in a coherent flow to identify homologues sequences, construct phylogenetic trees and execute evolutionary and structural experiments. This way, it can be modeled as scientific workflows. Pharmacophylogenomic analysis workflows are complex, computing and data intensive and may execute during weeks. This way, it benefits from parallel execution. We propose SciPPGx, a scientific workflow that aims at providing thorough inferring support for pharmacophylogenomic hypotheses. SciPPGx is executed in parallel in a cloud using SciCumulus workflow engine. Experiments show that SciPPGx considerably reduces the total execution time up to 97.1% when compared to a sequential execution. We also present representative biological results taking advantage of the inference covering several related bioinformatics overviews. Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
eScience | 2 |
| 2012 | An adaptive parallel execution strategy for cloud-based scientific workflowsabstractSUMMARY Many of the existing large‐scale scientific experiments modeled as scientific workflows are compute‐intensive. Some scientific workflow management systems already explore parallel techniques, such as parameter sweep and data fragmentation, to improve performance. In those systems, computing resources are used to accomplish many computational tasks in high performance environments, such as multiprocessor machines or clusters. Meanwhile, cloud computing provides scalable and elastic resources that can be instantiated on demand during the course of a scientific experiment, without requiring its users to acquire expensive infrastructure or to configure many pieces of software. In fact, because of these advantages some scientists have already adopted the cloud model in their scientific experiments. However, this model also raises many challenges. When scientists are executing scientific workflows that require parallelism, it is hard to decide a priori the amount of resources to use and how long they will be needed because the allocation of these resources is elastic and based on demand. In addition, scientists have to manage new aspects such as initialization of virtual machines and impact of data staging. SciCumulus is a middleware that manages the parallel execution of scientific workflows in cloud environments. In this paper, we introduce an adaptive approach for executing parallel scientific workflows in the cloud. This approach adapts itself according to the availability of resources during workflow execution. It checks the available computational power and dynamically tunes the workflow activity size to achieve better performance. Experimental evaluation showed the benefits of parallelizing scientific workflows using the adaptive approach of SciCumulus, which presented an increase of performance up to 47.1%. Copyright © 2011 John Wiley & Sons, Ltd. Daniel de Oliveira 0001, Eduardo S. Ogasawara, Kary A. C. S. Ocaña, Fernanda Baião, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | A Provenance-based Adaptive Scheduling Heuristic for Parallel Scientific Workflows in Clouds
Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Fernanda Baião, Marta Mattoso |
J. Grid Comput. | 1 |
| 2011 | A Performance Evaluation of X-Ray Crystallography Scientific Workflow Using SciCumulusabstractX-ray crystallography is an important field due to its role in drug discovery and its relevance in bioinformatics experiments of comparative genomics, phylogenomics, evolutionary analysis, ortholog detection, and three-dimensional structure determination. Managing these experiments is a challenging task due to the orchestration of legacy tools and the management of several variations of the same experiment. Workflows can model a coherent flow of activities that are managed by scientific workflow management systems (SWfMS). Due to the huge amount of variations of the workflow to be explored (parameters, input data) it is often necessary to execute X-ray crystallography experiments in High Performance Computing (HPC) environments. Cloud computing is well known for its scalable and elastic HPC model. In this paper, we present a performance evaluation for the X-ray crystallography workflow defined by the PC4 (Provenance Challenge series). The workflow was executed using the SciCumulus middleware at the Amazon EC2 cloud environment. SciCumulus is a layer for SWfMS that offers support for the parallel execution of scientific workflows in cloud environments with provenance mechanisms. Our results reinforce the benefits (total execution time × monetary cost) of parallelizing the X-ray crystallography workflow using SciCumulus. The results show a consistent way to execute X-ray crystallography workflows that need HPC using cloud computing. The evaluated workflow shares features of many scientific workflows and can be applied to other experiments. Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Eduardo S. Ogasawara, Jonas Dias, Fernanda Baião, Marta Mattoso |
IEEE CLOUD | 1 |
| 2011 | Optimizing Phylogenetic Analysis Using SciHmm Cloud-based Scientific WorkflowabstractPhylogenetic analysis and multiple sequence alignment (MSA) are closely related bioinformatics fields. Phylogenetic analysis makes extensive use of MSA in the construction of phylogenetic trees, which are used to infer the evolutionary relationships between homologous genes. These bioinformatics experiments are usually modeled as scientific workflows. There are many alternative workflows that use different MSA methods to conduct phylogenetic analysis and each one can produce MSA with different quality. Scientists have to explore which MSA method is the most suitable for their experiments. However, workflows for phylogenetic analysis are both computational and data intensive and they may run sequentially during weeks. Although there any many approaches that parallelize these workflows, exploring all MSA methods many become a burden and expensive task. If scientists know the most adequate MSA method a priori, it would spare time and money. To optimize the phylogenetic analysis workflow, we propose in this paper SciHmm, a bioinformatics scientific workflow based in profile hidden Markov models (pHMMs) that aims at determining the most suitable MSA method for a phylogenetic analysis prior than executing the phylogenetic workflow. SciHmm is also executed in parallel in a cloud environment using SciCumulus middleware. The results demonstrated that optimizing a phylogenetic analysis using SciHmm considerably reduce the total execution time of phylogenetic analysis (up to 80%). This optimization also demonstrates that the biological results presented more quality. In addition, the parallel execution of SciHmm demonstrates that this kind of bioinformatics workflow is suitable to be executed in the cloud. Kary A. C. S. Ocaña, Daniel de Oliveira 0001, Jonas Dias, Eduardo S. Ogasawara, Marta Mattoso |
eScience | 2 |
| 2011 | Towards a Cost Model for Scheduling Scientific Workflows Activities in Cloud EnvironmentsabstractCloud computing has emerged as a new paradigm that enables scientists to benefit from several distributed resources such as hardware and software. Clouds poses as an opportunity for scientists that need high performance computing infrastructure to execute their scientific experiments. Most of the experiments modeled as scientific workflows manage the execution of several activities and work with large amounts of data. In this way parallel techniques are often a key factor. Parallelizing a scientific workflow in the cloud environment is not trivial. One of the complex tasks is to define the number and types of virtual machines and to design the parallel execution strategy. Due to the number of options for configuring an environment it is a hard task to do it manually and it may produce negative impact on performance. This paper initially proposes a cost model based on concepts of quality of service (QoS) in clouds to help determining an adequate configuration of the environment according to restrictions imposed by scientists. Vitor Viana, Daniel de Oliveira 0001, Marta Mattoso |
SERVICES | 2 |
| 2011 | Many task computing for orthologous genes identification in protozoan genomes using HydraabstractSUMMARY One of the main advantages of using a scientific workflow management system (SWfMS) is to orchestrate data flows among scientific activities and register provenance of the whole workflow execution. Nevertheless, the execution control of distributed activities in high performance computing environments by SWfMS presents challenges such as steering control and provenance gathering. Such challenges may become a complex task to be accomplished in bioinformatics experiments, particularly in Many Task Computing scenarios. This paper presents a data parallelism solution for a bioinformatics experiment supported by Hydra, a middleware that bridges SWfMS and high performance computing to enable workflow parallelization with provenance gathering. Hydra Many Task Computing parallelization strategies can be registered and reused. Using Hydra, provenance may also be uniformly gathered. We have evaluated Hydra using an Orthologous Gene Identification workflow. Experimental results show that a systematic approach for distributing parallel activities is viable, sparing scientist time and diminishing operational errors, with the additional benefits of distributed provenance support. Copyright © 2011 John Wiley & Sons, Ltd. Fábio Coutinho, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Vanessa Braganholo, Alexandre A. B. Lima, Alberto M. R. Dávila, Marta Mattoso |
Concurr. Comput. Pract. Exp. | 3 |
| 2011 | An Algebraic Approach for Data-Centric Scientific Workflows
Eduardo S. Ogasawara, Daniel de Oliveira 0001, Patrick Valduriez, Jonas Dias, Fábio Porto 0001, Marta Mattoso |
Proc. VLDB Endow. | 2 |
| 2010 | SciCumulus: A Lightweight Cloud Middleware to Explore Many Task Computing Paradigm in Scientific WorkflowsabstractMost of the large-scale scientific experiments modeled as scientific workflows produce a large amount of data and require workflow parallelism to reduce workflow execution time. Some of the existing Scientific Workflow Management Systems (SWfMS) explore parallelism techniques - such as parameter sweep and data fragmentation. In those systems, several computing resources are used to accomplish many computational tasks in homogeneous environments, such as multiprocessor machines or cluster systems. Cloud computing has become a popular high performance computing model in which (virtualized) resources are provided as services over the Web. Some scientists are starting to adopt the cloud model in scientific domains and are moving their scientific workflows (programs and data) from local environments to the cloud. Nevertheless, it is still difficult for the scientist to express a parallel computing paradigm for the workflow on the cloud. Capturing distributed provenance data at the cloud is also an issue. Existing approaches for executing scientific workflows using parallel processing are mainly focused on homogeneous environments whereas, in the cloud, the scientist has to manage new aspects such as initialization of virtualized instances, scheduling over different cloud environments, impact of data transferring and management of instance images. In this paper we propose SciCumulus, a cloud middleware that explores parameter sweep and data fragmentation parallelism in scientific workflow activities (with provenance support). It works between the SWfMS and the cloud. SciCumulus is designed considering cloud specificities. We have evaluated our approach by executing simulated experiments to analyze the overhead imposed by clouds on the workflow execution time. Daniel de Oliveira 0001, Eduardo S. Ogasawara, Fernanda Baião, Marta Mattoso |
IEEE CLOUD | 1 |
| 2010 | Data parallelism in bioinformatics workflows using HydraabstractLarge scale bioinformatics experiments are usually composed by a set of data flows generated by a chain of activities (programs or services) that may be modeled as scientific workflows. Current Scientific Workflow Management Systems (SWfMS) are used to orchestrate these workflows to control and monitor the whole execution. It is very common in bioinformatics experiments to process very large datasets. In this way, data parallelism is a common approach used to increase performance and reduce overall execution time. However, most of current SWfMS still lack on supporting parallel executions in high performance computing (HPC) environments. Additionally keeping track of provenance data in distributed environments is still an open, yet important problem. Recently, Hydra middleware was proposed to bridge the gap between the SWfMS and the HPC environment, by providing a transparent way for scientists to parallelize workflow executions while capturing distributed provenance. This paper analyzes data parallelism scenarios in bioinformatics domain and presents an extension to Hydra middleware through a specific cartridge that promotes data parallelism in bioinformatics workflows. Experimental results using workflows with BLAST show performance gains with the additional benefits of distributed provenance support. Fábio Coutinho, Eduardo S. Ogasawara, Daniel de Oliveira 0001, Vanessa Braganholo, Alexandre A. B. Lima, Alberto M. R. Dávila, Marta Mattoso |
HPDC | 3 |
| 2010 | Adaptive Normalization: A novel data normalization approach for non-stationary time seriesabstractData normalization is a fundamental preprocessing step for mining and learning from data. However, finding an appropriated method to deal with time series normalization is not a simple task. This is because most of the traditional normalization methods make assumptions that do not hold for most time series. The first assumption is that all time series are stationary, i.e., their statistical properties, such as mean and standard deviation, do not change over time. The second assumption is that the volatility of the time series is considered uniform. None of the methods currently available in the literature address these issues. This paper proposes a new method for normalizing non-stationary heteroscedastic (with non-uniform volatility) time series. The method, named Adaptive Normalization (AN), was tested together with an Artificial Neural Network (ANN) in three forecast problems. The results were compared to other four traditional normalization methods, and showed AN improves ANN accuracy in both short- and long-term predictions. Eduardo S. Ogasawara, Leonardo C. Martinez, Daniel de Oliveira 0001, Geraldo Zimbrão, Gisele L. Pappa, Marta Mattoso |
IJCNN | 3 |