VLDB 2026 Research / reviewers in the wild / expert
Aletéia P. F. Araújo
dblp:38/2179 · also Aleteia Araujo, Aletéia Patrícia Favacho de Araújo
· DBLP profile ↗
50ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0003-4645-6700ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 since 2021Systems, architecture and hardware · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Orama++: Extending Serverless Benchmarking with Tiobe and Halstead Metrics for Improved Performance Prediction
Leonardo Rebouças de Carvalho, Geraldo P. R. Filho, Aletéia P. F. Araújo |
CLOSER | 3 |
| 2026 | Evaluating Scalability Using Open Table File Formats in Cloud Lakehouse Architectures with TPC-DS
Italo V. P. Guimaraes, Aletéia P. F. Araújo |
CLOSER | 2 |
| 2025 | Self-Tuning DBMS: A Data-Driven Approach to Buffer Pool Optimization in Enterprise SystemsabstractThis article tackles the critical challenge of optimizing the buffer pool, a core component of Database Management Systems (DBMS) that caches frequently accessed data pages, where manual configuration often proves inadequate in dynamic, high-demand environments. To address this gap, we present an automated, data-driven methodology that combines advanced Machine Learning techniques with Bayesian optimization. Our approach follows a systematic three-phase process: (1) Exploratory Factor Analysis (EFA) coupled with K-means clustering to uncover latent factors and reduce the dimensionality of performance metrics; (2) LASSO regression to identify and rank the most influential configuration parameters; and (3) Bayesian optimization using Gaussian Process modeling with acquisition functions (Expected Improvement, Probability of Improvement, and Upper Confidence Bound) to fine-tune buffer pool settings. The main contributions of this work include a novel automated framework for DBMS tuning that simplifies configuration, enhances memory management, and boosts performance efficiency. We validated the proposed solution using real workloads collected from a large-scale financial system in Latin America, achieving up to a 45% reduction in maximum data access wait times, confirming improvements in performance and scalability. Eduardo Mendizabal, Geraldo P. R. Filho, Marcelo Antonio Marotta, Marcos F. Caetano, João J. C. Gondim, Lucas Bondan, Aletéia P. F. Araújo |
CLEI | 7 |
| 2025 | An Experiment on Serialization of DIS Messages Using Protocol BuffersabstractThis paper explores Protocol Buffers as a serialization mechanism for Distributed Interactive Simulation (DIS) Protocol Data Units (PDUs) to enhance performance in distributed simulation communications. The DIS protocol is a widely used standard for connecting various simulators across geographically dispersed locations, utilizing PDUs for communication. The experiment compares the performance of Entity State PDUs (ESPDU) serialization using the standard DIS method against Protocol Buffers, focusing on encoding time, decoding time, packet transmission rate, and packet loss. A client/server application prototype was developed, and tests were conducted in a LAN environment with messages being generated at one message per millisecond. The results indicated the Protocol Buffers had good performance with 56.96 % less decoding time. However, DIS still showed better performance in encoding time, being 41.20 % faster. Future research could explore combining C-DIS strategies with Protocol Buffers to leverage compact serialization and improve military simulation networking performance. José Niuton da Nova, Flavio de Barros Vidal, Aletéia P. F. Araújo |
CLEI | 3 |
| 2025 | Introducing Computing in Brazilian Basic Education: Insights from Teachers' PerspectivesabstractThis paper results from a project for promoting students’ interest in the field of computing through teaching computer science, robotics, and programming in Brazilian public schools. Here, we present a mixed-methods study conducted with elementary and secondary school teachers involved with such a project. The study analyzes teachers’ perceptions regarding the project’s impact on their professional trajectories, on the students, and on the school and community contexts in which they operate. The findings highlight the importance of early exposure to computing as a strategy to bridge the gap between educational levels and as a tool for inclusion and social transformation. The results indicate that participation in the project has generated positive outcomes. Teachers report the development of new skills—particularly technical skills in computing—increased motivation for teaching, and the opportunity to make a meaningful impact on students’ lives. Furthermore, they observe positive changes among the girls, especially in terms of learning and engagement, as well as improvements in the school and community environments, including greater parental involvement in the students’ educational journey. Aline de Galés Silva, Renata Muniz Prado, Maristela Holanda, Maria Emília M. T. Walter, Mirella M. Moro, Aletéia P. F. Araújo |
CLEI | 6 |
| 2024 | Graduate Programs in Computing at the University of Brasilia: Comparison of Academic Papers and Collaborations by GenderabstractThe computing field has a low level of gender diver-sity, being predominantly male. This diversity gap is reflected at different academic levels, ranging from undergraduate degrees to master's and doctoral qualifications. As in other parts of the world, the University of Brasilia, one of the top 10 universities in Brazil, has a low rate of women in Computing in its graduate programs. According to data from CAPES (Coordination for the Improvement of Higher Education Personnel in Brazil), the computing area is one of the areas of Exact Sciences with the lowest number of women proportionally in Brazil. In this context, this paper aims to present an analysis of the number of publications and scientific collaborations related to the gender of researchers in the Graduate Program in Informatics (PPG I) at the University of Brasilia. To develop this research, technologies such as scraping were used to collect data from the program's professors, articles published (only full papers in conferences and academic journals) by them and names of the people who collaborated in the preparation of the articles, a database graph-based noSQL was used to generate the relationship networks. A relationship network was created, in which it was possible to analyze the relationship level of each professor in the program. As initial results, the average number of articles published is the same by gender, we did not find significant differences in the number of publications between men and women. In the study we only counted the number of publications, we did not analyze the impact of the publications. Regarding collaborations, initial results indicate that women are more collaborative than men in the program, with a higher degree of collaboration than the average for male researchers. In general, in the partial results, collaboration networks for journals and conferences present the same result, that female researchers proportionally collaborate more than men, both internally (publications co-authored with members of the University of Brasilia) and externally (publication with co-authorship outside the University of Brasilia), in academic publications. Mariana Alencar do Vale, Maristela Holanda, Célia Ghedini Ralha, Aletéia P. F. Araújo, Dilma Da Silva |
FIE | 4 |
| 2024 | SWPTMAC: Sleep Wake-up Power Transfer MAC ProtocolabstractWireless Underground Sensor Networks (WUSNs) are complex systems comprised of subterranean sensors interconnected through wireless communication technologies. These networks fulfill a crucial role in monitoring subsurface environments. However, they grapple with a formidable challenge concerning their Network Lifetime (NL), which can be defined as the maximum duration over which the network remains operational and thus connected to a designated observation area. Given the paramount significance of prolonging NL to ensure comprehensive coverage of the observed region, the deployment of wireless power transfer stands out as a preeminent solution for augmenting NL. Nonetheless, the existing sleep-wakeup protocols have not been originally engineered to support this paradigm, which has subsequently resulted in suboptimal network performance. Therefore, we present a study to introduce a novel sleep-wakeup protocol explicitly tailored for wireless power transfer in WUSNs with the overarching aim of optimizing the network’s operational lifetime called Sleep-Wakeup Power-Transfer Media Access Control (SWPTMAC). The evaluation of SWPTMAC has been conducted through comprehensive simulations leveraging the Castália simulator. The empirical findings disclosed an average improvement of approximately 24% when contrasted against incumbent protocols in the domain. Luan Borges Dos Santos, Geraldo P. R. Filho, Lucas Bondan, Marcos F. Caetano, Aletéia P. F. Araújo, Marcelo Antonio Marotta |
NOMS | 5 |
| 2024 | MAS-Cloud+: A novel multi-agent architecture with reasoning models for resource management in multiple providers
Aldo H. D. Mendes, Michel J. F. Rosa, Marcelo Antonio Marotta, Aletéia P. F. Araújo, Alba Cristina Magalhaes Alves de Melo, Célia Ghedini Ralha |
Future Gener. Comput. Syst. | 4 |
| 2023 | FaaS Benchmarking over Orama Framework's Distributed Architecture
Leonardo Rebouças de Carvalho, Bruno Kamienski, Aletéia P. F. Araújo |
CLOSER | 3 |
| 2023 | Achieving Observability on Fog Computing with the Use of Open-Source Tools
Breno G. S. Costa, Abhik Banerjee, Prem Prakash Jayaraman, Leonardo Rebouças de Carvalho, João Bachiega Jr., Aletéia P. F. Araújo |
MobiQuitous (2) | 6 |
| 2023 | AFMC: An alignment framework for multiple computing services and providersabstractSummary The Hirschberg algorithm is commonly used for protein sequence alignment, which is a very important task in bioinformatics. This article presents the AFMC framework for using the Hirschberg method to perform sequence alignment in multiple cloud computing services of different models, such as Infrastructure‐as‐a‐Service and Function‐as‐a‐Service (FaaS). Experiments were carried out in which several instances of AWS EC2, Azure VMs and Google Compute Engine as well as varied configurations of AWS Lambda, Azure Function, and Google Cloud Function were used to pairwise align COVID‐19 spike proteins. The services were submitted to different levels of simultaneity to align the genetic sequences. The findings reveal that there is a tradeoff between predicted execution time and cost for this application, for example, FaaS‐oriented cloud service models generally took less time to process the workloads. On the other hand, it was observed that, as the level of concurrence increased, there was a marked augmentation in cost. In this context, a framework that provides multi cloud solutions for bioinformatics such as AFMC is essential. Leonardo Rebouças de Carvalho, Alba Cristina Magalhaes Alves de Melo, Aletéia P. F. Araújo |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | From the Sky to the Ground: Comparing Fog Computing with Related Distributed Paradigms
João Bachiega Jr., Breno G. S. Costa, Leonardo Rebouças de Carvalho, Victor H. C. Oliveira, William X. Santos, Maria Clicia Stelling de Castro, Aletéia P. F. Araújo |
CLOSER | 7 |
| 2022 | Orama: A Benchmark Framework for Function-as-a-Service
Leonardo Rebouças de Carvalho, Aletéia P. F. Araújo |
CLOSER | 2 |
| 2022 | Comparison of FaaS Platform Performance in Private Clouds
Marcelo Augusto da Cruz Motta, Leonardo Rebouças de Carvalho, Michel J. F. Rosa, Aletéia P. F. Araújo |
CLOSER | 4 |
| 2022 | Gender Diversity in STEM Graduate Programs at the University of Brasília in BrazilabstractIncreasing gender diversity in STEM graduate programs is a challenge. In Brazil, the National Council for Scientific and Technological Development (CNPq) has classified knowledge into different "broad areas", one of which is Exact and Earth Sciences (EES). This area includes the STEM subjects: Physics, Computer Science, Mathematics, Statistics and Chemistry. These EES areas have a low representation of women. The University of Brasília, one of the top 10 universities in Brazil, has graduate programs (master’s and doctoral degrees) in all these subjects. In this context, this paper has the main research question: What is the level of gender diversity in each EES area at the University of Brasília in master’s and doctoral programs? This research question was analyzed with the indicators of student enrollment, number of graduations, and retention rates in the programs. The data used for analysis were the available Brazilian open public data of graduate programs for 11 years, 2007-2017. The findings include that women are in the minority in the total number of graduates in Computer Science and Physics. Despite the low number of women overall in EES, the Chemistry program stands out with the highest female participation, reaching more women than men at the doctorate level. The program that has the fewest women is Computer Science. This paper presents all the results of this study. Maristela Holanda, Thayanna Klysnney, Aletéia P. F. Araújo, Dilma Da Silva, Roberta B. Oliveira, Carla Koike, Carla Denise Castanho, Juliana Betini Fachini Gomes |
FIE | 3 |
| 2022 | Deep-vacuity: A Proposal of a Machine Learning Platform based on High-performance Computing Architecture for Insights on Government of Brazil Official Gazettes
Leonardo Rebouças de Carvalho, Felipe L. S. Mendes, Jefferson Chaves, Marcos C. Lima, Flavio E. de Deus, Aletéia P. F. Araújo, Flavio de Barros Vidal |
WEBIST | 6 |
| 2022 | Monitoring fog computing: A review, taxonomy and open challenges
Breno G. S. Costa, João Bachiega Jr., Leonardo Rebouças de Carvalho, Michel J. F. Rosa, Aletéia P. F. Araújo |
Comput. Networks | 5 |
| 2022 | Optimized Solutions for Deploying a Militarized 4G/LTE Network With Maximum Coverage and Minimum InterferenceabstractThis work proposes to solve the maximal covering location problem of the Mobile Operations Coordination Center (CCOp Mv), which aims to support the operational command of the Brazilian Army. This problem consists of selecting, in a limited region and with poor communication infrastructure to the ground troops’s operating area, the positions of vehicles equipped with Base Transceiver Station (BTS), the amount needed, and theirs transmission power to be set that maximizes the coverage area and reduce the interference due to the overlap of signals. For this reason, analytical modeling based on the mixed-integer linear problem was proposed that guided two optimization solutions: (i) E-ALLOCATOR – Exact ALLOCATiOn seRvice; and (ii) M-ALLOCATOR – Metaheuristic ALLOCATiOn seRvice. The solutions were evaluated in a scenario that employs CCOp Mv to support a rescue operation based on the tragedy in January 2019 in Brumadinho-MG and compared with a heuristic. The performance evaluation results show evidence of efficiencies in terms of quality and resource savings of the proposed solutions. Furthermore, E-ALLOCATOR has been proven to be suitable for a low workload on the network. At the same time, M-ALLOCATOR is suitable for scenarios with a high workload providing almost optimal solutions within the adequate computational time for all problem instances. Emerson de O. Antunes, Marcos F. Caetano, Marcelo Antonio Marotta, Aletéia P. F. Araújo, Lucas Bondan, Rodolfo I. Meneguette, Geraldo P. R. Filho |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2021 | Resource Prediction Service for Efficient Execution of Bioinformatics Workflows in Federated Cloud with Machine LearningabstractCloud federation emerged to extend the resources available between different interconnected cloud providers for transparent and unlimited availability to the end-user. Cloud orchestration platforms have become a way to centralize demands for high computational power in applications such as Bioinformatics workflows. The large quantity of resources available among several providers in a federation makes it challenging to choose a suitable one for particular workflows. This work proposes a Machine Learning Resource Prediction Service called sPCRAM. sPCRAM uses a machine learning model combined with a GRASP metaheuristic to transparently and adequately dimension the resources, determining the monetary cost and the runtime before the workflow execution. sPCRAM interactively allows the user to set the execution type, calibrate time and cost. Such executions can have, for example, long duration and low cost, as well as a shorter duration and a higher cost. The results demonstrate that sPCRAM can appropriately estimate runtime and cost for cloud federation resources on average 97,70% faster than the brute force technique for resource selection. Matheus Sobrinho, Michel J. F. Rosa, Waldeyr M. C. Silva, Aletéia P. F. Araújo |
BIBM | 4 |
| 2021 | Sense of Belonging of Female Undergraduate Students in Introductory Computer Science Courses at University of Brasília in BrazilabstractFull Paper - The field of Computer Science (CS) has been of little interest to women straight out of high school when considering undergraduate majors in Brazil. At the University of Brasília, a top-ten university in Brazil, female undergraduate students account for less than 15% of the students in the Department of Computer Science. According to Stout and Blaney, a sense of intellectual belonging is “the sense that one is believed to be a competent member of the community”. This perception may be especially challenging for members of underrepresented minority groups, such as female undergraduate students in CS majors. In this context, this paper addresses two research questions: i) “How does the intellectual sense of belonging of female students compare to the male students' in introduction to computer science courses?”; ii) Is it similar for female undergraduate students in both CS and non-CS majors?”. We devised a questionnaire for students in the introduction to computer science courses for different majors. We analyzed the responses and, in general, introductory programming courses are challenging for all students, however, female students feel worse about their computing competencies than male ones. Maristela Holanda, Aletéia P. F. Araújo, Dilma Da Silva, George von Borries, Roberta B. Oliveira, Carla Koike, Carla Denise Castanho |
FIE | 2 |
| 2021 | Computational resource and cost prediction service for scientific workflows in federated clouds
Michel J. F. Rosa, Célia Ghedini Ralha, Maristela Holanda, Aletéia P. F. Araújo |
Future Gener. Comput. Syst. | 4 |
| 2020 | Performance Comparison of Terraform and Cloudify as Multicloud OrchestratorsabstractStudies indicate that by 2022 multicloud models will reach 75% of the cloud computing market. To handle the high-featured set of computing capability on this paradigm, often called "Sky Computing", infrastructure tools such as cloud orchestrators has emerged. This paper analyzes the most referenced tools in the literature such as Cloudify, Heat, CloudFormation, Terraform and Cloud Assembly, as well as the TOSCA standard. The literature review, complemented by a practical experiment, revealed that Terraform and Cloudify presents great affinity with Sky Computing scenarios. In the experiment Terraform outperformed Cloudify in several aspects. Leonardo Rebouças de Carvalho, Aletéia P. F. Araújo |
CCGRID | 2 |
| 2020 | Remote Procedure Call Approach using the Node2FaaS Framework with Terraform for Function as a Service
Leonardo Rebouças de Carvalho, Aletéia P. F. Araújo |
CLOSER | 2 |
| 2020 | What do Female Students in Middle and High Schools Think about Computer Science Majors in Brasilia, Brazil? A Survey in 2011 and 2019abstractResearch Full Paper - Computer Science majors lack gender diversity in Brasília, Brazil. Women are an underrepresented minority group in these majors. At the University of Brasilia, one of the top ten universities in Brazil, female undergraduate students account for less than 15% of the students in the Department of Computer Science. In an effort to understand the lack of interest in Computer Science majors among women, this paper addresses the following research questions: 1)Are female students in high school aware that Computer Science majors are predominantly male? 2)Are families of girls from Brasília supportive of their enrolment in Computer Science majors? 3)Do girls from Brasília think that Computer Science majors need a lot of Math? 4)Do female students in high school think that it is difficult to get a job in the field of Computing, with a good salary, and sufficient leisure time? and 5)Which factors influence a female student's choice of a Computer Science major? We devised a questionnaire and applied it to female students in middle and high school on two occasions, in October 2011 (1391 responses) and in July 2019 (429 responses). This paper presents the analysis of the data from the responses, which indicates that the girls' perceptions of Computing have not changed in those years. Maristela Holanda, Roberto Nunes Mourão, George von Borries, Guilherme Novaes Ramos, Aletéia P. F. Araújo, Maria Emília M. T. Walter |
FIE | 5 |
| 2019 | Data Provenance Management of Bioinformatics Workflows in Federated CloudsabstractIn Bioinformatics, the reproducibility of experiments is a fundamental principle to which the data provenance significantly contributes through the acquisition and management of information about the trajectory of the data in a workflow. In addition to the data provenance, aspects such as program configurations and the entire computational environment must be considered to achieve this goal. Cloud computing can provide computational resources, hiding technical details and providing an accessible and configurable on-demand environment for researchers. Cloud federation enables the broad distribution of services and a flexible combination of computing power. Considering this particular scenario, we propose a management platform that collects the provenance data of Bioinformatics workflows allowing portability between different clouds in federated clouds using the Infrastructure as a Service model. This platform is composed of tools that support running these workflows with data provenance captured in NoSQL databases. The retrospective data provenance is captured according to the PROV-DM norms. The findings indicate that the proposed platform can be used for some of the most commoninly available clouds. Polyane Wercelens, Waldeyr M. C. Silva, Klayton Castro, Aletéia P. F. Araújo, Sérgio Lifschitz, Maristela Holanda |
BIBM | 4 |
| 2019 | Framework Node2FaaS: Automatic NodeJS Application Converter for Function as a ServiceabstractCloud computing emerged in the area of computer science as a means to achieve significant cost and time savings when starting projects. Among the various cloud models available, this work highlights Function as a Service - FaaS, and proposes the Node2FaaS framework for automatic conversion of applications written in NodeJS to work in a transparent way with the FaaS model. The experiments demonstrated significant gains of up to 170% at runtime for applications with high file I/O requirements. Applications with high CPU and RAM consumption also have benefits in adopting FaaS after conversion, but only when a threshold of competing processes is reached. Leonardo Rebouças de Carvalho, Aletéia P. F. Araújo |
CLOSER | 2 |
| 2019 | Cloud.Jus: Architecture for Provisioning Infrastructure as a Service in the Government Sector
Klayton Castro, Gabriel R. D. Macedo, Aletéia P. F. Araújo, Leonardo Rebouças de Carvalho |
CLOSER | 3 |
| 2019 | ArchaDIA: An Architecture for Big Data as a Service in Private Cloud
Marco Antonio Sousa Reis, Aletéia P. F. Araújo |
CLOSER | 2 |
| 2019 | Multiagent system for dynamic resource provisioning in cloud computing platforms
Célia Ghedini Ralha, Aldo H. D. Mendes, Luiz A. Laranjeira, Aletéia P. F. Araújo, Alba Cristina Magalhaes Alves de Melo |
Future Gener. Comput. Syst. | 4 |
| 2018 | Cost and Time Prediction for Efficient Execution of Bioinformatics Workflows in Federated Cloud
Michel J. F. Rosa, Aletéia P. F. Araújo, Felipe L. S. Mendes |
BIBM | 2 |
| 2018 | Performance and Cost Analysis Between On-Demand and Preemptive Virtual Machines
Breno G. S. Costa, Marco Antonio Sousa Reis, Aletéia P. F. Araújo, Priscila Solís Barreto |
CLOSER | 3 |
| 2018 | A Hadoop Open Source Backup Solution
Heitor Faria, Rodrigo Otávio Ribeiro Hagstrom, Marco Antonio Sousa Reis, Breno G. S. Costa, Edward de Oliveira Ribeiro, Maristela Holanda, Priscila Solís Barreto, Aletéia P. F. Araújo |
CLOSER | 8 |
| 2018 | An Investigative Analysis of Quality of Service Metrics on OpenStack
Edna Dias Canedo, Ítalo Paiva Batista, Emilie T. de Morais, Aletéia P. F. Araújo |
ICCSA (1) | 4 |
| 2017 | AProvBio: An architecture for data provenance in bioinformatics workflows using graph databaseabstractMany scientific experiments in Bioinformatics are executed as computational workflows. Frequently, it is necessary to re-run an experiment under the original circumstances in which it was run to recognize and validate it. Data provenance concerns the origin of data. Knowing the data source facilitates the understanding and analysis of the results, by detailing and documenting the history and the paths of the input data, from the beginning to the end of an experiment. Therefore, in this context, data provenance can be applied when experimenting traceability. This document presents AProvBio, an architecture that can perform the data provenance of scientific experiments in bioinformatics automatically, using the provenance data model PROV-DM and in a graph database. The architecture can perform the automatic provenance type prospectively, retrospectively and with user-defined data. Thus, the architecture stores and captures information obtained during the execution of the data generation processes with user-defined data information, such as features and versions of the programs used. A graph model, based on the PROV-DM model, was proposed for storing the data provenance. The PROV-DM can be represented by a graph, it allows for a more natural modelling, as well as expressing queries at a more natural level, and the implementation of efficient algorithms to perform specific operations. Rodrigo F. Almeida, Waldeyr M. C. Silva, Klayton Castro, Maria Emília M. T. Walter, Aletéia P. F. Araújo, Maristela Holanda, Sérgio Lifschitz |
BIBM | 5 |
| 2017 | Data provenance management for bioinformatics workflows using NoSQL database systems in a cloud computing environmentabstractComputer science solutions for molecular biology problems are often presented in the form of workflows. There is a set of activities performed by different processing entities through managed tasks. Knowledge about the data trajectory throughout a given workflow enables reproducibility by data provenance. In order to reproduce an in silico bioinformatics experiment one must consider other aspects besides those steps followed by a workflow. Indeed, the computational settings in which the involved programs run is a requirement for reproducibility. Cloud computing technology may hide the technical details and make it easier for the user to set up such an on-demand environment. NoSQL database systems have also gained popularity, particularly in the cloud. Considering this particular scenario, we have planned and executed a research study about a bioinformatics workflow running in an IaaS cloud computing environment. We have persisted provenance data according to the PROV-DM model, using different types of NoSQL database systems. We present in this paper some preliminary results from our research work, where we have explored the characteristics of several NoSQL database systems to persist provenance data. Fernanda Hondo, Polyane Wercelens, Waldeyr M. C. Silva, Klayton Castro, Ingrid Santana, Maria Emília M. T. Walter, Aletéia P. F. Araújo, Maristela Holanda, Sérgio Lifschitz |
BIBM | 7 |
| 2017 | Cost Optimization on Public Cloud Provider for Big Geospatial Data
João Bachiega Jr., Marco Antonio Sousa Reis, Aletéia P. F. Araújo, Maristela Holanda |
CLOSER | 3 |
| 2016 | An evaluation of data replication for bioinformatics workflows on NoSQL systemsabstractMany research projects in bioinformatics may be viewed as scientific workflows. Biologists often run multiple times the same workflow with different parameters in order to refine their data analysis. These executions generate a large volume of files with different formats, which need to be stored for future evaluations. New database models, like NoSQL systems, could be considered to deal with large volumes of data, particularly in distributed systems. This work presents a data replication impact assessment from the execution of scientific workflows for two NoSQL database management systems: Cassandra and MongoDB. Iasmini Lima, Matheus Oliveira, Diego Kieckbusch, Maristela Holanda, Maria Emília M. T. Walter, Aletéia P. F. Araújo, Márcio Victorino, Waldeyr M. C. Silva, Sérgio Lifschitz |
BIBM | 6 |
| 2016 | BioNimbuZ: A federated cloud platform for bioinformatics applicationsabstractChallenges in bioinformatics include tools to treat large-scale processing, mainly due to the large volumes of data generated by high-throughput sequencing machines. Besides, many of these tools are not user friendly, and do not distribute their workloads properly. In federated cloud environments, even though services and resources are shared and available online, the processes of a workflow execution are almost entirely not automated, and the majority of these processes do not efficiently balance their workloads. This paper presents the federated cloud platform, called BioNimbuZ, a hybrid platform designed to execute bioinformatics applications easily and efficiently, with good workload balance. Our tests were performed using a real bioinformatics workflow, with fragments generated by the Illumina sequencer, having achieved good performance in practice. Michel J. F. Rosa, Breno Moura, Guilherme Vergara, Lucas Santos, Edward de Oliveira Ribeiro, Maristela Holanda, Maria Emília M. T. Walter, Aletéia P. F. Araújo |
BIBM | 8 |
| 2016 | Hiding color watermarks in halftone images using maximum-similarity binary patterns
Pedro Garcia Freitas, Mylène C. Q. Farias, Aletéia P. F. Araújo |
Signal Process. Image Commun. | 3 |
| 2016 | Enhancing inverse halftoning via coupled dictionary training
Pedro Garcia Freitas, Mylène C. Q. Farias, Aletéia P. F. Araújo |
Signal Process. Image Commun. | 3 |
| 2015 | A study of genomic data provenance in NoSQL document-oriented database systemsabstractThis work considers a scientific experiment as a computational workflow. Provenance models store details of each workflow execution, including produced data, computational tools parameters and their versions, among others. This way, scientists can review details of a particular workflow execution, compare information generated among different executions and plan new ones efficiently. In the bioinformatics domain, particularly in the presence of large volumes of data, persistency of those data generated during the workflow execution is still a research challenge. In this article, we consider a study on provenance data storage for bioinformatics in a document-oriented NoSQL database system. We present data modeling issues and discuss an actual implementation into MongoDB. Valeria Guimarâes, Fernanda Hondo, Rodrigo F. Almeida, Harley Vera Olivera, Maristela Holanda, Aletéia P. F. Araújo, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 6 |
| 2015 | Improved performance of inverse halftoning algorithms via coupled dictionariesabstractInverse halftoning techniques are known to introduce visible distortions (typically, blurring or noise) into the reconstructed image. To reduce the severity of these distortions, we propose a novel training approach for inverse halftoning algorithms. The proposed technique uses a coupled dictionary (CD) to match distorted and original images via a sparse representation. This technique enforces similarities of sparse representations between distorted and non-distorted images. Results show that the proposed technique can improve the performance of different inverse halftone approaches. Images reconstructed with the proposed approach have a higher quality, showing less blur, noise, and chromatic aberrations. Pedro Garcia Freitas, Mylène C. Q. Farias, Aletéia P. F. Araújo |
ICME | 3 |
| 2014 | Storing provenance data of genome project workflows using graph databaseabstractMany scientific experiments are designed as computational workflows in bioinformatics. However, the amount of data generated increases at every phase of each execution, hindering the identification of the source and the transformation of data. Therefore, it has become necessary to create new tools to store data provenance, mainly which resources and parameters were used to generate the results, among other information, to validate and publish the experiment. In this paper, we propose to use graph database to store data provenance using the PROV-DM model of bioinformatics workflows. To validate the model, we developed a simulator that worked as a logbook to capture data provenance. A workflow with real genomic data showed that very little additional data should be stored, which means that our provenance model can be easily included in genome projects. Rodrigo Pinheiro, Bruno Aires, Aletéia P. F. Araújo, Maristela Holanda, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 3 |
| 2014 | A Storage Policy for a Hybrid Federated Cloud platform: A Case Study for BioinformaticsabstractBioinformatics tools require large-scale processing mainly due to very large databases achieving gigabytes of size. In federated cloud environments, although services and resources may be shared, storage is particularly difficult, due to distinct computational capabilities and data management policies of several separated clouds. In this work, we propose a storage policy for BioNimbuZ, a hybrid federated cloud platform designed to execute bioinformatics applications. Our storage policy, BioClouZ, aims to perform efficient choices to distribute and replicate files to the best available cloud resources in the federation in order to reduce computational time. BioClouZ uses four parameters - latency, uptime, free size and cost, weighted (according to ad hoc tests) to model their influences to data storage and recovery. Experiments were performed with real biological data executing a commonly used tool to map short reads in a reference genome in BioNimbuZ, composed of clouds executing in Amazon EC2, Azure and University of Brasilia. The results showed that, when compared to the greedy algorithm first used in BioNimbuZ, the BioClouZ policy significantly improved the total execution time due to more efficient choices of the clouds to store the files. Other bioinformatics applications can be used with BioClouZ in BioNimbuZ as well, since the platform was designed independently from particular tools and databases. Deric Lima, Breno Moura, Gabriel S. S. de Oliveira, Edward de Oliveira Ribeiro, Aletéia P. F. Araújo, Maristela Holanda, Roberto C. Togawa, Maria Emília M. T. Walter |
CCGRID | 5 |
| 2014 | Model to estimate the size of a Hadoop cluster - HCEmabstractThis paper describes a model which aims to estimate the size of a cluster running Hadoop framework for the processing of large datasets at a given timeframe. As main contributions it denes (i) a light layer of optimization for MapReduce jobs, (ii) presents a model to estimate the size cluster for a Hadoop framework and (iii) performs tests using a real environment - the Amazon Elastic MapReduce. The proposed approach works with the MapReduce to dene the main configuration parameters and determines computational resources of hosts in the cluster in order to meet the desired runtime for the requirements of a given workload requirement. Thus, the results show that the proposed model is able to avoid to over-allocation or sub-allocation of computing resources on a Hadoop cluster. Jose Benedito de Souza Brito, Aletéia P. F. Araújo |
ICPADS | 2 |
| 2013 | ACOsched: A scheduling algorithm in a federated cloud infrastructure for bioinformatics applicationsabstractTask scheduling in a federated cloud environment is a complex problem since there are several cloud providers presenting distinct memory and storage capacities that should be addressed. This article focus on the task scheduling problem in BioNimbuZ, a federated cloud infrastructure for executing bioinformatics applications, which was previously proposed by our group. We present a scheduling algorithm based on Load Balancing Ant Colony (LBACO), called ACOsched, to perform efficient distribution of tasks by finding the best cloud in the federation to execute these tasks. We developed experiments using real biological data, executing the Bowtie mapping tool on one instance of BioNimbuZ, composed by two cloud providers, Amazon EC2 and a bioinformatics laboratory at the University of Brasilia/Brazil. The obtained results show that ACOsched led to a significant improvement in the makespan time of Bowtie executing in BioNimbuZ, when compared to the simple round robin algorithm called DynamicAHP, previously developed in this federated cloud infrastrucutre. Gabriel S. S. de Oliveira, Edward de Oliveira Ribeiro, Diogo A. Ferreira, Aletéia P. F. Araújo, Maristela Holanda, Maria Emília M. T. Walter |
BIBM | 4 |
| 2013 | Automatic capture of provenance data in genome project workflowsabstractMany scientific experiments are designed as computational workflows in the bioinformatics domain, which facilitates implementation and analysis. However, the amount of data generated increases at every phase of each execution, hindering the identification of the source and the data transformation. Therefore, it has become necessary to create new tools to verify automatically which resources and parameters were used to generate the results, among other information to validate and publish the experiment. This functionality of automatically capturing data provenance has been receiving attention in the scientific community, primarily with regard to bioinformatics projects, due the fact that the same workflow is executed several times with different parameters and versions of the tools. In this paper, we propose to use relational schema to automatically store data provenance using the PROV-DM model for workflows in bioinformatics projects. Rodrigo Pinheiro, Maristela Holanda, Aletéia P. F. Araújo, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 3 |
| 2012 | Task Scheduling in a Federated Cloud Infrastructure for Bioinformatics Applications
C. A. L. Borges, Hugo Saldanha, Edward de Oliveira Ribeiro, Maristela Holanda, Aletéia P. F. Araújo, Maria Emília M. T. Walter |
CLOSER | 5 |
| 2011 | A Cloud Architecture for Bioinformatics Workflows
Hugo Saldanha, Edward de Oliveira Ribeiro, Maristela Holanda, Aletéia P. F. Araújo, Genaína Nunes Rodrigues, Maria Emília M. T. Walter, João Carlos Setubal, Alberto M. R. Dávila |
CLOSER | 4 |
| 2005 | Towards Grid Implementations of Metaheuristics for Hard Combinatorial Optimization ProblemsabstractMetaheuristics are approximate algorithms that are able to find very good solutions to hard combinatorial optimization problems. They do, however, offer a wide range of possibilities for implementations of effective robust parallel algorithms which run in much smaller computation times than their sequential counterparts. We present four slightly differing strategies for the parallelization of an extended GRASP with ILS heuristic for the mirrored traveling tournament problem. Computational results on widely used benchmark instances, using a varying number of processors, illustrate the effectiveness and the scalability of the different strategies. These low communication cost parallel heuristics not only find solutions faster, but also produce better quality solutions than the best known sequential algorithm. Aletéia P. F. Araújo, Sebastián Urrutia |
SBAC-PAD | 1 |