VLDB 2026 Research / reviewers in the wild / expert
Maristela Holanda
dblp:20/1662 · also Maristela T. Holanda, Maristela Terto de Holanda
· DBLP profile ↗
65ranked-venue papers
10as first author
23since 2021 · last 2025
0000-0002-0883-2579ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 25 · 9 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Security and privacy · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Introducing Computing in Brazilian Basic Education: Insights from Teachers' PerspectivesabstractThis paper results from a project for promoting students’ interest in the field of computing through teaching computer science, robotics, and programming in Brazilian public schools. Here, we present a mixed-methods study conducted with elementary and secondary school teachers involved with such a project. The study analyzes teachers’ perceptions regarding the project’s impact on their professional trajectories, on the students, and on the school and community contexts in which they operate. The findings highlight the importance of early exposure to computing as a strategy to bridge the gap between educational levels and as a tool for inclusion and social transformation. The results indicate that participation in the project has generated positive outcomes. Teachers report the development of new skills—particularly technical skills in computing—increased motivation for teaching, and the opportunity to make a meaningful impact on students’ lives. Furthermore, they observe positive changes among the girls, especially in terms of learning and engagement, as well as improvements in the school and community environments, including greater parental involvement in the students’ educational journey. Aline de Galés Silva, Renata Muniz Prado, Maristela Holanda, Maria Emília M. T. Walter, Mirella M. Moro, Aletéia P. F. Araújo |
CLEI | 3 |
| 2025 | Teaching Algorithms to Indigenous Students of Brazil's AmazonabstractThe Constitution of Brazil and its subsequent laws have established various rights and protections for Indigenous peoples, among them the right to Indigenous schools where their culture and native language must be taught, learned, and preserved as something alive and essential to their well-being. Brazil's National Digital Education Policy, which mandates the teaching of computing in K-12 education, is a recent development not yet implemented in indigenous schools, where access to computers and the Internet is still quite limited. To promote the inclusion of indigenous populations in higher education, the University of Brasília (UnB), in collaboration with Brazil's National Foundation for Indigenous Peoples, has created an admission pathway dedicated to students from these populations. In 2022, UNB's Computer Science Department welcomed its first three Indigenous students from the Ticuna community in the Amazon region of Brazil. The Ticuna people represent the largest indigenous ethnic group in Brazil. Ticuna students, computer science professors, and computer science students at UnB have collaborated to address the gap in the K-12 teaching of computing in Ticuna communities. This work describes the materials created by the indigenous students for teaching computing in their communities within the context of their culture and language. Maristela Holanda, Edison Ishikawa, Dilma Da Silva |
SIGCSE (2) | 1 |
| 2024 | A Quantitative Study of Publications about Underrepresented Minority Undergraduate Students in Computer Science Majors in the United StatesabstractThis is a full research paper. Including underrepresented minority groups (URM) in the computing field has been a challenge in the United States for a long time. Specifically, African American, Hispanic/Latinx and native American students continue to have low participation in computer science majors despite efforts to include these groups in Computer Science majors in the United States. In this context, the paper presents the research question: What types of publications are there in the literature about activities focused on including underrepresented minority groups in Computer Science majors in the United States? URM in this paper means minority African-American, Hispanic/Latinx and Native American students. To answer this question, a systematic literature mapping process covering the last 15 years (2009–2023) was applied using four academic sources: Scopus, Web of Science, IEEE Xplore, and ACM Digital Library. The inclusion criteria were: documents published in a conference or journal; documents published in the period 2009-2023; computer science and education areas; documents focusing on underrepresented minority groups, specifically African American and Latinx/Hispanic and Native American; and documents concerning the undergraduate level. The exclusion criteria were: documents with less than four pages (an attempt to map only full papers); documents in a language other than English; documents that do not mention African American and Latinx/Hispanic; and documents from educational institutions outside of the U.S. We found 521 relevant papers about URM and CS majors, most of which were published at ASEE, FIE, and SIGCSE conferences. There has been an increasing number of papers published in these 15 years. In the last five years, 329 papers were found, representing more than 63% of all papers found in this literature mapping. The universities with the most publications were Florida University, the University of Texas at El Paso, and Purdue University. California, Texas, and Florida are the top three states which the highest number of papers. We classified the paper and found 48 papers for EDM (educational data mining), 163 papers for Educational Model, Report 37, Perception 105, and Program 198. Maristela Holanda, Manuella Valadares, Suxia Cui, Dilma Da Silva |
FIE | 1 |
| 2024 | AI Generated Code Plagiarism Detection in Computer Science Courses: A Literature MappingabstractThis is a full research paper. Integrity in the detection of plagiarism in students' source codes in university programming courses is a research topic for instructors and institutions seeking to improve the quality of their teaching. In particular, introductory courses such as CS1, are of paramount importance, as this is when students gain fundamental knowledge to build their future on. With the latest developments in Large Language Models (LLM) such as ChatGPT, GitHub Copilot, etc., methods of plagiarism have evolved, however methods of detection may not be capable of accurately differentiating between code generated by human and artificial intelligence (AI). In this context, this paper seeks to answer the research question: What does the current literature report on AI generated code plagiarism detection in higher education? To expand on and formulate a comprehensive answer to our research question (RQ), we have formulated six sub-questions: RQ1) How many papers were published per year by country?; RQ2) Which conferences and journals have published most papers on this subject?; RQ3) Which plagiarism detection tools were most often used prior to common AI use?; RQ4) How are educators adapting assignments to minimize the use of AI?; RQ5) Which modern methods are being deployed to specifically detect AI?; RQ6) Which data sources and languages are most prevalent in the literature? The methodology was based on a systematic literature review. Initially, we confined our search for literature to Scopus and Web of Science, however additional literature was included from Google Scholar. Inclusion criteria were applied to include documents from the years 2023 and 2024 (after the launch of ChatGPT), and only published by conferences and journals. Exclusion criteria: papers that do not focus on plagiarism and programming courses; papers that are not about the undergraduate-level; papers not written in English. We found 165 papers via Scopus and WebScience, from which the metadata were collected, resulting in 17 relevant papers selected for this work. The second step was a search in Google Scholar, where we analyzed 200 documents from 2023 (100 relevant documents) and 2024 (100 relevant documents). We used the same inclusion and exclusion criteria, however, we included the ArXiv papers, and found 9 more papers. Following this process, we have identified 26 papers to include in this literary mapping. In this paper we present the answers to these research questions and discussions about this research topic. Archer Simmons, Maristela Holanda, Christiana Chamon, Dilma Da Silva |
FIE | 2 |
| 2024 | Graduate Programs in Computing at the University of Brasilia: Comparison of Academic Papers and Collaborations by GenderabstractThe computing field has a low level of gender diver-sity, being predominantly male. This diversity gap is reflected at different academic levels, ranging from undergraduate degrees to master's and doctoral qualifications. As in other parts of the world, the University of Brasilia, one of the top 10 universities in Brazil, has a low rate of women in Computing in its graduate programs. According to data from CAPES (Coordination for the Improvement of Higher Education Personnel in Brazil), the computing area is one of the areas of Exact Sciences with the lowest number of women proportionally in Brazil. In this context, this paper aims to present an analysis of the number of publications and scientific collaborations related to the gender of researchers in the Graduate Program in Informatics (PPG I) at the University of Brasilia. To develop this research, technologies such as scraping were used to collect data from the program's professors, articles published (only full papers in conferences and academic journals) by them and names of the people who collaborated in the preparation of the articles, a database graph-based noSQL was used to generate the relationship networks. A relationship network was created, in which it was possible to analyze the relationship level of each professor in the program. As initial results, the average number of articles published is the same by gender, we did not find significant differences in the number of publications between men and women. In the study we only counted the number of publications, we did not analyze the impact of the publications. Regarding collaborations, initial results indicate that women are more collaborative than men in the program, with a higher degree of collaboration than the average for male researchers. In general, in the partial results, collaboration networks for journals and conferences present the same result, that female researchers proportionally collaborate more than men, both internally (publications co-authored with members of the University of Brasilia) and externally (publication with co-authorship outside the University of Brasilia), in academic publications. Mariana Alencar do Vale, Maristela Holanda, Célia Ghedini Ralha, Aletéia P. F. Araújo, Dilma Da Silva |
FIE | 2 |
| 2024 | Gender Bias in Tech - Young People's Perception of STEM in Portugal
Helena Elias, Isabel Pedrosa, Maristela Holanda |
WorldCIST (4) | 3 |
| 2023 | Automating Source Code Plagiarism Detection in a Moodle-Based Programming CourseabstractPlagiarism in programming courses in college is an issue. Automatic plagiarism detection tools using source code similarities are important to combat this issue. For example, Moss and JPlag tools are used worldwide by several universities for the detection of similarity in students' source codes. However, the similarity could cause lead to notification of false positives. Given the large number of students in programming courses, considering only similarity could increase the number of students that the instructors should investigate to decide whether there is plagiarism. In this context, this paper presents the ProjPlag tool that seeks to combine similarity detection results with student behavior data, such as the coding process in the learning management system (LMS) and assignment scores during the programming courses. The ProjPlag automates the utilization of Jplag and Moss tools and also analyses students' behavior by means of data extracted from the Moodle platform, the LMS used in the programming courses at the Department of Computer Science at the University of Brasilia. This paper presents an analysis of correlations between the list of confirmed plagiarized assignments with the student behavior data on the Moodle platform, such as time to implement an assignment, the remaining time from the deadline, and the assignment's score. ProjPlag was tested in two different academic semesters at the University of Brasilia in 2022 for the first programming course, (Algorithms and Computer Programming course) in the Department of Computer Science. The findings show that the detection tools can help the instructor to ascertain whether there was plagiarism or not. However, the majority of student behavior data from Moodle had limited relevance in confirming plagiarism and can generate a false alarm. The correlations between confirmed plagiarism and student's behavior (plagiarism and assignments scores, time spent coding, and project scores) on the Moodle platform were low. Rodrigo Aniceto, Maristela Holanda, Dilma Da Silva |
FIE | 2 |
| 2023 | Automatic Formative and Motivational Feedback Personalized for Introductory Programming CourseabstractThe teaching and learning of the first programming language course, commonly called CS1 (Computer Science 1), is a reported challenge for undergraduate students in different majors and universities. In general, these courses have a high failure rate at institutions around the world. At the same time, there has been an increase in the usage of automated feedback tools for programming language courses. These tools present feedback to help the students know if their code is correct or incorrect while studying. This automatic feedback is important for student learning. In this context, this paper presents the AsPin plugin, developed for the virtual learning environment, which uses the student's profile to provide feedback (messages and advice) to personalize their learning. The automatic feedback helps the student understand where they made mistakes, delivers motivational messages so that the student continues to learn, and points to further academic materials so that the student spends more time studying what they need. This feedback is also personalized for the student's profile, as at the beginning of the course, the student fills out a form with basic information, for example, whether or not they already have programming experience and their major. The plugin was developed for the Moodle platform, a free and open-source virtual learning environment with CodeRunner, an automated grader for multi-programming languages. The AsPin plugin was developed for the CS1 course in Python language by students and instructors from the Department of Computer Science at the University of Brasilia. The plugin presented in this paper was evaluated by undergraduate students taking different majors, which has a high student acceptance rate. This paper presents the applied methodology, the development of the plugin, as well the performance evaluations and perceptions of students from the University of Brasilia in two academic semesters in 2022. Maristela Holanda, Lucas R. F. De Miranda, Fernanda Macedo, João Lucas Yamin, Camilo C. Dorea, Christiana Chamon, Dilma Da Silva |
FIE | 1 |
| 2023 | Learning to Teach: A Guide to Using Learning Theories in Computer Science EducationabstractThis full paper proposes an innovative framework for teaching computer science topics to undergraduate students in Computer Science and Teacher Training (CSTT) for k12 majors. A Computer Science major is required by k12 teachers in Computer Science. However, there has been a concern in Brazil over the lack of explicit pedagogical methods included in CSTT majors. In this context, our research question is: how to teach a CSTT major, so that knowledge of computer science is learned by learning how to teach it through the application of a learning theory? To address this research question, this paper presents the LLL framework, a practical approach in which undergraduate students design and plan computer science lessons using a learning theory approach with a focus on teaching computer science to schools. The acronym LLL comes from “Learn computer science and Learn how to teach computer science by applying a Learning theory”. The LLL framework was tested in two consecutive courses in a year-long study on the operating systems topic in the Department of Computer Science at the University of Brasilia (UnB). UnB was a pioneer, founding the first CSTT major for basic education in Brazil in 1997. The undergraduate students were asked to plan a lesson using this approach and teach the content to their peers, being assessed by the instructors and their peers. The findings showed that the LLL framework is effective for preparing future k12 teachers to teach this content to school students. By adopting the LLL framework, CSTT undergraduate students (teacher candidates) will be able to develop solid pedagogical skills and provide a more meaningful learning experience for k12 students. This methodology has the potential to improve the quality of computer science teaching in schools, contributing to the training of undergraduate students who will be more prepared and engaged in this field. Edison Ishikawa, Hanniel Fernando Lopes Saldanha, Hugo Hiroshi Silva Tutida, Maria de Fátima Brandão, Maristela Holanda |
FIE | 5 |
| 2023 | Study on Computer Science Undergraduate Students Dropout at the University of BrasiliaabstractDropout is a chronic problem that affects education in Brazil at all levels. When considering higher education, Brazil has one of the highest university dropout rates in public and private institutions. Such a problem is considered a social loss as well as a misuse of resources. This paper aims to identify academic and social performance factors that influence the dropout of undergraduate students from Computer Science major at the University of Brasilia (UnB). The data set consists of 879 observations and 16 variables with aspects of the students enrolled in Computer Science from the first semester of 2014 to the second semester of 2019. The survival analysis methodology was employed, more specifically, the Log-Normal regression model, in which the response variable was the time (number of semesters) until dropout (known as failure time) or the last follow-up of the student (either the completion of the course or the student is still enrolled by 2019/2). The model proved to be robust in the residual analysis and in presenting results consistent with the dropout literature. This work can assist in the development of policies and educational strategies to decrease the number of dropouts in the computer science undergraduate major. Mathews de N. S. Lisboa, Juliana Betini Fachini Gomes, Maristela Holanda, Carla Koike, Maria Teresa Leao Costa |
FIE | 3 |
| 2023 | Early Introduction to Computer Architecture in K-12abstractComputer science and engineering students in college get introduced to high-level language programming (Java, C++, Python) early in their first year and later to computer organization and architecture courses. Most students lack a clear understanding of the architecture of a computer before learning how to write code for the first time. This deficiency is due to the lack of courses focused on computer architecture and organization early in high school. Even though introductory computer science courses are now offered from 6th to 12th grade, in some schools, the curriculum lacks emphasis on the fundamentals of computer architecture. This work presents an educational framework suitable for K-12 and undergraduate college students to learn computer architecture by building custom processors, exploring computer subsystems, and observing how programs are simulated in real-time. David Kebo Houngninou, Maristela Holanda, Dilma Da Silva |
SIGCSE (2) | 2 |
| 2022 | Automatic Feedback in the Teaching of Programming in Undergraduate Courses: a Literature MappingabstractTeaching programming in the early years of undergraduate courses has been a challenge for students, institutions, and professors. In view of this, Learning Management Systems (LMSs) and other teaching platforms have emerged to address some of the difficulties in this process. In this context, the present work intends to answer the following research question (RQ): What does the literature tell us about the use of automatic feedback in teaching programming in undergraduate courses? To answer this question, a literature mapping was conducted based on 119 articles published between 2017 and 2021. The mapping showed that the research area is expanding and has related studies from all over the world. The papers have different origins, and 37 countries are represented in this survey. The main programming languages used are Java, Python, C and C++. Another finding was that it is common practice to develop specific platforms for automatic feedback in programming courses. This paper presents the findings and results obtained. Wanderson Conceição, Maristela Holanda, Fernanda Macedo, Edison Ishikawa, Vanessa Tavares Nunes, Dilma Da Silva |
FIE | 2 |
| 2022 | Visual Analysis of Educational Data: a Case Study of Introductory Programming courses at the University of BrasíliaabstractData visualization aims to graphically represent information from a given application domain. This approach helps in the analysis and understanding of a data set through the mechanisms of interaction and generation of graphical representations, which emphasize the observation of characteristics and patterns. In this way, this technique combined with visual learning analytic enables the detection of the expected and the discovery of the unexpected. Those have been used in numerous areas such as education, where the motivation is to understand and improve the teaching and learning processes. In the literature, data visualization is used within the educational area to predict performance and identify the student profiles, as well as monitoring educational systems in order to improve the quality of teaching. In this area, introductory computing courses stand out for the high number of students who fail or drop out of these courses. On average more than 30% of students, worldwide, drop out of introductory computing courses. At University of Brasília (UnB) this ratio is greater than 50%, which makes it an appropriate scenario for the use of data analysis and visualization techniques, in order to discover patterns related to the scenario and ways to improve the situation. This paper searches and implements the most used visualization algorithms, according to the literature, in order to assist instructors and educational managers to get information about historical and demographic data related to the course. To evaluate the visualizations, the algorithms were applied in a case study of three introductory computing courses at UnB and were evaluated through a questionnaire applied to instructors and educational managers. The results show that the respondents felt more secure when using familiar algorithms, such as pie charts and bar charts. Among the selected visualizations, sankey chart, treemap, and violin chart were the least known by the respondents. Furthermore, the bar chart was the algorithm where the information was identified quickly and correctly most of the time. Luiza Hansen, Maristela Holanda, Vinicius Ruela Pereira Borges, Dilma Da Silva |
FIE | 2 |
| 2022 | Gender Diversity in STEM Graduate Programs at the University of Brasília in BrazilabstractIncreasing gender diversity in STEM graduate programs is a challenge. In Brazil, the National Council for Scientific and Technological Development (CNPq) has classified knowledge into different "broad areas", one of which is Exact and Earth Sciences (EES). This area includes the STEM subjects: Physics, Computer Science, Mathematics, Statistics and Chemistry. These EES areas have a low representation of women. The University of Brasília, one of the top 10 universities in Brazil, has graduate programs (master’s and doctoral degrees) in all these subjects. In this context, this paper has the main research question: What is the level of gender diversity in each EES area at the University of Brasília in master’s and doctoral programs? This research question was analyzed with the indicators of student enrollment, number of graduations, and retention rates in the programs. The data used for analysis were the available Brazilian open public data of graduate programs for 11 years, 2007-2017. The findings include that women are in the minority in the total number of graduates in Computer Science and Physics. Despite the low number of women overall in EES, the Chemistry program stands out with the highest female participation, reaching more women than men at the doctorate level. The program that has the fewest women is Computer Science. This paper presents all the results of this study. Maristela Holanda, Thayanna Klysnney, Aletéia P. F. Araújo, Dilma Da Silva, Roberta B. Oliveira, Carla Koike, Carla Denise Castanho, Juliana Betini Fachini Gomes |
FIE | 1 |
| 2022 | Expanding the cybersecurity pipeline through early exposure in undergraduate programsabstractThe cybersecurity field has exponentially grown in recent history, with little to no general understanding of the requirements for professionals in the area. In 2018, it was estimated that 3.5 million cybersecurity jobs would be unfilled by 2021 globally. Many students associate the cybersecurity field with computer programming and hacking, unaware of the societal importance of this career path and its connection to many majors unrelated to computer science or computer engineering. Previous work in cybersecurity education has focused on tools that aid students in understanding specific concepts. None has taken the approach of clarifying the cybersecurity profession in the form of a short-duration seminar series. With this goal, we propose an introductory seminar and a follow-on optional seminar series. The implementation of such a strategy early in undergraduate education may allow students to connect the cybersecurity career paths to one of their societal benefits. This series is designed to provide the student with direct practical examples and thought-provoking questions for the applications of security in their daily life, emphasizing multidisciplinary aspects of the field. This paper describes our experience with this seminar series at a large public university in the Fall of 2021 and Spring of 2022. We evaluate the effectiveness of the proposed approach by assessing its impact on students’ perceived awareness of cybersecurity careers and their interest in cybersecurity minors. Through this study, we observe that the seminar series resulted in an increase in student confidence regarding their understanding of the cybersecurity profession. The data also revealed an increased interest in the cybersecurity minors offered in our institution. Our implementation of this new seminar-based intervention has been limited to providing it as an extracurricular activity among many, therefore reaching only a small part of the first-year engineering students at the university. Still, the data indicates that such a seminar series is a viable instrument to make students aware of the opportunities in a cybersecurity career, its significant demand for professionals, and its potential for addressing societal problems. We make the seminar materials publicly available to facilitate the adoption of the proposed intervention at other institutions. Dilma Da Silva, Maristela Holanda, Nina Miner |
FIE | 2 |
| 2022 | Analysis of Student Performance and Social-economic Data in Introductory Computer Science Courses at the University of BrasíliaabstractComputer Science 1 (CS1) courses introduce undergraduate students to computational thinking and their first programming language. As in most institutions, CS1 is a challenge for students at the University of Brasilia, one of the top 10 universities in Brazil. In 2012, the Brazilian higher education system changed with an affirmative-action policy to admit more students from the public K-12 system: the “Quota” Law was implemented at all federal public universities. This paper aims to answer two research questions: 1) What knowledge about the positive/negative impact of certain features on the success of a CS1 course can be discovered from mining educational data augmented by social-economic information? 2) Are these features different between quota and non-quota students? The analysis uses social-economic and academic performance data of undergraduate students from 2012 to 2019. Data mining algorithms such as generalized linear model, gradient boosting machine, and random forest were applied to the data. The findings include: (1) the relevance of indicators such as the consumption rate of university-subsidized meals, (2) that gender is not a determining factor in failure/success, and (3) a higher failure rate for quota students in the Computer Engineering and Mechatronics Engineering majors. Rodrigo da Fonseca Silveira, Maristela Holanda, Guilherme Novaes Ramos, Márcio Victorino, Dilma Da Silva |
FIE | 2 |
| 2022 | Analysis of Academic Databases for Literature Review in the Computer Science Education FieldabstractLiterature review is a fundamental part of a research process, and systematic protocols for this activity have been used for a long time, mainly in the field of health. Specifically in the Computer Science Education area, the use of systematic literature review has grown. One of the steps in a systematic literature review (SLR) is the selection of academic databases in which to search for articles. There are several databases with academic documents that may be relevant to SLR, for example: Google Scholar, which indexes different types of documents, such as articles, dissertations, theses, and others; Scopus and Web of Science are large databases that index articles from different conferences and journals. ACM Digital Library and IEEE Xplore are also important sources of information in the field of Computer Education. These tools have different characteristics, some charge a fee, others have only information about the title and authors and do not have access to the full article, others have advanced features, with many filters. In this context, this article presents the following research questions: RQ1) What metadata can be extracted automatically from the databases?; RQ2) What kind of visualization tools are available?; RQ3) Do the documents returned by the databases cover the research topic?; RQ4) Do the databases have papers from the main CSE venues?; and RQ5) How many databases are required to perform a literature review in CSE? To answer these questions we used five academic databases: Google Scholar, Scopus, Web of Science, ACM Digital Library, and IEEE xplore. Regarding the results, Scopus and Web of Science have the best visualization of the documents and a robust query engine, however those academic databases are not free. ACM Digital library, IEEE Xplore, Scopus and Web of Science allow the automatic download of the papers’ metadata (author, title, abstract, affiliation and others). Specifically in the field of Computer Science Education, the ACM Digital Library and the IEEE Xplore have important papers from conferences (SIGCSE and FIE) and journals (ACM Transaction on Education and IEEE Transaction on Education). In this full paper, the results will be presented to help researchers to choose the most appropriate academic databases based on their requirements and available options. Aline S. Oliveira Valente, Maristela Holanda, Ari Melo Mariano, Richard Furuta, Dilma Da Silva |
FIE | 2 |
| 2021 | Source Code Plagiarism Detection in an Educational Context: A Literature MappingabstractDetection of plagiarism in students' source codes in college-level programming courses is an important topic for instructors and institutions that seek to pursue project-based learning while enforcing honor codes and maintaining traditional grade-based skill assessment methods. There are different approaches for plagiarism detection currently being researched. This paper aims to answer the question: What does the literature report on source code plagiarism detection in university settings? To answer that, we used a systematic mapping process of recent literature. We selected 109 papers published between 2015 and 2020 that deal with this subject specifically in an educational context. We found that this research area is currently expanding and being studied worldwide. There were papers from 37 different countries, and the number of publications per year has been increasing since 2017. The most targeted programming languages are Java, C++, C, and Python. The most studied plagiarism detection tools are MOSS, JPlag, SIM, Plaggie, and Sherlock. Our study also identified new methodologies created to tackle this problem, such as the analysis of students' typing patterns or their coding style. We noticed that the proposed solutions are mainly based on static source code analysis instead of following the development process. This paper describes our findings. Rodrigo Aniceto, Maristela Holanda, Carla Denise Castanho, Dilma Da Silva |
FIE | 2 |
| 2021 | Sense of Belonging of Female Undergraduate Students in Introductory Computer Science Courses at University of Brasília in BrazilabstractFull Paper - The field of Computer Science (CS) has been of little interest to women straight out of high school when considering undergraduate majors in Brazil. At the University of Brasília, a top-ten university in Brazil, female undergraduate students account for less than 15% of the students in the Department of Computer Science. According to Stout and Blaney, a sense of intellectual belonging is “the sense that one is believed to be a competent member of the community”. This perception may be especially challenging for members of underrepresented minority groups, such as female undergraduate students in CS majors. In this context, this paper addresses two research questions: i) “How does the intellectual sense of belonging of female students compare to the male students' in introduction to computer science courses?”; ii) Is it similar for female undergraduate students in both CS and non-CS majors?”. We devised a questionnaire for students in the introduction to computer science courses for different majors. We analyzed the responses and, in general, introductory programming courses are challenging for all students, however, female students feel worse about their computing competencies than male ones. Maristela Holanda, Aletéia P. F. Araújo, Dilma Da Silva, George von Borries, Roberta B. Oliveira, Carla Koike, Carla Denise Castanho |
FIE | 1 |
| 2021 | Educational Initiatives to Increase Diversity in CS1 Courses: A Literature Mapping of U.S. effortsabstractFull paper. Introductory programming courses such as Computer Science 1 (CS1) are challenging for undergraduate students. Additional obstacles to attracting and retaining students from groups underrepresented in Computer Science majors (such as women, African-Americans, Latinx/Hispanic, and Native Americans) motivated several CS1 educational initiatives aiming at better support for students from these groups. In the context of broadening participation in computing, our paper aims to answer the following question: What does the literature tell us about educational initiatives in CS1 courses that focus on underrepresented minority groups in the United States? To answer this question, we deployed a systematic literature mapping process covering the last twelve years (2009–2020) using four academic databases: Scopus, Web of Science, IEEExplore, and ACM Digital Library. We found 67 academic documents published in conferences and journals, covering activities such as modifications to lesson plans, implementation of pedagogical changes, assessment of student sentiment, the establishment of learning communities, and mentoring events targeting to increase inclusion in CS1 courses. Maristela Holanda, Keishla D. Ortiz-Lopez, Dilma Da Silva, Richard Furuta |
FIE | 1 |
| 2021 | PyDash - A Framework Based Educational Tool for Adaptive Streaming Video Algorithms StudyabstractFull Paper in the Innovative Practice track - The pandemics caused by the spreading of the COVID 19 virus cornered the educational system worldwide, changing the classroom into remote class activities. This change in our social behavior has directly impacted the volume and shape of the Internet traffic data. A recent study shows 15% to 30% increases in Internet traffic caused, among other reasons, by educational video streaming traffic during few weeks in the 2020 lockdown period in Europe. To give some perspective, network providers usually work with a 30% data traffic increase per year. In 2021, it is expected that almost 82% of all Internet traffic will be video, according to CISCO annual forecast report. This scenario has a tremendous impact on the Internet bandwidth capacity, demanding optimized video streaming solutions, such as adaptive bitrate algorithms (ABR). On the other hand, considering the educational challenges in computer network courses, the core activities must be executed using specialized infrastructure to develop students' capabilities with networking equipment. As these types of equipment are costly to be obtained and forwarded to in-home students or simply e-students, a remote platform capable of reproducing an environment for networking applications is required. This is the scenario where PyDash was built. PyDash is a framework for the development of adaptive streaming video algorithms. It is a learning tool designed to abstract the networking communication details, allowing e-students to focus exclusively on developing and evaluating ABR protocols. This paper presents our practical experience developing and using PyDash as an educational tool for teaching ABR protocols at Computing Networking courses at the Department of Computer Science at the University of Brasilia, Brazil. Last semester, over 120 students, divided into four different undergraduate courses, had their first contact with PyDash. Even though this was their first contact with ABR concepts and the pyDash tool, they were able to perform the design, implementation, validation, and analysis of some state-of-the-art algorithms used by Netflix and Youtube. Marcelo Antonio Marotta, Gustavo C. Souza, Maristela Holanda, Marcos F. Caetano |
FIE | 3 |
| 2021 | C*DynaConf: An Apache Cassandra Auto-tuning Tool for Internet of Things Data
Lucas Benevides Dias, Dennis Sávio Silva, Rafael Timóteo de Sousa Júnior, Maristela Holanda |
IoTBDS | 4 |
| 2021 | Computational resource and cost prediction service for scientific workflows in federated clouds
Michel J. F. Rosa, Célia Ghedini Ralha, Maristela Holanda, Aletéia P. F. Araújo |
Future Gener. Comput. Syst. | 3 |
| 2020 | A terrain classifier to improve the inclusion of objects in the collaborative mapping tool OpenStreetMapabstractIn Geographic Information Systems (GIS) and, more specifically in the case of Volunteered Geographic Information (VGI), users actively participate in the processes of data inclusion, edition and deletion. Thus, the issue of data quality becomes a central topic, since it is essential to ascertain if some dimensions are being successfully achieved, such as accuracy, logical consistency, and completeness of the information within these types of system. In this context, the present paper develops the application QualiOSM, with the purpose of assisting users in the tasks of including and editing objects in the collaborative mapping tool OpenStreetMap (OSM). This paper focuses on the implementation of a terrain classifier, which was developed using the parallelepiped classifier technique and proved to be useful especially in identifying buildings of various shapes, contributing to a significant improvement in data consistency of the OSM platform. Gabriel F. B. de Medeiros, Rodrigo da Fonseca Silveira, Maristela Holanda |
EATIS | 3 |
| 2020 | ROLAP DW transformation proposal for OLAP architecture in NoSQL databaseabstractThis paper aims to present a case study related to the migration of a ROLAP architecture Data Warehouse of the University of Brasilia to an OLAP architecture in the NoSQL DB family of columns. This migration starts from the need to study new paradigms of architectures for decision support systems due to the new reality of the problems generated by the use of Big Data in the present moment at the university. We made two approaches: the first one by migrating to the Cassandra DBMS and the second one by migrating to Apache Hive. State of the art was made from Web of Science and we used the TEMAC methodology. We made a transformation of integrating all dimensions and table of facts into a single table for both Cassandra and Hive and we made a comparison by running two queries from these tables. Besides that, we also made a qualitative analysis, addressing the advantages and disadvantages of each approach. In the end, we concluded that Apache Hive should be the best choice for the University of Brasilia. Rodrigo da Fonseca Silveira, Márcio Victorino, Maristela Holanda |
EATIS | 3 |
| 2020 | A Literature Study of Visual Analysis in an Educational ContextabstractResearch Full Paper - Visual Analytics is an emerging field that enables detection of the expected and discovery of the unexpected. This technique has been used in numerous areas such as education, where the motivation is to understand and improve the teaching and learning processes. This paper presents a systematic literature study of the field of visual analysis in an educational context in order to provide more insights into this research area. Therefore, 128 papers were related to this topic. The majority of them were developed in the United States, although Spain, China, the United Kingdom and Brazil are in the ranking of countries that published the most. According to this literature study, 72% of the found papers were published after 2015, showing that Visual Learning Analytics is an emerging field, also, there are several conferences that have included this topic in their publications. Additional analyses were made, focusing on the discovery of the most used algorithms, which include bar charts, line charts, pie charts, Heatmap, and others; whether data mining is commonly combined with Visual Learning Analytics; which educational level is most analyzed in the literature: high school, undergraduate or graduate; and whether any subject has more approaches as well as whether the initial computing classes have been analyzed. Luiza Hansen, Vinicius Ruela Pereira Borges, Maristela Holanda |
FIE | 3 |
| 2020 | The Intellectual Sense of Belonging and Self-efficacy in the Introduction to Computer Science Courses at University of Brasilia in BrazilabstractResearch Full Paper Most top universities in Brazil are public government institutions and tuition free. However, until recently, access to these institutions has been limited by extremely difficult entrance exams. The high standards at public universities are in contrast to the k-12 educational system, where public schools fail to prepare students for the exams, with only some of the private schools offering adequate preparation. In 2012, the Higher Education System in Brazil changed: the Quota Law was implemented for all 59 federal public government universities. This law reserves 50% of the enrollments for the public high-school students with the best grades in the entrance exams. Also, from this 50% allocation of places for students from the public high-school system, half are allocated to students from low-income families (up to one and a half times the minimum monthly salary), black and indigenous students. In this context, this paper addresses the research question: "How does the intellectual sense of belonging and self-efficacy of the quota students compare to that the non-quota students' taking Introduction to Computer Science courses?" We devised a questionnaire for students enrolled in the first programming course of different majors at a top-10 Brazilian university. This paper presents an analysis of the responses that indicates some differences in self-efficacy perceptions between the students admitted through the quota system and the ones admitted exclusively by their placement in entrance exams. Maristela Holanda, George von Borries, Dilma Da Silva, Camilo C. Dorea, Roberta B. Oliveira, Edison Ishikawa |
FIE | 1 |
| 2020 | What do Female Students in Middle and High Schools Think about Computer Science Majors in Brasilia, Brazil? A Survey in 2011 and 2019abstractResearch Full Paper - Computer Science majors lack gender diversity in Brasília, Brazil. Women are an underrepresented minority group in these majors. At the University of Brasilia, one of the top ten universities in Brazil, female undergraduate students account for less than 15% of the students in the Department of Computer Science. In an effort to understand the lack of interest in Computer Science majors among women, this paper addresses the following research questions: 1)Are female students in high school aware that Computer Science majors are predominantly male? 2)Are families of girls from Brasília supportive of their enrolment in Computer Science majors? 3)Do girls from Brasília think that Computer Science majors need a lot of Math? 4)Do female students in high school think that it is difficult to get a job in the field of Computing, with a good salary, and sufficient leisure time? and 5)Which factors influence a female student's choice of a Computer Science major? We devised a questionnaire and applied it to female students in middle and high school on two occasions, in October 2011 (1391 responses) and in July 2019 (429 responses). This paper presents the analysis of the data from the responses, which indicates that the girls' perceptions of Computing have not changed in those years. Maristela Holanda, Roberto Nunes Mourão, George von Borries, Guilherme Novaes Ramos, Aletéia P. F. Araújo, Maria Emília M. T. Walter |
FIE | 1 |
| 2020 | What does a Literature Survey Reveal about the Initiatives to Attract and Retain Women into Computer Science Majors in Latin America?abstractAmong the many papers describing initiatives to recruit and re- tain women into Computer Science (CS) majors, the vast majority focus on the United States and Europe. This poster addresses the research question: "What does the literature tell us about inter- ventions for women in CS majors in Latin America?". We have analyzed papers indexed by Scopus, Web of Science (WoS), and the Latin America Women in Computing Conference (LAWCC) using a systematic literature review process. We found papers from ten countries covering initiatives at different educational levels to increase the participation of women in Computing majors in Latin America. Maristela Holanda, Dilma Da Silva |
SIGCSE | 1 |
| 2019 | A New Approach for De Bruijn Graph Construction in De Novo Genome AssemblingabstractFragment assembly is a current fundamental problem in bioinformatics. In the absence of a reference genome sequence that could guide the whole process, a de Bruijn Graph data structure has been considered to improve the computational processing. Notably, we need to count on a broad set of k-mers, biological sequences substrings. However, the construction of a de Bruijn Graph has a high computational cost, primarily due to main memory consumption. Some approaches use external memory processing to achieve feasibility. These solutions generate all k-mers with high redundancy, increasing the number of managed data and, consequently, the number of I/O operations. This work proposes a new approach for de Bruijn Graph construction that does not need to generate all k-mers. The solution enables to reduce computational requirements and execution feasibility. Elvismary Molina de Armas, Liester Cruz Castro, Maristela Holanda, Sérgio Lifschitz |
BIBM | 3 |
| 2019 | Big Data Trends in BioinformaticsabstractThe amount of biological data available for both the academic community and industry has increased due to the rise of high throughput omics technologies, biotechnology, and health monitoring. This scenario demands efficient storage and analysis of the massive amount of data involved. This work aimed to investigate the scientific literature to map topics related to the use of Big Data in Bioinformatics. Due to the large number of relevant papers to inspect, we employed a three-step data-driven systematic approach. We performed a literature search and selected works related to the theme in three search bases (Scopus, ACM, and Web of Science). Then, we proceeded to a text mining step to analyze terms frequently used in the documents and prepare the documents to be inspected. Afterward, we performed a topic modeling using the LDA (Latent Dirichlet Allocation) algorithm. Twenty groups of topics were obtained and drawn in twenty trend topics, which show a revealing state-of-the-art of the Big Data application in Bioinformatics. Dennis Sávio Silva, Waldeyr M. C. Silva, Guo RuiZhe, Ana Paula Bernardi, Ari Melo Mariano, Maristela Holanda |
BIBM | 6 |
| 2019 | Data Provenance Management of Bioinformatics Workflows in Federated CloudsabstractIn Bioinformatics, the reproducibility of experiments is a fundamental principle to which the data provenance significantly contributes through the acquisition and management of information about the trajectory of the data in a workflow. In addition to the data provenance, aspects such as program configurations and the entire computational environment must be considered to achieve this goal. Cloud computing can provide computational resources, hiding technical details and providing an accessible and configurable on-demand environment for researchers. Cloud federation enables the broad distribution of services and a flexible combination of computing power. Considering this particular scenario, we propose a management platform that collects the provenance data of Bioinformatics workflows allowing portability between different clouds in federated clouds using the Infrastructure as a Service model. This platform is composed of tools that support running these workflows with data provenance captured in NoSQL databases. The retrospective data provenance is captured according to the PROV-DM norms. The findings indicate that the proposed platform can be used for some of the most commoninly available clouds. Polyane Wercelens, Waldeyr M. C. Silva, Klayton Castro, Aletéia P. F. Araújo, Sérgio Lifschitz, Maristela Holanda |
BIBM | 6 |
| 2019 | Educational Data Mining: Analysis of Drop out of Engineering Majors at the UnB - BrazilabstractThis paper presents an analysis of data about the drop out of undergraduate engineering students at the University of Brasilia(UnB), Brazil. In Brazil, similar to other countries, there is a representative amount of engineering students that enroll in engineering majors, however, they don't get to graduate in those majors. Information about the reason for that phenomenon is important for action on the matter by university decisionmakers. This paper aims to answer the research question: What are the main factors that motivate engineering students to drop out of engineering majors at UnB? We have collected the social and performance data of engineering students from 2009 to 2019. Some of the data can be considered rare in similar studies, like students' distance from home to campus and factors like students' leave of absence requests rather than performance factors. We used three data mining techniques: Generalized Linear Model (GLM), Boosting algorithm (GBM) and Random Forest(RF). The results of the study showed that international students deserve some attention from the university and courses like Physics 1 can be challenging for engineering students. Rodrigo da Fonseca Silveira, Maristela Holanda, Márcio Victorino, Marcelo Ladeira |
ICMLA | 2 |
| 2019 | A Real Data Analysis in an Internet of Things Environment
João Victor Poletti, Lucas M. C. e Martins, Samuel Almeida, Maristela Holanda, Rafael Timóteo de Sousa Júnior |
IoTBDS | 4 |
| 2019 | Enhancing Open Government Data With Data ProvenanceabstractThe Brazilian Government has adhered to the Linked Open Government Data publication policy. Thus, promoting a more transparent and open administration, allowing greater participation of society, the strengthening of democracy, and combating corruption. All of these matters can be affected by how open data is published. Beyond the data itself, data provenance allows aggregate metadata such as when, how, and why the data were created and published. Given this scenario, we consider that the combination of data and its provenance enriches the traceability of the data exposing the methods and agents involved in its creation. This paper presents a technological solution in the context of Linked Open Government Data to enhance the public open government data publishing. It is delivered employing an information architecture that can provide the data provenance of public open government data using the PROV-DM and a graph database. In addition, we also present an implementation of the proposed information architecture for a public open linked data as a case study. Cleyton P. dos Reis, Waldeyr M. C. Silva, Luiz C. B. Martins, Rodrigo Pinheiro, Márcio Victorino, Maristela Holanda |
MEDES | 6 |
| 2019 | Solutions for Data Quality in GIS and VGI: A Systematic Literature Review
Gabriel F. B. de Medeiros, Maristela Holanda |
WorldCIST (1) | 2 |
| 2019 | A NoSQL Solution for Bioinformatics Data Provenance Storage
Ingrid Santana, Waldeyr M. C. Silva, Maristela Holanda |
WorldCIST (1) | 3 |
| 2019 | Designing Graph Databases With GRAPHEDabstractIn recent years, graph database systems have become very popular and been deployed mainly in situations where the relationship between data is significant, such as in social networks. Although they do not require a particular schema design, a data model contributes to their consistency. Designing diagrams is an approach to satisfying this demand for a conceptual data model. While researchers and companies have been developing concepts and notations for graph database modeling, their notations focus on their specific implementations. In this article, the authors propose a diagram to address this lack of a generic and comprehensive notation for graph databases modeling, named GRAPHED (Graph Description Diagram for Graph Databases). The authors verified the effectiveness and compatibility of GRAPHED in two case studies: fraud identification, and a biological network model. Gustavo Cordeiro Galvão Van Erven, Rommel N. Carvalho, Waldeyr M. C. Silva, Sérgio Lifschitz, Harley Vera Olivera, Maristela Holanda |
J. Database Manag. | 6 |
| 2018 | A Hadoop Open Source Backup Solution
Heitor Faria, Rodrigo Otávio Ribeiro Hagstrom, Marco Antonio Sousa Reis, Breno G. S. Costa, Edward de Oliveira Ribeiro, Maristela Holanda, Priscila Solís Barreto, Aletéia P. F. Araújo |
CLOSER | 6 |
| 2018 | NoSQL Database Performance Tuning for IoT Data - Cassandra Case Study
Lucas Benevides Dias, Maristela Holanda, Ruben Cruz Huacarpuma, Rafael Timóteo de Sousa Júnior |
IoTBDS | 2 |
| 2018 | UnBGOLD: UnB government open linked data: semantic enrichment of open data toolabstractIn accordance with current legislation designed to make public management more efficient and transparent, Brazilian Federal agencies have adhered to an open data publication policy, despite the challenge presented by datasets being published collectively rather than in isolation. Aiming to facilitate this process, this article presents the UnBGOLD, which adresses the need to connect the data in order to facilitate the publication of open semantically enhanced data. It is a tool that couples the architecture of open data publishing of the University of Brasilia and makes it possible to transform datasets into linked open data utilizing metadatas and ontologies in RDF formats, aside from making it possible for the data to be published automatically on the CKAN platform. Luiz C. B. Martins, Márcio Victorino, Maristela Holanda, George Ghinea, Tor-Morten Grønli |
MEDES | 3 |
| 2018 | NoSQL2: SQL to NoSQL Databases
Jane Adriana, Maristela Holanda |
WorldCIST (2) | 2 |
| 2018 | GRAPHED: A Graph Description Diagram for Graph Databases
Gustavo Cordeiro Galvão Van Erven, Waldeyr M. C. Silva, Rommel N. Carvalho, Maristela Holanda |
WorldCIST (1) | 4 |
| 2018 | An Evaluation of Data Model for NoSQL Document-Based Databases
Debora G. Reis, Fabio S. Gasparoni, Maristela Holanda, Márcio Victorino, Marcelo Ladeira, Edward de Oliveira Ribeiro |
WorldCIST (1) | 3 |
| 2017 | AProvBio: An architecture for data provenance in bioinformatics workflows using graph databaseabstractMany scientific experiments in Bioinformatics are executed as computational workflows. Frequently, it is necessary to re-run an experiment under the original circumstances in which it was run to recognize and validate it. Data provenance concerns the origin of data. Knowing the data source facilitates the understanding and analysis of the results, by detailing and documenting the history and the paths of the input data, from the beginning to the end of an experiment. Therefore, in this context, data provenance can be applied when experimenting traceability. This document presents AProvBio, an architecture that can perform the data provenance of scientific experiments in bioinformatics automatically, using the provenance data model PROV-DM and in a graph database. The architecture can perform the automatic provenance type prospectively, retrospectively and with user-defined data. Thus, the architecture stores and captures information obtained during the execution of the data generation processes with user-defined data information, such as features and versions of the programs used. A graph model, based on the PROV-DM model, was proposed for storing the data provenance. The PROV-DM can be represented by a graph, it allows for a more natural modelling, as well as expressing queries at a more natural level, and the implementation of efficient algorithms to perform specific operations. Rodrigo F. Almeida, Waldeyr M. C. Silva, Klayton Castro, Maria Emília M. T. Walter, Aletéia P. F. Araújo, Maristela Holanda, Sérgio Lifschitz |
BIBM | 6 |
| 2017 | Data provenance management for bioinformatics workflows using NoSQL database systems in a cloud computing environmentabstractComputer science solutions for molecular biology problems are often presented in the form of workflows. There is a set of activities performed by different processing entities through managed tasks. Knowledge about the data trajectory throughout a given workflow enables reproducibility by data provenance. In order to reproduce an in silico bioinformatics experiment one must consider other aspects besides those steps followed by a workflow. Indeed, the computational settings in which the involved programs run is a requirement for reproducibility. Cloud computing technology may hide the technical details and make it easier for the user to set up such an on-demand environment. NoSQL database systems have also gained popularity, particularly in the cloud. Considering this particular scenario, we have planned and executed a research study about a bioinformatics workflow running in an IaaS cloud computing environment. We have persisted provenance data according to the PROV-DM model, using different types of NoSQL database systems. We present in this paper some preliminary results from our research work, where we have explored the characteristics of several NoSQL database systems to persist provenance data. Fernanda Hondo, Polyane Wercelens, Waldeyr M. C. Silva, Klayton Castro, Ingrid Santana, Maria Emília M. T. Walter, Aletéia P. F. Araújo, Maristela Holanda, Sérgio Lifschitz |
BIBM | 8 |
| 2017 | Cost Optimization on Public Cloud Provider for Big Geospatial Data
João Bachiega Jr., Marco Antonio Sousa Reis, Aletéia P. F. Araújo, Maristela Holanda |
CLOSER | 4 |
| 2017 | Early Prediction of College Attrition Using Data MiningabstractCollege attrition is a chronic problem for institutions of higher education. In Brazilian public universities, attrition also accounts for the significant waste of public resources desperately needed in other sectors of society. Thus, given the severity and persistence of this problem, several studies have been conducted in an attempt to mitigate undergraduate dropout rates. Using H2O software as a data mining tool, our study employed parameter tuning to train 321 of three classification algorithms, and with Deep Learning, it was possible to predict 71.1% of the cases of dropout given these characteristics. With this result, it will be possible to identify the attrition profiles of students and implement corrective measures on initiating their studies. Luiz C. B. Martins, Rommel N. Carvalho, Ricardo S. Carvalho, Márcio Victorino, Maristela Holanda |
ICMLA | 5 |
| 2017 | Proposal of a Brazilian Database Government Open Linked Data: DBgoldbr: Invited PaperabstractThe Brazilian Government has made available on the Web a massive volume of public data. This data may be structured, semi-structured or non-structured in order to turn the administration as transparent as possible. Thus, we notice the great challenge in providing applications capable enough to handle this Big Data environment, and to make information available for decision making. In this environment, data processing is done via new approaches from Information Science and Computer Science areas, by involving Technologies and processes for collecting, representing, storing and disseminating information. This paper presents a conceptual model, the technical architecture and the prototype implementation of a tool DBgoldbr, designed to classify government public information with the help of ontologies, by transforming open data into open linked data. To fulfill the purpose of the solution, we used Soft System Methodology to identify problems, to collect users needs and to design solutions that fit the purpose of specific groups. The DBgoldbr tool was designed to ease up the search for open data made available by many Brazilian Government institutions, so that this data can be reused to support the evaluation and monitoring of social programs, in order to support the design and management of public policies. Márcio Victorino, Maristela Holanda, Edison Ishikawa, Edgard Costa Oliveira, George Ghinea, Sammohan Chhetri |
MEDES | 2 |
| 2017 | Detecting Evidence of Fraud in the Brazilian Government Using Graph Databases
Gustavo Cordeiro Galvão Van Erven, Maristela Holanda, Rommel N. Carvalho |
WorldCIST (2) | 2 |
| 2017 | Educational Data Mining: Discovery Standards of Academic Performance by Students in Public High Schools in the Federal District of Brazil
Eduardo Fernandes, Rommel N. Carvalho, Maristela Holanda, Gustavo Cordeiro Galvão Van Erven |
WorldCIST (1) | 3 |
| 2016 | K-mer Mapping and de Bruijn graphs: The case for velvet fragment assemblyabstractK-mer Mapping, an internal process for many de novo genome fragments assembly methods, constitutes a computational challenge due to its high main memory consumption. We present in this paper a study of indexing methods to deal with this problem, considering plant genome assembling. We propose an ad-hoc I/O cost model to analyze the performance of B+- tree and hashing index structures. We use indexes to detect duplicate k-mers and improve the execution time. An actual RDBMS implementation for experiments with a sugarcane data set shows that one can obtain considerable performance gains while reducing RAM requirements. Elvismary Molina de Armas, Edward Hermann Haeusler, Sérgio Lifschitz, Maristela Holanda, Waldeyr M. C. Silva, Paulo Cavalcanti Gomes Ferreira |
BIBM | 4 |
| 2016 | An evaluation of data replication for bioinformatics workflows on NoSQL systemsabstractMany research projects in bioinformatics may be viewed as scientific workflows. Biologists often run multiple times the same workflow with different parameters in order to refine their data analysis. These executions generate a large volume of files with different formats, which need to be stored for future evaluations. New database models, like NoSQL systems, could be considered to deal with large volumes of data, particularly in distributed systems. This work presents a data replication impact assessment from the execution of scientific workflows for two NoSQL database management systems: Cassandra and MongoDB. Iasmini Lima, Matheus Oliveira, Diego Kieckbusch, Maristela Holanda, Maria Emília M. T. Walter, Aletéia P. F. Araújo, Márcio Victorino, Waldeyr M. C. Silva, Sérgio Lifschitz |
BIBM | 4 |
| 2016 | BioNimbuZ: A federated cloud platform for bioinformatics applicationsabstractChallenges in bioinformatics include tools to treat large-scale processing, mainly due to the large volumes of data generated by high-throughput sequencing machines. Besides, many of these tools are not user friendly, and do not distribute their workloads properly. In federated cloud environments, even though services and resources are shared and available online, the processes of a workflow execution are almost entirely not automated, and the majority of these processes do not efficiently balance their workloads. This paper presents the federated cloud platform, called BioNimbuZ, a hybrid platform designed to execute bioinformatics applications easily and efficiently, with good workload balance. Our tests were performed using a real bioinformatics workflow, with fragments generated by the Illumina sequencer, having achieved good performance in practice. Michel J. F. Rosa, Breno Moura, Guilherme Vergara, Lucas Santos, Edward de Oliveira Ribeiro, Maristela Holanda, Maria Emília M. T. Walter, Aletéia P. F. Araújo |
BIBM | 6 |
| 2016 | 2Path: A terpenoid metabolic network modeled as graph databaseabstractTerpenoids are involved in interactions as signaling for communication intra/inter species, signal molecules to attract pollinating insects, and defense against herbivores and microbes. Due to their chemical composition, many terpenoids possess vast pharmacological applicability in medicine and biotechnology, besides important roles in ecology, industry and commerce. The biosynthesis of terpenes has been widely studied over the years, and it is well known that they can be synthesized from two metabolic pathways: mevalonate pathway (MVA) and non-mevalonate pathway (MEP). On the other hand, genome-scale reconstruction of metabolic networks faces many challenges, including organizational data storage and data modeling, to properly represent the complexity of systems biology. Recent NoSQL database paradigms have introduced new concepts of scalable storage and data queries. Among them graph databases, which are versatile enough to cope with biological data. In this paper, we propose 2Path, a graph database model designed to represent terpenoid metabolic networks, with thousands of secondary metabolism reactions, such that it preserves important terpenoid biosynthesis characteristics. Waldeyr M. C. Silva, Danilo Vilar, Daniel S. Souza, Maria Emília M. T. Walter, Marcelo M. Brigido, Maristela Holanda |
BIBM | 6 |
| 2015 | A study of genomic data provenance in NoSQL document-oriented database systemsabstractThis work considers a scientific experiment as a computational workflow. Provenance models store details of each workflow execution, including produced data, computational tools parameters and their versions, among others. This way, scientists can review details of a particular workflow execution, compare information generated among different executions and plan new ones efficiently. In the bioinformatics domain, particularly in the presence of large volumes of data, persistency of those data generated during the workflow execution is still a research challenge. In this article, we consider a study on provenance data storage for bioinformatics in a document-oriented NoSQL database system. We present data modeling issues and discuss an actual implementation into MongoDB. Valeria Guimarâes, Fernanda Hondo, Rodrigo F. Almeida, Harley Vera Olivera, Maristela Holanda, Aletéia P. F. Araújo, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 5 |
| 2014 | Genomic data persistency on a NoSQL database systemabstractPersistency of genomic data brings many research challenges. Among them, alternatives to the commonly used relational database systems should be evaluated, to deal with large volumes of data, in particular writing operations and specific modelling issues. In this paper, we present a NoSQL approach to treat genomic data (Cassandra database system), and evaluate both persistency and I/O operations in experiments with real data. Our results show that this approach enables to store, access and manage genomic data with performance gains, when compared to relational database systems. Rodrigo Aniceto, Rene Xavier, Maristela Holanda, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 3 |
| 2014 | Storing provenance data of genome project workflows using graph databaseabstractMany scientific experiments are designed as computational workflows in bioinformatics. However, the amount of data generated increases at every phase of each execution, hindering the identification of the source and the transformation of data. Therefore, it has become necessary to create new tools to store data provenance, mainly which resources and parameters were used to generate the results, among other information, to validate and publish the experiment. In this paper, we propose to use graph database to store data provenance using the PROV-DM model of bioinformatics workflows. To validate the model, we developed a simulator that worked as a logbook to capture data provenance. A workflow with real genomic data showed that very little additional data should be stored, which means that our provenance model can be easily included in genome projects. Rodrigo Pinheiro, Bruno Aires, Aletéia P. F. Araújo, Maristela Holanda, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 4 |
| 2014 | A Storage Policy for a Hybrid Federated Cloud platform: A Case Study for BioinformaticsabstractBioinformatics tools require large-scale processing mainly due to very large databases achieving gigabytes of size. In federated cloud environments, although services and resources may be shared, storage is particularly difficult, due to distinct computational capabilities and data management policies of several separated clouds. In this work, we propose a storage policy for BioNimbuZ, a hybrid federated cloud platform designed to execute bioinformatics applications. Our storage policy, BioClouZ, aims to perform efficient choices to distribute and replicate files to the best available cloud resources in the federation in order to reduce computational time. BioClouZ uses four parameters - latency, uptime, free size and cost, weighted (according to ad hoc tests) to model their influences to data storage and recovery. Experiments were performed with real biological data executing a commonly used tool to map short reads in a reference genome in BioNimbuZ, composed of clouds executing in Amazon EC2, Azure and University of Brasilia. The results showed that, when compared to the greedy algorithm first used in BioNimbuZ, the BioClouZ policy significantly improved the total execution time due to more efficient choices of the clouds to store the files. Other bioinformatics applications can be used with BioClouZ in BioNimbuZ as well, since the platform was designed independently from particular tools and databases. Deric Lima, Breno Moura, Gabriel S. S. de Oliveira, Edward de Oliveira Ribeiro, Aletéia P. F. Araújo, Maristela Holanda, Roberto C. Togawa, Maria Emília M. T. Walter |
CCGRID | 6 |
| 2013 | ACOsched: A scheduling algorithm in a federated cloud infrastructure for bioinformatics applicationsabstractTask scheduling in a federated cloud environment is a complex problem since there are several cloud providers presenting distinct memory and storage capacities that should be addressed. This article focus on the task scheduling problem in BioNimbuZ, a federated cloud infrastructure for executing bioinformatics applications, which was previously proposed by our group. We present a scheduling algorithm based on Load Balancing Ant Colony (LBACO), called ACOsched, to perform efficient distribution of tasks by finding the best cloud in the federation to execute these tasks. We developed experiments using real biological data, executing the Bowtie mapping tool on one instance of BioNimbuZ, composed by two cloud providers, Amazon EC2 and a bioinformatics laboratory at the University of Brasilia/Brazil. The obtained results show that ACOsched led to a significant improvement in the makespan time of Bowtie executing in BioNimbuZ, when compared to the simple round robin algorithm called DynamicAHP, previously developed in this federated cloud infrastrucutre. Gabriel S. S. de Oliveira, Edward de Oliveira Ribeiro, Diogo A. Ferreira, Aletéia P. F. Araújo, Maristela Holanda, Maria Emília M. T. Walter |
BIBM | 5 |
| 2013 | Automatic capture of provenance data in genome project workflowsabstractMany scientific experiments are designed as computational workflows in the bioinformatics domain, which facilitates implementation and analysis. However, the amount of data generated increases at every phase of each execution, hindering the identification of the source and the data transformation. Therefore, it has become necessary to create new tools to verify automatically which resources and parameters were used to generate the results, among other information to validate and publish the experiment. This functionality of automatically capturing data provenance has been receiving attention in the scientific community, primarily with regard to bioinformatics projects, due the fact that the same workflow is executed several times with different parameters and versions of the tools. In this paper, we propose to use relational schema to automatically store data provenance using the PROV-DM model for workflows in bioinformatics projects. Rodrigo Pinheiro, Maristela Holanda, Aletéia P. F. Araújo, Maria Emília M. T. Walter, Sérgio Lifschitz |
BIBM | 2 |
| 2013 | Provenance in bioinformatics workflowsabstractIn this work, we used the PROV-DM model to manage data provenance in workflows of genome projects. This provenance model allows the storage of details of one workflow execution, e.g., raw and produced data and computational tools, their versions and parameters. Using this model, biologists can access details of one particular execution of a workflow, compare results produced by different executions, and plan new experiments more efficiently. In addition to this, a provenance simulator was created, which facilitates the inclusion of provenance data of one genome project workflow execution. Finally, we discuss one case study, which aims to identify genes involved in specific metabolic pathways of Bacillus cereus, as well as to compare this isolate with other phylogenetic related bacteria from the Bacillus group. B. cereus is an extremophilic bacteria, collected in warm water in the Midwestern Region of Brazil, its DNA samples having been sequenced with an NGS machine. Renato de Paula, Maristela Holanda, Luciana S. A. Gomes, Sérgio Lifschitz, Maria Emília M. T. Walter |
BMC Bioinform. | 2 |
| 2012 | Task Scheduling in a Federated Cloud Infrastructure for Bioinformatics Applications
C. A. L. Borges, Hugo Saldanha, Edward de Oliveira Ribeiro, Maristela Holanda, Aletéia P. F. Araújo, Maria Emília M. T. Walter |
CLOSER | 4 |
| 2011 | A Cloud Architecture for Bioinformatics Workflows
Hugo Saldanha, Edward de Oliveira Ribeiro, Maristela Holanda, Aletéia P. F. Araújo, Genaína Nunes Rodrigues, Maria Emília M. T. Walter, João Carlos Setubal, Alberto M. R. Dávila |
CLOSER | 3 |
| 2008 | A self-adaptable scheduler for synchronizing transactions in dynamically configurable environments
Maristela Holanda, Angelo Brayner, Sergio Fialho |
Data Knowl. Eng. | 1 |