Leonardo Murta 0001

dblp:25/3645 · also Leonardo G. P. Murta, Leonardo Gresta Paulino Murta · DBLP profile ↗
← Back
57ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-5173-1247ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 42 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 1 since 2021Databases, data management, data science and information retrieval · 5Systems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 3Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Towards a Feasible Evaluation Function for Search-Based Merge Conflict Resolution
abstract
Resolving merge conflicts manually is a tedious and complex task. While automated approaches exist, many challenges persist. One promising yet underexplored solution is the use of search-based optimization algorithms, which require an evaluation function to measure the quality of intermediary solutions. However, using code compilation and test execution for this purpose is computationally expensive. This study investigates the relationship between conflict resolutions and conflicting content to identify a metric for guiding search-based optimization techniques. We analyzed 9,998 conflict chunks from 1,062 open source projects, focusing on the similarity of resolutions to their parents and the correlation between randomly generated candidates, parent versions, and the resolution. Our findings reveal that conflict resolutions are, on average, 70% similar to both parents. A strong median correlation ( \(\rho=0.791\) ) exists between candidate-parent and candidate-resolution similarities when aggregating parent similarities with the mean function. Based on these findings, we propose and evaluate SBCR, a Search-Based Conflict Resolution approach that uses parent similarity as a guiding function. We found that the resolution candidates generated by SBCR have a median of 86.5% similarity to the expected resolution, achieving 100% of similarity in 25.2% of the conflicts.
Heleno de S. Campos Junior, Gleiph Ghiotto, Márcio de Oliveira Barros, André van der Hoek, Leonardo Murta 0001
ACM Trans. Softw. Eng. Methodol.5
2025 Analyzing the adoption of database management systems throughout the history of open source projects
Camila A. Paiva, Raquel Maximino, Frederico Paiva, Rafael Accetta Vieira, Nicole Espanha, João Felipe Pimentel, Igor Scaliante Wiese, Marco Aurélio Gerosa, Igor Steinmacher, Leonardo Murta 0001, Vanessa Braganholo
Empir. Softw. Eng.10
2024 Prov-Dominoes: An approach for knowledge discovery from provenance data
Victor Alencar, Troy C. Kohwalter, Vanessa Braganholo, Jose Ricardo da Silva Jr., Leonardo Murta 0001
Expert Syst. Appl.5
2023 Do code refactorings influence the merge effort?
abstract
In collaborative software development, multiple contributors frequently change the source code in parallel to implement new features, fix bugs, refactor existing code, and make other changes. These simultaneous changes need to be merged into the same version of the source code. However, the merge operation can fail, and developer intervention is required to resolve the conflicts. Studies in the literature show that 10 to 20 percent of all merge attempts result in conflicts, which require the manual developer's intervention to complete the process. In this paper, we concern about a specific type of change that affects the structure of the source code and has the potential to increase the merge effort: code refactorings. We analyze the relationship between the occurrence of refactorings and the merge effort. To do so, we applied a data mining technique called association rule extraction to find patterns of behavior that allow us to analyze the influence of refactorings on the merge effort. Our experiments extracted association rules from 40,248 merge commits that occurred in 28 popular open-source projects. The results indicate that: (i) the occurrence of refactorings increases the chances of having merge effort; (ii) the more refactorings, the greater the chances of effort; (iii) the more refactorings, the greater the effort; and (iv) parallel refactorings increase even more the chances of having effort, as well as the intensity of it. The results obtained may suggest behavioral changes in the way refactorings are implemented by developer teams. In addition, they can indicate possible ways to improve tools that support code merging and those that recommend refactorings, considering the number of refactorings and merge effort attributes.
Vânia de Oliveira Neves, Alexandre Plastino 0001, Ana Carla Bibiano, Alessandro F. Garcia 0001, Leonardo Murta 0001
ICSE6
2023 On the assignment of commits to releases
Felipe Curty, Leonardo Murta 0001
Empir. Softw. Eng.2
2023 Towards accurate recommendations of merge conflicts resolution strategies
Paulo Elias, Heleno de S. Campos Junior, Eduardo S. Ogasawara, Leonardo Murta 0001
Inf. Softw. Technol.4
2022 How the adoption of feature toggles correlates with branch merges and defects in open-source projects?
abstract
Abstract Context Branching has been widely adopted in version control to enable collaborative software development. However, the isolation caused by branches may impose challenges on the upcoming merging process. Recently, companies like Google, Microsoft, Facebook, and Spotify, among others, have adopted trunk‐based development together with feature toggles. This strategy enables collaboration without the need for isolation through branches, potentially reducing the merging challenges. However, the literature lacks evidence about the benefits and limitations of feature toggles to collaborative software development. Objective/method In this article, we study the effects of applying feature toggles on 949 open‐source projects written in six different programming languages. We first identified the moment in which each project adopted a feature toggles framework. Then, we observed whether the adoption implied significant changes in the frequency or complexity of branch merges and the number of defects, and the average time to fix them. Finally, we compared the obtained results with results obtained from a set of control projects that do not use feature toggles frameworks. Results/conclusion We could observe a reduction in the average merge effort and an increase in the average total time needed to fix defects after adopting feature toggles frameworks. However, we could not confirm that this increase was influenced by the use of feature toggles.
Eduardo Smil Prutchi, Heleno de S. Campos Junior, Leonardo Murta 0001
Softw. Pract. Exp.3
2022 Dominoes: An Interactive Exploratory Data Analysis Tool for Software Relationships
abstract
Project comprehension questions, such as “which modified artifacts can affect my work?” and “how can I identify the developers who should be assigned to a given task?” are difficult to answer, require an analysis of the project and its data, are context specific, and cannot always be pre-defined. Current research approaches are restricted to post hoc analyses over software repositories. Very few interactive exploratory tools exist since the large amount of data that need to be analyzed prohibits its exploration at interactive rates. Moreover, such analyses typically require the user to create complex scripts or queries to extract the desired information from data. Here we present Dominoes, a tool for interactive data exploration aimed at end users (i.e., project managers or developers). Dominoes allows users to interact with different types and units of data to investigate project relationships and view intermediate results as charts, tables, and graphs. Additionally, it allows users to save the derived data as well as their exploration paths for later use. In a scenario-based evaluation study, participants achieved a success rate of 86 percent in their explorations, with a mean time of 7.25 minutes for answering a set of (project) exploration questions.
Jose Ricardo da Silva Jr., Daniel Prett Campagna, Esteban Walter Gonzalez Clua, Anita Sarma, Leonardo Murta 0001
IEEE Trans. Software Eng.5
2021 Assessing time-based and range-based strategies for commit assignment to releases
abstract
Release is a ubiquitous concept in software development, referring to grouping multiple independent changes into a deliverable piece of software. Mining releases can help developers understand the software evolution at coarse grain, identify which features were delivered or bugs were fixed, and pinpoint who contributed on a given release. A typical initial step of release mining consists of identifying which commits compose a given release. We could find two main strategies used in the literature to perform this task: time-based and range-based. Some release mining works recognize that those strategies are subject to misclassifications but do not quantify the impact of such a threat. This paper analyzed 13,419 releases and 1,414,997 commits from 100 relevant open source projects hosted at GitHub to assess both strategies in terms of precision and recall. We observed that, in general, the range-based strategy has superior results than the time-based strategy. Nevertheless, even when the range-based strategy is in place, some releases still show misclassifications. Thus, our paper also discusses some situations in which each strategy degrades, potentially leading to bias on the mining results if not adequately known and avoided.
Felipe Curty, Leonardo Murta 0001
SANER3
2021 Understanding and improving the quality and reproducibility of Jupyter notebooks
João Felipe Pimentel, Leonardo Murta 0001, Vanessa Braganholo, Juliana Freire
Empir. Softw. Eng.2
2021 Sequential coding patterns: How to use them effectively in code recommendation
Luiz Laerte Nunes da Silva Junior, Troy C. Kohwalter, Alexandre Plastino 0001, Leonardo Murta 0001
Inf. Softw. Technol.4
2021 Predicting the lifetime of pull requests in open-source projects
abstract
Abstract A recent survey using industrial projects has shown that providing an estimate of the lifetime of pull requests to developers helps to speed up their conclusion. Previous work has explored pull request lifetime prediction in open‐source projects using regression techniques but with a broad margin of error. The first objective of our work was to reduce the average error rate of the prediction obtained by the regression techniques so far. We performed experiments with different regression techniques and achieved a significant decrease in the mean error rate. The second objective of our work was to obtain a more effective and useful predictive model that can classify pull requests according to five discrete time intervals. We proposed new predictive attributes for the estimation of the time intervals and employed attribute selection strategies to identify subsets of attributes that could improve the predictive behavior of the classifiers. Our classification approach achieved the best accuracy in all the 20 projects evaluated in comparison with the literature. The average accuracy was of 45.28% to predict pull request lifetime, with an average normalized improvement of 14.68% in relation to the majority class and 6.49% in relation to the state‐of‐the‐art.
Manoel Limeira de Lima Júnior, Daricélio Moreira Soares, Alexandre Plastino 0001, Leonardo Murta 0001
J. Softw. Evol. Process.4
2021 What factors influence the lifetime of pull requests?
abstract
Summary When external contributors want to collaborate with an open‐source project, they fork the repository, make changes, and send a pull request to the core team. However, the lifetime of a pull request, defined by the time interval between its opening and its closing, has a high variation, potentially affecting the contributor engagement. In this context, understanding the root causes of pull request lifetime is important to both the external contributors and the core team. The former can adopt strategies that increase the chances of fast review, while the latter can establish priorities in the reviewing process, alleviating the pending tasks and improving the software quality. In this work, we mined association rules from 97,463 pull requests from 30 projects in order to find characteristics that have affected the pull requests lifetime. In addition, we present a qualitative analysis, helping to understand the patterns discovered from the association rules. The results indicate that: (i) contributions with shorter lifetimes tend to be accepted; (ii) structural characteristics, such as number of commits, changed files, and lines of code, have influence, in an isolated or combined way, on the pull request lifetime; (iii) the files changed and the directories to which they belong can be robust predictors for pull request lifetime; (iv) the profile of external contributors and their social relationships have influence on lifetime; and (v) the number of comments in a pull request, as well as the developer responsible for the review, are important predictors for its lifetime.
Daricélio Moreira Soares, Manoel Limeira de Lima Júnior, Leonardo Murta 0001, Alexandre Plastino 0001
Softw. Pract. Exp.3
2021 Recommending Participants for Collaborative Merge Sessions
abstract
Development of large projects often involves parallel work performed in multiple branches. Eventually, these branches need to be reintegrated through a merge operation. During merge, conflicts may arise and developers need to communicate to reach consensus about the desired resolution. For this reason, including the right developers to a collaborative merge session is fundamental. However, this task can be difficult especially when many different developers have made significant changes on each branch over a large number of files. In this paper, we present TIPMerge, an approach designed to recommend participants for collaborative merge sessions. TIPMerge analyzes the project history and builds a ranked list of developers who are the most appropriate to integrate a pair of branches (Developer Ranking) by considering developers' changes in the branches, in the previous history, and in the dependencies among files across branches. Simply selecting the top developers in such a ranking is easy, but is not effective for collaborative merge sessions as the top developers may have overlapping knowledge. To support collaborative merge, TIPMerge employs optimization techniques to recommend developers with complementary knowledge (Team Recommendation) aiming to maximize joint knowledge coverage. Our results show a mean normalized improvement of 49.5% (median 50.4%) for the joint knowledge coverage with the optimization techniques for assembling teams of three developers for collaborative merge in comparison to choosing the top-3 developers in the ranked list.
Catarina de Souza Costa, Jair Figueiredo, João Felipe Pimentel, Anita Sarma, Leonardo Murta 0001
IEEE Trans. Software Eng.5
2020 Player Behavior Profiling through Provenance Graphs and Representation Learning
abstract
Arguably, player behavior profiling is one of the most relevant tasks of Game Analytics. However, to fulfill the needs of this task, gameplay data should be handled so that the player behavior can be profiled and even understood. Usually, gameplay data is stored as raw log-like files, from which gameplay metrics are computed. However, gameplay metrics have been commonly used as input to classify player behavior with two drawbacks: (1) gameplay metrics are mostly handcrafted and (2) they might not be adequate for fine-grain analysis as they are just computed after key events, such as stage or game completion. In this paper, we present a novel approach for player profiling based on provenance graphs, an alternative to log-like files that model causal relationships between entities in game. Our approach leverages recent advances in deep learning over graph representation of player states and its neighboring contexts, requiring no handcrafted features. We perform clustering on learned nodes representations to profile at a fine-grain the player behavior in provenance data collected from a multiplayer battle game and assess the obtained profiles through statistical analysis and data visualization.
Sidney Araujo Melo, Troy C. Kohwalter, Esteban Walter Gonzalez Clua, Aline Paes, Leonardo Murta 0001
FDG5
2020 Provchastic: Understanding and Predicting Game Events Using Provenance
Troy C. Kohwalter, Leonardo Murta 0001, Esteban Walter Gonzalez Clua
ICEC2
2020 On the performance of hybrid search strategies for systematic literature reviews in software engineering
Érica Mourão, João Felipe Pimentel, Leonardo Murta 0001, Marcos Kalinowski, Emilia Mendes, Claes Wohlin
Inf. Softw. Technol.3
2020 XChange: A semantic diff approach for XML documents
Alessandreia Marta de Oliveira, Troy C. Kohwalter, Marcos Kalinowski, Leonardo Murta 0001, Vanessa Braganholo
Inf. Syst.4
2020 On the Nature of Merge Conflicts: A Study of 2, 731 Open Source Java Projects Hosted by GitHub
abstract
When multiple developers change a software system in parallel, these concurrent changes need to be merged to all appear in the software being developed. Numerous merge techniques have been proposed to support this task, but none of them can fully automate the merge process. Indeed, it has been reported that as much as 10 to 20 percent of all merge attempts result in a merge conflict, meaning that a developer has to manually complete the merge. To date, we have little insight into the nature of these merge conflicts. What do they look like, in detail? How do developers resolve them? Do any patterns exist that might suggest new merge techniques that could reduce the manual effort? This paper contributes an in-depth study of the merge conflicts found in the histories of 2,731 open source Java projects. Seeded by the manual analysis of the histories of five projects, our automated analysis of all 2,731 projects: (1) characterizes the merge conflicts in terms of number of chunks, size, and programming language constructs involved, (2) classifies the manual resolution strategies that developers use to address these merge conflicts, and (3) analyzes the relationships between various characteristics of the merge conflicts and the chosen resolution strategies. Our results give rise to three primary recommendations for future merge techniques, that - when implemented - could on one hand help in automatically resolving certain types of conflicts and on the other hand provide the developer with tool-based assistance to more easily resolve other types of conflicts that cannot be automatically resolved.
Gleiph Ghiotto, Leonardo Murta 0001, Márcio de Oliveira Barros, André van der Hoek
IEEE Trans. Software Eng.2
2019 A large-scale study about quality and reproducibility of jupyter notebooks
abstract
Jupyter Notebooks have been widely adopted by many different communities, both in science and industry. They support the creation of literate programming documents that combine code, text, and execution results with visualizations and all sorts of rich media. The self-documenting aspects and the ability to reproduce results have been touted as significant benefits of notebooks. At the same time, there has been growing criticism that the way notebooks are being used leads to unexpected behavior, encourage poor coding practices, and that their results can be hard to reproduce. To understand good and bad practices used in the development of real notebooks, we studied 1.4 million notebooks from GitHub. We present a detailed analysis of their characteristics that impact reproducibility. We also propose a set of best practices that can improve the rate of reproducibility and discuss open challenges that require further research and development.
João Felipe Pimentel, Leonardo Murta 0001, Vanessa Braganholo, Juliana Freire
MSR2
2019 An Efficient Algorithm for Combining Verification and Validation Methods
Isela Mendoza, Uéverton S. Souza, Marcos Kalinowski, Ruben Interian, Leonardo Murta 0001
SOFSEM5
2018 Filtering irrelevant sequential data out of game session telemetry though similarity collapses
Troy C. Kohwalter, Leonardo Murta 0001, Esteban Walter Gonzalez Clua
Future Gener. Comput. Syst.2
2018 What factors influence the reviewer assignment to pull requests?
Daricélio Moreira Soares, Manoel Limeira de Lima Júnior, Alexandre Plastino 0001, Leonardo Murta 0001
Inf. Softw. Technol.4
2018 An efficient similarity-based approach for comparing XML documents
Alessandreia Marta de Oliveira, Gabriel Tessarolli, Gleiph Ghiotto, Bruno Pinto, Fernando Campello, Matheus Marques, Carlos Roberto Carvalho Oliveira, Igor Rodrigues, Marcos Kalinowski, Uéverton S. Souza, Leonardo Murta 0001, Vanessa Braganholo
Inf. Syst.11
2018 Automatic assignment of integrators to pull requests: The importance of selecting appropriate attributes
Manoel Limeira de Lima Júnior, Daricélio Moreira Soares, Alexandre Plastino 0001, Leonardo Murta 0001
J. Syst. Softw.4
2017 Investigating the Use of a Hybrid Search Strategy for Systematic Reviews
abstract
[Background] Systematic Literature Reviews (SLRs) are one of the important pillars when employing an evidence-based paradigm in Software Engineering. To date most SLRs have been conducted using a search strategy involving several digital libraries. However, significant issues have been reported for digital libraries and applying such search strategy requires substantial effort. On the other hand, snowballing has recently arisen as a potentially more efficient alternative or complementary solution. Nevertheless, it requires a relevant seed set of papers. [Aims] This paper proposes and evaluates a hybrid search strategy combining searching in a specific digital library (Scopus) with backward and forward snowballing. [Method] The proposed hybrid strategy was applied to two previously published SLRs that adopted database searches. We investigate whether it is able to retrieve the same included papers with lower effort in terms of the number of analysed papers. The two selected SLRs relate respectively to elicitation techniques (not confined to Software Engineering (SE)) and to a specific SE topic on cost estimation. [Results] Our results provide preliminary support for the proposed hybrid search strategy as being suitable for SLRs investigating a specific research topic within the SE domain. Furthermore, it helps overcoming existing issues with using digital libraries in SE. [Conclusions] The hybrid search strategy provides competitive results, similar to using several digital libraries. However, further investigation is needed to evaluate the hybrid search strategy.
Érica Mourão, Marcos Kalinowski, Leonardo Murta 0001, Emilia Mendes, Claes Wohlin
ESEM3
2017 Deriving scientific workflows from algebraic experiment lines: A practical approach
Anderson Marinho, Daniel de Oliveira 0001, Eduardo S. Ogasawara, Vítor Silva 0003, Kary A. C. S. Ocaña, Leonardo Murta 0001, Vanessa Braganholo, Marta Mattoso
Future Gener. Comput. Syst.6
2017 noWorkflow: a Tool for Collecting, Analyzing, and Managing Provenance from Python Scripts
abstract
We present noWorkflow, an open-source tool that systematically and transparently collects provenance from Python scripts, including data about the script execution and how the script evolves over time. During the demo, we will show how noWorkflow collects and manages provenance, as well as how it supports the analysis of computational experiments. We will also encourage attendees to use noWorkflow for their own scripts.
João Felipe Pimentel, Leonardo Murta 0001, Vanessa Braganholo, Juliana Freire
Proc. VLDB Endow.2
2017 Managing Provenance of Implicit Data Flows in Scientific Experiments
abstract
Scientific experiments modeled as scientific workflows may create, change, or access data products not explicitly referenced in the workflow specification, leading to implicit data flows. The lack of knowledge about implicit data flows makes the experiments hard to understand and reproduce. In this article, we present ProvMonitor, an approach that identifies the creation, change, or access to data products even within implicit data flows. ProvMonitor links this information with the workflow activity that generated it, allowing for scientists to compare data products within and throughout trials of the same workflow, identifying side effects on data evolution caused by implicit data flows. We evaluated ProvMonitor and observed that it could answer queries for scenarios that demand specific knowledge related to implicit provenance.
Vitor C. Neves, Daniel de Oliveira 0001, Kary A. C. S. Ocaña, Vanessa Braganholo, Leonardo Murta 0001
ACM Trans. Internet Techn.5
2016 TIPMerge: recommending experts for integrating changes across branches
abstract
Parallel development in branches is a common software practice. However, past work has found that integration of changes across branches is not easy, and often leads to failures. Thus far, there has been little work to recommend developers who have the right expertise to perform a branch integration. We propose TIPMerge, a novel tool that recommends developers who are best suited to perform merges, by taking into consideration developers’ past experience in the project, their changes in the branches, and de-pendencies among modified files in the branches. We evaluated TIPMerge on 28 projects, which included up to 15,584 merges with at least two developers, and potentially conflicting changes. On average, 85% of the top-3 recommendations by TIPMerge correctly included the developer who performed the merge. Best (accuracy) results of recommendations were at 98%. Our inter-views with developers of two projects reveal that in cases where the TIPMerge recommendation did not match the actual merge developer, the recommended developer had the expertise to per-form the merge, or was involved in a collaborative merge session.
Catarina Costa, Jair Figueiredo, Leonardo Murta 0001, Anita Sarma
SIGSOFT FSE3
2016 TIPMerge: recommending developers for merging branches
abstract
Development in large projects often involves branches, where changes are performed in parallel and merged periodically. This merge process often combines two independent and long sequences of commits that may have been performed by multiple, different developers. It is nontrivial to identify the right developer to perform the merge, as the developer must have enough understanding of changes in both branches to ensure that the merged changes comply with the objective of both lines of work (branches), which may have been active for several months. We designed and developed TIPMerge, a novel tool that recommends developers who are best suited to perform the merge between two given branches. TIPMerge does so by taking into consideration developers’ past experience in the project, their changes in the branches, and the dependencies among modified files in the branches. In this paper we demonstrate TIPMerge over a real merge case from the Voldemort project.
Catarina Costa, Jair Figueiredo, Anita Sarma, Leonardo Murta 0001
SIGSOFT FSE4
2016 Efficient image-aware version control systems using GPU
abstract
Version control is considered to be a vital component for supporting professional software development. While it has been widely used for textual artifacts, such as source code or documentation, little attention has been given to binary artifacts. This omission can place huge restrictions on projects in the game and media industries as they contain large amounts of binary data, such as images, videos, three-dimensional models, and animations, along with their source code. For these kinds of artifacts, existing strategies such as storing the file as a whole for each revision or saving conventional binary deltas consume significant storage space with duplicate data and, even worse, do not provide any understandable information on which modifications were made. As a response to this problem, this paper introduces a change-set model infrastructure to support version control of image artifacts using a specialized data structure. Additionally, our approach can deal with the maintenance of duplicate nearly identical images through a merge operation. Because of the amount of data that has to be processed, we designed our solution based on a parallel architecture, which permits a massively parallel approach to version control. The paper also compares our approach with some popular open-source version control systems, showing their repository growth in relation to ours as well as the time required to process image artifacts. Finally, we demonstrate that our architecture requires less storage space and runs much faster than current methods. Copyright © 2015 John Wiley & Sons, Ltd.
Jose Ricardo da Silva Jr., Esteban Walter Gonzalez Clua, Leonardo Murta 0001
Softw. Pract. Exp.3
2015 Rejection Factors of Pull Requests Filed by Core Team Developers in Software Projects with High Acceptance Rates
abstract
When developers want to contribute to an opensource project, they fork the repository, make changes, and send a pull request to the core team to incorporate these changes back into the repository. However, some projects enforce this collaboration model even for changes made by core team developers. This potentially enhances the quality of the repository by adding an inspection step before accepting a contribution into the repository. In this context, though less frequently, the contributions may be rejected. The understanding of the factors that lead to the rejection of these internal contributions is crucial for the improvement of the ways core developers collaborate, having a direct impact on the team productivity. In this work we extract association rules from pull request data stored in software repositories in order to find factors that have influence over the decision of rejecting contributions made by core developers. In addition, we present a qualitative analysis of some cases, helping to understand the patterns that arose from the association rules. The results indicate that some key factors increase the changes of having internal contributions rejected: (i) the inexperience with pull requests, (ii) the complexity of contributions, as well as the locality of the artifacts that have been modified, and (iii) the contribution policy of the projects.
Daricélio Moreira Soares, Manoel Limeira de Lima Júnior, Leonardo Murta 0001, Alexandre Plastino 0001
ICMLA3
2015 Niche vs. breadth: Calculating expertise over time through a fine-grained analysis
abstract
Identifying expertise in a project is essential for task allocation, knowledge dissemination, and risk management, among other activities. However, keeping a detailed record of such expertise at class and method levels is cumbersome due to project size, evolution, and team turnover. Existing approaches that automate this task have limitations in terms of the number and granularity of elements that can be analyzed and the analysis timeframe. In this paper, we introduce a novel technique to identify expertise for a given project, package, file, class, or method by considering not only the total number of edits that a developer has made, but also the spread of their changes in an artifact over time, and thereby the breadth of their expertise. We use Dominoes - our GPU-based approach for exploratory repository analysis - for expertise identification over any given granularity and time period with a short processing time. We evaluated our approach through Apache Derby and observed that granularity and time can have significant influence on expertise identification.
Jose Ricardo da Silva Jr., Esteban Walter Gonzalez Clua, Leonardo Murta 0001, Anita Sarma
SANER3
2015 Multi-Perspective Exploratory Analysis of Software Development Data
abstract
In this paper, we present Dominoes, an approach for analyzing software repositories with thousands of artifacts by considering multiple perspectives of the software development data. In order to achieve computational power we model the data and its relationships as matrices, making possible to efficiently process them with a GPUs (Graphics Processing Unit) based architectures. Dominoes can support automated exploration of different relationships among project artifacts, where users have the flexibility to interactively combine and compose them. Our solution organizes data extracted from software repositories into multiple matrices that can be treated as domino pieces (e.g. [commit|method]). The connection of such pieces corresponds to a set of matrices operations, which derive additional domino pieces. These derived domino pieces represent specific project entity relationships (e.g. number of commits in which two methods co-occurred) and can be used for further explorations. As an evaluation of the Dominoes framework we present two exploratory case studies based on Apache Derby. First, we use Dominoes to show how dependencies among artifacts can be derived. Then, we identify expertise of developers by considering the commits that developers make to artifacts. We show that identifying relationships among 34,335 elements along 7,578 commits takes about 0.2 minutes in GPU, while the same processing in CPU takes about 413 minutes. Besides, identifying expertise of developer on a set of 34,335 files and 36 developers takes about 0.1 minute in GPU, whereas in CPU it takes 324 minutes.
Jose Ricardo da Silva Jr., Esteban Walter Gonzalez Clua, Leonardo Murta 0001, Anita Sarma
Int. J. Softw. Eng. Knowl. Eng.3
2015 Software rejuvenation via a multi-agent approach
Heliomar Santos, João Felipe Pimentel, Viviane Torres da Silva, Leonardo Murta 0001
J. Syst. Softw.4
2014 Collaborative Merge in Distributed Software Development: Who Should Participate?
Catarina Costa, José J. C. Figueiredo, Leonardo Murta 0001
SEKE3
2014 Detecting Semantic Equivalence in UML Class Diagrams
Valéria Oliveira Costa, Rodrigo Monteiro, Leonardo Murta 0001
SEKE3
2014 Exploratory Data Analysis of Software Repositories via GPU Processing
Jose Ricardo da Silva Jr., Esteban Walter Gonzalez Clua, Leonardo Murta 0001, Anita Sarma
SEKE3
2014 Characterizing the Problem of Developers' Assignment for Merging Branches
abstract
During the software development process, artifacts are constructed and manipulated by many developers working in parallel. A common practice to manage parallel development is the use of branches in the version control system. Usually, at some point, the merge of these branches may be necessary. This process can combine two independent and eventually long sequences of commits, which may have been performed by different developers. Conflicts resulting from the merge of parallel changes may arise. When these conflicts are not automatically solved by the version control system, the developers in charge of the merge process must act. Normally, the developers' knowledge regarding the changes performed in parallel is usually not taken into consideration when assigning developers to the merge task. With this in mind, the goal of this work is to characterize the problem of developers' assignment for merging branches. To do so, this work analyzed merge profiles of eight software projects and check if the development history is an appropriate source of information for identifying the key participants for collaborative merge. In addition, this work presents a survey on developers about what actions they take when they need to merge branches, and especially when a conflict arises during the merge.
Catarina de Souza Costa, José J. C. Figueiredo, Gleiph Ghiotto, Leonardo Murta 0001
Int. J. Softw. Eng. Knowl. Eng.4
2013 Game Flux Analysis with Provenance
Troy C. Kohwalter, Esteban Walter Gonzalez Clua, Leonardo Murta 0001
Advances in Computer Entertainment3
2013 Version Control in Distributed Software Development: A Systematic Mapping Study
abstract
Along the last decade, many companies started using Distributed Software Development (DSD). The distribution of the software development teams over the globe has become almost a rule in large companies. However, in this context, new problems arise, which mainly involve the physical and temporal distance among the participants. Some studies show that deploying a version control system to alleviate this problem is a big challenge for distributed teams. This paper presents a systematic mapping study about works about version control that focus on DSD. We found 29 studies related to DSD version control, published between 2002 and 2012. Using the systematically extracted data from these works, we present challenges, tools, and other solutions proposed to version control in DSD. These results can support practitioners and researchers to better understand and overcome the challenges related do DSD version control, and devise more effective solutions to improve version control in a distributed setting.
Catarina Costa, Leonardo Murta 0001
ICGSE2
2013 Runtime Monitoring and Auditing of Self-Adaptive Systems (S)
Daniel H. Carmo, Sérgio T. Carvalho, Leonardo Murta 0001, Orlando Loques
SEKE3
2013 Semantic Conflicts Detection in Model-driven Engineering
Valéria Oliveira Costa, João M. B. Oliveira Junior, Leonardo Murta 0001
SEKE3
2012 Optimal Variability Selection in Product Line Engineering
Rafael Pinto Medeiros, Uéverton S. Souza, Fábio Protti, Leonardo Murta 0001
SEKE4
2012 ProvManager: a provenance management system for scientific workflows
abstract
SUMMARY Running scientific workflows in distributed and heterogeneous environments has been a motivating approach for provenance management, which is loosely coupled to the workflow execution engine. This kind of approach is interesting because it allows both storage and access to provenance data in a homogeneous way, even in an environment where different workflow management systems work together. However, current approaches overload scientists with many ad hoc tasks, such as script adaptations and implementations of extra functionalities to provide provenance independence. This paper proposes ProvManager, a provenance management approach that eases the gathering, storage, and analysis of provenance information in a distributed and heterogeneous environment scenario, without putting the burden of adaptations on the scientist. ProvManager leverages the provenance management at the experiment level by integrating different workflow executions from multiple workflow management systems. Copyright © 2011 John Wiley & Sons, Ltd.
Anderson Marinho, Leonardo Murta 0001, Cláudia M. L. Werner, Vanessa Braganholo, Sérgio Manuel Serra da Cruz, Eduardo S. Ogasawara, Marta Mattoso
Concurr. Comput. Pract. Exp.2
2012 To lock, or not to lock: That is the question
João Gustavo Prudêncio, Leonardo Murta 0001, Cláudia M. L. Werner, Rafael da Silva Viterbo de Cepêda
J. Syst. Softw.2
2011 The Doctoral Symposium of the 12th International Conference of Software Reuse
Leonardo Murta 0001
ICSR1
2009 Extending a Software Component Repository to Provide Services
Anderson Marinho, Leonardo Murta 0001, Cláudia M. L. Werner
ICSR2
2009 Neural networks cartridges for data mining on time series
abstract
Neural networks is one of the techniques used for time series analysis. The performance of neural networks is affected by some parameters such as neural network structure and the quality of data preprocessing. These parameters need to be explored in order to obtain an optimal neural network. However, the manual establishment of different neural networks configurations for selecting the best ones may be error-prone and time-consuming. This paper proposes the creation of neural networks cartridges to systematically empower neural network performance by means of data mining activities, which obtain an optimal neural network structure. The experiments conducted in this paper use stock market and exchange rate series, and show that the usage of neural network cartridges can lead to configurations that double the performance of some ad-hoc neural network configuration.
Eduardo S. Ogasawara, Leonardo Murta 0001, Geraldo Zimbrão, Marta Mattoso
IJCNN2
2009 Experiment Line: Software Reuse in Scientific Workflows
Eduardo S. Ogasawara, Carlos Eduardo Paulino Silva, Leonardo Murta 0001, Cláudia M. L. Werner, Marta Mattoso
SSDBM3
2008 Odyssey-MEC: Model Evolution Control in the Context of Model-Driven Architecture
Chessman K. F. Corrêa, Leonardo Murta 0001, Cláudia M. L. Werner
SEKE2
2008 Feature Modeling for Context-Aware Software Product Lines
Paula Fernandes, Cláudia M. L. Werner, Leonardo Murta 0001
SEKE3
2008 Continuous and automated evolution of architecture-to-implementation traceability links
Leonardo Murta 0001, André van der Hoek, Cláudia M. L. Werner
Autom. Softw. Eng.1
2007 Odyssey-SCM: An integrated software configuration management infrastructure for UML models
Leonardo Murta 0001, Hamilton L. R. Oliveira, Cristine R. Dantas, Luiz Gustavo Lopes, Cláudia M. L. Werner
Sci. Comput. Program.1
2006 Odyssey-CCS: A Change Control System Tailored to Software Reuse
Luiz Gustavo Lopes, Leonardo Murta 0001, Cláudia M. L. Werner
ICSR2
2006 ArchTrace: Policy-Based Support for Managing Evolving Architecture-to-Implementation Traceability Links
abstract
Traditional techniques of traceability detection and management are not equipped to handle evolution. This is a problem for the field of software architecture, where it is critical to keep synchronized an evolving conceptual architecture with its realization in an evolving code base. ArchTrace is a new tool that addresses this problem through a policy-based infrastructure for automatically updating traceability links every time an architecture or its code base evolves. ArchTrace is pluggable, allowing developers to choose a set of traceability management policies that best match their situational needs and working styles. We discuss ArchTrace, its conceptual basis, its implementation, and our evaluation of its strengths and weaknesses in a retrospective analysis of data collected from a 20 month period of development of Odyssey, a large-scale software development environment. Results are promising: with respect to the ideal set of traceability links, the policies applied resulted in 95% precision at 89% recall
Leonardo Murta 0001, André van der Hoek, Cláudia M. L. Werner
ASE1