VLDB 2026 Research / reviewers in the wild / expert
Marco Túlio Valente
dblp:o/MarcoTuliodeOliveiraValente · also Marco Túlio de Oliveira Valente
· DBLP profile ↗
97ranked-venue papers
3as first author
25since 2021 · last 2027
0000-0002-8180-7548ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 94 · 2 first-author · 25 since 2021Artificial intelligence and machine learning · 5Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Property-based testing in Python: empirical insightsabstractAbstract Property-Based Testing (PBT) automatically generates test inputs to validate properties of programs, shifting developers’ effort from writing examples to specifying invariants. While the technique has gained popularity in Python through the Hypothesis framework, little is known about how developers adopt and use it in practice. This paper reports on three empirical studies. First, we analyzed 367 PBTs from 244 Python projects, classifying them into nine property categories and quantifying their use of Hypothesis constructs. We found that Test Oracle properties dominate (29.97%), and that PBTs are generally concise (median 14 LOC), relying heavily on built-in strategies (75.20%), but also on external (22.62%) and internal (17.17%) ones. Second, we studied 213 Stack Overflow posts tagged with PBT, revealing that the main challenges developers face concern data generation strategies (36.62%), especially for composite and tabular data (24.36%). Finally, we evaluated Ghostwriter, Hypothesis’s automated test generator, against 203 tests from our dataset; only 18.23% were fully automatable, while most required partial adaptation (30.05%) or were incompatible (51.72%). Together, our findings provide the largest empirical characterization of PBT in Python to date, highlight developers’ difficulties in adopting the technique, and expose limitations of current tool support. Isadora de Oliveira, Arthur Lisboa Corgozinho, Henrique Rocha, Marco Túlio Valente |
Empir. Softw. Eng. | 4 |
| 2025 | Understanding refactorings in Elixir functional language
Lucas Francisco da Matta Vegi, Marco Túlio Valente |
Empir. Softw. Eng. | 2 |
| 2025 | Detection of code smells in react with TypeScript applications
Maykon Nunes, Carla I. M. Bezerra, Fabio Ferreira, Bruno Gois Mateus, Marco Túlio Valente |
Inf. Softw. Technol. | 5 |
| 2025 | NoCodeGPT: A No-Code Interface for Building Web Apps With Language ModelsabstractABSTRACT Background Language models are increasingly used by software developers. However, it remains unclear whether their standard chat‐based interfaces are suitable for software development—especially for users with limited programming experience. Objective This work presents a tool, called NoCodeGPT, that provides a customized interface for language models aimed at enabling the implementation of small web applications without writing code. Methods We first conducted an exploratory study in which three participants used ChatGPT to implement a simple web‐based application. After that, we designed and implemented a customized GPT interface, called NoCodeGPT. To evaluate this new interface, we asked 14 students with limited web development experience to build two small web applications using only prompts. Results The exploratory study showed that general‐purpose chat interfaces like ChatGPT are not user‐friendly for application development. One participant, for instance, was unable to complete any proposed user stories. In contrast, results with NoCodeGPT were encouraging: 9 out of 14 participants completed all user stories, while the remaining five completed at least half. Conclusion The standard GPT interface is not well‐suited for novice web developers. In response, we proposed, designed, and implemented a new interface that offers a more accessible experience for building web applications with language models. Mauricio Monteiro, Bruno Castelo Branco, Samuel Silvestre, Guilherme Avelino 0001, Marco Túlio Valente |
Softw. Pract. Exp. | 5 |
| 2024 | Automatic Library Migration Using Large Language Models: First ResultsabstractDespite being introduced only a few years ago, Large Language Models (LLMs) are already widely used by developers for code generation. However, their application in automating other Software Engineering activities remains largely unexplored. Thus, in this paper, we report the first results of a study in which we are exploring the use of ChatGPT to support API migration tasks, an important problem that demands manual effort and attention from developers. Specifically, in the paper, we share our initial results involving the use of ChatGPT to migrate a client application to use a newer version of SQLAlchemy, an ORM (Object Relational Mapping) library widely used in Python. We evaluate the use of three types of prompts (Zero-Shot, One-Shot, and Chain Of Thoughts) and show that the best results are achieved by the One-Shot prompt, followed by the Chain Of Thoughts. Particularly, with the One-Shot prompt we were able to successfully migrate all columns of our target application and upgrade its code to use new functionalities enabled by SQLAlchemy’s latest version, such as Python’s asyncio and typing modules, while preserving the original code behavior. Aylton Almeida, Laerte Xavier, Marco Túlio Valente |
ESEM | 3 |
| 2024 | Detecting Code Smells using ChatGPT: Initial InsightsabstractThis paper presents initial insights into the effectiveness of ChatGPT in detecting code smells in Java projects. We utilize a large dataset comprising four code smells—Blob, Data Class, Feature Envy, and Long Method—classified into three severity levels. To assess ChatGPT’s proficiency, we employ two different prompts: (i) a generic prompt and (ii) a prompt specifying the smells selected for our research. We evaluate ChatGPT’s abilities using metrics such as precision, recall, and F-measure. Our results reveal that the odds of ChatGPT providing a correct outcome with a specific prompt are 2.54 times higher compared to a generic one. Furthermore, ChatGPT is more effective at detecting smells with critical severity (F-measure = 0.52) than those with minor severity (F-measure = 0.43). Finally, we discuss the implications of our findings and suggest future research directions for leveraging large language models to detect code smells. Luciana Lourdes Silva, Janio Rosa da Silva, João Eduardo Montandon, Marcus Andrade, Marco Túlio Valente |
ESEM | 5 |
| 2024 | Source code expert identification: Models and application
Otávio Cury, Guilherme Avelino 0001, Pedro de Alcântara dos Santos Neto, Marco Túlio Valente, Ricardo Britto 0001 |
Inf. Softw. Technol. | 4 |
| 2024 | Refactoring react-based Web apps
Fabio Ferreira, Hudson Borges, Marco Túlio Valente |
J. Syst. Softw. | 3 |
| 2024 | Towards a catalog of composite refactoringsabstractAbstract Catalogs of refactoring have key importance in software maintenance and evolution, since developers rely on such documents to understand and perform refactoring operations. Furthermore, these catalogs constitute a reference guide for communication between practitioners since they standardize a common refactoring vocabulary. Fowler's book describes the most popular catalog of refactorings, which documents single and well‐known refactoring operations. However, sometimes, refactorings are composite transformations, that is, a sequence of refactorings is performed over a given program element. For example, a sequence of Extract Method operations (a single refactoring) can be performed over the same method, in one or in multiple commits, to simplify its implementation, therefore, leading to a Method Decomposition operation (a composite refactoring). In this paper, we propose and document a catalog with eight composite refactorings. We also implement a set of scripts to mine composite refactorings by preprocessing the results of refactoring detection tools. Using such scripts, we search for composites in a representative refactoring oracle with hundreds of confirmed single refactoring operations. Next, to complement this first study, we also search for composites in the full history of 10 well‐known open‐source projects. We characterize the detected composite refactorings, under dimensions such as size and location. We conclude by addressing the applications and implications of the proposed catalog. Aline Brito, Andre Hora 0001, Marco Túlio Valente |
J. Softw. Evol. Process. | 3 |
| 2023 | How Developers Implement Property-Based TestsabstractProperty-based testing (PBT) is an interesting alternative to example-based testing where the inputs are randomly generated by the testing tool. In PBT, we check properties that always hold for any input. Despite being a promising testing category, to the best of our knowledge, we still lack studies that investigate in the wild how developers are using PBT in practice. In this paper, we report the preliminary results of a study we are conducting on the usage of PBT. We created a dataset of 30 popular Python repositories using Hypothesis (a PBT tool) and selected a random sample of 86 tests. We manually analyzed these tests to understand the most commonly implemented properties and also to reveal the most used features to create them. Arthur Lisboa Corgozinho, Marco Túlio Valente, Henrique Rocha |
ICSME | 2 |
| 2023 | Towards a Catalog of Refactorings for ElixirabstractElixir is an emerging functional programming language that is gaining popularity in the industry. However, to the best of our knowledge, no study has yet presented a specialized catalog of refactoring strategies specifically tailored for this language. Therefore, this paper aims to address this research gap by conducting a systematic literature review to explore whether there are existing refactoring strategies, compatible with Elixir, that have been proposed for other functional languages or languages that served as inspiration for the development of Elixir. Our preliminary results indicate that there are 54 refactoring strategies compatible with Elixir code, thus forming a comprehensive catalog of refactoring techniques tailored specifically for this language. To illustrate the application of these refactoring strategies, we have provided code examples that show the transformations resulting from each refactoring. Additionally, these refactorings have been categorized into three distinct groups based on the specific programming features required for the respective code transformations. Lucas Francisco da Matta Vegi, Marco Túlio Valente |
ICSME | 2 |
| 2023 | Unboxing Default Argument Breaking Changes in Scikit LearnabstractMachine Learning (ML) has revolutionized the field of computer software development, enabling data-based predictions and decision-making across several domains. Following modern software development practices, developers use third-party libraries—e.g., Scikit Learn, TensorFlow, and PyTorch—to integrate ML-based functionalities into their applications. Due to the complexity inherent in ML techniques, the models available in the APIs of these tools often require an extensive list of arguments to be set up. Library maintainers overcome this issue by defining default values for most of these arguments so developers can use ML models in their client applications effortlessly. By relying on these default arguments, the clients inadvertently depend on the value defined in these parameters to keep running as expected. We interpret this problem as a semantical breaking change variant, which we named Default Argument Breaking Change (DABC). In this work, we leverage 77 DABCs in Scikit Learn—a well-known ML library—and investigate how 194K client applications are vulnerable to them. Our results show that 72 DABCs (93%) are responsible for exposing 67,747 clients (35%). We also detected that most DABCs (61, 79%) involve APIs used in ML model training and model evaluation stages. Finally, we discuss the importance of managing DABCs in third-party ML libraries and provide insights for developers to mitigate the potential impact of these changes in their applications. João Eduardo Montandon, Luciana Lourdes Silva, Cristiano Politowski, Ghizlane El-Boussaidi, Marco Túlio Valente |
SCAM | 5 |
| 2023 | Understanding code smells in Elixir functional language
Lucas Francisco da Matta Vegi, Marco Túlio Valente |
Empir. Softw. Eng. | 2 |
| 2023 | Detecting code smells in React-based Web apps
Fabio Ferreira, Marco Túlio Valente |
Inf. Softw. Technol. | 2 |
| 2023 | Snapshot testing in practice: Benefits and drawbacks
Victor Pezzi Gazzinelli Cruz, Henrique Rocha, Marco Túlio Valente |
J. Syst. Softw. | 3 |
| 2022 | Identifying Source Code File ExpertsabstractBackground: In software development, the identification of source code file experts is an important task. Identifying these experts helps to improve software maintenance and evolution activities, such as developing new features, code reviews, and bug fixes. Although some studies have proposed repository-mining techniques to automatically identify source code experts, there are still gaps in this area that can be explored. For example, investigating new variables related to source code knowledge and applying machine learning aiming to improve the performance of techniques to identify source code experts. Aim: The goal of this study is to investigate opportunities to improve the performance of existing techniques to recommend source code files experts. Method: We built an oracle by collecting data from the development history and surveying developers of 113 software projects. Then, we use this oracle to: (i) analyze the correlation between measures extracted from the development history and the developers’ source code knowledge and (ii) investigate the use of machine learning classifiers by evaluating their performance in identifying source code files experts. Results:First Authorship and Recency of Modification are the variables with the highest positive and negative correlations with source code knowledge, respectively. Machine learning classifiers outperformed the linear techniques (F-Measure = 71% to 73%) in the public dataset, but this advantage is not clear in the private dataset, with F-Measure ranging from 55% to 68% for the linear techniques and 58% to 67% for ML techniques. Conclusion: Overall, the linear techniques and the machine learning classifiers achieved similar performance, particularly if we analyze F-Measure. However, machine learning classifiers usually get higher precision while linear techniques obtained the highest recall values. Therefore, the choice of the best technique depends on the user’s tolerance to false positives and false negatives. Otávio Cury, Guilherme Avelino 0001, Pedro de Alcântara dos Santos Neto, Ricardo Britto 0001, Marco Túlio Valente |
ESEM | 5 |
| 2022 | Code smells in Elixir: early results from a grey literature reviewabstractElixir is a new functional programming language whose popularity is rising in the industry. However, there are few works in the literature focused on studying the internal quality of systems implemented in this language. Particularly, to the best of our knowledge, there is currently no catalog of code smells for Elixir. Therefore, in this paper, through a grey literature review, we investigate whether Elixir developers discuss code smells. Our preliminary results indicate that 11 of the 22 traditional code smells cataloged by Fowler and Beck are discussed by Elixir developers. We also propose a list of 18 new smells specific for Elixir systems and investigate whether these smells are currently identified by Credo, a well-known static code analysis tool for Elixir. We conclude that only two traditional code smells and one Elixir-specific code smell are automatically detected by this tool. Thus, these early results represent an opportunity for extending tools such as Credo to detect code smells and then contribute to improving the internal quality of Elixir systems. Lucas Francisco da Matta Vegi, Marco Túlio Valente |
ICPC | 2 |
| 2022 | On the documentation of self-admitted technical debt in issues
Laerte Xavier, João Eduardo Montandon, Fabio Ferreira, Rodrigo Brito, Marco Túlio Valente |
Empir. Softw. Eng. | 5 |
| 2022 | On the (un-)adoption of JavaScript front-end frameworksabstractAbstract JavaScript is characterized by a rich ecosystem of libraries and frameworks. A key element in this ecosystem are frameworks used for implementing the front‐end of web‐based applications, such as Vue and React. However, despite their relevance, we have few works investigating the factors that drive the adoption—and un‐adoption—of front‐end‐based JavaScript frameworks. Therefore, in this article, we first report the results of a survey with 49 developers where we asked them to describe the factors they consider when selecting a front‐end framework. In the second part of the work, we focus on projects that migrate from one framework to another since JavaScript's ecosystem is also very dynamic. Finally, we provide a quantitative characterization of the migration effort and reveal the main barriers faced by the developers during this effort. Although not completely generalizable, our central findings are as follows: (a) popularity and learnability are the key factors that motivate the choice of front‐end frameworks in JavaScript; (b) from the 49 surveyed developers, one out of four have plans to migrate to another framework in the future; (c) the time spent performing the migration is greater than or equal to the time spent using the old framework in all studied projects. We conclude with a list of implications for practitioners, framework developers, tool builders, and researchers. Fabio Ferreira, Hudson Borges, Marco Túlio Valente |
Softw. Pract. Exp. | 3 |
| 2021 | RAID: Tool Support for Refactoring-Aware Code ReviewsabstractCode review is a key development practice that contributes to improve software quality and to foster knowledge sharing among developers. However, code review usually takes time and demands detailed and time-consuming analysis of textual diffs. Particularly, detecting refactorings during code reviews is not a trivial task, since they are not explicitly represented in diffs. For example, a Move Function refactoring is represented by deleted (-) and added lines (+) of code which can be located in different and distant source code files. To tackle this problem, we introduce RAID, a refactoring-aware and intelligent diff tool. Besides proposing an architecture for RAID, we implemented a Chrome browser plug-in that supports our solution. Then, we conducted a field experiment with eight professional developers who used RAID for three months. We concluded that RAID can reduce the cognitive effort required for detecting and reviewing refactorings in textual diff. Besides documenting refactorings in diffs, RAID reduces the number of lines required for reviewing such operations. For example, the median number of lines to be reviewed decreases from 14.5 to 2 lines in the case of move refactorings and from 113 to 55 lines in the case of extractions. Rodrigo Brito, Marco Túlio Valente |
ICPC | 2 |
| 2021 | Characterizing refactoring graphs in Java and JavaScript projects
Aline Brito, Andre Hora 0001, Marco Túlio Valente |
Empir. Softw. Eng. | 3 |
| 2021 | What skills do IT companies look for in new developers? A study with Stack Overflow jobs
João Eduardo Montandon, Cristiano Politowski, Luciana Lourdes Silva, Marco Túlio Valente, Fábio Petrillo, Yann-Gaël Guéhéneuc |
Inf. Softw. Technol. | 4 |
| 2021 | Mining the Technical Roles of GitHub Users
João Eduardo Montandon, Marco Túlio Valente, Luciana Lourdes Silva |
Inf. Softw. Technol. | 2 |
| 2021 | Are game engines software frameworks? A three-perspective study
Cristiano Politowski, Fábio Petrillo, João Eduardo Montandon, Marco Túlio Valente, Yann-Gaël Guéhéneuc |
J. Syst. Softw. | 4 |
| 2021 | RefDiff 2.0: A Multi-Language Refactoring Detection ToolabstractIdentifying refactoring operations in source code changes is valuable to understand software evolution. Therefore, several tools have been proposed to automatically detect refactorings applied in a system by comparing source code between revisions. The availability of such infrastructure has enabled researchers to study refactoring practice in large scale, leading to important advances on refactoring knowledge. However, although a plethora of programming languages are used in practice, the vast majority of existing studies are restricted to the Java language due to limitations of the underlying tools. This fact poses an important threat to external validity. Thus, to overcome such limitation, in this paper we propose RefDiff 2.0, a multi-language refactoring detection tool. Our approach leverages techniques proposed in our previous work and introduces a novel refactoring detection algorithm that relies on the Code Structure Tree (CST), a simple yet powerful representation of the source code that abstracts away the specificities of particular programming languages. Despite its language-agnostic design, our evaluation shows that RefDiff's precision (96 percent) and recall (80 percent) are on par with state-of-the-art refactoring detection approaches specialized in the Java language. Our modular architecture also enables one to seamlessly extend RefDiff to support other languages via a plugin system. As a proof of this, we implemented plugins to support two other popular programming languages: JavaScript and C. Our evaluation in these languages reveals that precision and recall ranges from 88 to 91 percent. With these results, we envision RefDiff as a viable alternative for breaking the single-language barrier in refactoring research and in practical applications of refactoring detection. Danilo Silva 0002, João Paulo da Silva, Gustavo Jansen de Souza Santos, Ricardo Terra, Marco Túlio Valente |
IEEE Trans. Software Eng. | 5 |
| 2020 | REST vs GraphQL: A Controlled ExperimentabstractGraphQL is a novel query language for implementing service-based software architectures. The language is gaining momentum and it is now used by major software companies, such as Facebook and GitHub. However, we still lack empirical evidence on the real gains achieved by GraphQL, particularly in terms of the effort required to implement queries in this language. Therefore, in this paper we describe a controlled experiment with 22 students (10 undergraduate and 12 graduate), who were asked to implement eight queries for accessing a web service, using GraphQL and REST. Our results show that GraphQL requires less effort to implement remote service queries when compared to REST (9 vs 6 minutes, median times). These gains increase when REST queries include more complex endpoints, with several parameters. Interestingly, GraphQL outperforms REST even among more experienced participants (as is the case of graduate students) and among participants with previous experience in REST, but no previous experience in GraphQL. Gleison Brito, Marco Túlio Valente |
ICSA | 2 |
| 2020 | Beyond the Code: Mining Self-Admitted Technical Debt in Issue Tracker SystemsabstractSelf-admitted technical debt (SATD) is a particular case of Technical Debt (TD) where developers explicitly acknowledge their sub-optimal implementation decisions. Previous studies mine SATD by searching for specific TD-related terms in source code comments. By contrast, in this paper we argue that developers can admit technical debt by other means, e.g., by creating issues in tracking systems and labelling them as referring to TD. We refer to this type of SATD as issue-based SATD or just SATD-I. We study a sample of 286 SATD-I instances collected from five open source projects, including Microsoft Visual Studio and GitLab Community Edition. We show that only 29% of the studied SATD-I instances can be tracked to source code comments. We also show that SATD-I issues take more time to be closed, compared to other issues, although they are not more complex in terms of code churn. Besides, in 45% of the studied issues TD was introduced to ship earlier, and in almost 60% it refers to DESIGN flaws. Finally, we report that most developers pay SATD-I to reduce its costs or interests (66%). Our findings suggest that there is space for designing novel tools to support technical debt management, particularly tools that encourage developers to create and label issues containing TD concerns. Laerte Xavier, Fabio Ferreira, Rodrigo Brito, Marco Túlio Valente |
MSR | 4 |
| 2020 | Refactoring Graphs: Assessing Refactoring over TimeabstractRefactoring is an essential activity during software evolution. Frequently, practitioners rely on such transformations to improve source code maintainability and quality. As a consequence, this process may produce new source code entities or change the structure of existing ones. Sometimes, the transformations are atomic, i.e., performed in a single commit. In other cases, they generate sequences of modifications performed over time. To study and reason about refactorings over time, in this paper, we propose a novel concept called refactoring graphs and provide an algorithm to build such graphs. Then, we investigate the history of 10 popular open-source Java-based projects. After eliminating trivial graphs, we characterize a large sample of 1,150 refactoring graphs, providing quantitative data on their size, commits, age, refactoring composition, and developers. We conclude by discussing applications and implications of refactoring graphs, for example, to improve code comprehension, detect refactoring patterns, and support software evolution studies. Aline Brito, Andre Hora 0001, Marco Túlio Valente |
SANER | 3 |
| 2020 | You broke my code: understanding the motivations for breaking changes in APIs
Aline Brito, Marco Túlio Valente, Laerte Xavier, Andre Hora 0001 |
Empir. Softw. Eng. | 2 |
| 2020 | Is this GitHub project maintained? Measuring the level of maintenance activity of open-source projects
Jailton Coelho, Marco Túlio Valente, Luciano Milen, Luciana Lourdes Silva |
Inf. Softw. Technol. | 2 |
| 2020 | Prioritizing versions for performance regression testing: The Pharo case
Juan Pablo Sandoval Alcocer, Alexandre Bergel, Marco Túlio Valente |
Sci. Comput. Program. | 3 |
| 2019 | On the abandonment and survival of open source projects: An empirical investigationabstractBackground: Evolution of open source projects frequently depends on a small number of core developers. The loss of such core developers might be detrimental for projects and even threaten their entire continuation. However, it is possible that new core developers assume the project maintenance and allow the project to survive. Aims: The objective of this paper is to provide empirical evidence on: 1) the frequency of project abandonment and survival, 2) the differences between abandoned and surviving projects, and 3) the motivation and difficulties faced when assuming an abandoned project. Method: We adopt a mixed-methods approach to investigate project abandonment and survival. We carefully select 1,932 popular GitHub projects and recover the abandoned and surviving projects, and conduct a survey with developers that have been instrumental in the survival of the projects. Results: We found that 315 projects (16%) were abandoned and 128 of these projects (41%) survived because of new core developers who assumed the project development. The survey indicates that (i) in most cases the new maintainers were aware of the project abandonment risks when they started to contribute; (ii) their own usage of the systems is the main motivation to contribute to such projects; (iii) human and social factors played a key role when making these contributions; and (iv) lack of time and the difficulty to obtain push access to the repositories are the main barriers faced by them. Conclusions: Project abandonment is a reality even in large open source projects and our work enables a better understanding of such risks, as well as highlights ways in avoiding them. Guilherme Avelino 0001, Eleni Constantinou, Marco Túlio Valente, Alexander Serebrenik |
ESEM | 3 |
| 2019 | Identifying experts in software libraries and frameworks among GitHub usersabstractSoftware development increasingly depends on libraries and frameworks to increase productivity and reduce time-to-market. Despite this fact, we still lack techniques to assess developers expertise in widely popular libraries and frameworks. In this paper, we evaluate the performance of unsupervised (based on clustering) and supervised machine learning classifiers (Random Forest and SVM) to identify experts in three popular JavaScript libraries: facebook/react, mongodb/node-mongodb, and socketio/socket.io. First, we collect 13 features about developers activity on GitHub projects, including commits on source code files that depend on these libraries. We also build a ground truth including the expertise of 575 developers on the studied libraries, as self-reported by them in a survey. Based on our findings, we document the challenges of using machine learning classifiers to predict expertise in software libraries, using features extracted from GitHub. Then, we propose a method to identify library experts based on clustering feature data from GitHub; by triangulating the results of this method with information available on Linkedin profiles, we show that it is able to recommend dozens of GitHub users with evidences of being experts in the studied JavaScript libraries. We also provide a public dataset with the expertise of 575 developers on the studied libraries. João Eduardo Montandon, Luciana Lourdes Silva, Marco Túlio Valente |
MSR | 3 |
| 2019 | GoCity: Code City for GoabstractGo is a statically typed and compiled language, which has been widely used to develop robust and popular projects. As other systems, these projects change over time. Developers commonly modify the source code to improve the quality or to implement new features. In this context, they use tools and approaches to support software maintenance tasks. However, there is a lack of tools to support Go developers in this process. To address these challenges, we introduce GoCity, a web-based implementation of the CodeCity program visualization metaphor for Go. The tool extracts source code metrics to create a software visualization in a automated way. We also report usage scenarios of GoCity involving three popular Go projects. Finally, we report the feedback of 12 developers about the GoCity of their projects. Rodrigo Brito, Aline Brito, Gleison Brito, Marco Túlio Valente |
SANER | 4 |
| 2019 | Migrating to GraphQL: A Practical AssessmentabstractGraphQL is a novel query language proposed by Facebook to implement Web-based APIs. In this paper, we present a practical study on migrating API clients to this new technology. First, we conduct a grey literature review to gain an in-depth understanding on the benefits and key characteristics normally associated to GraphQL by practitioners. After that, we assess such benefits in practice, by migrating seven systems to use GraphQL, instead of standard REST-based APIs. As our key result, we show that GraphQL can reduce the size of the JSON documents returned by REST APIs in 94% (in number of fields) and in 99% (in number of bytes), both median results. Gleison Brito, Thaís Mombach, Marco Túlio Valente |
SANER | 3 |
| 2019 | Co-change patterns: A large scale empirical study
Luciana Lourdes Silva, Marco Túlio Valente, Marcelo de Almeida Maia |
J. Syst. Softw. | 2 |
| 2019 | Measuring and analyzing code authorship in 1 + 118 open source projects
Guilherme Avelino 0001, Leonardo Teixeira Passos, Andre Hora 0001, Marco Túlio Valente |
Sci. Comput. Program. | 4 |
| 2019 | Algorithms for estimating truck factors: a comparative study
Mívian M. Ferreira, Thaís Mombach, Marco Túlio Valente, Kecia Aline M. Ferreira |
Softw. Qual. J. | 3 |
| 2018 | Identifying unmaintained projects in githubabstractBackground: Open source software has an increasing importance in modern software development. However, there is also a growing concern on the sustainability of such projects, which are usually managed by a small number of developers, frequently working as volunteers. Aims: In this paper, we propose an approach to identify GitHub projects that are not actively maintained. Our goal is to alert users about the risks of using these projects and possibly motivate other developers to assume the maintenance of the projects. Method: We train machine learning models to identify unmaintained or sparsely maintained projects, based on a set of features about project activity (commits, forks, issues, etc). We empirically validate the model with the best performance with the principal developers of 129 GitHub projects. Results: The proposed machine learning approach has a precision of 80%, based on the feedback of real open source developers; and a recall of 96%. We also show that our approach can be used to assess the risks of projects becoming unmaintained. Conclusions: The model proposed in this paper can be used by open source users and developers to identify GitHub projects that are not actively maintained anymore. Jailton Coelho, Marco Túlio Valente, Luciana Lourdes Silva, Emad Shihab |
ESEM | 2 |
| 2018 | Assessing the threat of untracked changes in software evolutionabstractWhile refactoring is extensively performed by practitioners, many Mining Software Repositories (MSR) approaches do not detect nor keep track of refactorings when performing source code evolution analysis. In the best case, keeping track of refactorings could be unnecessary work; in the worst case, these untracked changes could significantly affect the performance of MSR approaches. Since the extent of the threat is unknown, the goal of this paper is to assess whether it is significant. Based on an extensive empirical study, we answer positively: we found that between 10 and 21% of changes at the method level in 15 large Java systems are untracked. This results in a large proportion (25%) of entities that may have their histories split by these changes, and a measurable effect on at least two MSR approaches. We conclude that handling untracked changes should be systematically considered by MSR studies. Andre Hora 0001, Danilo Silva 0002, Marco Túlio Valente, Romain Robbes |
ICSE | 3 |
| 2018 | Feature location benchmark with argoUML SPLabstractFeature location is a traceability recovery activity to identify the implementation elements associated to a characteristic of a system. Besides its relevance for software maintenance of a single system, feature location in a collection of systems received a lot of attention as a first step to re-engineer system variants (created through clone-and-own) into a Software Product Line (SPL). In this context, the objective is to unambiguously identify the boundaries of a feature inside a family of systems to later create reusable assets from these implementation elements. Among all the case studies in the SPL literature, variants derived from ArgoUML SPL stands out as the most used one. However, the use of different settings, or the omission of relevant information (e.g., the exact configurations of the variants or the way the metrics are calculated), makes it difficult to reproduce or benchmark the different feature location techniques even if the same ArgoUML SPL is used. With the objective to foster the research area on feature location, we provide a set of common scenarios using ArgoUML SPL and a set of utils to obtain metrics based on the results of existing and novel feature location techniques. Jabier Martinez, Nicolas Ordoñez, Xhevahire Tërnava, Tewfik Ziadi, Jairo Aponte, Eduardo Figueiredo 0001, Marco Túlio Valente |
SPLC | 7 |
| 2018 | Why and how Java developers break APIsabstractModern software development depends on APIs to reuse code and increase productivity. As most software systems, these libraries and frameworks also evolve, which may break existing clients. However, the main reasons to introduce breaking changes in APIs are unclear. Therefore, in this paper, we report the results of an almost 4-month long field study with the developers of 400 popular Java libraries and frameworks. We configured an infrastructure to observe all changes in these libraries and to detect breaking changes shortly after their introduction in the code. After identifying breaking changes, we asked the developers to explain the reasons behind their decision to change the APIs. During the study, we identified 59 breaking changes, confirmed by the developers of 19 projects. By analyzing the developers' answers, we report that breaking changes are mostly motivated by the need to implement new features, by the desire to make the APIs simpler and with fewer elements, and to improve maintainability. We conclude by providing suggestions to language designers, tool builders, software engineering researchers and API developers. Aline Brito, Laerte Xavier, Andre Hora 0001, Marco Túlio Valente |
SANER | 4 |
| 2018 | APIDiff: Detecting API breaking changesabstractLibraries are commonly used to increase productivity. As most software systems, they evolve over time and changes are required. However, this process may involve breaking compatibility with previous versions, leading clients to fail. In this context, it is important that libraries creators and clients frequently assess API stability in order to better support their maintenance practices. In this paper, we introduce APIDIFF, a tool to identify API breaking and non-breaking changes between two versions of a Java library. The tool detects changes on three API elements: types, methods, and fields. We also report usage scenarios of APIDIFF with four real-world Java libraries. Aline Brito, Laerte Xavier, Andre Hora 0001, Marco Túlio Valente |
SANER | 4 |
| 2018 | What's in a GitHub Star? Understanding Repository Starring Practices in a Social Coding Platform
Hudson Borges, Marco Túlio Valente |
J. Syst. Softw. | 2 |
| 2018 | On the use of replacement messages in API deprecation: An empirical study
Gleison Brito, Andre Hora 0001, Marco Túlio Valente, Romain Robbes |
J. Syst. Softw. | 3 |
| 2018 | JMove: A novel heuristic and tool to detect move method refactoring opportunities
Ricardo Terra, Marco Túlio Valente, Sergio Miranda, Vitor Sales |
J. Syst. Softw. | 2 |
| 2018 | How do developers react to API evolution? A large-scale empirical study
Andre Hora 0001, Romain Robbes, Marco Túlio Valente, Nicolas Anquetil, Anne Etien, Stéphane Ducasse |
Softw. Qual. J. | 3 |
| 2017 | Refactoring Legacy JavaScript Code to Use Classes: The Good, The Bad and The Ugly
Leonardo Humberto Silva, Marco Túlio Valente, Alexandre Bergel |
ICSR | 2 |
| 2017 | A comparison of three algorithms for computing truck factorsabstractTruck Factor (also known as Bus Factor or Lottery Number) is the minimal number of developers that have to be hit by a truck (or leave) before a project is incapacitated. Therefore, it is a measure that reveals the concentration of knowledge and the key developers in a project. Due to the importance of this information to project managers, algorithms were proposed to automatically compute Truck Factors, using maintenance activity data extracted from version control systems. However, to the best of our knowledge, we still lack studies that compare the accuracy of the results produced by such algorithms. Therefore, in this paper, we evaluate and compare the results of three Truck Factor algorithms. To this end, we empirically determine the truck factors of 35 open-source systems by consulting their developers. Our results show that two algorithms are very accurate, especially when the systems have a small Truck Factor. We also evaluate the impact of different thresholds and configurations in algorithm results. Mívian M. Ferreira, Marco Túlio Valente, Kecia Aline M. Ferreira |
ICPC | 2 |
| 2017 | RefDiff: detecting refactorings in version historiesabstractRefactoring is a well-known technique that is widely adopted by software engineers to improve the design and enable the evolution of a system. Knowing which refactoring operations were applied in a code change is a valuable information to understand software evolution, adapt software components, merge code changes, and other applications. In this paper, we present RefDiff, an automated approach that identifies refactorings performed between two code revisions in a git repository. RefDiff employs a combination of heuristics based on static analysis and code similarity to detect 13 well-known refactoring types. In an evaluation using an oracle of 448 known refactoring operations, distributed across seven Java projects, our approach achieved precision of 100% and recall of 88%. Moreover, our evaluation suggests that RefDiff has superior precision and recall than existing state-of-the-art approaches. Danilo Silva 0002, Marco Túlio Valente |
MSR | 2 |
| 2017 | Why modern open source projects failabstractOpen source is experiencing a renaissance period, due to the appearance of modern platforms and workflows for developing and maintaining public code. As a result, developers are creating open source software at speeds never seen before. Consequently, these projects are also facing unprecedented mortality rates. To better understand the reasons for the failure of modern open source projects, this paper describes the results of a survey with the maintainers of 104 popular GitHub systems that have been deprecated. We provide a set of nine reasons for the failure of these open source projects. We also show that some maintenance practices---specifically the adoption of contributing guidelines and continuous integration---have an important association with a project failure or success. Finally, we discuss and reveal the principal strategies developers have tried to overcome the failure of the studied projects. Jailton Coelho, Marco Túlio Valente |
ESEC/SIGSOFT FSE | 2 |
| 2017 | Statically identifying class dependencies in legacy JavaScript systems: First resultsabstractIdentifying dependencies between classes is an essential activity when maintaining and evolving software applications. It is also known that JavaScript developers often use classes to structure their projects. This happens even in legacy code, i.e., code implemented in JavaScript versions that do not provide syntactical support to classes. However, identifying associations and other dependencies between classes remain a challenge due to the lack of static type annotations. This paper investigates the use of type inference to identify relations between classes in legacy JavaScript code. To this purpose, we rely on Flow, a state-of-the-art type checker and inferencer tool for JavaScript. We perform a study using code with and without annotating the class import statements in two modular applications. The results show that precision is 100% in both systems, and that the annotated version improves the recall, ranging from 37% to 51% for dependencies in general and from 54% to 85% for associations. Therefore, we hypothesize that these tools should also depend on dynamic analysis to cover all possible dependencies in JavaScript code. Leonardo Humberto Silva, Marco Túlio Valente, Alexandre Bergel |
SANER | 2 |
| 2017 | Historical and impact analysis of API breaking changes: A large-scale studyabstractChange is a routine in software development. Like any system, libraries also evolve over time. As a consequence, clients are compelled to update and, thus, benefit from the available API improvements. However, some of these API changes may break contracts previously established, resulting in compilation errors and behavioral changes. In this paper, we study a set of questions regarding API breaking changes. Our goal is to measure the amount of breaking changes on real-world libraries and its impact on clients at a large-scale level. We assess (i) the frequency of breaking changes, (ii) the behavior of these changes over time, (iii) the impact on clients, and (iv) the characteristics of libraries with high frequency of breaking changes. Our large-scale analysis on 317 real-world Java libraries, 9K releases, and 260K client applications shows that (i) 14.78% of the API changes break compatibility with previous versions, (ii) the frequency of breaking changes increases over time, (iii) 2.54% of their clients are impacted, and (iv) systems with higher frequency of breaking changes are larger, more popular, and more active. Based on these results, we provide a set of lessons to better support library and client developers in their maintenance tasks. Laerte Xavier, Aline Brito, Andre Hora 0001, Marco Túlio Valente |
SANER | 4 |
| 2017 | Why do we break APIs? First answers from developersabstractBreaking contracts have a major impact on API clients. Despite this fact, recent studies show that libraries are often backward incompatible and that the rate of breaking changes increase over time. However, the specific reasons that motivate library developers to break contracts with their clients are still unclear. In this paper, we describe a qualitative study with library developers and real instance of API breaking changes. Our goal is to (i) elicit the reasons why developers introduce breaking changes; and (ii) check if they are aware about the risks of such changes. Our survey with the top contributors of popular Java libraries contributes to reveal a list of five reasons why developers break API contracts. Moreover, it also shows that most of developers are aware of these risks and, in some cases, adopt strategies to mitigate them. We conclude by prospecting a future study to strengthen our current findings. With this study, we expect to contribute on delineating tools to better assess the risks and impacts of API breaking changes. Laerte Xavier, Andre Hora 0001, Marco Túlio Valente |
SANER | 3 |
| 2017 | Identifying Classes in Legacy JavaScript CodeabstractJavaScript is the most popular programming language for the Web. Although the language is prototype-based, developers can emulate class-based abstractions in JavaScript to master the increasing complexity of their applications. Identifying classes in legacy JavaScript code can support these developers at least in the following activities: (1) program comprehension; (2) migration to the new JavaScript syntax that supports classes; and (3) implementation of supporting tools, including IDEs with class-based views and reverse engineering tools. In this paper, we propose a strategy to detect class-based abstractions in the source code of legacy JavaScript systems. We report on a large and in-depth study to understand how class emulation is employed, using a dataset of 918 JavaScript applications available on GitHub. We found that almost 70% of the JavaScript systems we study make some usage of classes. We also performed a field study with the main developers of 60 popular JavaScript systems to validate our findings. The overall results range from 97% to 100% for precision, from 70% to 89% for recall, and from 82% to 94% for F-score. Leonardo Humberto Silva, Marco Túlio Valente, Alexandre Bergel, Nicolas Anquetil, Anne Etien |
J. Softw. Evol. Process. | 2 |
| 2017 | The shape of feature code: an analysis of twenty C-preprocessor-based systems
Rodrigo Queiroz, Leonardo Teixeira Passos, Marco Túlio Valente, Claus Hunsen, Sven Apel, Krzysztof Czarnecki 0001 |
Softw. Syst. Model. | 3 |
| 2016 | Understanding the Factors That Impact the Popularity of GitHub RepositoriesabstractSoftware popularity is a valuable information to modern open source developers, who constantly want to know if their systems are attracting new users, if new releases are gaining acceptance, or if they are meeting user's expectations. In this paper, we describe a study on the popularity of software systems hosted at GitHub, which is the world's largest collection of open source software. GitHub provides an explicit way for users to manifest their satisfaction with a hosted repository: the stargazers button. In our study, we reveal the main factors that impact the number of stars of GitHub projects, including programming language and application domain. We also study the impact of new features on project popularity. Finally, we identify four main patterns of popularity growth, which are derived after clustering the time series representing the number of stars of 2,279 popular GitHub repositories. We hope our results provide valuable insights to developers and maintainers, which could help them on building and evolving systems in a competitive software market. Hudson Borges, Andre Hora 0001, Marco Túlio Valente |
ICSME | 3 |
| 2016 | A novel approach for estimating Truck FactorsabstractTruck Factor (TF) is a metric proposed by the agile community as a tool to identify concentration of knowledge in software development environments. It states the minimal number of developers that have to be hit by a truck (or quit) before a project is incapacitated. In other words, TF helps to measure how prepared is a project to deal with developer turnover. Despite its clear relevance, few studies explore this metric. Altogether there is no consensus about how to calculate it, and no supporting evidence backing estimates for systems in the wild. To mitigate both issues, we propose a novel (and automated) approach for estimating TF-values, which we execute against a corpus of 133 popular project in GitHub. We later survey developers as a means to assess the reliability of our results. Among others, we find that the majority of our target systems (65%) have TF ≤ 2. Surveying developers from 67 target systems provides confidence towards our estimates; in 84% of the valid answers we collect, developers agree or partially agree that the TF's authors are the main authors of their systems; in 53% we receive a positive or partially positive answer regarding our estimated truck factors. Guilherme Avelino 0001, Leonardo Teixeira Passos, Andre Hora 0001, Marco Túlio Valente |
ICPC | 4 |
| 2016 | When should internal interfaces be promoted to public?abstractCommonly, software systems have public (and stable) interfaces, and internal (and possibly unstable) interfaces. Despite being discouraged, client developers often use internal interfaces, which may cause their systems to fail when they evolve. To overcome this problem, API producers may promote internal interfaces to public. In practice, however, API producers have no assistance to identify public interface candidates. In this paper, we study the transition from internal to public interfaces. We aim to help API producers to deliver a better product and API clients to benefit sooner from public interfaces. Our empirical investigation on five widely adopted Java systems present the following observations. First, we identified 195 promotions from 2,722 internal interfaces. Second, we found that promoted internal interfaces have more clients. Third, we predicted internal interface promotion with precision between 50%-80%, recall 26%-82%, and AUC 74%-85%. Finally, by applying our predictor on the last version of the analyzed systems, we automatically detected 382 public interface candidates. Andre Hora 0001, Marco Túlio Valente, Romain Robbes, Nicolas Anquetil |
SIGSOFT FSE | 2 |
| 2016 | Why we refactor? confessions of GitHub contributorsabstractRefactoring is a widespread practice that helps developers to improve the maintainability and readability of their code. However, there is a limited number of studies empirically investigating the actual motivations behind specific refactoring operations applied by developers. To fill this gap, we monitored Java projects hosted on GitHub to detect recently applied refactorings, and asked the developers to explain the reasons behind their decision to refactor the code. By applying thematic analysis on the collected responses, we compiled a catalogue of 44 distinct motivations for 12 well-known refactoring types. We found that refactoring activity is mainly driven by changes in the requirements and much less by code smells. Extract Method is the most versatile refactoring operation serving 11 different purposes. Finally, we found evidence that the IDE used by the developers affects the adoption of automated refactoring tools. Danilo Silva 0002, Nikolaos Tsantalis, Marco Túlio Valente |
SIGSOFT FSE | 3 |
| 2016 | Do Developers Deprecate APIs with Replacement Messages? A Large-Scale Analysis on Java SystemsabstractAs any other software system, frameworks and libraries evolve over time, and so their APIs. Consequently, client systems should be updated to benefit from improved APIs. To facilitate this task and preserve backward compatibility, API elements should always be deprecated with clear replacement messages. However, in practice, there are evidences that API elements are usually deprecated without such messages. In this paper, we study a set of questions regarding the adoption of deprecation messages. Our goal is twofold: to measure the usage of deprecation messages and to investigate whether a tool is needed to recommend such messages. Thus, we verify (i) the frequency of deprecated elements with replacement messages, (ii) the impact of software evolution on such frequency, and (iii) the characteristics of systems which deprecate API elements in a correct way. Our large-scale analysis on 661 real-world Java systems shows that (i) 64% of the API elements are deprecated with replacement messages per system, (ii) there is almost no major effort to improve deprecation messages over time, and (iii) systems that deprecated API elements in a correct way are statistically significantly different from the ones that do not in terms of size and developing community. As a result, we provide the basis for the design of a tool to support client developers on detecting missing deprecation messages. Gleison Brito, Andre Hora 0001, Marco Túlio Valente, Romain Robbes |
SANER | 3 |
| 2016 | Identifying Utility Functions Using Random ForestsabstractUtility functions are general purpose functions, which are useful in many parts of a system. To facilitate reuse, they are usually implemented in specific libraries. However, developers frequently miss opportunities to implement general-purpose functions in utility libraries, which decreases the chances of reuse. In this paper, we describe our ongoing investigation on using Random Forest classifiers to automatically identify utility functions. Using a list of static source code metrics we train a classifier to identify such functions, both in Java (using 84 projects from the Qualitas Corpus) and in JavaScript (using 22 popular projects from GitHub). We achieve the following median results for Java: 0.90 (AUC), 0.83 (precision), 0.88 (recall), and 0.84 (F-measure). For JavaScript, the median results are 0.80 (AUC), 0.75 (precision), 0.89 (recall), and 0.76 (F-measure). Tamara Mendes, Marco Túlio Valente, Andre Hora 0001, Alexander Serebrenik |
SANER | 2 |
| 2016 | An Empirical Study on Recommendations of Similar BugsabstractThe work to be performed on open source systems, whether feature developments or defects, is typically described as an issue (or bug). Developers self-select bugs from the many open bugs in a repository when they wish to perform work on the system. This paper evaluates a recommender, called NextBug, that considers the textual similarity of bug descriptions to predict bugs that require handling of similar code fragments. First, we evaluate this recommender using 69 projects in the Mozilla ecosystem. We show that for detecting similar bugs, a technique that considers just the bug components and short descriptions perform just as well as a more complex technique that considers other features. Second, we report a field study where we monitored the bugs fixed for Mozilla during a week. We sent mails to the developers who fixed these bugs, asking whether they would consider working on the recommendations provided by NextBug, 39 developers (59%) stated that they would consider working on these recommendations, 44 developers (67%) also expressed interest in seeing the recommendations in their bug tracking system. Henrique Rocha, Marco Túlio Valente, Humberto Torres Marques-Neto, Gail C. Murphy |
SANER | 2 |
| 2016 | Learning from Source Code History to Identify Performance FailuresabstractSource code changes may inadvertently introduce performance regressions. Benchmarking each software version is traditionally employed to identify performance regressions. Although effective, this exhaustive approach is hard to carry out in practice. This paper contrasts source code changes against performance variations. By analyzing 1,288 software versions from 17 open source projects, we identified 10 source code changes leading to a performance variation (improvement or regression). We have produced a cost model to infer whether a software commit introduces a performance variation by analyzing the source code and sampling the execution of a few versions. By profiling the execution of only 17% of the versions, our model is able to identify 83% of the performance regressions greater than 5% and 100% of the regressions greater than 50%. Juan Pablo Sandoval Alcocer, Alexandre Bergel, Marco Túlio Valente |
ICPE | 3 |
| 2016 | Mining architectural violations from version history
Cristiano Amaral Maffort, Marco Túlio Valente, Ricardo Terra, Mariza Andrade da Silva Bigonha, Nicolas Anquetil, Andre Hora 0001 |
Empir. Softw. Eng. | 2 |
| 2015 | How do developers react to API evolution? The Pharo ecosystem caseabstractSoftware engineering research now considers that no system is an island, but it is part of an ecosystem involving other systems, developers, users, hardware, ... When one system (e.g., a framework) evolves, its clients often need to adapt. Client developers might need to adapt to functionalities, client systems might need to be adapted to a new API, client users might need to adapt to a new User Interface. The consequences of such changes are yet unclear, what proportion of the ecosystem might be expected to react, how long might it take for a change to diffuse in the ecosystem, do all clients react in the same way? This paper reports on an exploratory study aimed at observing API evolution and its impact on a large-scale software ecosystem, Pharo, which has about 3,600 distinct systems, more than 2,800 contributors, and six years of evolution. We analyze 118 API changes and answer research questions regarding the magnitude, duration, extension, and consistency of such changes in the ecosystem. The results of this study help to characterize the impact of API evolution in large software ecosystems, and provide the basis to better understand how such impact can be alleviated. Andre Hora 0001, Romain Robbes, Nicolas Anquetil, Anne Etien, Stéphane Ducasse, Marco Túlio Valente |
ICSME | 6 |
| 2015 | Apiwave: Keeping track of API popularity and migrationabstractEvery day new frameworks and libraries are created and existing ones evolve. To benefit from such newer or improved APIs, client developers should update their applications. In practice, this process presents some challenges: APIs are commonly backward-incompatible (causing client applications to fail when updating) and multiple APIs are available (making it difficult to decide which one to use). To address these challenges, we propose apiwave, a tool that keeps track of API popularity and migration of major frameworks/libraries. The current version includes data about the evolution of top 650 GitHub Java projects, from which 320K APIs were extracted. We also report an experience using apiwave on real-world scenarios. Andre Hora 0001, Marco Túlio Valente |
ICSME | 2 |
| 2015 | Validating metric thresholds with developers: An early resultabstractThresholds are essential for promoting source code metrics as an effective instrument to control the internal quality of software applications. However, little is known about the relation between software quality as identified by metric thresholds and as perceived by real developers. In this paper, we report the first results of a study designed to validate a technique that extracts relative metric thresholds from benchmark data. We use this technique to extract thresholds from a benchmark of 79 Pharo/Smalltalk applications, which are validated with five experts and 25 developers. Our preliminary results indicate that good quality applications - as cited by experts - respect metric thresholds. In contrast, we observed that noncompliant applications are not largely viewed as requiring more effort to maintain than other applications. Paloma Oliveira, Marco Túlio Valente, Alexandre Bergel, Alexander Serebrenik |
ICSME | 2 |
| 2015 | System specific, source code transformationsabstractDuring its lifetime, a software system might undergo a major transformation effort in its structure, for example to migrate to a new architecture or bring some drastic improvements to the system. Particularly in this context, we found evidences that some sequences of code changes are made in a systematic way. These sequences are composed of small code transformations (e.g., create a class, move a method) which are repeatedly applied to groups of related entities (e.g., a class and some of its methods). A typical example consists in the systematic introduction of a Factory design pattern on the classes of a package. We define these sequences as transformation patterns. In this paper, we identify examples of transformation patterns in real world software systems and study their properties: (i) they are specific to a system; (ii) they were applied manually; (iii) they were not always applied to all the software entities which could have been transformed; (iv) they were sometimes complex; and (v) they were not always applied in one shot but over several releases. These results suggest that transformation patterns could benefit from automated support in their application. From this study, we propose as future work to develop a macro recorder, a tool with which a developer records a sequence of code transformations and then automatically applies them in other parts of the system as a customizable, large-scale transformation operator. Gustavo Santos, Nicolas Anquetil, Anne Etien, Stéphane Ducasse, Marco Túlio Valente |
ICSME | 5 |
| 2015 | Developers' perception of co-change patterns: An empirical studyabstractCo-change clusters are groups of classes that frequently change together. They are proposed as an alternative modular view, which can be used to assess the traditional decomposition of systems in packages. To investigate developer's perception of co-change clusters, we report in this paper a study with experts on six systems, implemented in two languages. We mine 102 co-change clusters from the version history of such systems, which are classified in three patterns regarding their projection to the package structure: Encapsulated, Crosscutting, and Octopus. We then collect the perception of expert developers on such clusters, aiming to ask two central questions: (a) what concerns and changes are captured by the extracted clusters? (b) do the extracted clusters reveal design anomalies? We conclude that Encapsulated Clusters are often viewed as healthy designs and that Crosscutting Clusters tend to be associated to design anomalies. Octopus Clusters are normally associated to expected class distributions, which are not easy to implement in an encapsulated way, according to the interviewed developers. Luciana Lourdes Silva, Marco Túlio Valente, Marcelo de Almeida Maia, Nicolas Anquetil |
ICSME | 2 |
| 2015 | Recording and replaying system specific, source code transformationsabstractDuring its lifetime, a software system is under continuous maintenance to remain useful. Maintenance can be achieved in activities such as adding new features, fixing bugs, improving the system's structure, or adapting to new APIs. In such cases, developers sometimes perform sequences of code changes in a systematic way. These sequences consist of small code changes (e.g., create a class, then extract a method to this class), which are applied to groups of related code entities (e.g., some of the methods of a class). This paper presents the design and proof-of-concept implementation of a tool called MacroRecorder. This tool records a sequence of code changes, then it allows the developer to generalize this sequence in order to apply it in other code locations. In this paper, we discuss MACRORECORDER's approach that is independent of both development and transformation tools. The evaluation is based on previous work on repetitive code changes related to rearchitecting. MacroRecorder was able to replay 92% of the examples, which consisted in up to seven code entities modified up to 66 times. The generation of a customizable, large-scale transformation operator has the potential to efficiently assist code maintenance. Gustavo Santos, Anne Etien, Nicolas Anquetil, Stéphane Ducasse, Marco Túlio Valente |
SCAM | 5 |
| 2015 | OrionPlanning: Improving modularization and checking consistency on software architectureabstractMany techniques have been proposed in the literature to support architecture definition, conformance, and analysis. However, there is a lack of adoption of such techniques by the industry. Previous work have analyzed this poor support. Specifically, former approaches lack proper analysis techniques (e.g., detection of architectural inconsistencies), and they do not provide extension and addition of new features. In this paper, we present ORIONPLANNING, a prototype tool to assist refactorings at large scale. The tool provides support for model-based refactoring operations. These operations are performed in an interactive visualization. The contributions of the tool consist in: (i) providing iterative modifications in the architecture, and (ii) providing an environment for architecture inspection and definition of dependency rules. We evaluate ORIONPLANNING against practitioners' requirements on architecture definition listed in a previous survey. We also evaluate the tool in a concrete example of software remodularization. Gustavo Santos, Nicolas Anquetil, Anne Etien, Stéphane Ducasse, Marco Túlio Valente |
VISSOFT | 5 |
| 2015 | Does JavaScript software embrace classes?abstractJavaScript is the de facto programming language for the Web. It is used to implement mail clients, office applications, or IDEs, that can weight hundreds of thousands of lines of code. The language itself is prototype based, but to master the complexity of their application, practitioners commonly rely on informal class abstractions. This practice has never been the target of empirical research in JavaScript. Yet, understanding it is key to adequately tuning programming environments and structure libraries such that they are accessible to programmers. In this paper we report on a large and in-depth study to understand how class emulation is employed in JavaScript applications. We propose a strategy to statically detect class-based abstractions in the source code of JavaScript systems. We used this strategy in a dataset of 50 popular JavaScript applications available from GitHub. We found four types of JavaScript software: class-free (systems that do not make any usage of classes), class-aware (systems that use classes, but marginally), class-friendly (systems that make a relevant usage of classes), and class-oriented (systems that have most of their data structures implemented as classes). The systems in these categories represent, respectively, 26%, 36%, 30%, and 8% of the systems we studied. Leonardo Humberto Silva, Miguel Ramos 0001, Marco Túlio Valente, Alexandre Bergel, Nicolas Anquetil |
SANER | 3 |
| 2015 | Automatic detection of system-specific conventions unknown to developers
Andre Hora 0001, Nicolas Anquetil, Anne Etien, Stéphane Ducasse, Marco Túlio Valente |
J. Syst. Softw. | 5 |
| 2015 | A recommendation system for repairing violations detected by static architecture conformance checkingabstractSummary This paper describes a recommendation system that provides refactoring guidelines for maintainers when tackling architectural erosion. The paper formalizes 32 refactoring recommendations to repair violations raised by static architecture conformance checking approaches; it describes a tool—called ArchFix—that triggers the proposed recommendations; and it evaluates the application of this tool in two industrial‐strength systems. For the first system—a 21 KLOC open‐source strategic management system—our approach has indicated correct refactoring recommendations for 31 out of 41 violations detected as the result of an architecture conformance process. For the second system—a 728 KLOC customer care system used by a major telecommunication company—our approach has triggered correct recommendations for 624 out of 787 violations, as asserted by the system's architect. Moreover, the architects have scored 82% of these recommendations as havingmoderateormajorcomplexity. Copyright © 2013 John Wiley & Sons, Ltd. Ricardo Terra, Marco Túlio Valente, Krzysztof Czarnecki 0001, Roberto da Silva Bigonha |
Softw. Pract. Exp. | 2 |
| 2014 | Object-Business Process Mapping Frameworks: Abstractions, Architecture, and ImplementationabstractThe integration between enterprise architectures and Business Process Management Systems (BPMS) is currently based on low-level programming interfaces that expose accidental complexities typical of process implementations. This paper describes an approach for integrating software architectures and BPMSs, based on mapping frameworks. Our inspiration are the Object-Relational Mapping (ORM) frameworks widely used to shield information systems from low-level structures exposed by relational database systems. The paper describes the central abstractions that should be provided by Object -- Business Process Mapping Frameworks (OBPM). We also propose a reference architecture for implementing OBPMs and a concrete OBPM implementation, called NextFlow. We evaluated our approach by comparing two implementations of the same system, one using NextFlow and another using the native API supported by jBPM, a popular BPMS. By using NextFlow, we achieved a reduction of 30% in terms of lines of code, 35% in terms of number of classes, and 90% in terms of import statements, when implementing this system. Rogel Garcia, Marco Túlio Valente |
EDOC | 2 |
| 2014 | RTTool: A Tool for Extracting Relative Thresholds for Source Code MetricsabstractMeaningful thresholds are essential for promoting source code metrics as an effective instrument to control the internal quality of software systems. Despite the increasing number of source code measurement tools, no publicly available tools support extraction of metric thresholds. Moreover, earlier studies suggest that in larger systems significant number of classes exceed recommended metric thresholds. Therefore, in our previous study we have introduced the notion of a relative threshold, i.e., a pair including an upper limit and a percentage of classes whose metric values should not exceed this limit. In this paper we propose RTTOOL, an open source tool for extracting relative thresholds from the measurement data of a benchmark of software systems. RTTOOL is publicly available at http://aserg.labsoft.dcc.ufmg.br/rttool. Paloma Oliveira, Fernando Paim Lima, Marco Túlio Valente, Alexander Serebrenik |
ICSME | 3 |
| 2014 | Recommending automated extract method refactoringsabstractExtract Method is a key refactoring for improving program comprehension. However, recent empirical research shows that refactoring tools designed to automate Extract Methods are often underused. To tackle this issue, we propose a novel approach to identify and rank Extract Method refactoring opportunities that are directly automated by IDE-based refactoring tools. Our approach aims to recommend new methods that hide structural dependencies that are rarely used by the remaining statements in the original method. We conducted an exploratory study to experiment and define the best strategies to compute the dependencies and the similarity measures used by the proposed approach. We also evaluated our approach in a sample of 81 extract method opportunities generated for JUnit and JHotDraw, achieving a precision of 48% (JUnit) and 38% (JHotDraw). Danilo Silva 0002, Ricardo Terra, Marco Túlio Valente |
ICPC | 3 |
| 2014 | Predicting software defects with causality tests
César Couto, Pedro Pires, Marco Túlio Valente, Roberto da Silva Bigonha, Nicolas Anquetil |
J. Syst. Softw. | 3 |
| 2013 | Mining Architectural Patterns Using Association Rules
Cristiano Amaral Maffort, Marco Túlio Valente, Roberto da Silva Bigonha, Andre Hora 0001, Nicolas Anquetil, Jonata Menezes |
SEKE | 2 |
| 2013 | Metrics-based Detection of Similar Software (S)
Paloma Oliveira, Hudson Borges, Marco Túlio Valente, Heitor A. X. Costa |
SEKE | 3 |
| 2013 | Measuring the Structural Similarity between Source Code Entities (S)
Ricardo Terra, João Brunet, Luis Fernando Miranda, Marco Túlio Valente, Dalton Serey Guerrero, Douglas Castilho 0001, Roberto da Silva Bigonha |
SEKE | 4 |
| 2013 | The crosscutting impact of the AOSD Brazilian research community
Uirá Kulesza, Sérgio Soares, Christina von Flach G. Chavez, Fernando Castor Filho, Paulo Borba, Carlos José Pereira de Lucena, Paulo César Masiero, Cláudio Sant'Anna, Fabiano Cutigi Ferrari, Vander Alves, Roberta Coelho, Eduardo Figueiredo 0001, Paulo F. Pires, Flávia Coimbra Delicato, Eduardo Piveta, Carla T. L. L. Silva, Valter Vieira de Camargo, Rosana T. V. Braga, Julio César Sampaio do Prado Leite, Otávio Augusto Lazzarini Lemos, Nabor das Chagas Mendonça, Thaís Vasconcelos Batista, Rodrigo Bonifácio, Nélio Cacho, Lyrene Fernandes da Silva, Arndt von Staa, Fábio Fagundes Silveira, Marco Túlio Valente, Fernanda M. R. Alencar, Jaelson Brelaz de Castro, Ricardo Argenton Ramos, Rosângela A. D. Penteado, Cecília M. F. Rubira |
J. Syst. Softw. | 28 |
| 2013 | Static correspondence and correlation between field defects and warnings reported by a bug finding tool
César Couto, João Eduardo Montandon, Christofer Silva, Marco Túlio Valente |
Softw. Qual. J. | 4 |
| 2013 | Mining the impact of evolution categories on object-oriented metrics
Henrique Rocha, César Couto, Cristiano Amaral Maffort, Rogel Garcia, Clarisse Simões, Leonardo Teixeira Passos, Marco Túlio Valente |
Softw. Qual. J. | 7 |
| 2012 | A Semi-Automatic Approach for Extracting Software Product LinesabstractThe extraction of nontrivial software product lines (SPL) from a legacy application is a time-consuming task. First, developers must identify the components responsible for the implementation of each program feature. Next, they must locate the lines of code that reference the components discovered in the previous step. Finally, they must extract those lines to independent modules or annotate them in some way. To speed up product line extraction, this paper describes a semi-automatic approach to annotate the code of optional features in SPLs. The proposed approach is based on an existing tool for product line development, called CIDE, that enhances standard IDEs with the ability to associate background colors with the lines of code that implement a feature. We have evaluated and successfully applied our approach to the extraction of optional features from three nontrivial systems: Prevayler (an in-memory database system), JFreeChart (a chart library), and ArgoUML (a UML modeling tool). Marco Túlio Valente, Virgilio Borges, Leonardo Teixeira Passos |
IEEE Trans. Software Eng. | 1 |
| 2011 | How Annotations are Used in Java: An Empirical Study
Henrique Rocha, Marco Túlio Valente |
SEKE | 2 |
| 2009 | Object-oriented transformations for extracting aspects
Marcelo Nassau Malta, Marco Túlio Valente |
Inf. Softw. Technol. | 2 |
| 2009 | A dependency constraint language to manage object-oriented software architecturesabstractAbstract This paper presents a domain‐specific dependency constraint language that allows software architects to restrict the spectrum of structural dependencies, which can be established in object‐oriented systems. The ultimate goal is to provide architects with means to define acceptable and unacceptable dependencies according to the planned architecture of their systems. Once defined, such restrictions are statically enforced by a tool, thus avoiding silent erosions in the architecture. The paper also presents results from applying the proposed approach to different versions of a real‐world human resource management system. Copyright © 2009 John Wiley & Sons, Ltd. Ricardo Terra, Marco Túlio Valente |
Softw. Pract. Exp. | 2 |
| 2008 | Towards a Dependency Constraint Language to Manage Software Architectures
Ricardo Terra, Marco Túlio Valente |
ECSA | 2 |
| 2008 | Non-invasive and non-scattered annotations for more robust pointcutsabstractAnnotations are often mentioned as a potential alternative to tackle the fragile nature of AspectJ pointcuts. However, annotations themselves can be considered crosscutting elements because they are normally pervasive and tangled with business-specific functionality. In this paper, we propose a solution to the fragile pointcut problem in aspect-oriented programming that relies on non-invasive and non-scattered annotations. The central components of the proposed solution are so-called annotator aspects, that superimpose annotations to the base code in a non-invasive way. Moreover, annotator aspects are generated semiautomatically, from a declarative annotation definition language. The paper presents examples of using the proposed solution in pointcut descriptors of two real-world aspect-oriented systems. We also describe a case study that evaluates the robustness of the proposed solution in face of possible changes to the classical Figure Editor system. Leonardo Humberto Silva, Samuel Domingues, Marco Túlio Valente |
ICSM | 3 |
| 2008 | A Loosely Coupled Aspect Language for SOA ApplicationsabstractThe aspect-oriented programming (AOP) paradigm offers software developers with powerful modularization abstractions to help them explicitly separate design concerns at the source code level. However, the impact of AOP in the service-oriented architecture (SOA) paradigm has been dwarfed by the fact that existing AOP solutions are tightly coupled to a particular programming language, middleware system or execution platform. Clearly, this not only restricts the implementation choices available to application developers, but it also clashes with the heterogeneous and loosely coupled nature of SOA. This paper presents the Web Service Aspect Language (WSAL) that seamlessly integrates AOP and SOA concepts, thus avoiding the drawbacks of existing solutions. In WSAL, aspects themselves are freely specified, implemented and executed as loosely coupled web services. This characteristic allows WSAL aspects to be easily woven into the message flow exchanged between service consumers and service providers, in a way that is completely independent from any particular implementation technology. This paper also reports on the implementation and preliminary evaluation of a prototype aspect weaver for WSAL, which is based on an existing web intermediary technology. Nabor das Chagas Mendonça, Clayton F. Silva, Ian G. Maia, Maria Andréia F. Rodrigues, Marco Túlio Valente |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2007 | Collocation optimizations in an aspect-oriented middleware system
Marco Túlio Valente, Rodrigo Palhares Silva |
J. Syst. Softw. | 1 |
| 2006 | Arcademis: a framework for object-oriented communication middleware developmentabstractThis paper presents Arcademis, a Java-based framework for communication middleware development. Arcademis consists of a set of abstract classes, interfaces and concrete components that define the general architecture of middleware systems. Its main objective is to support the implementation of non-monolithic and easily configurable middleware platforms. Arcademis can be used by middleware developers to deploy systems that meet the requirements of a particular network or technology. Instances of Arcademis can also be customized by distributed systems engineers to meet the requirements of a particular application. For example, new transport protocols, connection management policies, authentication algorithms or invocation semantics can be easily configured in middleware platforms derived from Arcademis. In order to illustrate the use of the framework, the paper describes the RME system, a middleware derived from Arcademis that adds a remote method invocation service to the CLDC configuration of Java 2 Micro Edition (J2ME). Copyright © 2006 John Wiley & Sons, Ltd. Fernando Magno Quintão Pereira, Marco Túlio Valente, Roberto da Silva Bigonha, Mariza Andrade da Silva Bigonha |
Softw. Pract. Exp. | 2 |
| 2004 | Personalizing Web Sites for Mobile Devices Using a Graphical User Interface
Leonardo Teixeira Passos, Marco Túlio Valente |
ICWE | 2 |
| 2004 | Coordination and mobility in CoreLimeabstractThe choice of suitable high-level communication primitives for wide area network programming languages remains an open problem. This paper is driven by the practical consideration of providing an efficient and secure communication infrastructure for mobile agent systems. This has led us to formalise the Lime coordination middleware and propose a simplified model, which we call CoreLime, that addresses some of the main shortcomings of Lime while retaining its distinguishing feature, namely transient sharing of tuple spaces. We further discuss a prototype implementation along with security extensions. Our contribution is thus an exploration of the language design space rather than a theoretical investigation of properties of these models. Bogdan Carbunar, Marco Túlio Valente, Jan Vitek |
Math. Struct. Comput. Sci. | 2 |
| 2003 | A Coordination Model for ad hoc Mobile Systems
Marco Túlio Valente, Fernando Magno Quintão Pereira, Roberto da Silva Bigonha, Mariza Andrade da Silva Bigonha |
Euro-Par | 1 |