Andre Hora 0001

dblp:54/7440 · also André C. Hora · DBLP profile ↗
← Back
52ranked-venue papers
19as first author
25since 2021 · last 2026
0000-0003-4900-1330ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 51 · 18 first-author · 25 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Testing Framework Migration with Large Language Models
abstract
Python developers rely on two major testing frameworks: unittest and Pytest. While Pytest offers simpler assertions, reusable fixtures, and better interoperability, migrating existing suites from unittest remains a manual and time-consuming process. Automating this migration could substantially reduce effort and accelerate test modernization. In this paper, we investigate the capability of Large Language Models (LLMs) to automate test framework migrations from unittest to Pytest. We evaluate GPT 4o and Claude Sonnet 4 under three prompting strategies (Zero-shot, One-shot, and Chain-of-Thought) and two temperature settings (0.0 and 1.0). To support this analysis, we first introduce a curated dataset of real-world migrations extracted from the top 100 Python open-source projects. Next, we actually execute the LLM-generated test migrations in their respective test suites. Overall, we find that 51.5% of the LLM-generated test migrations failed, while 48.5% passed. The results suggest that LLMs can accelerate test migration, but there are often caveats. For example, Claude Sonnet 4 exhibited more conservative migrations (e.g., preserving class-based tests and legacy unittest references), while GPT-4o favored more transformations (e.g., to function-based tests). We conclude by discussing multiple implications for practitioners and researchers.
Altino Alves, João Eduardo Montandon, Andre Hora 0001
AST3
2026 Understanding Bug-Reproducing Tests: A First Empirical Study
abstract
Developers create bug-reproducing tests that support debugging by failing as long as the bug is present, and passing once the bug has been fixed. These tests are usually integrated into existing test suites and executed regularly alongside all other tests to ensure that future regressions are caught. Despite this co-existence with other types of tests, the properties of bug-reproducing tests are scarcely researched, and it remains unclear whether they differ fundamentally. In this short paper, we provide an initial empirical study to understand bug-reproducing tests better. We analyze 642 bug-reproducing tests of 15 real-world Python systems. Overall, we find that bug-reproducing tests are not (statistically significantly) different from other tests regarding LOC, number of assertions, and complexity. However, bug-reproducing tests contain slightly more try/except blocks and “weak assertions” (e.g., assertNotEqual). Lastly, we detect that the majority (95%) of the bug-reproducing tests reproduce a single bug, while 5% reproduce multiple bugs. We conclude by discussing implications and future research directions.
Andre Hora 0001, Gordon Fraser 0001
AST1
2026 Test Behaviors, Not Methods! Detecting Tests Obsessed by Methods
abstract
Best testing practices state that tests should verify a single functionality or behavior of the system. Tests that verify multiple behaviors are harder to understand, lack focus, and are more coupled to the production code. An attempt to identify this issue is the test smell Eager Test, which aims to capture tests that verify too much functionality based on the number of production method calls. Unfortunately, prior research suggests that counting production method calls is an inaccurate measure, as these calls do not reliably serve as a proxy for functionality. We envision a complementary solution based on runtime analysis: we hypothesize that some tests that verify multiple behaviors will likely cover multiple paths of the same production methods. Thus, we propose a novel test smell named Test Obsessed by Method, a test method that covers multiple paths of a single production method. We provide an initial empirical study to explore the presence of this smell in 2,054 tests provided by 12 test suites of the Python Standard Library. (1) We detect 44 Tests Obsessed by Methods in 11 of the 12 test suites. (2) Each smelly test verifies a median of two behaviors of the production method. (3) The 44 smelly tests could be split into 118 novel tests. (4) 23% of the smelly tests have code comments recognizing that distinct behaviors are being tested. We conclude by discussing benefits, limitations, and further research.
Andre Hora 0001, Andy Zaidman
ICPC1
2026 GitEvo: Code Evolution Analysis for Git Repositories
abstract
Analyzing the code evolution of software systems is relevant for practitioners, researchers, and educators. It can help practitioners identify design trends and maintenance challenges, provide researchers with empirical data to study changes over time, and give educators real-world examples that enhance the teaching of software evolution concepts. Unfortunately, we lack tools specifically designed to support code evolution analysis. In this paper, we propose GitEvo, a multi-language and extensible tool for analyzing code evolution in Git repositories. GitEvo leverages Git frameworks and code parsing tools to integrate both Git-level and code-level analysis. We conclude by describing how GitEvo can support the development of novel empirical studies on code evolution and act as a learning tool for educators and students. GitEvo is available at: https://github.com/andrehora/gitevo.
Andre Hora 0001
MSR1
2026 Are Coding Agents Generating Over-Mocked Tests? An Empirical Study
abstract
Coding agents have received significant adoption in software development recently. Unlike traditional LLM-based code completion tools, coding agents work with autonomy (e.g., invoking external tools) and leave visible traces in software repositories, such as authoring commits. Among their tasks, coding agents may autonomously generate software tests; however, the quality of these tests remains uncertain. In particular, excessive use of mocking can make tests harder to understand and maintain. This paper presents the first study to investigate the presence of mocks in agent-generated tests of real-world software systems. We analyzed over 1.2 million commits made in 2025 in 2,168 TypeScript, JavaScript, and Python repositories, including 48,563 commits by coding agents, 169,361 commits that modify tests, and 44,900 commits that add mocks to tests. Overall, we find that coding agents are more likely to modify tests and to add mocks to tests than non-coding agents. We detect that (1) 60% of the repositories with agent activity also contain agent test activity; (2) 23% of commits made by coding agents add/change test files, compared with 13% by non-agents; (3) 68% of the repositories with agent test activity also contain agent mock activity; (4) 36% of commits made by coding agents add mocks to tests, compared with 26% by non-agents; and (5) repositories created recently contain a higher proportion of test and mock commits made by agents. Finally, we conclude by discussing implications for developers and researchers. We call attention to the fact that tests with mocks may be potentially easier to generate automatically (but less effective at validating real interactions), and the need to include guidance on mocking practices in agent configuration files.
Andre Hora 0001, Romain Robbes
MSR1
2026 Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity
abstract
In 2025, coding agents have seen a very rapid adoption. Coding agents leverage Large Language Models (LLMs) in ways that are markedly different from LLM-based code completion, making their study critical. Moreover, unlike LLM-based completion, coding agents leave visible traces in software repositories, enabling the use of MSR techniques to study their impact on SE practices. This paper documents the promises, perils, and heuristics that we have gathered from studying coding agent activity on GitHub.
Romain Robbes, Théo Matricon, Thomas Degueule, Andre Hora 0001, Stefano Zacchiroli
MSR4
2026 Assessing reusable web applications: The Django ecosystem case
abstract
Modern web applications are built on top of web frameworks, which offer benefits such as feature reuse and productivity improvement. By relying on web frameworks, developers can safely handle the multiple needs of real-world web applications. Despite the benefits of web frameworks, they are not intended to address all requirements to create web applications. In particular, the Django web framework provides a rich software ecosystem of reusable applications to support developers with advanced development aspects, such as modern authentication, caching, and deployment. However, currently, it is not clear what solutions developers can reuse or how active and reliable this ecosystem really is. To overcome this gap, this paper proposes an empirical study to assess the Django ecosystem of reusable applications. We analyze 487 reusable applications to understand their target domain, activity level, frequency of tests, and compatibility. We find that reusable applications are mostly concentrated in three categories: developer tools (27.8%), user interface (17.1%), and persistence (15.9%). Overall, these categories have distinct characteristics; for example, developer tools have the highest concentration of young and inactive applications, while persistence applications are the oldest. The reusable applications are 7.7 years old but have 10.17 months without commits, on the median. We also find that 80% of the reusable applications include the supported Python versions, 64% include the Django versions, and only 61% include both Python and Django versions. We detect that upgrading to recent versions of Django and Python may not be supported in some reusable applications, but compatibility issues become less problematic in recent Django and Python versions. Finally, based on our findings, we provide implications for practitioners and researchers.
Andre Hora 0001
Softw. Qual. J.2
2025 Exceptional Behaviors: How Frequently Are They Tested?
abstract
Exceptions allow developers to handle error cases expected to occur infrequently. Ideally, good test suites should test both normal and exceptional behaviors to catch more bugs and avoid regressions. While current research analyzes exceptions that propagate to tests, it does not explore other exceptions that do not reach the tests. In this paper, we provide an empirical study to explore how frequently exceptional behaviors are tested in real-world systems. We consider both exceptions that propagate to tests and the ones that do not reach the tests. For this purpose, we run an instrumented version of test suites, monitor their execution, and collect information about the exceptions raised at runtime. We analyze the test suites of 25 Python systems, covering 5,372 executed methods, 17.9M calls, and 1.4M raised exceptions. We find that 21.4% of the executed methods do raise exceptions at runtime. In methods that raise exceptions, on the median, 1 in 10 calls exercise exceptional behaviors. Close to 80% of the methods that raise exceptions do so infrequently, but about 20% raise exceptions more frequently. Finally, we provide implications for researchers and practitioners. We suggest developing novel tools to support exercising exceptional behaviors and refactoring expensive try/except blocks. We also call attention to the fact that exception-raising behaviors are not necessarily "abnormal" or rare.
Andre Hora 0001, Gordon Fraser 0001
AST1
2025 TestMigrationsInPy: A Dataset of Test Migrations from Unittest to Pytest
abstract
Unittest and pytest are the most popular testing frameworks in Python. Overall, pytest provides some advantages, including simpler assertion, reuse of fixtures, and interoperability. Due to such benefits, multiple projects in the Python ecosystem have migrated from unittest to pytest. To facilitate the migration, pytest can also run unittest tests, thus, the migration can happen gradually over time. However, the migration can be time-consuming and take a long time to conclude. In this context, projects would benefit from automated solutions to support the migration process. In this paper, we propose TestMigrationsInPy, a dataset of test migrations from unittest to pytest. TestMigrationsInPy contains 923 real-world migrations performed by developers. Future research proposing novel solutions to migrate frameworks in Python can rely on TestMigrationsInPy as a ground truth. Moreover, as TestMigrationsInPy includes information about the migration type (e.g., changes in assertions or fixtures), our dataset enables novel solutions to be verified effectively, for instance, from simpler assertion migrations to more complex fixture migrations. TestMigrationsInPy is publicly available at: https://github.com/altinoalvesjunior/TestMigrationsInPy.
Altino Alves, Andre Hora 0001
MSR2
2025 What Do Contribution Guidelines Say About Software Testing?
abstract
Software testing plays a crucial role in the contribution process of open-source projects. For example, contributions introducing new features are expected to include tests, and contributions with tests are more likely to be accepted. Although most real-world projects require contributors to write tests, the specific testing practices communicated to contributors remain unclear. In this paper, we present an empirical study to understand better how software testing is approached in contribution guidelines. We analyze the guidelines of 200 Python and JavaScript open-source software projects. We find that $78 \%$ of the projects include some form of test documentation for contributors. Test documentation is located in multiple sources, including CONTRIBUTING files ($\mathbf{5 8 \%}$), external documentation ($\mathbf{2 4 \%}$), and README files ($\mathbf{8 \%}$). Furthermore, test documentation commonly explains how to run tests ($83.5 \%$), but less often provides guidance on how to write tests $(37 \%)$. It frequently covers unit tests $(71 \%)$, but rarely addresses integration ($\mathbf{2 0. 5 \%}$) and end-to-end tests ($\mathbf{1 5. 5 \%}$). Other key testing aspects are also less frequently discussed: test coverage ($\mathbf{2 5. 5 \%}$) and mocking ($\mathbf{9. 5 \%}$). We conclude by discussing implications and future research.
Bruna Falcucci, Felipe Gomide, Andre Hora 0001
MSR3
2025 How and why developers implement OS-specific tests
Ricardo de Sousa Job, Andre Hora 0001
Empir. Softw. Eng.2
2024 TestDossier: A Dataset of Tested Values Automatically Extracted from Test Execution
abstract
Real-world test suites are often complex and may have thousands of test cases. In this scenario, it is not easy to spot what values are actually covered by the tests. Having access to every tested value of a test suite would provide the basis to (1) assess the quality of tested data and (2) have actionable information to improve them. In this paper, we propose TestDossier, a dataset of tested values automatically extracted from the execution of Python tests. To collect runtime data, we run an instrumented version of the tests, monitoring the test execution, and extracting argument and variable values. We monitored the test suites of 15 Python Standard Libraries to create the dataset. TestDossier contains 1,234 distinct argument/variable names and 133,169 distinct values, leading to a total of 12,9M individual values. We envision that our dataset can help developers detect rarely tested values, untested values, and variations of tested values. We also foresee that our dataset can support novel empirical studies in the context of software testing, for example, it can expose the diversity of the tested data.
Andre Hora 0001
MSR1
2024 Availability and Usage of Platform-Specific APIs: A First Empirical Study
abstract
A platform-specific API is an API implemented for a particular platform (e.g., operating system), therefore, it may not work on other platforms than the target one. In this paper, we propose a first empirical study to assess the availability and usage of platform-specific APIs. We analyze the platform-specific APIs provided by the Python Standard Library and mine their usage in 100 popular systems. We find that 21% of the Python Standard Library APIs are platform-specific and that 15% of the modules contain at least one. The platforms with the most availability restrictions are WASI (43.69%), Emscripten (43.64%), Unix (6.76%), and Windows (2.12%). Moreover, we find that platform-specific APIs are largely used in Python. We detect over 19K API usages in all 100 projects, in both production (52.6%) and test code (47.4%). We conclude by discussing practical implications for practitioners and researchers.
Ricardo de Sousa Job, Andre Hora 0001
MSR2
2024 Towards a catalog of composite refactorings
abstract
Abstract Catalogs of refactoring have key importance in software maintenance and evolution, since developers rely on such documents to understand and perform refactoring operations. Furthermore, these catalogs constitute a reference guide for communication between practitioners since they standardize a common refactoring vocabulary. Fowler's book describes the most popular catalog of refactorings, which documents single and well‐known refactoring operations. However, sometimes, refactorings are composite transformations, that is, a sequence of refactorings is performed over a given program element. For example, a sequence of Extract Method operations (a single refactoring) can be performed over the same method, in one or in multiple commits, to simplify its implementation, therefore, leading to a Method Decomposition operation (a composite refactoring). In this paper, we propose and document a catalog with eight composite refactorings. We also implement a set of scripts to mine composite refactorings by preprocessing the results of refactoring detection tools. Using such scripts, we search for composites in a representative refactoring oracle with hundreds of confirmed single refactoring operations. Next, to complement this first study, we also search for composites in the full history of 10 well‐known open‐source projects. We characterize the detected composite refactorings, under dimensions such as size and location. We conclude by addressing the applications and implications of the proposed catalog.
Aline Brito, Andre Hora 0001, Marco Túlio Valente
J. Softw. Evol. Process.2
2023 How do developers collaborate? Investigating GitHub heterogeneous networks
Gabriel P. Oliveira, Ana Flávia C. Moura, Natércia A. Batista, Michele A. Brandão, Andre Hora 0001, Mirella M. Moro
Softw. Qual. J.5
2022 Characterizing High-Quality Test Methods: A First Empirical Study
abstract
To assess the quality of a test suite, one can rely on mutation testing, which computes whether the overall test cases are adequately exercising the covered lines. However, this high level of granularity may overshadow the quality of individual test methods. In this paper, we propose an empirical study to assess the quality of test methods by relying on mutation testing at the method level. We find no major differences between high-quality and low-quality test methods in terms of size, number of asserts, and modifications. In contrast, high-quality test methods are less affected by critical test smells. Finally, we discuss practical implications for researchers and practitioners.
Victor Veloso, Andre Hora 0001
MSR2
2022 How and Why Developers Migrate Python Tests
abstract
Nowadays, Python developers can rely on two major testing frameworks: unittest and pytest. Due to the benefits of pytest (e.g., fixture reuse), several relevant projects in the Python ecosystem have migrated from unittest to pytest. Despite being performed by the Python community, we are not yet aware of how systems are migrated from unittest to pytest nor the major reasons behind the migration. In this paper, we provide the first empirical study to assess testing framework migration. We analyze how and why developers migrate from unittest to pytest. We mine 100 popular Python systems and assess their migration status. We find that 34% of the systems rely on both testing frameworks and that Python projects are moving to pytest. While some systems have fully migrated, others are still migrating after a long period, suggesting that the migration is not always straightforward. Overall, the migrated test code is smaller than the original one. Furthermore, developers migrate to pytest due to several reasons, such as the easier syntax, interoperability, easier maintenance, and fixture flexibility/reuse, however, the implicit mechanics of pytest is a concern. We conclude by discussing practical implications for practitioners and researchers.
Lívia Barbosa, Andre Hora 0001
SANER2
2022 Exploring API Deprecation Evolution in JavaScript
abstract
Building an application using third-party libraries is a common practice in software development. As any other system, software libraries and their APIs evolve. To support version migration and ensure backward compatibility, a recommended practice during development is to deprecate API. Different from other popular programming languages such as Java and C#, JavaScript has no native support to deprecate API elements. However, several strategies are commonly adopted to communicate that an API should be avoided, such as the project documentation, JSDoc annotation, code comment, console message, and deprecation utility. Indeed, there have been many studies on deprecation strategies and evolution mostly on Java, C#, and Python. However, to the best of our knowledge, there are no detailed studies aiming at analyzing how API deprecation changes over time in the JavaScript ecosystem. This paper provides an empirical study on how API deprecation evolves in JavaScript by analyzing 1,918 releases of 50 popular packages. Results show that close to 60% have rising trends in the number of deprecated APIs, while only 9.4% indicate a downward trend. Also, most deprecation occurrences are both added and removed on minor releases instead of removed on major releases, as recommended by best practices.
Romulo Nascimento, Andre Hora 0001, Eduardo Figueiredo 0001
SANER2
2022 How and why we end up with complex methods: a multi-language study
Mateus Lopes, Andre Hora 0001
Empir. Softw. Eng.2
2022 How are framework code samples maintained and used by developers? The case of Android and Spring Boot
Gabriel Menezes, Bruno B. P. Cafeo, Andre Hora 0001
J. Syst. Softw.3
2021 Googling for Software Development: What Developers Search For and What They Find
abstract
Developers often search for software resources on the web. In practice, instead of going directly to websites (e.g., Stack Overflow), they rely on search engines (e.g., Google). Despite this being a common activity, we are not yet aware of what developers search from the perspective of popular software development websites and what search results are returned. With this knowledge, we can understand real-world queries, developers' needs, and the query impact on the search results. In this paper, we provide an empirical study to understand what developers search on the web and what they find. We assess 1.3M queries to popular programming websites and we perform thousands of queries on Google to explore search results. We find that (i) developers' queries typically start with keywords (e.g., Python, Android, etc.), are short (3 words), tend to omit functional words, and are similar among each other; (ii) minor changes to queries do not largely affect the Google search results, however, some cosmetic changes may have a non-negligible impact; and (iii) search results are dominated by Stack Overflow, but YouTube is also a relevant source nowadays. We conclude by presenting detailed implications for researchers and developers.
Andre Hora 0001
MSR1
2021 What Code Is Deliberately Excluded from Test Coverage and Why?
abstract
Test coverage is largely used to assess test effectiveness. In practice, not all code is equally important for coverage analysis, for instance, code that will not be executed during tests is irrelevant and can actually harm the analysis. Some coverage tools provide support for code exclusion from coverage reports, however, we are not yet aware of what code tends to be excluded nor the reasons behind it. This can support the creation of more accurate coverage reports and reveal novel and harmful usage cases. In this paper, we provide the first empirical study to understand code exclusion practices in test coverage. We mine 55 Python projects and assess commit messages and code comments to detect rationales for exclusions. We find that (1) over 1/3 of the projects perform deliberate coverage exclusion; (2) 75% of the code are already created using the exclusion feature, while 25% add it over time; (3) developers exclude non-runnable, debug-only, and defensive code, but also platform-specific and conditional importing; and (4) most code is excluded because it is already untested, low-level, or complex. Finally, we discuss implications to improve coverage analysis and shed light on the existence of biased coverage reports.
Andre Hora 0001
MSR1
2021 Characterizing refactoring graphs in Java and JavaScript projects
Aline Brito, Andre Hora 0001, Marco Túlio Valente
Empir. Softw. Eng.2
2021 Characterizing top ranked code examples in Google
Andre Hora 0001
J. Syst. Softw.1
2021 APISonar: Mining API usage examples
abstract
Abstract Developers spend a significant part of their time searching for code examples on the web. Often, they look for Application Programming Interface (API) usage examples, that is, how to use APIs provided by libraries and frameworks. For this purpose, several programming websites are available. Some programming websites provide manually created examples: unfortunately, as millions of APIs are available nowadays, they do not cover the majority of the APIs. To alleviate this limitation, other programming websites focus on automatically mining API usage examples from code repositories. To the best of our knowledge, however, these solutions are still very limited: they often present poor, duplicated, and similar API usage examples. In this article, we propose an approach, APISonar, to automatically mine API usage examples from code repositories. Our approach aims to overcome the limitations of current solutions: we focus on presenting readable and reusable API usage examples. We analyze millions of source files provided by 4486 software projects hosted on GitHub. Based on this data, we extract 11 million API usage examples about 1.5 million distinct APIs. We evaluate APISonar by assessing its quality and usage. We show that APISonar is a competitive solution, providing the best API examples in terms of readability and reusability, as compared with popular programming websites. Moreover, despite being a novel website, APISonar attracted a significant amount of users in a short period (3.7K users from 119 countries during 5 months). APISonar is available at www.apisonar.com.
Andre Hora 0001
Softw. Pract. Exp.1
2020 Assessing Mock Classes: An Empirical Study
abstract
During testing activities, developers frequently rely on dependencies (e.g., web services, etc) that make the test harder to be implemented. In this scenario, they can use mock objects to emulate the dependencies' behavior, which contributes to make the test fast and isolated. In practice, the emulated dependency can be dynamically created with the support of mocking frameworks or manually hand-coded in mock classes. While the former is well-explored by the research literature, the latter has not yet been studied. Assessing mock classes would provide the basis to better understand how those mocks are created and consumed by developers and to detect novel practices and challenges. In this paper, we provide the first empirical study to assess mock classes. We analyze 12 popular software projects, detect 604 mock classes, and assess their content, design, and usage. We find that mock classes: often emulate domain objects, external dependencies, and web services; are typically part of a hierarchy; are mostly public, but 1/3 are private; and are largely consumed by client projects, particularly to support web testing. Finally, based on our results, we provide implications and insights to researchers and practitioners working with mock classes.
Gustavo Pereira, Andre Hora 0001
ICSME2
2020 Refactoring Graphs: Assessing Refactoring over Time
abstract
Refactoring is an essential activity during software evolution. Frequently, practitioners rely on such transformations to improve source code maintainability and quality. As a consequence, this process may produce new source code entities or change the structure of existing ones. Sometimes, the transformations are atomic, i.e., performed in a single commit. In other cases, they generate sequences of modifications performed over time. To study and reason about refactorings over time, in this paper, we propose a novel concept called refactoring graphs and provide an algorithm to build such graphs. Then, we investigate the history of 10 popular open-source Java-based projects. After eliminating trivial graphs, we characterize a large sample of 1,150 refactoring graphs, providing quantitative data on their size, commits, age, refactoring composition, and developers. We conclude by discussing applications and implications of refactoring graphs, for example, to improve code comprehension, detect refactoring patterns, and support software evolution studies.
Aline Brito, Andre Hora 0001, Marco Túlio Valente
SANER2
2020 JavaScript API Deprecation in the Wild: A First Assessment
abstract
Building an application using third-party libraries is a common practice in software development. As any other software system, code libraries and their APIs evolve over time. In order to help version migration and ensure backward compatibility, a recommended practice during development is to deprecate API. Although studies have been conducted to investigate deprecation in some programming languages, such as Java and C#, there are no detailed studies on API deprecation in the JavaScript ecosystem. This paper provides an initial assessment of API deprecation in JavaScript by analyzing 50 popular software projects. Initial results suggest that the use of deprecation mechanisms in JavaScript packages is low. However, we find five different ways that developers use to deprecate API in the studied projects. Among these solutions, deprecation utility (i.e., any sort of function specially written to aid deprecation) and code comments are the most common practices in JavaScript. Finally, we find that the rate of helpful message is high: 67% of the deprecations have replacement messages to support developers when migrating APIs.
Romulo Nascimento, Aline Brito, Andre Hora 0001, Eduardo Figueiredo 0001
SANER3
2020 You broke my code: understanding the motivations for breaking changes in APIs
Aline Brito, Marco Túlio Valente, Laerte Xavier, Andre Hora 0001
Empir. Softw. Eng.4
2020 Characteristics of method extractions in Java: a large scale empirical study
Andre Hora 0001, Romain Robbes
Empir. Softw. Eng.1
2020 What are the characteristics of popular APIs? A large-scale study on Java, Android, and 165 libraries
Natalia Caroline Lima, Andre Hora 0001
Softw. Qual. J.2
2019 Framework Code Samples: How Are They Maintained and Used by Developers?
abstract
Background: Modern software systems are commonly built on the top of frameworks. To accelerate the learning process of features provided by frameworks, code samples are made available to assist developers. However, we know little about how code samples are actually developed. Aims: In this paper, we aim to fill this gap by assessing the characteristics of framework code samples. We provide insights on how code samples are maintained and used by developers. Method: We analyze 233 code samples of Android and SpringBoot, and assess aspects related to their source code, evolution, popularity, and client usage. Results: We find that most code samples are small and simple, provide a working environment to the clients, and rely on automated build tools. They change frequently over time, for example, to adapt to new framework versions. We also detect that clients commonly fork the code samples, however, they rarely modify them. Conclusions: We provide a set of lessons learned and implications to creators and clients of code samples to improve maintenance and usage activities.
Gabriel Menezes, Bruno B. P. Cafeo, Andre Hora 0001
ESEM3
2019 Measuring and analyzing code authorship in 1 + 118 open source projects
Guilherme Avelino 0001, Leonardo Teixeira Passos, Andre Hora 0001, Marco Túlio Valente
Sci. Comput. Program.3
2018 Assessing the threat of untracked changes in software evolution
abstract
While refactoring is extensively performed by practitioners, many Mining Software Repositories (MSR) approaches do not detect nor keep track of refactorings when performing source code evolution analysis. In the best case, keeping track of refactorings could be unnecessary work; in the worst case, these untracked changes could significantly affect the performance of MSR approaches. Since the extent of the threat is unknown, the goal of this paper is to assess whether it is significant. Based on an extensive empirical study, we answer positively: we found that between 10 and 21% of changes at the method level in 15 large Java systems are untracked. This results in a large proportion (25%) of entities that may have their histories split by these changes, and a measurable effect on at least two MSR approaches. We conclude that handling untracked changes should be systematically considered by MSR studies.
Andre Hora 0001, Danilo Silva 0002, Marco Túlio Valente, Romain Robbes
ICSE1
2018 Why and how Java developers break APIs
abstract
Modern software development depends on APIs to reuse code and increase productivity. As most software systems, these libraries and frameworks also evolve, which may break existing clients. However, the main reasons to introduce breaking changes in APIs are unclear. Therefore, in this paper, we report the results of an almost 4-month long field study with the developers of 400 popular Java libraries and frameworks. We configured an infrastructure to observe all changes in these libraries and to detect breaking changes shortly after their introduction in the code. After identifying breaking changes, we asked the developers to explain the reasons behind their decision to change the APIs. During the study, we identified 59 breaking changes, confirmed by the developers of 19 projects. By analyzing the developers' answers, we report that breaking changes are mostly motivated by the need to implement new features, by the desire to make the APIs simpler and with fewer elements, and to improve maintainability. We conclude by providing suggestions to language designers, tool builders, software engineering researchers and API developers.
Aline Brito, Laerte Xavier, Andre Hora 0001, Marco Túlio Valente
SANER3
2018 APIDiff: Detecting API breaking changes
abstract
Libraries are commonly used to increase productivity. As most software systems, they evolve over time and changes are required. However, this process may involve breaking compatibility with previous versions, leading clients to fail. In this context, it is important that libraries creators and clients frequently assess API stability in order to better support their maintenance practices. In this paper, we introduce APIDIFF, a tool to identify API breaking and non-breaking changes between two versions of a Java library. The tool detects changes on three API elements: types, methods, and fields. We also report usage scenarios of APIDIFF with four real-world Java libraries.
Aline Brito, Laerte Xavier, Andre Hora 0001, Marco Túlio Valente
SANER3
2018 On the use of replacement messages in API deprecation: An empirical study
Gleison Brito, Andre Hora 0001, Marco Túlio Valente, Romain Robbes
J. Syst. Softw.2
2018 How do developers react to API evolution? A large-scale empirical study
Andre Hora 0001, Romain Robbes, Marco Túlio Valente, Nicolas Anquetil, Anne Etien, Stéphane Ducasse
Softw. Qual. J.1
2017 Historical and impact analysis of API breaking changes: A large-scale study
abstract
Change is a routine in software development. Like any system, libraries also evolve over time. As a consequence, clients are compelled to update and, thus, benefit from the available API improvements. However, some of these API changes may break contracts previously established, resulting in compilation errors and behavioral changes. In this paper, we study a set of questions regarding API breaking changes. Our goal is to measure the amount of breaking changes on real-world libraries and its impact on clients at a large-scale level. We assess (i) the frequency of breaking changes, (ii) the behavior of these changes over time, (iii) the impact on clients, and (iv) the characteristics of libraries with high frequency of breaking changes. Our large-scale analysis on 317 real-world Java libraries, 9K releases, and 260K client applications shows that (i) 14.78% of the API changes break compatibility with previous versions, (ii) the frequency of breaking changes increases over time, (iii) 2.54% of their clients are impacted, and (iv) systems with higher frequency of breaking changes are larger, more popular, and more active. Based on these results, we provide a set of lessons to better support library and client developers in their maintenance tasks.
Laerte Xavier, Aline Brito, Andre Hora 0001, Marco Túlio Valente
SANER3
2017 Why do we break APIs? First answers from developers
abstract
Breaking contracts have a major impact on API clients. Despite this fact, recent studies show that libraries are often backward incompatible and that the rate of breaking changes increase over time. However, the specific reasons that motivate library developers to break contracts with their clients are still unclear. In this paper, we describe a qualitative study with library developers and real instance of API breaking changes. Our goal is to (i) elicit the reasons why developers introduce breaking changes; and (ii) check if they are aware about the risks of such changes. Our survey with the top contributors of popular Java libraries contributes to reveal a list of five reasons why developers break API contracts. Moreover, it also shows that most of developers are aware of these risks and, in some cases, adopt strategies to mitigate them. We conclude by prospecting a future study to strengthen our current findings. With this study, we expect to contribute on delineating tools to better assess the risks and impacts of API breaking changes.
Laerte Xavier, Andre Hora 0001, Marco Túlio Valente
SANER2
2016 Understanding the Factors That Impact the Popularity of GitHub Repositories
abstract
Software popularity is a valuable information to modern open source developers, who constantly want to know if their systems are attracting new users, if new releases are gaining acceptance, or if they are meeting user's expectations. In this paper, we describe a study on the popularity of software systems hosted at GitHub, which is the world's largest collection of open source software. GitHub provides an explicit way for users to manifest their satisfaction with a hosted repository: the stargazers button. In our study, we reveal the main factors that impact the number of stars of GitHub projects, including programming language and application domain. We also study the impact of new features on project popularity. Finally, we identify four main patterns of popularity growth, which are derived after clustering the time series representing the number of stars of 2,279 popular GitHub repositories. We hope our results provide valuable insights to developers and maintainers, which could help them on building and evolving systems in a competitive software market.
Hudson Borges, Andre Hora 0001, Marco Túlio Valente
ICSME2
2016 A novel approach for estimating Truck Factors
abstract
Truck Factor (TF) is a metric proposed by the agile community as a tool to identify concentration of knowledge in software development environments. It states the minimal number of developers that have to be hit by a truck (or quit) before a project is incapacitated. In other words, TF helps to measure how prepared is a project to deal with developer turnover. Despite its clear relevance, few studies explore this metric. Altogether there is no consensus about how to calculate it, and no supporting evidence backing estimates for systems in the wild. To mitigate both issues, we propose a novel (and automated) approach for estimating TF-values, which we execute against a corpus of 133 popular project in GitHub. We later survey developers as a means to assess the reliability of our results. Among others, we find that the majority of our target systems (65%) have TF ≤ 2. Surveying developers from 67 target systems provides confidence towards our estimates; in 84% of the valid answers we collect, developers agree or partially agree that the TF's authors are the main authors of their systems; in 53% we receive a positive or partially positive answer regarding our estimated truck factors.
Guilherme Avelino 0001, Leonardo Teixeira Passos, Andre Hora 0001, Marco Túlio Valente
ICPC3
2016 When should internal interfaces be promoted to public?
abstract
Commonly, software systems have public (and stable) interfaces, and internal (and possibly unstable) interfaces. Despite being discouraged, client developers often use internal interfaces, which may cause their systems to fail when they evolve. To overcome this problem, API producers may promote internal interfaces to public. In practice, however, API producers have no assistance to identify public interface candidates. In this paper, we study the transition from internal to public interfaces. We aim to help API producers to deliver a better product and API clients to benefit sooner from public interfaces. Our empirical investigation on five widely adopted Java systems present the following observations. First, we identified 195 promotions from 2,722 internal interfaces. Second, we found that promoted internal interfaces have more clients. Third, we predicted internal interface promotion with precision between 50%-80%, recall 26%-82%, and AUC 74%-85%. Finally, by applying our predictor on the last version of the analyzed systems, we automatically detected 382 public interface candidates.
Andre Hora 0001, Marco Túlio Valente, Romain Robbes, Nicolas Anquetil
SIGSOFT FSE1
2016 Do Developers Deprecate APIs with Replacement Messages? A Large-Scale Analysis on Java Systems
abstract
As any other software system, frameworks and libraries evolve over time, and so their APIs. Consequently, client systems should be updated to benefit from improved APIs. To facilitate this task and preserve backward compatibility, API elements should always be deprecated with clear replacement messages. However, in practice, there are evidences that API elements are usually deprecated without such messages. In this paper, we study a set of questions regarding the adoption of deprecation messages. Our goal is twofold: to measure the usage of deprecation messages and to investigate whether a tool is needed to recommend such messages. Thus, we verify (i) the frequency of deprecated elements with replacement messages, (ii) the impact of software evolution on such frequency, and (iii) the characteristics of systems which deprecate API elements in a correct way. Our large-scale analysis on 661 real-world Java systems shows that (i) 64% of the API elements are deprecated with replacement messages per system, (ii) there is almost no major effort to improve deprecation messages over time, and (iii) systems that deprecated API elements in a correct way are statistically significantly different from the ones that do not in terms of size and developing community. As a result, we provide the basis for the design of a tool to support client developers on detecting missing deprecation messages.
Gleison Brito, Andre Hora 0001, Marco Túlio Valente, Romain Robbes
SANER2
2016 Identifying Utility Functions Using Random Forests
abstract
Utility functions are general purpose functions, which are useful in many parts of a system. To facilitate reuse, they are usually implemented in specific libraries. However, developers frequently miss opportunities to implement general-purpose functions in utility libraries, which decreases the chances of reuse. In this paper, we describe our ongoing investigation on using Random Forest classifiers to automatically identify utility functions. Using a list of static source code metrics we train a classifier to identify such functions, both in Java (using 84 projects from the Qualitas Corpus) and in JavaScript (using 22 popular projects from GitHub). We achieve the following median results for Java: 0.90 (AUC), 0.83 (precision), 0.88 (recall), and 0.84 (F-measure). For JavaScript, the median results are 0.80 (AUC), 0.75 (precision), 0.89 (recall), and 0.76 (F-measure).
Tamara Mendes, Marco Túlio Valente, Andre Hora 0001, Alexander Serebrenik
SANER3
2016 Mining architectural violations from version history
Cristiano Amaral Maffort, Marco Túlio Valente, Ricardo Terra, Mariza Andrade da Silva Bigonha, Nicolas Anquetil, Andre Hora 0001
Empir. Softw. Eng.6
2015 How do developers react to API evolution? The Pharo ecosystem case
abstract
Software engineering research now considers that no system is an island, but it is part of an ecosystem involving other systems, developers, users, hardware, ... When one system (e.g., a framework) evolves, its clients often need to adapt. Client developers might need to adapt to functionalities, client systems might need to be adapted to a new API, client users might need to adapt to a new User Interface. The consequences of such changes are yet unclear, what proportion of the ecosystem might be expected to react, how long might it take for a change to diffuse in the ecosystem, do all clients react in the same way? This paper reports on an exploratory study aimed at observing API evolution and its impact on a large-scale software ecosystem, Pharo, which has about 3,600 distinct systems, more than 2,800 contributors, and six years of evolution. We analyze 118 API changes and answer research questions regarding the magnitude, duration, extension, and consistency of such changes in the ecosystem. The results of this study help to characterize the impact of API evolution in large software ecosystems, and provide the basis to better understand how such impact can be alleviated.
Andre Hora 0001, Romain Robbes, Nicolas Anquetil, Anne Etien, Stéphane Ducasse, Marco Túlio Valente
ICSME1
2015 Apiwave: Keeping track of API popularity and migration
abstract
Every day new frameworks and libraries are created and existing ones evolve. To benefit from such newer or improved APIs, client developers should update their applications. In practice, this process presents some challenges: APIs are commonly backward-incompatible (causing client applications to fail when updating) and multiple APIs are available (making it difficult to decide which one to use). To address these challenges, we propose apiwave, a tool that keeps track of API popularity and migration of major frameworks/libraries. The current version includes data about the evolution of top 650 GitHub Java projects, from which 320K APIs were extracted. We also report an experience using apiwave on real-world scenarios.
Andre Hora 0001, Marco Túlio Valente
ICSME1
2015 Automatic detection of system-specific conventions unknown to developers
Andre Hora 0001, Nicolas Anquetil, Anne Etien, Stéphane Ducasse, Marco Túlio Valente
J. Syst. Softw.1
2013 Mining Architectural Patterns Using Association Rules
Cristiano Amaral Maffort, Marco Túlio Valente, Roberto da Silva Bigonha, Andre Hora 0001, Nicolas Anquetil, Jonata Menezes
SEKE4
2012 Domain specific warnings: Are they any better?
abstract
Tools to detect coding standard violations in source code are commonly used to improve code quality. One of their original goals is to prevent bugs, yet, a high number of false positives is generated by the rules of these tools, i.e., most warnings do not indicate real bugs. There are empirical evidences supporting the intuition that the rules enforced by such tools do not prevent the introduction of bugs in software. This may occur because the rules are too generic and do not focus on domain specific problems of the software under analysis. We underwent an investigation of rules created for a specific domain based on expert opinion to understand if such rules are worthwhile enforcing in the context of defect prevention. In this paper, we performed a systematic study to investigate the relation between generic and domain specific warnings and observed defects. From our experiment on a real case, long term evolution, software, we have found that domain specific rules provide better defect prevention than generic ones.
Andre Hora 0001, Nicolas Anquetil, Stéphane Ducasse, Simon Allier
ICSM1
2008 Using easy API to develop multimedia applications for maemo platform
abstract
This paper introduces the Easy API as a tool to ease the development of multimedia applications that run into mobile devices. Easy is implemented for maemo, a Linux-based open source development platform for Internet Tablets. In this paper we describe how to use its internal multimedia modules such as Audio, Camera and Radio, providing useful examples implemented using Easy API. We conclude this paper with a comparison between Easy's approach and other similar mechanisms largely used to create multimedia applications for the maemo platform.
Andre Hora 0001, Miguel A. de S. Falcão, Emanuel Dantas Filho
EATIS1