VLDB 2026 Research / reviewers in the wild / expert
Muhammad Firhard Roslan
dblp:333/5376
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0002-7177-1702ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Do Automatic Test Generation Tools Generate Flaky Tests?abstractNon-deterministic test behavior, or flakiness, is common and dreaded among developers. Researchers have studied the issue and proposed approaches to mitigate it. However, the vast majority of previous work has only considered developer-written tests. The prevalence and nature of flaky tests produced by test generation tools remain largely unknown. We ask whether such tools also produce flaky tests and how these differ from developer-written ones. Furthermore, we evaluate mechanisms that suppress flaky test generation. We sample 6 356 projects written in Java or Python. For each project, we generate tests using EvoSuite (Java) and Pynguin (Python), and execute each test 200 times, looking for inconsistent outcomes. Our results show that flakiness is at least as common in generated tests as in developer-written tests. Nevertheless, existing flakiness suppression mechanisms implemented in EvoSuite are effective in alleviating this issue (71.7 % fewer flaky tests). Compared to developer-written flaky tests, the causes of generated flaky tests are distributed differently. Their non-deterministic behavior is more frequently caused by randomness, rather than by networking and concurrency. Using flakiness suppression, the remaining flaky tests differ significantly from any flakiness previously reported, where most are attributable to runtime optimizations and EvoSuite-internal resource thresholds. These insights, with the accompanying dataset, can help maintainers to improve test generation tools, give recommendations for developers using these tools, and serve as a foundation for future research in test flakiness or test generation. Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck, Phil McMinn, Gordon Fraser 0001 |
ICSE | 2 |
| 2024 | Private-Keep Out? Understanding How Developers Account for Code Visibility in Unit TestingabstractRegression test maintenance costs can be reduced by striving to write tests that will require as few changes as possible in the future. Writing unit tests against behavior, as opposed to implementation, is one way to try to achieve this, because as long as the public API remains constant, units can be safely refactored without the need to also change the tests. However, in a study on 4,801 open-source Java projects reported in this paper, we found that 28% of projects contradict this advice, with tests that side-step the public API by directly calling non-public methods. We investigated why developers do not solely test public APIs-potentially increasing future test maintenance costs-by surveying 73 developers and conducting a systematic review of 60 StackOverflow posts dating from 2008–2023. Through numerical and thematic analyses, we uncover several findings, including (1) developers are disunited on whether to test only through public APIs or not; (2) those in favor of only testing through the public API tend to be more experienced and believe the need or desire to break with this is borne out of poor software design; while (3) those that test non-public methods directly are concerned about untested code complexity and overly intricate tests. Our findings provide multiple implications for future work, including automated developer support in the form of automated non-public method sequence replacement, and automated refactoring of production code using problematic public API-avoiding tests. Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
ICSME | 1 |
| 2024 | Viscount: A Direct Method Call Coverage Tool for JavaabstractWriting unit tests against implementation detail in production code, often embodied in non-public methods, is considered bad practice in formal and gray literature. This is because it leads to fragile tests that break easily when underlying implementation details change. For this reason, tests that focus on behavior are encouraged. One way to achieve this is to test units exclusively through their public API. However, our recent developer survey shows that this advice is not always followed in practice. Moreover, code coverage tools do not provide a way to determine which methods were called directly from tests, meaning there is no easy way to identify whether units make calls to non-public methods, other than through manual examination. To address this problem, we developed Viscount, a tool that can determine direct method call coverage for Java tests written in JUnit. Viscount reports the percentage of methods invoked directly from tests, according to their visibility - i.e., public or non-public (protected, package-private, or private). This can help developers and researchers identify tests that potentially need to be refactored or rewritten. In this paper, we describe Viscount's overall architecture, its core features, and how to use it. Viscount is also publicly available on GitHub: https://github.com/unittesting-nonpublic/viscount. A demo video of Viscount is available at: https://youtu.be/ZUyRtiUnbsU. Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
ICSME | 1 |
| 2022 | An Empirical Comparison of EvoSuite and DSpot for Improving Developer-Written Test Suites with Respect to Mutation Score
Muhammad Firhard Roslan, José Miguel Rojas, Phil McMinn |
SSBSE | 1 |