Jeanderson Cândido

dblp:207/7152 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0003-0846-040XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Software testing · 50% Empirical software engineering · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Empirical software engineering
mining software repositories
0.312017
Test suite parallelization in open-source projects: a study on its usage and impact · ASE 2017
Empirical software engineering › mining software repositories
open-source project analysis
0.312017
Test suite parallelization in open-source projects: a study on its usage and impact · ASE 2017
Software testing › test execution
parallel testing
0.312017
Test suite parallelization in open-source projects: a study on its usage and impact · ASE 2017
Software testing
test execution
0.312017
Test suite parallelization in open-source projects: a study on its usage and impact · ASE 2017

Methods — techniques the papers use, named apart from their topics

empirical study · 0.3
YearPublicationVenuePosition
2021 An Exploratory Study of Log Placement Recommendation in an Enterprise System
abstract
Logging is a development practice that plays an important role in the operations and monitoring of complex systems. Developers place log statements in the source code and use log data to understand how the system behaves in production. Unfortunately, anticipating where to log during development is challenging. Previous studies show the feasibility of leveraging machine learning to recommend log placement despite the data imbalance since logging is a fraction of the overall code base. However, it remains unknown how those techniques apply to an industry setting, and little is known about the effect of imbalanced data and sampling techniques. In this paper, we study the log placement problem in the code base of Adyen, a large-scale payment company. We analyze 34,526 Java files and 309,527 methods that sum up +2M SLOC. We systematically measure the effectiveness of five models based on code metrics, explore the effect of sampling techniques, understand which features models consider to be relevant for the prediction, and evaluate whether we can exploit 388,086 methods from 29 Apache projects to learn where to log in an industry setting. Our best performing model achieves 79% of balanced accuracy, 81% of precision, 60% of recall. While sampling techniques improve recall, they penalize precision at a prohibitive cost. Experiments with open-source data yield under-performing models over Adyen's test set; nevertheless, they are useful due to their low rate of false positives. Our supporting scripts and tools are available to the community.
Jeanderson Cândido, Jan Haesen, Mauricio Finavaro Aniche, Arie van Deursen
MSR1
2020 Practical detection of CMS plugin conflicts in large plugin sets
Igor Lima, Jeanderson Cândido, Marcelo d'Amorim
Inf. Softw. Technol.2
2017 Test suite parallelization in open-source projects: a study on its usage and impact
abstract
Dealing with high testing costs remains an important problem in Software Engineering. Test suite parallelization is an important approach to address this problem. This paper reports our findings on the usage and impact of test suite parallelization in open-source projects. It provides recommendations to practitioners and tool developers to speed up test execution. Considering a set of 468 popular Java projects we analyzed, we found that 24% of the projects contain costly test suites but parallelization features still seem underutilized in practice - only 19.1% of costly projects use parallelization. The main reported reason for adoption resistance was the concern to deal with concurrency issues. Results suggest that, on average, developers prefer high predictability than high performance in running tests.
Jeanderson Cândido, Luis Melo, Marcelo d'Amorim
ASE1