VLDB 2026 Research / reviewers in the wild / expert
Plínio de Sá Leitão Júnior
dblp:134/0939
· DBLP profile ↗
11ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empirical Evaluation of the uSV Ratio Metric for Fault Localization
Willian de Jesus Ferreira, Plínio de Sá Leitão Júnior, Thamer H. Nascimento, Deuslirio Silva-Junior, Rachel Harrison |
COMPSAC | 2 |
| 2026 | Metric-Driven Analysis of SMOTE Efficacy in Colorectal Metastasis Prediction
Áurea Valéria Pereira Silva, Danilo Zuccati De Oliveira, Juliana Paula Felix, Plínio de Sá Leitão Júnior |
COMPSAC | 4 |
| 2025 | A Comparative Study of Data Balancing Techniques for Predicting Metastases in Colorectal Cancer Using the SEER DatabaseabstractPredicting liver and/or lung metastases in colorectal cancer (CRC) patients remains a critical challenge, especially due to the strong class imbalance commonly found in clinical datasets such as SEER (Surveillance, Epidemiology, and End Results Program). To address this issue, this study presents a comparative analysis of eight data balancing techniques combined with seven machine learning algorithms for metastasis prediction using SEER data. The evaluated techniques include traditional SMOTE, ADASYN, Borderline-SMOTE, SMOTE-LOF, Radius-SMOTE, RDSMOTE, and CUSS, along with a baseline configuration without any balancing. A total of 53,463 CRC patients were analyzed, and model performance was assessed using F1-score and AUC metrics under five-fold stratified cross-validation. Among the techniques, SMOTE-LOF achieved the best overall results, particularly when combined with XGBoost, reaching an F1-score of 0.6642 and AUC of 0.8988. The results indicate that combining oversampling with noise filtering or structural awareness significantly enhances model sensitivity. This study highlights the importance of selecting appropriate resampling techniques for clinical prediction tasks and offers insights for improving the robustness and fairness of machine learning applications in oncology. Áurea Valéria Pereira Silva, Plínio de Sá Leitão Júnior, Juliana Paula Felix |
COMPSAC | 2 |
| 2025 | Evolutionary Fault Localization Based on the Diversity of Suspiciousness ValuesabstractContext.Fault localization (FL) is a software lifecycle activity and its automation is a challenge for researchers and practitioners.Method.The study focuses on evolutionary fault localization and introduces a novel Genetic Programming (GP) approach that evolves FL heuristics based on the diversity of the suspiciousness score of program statements -a score to grade how faulty a statement is.Experimental analysis.The approach was evaluated against baselines, which include the canonical GP, in benchmarks with real programs and real faults.Conclusion.The results showed the competitiveness of the approach through evaluation metrics commonly used in the research field. Willian de Jesus Ferreira, Plínio de Sá Leitão Júnior, Deuslirio Silva-Junior, Rachel Harrison |
ESANN | 2 |
| 2020 | Search-based fault localisation: A systematic mapping study
Plínio de Sá Leitão Júnior, Diogo M. De-Freitas, Silvia Regina Vergilio, Celso G. Camilo-Junior, Rachel Harrison |
Inf. Softw. Technol. | 1 |
| 2020 | From keywords to relational database content: A semantic mapping method
Mariana S. Ramada, João Carlos da Silva, Plínio de Sá Leitão Júnior |
Inf. Syst. | 3 |
| 2018 | Mutation-Based Evolutionary Fault LocalisationabstractFault localisation is an expensive and time-consuming stage of software maintenance. Research is continuing to develop new techniques to automate the process of reducing the effort needed for fault localisation without losing quality. For instance, spectrum-based techniques use execution information from testing to formulate measures for ranking a list of suspicious code locations at which the program may be defective: the suspiciousness formulae mainly combine variables related to code coverage and test results (pass or fail). Moreover previous research has evaluated mutation analysis data (mutation spectra) instead of coverage traces, to yield promising results. This paper reports on a Genetic Programming (GP) solution for the fault localisation problem together with a set of experiments to evaluate the GP solution with respect to baselines and benchmarks. The innovative aspects are the joint investigation of: (i) specialisation of suspiciousness formulae for certain contexts; (ii) the application of mutation spectra to GP-evolved formulae, i.e. signals other than program coverage; (iii) a comparison of the effectiveness of coverage spectra and mutation spectra in the context of evolutionary approaches; and (iv) an analysis of the mutation spectra quality. The results show the competitiveness of GP-evolved mutation spectra heuristics over coverage traces as well as over a number of baselines, and suggest that the quality of mutation-related variables increases the effectiveness of fault localisation heuristics. Diogo M. De-Freitas, Plínio de Sá Leitão Júnior, Celso G. Camilo-Junior, Rachel Harrison |
CEC | 2 |
| 2018 | Evolutionary Composition of Customized Fault Localization Heuristics
Diogo M. De-Freitas, Plínio de Sá Leitão Júnior, Celso G. Camilo-Junior, Rachel Harrison |
ESANN | 2 |
| 2013 | Shrinking a database to perform SQL mutation tests using an evolutionary algorithmabstractThis paper tries to combine SQL mutation testing techniques with evolutionary computation aiming to improve the test data to SQL instructions. Based on a heuristic perspective it presents an approach that uses Genetic Algorithms (GA) to select tuples from an original database trying to reduce this one in an effective data set. The goal is to find a reduced data set which is able to detect a large number of faults in the SQL instructions of a given application. During the evolutionary process, the analysis of mutants is used to assess each set of data test selected by GA. The results obtained from experiments reveal a good performance using GA metaheuristic. Ana Claudia Bastos Loureiro Monção, Celso G. Camilo-Junior, Leonardo T. Queiroz, Cássio L. Rodrigues, Plínio de Sá Leitão Júnior, Auri M. R. Vincenzi |
IEEE Congress on Evolutionary Computation | 5 |
| 2009 | Mutation Analysis for SQL Database ApplicationsabstractTesting database applications is crucial for ensuring good quality software as undetected faults can result in unrecoverable data corruption. SQL is the most widely used interface language for relational database systems. Our approach aims to achieve better testing by selecting fault revealing databases. We propose the use of mutation analysis on SQL statements and discuss two scenarios for applying strong and weak mutation techniques. Experiments using real applications, real faults and real data were performed to: (i) evaluate the applicability of the approach, and (ii) compare fault-revealing abilities of input databases. Andrea Goncalves Cabeca, Mário Jino, Plínio de Sá Leitão Júnior |
ICSEA | 3 |
| 2008 | Data Flow Testing of SQL-Based Active Database ApplicationsabstractThe relevance of reactive capabilities as a unifying paradigm for handling a number of database features and applications is well-established. Active database systems have been used to implement the persistent data requirements of applications on several knowledge domains. They extend passive ones by automatically performing predefined actions in response to events that they monitor. These reactive abilities are generally expressed with active rules defined within the database itself. We investigate the use of data flow-based testing to identify the presence of faults in active rules written in SQL. The goal is to improve reliability and overall quality in this realm. Our contribution is the definition of a family of adequacy criteria, which require the coverage of inter-rule persistent data flow associations, and its effectiveness in various data flow analysis precisions. Both theoretical and empirical investigations show that the criteria have strong fault detecting ability at a polynomial complexity cost. Plínio de Sá Leitão Júnior, Plínio R. S. Vilela, Mário Jino |
ICSEA | 1 |