VLDB 2026 Research / reviewers in the wild / expert
George G. Cabral
dblp:35/5737 · also George Gomes Cabral
· DBLP profile ↗
18ranked-venue papers
11as first author
5since 2021 · last 2025
0000-0003-2831-4274ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-authorSoftware engineering, systems software and programming languages · 5 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Offline and Continual Just-in-Time Software Defect Prediction with Pre-trained Language ModelsabstractJust-in-time Software Defect Prediction (JIT-SDP) aims to detect potential defects early, helping to prevent risky code from entering the repository during development. This study evaluates JIT-SDP using pre-trained language models in different architectures and settings. It compares open-source fine-tuned models such as CodeT5+ and UniXCoder with closed LLMs such as GPT and Gemini. This is the first known study to compare trainable open and prompt-based closed decoder-only models for JIT-SDP. The main results show that fine-tuned open models outperform closed models in zero-shot and few-shot scenarios without advanced prompt engineering techniques, and in cross-project tasks, CodeT5+ and UniXCoder surpass previous state-of-the-art results. The findings underscore the value of model architecture, fine-tuning, and expert features for effective defect prediction. Finally, we introduce CodeFlowLM – to our knowledge, the first framework for continual JIT-SDP using pre-trained language models. Monique Louise Monteiro, George G. Cabral, Adriano Lorena Inácio de Oliveira |
SMC | 2 |
| 2025 | Semantic SZZ: Mitigating the Impact of Misclassified Corrective Changes in Just-in-Time Software Defect PredictionabstractIn the evolving landscape of software engineering, accurate identification of defect-inducing commits is critical to improving software quality and reducing development costs. This paper revisits the widely adopted SZZ algorithm, which is utilized for labeling commits as clean or defect-inducing, to address one of its main limitations, i.e., its reliance on outdated corrective commits identification strategies. We propose an innovative approach that integrates the semantic understanding capability of the GPT OpenAI model into the SZZ flow to better interpret commit messages. Our experiments reveal, for some projects, a large number of commits incorrectly interpreted as defect-fixing, consequently, leading to the misclassification of commits as defect-inducing. As an example, for the Postgresql dataset, the number of defect-inducing commits was reduced in 21% when compared to the original SZZ. Furthermore, results of our experiments strongly suggest that, as a result of the proposed SZZ labeling process, the JIT-SDP problem has been shown to be more challenging than originally reported by previous works. Ronaldo C. Veras, George G. Cabral, Adriano Lorena Inácio de Oliveira |
SMC | 2 |
| 2024 | Correction to: An investigation of online and offline learning models for online just-in-time software defect predictionabstractWhile the University of Birmingham exercises care and attention in making items available there are rare occasions when an item has been uploaded in error or has been deemed to be commercially or otherwise sensitive.If you believe that this is the case for this document, please contact [email protected] providing details and we will remove George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira, Dinaldo A. Pessoa, Sadia Tabassum |
Empir. Softw. Eng. | 1 |
| 2023 | An investigation of online and offline learning models for online Just-in-Time Software Defect PredictionabstractAbstract Just-in-Time Software Defect Prediction (JIT-SDP) operates in an online scenario where additional training data is received over time. Existing online JIT-SDP studies used online Oza ensemble learning methods with Hoeffding Trees as base learners to learn and update JIT-SDP models over time in this scenario. However, it is unknown how these approaches compare against offline learning approaches adapted to operate in online scenarios, and how the use of any other online or offline base learners would affect online JIT-SDP in terms of predictive performance and computational cost. We therefore propose a new approach called Batch Oversampling Rate Boosting (BORB) that is able to use offline base learners in an online JIT-SDP scenario. Based on 10 open source projects, we provide a comprehensive evaluation of BORB with 5 different base learners and the existing online approach Oversampling Rate Boosting with 4 different base learners, both in within-project and cross-project online JIT-SDP scenarios. The results show that offline learning can lead to better predictive performance than the top performing online learning approaches considered in our study, at a higher computational cost. Cross-project data was helpful to improve predictive performance both for offline and online learning, but especially for online learning. George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira, Dinaldo A. Pessoa, Sadia Tabassum |
Empir. Softw. Eng. | 1 |
| 2023 | Towards Reliable Online Just-in-Time Software Defect PredictionabstractThroughout its development period, a software project experiences different phases, comprises modules with different complexities and is touched by many different developers. Hence, it is natural that problems such as Just-in-Time Software Defect Prediction (JIT-SDP) are affected by changes in the defect generating process (concept drifts), potentially hindering predictive performance. JIT-SDP also suffers from delays in receiving the labels of training examples (verification latency), potentially exacerbating the challenges posed by concept drift and further hindering predictive performance. However, little is known about what types of concept drift affect JIT-SDP and how they affect JIT-SDP classifiers in view of verification latency. This work performs the first detailed analysis of that. Among others, it reveals that different types of concept drift together with verification latency significantly impair the stability of the predictive performance of existing JIT-SDP approaches, drastically affecting their reliability over time. Based on the findings, a new JIT-SDP approach is proposed, aimed at providing higher and more stable predictive performance (i.e., reliable) over time. Experiments based on ten GitHub open source projects show that our approach was capable of produce significantly more stable predictive performances in all investigated datasets while maintaining or improving the predictive performance obtained by state-of-art methods. George G. Cabral, Leandro L. Minku |
IEEE Trans. Software Eng. | 1 |
| 2020 | An investigation of cross-project learning in online just-in-time software defect predictionabstractJust-In-Time Software Defect Prediction (JIT-SDP) is concerned with predicting whether software changes are defect-inducing or clean based on machine learning classifiers. Building such classifiers requires a sufficient amount of training data that is not available at the beginning of a software project. Cross-Project (CP) JIT-SDP can overcome this issue by using data from other projects to build the classifier, achieving similar (not better) predictive performance to classifiers trained on Within-Project (WP) data. However, such approaches have never been investigated in realistic online learning scenarios, where WP software changes arrive continuously over time and can be used to update the classifiers. It is unknown to what extent CP data can be helpful in such situation. In particular, it is unknown whether CP data are only useful during the very initial phase of the project when there is little WP data, or whether they could be helpful for extended periods of time. This work thus provides the first investigation of when and to what extent CP data are useful for JIT-SDP in a realistic online learning scenario. For that, we develop three different CP JIT-SDP approaches that can operate in online mode and be updated with both incoming CP and WP training examples over time. We also collect 2048 commits from three software repositories being developed by a software company over the course of 9 to 10 months, and use 19,8468 commits from 10 active open source GitHub projects being developed over the course of 6 to 14 years. The study shows that training classifiers with incoming CP+WP data can lead to improvements in G-mean of up to 53.90% compared to classifiers using only WP data at the initial stage of the projects. For the open source projects, which have been running for longer periods of time, using CP data to supplement WP data also helped the classifiers to reduce or prevent large drops in predictive performance that may occur over time, leading to up to around 40% better G-Mean during such periods. Such use of CP data was shown to be beneficial even after a large number of WP data were received, leading to overall G-means up to 18.5% better than those of WP classifiers. Sadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral, Liyan Song |
ICSE | 4 |
| 2019 | Class imbalance evolution and verification latency in just-in-time software defect predictionabstractJust-in-Time Software Defect Prediction (JIT-SDP) is an SDP approach that makes defect predictions at the software change level. Most existing JIT-SDP work assumes that the characteristics of the problem remain the same over time. However, JIT-SDP may suffer from class imbalance evolution. Specifically, the imbalance status of the problem (i.e., how much underrepresented the defect-inducing changes are) may be intensified or reduced over time. If occurring, this could render existing JIT-SDP approaches unsuitable, including those that re-build classifiers over time using only recent data. This work thus provides the first investigation of whether class imbalance evolution poses a threat to JIT-SDP. This investigation is performed in a realistic scenario by taking into account verification latency -- the often overlooked fact that labeled training examples arrive with a delay. Based on 10 GitHub projects, we show that JIT-SDP suffers from class imbalance evolution, significantly hindering the predictive performance of existing JIT-SDP approaches. Compared to state-of-the-art class imbalance evolution learning approaches, the predictive performance of JIT-SDP approaches was up to 97.2% lower in terms of g-mean. Hence, it is essential to tackle class imbalance evolution in JIT-SDP. We then propose a novel class imbalance evolution approach for the specific context of JIT-SDP. While maintaining top ranked g-means, this approach managed to produce up to 63.59% more balanced recalls on the defect-inducing and clean classes than state-of-the-art class imbalance evolution approaches. We thus recommend it to avoid overemphasizing one class over the other in JIT-SDP. George G. Cabral, Leandro L. Minku, Emad Shihab, Suhaib Mujahid |
ICSE | 1 |
| 2017 | Time Series Forecasting in the Presence of Concept Drift: A PSO-based ApproachabstractTime series forecasting is a problem with many applications. However, in many domains, such as stock market, the underlying generating process of the time series observations may change, making forecasting models obsolete. This problem is known as Concept Drift. Approaches for time series forecasting should be able to detect and react to concept drift in a timely manner, so that the forecasting model can be updated as soon as possible. Despite the fact that the concept drift problem is well investigated in the literature, little effort has been made to solve this problem for time series forecasting so far. This work proposes two novel methods for dealing with the time series forecasting problem in the presence of concept drift. The proposed methods benefit from the Particle Swarm Optimization (PSO) technique to detect and react to concept drifts in the time series data stream. It is expected that the use of collective intelligence of PSO makes the proposed method more robust to false positive drift detections while maintaining a low error rate on the forecasting task. Experiments show that the methods achieved competitive results in comparison to state-of-the-art methods. Gustavo H. F. M. Oliveira, Rodolfo Carneiro Cavalcante, George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira |
ICTAI | 3 |
| 2014 | One-class Classification for heart disease diagnosisabstractAs has been shown by the recent literature, machine learning techniques are important tools for diagnosing a number of diseases. Hospitals and medical clinics store a large amount of data with respect to the treatment of their patients. However, rarely an analysis of these data is conducted in order to extract intrinsic information for modeling a specific problem. This work presents an analysis of medical data aimed at determining whether or not patients are cardiac. To this end, raw data was collected and preprocessed at a Brazilian local hospital in order to build a new dataset containing only non-invasive information of children with heart murmur symptoms. The gathered data contain information, such as height, weight, gender and birthday date. The collected data was shown to be very imbalanced. Due to this imbalance, we employ the One-class Classification (OCC) paradigm to solve the problem by experimenting five methods; including the FBDOCC, that we proposed in a previous paper. Furthermore, two additional datasets were experimented in order to assess effectiveness of One-Class classifiers on the domain of heart disease detection. The overall results show that the FBDOCC succeeded in this task, yielding, statistically, the best performance for the gathered dataset as well as the other two heart disease datasets. George G. Cabral, Adriano Lorena Inácio de Oliveira |
SMC | 1 |
| 2014 | One-Class Classification based on searching for the problem features limits
George G. Cabral, Adriano Lorena Inácio de Oliveira |
Expert Syst. Appl. | 1 |
| 2013 | Preprocessing unbalanced data using weighted support vector machines for prediction of heart disease in childrenabstractMachine learning techniques are an important tool for diagnosing a number of diseases, as has been shown by the recent literature. Hospitals and medical clinics have a huge amount of data about the treatment of their patients, however, rarely analysis of these data is performed in order to extract intrinsic information aimed at modeling a specific problem. This work presents an analysis of medical data aimed at determining whether children patients are cardiac or not. To this end, raw data was collected at a Brazilian local hospital to be preprocessed in order to build the classification models. Only non invasive information were used, such as height, weight, gender and birthday date to create another set of derived variables such as BMI (Body Mass Index) to support the classification phase. However, the collected data was shown to be very imbalanced. Aimed at treat this problem, many tecniques were employed and one new approach was proposed. The results shown that the proposed approach outperforms the other methods in three out of four evaluation metrics. Thiago Tavares, Adriano Lorena Inácio de Oliveira, George G. Cabral, Sandra da Silva Mattos, Renata Grigorio |
IJCNN | 3 |
| 2012 | One-Class Classification through Optimized Feature Boundaries Detection and Prototype Reduction
George G. Cabral, Adriano Lorena Inácio de Oliveira |
ICANN (1) | 1 |
| 2012 | Extreme Learning Machines for Intrusion Detection Systems
Gilles Paiva M. de Farias, Adriano Lorena Inácio de Oliveira, George G. Cabral |
ICONIP (4) | 3 |
| 2011 | A novel one-class classification method based on feature analysis and prototype reductionabstractOne-class classification is an important problem with applications in several different areas such as outlier detection and machine monitoring. In this paper we propose a novel method for one-class classification which also implements prototype reduction. The main feature of the proposed method is to analyze every limit of all the feature dimensions to find the true border which describes the normal class. To this end, the proposed method simulates the novelty class by creating artificial prototypes outside the normal description. The method is able to describe data distributions with complex shapes. Aiming to assess the proposed method, we carried out experiments with synthetic and real datasets to compare it with the Support Vector Domain Description (SVDD), kMeansDD, ParzenDD and kNNDD methods. The experimental results show that our one-class classification approach outperformed the other methods in terms of the area under the receiver operating characteristic (ROC) curve in three out of six data sets. The results also show that the proposed method remarkably outperformed the SVDD regarding training time and reduction of prototypes. George G. Cabral, Adriano Lorena Inácio de Oliveira |
SMC | 1 |
| 2010 | A hybrid method for novelty detection in time series based on states transitions and swarm intelligenceabstractThis paper introduces a novel instance-based one-class classification method for novelty detection in time series based on its states transition. The main feature of our work is to generate an efficient method which automatically finds the parameters (whose yields the best model) according with the quality of the discovered time series states and the validation error. This method involves clustering and reducing the number of samples in a training dataset which does not contain novelty samples. Experiments carried out using three real-world time series show that the proposed method is able to build models with a reduced number of stored prototypes. The results obtained by our method were compared with the results of the SAX and both methods have successfully detected the novelties, however, the parameters which resulted in the best SAX model were achieved without validation phase (i.e. analyzing the results obtained for the test set). George G. Cabral, Adriano Lorena Inácio de Oliveira |
IJCNN | 1 |
| 2009 | Combining nearest neighbor data description and structural risk minimization for one-class classification
George G. Cabral, Adriano Lorena Inácio de Oliveira, Carlos B. G. Cahu |
Neural Comput. Appl. | 1 |
| 2008 | A Comparative Study of Machine Learning Techniques for Caries PredictionabstractThere are striking disparities in the prevalence of dental disease by income. Poor children suffer twice as much dental caries as their more affluent peers, but are less likely to receive treatment. This paper presents an experimental study of the application of machine learning methods to the problem of caries prediction. For this paper a data set collected from interviews with children under five years of age, in 2006, in Recife, the capital of Pernambuco, a state in northeast Brazil, was built. Four different data mining techniques were applied to this problem and their results were confronted in terms of the classification error and area under the ROC curve (AUC). Results showed that the MLP neural network classifier out performed the other machine learning methods employed in the experiments, followed by the support vector machine (SVM) predictor. In addition, the results also show that some rules (extracted by decision tress) may be useful for understanding the most important factors that influence the occurrence of caries in children. Robson D. Montenegro, Adriano Lorena Inácio de Oliveira, George G. Cabral, Cintia R. T. Katz, Aronita Rosenblatt |
ICTAI (2) | 3 |
| 2007 | A Novel Method for One-Class Classification Based on the Nearest Neighbor Data Description and Structural Risk MinimizationabstractOne-class classification is an important problem with applications in several different areas such as novelty detection, outlier detection and machine monitoring. In this paper we propose a novel method for one-class classification, referred to as NNDDSRM. It is based on the principle of structural risk minimization and the nearest neighbor data description (NNDD) method. Experiments carried out using both artificial and real-world datasets show that the proposed method is able to significantly reduce the number of stored prototypes in comparison to NNDD. The experimental results also show that the proposed method outperformed NNDD - in terms of the area under the receiver operating characteristic (ROC) curve - on four of the five datasets considered in the experiments and had a similar performance on the remaining one. George G. Cabral, Adriano Lorena Inácio de Oliveira, Carlos B. G. Cahu |
IJCNN | 1 |