Sadia Tabassum

dblp:276/3315 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-5096-7100ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Fine-Tuning for Efficient Long-Context Learning in Large Language Models
Sadia Tabassum, Mussammat Maimuna Faria, Md. Nurul Ahad Tawhid
ENASE (1)1
2025 Hyperon: An Online Hyperparameter Tuning Approach for Data Stream Learning
abstract
Predictive models built using machine learning algorithms usually involve a number of hyperparameters that can significantly affect their performance. While many approaches for hyperparameter tuning have been investigated for offline learning, there is little work in the context of online data stream learning. Hyperparameter tuning for online data stream learning can be particularly challenging, due to possible changes in the underlying distribution of the problem. Such changes can result in the best hyperparameter choice varying over time, requiring efficient, real-time online adaptation. However, existing online hyperparameter tuning approaches are limited to specific models, are susceptible to local optima, rely on fixed hyperparameter grids, or on concept drift detection methods. We propose a novel online hyperparameter tuning approch for data stream learning called Hyperon to overcome these issues. Hyperon intertwines online data stream learning with a steady-state evolutionary algorithm, enabling efficient and effective hyperparameter optimisation over time. Experiments on 10 real world data streams show that Hyperon is able to significantly improve the predictive performance of the underlying online data stream learning approach in a computationally efficient manner.
Sadia Tabassum, Leandro L. Minku
DSAA1
2024 Correction to: An investigation of online and offline learning models for online just-in-time software defect prediction
abstract
While the University of Birmingham exercises care and attention in making items available there are rare occasions when an item has been uploaded in error or has been deemed to be commercially or otherwise sensitive.If you believe that this is the case for this document, please contact [email protected] providing details and we will remove
George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira, Dinaldo A. Pessoa, Sadia Tabassum
Empir. Softw. Eng.5
2023 An investigation of online and offline learning models for online Just-in-Time Software Defect Prediction
abstract
Abstract Just-in-Time Software Defect Prediction (JIT-SDP) operates in an online scenario where additional training data is received over time. Existing online JIT-SDP studies used online Oza ensemble learning methods with Hoeffding Trees as base learners to learn and update JIT-SDP models over time in this scenario. However, it is unknown how these approaches compare against offline learning approaches adapted to operate in online scenarios, and how the use of any other online or offline base learners would affect online JIT-SDP in terms of predictive performance and computational cost. We therefore propose a new approach called Batch Oversampling Rate Boosting (BORB) that is able to use offline base learners in an online JIT-SDP scenario. Based on 10 open source projects, we provide a comprehensive evaluation of BORB with 5 different base learners and the existing online approach Oversampling Rate Boosting with 4 different base learners, both in within-project and cross-project online JIT-SDP scenarios. The results show that offline learning can lead to better predictive performance than the top performing online learning approaches considered in our study, at a higher computational cost. Cross-project data was helpful to improve predictive performance both for offline and online learning, but especially for online learning.
George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira, Dinaldo A. Pessoa, Sadia Tabassum
Empir. Softw. Eng.5
2023 Cross-Project Online Just-In-Time Software Defect Prediction
abstract
Cross-Project (CP) Just-In-Time Software Defect Prediction (JIT-SDP) makes use of CP data to overcome the lack of data necessary to train well performing JIT-SDP classifiers at the beginning of software projects. However, such approaches have never been investigated in realistic online learning scenarios, where Within-Project (WP) software changes naturally arrive over time and can be used to automatically update the classifiers. We provide the first investigation of when and to what extent CP data are useful for JIT-SDP in such realistic scenarios. For that, we propose three different online CP JIT-SDP approaches that can be updated with incoming CP and WP training examples over time. We also collect data on 9 proprietary software projects and use 10 open source software projects to analyse these approaches. We find that training classifiers with incoming CP+WP data can lead to absolute improvements in G-mean of up to 53.89% and up to 35.02% at the initial stage of the projects compared to classifiers using WP-only and CP-only data, respectively. Using CP+WP data was also shown to be beneficial after a large number of WP data were received. Using CP data to supplement WP data helped the classifiers to reduce or prevent large drops in predictive performance that may occur over time, leading to absolute G-Mean improvements of up to 37.35% and 48.16% compared to WP-only and CP-only data during such periods, respectively. During periods of stable predictive performance, absolute improvements were of up to 29.03% and up to 41.25% compared to WP-only and CP-only classifiers, respectively. Our results highlight the importance of using both CP and WP data together in realistic online JIT-SDP scenarios.
Sadia Tabassum, Leandro L. Minku, Danyi Feng
IEEE Trans. Software Eng.1
2020 An investigation of cross-project learning in online just-in-time software defect prediction
abstract
Just-In-Time Software Defect Prediction (JIT-SDP) is concerned with predicting whether software changes are defect-inducing or clean based on machine learning classifiers. Building such classifiers requires a sufficient amount of training data that is not available at the beginning of a software project. Cross-Project (CP) JIT-SDP can overcome this issue by using data from other projects to build the classifier, achieving similar (not better) predictive performance to classifiers trained on Within-Project (WP) data. However, such approaches have never been investigated in realistic online learning scenarios, where WP software changes arrive continuously over time and can be used to update the classifiers. It is unknown to what extent CP data can be helpful in such situation. In particular, it is unknown whether CP data are only useful during the very initial phase of the project when there is little WP data, or whether they could be helpful for extended periods of time. This work thus provides the first investigation of when and to what extent CP data are useful for JIT-SDP in a realistic online learning scenario. For that, we develop three different CP JIT-SDP approaches that can operate in online mode and be updated with both incoming CP and WP training examples over time. We also collect 2048 commits from three software repositories being developed by a software company over the course of 9 to 10 months, and use 19,8468 commits from 10 active open source GitHub projects being developed over the course of 6 to 14 years. The study shows that training classifiers with incoming CP+WP data can lead to improvements in G-mean of up to 53.90% compared to classifiers using only WP data at the initial stage of the projects. For the open source projects, which have been running for longer periods of time, using CP data to supplement WP data also helped the classifiers to reduce or prevent large drops in predictive performance that may occur over time, leading to up to around 40% better G-Mean during such periods. Such use of CP data was shown to be beneficial even after a large number of WP data were received, leading to overall G-means up to 18.5% better than those of WP classifiers.
Sadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral, Liyan Song
ICSE1