Liyan Song

dblp:129/2667 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0003-1172-8825ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 9 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Coevolutionary software multi-project scheduling with matrix management and online skill training
Xiaoning Shen, Liyan Song, Chengbin Yao
Eng. Appl. Artif. Intell.3
2026 Stable region enhanced online learning method for intermediate verification latency and concept drift
Zixin Zhong, Liyan Song, Fengzhen Tang, Bo Yuan 0006
Pattern Recognit.2
2024 Addressing Intermediate Verification Latency in Online Learning Through Immediate Pseudo-labeling and Oriented Synthetic Correction
abstract
In non-stationary data streams, the challenges of concept drift are further compounded by the issue of Intermediate Verification Latency (IVL), which can impede timely model adaptation. IVL refers to the finite delay between the arrival of data features and their corresponding labels. This delay could pose a significant challenge in adapting models to new concepts, ultimately hindering predictive performance. However, existing IVL approaches exhibit certain limitations. Some approaches passively wait for delayed labels, thereby overlooking temporarily unlabeled data. Other approaches employ pseudo-labeling for immediate model updates, but may risk losing valuable information when reverting model states to rectify previous pseudo-labeling mistakes. To overcome these limitations, we propose a novel approach called Micro-cluster based Immediate Pseudo-Labeling with Oriented Synthetic Correction (MIPLOSC). MIPLOSC leverages micro-cluster systems to effectively capture data distributions, thus facilitating its two core components: immediate pseudo-labeling and oriented synthetic correction. The immediate pseudo-labeling mechanism facilitates immediate utilization of temporarily unlabeled data, and the oriented synthetic correction mechanism enables finergrained rectification from previous erroneous pseudo-labels and concept drift, minimizing the loss of learned information. Experimental studies validated the effectiveness of MIPLOSC in addressing IVL, demonstrating its superiority over competing methods in both space consumption and predictive performance across varying degrees of label delay.
Zixin Zhong, Liyan Song, Fengzhen Tang, Bo Yuan 0006
IJCNN2
2024 Coevolutionary scheduling of dynamic software project considering the new skill learning
Xiaoning Shen, Chengbin Yao, Liyan Song, Jiyong Xu, Mingjian Mao
Autom. Softw. Eng.3
2024 Online cross-project approach with project-level similarity for just-in-time software defect prediction
Cong Teng, Liyan Song, Xin Yao 0001
Empir. Softw. Eng.2
2024 Feature mapping based on heterogeneous cross-company effort estimation
Xiaoning Shen, Liyan Song
Softw. Qual. J.4
2024 Multi-Class Imbalance Classification Based on Data Distribution and Adaptive Weights
abstract
AdaBoost approaches have been used for multi-class imbalance classification with an imbalance ratio measured on class sizes. However, such ratio would assign each training sample of the same class with the same weight, thus failing to reflect the data distribution within a class. We propose to incorporate the density information of training samples into the class imbalance ratio so that samples of the same class could have different weights. As one could use the entire training set to calculate the imbalance and density factors, the weight of a training sample resulting from the two factors remains static throughout the training epochs. However, static weights could not reflect the up-to-date training status of base learners. To deal with this, we propose to design an adaptive weighting mechanism by making use of up-to-date training status to further alleviate the multi-class imbalance issue. Ultimately, we incorporate the class imbalance ratio, the density-based factor, and the adaptive weighting mechanism into a single variable, based on which the adaptive weights of all training samples are computed. Experimental studies are carried out to investigate the effectiveness of the proposed approach and each of the three components in dealing with multi-class imbalance classification problem.
Liyan Song, Zheng Hu 0002, Yiu-Ming Cheung, Xin Yao 0001
IEEE Trans. Knowl. Data Eng.2
2023 BEDCOE: Borderline Enhanced Disjunct Cluster Based Oversampling Ensemble for Online Multi-Class Imbalance Learning
abstract
Multi-class imbalance learning usually confronts more challenges especially when learning from streaming data. Most existing methods focus on manipulating class imbalance ratios, disregarding other data properties such as the borderline and the disjunct. Recent studies have shown non-negligible impact of disregarding these properties on deteriorating predictive performance. Online multi-class imbalance would further exacerbate such negative impact. To abridge the research gap of online multi-class imbalance learning, we propose to enhance the number of training times of borderline samples based on the disjunct class-wise clusters that are adaptively constructed over time for each class individually. Specifically, we propose a borderline enhanced strategy for ensemble aiming to increase the number of training times of samples neighboring to borderline areas of different classes. We also propose to generate synthetic samples for training based on the adaptively learned disjunct clusters that are maintained for each class individually online, catering for online multi-class imbalance problem directly. These two components construct the Borderline Enhanced Disjunct Cluster Based Oversampling Ensemble (BEDCOE). Experimental studies are conducted and demonstrate the effectiveness of BEDCOE and each of its components in dealing with online multi-class imbalance.
Liyan Song, Yiu-Ming Cheung, Xin Yao 0001
ECAI2
2023 ARConvL: Adaptive Region-Based Convolutional Learning for Multi-class Imbalance Classification
Liyan Song, Yiu-Ming Cheung, Xin Yao 0001
ECML/PKDD (2)2
2023 A Practical Human Labeling Method for Online Just-in-Time Software Defect Prediction
abstract
Just-in-Time Software Defect Prediction (JIT-SDP) can be seen as an online learning problem where additional software changes produced over time may be labeled and used to create training examples. These training examples form a data stream that can be used to update JIT-SDP models in an attempt to avoid models becoming obsolete and poorly performing. However, labeling procedures adopted in existing online JIT-SDP studies implicitly assume that practitioners would not inspect software changes upon a defect-inducing prediction, delaying the production of training examples. This is inconsistent with a real-world scenario where practitioners would adopt JIT-SDP models and inspect certain software changes predicted as defect-inducing to check whether they really induce defects. Such inspection means that some software changes would be labeled much earlier than assumed in existing work, potentially leading to different JIT-SDP models and performance results. This paper aims at formulating a more practical human labeling procedure that takes into account the adoption of JIT-SDP models during the software development process. It then analyses whether and to what extent it would impact the predictive performance of JIT-SDP models. We also propose a new method to target the labeling of software changes with the aim of saving human inspection effort. Experiments based on 14 GitHub projects revealed that adopting a more realistic labeling procedure led to significantly higher predictive performance than when delaying the labeling process, meaning that existing work may have been underestimating the performance of JIT-SDP. In addition, our proposed method to target the labeling process was able to reduce human effort while maintaining predictive performance by recommending practitioners to inspect software changes that are more likely to induce defects. We encourage the adoption of more realistic human labeling methods in research studies to obtain an evaluation of JIT-SDP predictive performance that is closer to reality.
Liyan Song, Leandro L. Minku, Cong Teng, Xin Yao 0001
ESEC/SIGSOFT FSE1
2023 On the validity of retrospective predictive performance evaluation procedures in just-in-time software defect prediction
abstract
Abstract Just-In-Time Software Defect Prediction (JIT-SDP) is concerned with predicting whether software changes are defect-inducing or clean. It operates in scenarios where labels of software changes arrive over time with delay, which in part corresponds to the time we wait to label software changes as clean (waiting time). However, clean labels decided based on waiting time may be different from the true labels of software changes, i.e., there may be label noise. This typically overlooked issue has recently been shown to affect the validity of continuous performance evaluation procedures used to monitor the predictive performance of JIT-SDP models during the software development process. It is still unknown whether this issue could potentially also affect evaluation procedures that rely on retrospective collection of software changes such as those adopted in JIT-SDP research studies, affecting the validity of the conclusions of a large body of existing work. We conduct the first investigation of the extent with which the choice of waiting time and its corresponding label noise would affect the validity of retrospective performance evaluation procedures. Based on 13 GitHub projects, we found that the choice of waiting time did not have a significant impact on the validity and that even small waiting times resulted in high validity. Therefore, (1) the estimated predictive performances in JIT-SDP studies are likely reliable in view of different waiting times, and (2) future studies can make use of not only larger (5k+ software changes), but also smaller (1k software changes) projects for evaluating performance of JIT-SDP models.
Liyan Song, Leandro L. Minku, Xin Yao 0001
Empir. Softw. Eng.1
2023 STCM: A spatio-temporal calibration model for low-cost air monitoring sensors
Chang Ju, Jiahu Qin, Liyan Song, Zongxi Li
Inf. Sci.4
2023 A Procedure to Continuously Evaluate Predictive Performance of Just-In-Time Software Defect Prediction Models During Software Development
abstract
Just-In-Time Software Defect Prediction (JIT-SDP) uses machine learning to predict whether software changes are defect-inducing or clean. When adopting JIT-SDP, changes in the underlying defect generating process may significantly affect the predictive performance of JIT-SDP models over time. Therefore, being able to continuously track the predictive performance of JIT-SDP models during the software development process is of utmost importance for software companies to decide whether or not to trust the predictions provided by such models over time. However, there has been little discussion on how to continuously evaluate predictive performance in practice, and such evaluation is not straightforward. In particular, labeled software changes that can be used for evaluation arrive over time with a delay, which in part corresponds to the time we have to wait to label software changes as ‘clean’ (waiting time). A clean label assigned based on a given waiting time may not correspond to the true label of the software changes. This can potentially hinder the validity of any continuous predictive performance evaluation procedure for JIT-SDP models. This paper provides the first discussion of how to continuously evaluate predictive performance of JIT-SDP models over time during the software development process, and the first investigation of whether and to what extent waiting time affects the validity of such continuous performance evaluation procedure in JIT-SDP. Based on 13 GitHub projects, we found that waiting time had a significant impact on the validity. Though typically small, the differences in estimated predicted performance were sometimes large, and thus inappropriate choices of waiting time can lead to misleading estimations of predictive performance over time. Such impact did not normally change the ranking between JIT-SDP models, and thus conclusions in terms of which JIT-SDP model performs better are likely reliable independent of the choice of waiting time, especially when considered across projects.
Liyan Song, Leandro L. Minku
IEEE Trans. Software Eng.1
2022 A Novel Data Stream Learning Approach to Tackle One-Sided Label Noise From Verification Latency
abstract
Many real-world data stream applications suffer from verification latency, where the labels of the training examples arrive with a delay. In binary classification problems, the labeling process frequently involves waiting for a pre-determined period of time to observe an event that assigns the example to a given class. Once this time passes, if such labeling event does not occur, the example is labeled as belonging to the other class. For example, in software defect prediction, one may wait to see if a defect is associated to a software change implemented by a developer, producing a defect-inducing training example. If no defect is found during the waiting time, the training example is labeled as clean. Such verification latency inherently causes label noise associated to insufficient waiting time. For example, a defect may be observed only after the pre-defined waiting time has passed, resulting in a noisy example of the clean class. Due to the nature of the waiting time, such noise is frequently one-sided, meaning that it only occurs to examples of one of the classes. However, no existing work tackles label noise associated to verification latency. This paper proposes a novel data stream learning approach that estimates the confidence in the labels assigned to the training examples and uses this to improve predictive performance in problems with one-sided label noise. Our experiments with 14 real-world datasets from the domain of software defect prediction demonstrate the effectiveness of the proposed approach compared to existing ones.
Liyan Song, Leandro L. Minku, Xin Yao 0001
IJCNN1
2022 Direct ICA on data tensor via random matrix modeling
abstract
Independent Component Analysis (ICA) is a fundamental method for Blind Source Separation (BSS). Classical ICA takes data matrix input formed by vector data. This paper focuses on ICA for BSS with third-order data tensor input formed by matrix data, such as 2D images. Two approaches exist for this problem. The first approach reshapes each matrix into a vector to apply classical ICA, with structural information lost. The second approach unfolds a data tensor into a data matrix along different modes to perform classical ICA mode-wise, which partially preserves structures but has strong or ill BSS assumptions. This paper proposes a third approach via RAndom Matrix ICA (RAMICA) modeling. RAMICA works on data tensor directly, without vectorization or unfolding, and preserves row or column structures under more general BSS assumptions. We develop the RAMICA model, algorithm, and related theories via defining new statistics for random matrices and new procedures for whitening and independent component estimation. We study the identifiability, higher-order extension, and relationships with existing methods. Experiments on both synthetic and real data show superior BSS performance of RAMICA over competing methods and offer insights on the trade-offs between different factors.
Liyan Song, Shuo Zhou 0008, Haiping Lu
Signal Process.1
2021 Label-Assisted Memory Autoencoder for Unsupervised Out-of-Distribution Detection
Chao Pan 0005, Liyan Song, Ke Pei, Peter Tiño, Xin Yao 0001
ECML/PKDD (3)3
2020 An investigation of cross-project learning in online just-in-time software defect prediction
abstract
Just-In-Time Software Defect Prediction (JIT-SDP) is concerned with predicting whether software changes are defect-inducing or clean based on machine learning classifiers. Building such classifiers requires a sufficient amount of training data that is not available at the beginning of a software project. Cross-Project (CP) JIT-SDP can overcome this issue by using data from other projects to build the classifier, achieving similar (not better) predictive performance to classifiers trained on Within-Project (WP) data. However, such approaches have never been investigated in realistic online learning scenarios, where WP software changes arrive continuously over time and can be used to update the classifiers. It is unknown to what extent CP data can be helpful in such situation. In particular, it is unknown whether CP data are only useful during the very initial phase of the project when there is little WP data, or whether they could be helpful for extended periods of time. This work thus provides the first investigation of when and to what extent CP data are useful for JIT-SDP in a realistic online learning scenario. For that, we develop three different CP JIT-SDP approaches that can operate in online mode and be updated with both incoming CP and WP training examples over time. We also collect 2048 commits from three software repositories being developed by a software company over the course of 9 to 10 months, and use 19,8468 commits from 10 active open source GitHub projects being developed over the course of 6 to 14 years. The study shows that training classifiers with incoming CP+WP data can lead to improvements in G-mean of up to 53.90% compared to classifiers using only WP data at the initial stage of the projects. For the open source projects, which have been running for longer periods of time, using CP data to supplement WP data also helped the classifiers to reduce or prevent large drops in predictive performance that may occur over time, leading to up to around 40% better G-Mean during such periods. Such use of CP data was shown to be beneficial even after a large number of WP data were received, leading to overall G-means up to 18.5% better than those of WP classifiers.
Sadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral, Liyan Song
ICSE5
2019 Software Effort Interval Prediction via Bayesian Inference and Synthetic Bootstrap Resampling
abstract
Software effort estimation (SEE) usually suffers from inherent uncertainty arising from predictive model limitations and data noise. Relying on point estimation only may ignore the uncertain factors and lead project managers (PMs) to wrong decision making. Prediction intervals (PIs) with confidence levels (CLs) present a more reasonable representation of reality, potentially helping PMs to make better-informed decisions and enable more flexibility in these decisions. However, existing methods for PIs either have strong limitations or are unable to provide informative PIs. To develop a “better” effort predictor, we propose a novel PI estimator called Synthetic Bootstrap ensemble of Relevance Vector Machines (SynB-RVM) that adopts Bootstrap resampling to produce multiple RVM models based on modified training bags whose replicated data projects are replaced by their synthetic counterparts. We then provide three ways to assemble those RVM models into a final probabilistic effort predictor, from which PIs with CLs can be generated. When used as a point estimator, SynB-RVM can either significantly outperform or have similar performance compared with other investigated methods. When used as an uncertain predictor, SynB-RVM can achieve significantly narrower PIs compared to its base learner RVM. Its hit rates and relative widths are no worse than the other compared methods that can provide uncertain estimation.
Liyan Song, Leandro L. Minku, Xin Yao 0001
ACM Trans. Softw. Eng. Methodol.1
2018 A novel automated approach for software effort estimation based on data augmentation
abstract
Software effort estimation (SEE) usually suffers from data scarcity problem due to the expensive or long process of data collection. As a result, companies usually have limited projects for effort estimation, causing unsatisfactory prediction performance. Few studies have investigated strategies to generate additional SEE data to aid such learning. We aim to propose a synthetic data generator to address the data scarcity problem of SEE. Our synthetic generator enlarges the SEE data set size by slightly displacing some randomly chosen training examples. It can be used with any SEE method as a data preprocessor. Its effectiveness is justified with 6 state-of-the-art SEE models across 14 SEE data sets. We also compare our data generator against the only existing approach in the SEE literature. Experimental results show that our synthetic projects can significantly improve the performance of some SEE methods especially when the training data is insufficient. When they cannot significantly improve the prediction performance, they are not detrimental either. Besides, our synthetic data generator is significantly superior or perform similarly to its competitor in the SEE literature. Therefore, our data generator plays a non-harmful if not significantly beneficial effect on the SEE methods investigated in this paper. Therefore, it is helpful in addressing the data scarcity problem of SEE.
Liyan Song, Leandro L. Minku, Xin Yao 0001
ESEC/SIGSOFT FSE1
2016 Proper Inner Product with Mean Displacement for Gaussian Noise Invariant ICA
abstract
Independent Component Analysis (ICA) is a classical method for Blind Source Separation (BSS). In this paper, we are interested in ICA in the presence of noise, i.e., the noisy ICA problem. Pseudo-Euclidean Gradient Iteration (PEGI) is a recent cumulant-based method that defines a pseudo Euclidean inner product to replace a quasi-whitening step in Gaussian noise invariant ICA. However, PEGI has two major limitations: 1) the pseudo Euclidean inner product is improper because it violates the positive definiteness of inner product; 2) the inner product matrix is orthogonal by design but it has gross errors or imperfections due to sample-based estimation. This paper proposes a new cumulant-based ICA method named as PIMD to address these two problems. We first define a Proper Inner product (PI) with proved positive definiteness and then relax the centering preprocessing step to a mean displacement (MD) step. Both PI and MD aim to improve the orthogonality of inner product matrix and the recovery of independent components (ICs) in sample-based estimation. We adopt a gradient iteration step to find the ICs for PIMD. Experiments on both synthetic and real data show the respective effectiveness of PI and MD as well as the superiority of PIMD over competing ICA methods. Moreover, MD can improve the performance of other ICA methods as well.
Liyan Song, Haiping Lu
ACML1
2016 EcoICA: Skewness-based ICA via Eigenvectors of Cumulant Operator
abstract
Independent component analysis (ICA) is an important unsupervised learning method. Most popular ICA methods use kurtosis as a metric of non-Gaussianity to maximize, such as FastICA and JADE.However, their assumption of kurtosic sources may not always be satisfied in practice. For weak-kurtosic but skewed sources, kurtosis-based methods could fail while skewness-based methods seem more promising, where skewness is another non-Gaussianity metric measuring the non-symmetry of signals. Partly due to the common assumption of signal symmetry, skewness-based ICA has not been systematically studied in spite of some existing works. In this paper, we take a systematic approach to develop EcoICA, a new skewness-based ICA method for weak-kurtosic but skewed sources. Specifically, we design a new cumulant operator, define its eigenvalues and eigenvectors, reveal their connections with the ICA model to formulate the EcoICA problem, and use Jacobi method to solve it. Experiments on both synthetic and real data show the superior performance of EcoICA over existing kurtosis-based and skewness-based methods for skewed sources. In particular, EcoICA is less sensitive to sample size, noise, and outlier than other methods. Studies on face recognition further confirm the usefulness of EcoICA in classification.
Liyan Song, Haiping Lu
ACML1
2013 Spectral Decomposition for Optimal Graph Index Prediction
Liyan Song, Yun Peng 0002, Byron Choi, Jianliang Xu, Bingsheng He
PAKDD (1)1