EDBT 2026 Demo / reviewers in the wild / expert
Zhiqiang Li 0003
dblp:14/5778-3
· DBLP profile ↗
29ranked-venue papers
12as first author
22since 2021 · last 2027
0000-0001-5999-3658ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | An empirical study of neural network based graph representations for software defect prediction
Zhiqiang Li 0003, Yanwei Xiang, Jie Ren 0007, Hongyu Zhang 0002, Fumin Qi, Xiaoyuan Jing |
Sci. Comput. Program. | 1 |
| 2026 | Lifting Optimized Binaries to Canonical Compiler IR via Structure-Aware Retrieval and Iterative VerificationabstractLifting stripped and highly optimized binaries to the canonical compiler intermediate representation (IR) enables program analysis when source code is unavailable.However, compiler optimizations severely distort controlflow and data-flow structure, making existing rule-based and LLM-based decompilation approaches brittle.We present BRIDGE, a system that reliably lifts optimized binaries to analysis-friendly compiler IR.BRIDGE combines control-flow-aware retrieval-augmented generation with feedback-driven verification.It uses pseudo-probe instrumentation to align optimized binary fragments with normalized IR semantics, and then employs an iterative refinement loop guided by static analysis and runtime feedback to improve executability and semantic consistency.We evaluate BRIDGE on HumanEval-Decompile and MBPP, lifting x86-64 and ARM64 binaries to LLVM IR.BRIDGE outperforms seven baselines, achieving an average of over 30% higher re-executability than the strongest general-purpose LLM baseline.Void func (){ PROBE(1); If else branch …… PROBE(2); PROBE(3); for Loop … PROBE(4);} Xiaoao Zhu, Jie Ren 0007, Zhiqiang Li 0003, Jie Zheng 0005, Zhanyong Tang, Zheng Wang 0001 |
ACL (1) | 3 |
| 2026 | Line-level defect prediction based on preceding line-aware and inter-line semantics enhancement
Xiaoke Zhu, Xiaopan Chen, Zhiqiang Li 0003, Xiaoyuan Jing |
Inf. Softw. Technol. | 4 |
| 2026 | An enhancing framework with an emphasis on decision balance in ensemble regression
Qiancheng Yu, Kaiguang Wang, Cai Dai, Zhiqiang Li 0003 |
Neural Networks | 6 |
| 2026 | Metric information mining with metric attention to boost software defect prediction performanceabstractIn the field of software engineering, defect prediction has always been a popular research direction. Currently, the research on traditional software defect prediction mainly focuses on metric features, which are derived from various descriptive rules. Many researchers have proposed a large number of defect prediction models based on these metric features and various framework models. However, the problem of data scarcity has severely hindered the development of the field. Therefore, this work proposes a new method, namely the Metric Attention Module (MAM), which excavates the correlations within the metric data features, between features, within modules, and between modules. By learning new data representations, MAM guides the model's learning process and ultimately improves the model's performance without changing the network framework structure. Additionally, the method is interpretable. In this work, experiments were conducted in various task environments and on different datasets, all resulting in varying degrees of improvement. In the context of within-project defect prediction (WPDP), experiments with the MAM data model showed an average improvement of 14.7% in Accuracy, 15.9% in F1 score, 23.7% in AUC, and 65.1% in MCC. In cross-project defect prediction (CPDP), under more complex task environments, the model demonstrated excellent performance across multiple standard datasets. Compared to the baseline models and training results, the F1, Accuracy, and MCC scores improved by approximately 40%, 20%, and 50%, respectively. Yongchang Ding, Zhiqiang Li 0003, Linjun Chen, Rong Peng, Xiaoyuan Jing |
Sci. Comput. Program. | 3 |
| 2025 | An empirical comparison of data transformation techniques for clustering-based unsupervised software defect predictionabstractAs a key pre-processing step of software defect prediction, data transformation techniques aim to eliminate dimensional inconsistencies among software metrics and improve model performance. However, the impact of data transformation techniques on the performance of clustering-based unsupervised defect prediction (UDP) has not been fully explored. To this end, our aim is to empirically investigate the effect of data transformation techniques on the performance of clustering-based UDP models. In this paper, we systematically evaluated 23 unsupervised clustering models under 4 data conditions: (1) raw untransformed data (original), (2) logarithmic transformation, (3) z-score standardization, and (4) min-max normalization. Our experimental design consists of two dimensions: comparative analysis of individual models’ performance before and after data transformation and evaluation of relative performance across different models under various transformations. Extensive experiments conducted on 22 software projects indicate that: (1) Data transformation techniques have significant potential to improve the performance of unsupervised clustering models. They achieve improvements of 6.9% -22.9% in AUC,10.7% -41.6% in MCC, and 3.1% -12.27% in F_measure@20%. However, these transformations lead to a degradation in IFA performance, which inevitably increases code inspection efforts but within tolerable operational thresholds. (2) The performance of unsupervised clustering models varies depending on the data transformation techniques used, with no single clustering model showing consistent superiority. Empirical findings indicate that data transformation techniques exhibit general effectiveness in clustering-based UDP models, and the choice of an unsupervised clustering model should align with the specific data transformation technique employed. Zhengxiang Chen, Zhiqiang Li 0003, Hongyu Zhang 0002, Jie Ren 0007, Feng Tian 0005 |
APSEC | 2 |
| 2025 | LLVMTuner: Predictive Compiler Optimization for LLVM IR Across Heterogeneous PlatformsabstractThe Low Level Virtual Machine Intermediate Representation (LLVM IR) is a key component of modern compilers, valued for its universality and cross-platform adaptability. However, identifying optimal optimization passes across platforms remains challenging, as current autotuning frameworks struggle on mobile devices due to resource constraints and network variability. This paper introduces a novel LLVM IR performance optimization framework, LL VMTUNER. At its core is a deep neural network-based predictive model specifically designed to forecast the execution time of input passes on various platforms by analyzing IR features. By integrating this predictive model with the advanced autotuning framework, we enable a rapid and precise search for optimal pass list, removing the need for real-time latency measurements on target platforms and reducing data transmission between devices and cloud-based autotuning frameworks. LLVMTuner significantly enhances the performance of LLVM IR, providing a robust solution for efficient compiler optimizations across a spectrum of computing environments, from high-performance laptops to resource-constrained mobile devices. We evaluate LL VMTuNER on three heterogeneous platforms using over 600 LLVM IR benchmarks. The results show that LLVMTuner achieves an average speedup of 2.41x over the -03 optimization configuration. Additionally, LL VMTUNER reduces search overhead by 83.68% compared to OpenTuner. Xiaoao Zhu, Jie Ren 0007, Zhiqiang Li 0003, Feng Tian 0005, Jie Zheng 0005 |
CSCWD | 3 |
| 2025 | Optimizing Personalized Federated Learning Through Adaptive Layer-Wise LearningabstractReal-life deployment of federated Learning (FL) often faces non-IID data, which leads to poor accuracy and slow convergence. Personalized FL (pFL) tackles these issues by tailoring local models to individual data sources and using weighted aggregation methods for client-specific learning. However, existing pFL methods often fail to provide each local model with global knowledge on demand while maintaining low computational overhead. Additionally, local models tend to over-personalize their data during the training process, potentially dropping previously acquired global information. We propose FLAYER, a novel layer-wise learning method for pFL that optimizes local model personalization performance. FLAYER considers the different roles and learning abilities of neural network layers of individual local models. It incorporates global information for each local model as needed to initialize the local model cost-effectively. It then dynamically adjusts learning rates for each layer during local training, optimizing the personalized learning process for each local model while preserving global knowledge. Additionally, to enhance global representation in pFL, FLAYER selectively uploads parameters for global aggregation in a layer-wise manner. We evaluate FLAYER on four representative datasets in computer vision and natural language processing domains. Compared to eight state-of-the-art pFL methods, FLAYER improves the inference accuracy, on average, by 5.20% (up to 14.29%). Code is available at https://github.com/lancasterJie/FLAYER/. Weihang Chen, Jie Ren 0007, Zhiqiang Li 0003, Zheng Wang 0001 |
IJCAI | 4 |
| 2025 | The impact of unsupervised feature selection techniques on the performance and interpretation of defect prediction models
Zhiqiang Li 0003, Wenzhi Zhu, Hongyu Zhang 0002, Yuantian Miao, Jie Ren 0007 |
Autom. Softw. Eng. | 1 |
| 2025 | Decision Preference Networks in Ensemble Classification learning: Focusing on decision preferences and influences
Kaiguang Wang, Cai Dai, Zhiqiang Li 0003 |
Inf. Process. Manag. | 6 |
| 2025 | Unsupervised Software Defect Prediction Through Multiview ClusteringabstractThe core goal of software defect prediction (SDP) is to identify modules with a high likelihood of defects, thereby enabling prioritization of quality assurance activities with low inspection effort. There are many supervised defect prediction models that are extensively studied. However, these methods require the need for labeling data to get enough training modules, which will cause a lot of waste of human resources. Cross-project defect prediction primarily reuses models trained on other projects with enough historical data. However, this strategy is often hindered by large distribution differences across different projects and privacy concerns of data. Unsupervised learning technique is an alternative solution to the unlabeled data, but it mainly focuses on single-view prediction by concatenating all the software metrics. This ignores the diversity and complementarity of different types of metrics. This study proposes a novel approach, namely, multiview unsupervised software defect prediction (MUSDP). It aims to collaboratively learn the diversity and complementarity of different views to build a robust and reliable defect prediction model. Extensive experiments on$ 28$releases from eight software projects indicate that MUSDP exhibits superior or comparable results regardingG-mean,AUC,$P_{\text{opt}}$, andRecall@20%compared to competing supervised and unsupervised methods. For the interpretation of MUSDP, the number of added and deleted lines significantly influence its predictions. Zhiqiang Li 0003, Hongyu Zhang 0002, Xiaoyuan Jing, Wangyang Yu 0001, Yueyue Liu 0002 |
IEEE Trans. Reliab. | 1 |
| 2025 | CFG2AT: Control Flow Graph and Graph Attention Network-Based Software Defect PredictionabstractSoftware defect prediction (SDP) plays a pivotal role in ensuring high-quality software development by aiding in the early identification of potential defects. This practice has gained substantial attention in the field of software engineering over the years. Recent advancements in deep learning have primarily focused on extracting general syntactic features from abstract syntax trees (ASTs) for SDP. However, AST-based neural network models might overlook important structural information related to control flows embedded within the source code. Given that software defects are often influenced by control flow patterns, this article proposes a novel SDP approach called control flow graph and graph attention (CFG2AT) network-based SDP. CFG2AT is specifically designed to automatically identify software defects and contains a graph-structured attention unit to effectively capture control flow information. To evaluate the effectiveness of CFG2AT, we carried out extensive experiments using data from 15 versions of six different open-source software projects under both within-project and cross-project defect prediction settings. Experimental results demonstrate that our proposed CFG2AT approach generally outperforms a range of competing methods for defect prediction. The improvement is 7.09%–12.80% inF1, 1.30%–4.15% in area under curve (AUC), and 6.78%–17.54% in Matthews correlation coefficient (MCC) under within-project defect prediction, and 23.76%–44.79% inF1, 8.93%–13.27% in AUC, and 36.92%–94.89% in MCC under CPDP, respectively. Zhiqiang Li 0003, Hongyu Zhang 0002, Xiaoyuan Jing |
IEEE Trans. Reliab. | 2 |
| 2024 | Optimizing the Utilization of Large Language Models via Schedule Optimization: An Exploratory StudyabstractBackground: Large Language Models (LLMs) have gained significant attention in machine-learning-as-a-service (MLaaS) offerings. In-context learning (ICL) is a technique that guides LLMs towards accurate query processing by providing additional information. However, longer prompts lead to higher costs of LLM service, creating a performance-cost trade-off. Aims: We aim to investigate the potential of combining schedule optimization with ICL to optimize LLM utilization. Method: We conduct an exploratory study. First, we consider the performance-cost trade-off in LLM utilization as a multi-objective optimization problem, aiming to select the most suitable prompt template for each LLM job to maximize accuracy (the percentage of correctly processed jobs) and minimize invocation cost. Next, we investigate three methods for prompt performance prediction to address the challenge of evaluating the accuracy objective in the fitness function, as the result can only be determined after submitting the job to the LLM. Finally, we apply widely used search-based techniques and evaluate their effectiveness. Results: The results indicate that the machine learning-based technique is an effective approach for prompt performance prediction and fitness function calculation. Schedule optimization can achieve higher accuracy or lower cost by selecting a suitable prompt template for each job, compared to simply submitting all jobs using a single prompt template, e.g., saving costs from 21.33% to 86.92% in our experiments on LLM-based log parsing. However, the performance of the evaluated search-based techniques varies across different instances and metrics, with no single technique consistently outperforming the others. Conclusions: This study demonstrates the potential of combining schedule optimization with ICL to improve the utilization of LLMs. However, there is still ample room for improving the searched-based techniques and prompt performance prediction techniques for more cost-effective LLM utilization. Yueyue Liu 0002, Hongyu Zhang 0002, Zhiqiang Li 0003, Yuantian Miao |
ESEM | 3 |
| 2024 | CPLS: Optimizing the Assignment of LLM QueriesabstractLarge Language Models (LLMs) like ChatGPT have gained significant attention because of their impressive capabilities, leading to a dramatic increase in their integration into intelligent software engineering. However, their usage as a service with varying performance and price options presents a challenging trade-off between desired performance and the associated cost. To address this challenge, we propose CPLS, a framework that utilizes transfer learning and local search techniques for assigning intelligent software engineering jobs to LLM-based services. CPLS aims to minimize the total cost of LLM invocations while maximizing the overall accuracy. The framework first leverages knowledge from historical data across different projects to predict the probability of an LLM processing a query correctly. Then, CPLS incorporates problem-specific rules into a local search algorithm to effectively generate Pareto optimal solutions based on the predicted accuracy and cost. To evaluate the proposed approach, we conduct extensive experiments on LLM-based log parsing, a typical software maintenance task. Our experimental results demonstrate that CPLS outperforms the baseline methods, providing solutions with the highest accuracy in 14 out of 16 instances. Compared to the baselines, CPLS achieves an accuracy improvement ranging from 1.24% to 485.54%, or reduces costs by 15.21% to 89.09% while maintaining the highest accuracy achieved by the baselines. Yueyue Liu 0002, Hongyu Zhang 0002, Zhiqiang Li 0003, Yuantian Miao |
ICSME | 3 |
| 2024 | OptLLM: Optimal Assignment of Queries to Large Language ModelsabstractLarge Language Models (LLMs) have garnered considerable attention owing to their remarkable capabilities, leading to an increasing number of companies offering LLMs as services. Different LLMs achieve different performance at different costs. A challenge for users lies in choosing the LLMs that best fit their needs, balancing cost and performance. In this paper, we propose a framework for addressing the cost-effective query allocation problem for LLMs. Given a set of input queries and candidate LLMs, our framework, named OptLLM, provides users with a range of optimal solutions to choose from, aligning with their budget constraints and performance preferences, including options for maximizing accuracy and minimizing cost. OptLLM predicts the performance of candidate LLMs on each query using a multi-label classification model with uncertainty estimation and then iteratively generates a set of non-dominated solutions by destructing and reconstructing the current solution. To evaluate the effectiveness of OptLLM, we conduct extensive experiments on various types of tasks, including text classification, question answering, sentiment analysis, reasoning, and log parsing. Our experimental results demonstrate that OptLLM substantially reduces costs by 2.40% to 49.18% while achieving the same accuracy as the best LLM. Compared to other multi-objective optimization algorithms, OptLLM improves accuracy by 2.94% to 69.05% at the same cost or saves costs by 8.79% and 95.87% while maintaining the highest attainable accuracy. Yueyue Liu 0002, Hongyu Zhang 0002, Yuantian Miao, Van-Hoang Le, Zhiqiang Li 0003 |
ICWS | 5 |
| 2024 | An empirical study of data sampling techniques for just-in-time software defect prediction
Zhiqiang Li 0003, Qiannan Du, Hongyu Zhang 0002, Xiaoyuan Jing, Fei Wu 0004 |
Autom. Softw. Eng. | 1 |
| 2024 | Software defect prediction: future directions and challenges
Zhiqiang Li 0003, Jingwen Niu, Xiaoyuan Jing |
Autom. Softw. Eng. | 1 |
| 2024 | Modeling and Analysis of ETC Control System with Colored Petri Net and Dynamic SlicingabstractNowadays, Electronic Toll Collection (ETC) control systems have been widely adopted to smoothen traffic flow on highways. However, as it is a complex business interaction system, there are inevitably flaws in its control logic process, such as the problem of vehicle fee evasion. We find that there is more than one way for vehicles to evade fees. This shows that it is difficult to ensure the completeness of its design. Therefore, it is necessary to adopt a novel formal method to model and analyze its design, detect flaws, and modify it. In this article, a Colored Petri net (CPN) is introduced to establish its model. To analyze and modify the system model more efficiently, a dynamic slicing method of CPN is proposed. First, a static slice is obtained from the static slicing criterion by backtracking. Second, considering all binding elements that can be enabled under the initial marking, a forward slice is obtained from the dynamic slicing criterion by traversing. Third, the dynamic slicing of CPN is obtained by taking the intersection of both slices. The proposed dynamic slicing method of CPN can be used to formalize and verify the behavior properties of an ETC control system, and the flaws can be detected effectively. As a case study, the flaw about a vehicle that has not completed the payment following the previous vehicle to pass the railing is detected by the proposed method. Wangyang Yu 0001, Jinming Kong, Zhijun Ding, Xiaojun Zhai, Zhiqiang Li 0003 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | DSSDPP: Data Selection and Sampling Based Domain Programming Predictor for Cross-Project Defect PredictionabstractCross-project defect prediction (CPDP) refers to recognizing defective software modules in one project (i.e., target) using historical data collected from other projects (i.e., source), which can help developers find defects and prioritize their testing efforts. Unfortunately, there often exists large distribution difference between the source and target data. Most CPDP methods neglect to select the appropriate source data for a given target at the project level. More importantly, existing CPDP models are parametric methods, which usually require intensive parameter selection and tuning to achieve better prediction performance. This would hinder wide applicability of CPDP in practice. Moreover, most CPDP methods do not address the cross-project class imbalance problem. These limitations lead to suboptimal CPDP results. In this paper, we propose a novel data selection and sampling based domain programming predictor (DSSDPP) for CPDP, which addresses the above limitations. DSSDPP is a non-parametric CPDP method, which can perform knowledge transfer across projects without the need for parameter selection and tuning. By exploiting the structures of source and target data, DSSDPP can learn a discriminative transfer classifier for identifying defects of the target project. Extensive experiments on 22 projects from four datasets indicate that DSSDPP achieves betterMCCandAUCresults against a range of competing methods both in the single-source and multi-source scenarios. Since DSSDPP is easy, effective, extensible, and efficient, we suggest that future work can use it with the well-chosen source data to conduct CPDP especially for the projects with limited computational budget. Zhiqiang Li 0003, Hongyu Zhang 0002, Xiaoyuan Jing, Juanying Xie, Jie Ren 0007 |
IEEE Trans. Software Eng. | 1 |
| 2022 | Data sampling and kernel manifold discriminant alignment for mixed-project heterogeneous defect prediction
Jingwen Niu, Zhiqiang Li 0003, Xiwei Dong, Xiaoyuan Jing |
Softw. Qual. J. | 2 |
| 2021 | Cross-Project Defect Prediction via Landmark Selection-Based Kernelized Discriminant Subspace AlignmentabstractCross-project defect prediction (CPDP) refers to identifying defect-prone software modules in one project (target) using historical data collected from other projects (source), which can help developers find bugs and prioritize their testing efforts. Recently, CPDP has attracted great research interest. However, the source and target data usually exist redundancy and nonlinearity characteristics. Besides, most CPDP methods do not exploit source label information to uncover the underlying knowledge for label propagation. These factors usually lead to unsatisfactory CPDP performance. To address the above limitations, we propose a landmark selection-based kernelized discriminant subspace alignment (LSKDSA) approach for CPDP. LSKDSA not only reduces the discrepancy of the data distributions between the source and target projects, but also characterizes the complex data structures and increases the probability of linear separability of the data. Moreover, LSKDSA encodes label information of the source data into domain adaptation learning process and makes itself with good discriminant ability. Extensive experiments on 13 public projects from three benchmark datasets demonstrate that LSKDSA performs better than a range of competing CPDP methods. The improvement is 3.44%-11.23% in g-measure, 5.75%-11.76% in AUC, and 9.34%-33.63% in MCC, respectively. Zhiqiang Li 0003, Jingwen Niu, Xiaoyuan Jing, Wangyang Yu 0001 |
IEEE Trans. Reliab. | 1 |
| 2021 | An Empirical Study on Heterogeneous Defect Prediction ApproachesabstractSoftware defect prediction has always been a hot research topic in the field of software engineering owing to its capability of allocating limited resources reasonably. Compared with cross-project defect prediction (CPDP), heterogeneous defect prediction (HDP) further relaxes the limitation of defect data used for prediction, permitting different metric sets to be contained in the source and target projects. However, there is still a lack of a holistic understanding of existing HDP studies due to different evaluation strategies and experimental settings. In this paper, we provide an empirical study on HDP approaches. We review the research status systematically and compare the HDP approaches proposed from 2014 to June 2018. Furthermore, we also investigate the feasibility of HDP approaches in CPDP. Through extensive experiments on 30 projects from five datasets, we have the following findings: (1) metric transformation-based HDP approaches usually result in better prediction effects, while metric selection-based approaches have better interpretability. Overall, the HDP approach proposed by Liet al.(CTKCCA) currently has the best performance. (2) Handling class imbalance problems can boost the prediction effects, but the improvements are usually limited. In addition, utilizing mixed project data cannot improve the performance of HDP approaches consistently since the label information in the target project is not used effectively. (3) HDP approaches are feasible for cross-project defect prediction in which the source and target projects have the same metric set. Xiaoyuan Jing, Zhiqiang Li 0003, Di Wu 0014, Zhiguo Huang |
IEEE Trans. Software Eng. | 3 |
| 2019 | Heterogeneous defect prediction with two-stage ensemble learning
Zhiqiang Li 0003, Xiaoyuan Jing, Xiaoke Zhu, Hongyu Zhang 0002, Baowen Xu |
Autom. Softw. Eng. | 1 |
| 2019 | On the Multiple Sources and Privacy Preservation Issues for Heterogeneous Defect PredictionabstractHeterogeneous defect prediction (HDP) refers to predicting defect-proneness of software modules in a target project using heterogeneous metric data from other projects. Existing HDP methods mainly focus on predicting target instances with single source. In practice, there exist plenty of external projects. Multiple sources can generally provide more information than a single project. Therefore, it is meaningful to investigate whether the HDP performance can be improved by employing multiple sources. However, a precondition of conducting HDP is that the external sources are available. Due to privacy concerns, most companies are not willing to share their data. To facilitate data sharing, it is essential to study how to protect the privacy of data owners before they release their data. In this paper, we study the above two issues in HDP. Specifically, to utilize multiple sources effectively, we propose a multi-source selection based manifold discriminant alignment (MSMDA) approach. To protect the privacy of data owners, a sparse representation based double obfuscation algorithm is designed and applied to HDP. Through a case study of 28 projects, our results show that MSMDA can achieve better performance than a range of baseline methods. The improvement is 3.4-15.3 percent in g-measure and 3.0-19.1 percent in AUG. Zhiqiang Li 0003, Xiaoyuan Jing, Xiaoke Zhu, Hongyu Zhang 0002, Baowen Xu |
IEEE Trans. Software Eng. | 1 |
| 2018 | Cost-sensitive transfer kernel canonical correlation analysis for heterogeneous defect prediction
Zhiqiang Li 0003, Xiaoyuan Jing, Fei Wu 0004, Xiaoke Zhu, Baowen Xu |
Autom. Softw. Eng. | 1 |
| 2018 | Heterogeneous fault prediction with cost-sensitive domain adaptationabstractSummary In the early phases of software testing, projects may have only limited historical defect data. Learning prediction model with such insufficient training data will limit the efficacy of learned predictor. In practice, there are usually many publicly available fault prediction datasets. Recently, heterogeneous fault prediction (HFP) has been proposed. However, existing HFP models do not investigate how to use mixed project data to predict target. Furthermore, defect data are often imbalanced. The imbalanced data distribution of source usually leads to serious misclassification of fault‐prone instances, which will degrade the predictor's performance. Existing HFP methods do not consider the class imbalance problem in the training stages. In this paper, we propose a novel Cost‐sensitive Label and Structure‐consistent Unilateral Projection (CLSUP) approach for HFP. CLSUP can not only make better use of the within‐project and cross‐project data but also alleviate the class imbalance problem by setting different misclassification costs for fault‐prone and non–fault‐prone instances. Extensive experiments on 30 projects demonstrate the effectiveness of CLSUP. Zhiqiang Li 0003, Xiaoyuan Jing, Xiaoke Zhu |
Softw. Test. Verification Reliab. | 1 |
| 2017 | Heterogeneous Defect Prediction Through Multiple Kernel Learning and Ensemble LearningabstractHeterogeneous defect prediction (HDP) aims to predict defect-prone software modules in one project using heterogeneous data collected from other projects. Recently, several HDP methods have been proposed. However, these methods do not sufficiently incorporate the two characteristics of the defect prediction data: (1) data could be linearly inseparable, and (2) data could be highly imbalanced. These two data characteristics make it challenging to build an effective HDP model. In this paper, we propose a novel Ensemble Multiple Kernel Correlation Alignment (EMKCA) based approach to HDP, which takes into consideration the two characteristics of the defect prediction data. Specifically, we first map the source and target project data into high dimensional kernel space through multiple kernel leaning, where the defective and non-defective modules can be better separated. Then, we design a kernel correlation alignment method to make the data distribution of the source and target projects similar in the kernel space. Finally, we integrate multiple kernel classifiers with ensemble learning to relieve the influence caused by class imbalance problem, which can improve the accuracy of the defect prediction model. Consequently, EMKCA owns the advantages of both multiple kernel learning and ensemble learning. Extensive experiments on 30 public projects show that EMKCA outperforms the related competing methods. Zhiqiang Li 0003, Xiaoyuan Jing, Xiaoke Zhu, Hongyu Zhang 0002 |
ICSME | 1 |
| 2016 | Multi-spectral low-rank structured dictionary learning for face recognition
Xiaoyuan Jing, Fei Wu 0004, Xiaoke Zhu, Xiwei Dong, Fei Ma 0004, Zhiqiang Li 0003 |
Pattern Recognit. | 6 |
| 2016 | Multi-Label Dictionary Learning for Image AnnotationabstractImage annotation has attracted a lot of research interest, and multi-label learning is an effective technique for image annotation. How to effectively exploit the underlying correlation among labels is a crucial task for multi-label learning. Most existing multi-label learning methods exploit the label correlation only in the output label space, leaving the connection between the label and the features of images untouched. Although, recently some methods attempt toward exploiting the label correlation in the input feature space by using the label information, they cannot effectively conduct the learning process in both the spaces simultaneously, and there still exists much room for improvement. In this paper, we propose a novel multi-label learning approach, named multi-label dictionary learning (MLDL) with label consistency regularization and partial-identical label embedding MLDL, which conducts MLDL and partial-identical label embedding simultaneously. In the input feature space, we incorporate the dictionary learning technique into multi-label learning and design the label consistency regularization term to learn the better representation of features. In the output label space, we design the partial-identical label embedding, in which the samples with exactly same label set can cluster together, and the samples with partial-identical label sets can collaboratively represent each other. Experimental results on the three widely used image datasets, including Corel 5K, IAPR TC12, and ESP Game, demonstrate the effectiveness of the proposed approach. Xiaoyuan Jing, Fei Wu 0004, Zhiqiang Li 0003, Ruimin Hu, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |