VLDB 2026 Research / reviewers in the wild / expert
Lina Gong
dblp:04/10049
· DBLP profile ↗
42ranked-venue papers
8as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCA-LUT: Deep Chromatic Alignment with 5D LUT for Purple Fringing RemovalabstractPurple fringing, a persistent artifact caused by Longitudinal Chromatic Aberration (LCA) in camera lenses, has long degraded the clarity and realism of digital imaging. Traditional solutions rely on complex and expensive apochromatic (APO) lens hardware and the extraction of handcrafted features, ignoring the data-driven approach. To fill this gap, we introduce DCA-LUT, the first deep learning framework for purple fringing removal. Inspired by the physical root of the problem-the spatial misalignment of RGB color channels due to lens dispersion, we introduce a novel Chromatic-Aware Coordinate Transformation (CA-CT) module, learning an image-adaptive color space to decouple and isolate fringing into a dedicated dimension. This targeted separation allows the network to learn a precise "purple fringe channel," which then guides the accurate restoration of the luminance channel. The final color correction is performed by a learned 5D Look-Up Table (5D LUT), enabling efficient and powerful non-linear color mapping. To enable robust training and fair evaluation, we constructed a large-scale synthetic purple fringing dataset (PF-Synth). Extensive experiments in synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance in purple fringing removal. Jialang Lu, Shuning Sun, Pu Wang 0008, Chen Wu 0006, Feng Gao 0005, Lina Gong, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng |
AAAI | 6 |
| 2026 | Evaluating the Interactions between class overlap and class imbalance for software defect prediction
Ningzhong Liu, Lina Gong |
Expert Syst. Appl. | 5 |
| 2026 | Low-Cost Testing for Path Coverage of MPI Programs Using Surrogate-Assisted Changeable Multi-Objective OptimizationabstractA target path of Message Passing Interface (MPI) programs typically consists of several target sub-paths. During solving a test case that cover the target path using an intelligent optimization algorithm, we often find that there are some hard-to-cover target sub-paths, which limit the testing efficiency of the entire target path. Therefore, this paper proposes an approach of low-cost testing for path coverage of MPI programs using surrogate-assisted changeable multi-objective optimization, which is used to further improve the effectiveness and efficiency of test case generation. The proposed approach first establishes a changeable multi-objective optimization model, which is used to guide the generation of test cases. During solving the changeable multi-objective optimization model using an intelligent optimization algorithm, we then determine each hard-to-cover target sub-path and form a corresponding sample set. Finally, we manage the surrogate model corresponding to each hard-to-cover target sub-path based on the formed sample set, and select superior evolutionary individuals to really execute the MPI program under test, thus reducing the cost and times of program execution. The proposed approach has been applied to path coverage testing of several benchmark MPI programs, and compared with several state-of-the-art approaches. The experimental results show that the proposed approach significantly improves the effectiveness and efficiency of generating test cases. Baicai Sun, Lina Gong, Yinan Guo 0001, Dun-Wei Gong, Gaige Wang |
IEEE Trans. Software Eng. | 2 |
| 2025 | Dialogue Framework for Bug Issue Types Classification in Deep Learning-oriented Projects Based on Large Language ModelabstractOpen-source repository platforms have become essential for the development and collaboration of modern deep learning (DL) projects. Efficient and accurate classification of issue reports submitted during the development process is critical for enhancing project quality and development efficiency. However, compared to traditional software projects, issue reports in DL projects exhibit substantial differences in terms of error causes and symptom manifestations, making conventional classification approaches less effective and harder to deal with. To address this challenge, we propose a novel issue classification dialogue framework based on large language models (llMs), which aligns with the full lifecycle of issue handling, i.e., from issue proposal to label assignment in open-source repositories. Specifically, our framework is built upon Qwen2.5 and incorporates a multi-turn dialogue mechanism that leverages comment information from different roles to enhance context understanding through multi-source signals. It not only relies on the semantic reasoning capabilities of LLMs to analyze each step of the conversation but also introduces a built-in selfreflection mechanism to verify and refine classification decisions. We conduct extensive experiments on datasets that comprise $\mathbf{9, 0 7 3}$ issue reports from TensorFlow, PyTorch, and Caffe. The evaluation involves four LLMs and six fine-tuned pre-trained models (PTMs) and the experimental results demonstrate that the proposed dialogue framework based on Qwen2.5 significantly outperforms state-of-the-art fine-tuned PTMs in both coarse- and fine-grained issue classification, particularly in identifying specific issue types. Our findings highlight the effectiveness of LLMbased dialogue frameworks in issue classification and open up new directions for applying LLMs in software engineering tasks, encouraging further exploration into their generalizability and robustness. Zixuan Zeng, Lina Gong |
APSEC | 4 |
| 2025 | TGGI: Text-Guided Object Removal in 3D Gaussian Scenes via Multi-View Image Inpainting
Liangyuan Zhang, Zhe Zhu, Lina Gong |
CW | 3 |
| 2025 | UITrans: Seamless UI Translation from Android to HarmonyOSabstractSeamless user interface (i.e., UI) translation has emerged as a pivotal technique for modern mobile developers, addressing the challenge of developing separate UI applications for Android and HarmonyOS platforms due to fundamental differences in layout structures and development paradigms.In this paper, we present UITrans, the first automated UI translation tool designed for Android to HarmonyOS.UITrans leverages an LLM-driven multi-agent reflective collaboration framework to convert Android XML layouts into HarmonyOS ArkUI layouts.It not only maps component-level and page-level elements to ArkUI equivalents but also handles project-level challenges, including complex layouts and interaction logic.Our evaluation of six Android applications demonstrates that our UITrans achieves translation success rates of over 90.1%, 89.3%, and 89.2% at the component, page, and project levels, respectively.UITrans is available at https://github.com/OpenSELab/UITransand the demo video can be viewed at https://www.youtube.com/watch?v=iqKOSm CnJG0. Lina Gong, Yujun Huang, Mingqiang Wei |
Internetware | 1 |
| 2025 | Unraveling the Characterization and Propagation of Security Vulnerabilities in TensorFlow-based Deep Learning Software Supply ChainabstractAs a widely adopted deep learning (DL) framework, TensorFlow's vulnerabilities have the potential to affect a substantial number of DL applications throughout the software supply chain (SSC).However, Existing research lacks a comprehensive exploration of the characterization of TensorFlow vulnerabilities and the propagation of vulnerabilities on SSC.To help TensorFlow-based developers and security specialists in understanding the security risks, we construct an empirical study on 429 vulnerabilities of TensorFlow across a TensorFlow-based vulnerability SSC comprising 5790 versions across 691 affected packages constructed through the GitHub dependency graph.We observe that: 1) A predominant share (79.6%) of vulnerabilities occur in the TensorFlow's Core modules (e.g., kernels, and ops), featuring prevalent vulnerabilities such as Reachable Assertion, Out-of-bounds Read, Improper Input Validation, and NULL Pointer Dereference.Notably, Divide By Zero vulnerabilities pose significant risks in the Lite module; 2) Vulnerability co-occurrence is observed in 20 pairs involving 50 vulnerabilities, with Heap-based Buffer Overflow vulnerabilities particularly likely to coexist with other types of vulnerabilities; 3) Packages within the Large Language Models (LLM) domain, often distributed across the third and fourth layers of the SSC, are vulnerable to TensorFlow's security issues; 4) Many commonly utilized APIs (i.e.tf.constant, tf.concat, and tf.range), are implicated in TensorFlow vulnerabilities, affecting over half of all packages and, by extension, a significant portion of software within the SSC.Our findings suggest that: i)It would be better for Tensorflow developers take input validation, bounds checking, and assertion and exception handling before performing tensor operations, division operations, and pointer accesses to avoid Reachable Assertion, Divide By Zero, and memory-related vulnerabilities.ii) TensorFlow developers would be better to scrutinize the presence of memory-related vulnerabilities when encountering a vulnerability stemming from improper input validation.iii)TensorFlow-based Developers would be better to be aware of * Yiren Zhou, Lina Gong |
Internetware | 2 |
| 2025 | Coding-Fuse: Efficient Fusion of Code Pre-Trained Models for Classification TasksabstractSoftware engineering (SE) classification tasks play a vital role in improving software quality. Nevertheless, SE researchers and practitioners tend to rely on a single code pre-trained model (PTM) for downstream classification tasks. Previous studies have found that different code PTMs yield different performance in SE classification tasks, which triggers our thinking of whether the integration of multiple code PTMs improves the performance of classification tasks. Therefore, we first conduct preliminary exploratory research to analyze the impact of fusing multiple PTMs on code classification tasks. The result shows that compared to the single code PTM, the fusion of multiple code PTMs can improve the performance of SE classification tasks. However, the performance improvement also brings about the problem of increased finetuning resources and reduced application efficiency, which does not meet the greenness requirements. In order to address these issues, we propose Coding-Fuse, a framework of efficient fusion of code PTMs for SE classification tasks. Coding-Fuse first introduces evidence theory to evaluate the adaptability of the output features of each layer of code PTMs and data labels, and locates the potential best performance layer of different code PTMs. Then, Coding-Fuse uses a soft voting strategy to fuse the outputs of these layers to obtain a new model. We conduct experiments for effectiveness by comparing Coding-Fuse with the full PTM fusion method and the original single PTM using five different code PTMs on three different SE classification tasks and two task scenarios. The results show that Coding-Fuse can achieve better performance than the full PTM fusion method with higher efficiency and fewer hardware resources, and can achieve better performance than the original single PTM at the same efficiency and hardware resource level. We encourage SE practitioners to use our Coding-Fuse method in practice to fully utilize the advantages of each code PTM in the PTM repository according to task requirements to easily create new SE intelligent PTMs to achieve performance and greenness improvements. Lina Gong, Mingqiang Wei |
ASE | 2 |
| 2025 | MCL-VD: Multi-modal contrastive learning with LoRA-enhanced GraphCodeBERT for effective vulnerability detection
Xiaolin Ju, Xiang Chen 0005, Lina Gong |
Autom. Softw. Eng. | 4 |
| 2025 | An empirical study of best practices for code pre-trained models on software engineering classification tasks
Lina Gong, Yaoshen Yu, Mingqiang Wei |
Expert Syst. Appl. | 2 |
| 2025 | Software Defect Prediction Based on Fuzzy Cost Broad Learning SystemabstractSoftware defect prediction (SDP) is an effective approach to ensure software reliability. Machine learning models have been widely employed in SDP, but they ignore the impact of class imbalance, noise and outliers on the prediction performance. This study proposes a fuzzy cost broad learning system (FC‐BLS). FC‐BLS not only handles class imbalance problems but also considers the specific sample distribution to address noise and outliers in software defect datasets. Our approach draws fully on the idea of the cost matrix and fuzzy membership functions. It introduces them to BLS, where the cost matrix prioritises the training errors on the minority samples. Hence, the classification hyperplane position is more reasonable, and fuzzy membership functions calculate the membership degree of the sample in a feature mapping space to remove the prediction error caused by noise and outlier samples. Then, the optimisation problem is constructed based on the idea that the minority class and normal instances have relatively high costs. By contrast, the majority class and noise and outlier instances have relatively small costs. This study conducted experiments on nine NASA SDP datasets, and the experimental findings demonstrated the effectiveness of the proposed methodology on most datasets. Heling Cao, Zhiying Cui, Yonghe Chu, Lina Gong, Guangen Liu, Yun Wang 0009, Fangchao Tian, Haoyang Ge |
Int. J. Intell. Syst. | 4 |
| 2025 | JIT-CF: Integrating contrastive learning with feature fusion for enhanced just-in-time defect prediction
Xiaolin Ju, Xiang Chen 0005, Lina Gong, Vaskar Chakma, Xin Zhou 0014 |
Inf. Softw. Technol. | 4 |
| 2025 | FedMVA: Enhancing software vulnerability assessment via federated multimodal learning
Qingyun Liu 0014, Xiaolin Ju, Xiang Chen 0005, Lina Gong |
J. Syst. Softw. | 4 |
| 2025 | Rethinking mixture of rain removal via depth-guided adversarial learning
Yongzhen Wang 0001, Xuefeng Yan 0001, Yanbiao Niu, Lina Gong, Yanwen Guo 0001, Mingqiang Wei |
Neural Networks | 4 |
| 2025 | Demystifying the Impact of Open-Source Machine Learning Libraries on Software AnalyticsabstractMachine learning (ML) classification techniques from various libraries have been widely introduced into software engineering (SE) to mine instructive insights, which help developers guarantee software quality. However, researchers would report instructive insights with no clear stability and consensus due to the indiscriminate use of various ML libraries. Such a lack of directive on using various ML libraries prevents developers from effectively applying instructive insights in practice. Therefore, through a case study of 23 popular software datasets across three task domains (i.e., software defect prediction, issue lifetime estimation, and code smell detection) in SE, in this article, we systematically study the impact of open-source ML libraries on performance consistency, performance stability, and model interpretation of six commonly used classifiers across two commonly used ML programming language (i.e., Python and R). We find that for a given classification technique: ML libraries with the tune setting cannot generate stable and consistent performance; ML libraries are sensitive in the interpretation of the model, even in case of the same parameter settings; and ML libraries from R would generate more stable and higher performance. Based on these findings, we suggest that future work in SE should indicate the specific ML libraries with the specific parameter settings that are used to discover the instructive insights for their tasks; try the ML libraries from R to build the models with high and stable performance for the software tasks; and use the same ML libraries that are used to select important features to construct the classification model. Yihui Gong, Lina Gong, Shujuan Jiang |
IEEE Trans. Reliab. | 3 |
| 2025 | PointCG: Self-Supervised Point Cloud Learning via Joint Completion and GenerationabstractThe core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points' representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance. Yun Liu 0002, Peng Li 0064, Xuefeng Yan 0001, Liangliang Nan, Bing Wang 0013, Honghua Chen, Lina Gong, Wei Zhao 0039, Mingqiang Wei |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Classifying Bug Issue Types for Deep Learning-Oriented Projects with Pre-Trained ModelabstractClassifying the bug issue types correctly plays a vital role in improving the quality of the deep learning (DL)-oriented projects. Although prior studies have proposed different approaches based on Pre-Trained Models (PTMs) for issue type classification in traditional GitHub repositories, DL-oriented projects are different from traditional software, especially in terms of bugs with different causes and symptoms. More importantly, these PTMs-based approaches trained on the issue reports are labeled when software users submit, which would be wrong and non-subdivided bug issue types. Therefore, an automated approach with the ground-truth bug issue types for labeling issues in DL-oriented projects is necessary for DL software repositories. To fill these gaps, we first manually labeled 9,073 issue reports from 11 DL-oriented projects as the ground truths to establish authentic labels. We then explore the effectiveness of six PTMs on the bug issues identification for the DL software repository. Our findings indicate that i) PTMs (especially BERT) could identify more precise bug issue types of DL software than prior DL approaches in all the datasets. ii) contrary to their performance in traditional software bug classification tasks, Software Engineering (SE) domain-specific PTMs cannot achieve significantly better performance than our compared general PTMs and may even perform worse for the DL bug issue classification. iii) in the cross-framework scenarios, the Fl-score of PTMs declined by 18.5% to 19.8%. Despite that the performance is suffered, BERT can still achieve the best results. Conclusively, we propose that PTM-based bug issue classification offers potential for more widespread applications and prompt future studies to further examine and verify the generalizability of PTM-based methods in software engineering. Zixuan Zeng, Lina Gong |
APSEC | 3 |
| 2024 | Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?abstractVulnerability detection is garnering increasing attention in software engineering, since code vulnerabilities possibly pose significant security. Recently, reusing various code pre-trained models (e.g., CodeBERT, CodeT5, and CodeGen) has become common for code embedding without providing reasonable justifications in vulnerability detection. The premise for casually utilizing pre-trained models (PTMs) is that the code embeddings generated by different PTMs would generate a similar impact on the performance. Is that TRUE? To answer this important question, we systematically investigate the effects of code embedding generated by ten different code PTMs on the performance of vulnerability detection, and get the answer, i.e., that is NOT true. We observe that code embedding generated by various code PTMs can indeed influence the performance and selecting an embedding technique based on parameter scales and embedding dimension is not reliable. Our findings highlight the necessity of quantifying and evaluating the characteristics of code embedding generated by various code PTMs to understand the effects. To achieve this goal, we analyze the numerical representation and data distribution of code embedding generated by different PTMs to evaluate differences and characteristics. Based on these insights, we propose Coding-PTMs, a recommendation framework to assist engineers in selecting optimal code PTMs for their specific vulnerability detection tasks. Specifically, we define thirteen code embedding metrics across three dimensions (i.e., statistics, norm, and distribution) for constructing a specialized code PTM recommendation dataset. We then employ a Random Forest classifier to train a recommendation model and identify the optimal code PTMs from the candidate model zoo. We encourage engineers to use our Coding-PTMs to evaluate the characteristics of code embeddings generated by candidate code PTMs on the performance and recommend optimal code PTMs for code embedding in their vulnerability detection tasks. Lina Gong, Mingqiang Wei, Fei Wu 0001 |
ASE | 2 |
| 2024 | An extensive study of the effects of different deep learning models on code vulnerability detection in Python code
Rongcun Wang, Senlei Xu, Xingyu Ji, Yuan Tian 0008, Lina Gong |
Autom. Softw. Eng. | 5 |
| 2024 | Basis path coverage testing of MPI programs based on multi-task evolutionary optimization
Baicai Sun, Lina Gong, Yinan Guo 0001, Dun-Wei Gong |
Expert Syst. Appl. | 2 |
| 2024 | How accessibility affects other quality attributes of software? A case study of GitHub
Yaxin Zhao, Lina Gong, Wenhua Yang 0001, Yu Zhou 0010 |
Sci. Comput. Program. | 2 |
| 2024 | MR${}^{2}$ 2-KG: A Multi-Relation Multi-Rationale Knowledge Graph for Modeling Software Engineering Knowledge on Stack OverflowabstractStack Overflow is a knowledge sharing platform where its users create and share informative content from both inside and outside the site. Prior studies have leveraged the relation across Stack Overflow posts through internal links to build services and applications to enhance the accessibility of knowledge. However, they focused on studying a knowledge unit that consists of a question post and all the associated answer posts to represent the relation. It is unknown whether such representation of knowledge on Stack Overflow could comprehensively model various complex relations among webpages, such as questions, answers, internal and external links. In addition, the rationales behind sharing knowledge on Stack Overflow have yet to be explored among distinct user groups, such as askers, answerers, readers who wish to learn. Thus, in this study, we first investigate the real-world characteristics of Stack Overflow knowledge by abstracting the complex knowledge representation into relations among its building blocks. We observe that a question thread includes three basic knowledge relations to reassemble into complex knowledge, that is, the hierarchy relation within the associated answers in a question, the coupling relation between knowledge artifacts (i.e., question or answer posts) through internal links, and the complimentary relation between Stack Overflow posts and external websites. All these three basic knowledge relations are informative and could be caused by different rationales when the crowdsourced knowledge is shared on Stack Overflow. Our findings highlight that it is necessary to propose a comprehensive knowledge graph to represent the real-world knowledge on Stack Overflow. Therefore, we further propose a Multi-Relation Multi-Rationale Knowledge Graph (MR2-KG), whose nodes represent questions, answers, and external webpages. Edges in the MR2-KG represent the rationales included in the three structures (i.e., question answering, duplicate, priori, posterior, parallelism, containment, and working examples knowledge). In addition, we develop an automated approach to model the nodes and edges to represent Stack Overflow knowledge associated with a question thread. Our case study shows that the automated knowledge representation generation can achieve an ROC AUC of 96% and MCC of 89% to identify edges in the MR2-KG. To further evaluate the applicability of MR2-KG, we develop an answer generator to help developers efficiently identify the answers that meet their intent. Our user study of 100 real-world Java questions indicates the usefulness of MR2-KG. Finally, we discuss the implications of our findings for developers, researchers, and Stack Overflow moderators. Lina Gong, Haoxiang Zhang 0001 |
IEEE Trans. Software Eng. | 1 |
| 2024 | PointSee: Image Enhances Point CloudabstractThere is a prevailing trend towards fusing multi-modal information for 3D object detection (3OD). However, challenges related to computational efficiency, plug-and-play capabilities, and accurate feature alignment have not been adequately addressed in the design of multi-modal fusion networks. In this paper, we present PointSee, a lightweight, flexible, and effective multi-modal fusion solution to facilitate various 3OD networks by semantic feature enhancement of point clouds (e.g., LiDAR or RGB-D data) assembled with scene images. Beyond the existing wisdom of 3OD, PointSee consists of a hidden module (HM) and a seen module (SM): HM decorates point clouds using 2D image information in an offline fusion manner, leading to minimal or even no adaptations of existing 3OD networks; SM further enriches the point clouds by acquiring point-wise representative semantic features, leading to enhanced performance of existing 3OD networks. Besides the new architecture of PointSee, we propose a simple yet efficient training strategy, to ease the potential inaccurate regressions of 2D object detection networks. Extensive experiments on the popular outdoor/indoor benchmarks show quantitative and qualitative improvements of our PointSee over thirty-five state-of-the-art methods. Lipeng Gu, Xuefeng Yan 0001, Peng Cui 0013, Lina Gong, Haoran Xie 0001, Fu Lee Wang, Harry Qin, Mingqiang Wei |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | ifUNet++: Iterative Feedback UNet++ for Infrared Small Target DetectionabstractSmall targets are often submerged in the cluttered backgrounds of infrared images. In this paper, we propose an iterative feedback UNet++ for infrared small target detection, dubbed ifUNet++. Unlike most of existing methods, ifU-Net++ enables to concentrate on small targets while weakening the interference of clutter backgrounds. ifUNet++ contains two parts: a simplified UNet++ and an iterative feedback strategy. We reduce the unnecessary nodes of UNet++ and have the simplified UNet++ as our backbone network, avoiding the loss of infrared small targets. Based on the simplified network, we search the infrared small targets in an iterative feedback manner, avoiding the interference of cluttered backgrounds. Besides, to optimize the iterative results, we propose Contextual Multiple Attention (CMA) to enhance the features in each iteration. Experimental results exhibit the clear promotion of ifUNet++ over eight state-of-the-art methods, in terms of noise-robustness and detection accuracy. Zhangying Weng, Peng Li 0064, Xin Zhuang, Xuefeng Yan 0001, Lina Gong, Haoran Xie 0001, Mingqiang Wei |
ICASSP | 5 |
| 2023 | BTLink : automatic link recovery between issues and commits based on pre-trained BERT model
Jinpeng Lan, Lina Gong, Haoxiang Zhang 0001 |
Empir. Softw. Eng. | 2 |
| 2023 | The impact of feature selection techniques on effort-aware defect prediction: An empirical studyabstractAbstract Effort‐Aware Defect Prediction (EADP) methods sort software modules based on the defect density and guide the testing team to inspect the modules with high defect density first. Previous studies indicated that some feature selection methods could improve the performance of Classification‐Based Defect Prediction (CBDP) models, and the Correlation‐based feature subset selection method with the Best First strategy (CorBF) performed the best. However, the practical benefits of feature selection methods on EADP performance are still unknown, and blindly employing the best‐performing CorBF method in CBDP to pre‐process the defect datasets may not improve the performance of EADP models but possibly result in performance degradation. To assess the impact of the feature selection techniques on EADP, a total of 24 feature selection methods with 10 classifiers embedded in a state‐of‐the‐art EADP model (CBS+) on the 41 PROMISE defect datasets were examined. We employ six evaluation metrics to assess the performance of EADP models comprehensively. The results show that (1) The impact of the feature selection methods varies in classifiers and datasets. (2) The four wrapper‐based feature subset selection methods with forwards search, that is, AdaBoost with Forwards Search, Deep Forest with Forwards Search, Random Forest with Forwards Search, and XGBoost with Forwards Search (XGBF) are better than other methods across the studied classifiers and the used datasets. And XGBF with XGBoost as the embedded classifier in CBS+ performs the best on the datasets. (3) The best‐performing CorBF method in CBDP does not perform well on the EADP task. (4) The selected features vary with different feature selection methods and different datasets, and the features noc (number of children), ic (inheritance coupling), cbo (coupling between object classes), and cbm (coupling between methods) are frequently selected by the four wrapper‐based feature subset selection methods with forwards search. (5) Using AdaBoost, deep forest, random forest, and XGBoost as the base classifiers embedded in CBS+ can achieve the best performance. In summary, we recommend the software testing team should employ XGBF with XGBoost as the embedded classifier in CBS+ to enhance the EADP performance. Wanpeng Lu, Jacky W. Keung, Xiao Yu 0008, Lina Gong |
IET Softw. | 5 |
| 2023 | ImLiDAR: Cross-Sensor Dynamic Message Propagation Network for 3-D Object DetectionabstractLiDAR and camera, as two different sensors, supply geometric (point clouds) and semantic (RGB images) information of 3-D scenes. However, it is still challenging for existing methods to fuse data from the two cross sensors, making them complementary for quality 3-D object detection (3OD). We propose ImLiDAR, a new 3OD paradigm to narrow the cross-sensor discrepancies by progressively fusing the multiscale features of camera Images and LiDAR point clouds. ImLiDAR enables to provide the detection head with cross-sensor yet robustly fused features. To achieve this, two core designs exist in ImLiDAR. First, we propose a cross-sensor dynamic message propagation (CDMP) module to combine the best of the multiscale image and point features. Second, we raise a direct set prediction problem that allows designing an effective set-based detector (SD) to tackle the inconsistency of the classification and localization confidences, and the sensitivity of hand-tuned hyperparameters. Besides, the novel SD can be detachable and easily integrated into various detection networks. Comparisons on the KITTI, nuScenes, and SUN-RGBD datasets all show clear visual and numerical improvements of our ImLiDAR over 45 state-of-the-art 3OD methods. Yiyang Shen, Rongwei Yu, Haoran Xie 0001, Lina Gong, Harry Qin, Mingqiang Wei |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | What Is the Intended Usage Context of This Model? An Exploratory Study of Pre-Trained Models on Various Model RepositoriesabstractThere is a trend of researchers and practitioners to directly apply pre-trained models to solve their specific tasks. For example, researchers in software engineering (SE) have successfully exploited the pre-trained language models to automatically generate the source code and comments. However, there are domain gaps in different benchmark datasets. These data-driven (or machine learning based) models trained on one benchmark dataset may not operate smoothly on other benchmarks. Thus, the reuse of pre-trained models introduces large costs and additional problems of checking whether arbitrary pre-trained models are suitable for the task-specific reuse or not. To our knowledge, software engineers can leverage code contracts to maximize the reuse of existing software components or software services. Similar to the software reuse in the SE field, reuse SE could be extended to the area of pre-trained model reuse. Therefore, according to the model card’s and FactSheet’s guidance for suppliers of pre-trained models on what information they should be published, we propose model contracts including the pre- and post-conditions of pre-trained models to enable better model reuse. Furthermore, many non-trivial yet challenging issues have not been fully investigated, although many pre-trained models are readily available on the model repositories. Based on our model contract, we conduct an exploratory study of 1908 pre-trained models on six mainstream model repositories (i.e., the TensorFlow Hub, PyTorch Hub, Model Zoo, Wolfram Neural Net Repository, Nvidia, and Hugging Face) to investigate the gap between necessary pre- and post-condition information and actual specifications. Our results clearly show that (1) the model repositories tend to provide confusing information of the pre-trained models, especially the information about the task’s type, model, training set, and (2) the model repositories cannot provide all of our proposed pre/post-condition information, especially the intended use, limitation, performance, and quantitative analysis. On the basis of our new findings, we suggest that (1) the developers of model repositories shall provide some necessary options (e.g., the training dataset, model algorithm, and performance measures) for each of pre/post-conditions of pre-trained models in each task type, (2) future researchers and practitioners provide more efficient metrics to recommend suitable pre-trained model, and (3) the suppliers of pre-trained models should report their pre-trained models in strict accordance with our proposed pre/post-condition and report their models according to the characteristics of each condition that has been reported in the model repositories. Lina Gong, Mingqiang Wei, Haoxiang Zhang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | An Accurate Identifier Renaming Prediction and Suggestion ApproachabstractIdentifiers play an important role in helping developers analyze and comprehend source code. However, many identifiers exist that are inconsistent with the corresponding code conventions or semantic functions, leading to flawed identifiers. Hence, identifiers need to be renamed regularly. Even though researchers have proposed several approaches to identify identifiers that need renaming and further suggest correct identifiers for them, these approaches only focus on a single or a limited number of granularities of identifiers without universally considering all the granularities and suggest a series of sub-tokens for composing identifiers without completely generating new identifiers. In this article, we propose a novel identifier renaming prediction and suggestion approach. Specifically, given a set of training source code, we first extract all the identifiers in multiple granularities. Then, we design and extract five groups of features from identifiers to capture inherent properties of identifiers themselves and the relationships between identifiers and code conventions, as well as other related code entities, enclosing files, and change history. By parsing the change history of identifiers, we can figure out whether specific identifiers have been renamed or not. These identifier features and their renaming history are used to train a Random Forest classifier, which can be further used to predict whether a given new identifier needs to be renamed or not. Subsequently, for the identifiers that need renaming, we extract all the related code entities and their renaming change history. Based on the intuition that identifiers are co-evolved as their relevant code entities with similar patterns and renaming sequences, we could suggest and recommend a series of new identifiers for those identifiers. We conduct extensive experiments to validate our approach in both the Java projects and the Android projects. Experimental results demonstrate that our approach could identify identifiers that need renaming with an average F-measure of more than 89%, which outperforms the state-of-the-art approach by 8.30% in the Java projects and 21.38% in the Android projects. In addition, our approach achieves a Hit@10 of 48.58% and 40.97% in the Java and Android projects in suggesting correct identifiers and outperforms the state-of-the-art approach by 29.62% and 15.75%, respectively. Junpeng Luo, Jiahui Liang, Lina Gong |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | A Comprehensive Investigation of the Impact of Class Overlap on Software Defect PredictionabstractSoftware Defect Prediction (SDP) is one of the most vital and cost-efficient operations to ensure the software quality. However, there exists the phenomenon of class overlap in the SDP datasets (i.e., defective and non-defective modules are similar in terms of values of metrics), which hinders the performance as well as the use of SDP models. Even though efforts have been made to investigate the impact of removing overlapping technique on the performance of SDP, many open issues are still challenging yet unknown. Therefore, we conduct an empirical study to comprehensively investigate the impact of class overlap on SDP. Specifically, we first propose an overlapping instances identification approach by analyzing the class distribution in the local neighborhood of a given instance. We then investigate the impact of class overlap and two common overlapping instance handling techniques on the performance and the interpretation of seven representative SDP models. Through an extensive case study on 230 diversity datasets, we observe that: i) 70.0% of SDP datasets contain overlapping instances; ii) different levels of class overlap have different impacts on the performance of SDP models; iii) class overlap affects the rank of the important feature list of SDP models, particularly the feature lists at the top 2 and top 3 ranks; IV) Class overlap handling techniques could statistically significantly improve the performance of SDP models trained on datasets with over 12.5% overlap ratios. We suggest that future work should apply our KNN method to identify the overlap ratios of datasets before building SDP models. Lina Gong, Haoxiang Zhang 0001, Mingqiang Wei |
IEEE Trans. Software Eng. | 1 |
| 2023 | BEQAIN: An Effective and Efficient Identifier Normalization Approach With BERT and the Question Answering SystemabstractAs one of the most important resources to express the semantics of source code, identifiers are usually composed of several common or domain-specific terms and abbreviations, thus heavily hindering developers from analyzing and comprehending source code. Hence, it is very necessary to normalize identifiers, which aims to align the vocabulary found in identifiers with natural language words found in other software artifacts. Even though researchers have proposed several identifier normalization approaches in the literature, these approaches only rely on the lexical information in identifiers and related source code entities to normalize identifiers, suffering from the lack of deep semantic understanding of identifiers. In this paper, we propose an effective and efficient identifier normalization approach BEQAIN to split identifiers into their composing words and expand the enclosed abbreviations. Specifically, BEQAIN employs a deep learning model, which is mainly composed of a Bidirectional Encoder Representation from Transformers (BERT) layer and a Conditional Random Fields (CRF) layer to embed identifiers into low-level vectors and learn the identifier splitting patterns. The BERT-CRF network is also combined with a pre-processing component and a post-processing component to resolve the problems of over-splitting and under-splitting so as to improve the identifier splitting performance. Furthermore, BEQAIN also employs a Question Answering (Q&A) system to learn the abbreviation expansion mappings and leverages the current programming context to determine the exactly correct expansion when there are multiple expansions for specific abbreviations. After BEQAIN is fully trained, it can be used to normalize identifiers. We conduct extensive experiments to validate the effectiveness and efficiency of BEQAIN over two publicly available datasets with nine projects. Experimental results show that BEQAIN achieves the overall average Accuracy of 80.20% and outperforms the existing state-of-the-art approach by 9.88% in normalizing identifiers. The pre-processing and post-processing components could improve the Accuracy of BEQAIN in identifier splitting by 11.70%. Employing the programming context information could improve the Accuracy of BEQAIN in abbreviation expansion by 11.15% on average. In addition, the average normalization time of BEQAIN is less than one second. Finally, we also discuss some observations for the road ahead for identifier normalization to inspire other researchers. Lina Gong, Haoxiang Zhang 0001, He Jiang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2022 | UTOPIC: Uncertainty-aware Overlap Prediction Network for Partial Point Cloud RegistrationabstractAbstract High‐confidence overlap prediction and accurate correspondences are critical for cutting‐edge models to align paired point clouds in a partial‐to‐partial manner. However, there inherently exists uncertainty between the overlapping and non‐overlapping regions, which has always been neglected and significantly affects the registration performance. Beyond the current wisdom, we propose a novel uncertainty‐aware overlap prediction network, dubbed UTOPIC, to tackle the ambiguous overlap prediction problem; to our knowledge, this is the first to explicitly introduce overlap uncertainty to point cloud registration. Moreover, we induce the feature extractor to implicitly perceive the shape knowledge through a completion decoder, and present a geometric relation embedding for Transformer to obtain transformation‐invariant geometry‐aware feature representations. With the merits of more reliable overlap scores and more precise dense correspondences, UTOPIC can achieve stable and accurate registration results, even for the inputs with limited overlapping areas. Extensive quantitative and qualitative experiments on synthetic and real benchmarks demonstrate the superiority of our approach over state‐of‐the‐art methods. Zhilei Chen, Honghua Chen, Lina Gong, Xuefeng Yan 0001, Jun Wang 0039, Yanwen Guo 0001, Harry Qin, Mingqiang Wei |
Comput. Graph. Forum | 3 |
| 2022 | Contrastive Semantic-Guided Image Smoothing NetworkabstractAbstract Image smoothing is a fundamental low‐level vision task that aims to preserve salient structures of an image while removing insignificant details. Deep learning has been explored in image smoothing to deal with the complex entanglement of semantic structures and trivial details. However, current methods neglect two important facts in smoothing: 1) naive pixel‐level regression supervised by the limited number of high‐quality smoothing ground‐truth could lead to domain shift and cause generalization problems towards real‐world images; 2) texture appearance is closely related to object semantics, so that image smoothing requires awareness of semantic difference to apply adaptive smoothing strengths. To address these issues, we propose a novel Contrastive Semantic‐Guided Image Smoothing Network (CSGIS‐Net) that combines both contrastive prior and semantic prior to facilitate robust image smoothing. The supervision signal is augmented by leveraging undesired smoothing effects as negative teachers, and by incorporating segmentation tasks to encourage semantic distinctiveness. To realize the proposed network, we also enrich the original VOC dataset with texture enhancement and smoothing labels, namely VOC‐smooth, which first bridges image smoothing and semantic segmentation. Extensive experiments demonstrate that the proposed CSGIS‐Net outperforms state‐of‐the‐art algorithms by a large margin. Code and dataset are available at https://github.com/wangjie6866/CSGIS-Net . Jie Wang 0069, Yongzhen Wang 0001, Yidan Feng, Lina Gong, Xuefeng Yan 0001, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei |
Comput. Graph. Forum | 4 |
| 2022 | TogetherNet: Bridging Image Restoration and Object Detection Together via Dynamic Enhancement LearningabstractAbstract Adverse weather conditions such as haze, rain, and snow often impair the quality of captured images, causing detection networks trained on normal images to generalize poorly in these scenarios. In this paper, we raise an intriguing question – if the combination of image restoration and object detection, can boost the performance of cutting‐edge detectors in adverse weather conditions. To answer it, we propose an effective yet unified detection paradigm that bridges these two subtasks together via dynamic enhancement learning to discern objects in adverse weather conditions, called TogetherNet. Different from existing efforts that intuitively apply image dehazing/deraining as a pre‐processing step, TogetherNet considers a multi‐task joint learning problem. Following the joint learning scheme, clean features produced by the restoration network can be shared to learn better object detection in the detection network, thus helping TogetherNet enhance the detection capacity in adverse weather conditions. Besides the joint learning architecture, we design a new Dynamic Transformer Feature Enhancement module to improve the feature extraction and representation capabilities of TogetherNet. Extensive experiments on both synthetic and real‐world datasets demonstrate that our TogetherNet outperforms the state‐of‐the‐art detection approaches by a large margin both quantitatively and qualitatively. Source code is available at https://github.com/yz-wang/TogetherNet . Yongzhen Wang 0001, Xuefeng Yan 0001, Kaiwen Zhang 0011, Lina Gong, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei |
Comput. Graph. Forum | 4 |
| 2022 | Solving the last mile problem in logistics: A mobile edge computing and blockchain-based unmanned aerial vehicle delivery systemabstractSummary The “last mile” problem in logistics is challenging due to its low efficiency and high cost. To address this problem, Unmanned Aerial Vehicle (UAV) delivery such as drone delivery has been proposed and widely accepted as a promising solution. However, currently most of the existing UAV delivery systems are based on Cloud Computing which cannot efficiently meet the requirements of many real‐time services in UAV delivery systems. Meanwhile, the security issues in UAV delivery systems also raise critical concerns due to the existence of multiple participants (such as the sender, middler, and receiver) who may not maintain a mutual trust relationship among them. How to secure the UAV delivery process in such an untrusted environment is still a challenging issue. In this paper, we propose a Mobile Edge Computing (MEC) and blockchain‐based UAV delivery system to resolve the “last mile” problem in logistics. Specifically, based on the MEC architecture, the blockchain nodes are deployed on the edge nodes to facilitate and secure the UAV delivery process. To verify the effectiveness of our proposed solution, a MEC‐based UAV delivery system prototype with a private blockchain on the Ethereum platform is implemented. Through the security analysis and performance evaluation, it is proven that our proposed solution can effectively solve the “last mile” problem and address the security issues in UAV delivery systems. Xuejun Li 0001, Lina Gong, Xiao Liu 0004, Frank Jiang 0001, Wenyu Shi, Lingmin Fan, Rui Li 0013, Jia Xu 0010 |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Detecting Occluded and Dense Trees in Urban Terrestrial Views With a High-Quality Tree Detection DatasetabstractUrban trees are often densely planted along the two sides of a street. When observing these trees from a fixed view, they are inevitably occluded with each other and the passing vehicles. The high density and occlusion of urban tree scenes significantly degrade the performance of object detectors. This paper raises an intriguing learning-related question – if a module is developed to enable the network to adaptively cope with occluded and un-occluded regions while enhancing its feature extraction capabilities, can the performance of a cutting-edge detection model be improved? To answer it, a lightweight yet effective object detection network is proposed for discerning occluded and dense urban trees, called OD-UTDNet. The main contribution is a newly-designed Dilated Attention Cross Stage Partial (DACSP) module. DACSP can expand the fields-of-view of OD-UTDNet for paying more attention to the un-occluded region, while enhancing the network’s feature extraction ability in the occluded region. This work further explores both the self-calibrated convolution module and GFocal loss, which enhance the OD-UTDNet’s ability to resolve the challenging problem of high densities and occlusions. Finally, to facilitate the detection task of urban trees, a high-quality urban tree detection dataset is established, named UTD; to our knowledge, this is the first time. Extensive experiments show clear improvements of the proposed OD-UTDNet over twelve representative object detectors on UTD. The code and dataset are available at https://github.com/yzwang/OD-UTDNet. Yongzhen Wang 0001, Xuefeng Yan 0001, Hexiang Bao, Yiping Chen 0002, Lina Gong, Mingqiang Wei, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Revisiting the Impact of Dependency Network Metrics on Software Defect PredictionabstractSoftware dependency network metrics extracted from the dependency graph of the software modules by the application of Social Network Analysis (SNA metrics) have been shown to improve the performance of the Software Defect prediction (SDP) models. However, the relative effectiveness of these SNA metrics over code metrics in improving the performance of the SDP models has been widely debated with no clear consensus. Furthermore, some of the common SDP scenarios like predicting the number of defects in a module (Defect-count) in Cross-version and Cross-project SDP contexts remain unexplored. Such lack of clear directive on the effectiveness of SNA metrics when compared to the widely used code metrics prevents us from potentially building better performing SDP models. Therefore, through a case study of 9 open source software projects across 30 versions, we study the relative effectiveness of SNA metrics when compared to code metrics across 3 commonly used SDP contexts (Within-project, Cross-version and Cross-project) and scenarios (Defect-count, Defect-classification (classifying if a module is defective) and Effort-aware (ranking the defective modules w.r.t to the involved effort)). We find the SNA metrics by themselves or along with code metrics improve the performance of SDP models over just using code metrics on 5 out of the 9 studied SDP scenarios (three SDP scenarios across three SDP contexts). However, we note that in some cases the improvements afforded by considering SNA metrics over or alongside code metrics might only be marginal, whereas in other cases the improvements could be potentially large. Based on these findings we suggest that the future work should: consider SNA metrics alongside code metrics in their SDP models; as well as consider Ego metrics and Global metrics, the two different types of the SNA metrics separately when training SDP models as they behave differently. Lina Gong, Gopi Krishnan Rajbahadur, Ahmed E. Hassan, Shujuan Jiang |
IEEE Trans. Software Eng. | 1 |
| 2020 | A Novel Class-Imbalance Learning Approach for Both Within-Project and Cross-Project Defect PredictionabstractSoftware defect prediction (SDP) is an available way to enhance test efficiency and guarantee software reliability. However, there are more clean instances than defective instances in real software projects, and this results in severe class distribution skews and gets the poor performance of classifiers. So solving the class-imbalance problem in SDP has attracted growing attention from industry and academia in software engineering. In this paper, we propose a novel class-imbalance learning approach for both within-project and cross-project class-imbalance problem. We utilize the thought of stratification embedded in nearest neighbor (STr-NN) to produce evolving training datasets with balanced data. For within-project, we directly employ the STr-NN approach for defect prediction. For cross-project, we first introduce transfer component analysis to mitigate the distribution differences between source and target dataset, and then employ the STr-NN approach on the transferred data. We conduct experiments on PROMISE and NASA datasets using ensemble learning based on weight vote. Experimental results indicate that our approach has higher area under curve (AUC), Recall and comparable probability of a false alarm (pf), and F-measure than some existing methods for the class-imbalance problem. Lina Gong, Shujuan Jiang, Lili Bo, Li Jiang 0015, Junyan Qian |
IEEE Trans. Reliab. | 1 |
| 2019 | Mobility-Aware Workflow Offloading and Scheduling Strategy for Mobile Edge Computing
Jia Xu 0010, Xuejun Li 0001, Xiao Liu 0004, Chong Zhang 0007, Lingmin Fan, Lina Gong |
ICA3PP (2) | 6 |
| 2019 | Empirical Evaluation of the Impact of Class Overlap on Software Defect PredictionabstractSoftware defect prediction (SDP) utilizes the learning models to detect the defective modules in project, and their performance depends on the quality of training data. The previous researches mainly focus on the quality problems of class imbalance and feature redundancy. However, training data often contains some instances that belong to different class but have similar values on features, and this leads to class overlap to affect the quality of training data. Our goal is to investigate the impact of class overlap on software defect prediction. At the same time, we propose an improved K-Means clustering cleaning approach (IKMCCA) to solve both the class overlap and class imbalance problems. Specifically, we check whether K-Means clustering cleaning approach (KMCCA) or neighborhood cleaning learning (NCL) or IKMCCA is feasible to improve defect detection performance for two cases (i) within-project defect prediction (WPDP) (ii) cross-project defect prediction (CPDP). To have an objective estimate of class overlap, we carry out our investigations on 28 open source projects, and compare the performance of state-of-the-art learning models for the above-mentioned cases by using IKMCCA or KMCCA or NCL VS. without cleaning data. The experimental results make clear that learning models obtain significantly better performance in terms of balance, Recall and AUC for both WPDP and CPDP when the overlapping instances are removed. Moreover, it is better to consider both class overlap and class imbalance. Lina Gong, Shujuan Jiang, Rongcun Wang |
ASE | 1 |
| 2019 | FogWorkflowSim: An Automated Simulation Toolkit for Workflow Performance Evaluation in Fog ComputingabstractWorkflow underlies most process automation software, such as those for product lines, business processes, and scientific computing. However, current Cloud Computing based workflow systems cannot support real-time applications due to network latency, which limits their application in many IoT systems such as smart healthcare and smart traffic. Fog Computing extends the Cloud by providing virtualized computing resources close to the End Devices so that the response time of accessing computing resources can be reduced significantly. However, how to most effectively manage heterogeneous resources and different computing tasks in the Fog is a big challenge. In this paper, we introduce "FogWorkflowSim" an efficient and extensible toolkit for automatically evaluating resource and task management strategies in Fog Computing with simulated user-defined workflow applications. Specifically, FogWorkflowSim is able to: 1) automatically set up a simulated Fog Computing environment for workflow applications; 2) automatically execute user submitted workflow applications; 3) automatically evaluate and compare the performance of different computation offloading and task scheduling strategies with three basic performance metrics, including time, energy and cost. FogWorkflowSim can serve as an effective experimental platform for researchers in Fog based workflow systems as well as practitioners interested in adopting Fog Computing and workflow systems for their new software projects. (Demo video: https://youtu.be/AsMovcuSkx8). Xiao Liu 0004, Lingmin Fan, Jia Xu 0010, Xuejun Li 0001, Lina Gong, John C. Grundy, Yun Yang 0001 |
ASE | 5 |
| 2019 | An improved transfer adaptive boosting approach for mixed-project defect predictionabstractAbstract Software defect prediction (SDP) has been a very important research topic in software engineering, since it can provide high‐quality results when given sufficient historical data of the project. Unfortunately, there are not abundant data to bulid the defect prediction model at the beginning of a project. For this scenario, one possible solution is to use data from other projects in the same company. However, using these data practically would get poor performance because of different distributional characteristics among projects. Also, software has more non‐defective instances than defective instances that may cause a significant bias towards defective instances. Considering these two problems, we propose an improved transfer adaptive boosting (ITrAdaBoost) approach for being given a small number of labeled data in the testing project. In our approach, ITrAdaBoost can not only employ the Matthews correlation coefficient (MCC) as the measure instead of accuracy rate but also use the asymmetric misclassification costs for non‐defective and defective instances. Extensive experiments on 18 public projects from four datasets indicate that: (a) our approach significantly outperforms state‐of‐the‐art cross‐project defect prediction (CPDP) approaches, and (b) our approach can obtain comparable prediction performances in contrast with within project prediction results. Consequently, the proposed approach can build an effective prediction model with a small number of labeled instances for mixed‐project defect prediction (MPDP). Lina Gong, Shujuan Jiang |
J. Softw. Evol. Process. | 1 |