EDBT 2026 Demo / reviewers in the wild / expert
Lu Lu 0011
dblp:01/2086-11
· DBLP profile ↗
33ranked-venue papers
0as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 11 since 2021Software engineering, systems software and programming languages · 15 · 10 since 2021Systems, architecture and hardware · 8 · 8 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KDSA - VD : Optimizing Cross-Project Vulnerability Detection Through Dual-Driven Fusion of Knowledge Distillation and Structural AlignmentabstractABSTRACT Supervised deep learning techniques have demonstrated potential for software vulnerability detection, yet their practical deployment is often constrained by the limited availability of high‐quality labelled data. Cross‐project vulnerability detection seeks to overcome this constraint by enabling knowledge transfer from labelled source projects to target projects with insufficient labels. Nevertheless, existing studies fail to fully leverage the discriminative knowledge of source‐domain models and overlook cross‐project local semantic inconsistencies. To tackle these issues, we propose KDSA‐VD, a cross‐project vulnerability detection approach that employs a dual‐driven fusion of knowledge distillation and structural alignment. Built on CodeBERT for code representation learning, the proposed method first acquires discriminative features through supervised training on source projects. It then enhances cross‐domain adaptation by integrating soft‐label distillation, consistency regularization, and structural alignment to better exploit unlabelled target‐project data. Experimental results on 12 cross‐project vulnerability detection tasks indicate that KDSA‐VD outperforms existing baseline methods in most settings and exhibits stable performance under varying target‐domain labelling conditions. Chao Hong, Shuqin Gan, Linglun Luo, Shaojian Qiu, Lu Lu 0011 |
Expert Syst. J. Knowl. Eng. | 5 |
| 2026 | Integrating Retrieval Augmentation and Decoding Intervention for Automated Program RepairabstractABSTRACT Automated program repair (APR) aims to automatically detect and fix software defects, thereby improving software reliability and reducing debugging effort. Recently, researchers have explored the retrieval augmentation techniques to enhance large code models' performance in program repair. Existing retrieval augmentation models often inject retrieved information at the input layer, which can lead to input sequence inflation and interfere with the encoder's ability to focus on the core repair task. Meanwhile, learning‐based methods frequently produce unreliable patches, lacking mechanisms to verify or refine low‐confidence outputs during generation. To address these challenges, this paper proposes RADI‐PR, a novel approach that integrates retrieval augmentation and decoding intervention at the model's output layer. RADI‐PR dynamically incorporates relevant repair patterns based on historical fixes and intervenes in low‐confidence generations to enhance both the accuracy and reliability of generated patches. Comprehensive evaluations on four benchmark datasets, including Java, FixJS, Codeflaws and TSSB‐3 M, show that RADI‐PR consistently outperforms baseline methods. RADI‐PR achieves improvements of up to 5.5% in Precision, 3.2% in F1‐score and 2.2% in Accuracy. Shaosheng Wang, Lu Lu 0011, Shaojian Qiu, Siliang Suo |
Expert Syst. J. Knowl. Eng. | 2 |
| 2026 | AscQLUT: A decode-fused INT4 GEMM Kernel for Ascend NPUs
Lu Lu 0011 |
J. Syst. Archit. | 2 |
| 2026 | An architecture-adaptive optimization strategy for high-performance SYMV on a heterogeneous AI accelerator
Lu Lu 0011 |
J. Syst. Archit. | 2 |
| 2025 | Unlocking NPU Performance: A Comparative Analysis of Framework-Hardware Interaction on AscendabstractThe growing use of specialized AI accelerators like Neural Processing Units (NPUs) requires a deep understanding of framework-hardware interactions to achieve optimal performance and energy efficiency. This study investigates these interactions through a comprehensive comparative analysis of TensorFlow, a general-purpose framework, and MindSpore, a hardware-optimized framework, on the Huawei Ascend 910 NPU. We evaluate performance across representative deep learning workloads (ResNet50, YOLOv5, Transformer) at both model and operator levels, analyzing throughput, hardware resource utilization (AI Core, HBM, DDR), and average power consumption. Our findings reveal that MindSpore significantly outperforms TensorFlow in throughput and energy efficiency, primarily driven by its substantially higher utilization of the NPU's High-Bandwidth Memory (HBM). Operator-level analysis confirms that MindSpore's advantage grows with scale, correlating strongly with memory demands. While acknowledging MindSpore's targeted optimizations, this work provides quantitative insights into the critical role of memory hierarchy management for NPU performance and offers practical guidance for optimizing AI systems on specialized hardware accelerators, providing valuable insights for both AI practitioners and the HPC community. Lu Lu 0011, Siliang Suo |
HPCC | 3 |
| 2025 | LE-GEMM: A lightweight emulation-based GEMM with precision refinement on GPU
Lu Lu 0011, Zhanyu Yang, Siliang Suo |
J. Syst. Archit. | 2 |
| 2025 | A load-balanced acceleration method for small and irregular batch matrix multiplication on GPU
Lu Lu 0011, Zhanyu Yang, Siliang Suo |
J. Syst. Archit. | 2 |
| 2025 | DALO-APR: LLM-based automatic program repair with data augmentation and loss function optimization
Shaosheng Wang, Lu Lu 0011, Shaojian Qiu, Qingyan Tian, Haishan Lin |
J. Supercomput. | 2 |
| 2025 | An efficient quantized GEMV implementation for large language models inference with matrix core
Lu Lu 0011, Yijie Guo, Zhanyu Yang |
J. Supercomput. | 2 |
| 2024 | Vulnerability Detection Based on Pre-trained Code Language Model and Convolutional Neural NetworkabstractSoftware vulnerabilities damage the reliability of software systems.Recently, many methods based on deep learning have been proposed for vulnerability detection by learning features from code sequences or various property graphs.However, sequence-based methods frequently fail to consider the structural information of the code, whereas graph-based methods often encounter difficulties in capturing long-distance contextual semantic information.To overcome these limitations, we propose PTCN, a novel vulnerability detection method that integrates Pre-Trained code language model and Convolutional Neural Network (CNN).First, the code tokens sequence and the propagation chains are constructed from the source code in preparation for the subsequent feature extraction.Next, GraphCodeBERT, a graph-based pre-trained code language model, is introduced as a means of learning both the semantic and structural information contained in the code tokens sequence and the propagation chains.Then, a CNN architecture is designed to effectively extract the crucial vulnerability features from the hidden states of the last Transformer layer of GraphCodeBERT for vulnerability detection.The experiments on the public benchmark dataset REVEAL indicate that PTCN outperforms the state-of-the-art methods with respect to the Accuracy and F1-score metrics by 3.41%-15.80%and 22.68%-104.27%,respectively. Tingfeng Liao, Lu Lu 0011, Siliang Suo |
SEKE | 2 |
| 2024 | Vulnerability Detection Based on Adapter Tuning and Enhanced Feature LearningabstractPre-trained code models have achieved promising results in the vulnerability detection field.The prevailing approach is to adapt these models with vulnerability datasets using inefficient full-model fine-tuning.These pre-trained models primarily capture the semantic features of code, while neglecting its structural characteristics.To address these issues, this paper proposes a vulnerability detection method with adapter tuning and enhanced feature learning.First, adapter modules are introduced to UniXcoder and tune only the parameters in the adapters to extract semantic features.This significantly reduces the number of training parameters and better adapts the model to downstream tasks.Then, structural features, including control and data flow information, are extracted from the Program Dependence Graph (PDG) to compensate for the limitations of pre-trained models that rely solely on semantic features.Finally, the semantic and structural features are integrated to train the detection model.The experimental results demonstrate that our proposed method outperforms state-of-the-art approaches and is highly efficient in terms of both training parameters and training data. Lu Lu 0011, Siliang Suo |
SEKE | 2 |
| 2024 | Multi-Scale Feature Extraction with Supervised Contrastive Learning for Vulnerability DetectionabstractThe capacity of Deep Learning to automatically learn features from source code has facilitated its extensive utilization in detecting software vulnerabilities.However, existing pre-trained models regard code snippets as token sequences, neglecting the inherent structure of the code.The utilization of graph-based code representations is constrained by the limitations of Graph Neural Networks, particularly the difficulty in capturing long-range dependencies, which renders it challenging to learn grammatical and semantic information from complex code graph representations.This paper proposes a novel multiscale software vulnerability detection method based on supervised contrastive learning.The method integrates local features obtained from code paths with structural features extracted from the Code Property Graph by GraphTrans, thereby achieving a multi-scale feature representation of the source code.Additionally, a supervised contrastive loss function is employed during training in order to fully utilize label information and address the class imbalance problem.The experimental results demonstrate that the proposed method outperforms existing state-of-the-art methods, achieving the highest accuracy, precision, and F1 on the real-world benchmark dataset CodeXGLUE for software vulnerability detection. Yahui Zhao, Lu Lu 0011, Siliang Suo |
SEKE | 2 |
| 2024 | IC-GraF: An Improved Clustering with Graph-Embedding-Based Features for Software Defect PredictionabstractSoftware defect prediction (SDP) has been a prominent area of research in software engineering. Previous SDP methods often struggled in industrial applications, primarily due to the need for sufficient historical data. Thus, clustering‐based unsupervised defect prediction (CUDP) and cross‐project defect prediction (CPDP) emerged to address this challenge. However, the former exhibited limitations in capturing semantic and structural features, while the latter encountered constraints due to differences in data distribution across projects. Therefore, we introduce a novel framework called improved clustering with graph‐embedding‐based features (IC‐GraF) for SDP without the reliance on historical data. First, a preprocessing operation is performed to extract program dependence graphs (PDGs) and mark distinct dependency relationships within them. Second, the improved deep graph infomax (IDGI) model, an extension of the DGI model specifically for SDP, is designed to generate graph‐level representations of PDGs. Finally, a heuristic‐based k‐means clustering algorithm is employed to classify the features generated by IDGI. To validate the efficacy of IC‐GraF, we conduct experiments based on 24 releases of the PROMISE dataset, using F‐measure and G‐measure as evaluation criteria. The findings indicate that IC‐GraF achieves 5.0%−42.7% higher F‐measure, 5%−39.4% higher G‐measure, and 2.5%−11.4% higher AUC over existing CUDP methods. Even when compared with eight supervised learning‐based SDP methods, IC‐GraF maintains a superior competitive edge. Xuanye Wang, Lu Lu 0011, Qingyan Tian, Haishan Lin |
IET Softw. | 2 |
| 2024 | Ensemble Kernel-Mapping-Based Ranking Support Vector Machine for Software Defect PredictionabstractRank-oriented software defect prediction (ROSDP) aims to establish a model to predict the testing priority of software modules according to defect severity for the reasonable allocation of limited testing resources. Some ROSDP methods construct the prediction model by a linear model with respect to software features. However, in software repositories, the linear condition between testing priority and software feature is not satisfied, and the ranking performance of the linear prediction model is limited. Thus, in order to relax the limitation of the linear prediction model and improve the ranking performance, ensemble kernel-mapping-based ranking support vector machine (EKMRSVM) is developed based on the theories of ranking SVM, which builds a nonlinear ranking function approximated by the kernel-mapping-based method. Furthermore, the sequential minimal optimization algorithm is developed to derive the ideal parameters of the nonlinear ranking function, and ensemble learning is introduced to reduce time costs and guarantee ranking performance. Experimental results on 20 open source datasets indicate that introducing the kernel mapping method in EKMRSVM is very effective in performance improvement, and ensemble learning makes the proposed ranking algorithm very competitive in terms of time costs. Thus, based on the comparative results of some baseline methods, EKMRSVM with the appropriate kernel function can achieve better ranking performance. Zhanyu Yang, Lu Lu 0011, Quanyi Zou |
IEEE Trans. Reliab. | 2 |
| 2023 | A Software Defect Prediction Method based on Multi-type Features and Feature SelectionabstractNumerous software defect prediction methods utilize semantic information and software metrics as code features, neglecting the structural knowledge inherent in the source code.Other studies improve feature completeness by simply combining different types of defect indicators, which causes information redundancy.To address these challenges, this paper proposes a novel software defect prediction method that incorporates multitype features and performs feature selection.Firstly, semantic and structural features are extracted by Text Convolutional Neural Network (TextCNN) and Graph Isomorphism Network (GIN) from Abstract Syntax Tree (AST) and Program Dependency Graph (PDG), respectively, which are combined with software metrics to build a multi-type feature set.Then, Recursive Feature Elimination with Cross-Validation (RFECV) integrating a novel feature importance measure is utilized to remove redundant features and generate a feature subset.Finally, a prediction model for classification is established based on the feature subset.The experiments validated the effectiveness of multi-type features and the improved RFECV.Overall our proposed method outperforms state-of-the-art techniques on nine Java open-source projects. Lu Lu 0011, Quanyi Zou, Zhanyu Yang |
SEKE | 3 |
| 2023 | Software Defect Prediction via Positional Hierarchical Attention Network (S)abstractSoftware Defect Prediction (SDP) aims to identify defect-prone modules in advance to ensure software quality.In SDP research based on deep learning, the mainstream approach is to extract deep semantic features from an Abstract Syntax Tree (AST).Theoretically, the AST as a bi-dimensional structure encloses information at the node level, fragment level, and entire tree level.However, most existing research serializes the whole AST without considering the expression at different granularities.To address this limitation, we introduce a positional hierarchical attention network (PHAN) that acquires semantic features by simultaneously considering contexts between nodes and paths.Specifically, our model incorporates attention mechanisms to capture information of varying importance at separate hierarchies, and relative position representations to distinguish the contributions of different paths.Experimental results demonstrate that PHAN significantly outperforms existing baseline methods. Xinyan Yi, Lu Lu 0011, Quanyi Zou, Zhanyu Yang |
SEKE | 3 |
| 2023 | Characterization of heart rate variability and oxygen saturation in sepsis patientsabstractAbstract One of the most major and common health crises which occur across all the hospitals, worldwide, is seen to be sepsis that occurs in patients. However, despite its wide prevalence no novel tool has been devised for predicting its occurrence. An accurate and early prediction of sepsis in the patients could significantly help the physicians administer proper treatment and decrease the uncertain diagnosis. Some machine‐learning‐based models or schemes can help in identifying the potential clinical variables and display a better performance compared to the prevailing conventional low‐performance models. In this study, a machine learning‐based scheme for fast and accurate sepsis identification was proposed. This scheme employed the power spectrum and mean estimation for data record intervals, which were then classified for reaching the final decision. For a 72‐h interval, the obtained detection accuracy was 94.2% that shows very good sign to use it as a fast and robust sepsis identification. Bilal Yaseen Al-Mualemi, Lu Lu 0011 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2022 | ASD-ISSPA: Adaptive Stochastic Droppath and Interactive Slow Semi-Polarized Attention based on FPN for Object DetectionabstractFeature Pyramid Network (FPN) has been widely used to combine features at different scales for extracting semantically strong representations in object detection. However, most of the previous approaches only focus on the original image resolution to allocate pyramidal layers without considering the image size change in the processing network. In this work, an algorithm is proposed to determine the feature layer based on the ratio of the target box area to the original image area, which improves the network adaptability to the scale change. Further-more, a novel Adaptive Stochastic Droppath and Interactive Slow Semi-Polarized Attention (ASD-ISSPA) is proposed to retain the valuable information in the feature maps of FPN. ASD-ISSPA can be split into two parts, Adaptive Stochastic Droppath Attention (ASDA) and Interactive Slow Semi-Polarized Attention (ISSPA). The ASDA module acts on the bottom-up pathway in the pyramid to collect the loss information after convolution, while the ISSPA module deals with the top-down path in the pyramid to capture the semantics before upsampling and the interaction details after upsampling. The integrated two-path attention mechanism effectively regains the loss information of the convolution and removes the redundancy of the deconvolution so that our model can improve the category prediction accuracy and reduce the discrepancy between the predicted and ground-truth boxes. Compared with existing methods, the proposed attention extract network, called ASD-ISSPA, achieves competitive results on the PASCAL VOC dataset. Yukang Yang, Lu Lu 0011 |
IJCNN | 3 |
| 2022 | Two-Stage AST Encoding for Software Defect PredictionabstractSoftware defect prediction (SDP) can find potential containing defect modules, which assists software developers in allocating limited test resources more efficiently.Because traditional software features fail to capture the semantics of source code, various studies have turned to extracting deep learning features.Existing related approaches often parse the program source code into Abstract Syntax Trees (ASTs) for further processing.However, most of these approaches ignore AST nodes' hierarchical and position-sensitive structure.To overcome the aforementioned issues, a two-stage AST encoding (TSE) method is proposed in this paper for software defect prediction.Experiments on eight Java open-source projects showed that our proposed SDP method outperforms several traditional methods and state-of-the-art deep learning methods in terms of F-measure and MCC. Yanwu Zhou, Lu Lu 0011, Quanyi Zou, Cuixu Li |
SEKE | 2 |
| 2022 | A high-performance batched matrix multiplication framework for GPUs under unbalanced input distribution
Lu Lu 0011 |
J. Supercomput. | 4 |
| 2021 | CcGL-GAN: Criss-Cross Attention and Global-Local Discriminator Generative Adversarial Networks for text-to-image synthesisabstractText-to-image synthesis aims to generate a visually realistic image according to a linguistic text description. Visual quality and semantic consistency are two key objectives. Although remarkable progress has been made in improving visual resolutions leveraging Generative Adversarial Networks (GANs), guaranteeing the semantic conformity remains challenging. In this paper, we address it by proposing a novel Criss-Cross Attention and Global-Local Discriminator Generative Adversarial Networks(CcGL-GAN). CcGL-GAN exploits a Criss-Cross Attention mechanism to capture the variation of contextual description, which enables back generators to generate images more efficiently. Moreover, it utilizes Global-Local discriminators to project low-resolution images onto global linguistic representations, and high-resolution images onto local linguistic representations, which ensures that our model narrows the gap between images and descriptions. Experiments conducted on two publicly available datasets, the CUB and Oxford-102, demonstrate the effectiveness of the proposed CcGL-GAN model. Xihong Ye, Lu Lu 0011 |
IJCNN | 2 |
| 2021 | Multi-source Cross Project Defect Prediction with Joint Wasserstein Distance and Ensemble LearningabstractCross-Project Defect Prediction (CPDP) refers to transferring knowledge from source software projects to a target software project. Previous research has shown that the impacts of knowledge transferred from different source projects differ on the target task. Therefore, one of the fundamental challenges in CPDP is how to measure the amount of knowledge transferred from each source project to the target task. This article proposed a novel CPDP method called Multi-source defect prediction with Joint Wasserstein Distance and Ensemble Learning (MJWDEL) to learn transferred weights for evaluating the importance of each source project to the target task. In particular, first of all, applying the TCA technique and Logistic Regression (LR) train a sub-model for each source project and the target project. Moreover, the article designs joint Wassertein distance to understand the source-target relationship and then uses this as a basis to compute the transferred weights of different sub-models. After that, the transferred weights can be used to reweight these sub-models to determine their importance in knowledge transfer to the target task. We conducted experiments on 19 software projects from PROMISE, NASA and AEEEM datasets. Compared with several state-of-the-art CPDP methods, the proposed method substantially improves CPDP performance in terms of four evaluation indicators (i.e., F-measure, Balance, G-measure and MMC). Quanyi Zou, Lu Lu 0011, Zhanyu Yang |
ISSRE | 2 |
| 2021 | Correlation feature and instance weights transfer learning for cross project software defect predictionabstractAbstract Due to the differentiation between training and testing data in the feature space, cross‐project defect prediction (CPDP) remains unaddressed within the field of traditional machine learning. Recently, transfer learning has become a research hot‐spot for building classifiers in the target domain using the data from the related source domains. To implement better CPDP models, recent studies focus on either feature transferring or instance transferring to weaken the impact of irrelevant cross‐project data. Instead, this work proposes a dual weighting mechanism to aid the learning process, considering both feature transferring and instance transferring. In our method, a local data gravitation between source and target domains determines instance weight, while features that are highly correlated with the learning task, uncorrelated with other features and minimizing the difference between the domains are rewarded with a higher feature weight. Experiments on 25 real‐world datasets indicate that the proposed approach outperforms the existing CPDP methods in most cases. By assigning weights based on the different contribution of features and instances to the predictor, the proposed approach is able to build a better CPDP model and demonstrates substantial improvements over the state‐of‐the‐art CPDP models. Quanyi Zou, Lu Lu 0011, Shaojian Qiu, Xiaowei Gu 0002 |
IET Softw. | 2 |
| 2021 | Joint feature representation learning and progressive distribution matching for cross-project defect prediction
Quanyi Zou, Lu Lu 0011, Zhanyu Yang, Xiaowei Gu 0002, Shaojian Qiu |
Inf. Softw. Technol. | 2 |
| 2021 | Verifiable, Reliable, and Privacy-Preserving Data Aggregation in Fog-Assisted Mobile CrowdsensingabstractFog-assisted mobile crowdsensing (FA-MCS) alleviates challenges with respect to computation, communication, and storage from the traditional model of mobile crowdsensing (MCS) “requester-server-users.” Data aggregation, as a specific MCS task, has attracted a lot of attentions in mining the potential value of the massive crowdsensing data. However, the process of data aggregation in FA-MCS may threaten the privacies of both users' data and aggregation results. The untrusted server and fog nodes (FNs) may damage the correctness of aggregation results. Moreover, bad FNs, which do not upload data to server or fail to verify successfully, can endanger the reliability of FA-MCS and the accuracy of aggregation results. To tackle these problems, we propose a verifiable, reliable, and privacy-preserving data aggregation scheme for FA-MCS. Specifically, the proposed scheme preserves privacies of both users' data and aggregation results, enables requester to verify the correctness of aggregation result, and is able to tolerate several bad FNs without affecting the data aggregation result. Through formal security analysis, the proposed scheme is shown to be secure and privacy preserving. Extensive experiments also show the proposed scheme is efficient and reliable. Xingfu Yan, Wing W. Y. Ng, Changlu Lin, Yuxian Liu, Lu Lu 0011, Ying Gao 0004 |
IEEE Internet Things J. | 6 |
| 2020 | Defect Prediction via LSTM Based on Sequence and Tree StructureabstractWith the ever-expanding spread of contemporary software, software defect prediction (SDP) is attracting more and more attention. However, sequential networks used in previous studies, weaken syntactic information and fail to capture longdistance dependencies. To solve these problems, we develop a long short-term memory network based on bidirectional and tree structure (LSTM-BT). Specifically, LSTM-BT combines bidirectional long short-term memory networks (BI-LSTM) and tree long short-term memory networks (Tree-LSTM) to capture semantic and syntactic features from source codes. First, token vectors are captured from the abstract syntax tree (AST). Second, an embedding layer is used to extract semantic information hidden inside the AST nodes. Last, features are fed to the LSTM- BT, which is used to conduct predictions of defect-proneness. To validate our method, we carried out experiments on 8 pairs of Java open-source projects and the results show that LSTM- BT performs better compared to several state-of-the-art defect prediction models. Lu Lu 0011 |
QRS | 2 |
| 2020 | Software defect prediction via LSTMabstractSoftware quality plays an important role in the software lifecycle. Traditional software defect prediction approaches mainly focused on using hand‐crafted features to detect defects. However, like human languages, programming languages contain rich semantic and structural information, and the cause of defective code is closely related to its context. Failing to catch this significant information, the performance of traditional approaches is far from satisfactory. In this study, the authors leveraged a long short‐term memory (LSTM) network to automatically learn the semantic and contextual features from the source code. Specifically, they first extract the program's Abstract Syntax Trees (ASTs), which is made up of AST nodes, and then evaluate what and how much information they can preserve for several node types. They traverse the AST of each file and fed them into the LSTM network to automatically the semantic and contextual features of the program, which is then used to determine whether the file is defective. Experimental results on several opensource projects showed that the proposed LSTM method is superior to the state‐of‐the‐art methods. Jiehan Deng, Lu Lu 0011, Shaojian Qiu |
IET Softw. | 2 |
| 2020 | Sentiment key frame extraction in user-generated micro-videos via low-rank and sparse representation
Xiaowei Gu 0002, Lu Lu 0011, Shaojian Qiu, Quanyi Zou, Zhanyu Yang |
Neurocomputing | 2 |
| 2019 | Cross-Project Defect Prediction via Transferable Deep Learning-Generated and Handcrafted FeaturesabstractAlthough the machine learning-based software defect prediction (SDP) method has shown promising value in software engineering, yet challenges remain.To improve the performance of SDP, some researchers have used deep learning algorithms to extract the semantic and structural features of the program.However, in more practical cross-project defect prediction (CPDP) tasks, whether deep learning-generated features can be directly used should be explored due to the data distribution shift that usually exists in different projects.In this paper, we propose a Transferable Hybrid Features Learning with Convolutional Neural Network (CNN-THFL) framework to conduct CPDP.Specially, CNN-THFL mines deep learning-generated features from token vectors extracted from programs' abstract syntax trees via convolutional neural network.Furthermore, CNN-THFL learns the transferable joint features simultaneously considering deep learning-generated and handcrafted features by applying a transfer component analysis algorithm.Finally, the features generated by CNN-THFL are fed to the classifier to train a defect prediction model.Extensive experiments verify that CNN-THFL can outperform referential methods on 72 pairs of CPDP tasks formed by 9 open-source projects. Shaojian Qiu, Lu Lu 0011, Siyu Jiang |
SEKE | 2 |
| 2019 | Joint distribution matching model for distribution-adaptation-based cross-project defect predictionabstractUsing classification methods to predict software defect is receiving a great deal of attention and most of the existing studies primarily conduct prediction under the within‐project setting. However, there usually had no or very limited labelled data to train an effective prediction model at an early phase of the software lifecycle. Thus, cross‐project defect prediction (CPDP) is proposed as an alternative solution, which is learning a defect predictor for a target project by using labelled data from a source project. Differing from previous CPDP methods that mainly apply instances selection and classifiers adjustment to improve the performance, in this study, the authors put forward a novel distribution–adaptation‐based CPDP approach, joint distribution matching (JDM). Specifically, JDM aims to minimise the joint distribution divergence between the source and target project to improve the CPDP performance. By constructing an adaptive weight vector for the instances of the source project, JDM can be effective and robust at reducing marginal distribution discrepancy and conditional distribution discrepancy simultaneously. Extensive experiments verify that JDM can outperform related distribution–adaptation‐based methods on 15 open‐source projects that are derived from two types of repositories. Shaojian Qiu, Lu Lu 0011, Siyu Jiang |
IET Softw. | 2 |
| 2019 | Application of a Traffic Flow Prediction Model Based on Neural Network in Intelligent Vehicle ManagementabstractThe ultimate direction of intelligent vehicle management is to achieve artificial intelligence (AI), and data mining is an important supporting technology for AI. The adoption of new AI technology can effectively improve operational efficiency and safety, especially in terms of performance. This paper takes the researches on traffic jam as an example and proposes one algorithm for combination forecasting model based on a segmentation algorithm for traffic flow sequence and BP neural network prediction. In this paper, it also introduces the traffic flow clustering analysis and mining algorithms for congestion events at the intersections. The blocking point algorithm is improved, and experimental analysis is performed through samples. Experimental results show that the algorithm use for combination forecasting model can greatly improve the real-time performance of short-term traffic flow prediction and significantly reduce the prediction error rate. Therefore, this algorithm has practical and innovative significance in the field of intelligent vehicle management. Yang Guo 0006, Lu Lu 0011 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2019 | An Investigation of Imbalanced Ensemble Learning Methods for Cross-Project Defect PredictionabstractMachine-learning-based software defect prediction (SDP) methods are receiving great attention from the researchers of intelligent software engineering. Most existing SDP methods are performed under a within-project setting. However, there usually is little to no within-project training data to learn an available supervised prediction model for a new SDP task. Therefore, cross-project defect prediction (CPDP), which uses labeled data of source projects to learn a defect predictor for a target project, was proposed as a practical SDP solution. In real CPDP tasks, the class imbalance problem is ubiquitous and has a great impact on performance of the CPDP models. Unlike previous studies that focus on subsampling and individual methods, this study investigated 15 imbalanced learning methods for CPDP tasks, especially for assessing the effectiveness of imbalanced ensemble learning (IEL) methods. We evaluated the 15 methods by extensive experiments on 31 open-source projects derived from five datasets. Through analyzing a total of 37504 results, we found that in most cases, the IEL method that combined under-sampling and bagging approaches will be more effective than the other investigated methods. Shaojian Qiu, Lu Lu 0011, Siyu Jiang, Yang Guo 0006 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2018 | Multiple-components weights model for cross-project software defect predictionabstractSoftware defect prediction (SDP) technology is receiving widely attention and most of SDP models are trained on data from the same project. However, at an early phase of the software lifecycle, there are little to no within‐project training data to learn an available supervised defect‐prediction model. Thus, cross‐project defect prediction (CPDP), which is learning a defect predictor for a target project by using labelled data from a source project, has shown promising value in SDP. To better perform the CPDP, most current studies focus on filtering instances or selecting features to weaken the impact of irrelevant cross‐project data. Instead, the authors propose a novel multiple‐components weights (MCWs) learning model to analyse the varying auxiliary power of multiple components in a source project to construct a more precise ensemble classifiers for a target project. By combining the MCW model with kernel mean matching algorithm, their proposed approach adjusts the source‐instance weights and source‐component weights to jointly alleviate the negative impacts of irrelevant cross‐project data. They conducted comprehensive experiments by employing 15 real‐world datasets to demonstrate the advantages and effectiveness of their proposed approach. Shaojian Qiu, Lu Lu 0011, Siyu Jiang |
IET Softw. | 2 |