Yongchang Ding

dblp:301/4992 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-4736-1845ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Metric information mining with metric attention to boost software defect prediction performance
abstract
In the field of software engineering, defect prediction has always been a popular research direction. Currently, the research on traditional software defect prediction mainly focuses on metric features, which are derived from various descriptive rules. Many researchers have proposed a large number of defect prediction models based on these metric features and various framework models. However, the problem of data scarcity has severely hindered the development of the field. Therefore, this work proposes a new method, namely the Metric Attention Module (MAM), which excavates the correlations within the metric data features, between features, within modules, and between modules. By learning new data representations, MAM guides the model's learning process and ultimately improves the model's performance without changing the network framework structure. Additionally, the method is interpretable. In this work, experiments were conducted in various task environments and on different datasets, all resulting in varying degrees of improvement. In the context of within-project defect prediction (WPDP), experiments with the MAM data model showed an average improvement of 14.7% in Accuracy, 15.9% in F1 score, 23.7% in AUC, and 65.1% in MCC. In cross-project defect prediction (CPDP), under more complex task environments, the model demonstrated excellent performance across multiple standard datasets. Compared to the baseline models and training results, the F1, Accuracy, and MCC scores improved by approximately 40%, 20%, and 50%, respectively.
Yongchang Ding, Zhiqiang Li 0003, Linjun Chen, Rong Peng, Xiaoyuan Jing
Sci. Comput. Program.1
2026 Explainability-guided augmentation for data-scarce CPDP via semantic anchors
Yongchang Ding, Haihao Zhou, Yunkun Cheng, Yao Lihaosheng, Rong Peng, Xiaoyuan Jing
Sci. Comput. Program.1
2025 Logic-infused knowledge graph QA: Enhancing large language models for specialized domains through Prolog integration
Aneesa Bashir, Rong Peng, Yongchang Ding
Data Knowl. Eng.3
2025 Sample-pair learning network for extremely imbalanced classification
abstract
In data classification, class-balanced data is ideal, but real datasets are often imbalanced, necessitating rebalancing through methods like resampling. In recent years, some new generative model-based resampling methods have been proposed. However, when facing extreme class imbalance, where the minority class is strongly underrepresented and on its own does not contain enough information to conduct the generative process. Some deep learning methods have been proposed to solve extremely imbalanced classification problems, but some of them are only used for specific datasets. Therefore, we proposed a novel deep learning method that combines a generative strategy with multi-task joint learning, termed sample-pair learning network (SPLN), for extremely imbalanced classification. The network consists of data preprocessing and multi-task joint learning modules. During data preprocessing, the training set is expanded by constructing positive and negative sample-pairs, then rebalanced using a strategy combining attention and resampling, termed undersampling based on attention power values (APVUS). The multi-task joint learning module employs a Siamese convolutional subnetwork to measure the similarity between sample-pairs and a multi-layer perceptron to recognize the category of single samples. The module can reduce the risk of overfitting caused by excessive noise in the training set. Finally, we designed a voting model based on the Siamese convolutional subnetwork to infer the categories of test samples. Experimental results demonstrate that our approach outperforms state-of-the-art generative model-based methods and is effective and general for extremely imbalanced classification.
Linjun Chen, Xiaoyuan Jing, Runhang Chen, Fei Wu 0004, Yongchang Ding, Changhui Hu 0001, Ziyun Cai
Neurocomputing5
2024 An interpretable logic KBQA method based on open-source large language models
abstract
Knowledge Base Question Answering (KBQA) aims to find correct answers to natural language questions by reasoning over large-scale knowledge bases.The main challenge is multi-hop reasoning, which requires inferring answers through multiple edges and nodes.Due to high data annotation costs, most datasets only provide final answer annotations, leaving the reasoning process unknown.The model cannot provide the reasoning path along with the answer, making the answer uninterpretable.To address this, we propose a generation-retrieval multi-hop KBQA method combining large language models (LLMs) and the logic programming language Prolog.In the generation phase, we fine-tune the open-source LLMs to generate the logical form of the corresponding prolog representation for the natural language question.In the retrieval phase, the model uses prolog query to reasoning over the KB to get the final answer and the reasoning path.A transparent reasoning path also helps identify and correct errors, enhancing the model's reliability and practicality.Experimental results on two standard KBQA datasets, MetaQA and WebQSP, demonstrate that our method outperforms existing models in both performance and interpretability.
Bicheng Xu, Rong Peng, Yongchang Ding
SEKE3
2023 A supervised data augmentation strategy based on random combinations of key features
abstract
Data augmentation strategies have always been important in machine learning techniques and play a unique role in model performance optimization processes. Therefore, in recent years, these techniques have become popular in the artificial intelligence field. In this paper, a new data augmentation strategy is proposed based on the interpretation algorithm of deep convolutional neural networks, i.e., constructing new training samples by deeply exploiting key features extracted from interpretable networks to achieve sample augmentation. Thus, a novel supervised data augmentation approach known as Supervised Data Augmentation–Key Feature Extraction (SDA-KFE) was proposed. By introducing the Neural Network Interpreter-Segmentation Recognition and Interpretation (NNI-SRI) algorithm, an augmentation strategy is proposed that can balance the high accuracy and high robustness of the final model while ensuring a large amount of data augmentation. The advantages of the SDA-KFE algorithm are mainly reflected in the following aspects. First, it is easy to implement. This algorithm is implemented based on the lightweight NNI-SRI algorithm, which lays the foundation for the implementation of SDA-KFE so that it can be easily implemented on convolutional neural networks. Second, this model, which is widely applicable, can be applied to almost any deep convolutional network. Through research and experiments on this proposed algorithm, SDA-KFE can be applied in graphical image binary classification and multiclassification models. Third, SDA-KFE can rapidly construct data samples with diverse variations. Under the premise of determining the classification labels of the generated samples, the distribution of the feature unit composition of the samples can be controlled. Compared with traditional data augmentation methods, SDA-KFE can control the direction of the model performance, i.e., the balance between the pursuit of high accuracy and robust performance of the model. Therefore, the novel supervised augmentation approach proposed in this paper is relevant for optimizing deep convolutional neural networks, solving model overfitting, augmenting data types, etc. The data augmentation algorithm proposed in this paper can be regarded as a useful supplement to traditional data augmentation methods, such as horizontal or vertical image flipping, cropping, color transformation, extension and rotation.
Yongchang Ding, Chang Liu 0123, Qianjun Chen
Inf. Sci.1
2022 Visualizing deep networks using segmentation recognition and interpretation algorithm
abstract
Owing to the advancements in artificial intelligence (AI) technology, deep learning has found applications in data analysis and problem-related decision-making processes. However, because of the “black-box” characteristics of deep neural networks, the explanations of decision-making processes are ambiguous. Therefore, model developers face challenges while establishing users’ trust in these processes and the associated decision results. Understanding and explaining the decision-making processes implemented by deep learning-based predictive models has importance in various fields. Based on existing interpretation algorithms, this study proposed a novel interpretation algorithm for deep convolutional neural network (CNN) models. The proposed algorithm, known as the neural network interpreter-segmentation recognition and interpretation (NNI-SRI) algorithm, uses the general ideas of segmentation and recognition to interpret models. Compared with other mainstream interpretation algorithms, this algorithm has the following advantages. First, it has a higher model interpretation speed. Second, it can label the positive and negative features observed during the model prediction process. Third, fewer parameters and hardware resources are required because only the forward propagation is used. Furthermore, the proposed algorithm showed better scene adaptability under the premise of ensuring the accuracy of interpretation, which is consistent with the CNN convolution process. Theoretically, the NNI-SRI algorithm can be applied to any model, particularly CNN models, such as InceptionV3, Xception, and ResNet. In this study, the interpretation algorithm was applied to the standard Inception deep network model and the spider sex recognition model developed by us, which achieved excellent results and verified the feasibility of the algorithm.
Yongchang Ding, Chang Liu 0123, Qianjun Chen
Inf. Sci.1