VLDB 2026 Research / reviewers in the wild / expert
Linjun Chen
dblp:259/0009
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Metric information mining with metric attention to boost software defect prediction performanceabstractIn the field of software engineering, defect prediction has always been a popular research direction. Currently, the research on traditional software defect prediction mainly focuses on metric features, which are derived from various descriptive rules. Many researchers have proposed a large number of defect prediction models based on these metric features and various framework models. However, the problem of data scarcity has severely hindered the development of the field. Therefore, this work proposes a new method, namely the Metric Attention Module (MAM), which excavates the correlations within the metric data features, between features, within modules, and between modules. By learning new data representations, MAM guides the model's learning process and ultimately improves the model's performance without changing the network framework structure. Additionally, the method is interpretable. In this work, experiments were conducted in various task environments and on different datasets, all resulting in varying degrees of improvement. In the context of within-project defect prediction (WPDP), experiments with the MAM data model showed an average improvement of 14.7% in Accuracy, 15.9% in F1 score, 23.7% in AUC, and 65.1% in MCC. In cross-project defect prediction (CPDP), under more complex task environments, the model demonstrated excellent performance across multiple standard datasets. Compared to the baseline models and training results, the F1, Accuracy, and MCC scores improved by approximately 40%, 20%, and 50%, respectively. Yongchang Ding, Zhiqiang Li 0003, Linjun Chen, Rong Peng, Xiaoyuan Jing |
Sci. Comput. Program. | 5 |
| 2025 | Sample-pair learning network for extremely imbalanced classificationabstractIn data classification, class-balanced data is ideal, but real datasets are often imbalanced, necessitating rebalancing through methods like resampling. In recent years, some new generative model-based resampling methods have been proposed. However, when facing extreme class imbalance, where the minority class is strongly underrepresented and on its own does not contain enough information to conduct the generative process. Some deep learning methods have been proposed to solve extremely imbalanced classification problems, but some of them are only used for specific datasets. Therefore, we proposed a novel deep learning method that combines a generative strategy with multi-task joint learning, termed sample-pair learning network (SPLN), for extremely imbalanced classification. The network consists of data preprocessing and multi-task joint learning modules. During data preprocessing, the training set is expanded by constructing positive and negative sample-pairs, then rebalanced using a strategy combining attention and resampling, termed undersampling based on attention power values (APVUS). The multi-task joint learning module employs a Siamese convolutional subnetwork to measure the similarity between sample-pairs and a multi-layer perceptron to recognize the category of single samples. The module can reduce the risk of overfitting caused by excessive noise in the training set. Finally, we designed a voting model based on the Siamese convolutional subnetwork to infer the categories of test samples. Experimental results demonstrate that our approach outperforms state-of-the-art generative model-based methods and is effective and general for extremely imbalanced classification. Linjun Chen, Xiaoyuan Jing, Runhang Chen, Fei Wu 0004, Yongchang Ding, Changhui Hu 0001, Ziyun Cai |
Neurocomputing | 1 |
| 2024 | Prompt enhance API recommendation: visualize the user's real intention behind this query
Linjun Chen, Yingtao Fang |
Autom. Softw. Eng. | 2 |
| 2024 | API Recommendation for Novice Programmers: Build a Bridge of Query-Task Knowledge GapabstractDuring software development, programmers often rely on a wide range of application programming interfaces (APIs) to facilitate their tasks. However, APIs have been growing rapidly in recent years, making it difficult for developers to choose among the many APIs that suit their programming needs. To facilitate the development process, automatic API recommendation is becoming increasingly important. Although there have been many effective research methods, these methods have a high dependence on the accuracy of the user's description of his own task, and there is a knowledge difference between the user's query and the user's actual task, increasing the difficulty of accurate API recommendation. In this article, we propose REAPI, a method to bridge the knowledge gap between the user's query and the user's actual task to improve the recommendation accuracy. The REAPI approach involves reconstructing query by tapping into Stack Overflow data to glean user intentions. Refactoring the user's query to display implicit information can better capture the user's true intentions. Specifically, we generate three candidate reconstruction statements based on natural language queries and Stack Overflow data and incorporate user feedback to refine and select the final statement. To evaluate the effectiveness of REAPI, we conducted experiments at both the class-level and method-level. Our results show that REAPI outperforms state-of-the-art baselines across key evaluation metrics such as S@1, S@3, S@10, MRR, and MAP. Yong Wang 0008, Yingtao Fang, Cuiyun Gao 0001, Linjun Chen |
IEEE Trans. Reliab. | 4 |
| 2023 | A novel two-way rebalancing strategy for identifying carbonylation sitesabstractBACKGROUND: As an irreversible post-translational modification, protein carbonylation is closely related to many diseases and aging. Protein carbonylation prediction for related patients is significant, which can help clinicians make appropriate therapeutic schemes. Because carbonylation sites can be used to indicate change or loss of protein function, integrating these protein carbonylation site data has been a promising method in prediction. Based on these protein carbonylation site data, some protein carbonylation prediction methods have been proposed. However, most data is highly class imbalanced, and the number of un-carbonylation sites greatly exceeds that of carbonylation sites. Unfortunately, existing methods have not addressed this issue adequately. RESULTS: In this work, we propose a novel two-way rebalancing strategy based on the attention technique and generative adversarial network (Carsite_AGan) for identifying protein carbonylation sites. Specifically, Carsite_AGan proposes a novel undersampling method based on attention technology that allows sites with high importance value to be selected from un-carbonylation sites. The attention technique can obtain the value of each sample's importance. In the meanwhile, Carsite_AGan designs a generative adversarial network-based oversampling method to generate high-feasibility carbonylation sites. The generative adversarial network can generate high-feasibility samples through its generator and discriminator. Finally, we use a classifier like a nonlinear support vector machine to identify protein carbonylation sites. CONCLUSIONS: Experimental results demonstrate that our approach significantly outperforms other resampling methods. Using our approach to resampling carbonylation data can significantly improve the effect of identifying protein carbonylation sites. Linjun Chen, Xiaoyuan Jing, Yaru Hao, Wei Liu 0200, Xiaoke Zhu |
BMC Bioinform. | 1 |
| 2022 | Collaborative filtering recommendation using fusing criteria against shilling attacksabstractThe collaborative filtering recommendation technique (CFR) is one of the techniques used in recommended systems, in which the most proximal neighbours to a target user are selected. Their profiles are used to predict rating for items as yet unrated by that target user. However, malicious users inject fake user profiles to destroy the security and reliability of the recommender systems, which is called shilling attacks. Therefore, it is crucial to improve the recommendation technique against shilling attacks. Malicious users use a single method to perform shilling attacks. Intuitively, fusing multiple criteria to construct CFR can effectively resist shilling attacks. A novel CFR is proposed against shilling attacks (called CFR-F). In our approach, a similar interest users’ resource set is obtained first by integrating users’ dynamic interest model and social tags. Then, a similar interest user resource set is selected according to a strategy that selects preference influence weight based on user background. Our experimental results show that our approach can recommend accurate information resources and has a lower Mean Absolute Error (MAE) and Average Prediction Shift (APS) than traditional techniques by 50% and 20%, respectively. Zhongqun Wang, Linjun Chen, Yong Wang 0008 |
Connect. Sci. | 4 |
| 2022 | Nonlinear Graph Learning-Convolutional Networks for Node Classification
Linjun Chen |
Neural Process. Lett. | 1 |
| 2022 | One-step spectral rotation clustering with balanced constrains
Guoqiu Wen, Yonghua Zhu, Linjun Chen, Shichao Zhang 0001 |
World Wide Web | 3 |
| 2021 | Global and Local Structure Preservation for Nonlinear High-dimensional Spectral ClusteringabstractAbstract Spectral clustering is widely applied in real applications, as it utilizes a graph matrix to consider the similarity relationship of subjects. The quality of graph structure is usually important to the robustness of the clustering task. However, existing spectral clustering methods consider either the local structure or the global structure, which can not provide comprehensive information for clustering tasks. Moreover, previous clustering methods only consider the simple similarity relationship, which may not output the optimal clustering performance. To solve these problems, we propose a novel clustering method considering both the local structure and the global structure for conducting nonlinear clustering. Specifically, our proposed method simultaneously considers (i) preserving the local structure and the global structure of subjects to provide comprehensive information for clustering tasks, (ii) exploring the nonlinear similarity relationship to capture the complex and inherent correlation of subjects and (iii) embedding dimensionality reduction techniques and a low-rank constraint in the framework of adaptive graph learning to reduce clustering biases. These constraints are considered in a unified optimization framework to result in one-step clustering. Experimental results on real data sets demonstrate that our method achieved competitive clustering performance in comparison with state-of-the-art clustering methods. Guoqiu Wen, Yonghua Zhu, Linjun Chen, Mengmeng Zhan, Yangcai Xie |
Comput. J. | 3 |
| 2021 | One-step spectral rotation clustering for imbalanced high-dimensional data
Guoqiu Wen, Xianxian Li, Yonghua Zhu, Linjun Chen, Qimin Luo, Malong Tan |
Inf. Process. Manag. | 4 |
| 2020 | Local Structure Preservation for Nonlinear Clustering
Linjun Chen, Guangquan Lu, Yangding Li, Jiaye Li 0001, Malong Tan |
Neural Process. Lett. | 1 |