VLDB 2026 Research / reviewers in the wild / expert
Lei Chen 0007
dblp:c/LeiChen0007
· DBLP profile ↗
20ranked-venue papers
9as first author
7since 2021 · last 2026
0000-0003-3068-1583ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Theory of computation · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PLysPTM-HGNN: predicting lysine PTM sites of proteins using hybrid graph neural networksabstractBACKGROUND: Protein post-translational modification (PTM) is a critical biological process that occurs after protein synthesis and has key roles in several biological processes. Among these, lysine modifications include multiple types and have received considerable attention. Most existing computational models predict whether a specific lysine site in a protein sequence corresponds to a lysine modification type by extracting features from a short peptide segment centered on that site. Therefore, information from the full protein sequence is not used. RESULTS: In this study, we gave a different direction for investigating lysine modifications. A computational model, PLysPTM-HGNN, was designed to identify lysine modification types at the protein level. Full protein sequence information was used to derive three feature types: gene ontology features, large language model features, and position-specific scoring matrix features. These features were refined separately through a linear transformation, a hybrid graph neural network, and a convolutional neural network combiner, after which they were concatenated and passed into a fully connected layer for prediction. Cross-validation results showed that the AUROC and AUPR were approximately 0.84 and 0.68, respectively, indicating strong predictive performance. CONCLUSIONS: PLysPTM-HGNN outperformed several existing protein subcellular localization models and models based on traditional multi-label classification algorithms. This model provides a useful tool for studies of lysine modifications. Lei Chen 0007, Jing Yang 0009, Yu-Dong Cai 0001 |
BMC Bioinform. | 1 |
| 2025 | STARFormer: A novel spatio-temporal aggregation reorganization transformer of FMRI for brain disorder diagnosis
Yueyang Li 0004, Weiming Zeng, Lei Chen 0007, Hongjie Yan, Wai Ting Siok, Nizhuan Wang 0001 |
Neural Networks | 4 |
| 2024 | Data-Driven Preference Sampling for Pareto Front LearningabstractPareto front learning is a technique that introduces preference vectors in a neural network to approximate the Pareto front. Previous Pareto front learning methods have demonstrated high performance in approximating simple Pareto fronts. These methods often sample preference vectors from a fixed Dirichlet distribution. However, no fixed sampling distribution can be adapted to diverse Pareto fronts. Efficiently sampling preference vectors and accurately estimating the Pareto front is a challenge. To address this challenge, we propose a data-driven preference vector sampling framework for Pareto front learning. We utilize the posterior information of the objective functions to adjust the parameters of the sampling distribution flexibly. In this manner, the proposed method can sample preference vectors from the location of the Pareto front with a high probability. Moreover, we design the distribution of the preference vector as a mixture of Dirichlet distributions to improve the performance of the model in disconnected Pareto fronts. Extensive experiments validate the superiority of the proposed method compared with state-of-the-art algorithms. Rongguang Ye, Lei Chen 0007, Weiduo Liao, Hisao Ishibuchi |
IJCNN | 2 |
| 2024 | RMTLysPTM: recognizing multiple types of lysine PTM sites by deep analysis on sequencesabstractPost-translational modification (PTM) occurs after a protein is translated from ribonucleic acid. It is an important living creature life phenomenon because it is implicated in almost all cellular processes. Identification of PTM sites from a given protein sequence is a hot topic in bioinformatics. Lots of computational methods have been proposed, and they provide good performance. However, most previous methods can only tackle one PTM type. Few methods consider multiple PTM types. In this study, a multi-label classification model, named RMTLysPTM, was developed to recognize four types of lysine (K) PTM sites, including acetylation, crotonylation, methylation and succinylation. The surrounding sites of a lysine site were selected to constitute a peptide segment, representing the lysine at the center. Deep analysis was conducted to count the distribution of 2-residues with fixed location across the four types of lysine PTM sites. By aggregating the distribution information of 2-residues in one peptide segment, the peptide segment was encoded by informative features. Furthermore, a prediction engine that can precisely capture the traits of the above representations was designed to recognize the types of lysine PTM sites. The cross-validation results on two datasets (Qiu and CPLM training datasets) suggested that the model had extremely high performance and RMTLysPTM had strong generalization ability by testing it on protein Q16778 and CPLM testing datasets. The model was found to be generally superior to all previous models and those using popular methods and features. A web server was set up for RMTLysPTM, and it can be accessed at http://119.3.127.138/. Lei Chen 0007 |
Briefings Bioinform. | 1 |
| 2024 | PMiSLocMF: predicting miRNA subcellular localizations by incorporating multi-source features of miRNAsabstractThe microRNAs (miRNAs) play crucial roles in several biological processes. It is essential for a deeper insight into their functions and mechanisms by detecting their subcellular localizations. The traditional methods for determining miRNAs subcellular localizations are expensive. The computational methods are alternative ways to quickly predict miRNAs subcellular localizations. Although several computational methods have been proposed in this regard, the incomplete representations of miRNAs in these methods left the room for improvement. In this study, a novel computational method for predicting miRNA subcellular localizations, named PMiSLocMF, was developed. As lots of miRNAs have multiple subcellular localizations, this method was a multi-label classifier. Several properties of miRNA, such as miRNA sequences, miRNA functional similarity, miRNA-disease, miRNA-drug, and miRNA-mRNA associations were adopted for generating informative miRNA features. To this end, powerful algorithms [node2vec and graph attention auto-encoder (GATE)] and one newly designed scheme were adopted to process above properties, producing five feature types. All features were poured into self-attention and fully connected layers to make predictions. The cross-validation results indicated the high performance of PMiSLocMF with accuracy higher than 0.83, average area under the receiver operating characteristic curve (AUC) and area under the precision-recall curve (AUPR) exceeding 0.90 and 0.77, respectively. Such performance was better than all previous methods based on the same dataset. Further tests proved that using all feature types can improve the performance of PMiSLocMF, and GATE and self-attention layer can help enhance the performance. Finally, we deeply analyzed the influence of miRNA associations with diseases, drugs, and mRNAs on PMiSLocMF. The dataset and codes are available at https://github.com/Gu20201017/PMiSLocMF. Lei Chen 0007, Jiahui Gu |
Briefings Bioinform. | 1 |
| 2024 | CMAGN: circRNA-miRNA association prediction based on graph attention auto-encoder and network consistency projectionabstractBACKGROUND: As noncoding RNAs, circular RNAs (circRNAs) can act as microRNA (miRNA) sponges due to their abundant miRNA binding sites, allowing them to regulate gene expression and influence disease development. Accurately identifying circRNA-miRNA associations (CMAs) is helpful to understand complex disease mechanisms. Given that biological experiments are time consuming and labor intensive, alternative computational methods to predict CMAs are urgently needed. RESULTS: This study proposes a novel computational model named CMAGN, which incorporates several advanced computational methods, for predicting CMAs. First, similarity networks for circRNAs and miRNAs are constructed according to their sequences. Graph attention autoencoder is then applied to these networks to generate the first representations of circRNAs and miRNAs. The second representations of circRNAs and miRNAs are obtained from the CMA network via node2vec. The similarity networks of circRNAs and miRNAs are reconstructed on the basis of these new representations. Finally, network consistency projection is applied to the reconstructed similarity networks and the CMA network to generate a recommendation matrix. CONCLUSION: Five-fold cross-validation of CMAGN reveals that the area under ROC and PR curves exceed 0.96 on two widely used CMA datasets, outperforming several existing models. Additional tests elaborate the reasonability of the architecture of CMAGN and uncover its strengths and weaknesses. Anhui Yin, Lei Chen 0007, Yu-Dong Cai 0001 |
BMC Bioinform. | 2 |
| 2022 | Identifying Protein Subcellular Locations With Embeddings-Based node2locabstractIdentifying protein subcellular locations is an important topic in protein function prediction. Interacting proteins may share similar locations. Thus, it is imperative to infer protein subcellular locations by taking protein-protein interactions (PPIs)into account. In this study, we present a network embedding-based method, node2loc, to identify protein subcellular locations. node2loc first learns distributed embeddings of proteins in a protein-protein interaction (PPI)network using node2vec. Then the learned embeddings are further fed into a recurrent neural network (RNN). To resolve the severe class imbalance of different subcellular locations, Synthetic Minority Over-sampling Technique (SMOTE)is applied to artificially synthesize proteins for minority classes. node2loc is evaluated on our constructed human benchmark dataset with 16 subcellular locations and yields a Matthews correlation coefficient (MCC)value of 0.800, which is superior to baseline methods. In addition, node2loc yields a better performance on a Yeast benchmark dataset with 17 locations. The results demonstrate that the learned representations from a PPI network have certain discriminative ability for classifying protein subcellular locations. However, node2loc is a transductive method, it only works for proteins connected in a PPI network, and it needs to be retrained for new proteins. In addition, the PPI network needs be annotated to some extent with location information. node2loc is freely available at https://github.com/xypan1232/node2loc. Xiaoyong Pan, Lei Chen 0007, Min Liu 0020, Zhibin Niu, Tao Huang 0004, Yu-Dong Cai 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | iATC-FRAKEL: a simple multi-label web server for recognizing anatomical therapeutic chemical classes of drugs with their fingerprints onlyabstractMOTIVATION: Anatomical therapeutic chemical (ATC) classification system is very important for drug utilization and studies. Correct prediction of the 14 classes in the first level for given drugs is an essential problem for the study on such system. Several multi-label classifiers have been proposed in this regard. However, only two of them provided the web servers and their performance was not very high. On the other hand, although some rest classifiers can provide better performance, they were built based on some prior knowledge on drugs, such as information of chemical-chemical interaction and chemical ontology, leading to limited applications. Furthermore, provided codes of these classifiers are almost inaccessible for pharmacologists. RESULTS: In this study, we built a simple web server, namely iATC-FRAKEL. This web server only required the SMILES format of drugs as input and extracted their fingerprints for making prediction. The performance of the iATC-FRAKEL was much higher than all existing web servers and was comparable to the best multi-label classifier but had much wider applications. Such web server can be visited at http://cie.shmtu.edu.cn/iatc/index. AVAILABILITY AND IMPLEMENTATION: The web server is available at http://cie.shmtu.edu.cn/iatc/index. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jian-Peng Zhou, Lei Chen 0007, Tianyun Wang, Min Liu 0020 |
Bioinform. | 2 |
| 2020 | iATC-NRAKEL: an efficient multi-label classifier for recognizing anatomical therapeutic chemical classes of drugsabstractMOTIVATION: The anatomical therapeutic chemical (ATC) classification system plays an increasingly important role in drug repositioning and discovery. The correct identification of classes in each level of such system that a given drug may belong to is an essential problem. Several multi-label classifiers have been proposed in this regard. Although they provided satisfactory performance, the feature extraction procedures were still rough. More refined features may further improve the predicted quality. RESULTS: In this article, we provide a novel multi-label classifier, called iATC-NRAKEL, to predict drug ATC classes in the first level. To obtain more informative drug features, we employed the drug association information in STITCH and KEGG, which was organized by seven drug networks. The powerful network embedding algorithm, Mashup, was adopted to extract informative drug features. The obtained features were fed into the RAndom k-labELsets (RAKEL) algorithm with support vector machine as the basic classification algorithm to construct the classifier. The 10-fold cross-validation of the benchmark dataset with 3883 drugs showed that the accuracy and absolute true were 76.56 and 74.51%, respectively. The comparison results indicated that iATC-NRAKEL was much superior to all previous reported classifiers. Finally, the contribution of each network was analyzed. AVAILABILITY AND IMPLEMENTATION: The codes of iATC-NRAKEL are available at https://github.com/zhou256/iATC-NRAKEL. Jian-Peng Zhou, Lei Chen 0007, Zi-Han Guo |
Bioinform. | 2 |
| 2017 | Analysis of cancer-related lncRNAs using gene ontology and KEGG pathways
Lei Chen 0007, Yu-Hang Zhang, Guohui Lu, Tao Huang 0004, Yu-Dong Cai 0001 |
Artif. Intell. Medicine | 1 |
| 2017 | Machine learning and graph analytics in computational biomedicine
Quan Zou 0001, Lei Chen 0007, Tao Huang 0004, Yungang Xu |
Artif. Intell. Medicine | 2 |
| 2016 | Identification of novel proliferative diabetic retinopathy related genes on protein-protein interaction network
Jing Yang 0009, Tao Huang 0004, Lei Chen 0007 |
Neurocomputing | 5 |
| 2016 | A Novel Brain Networks Enhancement Model (BNEM) for BOLD fMRI Data Analysis With Highly Spatial ReproducibilityabstractIndependent component analysis aiming at detecting the functional connectivity among discrete cortical brain regions has been extensively used to explore the functional magnetic resonance imaging data. Although the independent components (ICs) were with relatively high quality, the noise embedding in ICs has a great impact on the true active/inactive region inference and the reproducibility, in postprocessing stage, e.g., the extraction of statistical parametrical maps (SPMs). In this paper, a novel brain network enhancement model (BNEM) is proposed, which mainly consists of two key techniques: 1) 3-D wavelet noise filter (3DWNF) for the meaningful ICs, which greatly suppresses noise and enforces the real activation inference of SPMs; and 2) a spatial reproducibility enhancement algorithm (SREA), aiming to improve the reproducibility of SPMs. The simulated experiment demonstrated that the postfiltering signals by 3DWNF were with higher correlation and less normalized mean square error to the ground truths than the prefiltering ones; SREA could further enhance the quality of most postfiltering ones, preserving the consistency with 3DWNF. The real data experiments also revealed that 1) 3DWNF could lead to more accurate preservation of the true positive voxels by correctly identifying the high proportionally misclassified voxels of the nonenhanced SPMs; 2) SREA could further improve the classification accuracy of the active/inactive voxels of SPMs corresponding to the 3DWNF denoised ICs; and 3) both 3DWNF and SREA contribute to the reproducibility enhancement of the reproduced SPMs by BNEM. Thus, BNEM is expected to have wide applicability in the neuroscience and clinical domain. Nizhuan Wang 0001, Weiming Zeng, Dongtailang Chen, Jun Yin 0003, Lei Chen 0007 |
IEEE J. Biomed. Health Informatics | 5 |
| 2012 | NP-completeness and APX-completeness of restrained domination in graphs
Lei Chen 0007, Weiming Zeng, Changhong Lu |
Theor. Comput. Sci. | 1 |
| 2011 | Predicting triplet of transcription factor - mediating enzyme - target gene by functional profiles
Tao Huang 0004, Lei Chen 0007, Xiao-Jun Liu, Yu-Dong Cai 0001 |
Neurocomputing | 2 |
| 2010 | Predicting the network of substrate-enzyme-product triads by combining compound similarity and functional domain compositionabstractBACKGROUND: Metabolic pathway is a highly regulated network consisting of many metabolic reactions involving substrates, enzymes, and products, where substrates can be transformed into products with particular catalytic enzymes. Since experimental determination of the network of substrate-enzyme-product triad (whether the substrate can be transformed into the product with a given enzyme) is both time-consuming and expensive, it would be very useful to develop a computational approach for predicting the network of substrate-enzyme-product triads. RESULTS: A mathematical model for predicting the network of substrate-enzyme-product triads was developed. Meanwhile, a benchmark dataset was constructed that contains 744,192 substrate-enzyme-product triads, of which 14,592 are networking triads, and 729,600 are non-networking triads; i.e., the number of the negative triads was about 50 times the number of the positive triads. The molecular graph was introduced to calculate the similarity between the substrate compounds and between the product compounds, while the functional domain composition was introduced to calculate the similarity between enzyme molecules. The nearest neighbour algorithm was utilized as a prediction engine, in which a novel metric was introduced to measure the "nearness" between triads. To train and test the prediction engine, one tenth of the positive triads and one tenth of the negative triads were randomly picked from the benchmark dataset as the testing samples, while the remaining were used to train the prediction model. It was observed that the overall success rate in predicting the network for the testing samples was 98.71%, with 95.41% success rate for the 1,460 testing networking triads and 98.77% for the 72,960 testing non-networking triads. CONCLUSIONS: It is quite promising and encouraged to use the molecular graph to calculate the similarity between compounds and use the functional domain composition to calculate the similarity between enzymes for studying the substrate-enzyme-product network system. The software is available upon request. Lei Chen 0007, Kai-Yan Feng, Yu-Dong Cai 0001, Kuo-Chen Chou |
BMC Bioinform. | 1 |
| 2009 | A linear-time algorithm for paired-domination problem in strongly chordal graphs
Lei Chen 0007, Changhong Lu, Zhenbing Zeng |
Inf. Process. Lett. | 1 |
| 2009 | Hardness results and approximation algorithms for (weighted) paired-domination in graphs
Lei Chen 0007, Changhong Lu, Zhenbing Zeng |
Theor. Comput. Sci. | 1 |
| 2009 | Distance paired-domination problems on subclasses of chordal graphs
Lei Chen 0007, Changhong Lu, Zhenbing Zeng |
Theor. Comput. Sci. | 1 |
| 2007 | Extremal problems on consecutive L(2, 1)-labelling
Changhong Lu, Lei Chen 0007, Mingqing Zhai |
Discret. Appl. Math. | 2 |