EDBT 2026 Demo / reviewers in the wild / expert
Cui-Xiang Lin
dblp:307/2576
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0004-2391-9164ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MFC-GCN: Identifying Disease Driver Genes via Graph Neural Networks and Multi-Network Fusion Contrastive LearningabstractThe identification of disease-driving genes is crucial for understanding disease mechanisms and advancing precision medicine. However, current methods for integrating heterogeneous biological data face challenges, including suboptimal feature representation and inadequate modeling of cross-modal dependencies, particularly in capturing latent multi-level inter-actions within complex biological networks. To address this, we propose MFC-GCN, a framework that integrates five biological networks—Protein-Protein Interaction, KEGG pathway co-occurrence, Gene Ontology semantic similarity, gene sequence similarity, and gene co-expression—using multidimensional enhancement, cross-network contrastive learning, and adaptive expert selection. A hierarchical graph convolutional encoder extracts deep representations from each modality, which are fused via an attention-guided mixture-of-experts gating network. A bi-directional contrastive learning mechanism further enhances the model's discriminative power. In five-fold cross-validation, our model achieved an average AUROC of 91% and AUPRC of 84%. Ablation studies confirm its ability to capture gene network dependencies and interactions, aiding in the understanding of disease driver genes. MFC-GCN is available at https://github.com/HaoTongXueWang/MFC-GCN. Cui-Xiang Lin, Si-Han Zhu, Hong-Dong Li |
BIBM | 2 |
| 2025 | LIMO-GCN: a linear model-integrated graph convolutional network for predicting Alzheimer disease genesabstractAlzheimer's disease (AD) is a complex disease with its genetic etiology not fully understood. Gene network-based methods have been proven promising in predicting AD genes. However, existing approaches are limited in their ability to model the nonlinear relationship between networks and disease genes, because (i) any data can be theoretically decomposed into the sum of a linear part and a nonlinear part, (ii) the linear part can be best modeled by a linear model since a nonlinear model is biased and can be easily overfit, and (iii) existing methods do not separate the linear part from the nonlinear part when building the disease gene prediction model. To address the limitation, we propose linear model-integrated graph convolutional network (LIMO-GCN), a generic disease gene prediction method that models the data linearity and nonlinearity by integrating a linear model with GCN. The reason to use GCN is that it is by design naturally suitable to dealing with network data, and the reason to integrate a linear model is that the linearity in the data can be best modeled by a linear model. The weighted sum of the prediction of the two components is used as the final prediction of LIMO-GCN. Then, we apply LIMO-GCN to the prediction of AD genes. LIMO-GCN outperforms the state-of-the-art approaches including GCN, network-wide association studies, and random walk. Furthermore, we show that the top-ranked genes are significantly associated with AD based on molecular evidence from heterogeneous genomic data. Our results indicate that LIMO-GCN provides a novel method for prioritizing AD genes. Cui-Xiang Lin, Hong-Dong Li, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2023 | Computational approaches for detecting disease-associated alternative splicing eventsabstractAlternative splicing (AS) is a key transcriptional regulation pathway. Recent studies have shown that AS events are associated with the occurrence of complex diseases. Various computational approaches have been developed for the detection of disease-associated AS events. In this review, we first describe the metrics used for quantitative characterization of AS events. Second, we review and discuss the three types of methods for detecting disease-associated splicing events, which are differential splicing analysis, aberrant splicing detection and splicing-related network analysis. Third, to further exploit the genetic mechanism of disease-associated AS events, we describe the methods for detecting genetic variants that potentially regulate splicing. For each type of methods, we conducted experimental comparison to illustrate their performance. Finally, we discuss the limitations of these methods and point out potential ways to address them. We anticipate that this review provides a systematic understanding of computational approaches for the analysis of disease-associated splicing. Jiashu Liu, Cui-Xiang Lin, Zongxuan Li, Wenkui Huang, Yuanfang Guan, Hong-Dong Li |
Briefings Bioinform. | 2 |
| 2023 | CellBRF: a feature selection method for single-cell clustering using cell balance and random forestabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) offers a powerful tool to dissect the complexity of biological tissues through cell sub-population identification in combination with clustering approaches. Feature selection is a critical step for improving the accuracy and interpretability of single-cell clustering. Existing feature selection methods underutilize the discriminatory potential of genes across distinct cell types. We hypothesize that incorporating such information could further boost the performance of single cell clustering. RESULTS: We develop CellBRF, a feature selection method that considers genes' relevance to cell types for single-cell clustering. The key idea is to identify genes that are most important for discriminating cell types through random forests guided by predicted cell labels. Moreover, it proposes a class balancing strategy to mitigate the impact of unbalanced cell type distributions on feature importance evaluation. We benchmark CellBRF on 33 scRNA-seq datasets representing diverse biological scenarios and demonstrate that it substantially outperforms state-of-the-art feature selection methods in terms of clustering accuracy and cell neighborhood consistency. Furthermore, we demonstrate the outstanding performance of our selected features through three case studies on cell differentiation stage identification, non-malignant cell subtype identification, and rare cell identification. CellBRF provides a new and effective tool to boost single-cell clustering accuracy. AVAILABILITY AND IMPLEMENTATION: All source codes of CellBRF are freely available at https://github.com/xuyp-csu/CellBRF. Yunpei Xu, Hong-Dong Li, Cui-Xiang Lin, Ruiqing Zheng, Yaohang Li, Jinhui Xu 0001, Jianxin Wang 0001 |
Bioinform. | 3 |
| 2023 | A Gene Set-Integrated Approach for Predicting Disease-Associated GenesabstractIt is important to identify disease-associated genes for studying the pathogenic mechanism of complex diseases. Recently, models for disease gene prediction are dominantly based on molecular expression data and networks, including gene expression, protein expression, co-expression networks, protein-protein interaction networks, etc. One limitation of these methods is that they do not consider the knowledge of annotated gene sets representing known pathways or functionally-related sets of genes. In this study, we propose a new approach to predict disease-associated genes by integrating annotated gene sets data from the Molecular Signature Database (MSigDB). It first represents and integrates the different types of annotated gene sets in the MSigDB database in the form of the signal matrix. It then uses the signal matrix as the gene feature to train the disease gene prediction model. We compare our method with existing methods in predicting genes for five complex diseases. The results show that our method is superior to other methods. Further, we perform a case study on autism spectrum disorder (ASD). We find that ASD predictions are associated with ASD based on the statistical analysis of biological networks and independent ASD studies. The source code, prediction results and datasets are publicly available on https://github.com/genemine/GSI.git. Hong-Dong Li, Cui-Xiang Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | An Alzheimer's disease gene prediction method based on ensemble of genome-wide association study summary statisticsabstractThe hitherto unknown specific etiology of Alzheimer’s disease (AD) poses a challenge for its prevention, diagnosis and treatment. Although genome-wide association studies (GWAS) are currently making rapid progress in identifying genetic variants associated with AD, the pathogenic mechanisms of the genetic loci identified are largely unknown. Transcriptome-wide association studies (TWAS) are an important class of methods for predicting disease genes. TWAS can explore the association of genes with the disease in relevant tissues by integrating genome-wide genetic regulatory data from specific tissues and disease-associated GWAS summary statistics. We found that TWAS analysis using different GWAS summary statistics may produce inconsistent results. To address this issue, we used ensemble summary statistics for AD-associated gene prediction considering the complementary nature of different datasets and the comparative nature between the results generated from different datasets. The prediction results were compared and analyzed to identify AD associated genes. The predicted genes were validated. In case study of an individual genes, we identified a potential association between AZGP1 and AD disease by this method. Jia-Hao Song, Cui-Xiang Lin, Hong-Dong Li |
BIBM | 2 |
| 2022 | An integrated brain-specific network identifies genes associated with neuropathologic and clinical traits of Alzheimer's diseaseabstractAlzheimer's disease (AD) has a strong genetic predisposition. However, its risk genes remain incompletely identified. We developed an Alzheimer's brain gene network-based approach to predict AD-associated genes by leveraging the functional pattern of known AD-associated genes. Our constructed network outperformed existing networks in predicting AD genes. We then systematically validated the predictions using independent genetic, transcriptomic, proteomic data, neuropathological and clinical data. First, top-ranked genes were enriched in AD-associated pathways. Second, using external gene expression data from the Mount Sinai Brain Bank study, we found that the top-ranked genes were significantly associated with neuropathological and clinical traits, including the Consortium to Establish a Registry for Alzheimer's Disease score, Braak stage score and clinical dementia rating. The analysis of Alzheimer's brain single-cell RNA-seq data revealed cell-type-specific association of predicted genes with early pathology of AD. Third, by interrogating proteomic data in the Religious Orders Study and Memory and Aging Project and Baltimore Longitudinal Study of Aging studies, we observed a significant association of protein expression level with cognitive function and AD clinical severity. The network, method and predictions could become a valuable resource to advance the identification of risk genes for AD. Cui-Xiang Lin, Hong-Dong Li, Weisheng Liu, Shannon Erhardt, Fang-Xiang Wu, Xing-Ming Zhao, Yuanfang Guan, Jun Wang 0153, Daifeng Wang, Bin Hu 0001, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2022 | GTFtools: a software package for analyzing various features of gene modelsabstractMOTIVATION: Gene-centric bioinformatics studies frequently involve the calculation or the extraction of various features of genes such as splice sites, promoters, independent introns and untranslated regions (UTRs) through manipulation of gene models. Gene models are often annotated in gene transfer format (GTF) files. The features are essential for subsequent analysis such as intron retention detection, DNA-binding site identification and computing splicing strength of splice sites. Some features such as independent introns and splice sites are not provided in existing resources including the commonly used BioMart database. A package that implements and integrates functions to analyze various features of genes will greatly ease routine analysis for related bioinformatics studies. However, to the best of our knowledge, such a package is not available yet. RESULTS: We introduce GTFtools, a stand-alone command-line software that provides a set of functions to calculate various gene features, including splice sites, independent introns, transcription start sites (TSS)-flanking regions, UTRs, isoform coordination and length, different types of gene lengths, etc. It takes the ENSEMBL or GENCODE GTF files as input and can be applied to both human and non-human gene models like the lab mouse. We compare the utilities of GTFtools with those of two related tools: Bedtools and BioMart. GTFtools is implemented in Python and not dependent on any third-party software, making it very easy to install and use. AVAILABILITY AND IMPLEMENTATION: GTFtools is freely available at www.genemine.org/gtftools.php as well as pyPI and Bioconda. Hong-Dong Li, Cui-Xiang Lin, Jiantao Zheng |
Bioinform. | 2 |
| 2022 | AlzCode: a platform for multiview analysis of genes related to Alzheimer's diseaseabstractMOTIVATION: Alzheimer's disease (AD) is a complex brain disorder with risk genes incompletely identified. The candidate genes are dominantly obtained by computational approaches. In order to obtain biological insights of candidate genes or screen genes for experimental testing, it is essential to assess their relevance to AD. A platform that integrates different types of omics data and approaches would facilitate the analysis of candidate genes and is in great need. RESULTS: We report AlzCode, a platform for multiview analysis of genes related to AD. First, this platform integrates a rich collection of functional genomic data, including expression data of AD samples (gene expression, single-cell RNA-seq data and protein expression), AD-specific biological networks (co-expression networks and functional gene networks), neuropathological and clinical traits (CERAD score, Braak staging score, Clinical Dementia Rating, cognitive function and clinical severity) and general data such as protein-protein interaction, regulatory networks, sequence similarity and miRNA-target interactions. These data provide basis for analyzing genes from different views. Second, the platform integrates multiple approaches designed for the various types of data. We implement functions to analyze both individual genes and gene sets. We also compare AlzCode with two existing platforms for AD analysis, which are Agora and AD Atlas. We pinpoint the features of each platform and highlight their differences. This platform would be valuable to the understanding of AD genetics and pathological mechanisms. AVAILABILITY AND IMPLEMENTATION: AlzCode is freely available at: http://www.alzcode.xyz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cui-Xiang Lin, Hong-Dong Li, Shannon Erhardt, Jun Wang 0153, Xiaoqing Peng, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2021 | Improving the Prediction of Disease-associated Genes by Integrating Annotated Gene SetsabstractIdentifying disease-associated genes is key to studying the pathogenic mechanism of complex diseases. Existing methods for predicting disease genes are dominantly based on molecular networks or omics data, including gene expression, protein expression, co-expression networks, protein-protein interaction, etc. The annotated gene sets such as Gene Ontology (GO) terms and Kyoto Encyclopedia of Genes and Genomes pathways represent knowledge about the functional association between genes, which may be of predictive value for disease gene identification. However, the knowledge is barely explored in existing approaches. In this study, we propose a new method for predicting disease-associated genes by integrating annotated gene sets. It first constructs a signal matrix by integrating annotated gene sets in the MSigDB data, a comprehensive database of curated gene sets. Then it uses the signal matrix as input features to build a prediction model with machine learning approaches. We compared our method with existing disease gene prediction methods on five complex diseases. The results showed that our method is superior to other methods. Cui-Xiang Lin, Hong-Dong Li |
BIBM | 2 |
| 2021 | REBET: a method to determine the number of cell clusters based on batch effect removalabstractIn single-cell RNA-seq (scRNA-seq) data analysis, a fundamental problem is to determine the number of cell clusters based on the gene expression profiles. However, the performance of current methods is still far from satisfactory, presumably due to their limitations in capturing the expression variability among cell clusters. Batch effects represent the undesired variability between data measured in different batches. When data are obtained from different labs or protocols batch effects occur. Motivated by the practice of batch effect removal, we considered cell clusters as batches. We hypothesized that the number of cell clusters (i.e. batches) could be correctly determined if the variances among clusters (i.e. batch effects) were removed. We developed a new method, namely, removal of batch effect and testing (REBET), for determining the number of cell clusters. In this method, cells are first partitioned into k clusters. Second, the batch effects among these k clusters are then removed. Third, the quality of batch effect removal is evaluated with the average range of normalized mutual information (ARNMI), which measures how uniformly the cells with batch-effects-removal are mixed. By testing a range of k values, the k value that corresponds to the lowest ARNMI is determined to be the optimal number of clusters. We compared REBET with state-of-the-art methods on 32 simulated datasets and 14 published scRNA-seq datasets. The results show that REBET can accurately and robustly estimate the number of cell clusters and outperform existing methods. Contact: H.D.L. ([email protected]) or Q.S.X. ([email protected]). Zhaoyu Fang, Cui-Xiang Lin, Yun-Pei Xu, Hong-Dong Li |
Briefings Bioinform. | 2 |