EDBT 2026 Demo / reviewers in the wild / expert
Hong-Dong Li
dblp:81/7927
· DBLP profile ↗
34ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0003-3438-739XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 31 · 6 first-author · 24 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pathogenicity Prediction of Missense Mutations Through Attention-Based Multi-omics Feature Integration
Chen-Lu Wang, Zong-Xuan Li, Shu-Qi Chen, Hong-Dong Li |
ISBRA (2) | 4 |
| 2025 | MFC-GCN: Identifying Disease Driver Genes via Graph Neural Networks and Multi-Network Fusion Contrastive LearningabstractThe identification of disease-driving genes is crucial for understanding disease mechanisms and advancing precision medicine. However, current methods for integrating heterogeneous biological data face challenges, including suboptimal feature representation and inadequate modeling of cross-modal dependencies, particularly in capturing latent multi-level inter-actions within complex biological networks. To address this, we propose MFC-GCN, a framework that integrates five biological networks—Protein-Protein Interaction, KEGG pathway co-occurrence, Gene Ontology semantic similarity, gene sequence similarity, and gene co-expression—using multidimensional enhancement, cross-network contrastive learning, and adaptive expert selection. A hierarchical graph convolutional encoder extracts deep representations from each modality, which are fused via an attention-guided mixture-of-experts gating network. A bi-directional contrastive learning mechanism further enhances the model's discriminative power. In five-fold cross-validation, our model achieved an average AUROC of 91% and AUPRC of 84%. Ablation studies confirm its ability to capture gene network dependencies and interactions, aiding in the understanding of disease driver genes. MFC-GCN is available at https://github.com/HaoTongXueWang/MFC-GCN. Cui-Xiang Lin, Si-Han Zhu, Hong-Dong Li |
BIBM | 4 |
| 2025 | A Survey on the Feedback Mechanism of LLM-based AI AgentsabstractLarge language models (LLMs) are increasingly being adopted to develop general-purpose AI agents. However, it remains challenging for these LLM-based AI agents to efficiently learn from feedback and iteratively optimize their strategies. To address this challenge, tremendous efforts have been dedicated to designing diverse feedback mechanisms for LLM-based AI agents. To provide a comprehensive overview of this rapidly evolving field, this paper presents a systematic review of these studies, offering a holistic perspective on the feedback mechanisms in LLM-based AI agents. We begin by discussing the construction of LLM-based AI agents, introducing a generalized framework that encapsulates much of the existing work. Next, we delve into the exploration of feedback mechanisms, categorizing them into four distinct types: internal feedback, external feedback, multi-agent feedback, and human feedback. Additionally, we provide an overview of evaluation protocols and benchmarks specifically tailored for LLM-based AI agents. Finally, we highlight the significant challenges and identify potential directions for future studies. The relevant papers are summarized and will be consistently updated at https://github.com/kevinson7515/Agents-Feedback-Mechanisms. Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Xiucheng Li, Yang Xiang 0003, Jin Liu 0012, Hong-Dong Li, Yaowei Wang 0001, Liqiang Nie, Min Zhang 0005 |
IJCAI | 8 |
| 2025 | LIMO-GCN: a linear model-integrated graph convolutional network for predicting Alzheimer disease genesabstractAlzheimer's disease (AD) is a complex disease with its genetic etiology not fully understood. Gene network-based methods have been proven promising in predicting AD genes. However, existing approaches are limited in their ability to model the nonlinear relationship between networks and disease genes, because (i) any data can be theoretically decomposed into the sum of a linear part and a nonlinear part, (ii) the linear part can be best modeled by a linear model since a nonlinear model is biased and can be easily overfit, and (iii) existing methods do not separate the linear part from the nonlinear part when building the disease gene prediction model. To address the limitation, we propose linear model-integrated graph convolutional network (LIMO-GCN), a generic disease gene prediction method that models the data linearity and nonlinearity by integrating a linear model with GCN. The reason to use GCN is that it is by design naturally suitable to dealing with network data, and the reason to integrate a linear model is that the linearity in the data can be best modeled by a linear model. The weighted sum of the prediction of the two components is used as the final prediction of LIMO-GCN. Then, we apply LIMO-GCN to the prediction of AD genes. LIMO-GCN outperforms the state-of-the-art approaches including GCN, network-wide association studies, and random walk. Furthermore, we show that the top-ranked genes are significantly associated with AD based on molecular evidence from heterogeneous genomic data. Our results indicate that LIMO-GCN provides a novel method for prioritizing AD genes. Cui-Xiang Lin, Hong-Dong Li, Jianxin Wang 0001 |
Briefings Bioinform. | 2 |
| 2025 | CrossIsoFun: predicting isoform functions using the integration of multi-omics dataabstractMOTIVATION: Isoforms spliced from the same gene may carry distinct biological functions. Therefore, annotating functions at the isoform level provides valuable insights into the functional diversity of genomes. Since experimental approaches for determining isoform functions are time- and cost-demanding, computational methods have been proposed. In this case, multi-omics data integration helps enhance the model performance, providing complementary insights for isoform functions. However, current methods underperform in leveraging diverse omics data, primarily due to the limited power to integrate the heterogeneous feature domains. Besides, among the multi-omics data, isoform-isoform interactions (IIIs) are a key data source, as isoforms interact with each other to perform functions. Unfortunately, IIIs remain largely underutilized in isoform function predictions until now. RESULTS: We introduce CrossIsoFun, a multi-omics data analysis framework for isoform function prediction. CrossIsoFun combines omics-specific and cross-omics learning for data integration and function prediction. In detail, CrossIsoFun uses a graph convolutional network (GCN) as the omics-specific classifier for each data source. The initial label predictions from GCNs are forwarded to the View Correlation Discovery Network (VCDN) and processed as a cross-omics integrative representation. The representation is then used to produce final predictions of isoform functions. In addition, an antoencoder within a cycle-consistency generative adversarial network (cycleGAN) is designed to generate IIIs from PPIs and thereby enrich the interactomics data. Our method outperforms the state-of-the-art methods on three tissue-naive datasets and 15 tissue-specific datasets with mRNA expression, sequence, and PPI data. The prediction of CrossIsoFun is further validated by its consistency with subcellular localization and isoform-level annotations with literature support. AVAILABILITY AND IMPLEMENTATION: CrossIsoFun is freely available at https://github.com/genemine/CrossIsoFun. Hong-Dong Li, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2025 | Inspired by pathogenic mechanisms: A novel gradual multi-modal fusion framework for mild cognitive impairment diagnosis
Hong-Dong Li, Hanhe Lin, Chao Li 0031, Harrison X. Bai, Wei Lan 0001, Jin Liu 0012 |
Neural Networks | 2 |
| 2024 | Pre-trained Feature Fusion and Matching for Mild Cognitive Impairment DetectionabstractEffective diagnosis of Mild Cognitive Impairment (MCI), a preclinical stage of cognitive decline, is significant for delaying disease progression.While most current spontaneous speechbased diagnostic methods focus on English speech, the Interspeech 2024 TAUKADIAL Challenge proposed an innovative research direction to develop a language-agnostic approach to diagnose MCI.This paper proposes an MCI diagnosis method by analyzing and combining linguistic and acoustic features using the bilingual Chinese-English speech dataset provided by the challenge.We employed a pre-trained multilingual model and expressivity encoder to extract language-agnostic speech features.To overcome the challenges of data scarcity and language diversity, we implemented data augmentation and alignment to enhance the model's generalization.Our approach achieved 77.5% accuracy, demonstrating its effectiveness and potential on cross-lingual data. Junwen Duan, Fangyuan Wei, Hong-Dong Li, Jin Liu 0012 |
INTERSPEECH | 3 |
| 2024 | Identifying new cancer genes based on the integration of annotated gene sets via hypergraph neural networksabstractMOTIVATION: Identifying cancer genes remains a significant challenge in cancer genomics research. Annotated gene sets encode functional associations among multiple genes, and cancer genes have been shown to cluster in hallmark signaling pathways and biological processes. The knowledge of annotated gene sets is critical for discovering cancer genes but remains to be fully exploited. RESULTS: Here, we present the DIsease-Specific Hypergraph neural network (DISHyper), a hypergraph-based computational method that integrates the knowledge from multiple types of annotated gene sets to predict cancer genes. First, our benchmark results demonstrate that DISHyper outperforms the existing state-of-the-art methods and highlight the advantages of employing hypergraphs for representing annotated gene sets. Second, we validate the accuracy of DISHyper-predicted cancer genes using functional validation results and multiple independent functional genomics data. Third, our model predicts 44 novel cancer genes, and subsequent analysis shows their significant associations with multiple types of cancers. Overall, our study provides a new perspective for discovering cancer genes and reveals previously undiscovered cancer genes. AVAILABILITY AND IMPLEMENTATION: DISHyper is freely available for download at https://github.com/genemine/DISHyper. Hong-Dong Li, Li-Shen Zhang, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2024 | Graph-Based Fusion of Imaging, Genetic and Clinical Data for Degenerative Disease DiagnosisabstractGraph learning methods have achieved noteworthy performance in disease diagnosis due to their ability to represent unstructured information such as inter-subject relationships. While it has been shown that imaging, genetic and clinical data are crucial for degenerative disease diagnosis, existing methods rarely consider how best to use their relationships. How best to utilize information from imaging, genetic and clinical data remains a challenging problem. This study proposes a novel graph-based fusion (GBF) approach to meet this challenge. To extract effective imaging-genetic features, we propose an imaging-genetic fusion module which uses an attention mechanism to obtain modality-specific and joint representations within and between imaging and genetic data. Then, considering the effectiveness of clinical information for diagnosing degenerative diseases, we propose a multi-graph fusion module to further fuse imaging-genetic and clinical features, which adopts a learnable graph construction strategy and a graph ensemble method. Experimental results on two benchmarks for degenerative disease diagnosis (Alzheimers Disease Neuroimaging Initiative and Parkinson's Progression Markers Initiative) demonstrate its effectiveness compared to state-of-the-art graph-based methods. Our findings should help guide further development of graph-based models for dealing with imaging, genetic and clinical data. Rui Guo 0009, Hanhe Lin, Stephen J. McKenna, Hong-Dong Li, Fei Guo 0001, Jin Liu 0012 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Computational approaches for detecting disease-associated alternative splicing eventsabstractAlternative splicing (AS) is a key transcriptional regulation pathway. Recent studies have shown that AS events are associated with the occurrence of complex diseases. Various computational approaches have been developed for the detection of disease-associated AS events. In this review, we first describe the metrics used for quantitative characterization of AS events. Second, we review and discuss the three types of methods for detecting disease-associated splicing events, which are differential splicing analysis, aberrant splicing detection and splicing-related network analysis. Third, to further exploit the genetic mechanism of disease-associated AS events, we describe the methods for detecting genetic variants that potentially regulate splicing. For each type of methods, we conducted experimental comparison to illustrate their performance. Finally, we discuss the limitations of these methods and point out potential ways to address them. We anticipate that this review provides a systematic understanding of computational approaches for the analysis of disease-associated splicing. Jiashu Liu, Cui-Xiang Lin, Zongxuan Li, Wenkui Huang, Yuanfang Guan, Hong-Dong Li |
Briefings Bioinform. | 8 |
| 2023 | FEED: a feature selection method based on gene expression decomposition for single cell clusteringabstractSingle-cell clustering is a critical step in biological downstream analysis. The clustering performance could be effectively improved by extracting cell-type-specific genes. The state-of-the-art feature selection methods usually calculate the importance of a single gene without considering the information contained in the gene expression distribution. Moreover, these methods ignore the intrinsic expression patterns of genes and heterogeneity within groups of different mean expression levels. In this work, we present a Feature sElection method based on gene Expression Decomposition (FEED) of scRNA-seq data, which selects informative genes to enhance clustering performance. First, the expression levels of genes are decomposed into multiple Gaussian components. Then, a novel gene correlation calculation method is proposed to measure the relationship between genes from the perspective of distribution. Finally, a permutation-based approach is proposed to determine the threshold of gene importance to obtain marker gene subsets. Compared with state-of-the-art feature selection methods, applying FEED on various scRNA-seq datasets including large datasets followed by different common clustering algorithms results in significant improvements in the accuracy of cell-type identification. The source codes for FEED are freely available at https://github.com/genemine/FEED. Zhi-Wei Duan, Yun-Pei Xu, Hong-Dong Li |
Briefings Bioinform. | 5 |
| 2023 | IsoFrog: a reversible jump Markov Chain Monte Carlo feature selection-based method for predicting isoform functionsabstractMOTIVATION: A single gene may yield several isoforms with different functions through alternative splicing. Continuous efforts are devoted to developing machine-learning methods to predict isoform functions. However, existing methods do not consider the relevance of each feature to specific functions and ignore the noise caused by the irrelevant features. In this case, we hypothesize that constructing a feature selection framework to extract the function-relevant features might help improve the model accuracy in isoform function prediction. RESULTS: In this article, we present a feature selection-based approach named IsoFrog to predict isoform functions. First, IsoFrog adopts a reversible jump Markov Chain Monte Carlo (RJMCMC)-based feature selection framework to assess the feature importance to gene functions. Second, a sequential feature selection procedure is applied to select a subset of function-relevant features. This strategy screens the relevant features for the specific function while eliminating irrelevant ones, improving the effectiveness of the input features. Then, the selected features are input into our proposed method modified domain-invariant partial least squares, which prioritizes the most likely positive isoform for each positive MIG and utilizes diPLS for isoform function prediction. Tested on three datasets, our method achieves superior performance over six state-of-the-art methods, and the RJMCMC-based feature selection framework outperforms three classic feature selection methods. We expect this proposed methodology will promote the identification of isoform functions and further inspire the development of new methods. AVAILABILITY AND IMPLEMENTATION: IsoFrog is freely available at https://github.com/genemine/IsoFrog. Changhuo Yang, Hong-Dong Li, Jianxin Wang 0001 |
Bioinform. | 3 |
| 2023 | CellBRF: a feature selection method for single-cell clustering using cell balance and random forestabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) offers a powerful tool to dissect the complexity of biological tissues through cell sub-population identification in combination with clustering approaches. Feature selection is a critical step for improving the accuracy and interpretability of single-cell clustering. Existing feature selection methods underutilize the discriminatory potential of genes across distinct cell types. We hypothesize that incorporating such information could further boost the performance of single cell clustering. RESULTS: We develop CellBRF, a feature selection method that considers genes' relevance to cell types for single-cell clustering. The key idea is to identify genes that are most important for discriminating cell types through random forests guided by predicted cell labels. Moreover, it proposes a class balancing strategy to mitigate the impact of unbalanced cell type distributions on feature importance evaluation. We benchmark CellBRF on 33 scRNA-seq datasets representing diverse biological scenarios and demonstrate that it substantially outperforms state-of-the-art feature selection methods in terms of clustering accuracy and cell neighborhood consistency. Furthermore, we demonstrate the outstanding performance of our selected features through three case studies on cell differentiation stage identification, non-malignant cell subtype identification, and rare cell identification. CellBRF provides a new and effective tool to boost single-cell clustering accuracy. AVAILABILITY AND IMPLEMENTATION: All source codes of CellBRF are freely available at https://github.com/xuyp-csu/CellBRF. Yunpei Xu, Hong-Dong Li, Cui-Xiang Lin, Ruiqing Zheng, Yaohang Li, Jinhui Xu 0001, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2023 | A Gene Set-Integrated Approach for Predicting Disease-Associated GenesabstractIt is important to identify disease-associated genes for studying the pathogenic mechanism of complex diseases. Recently, models for disease gene prediction are dominantly based on molecular expression data and networks, including gene expression, protein expression, co-expression networks, protein-protein interaction networks, etc. One limitation of these methods is that they do not consider the knowledge of annotated gene sets representing known pathways or functionally-related sets of genes. In this study, we propose a new approach to predict disease-associated genes by integrating annotated gene sets data from the Molecular Signature Database (MSigDB). It first represents and integrates the different types of annotated gene sets in the MSigDB database in the form of the signal matrix. It then uses the signal matrix as the gene feature to train the disease gene prediction model. We compare our method with existing methods in predicting genes for five complex diseases. The results show that our method is superior to other methods. Further, we perform a case study on autism spectrum disorder (ASD). We find that ASD predictions are associated with ASD based on the statistical analysis of biological networks and independent ASD studies. The source code, prediction results and datasets are publicly available on https://github.com/genemine/GSI.git. Hong-Dong Li, Cui-Xiang Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | A Comparison of Topologically Associating Domain Callers Based on Hi-C DataabstractTopologically associating domains (TADs) are local chromatin interaction domains, which have been shown to play an important role in gene expression regulation. TADs were originally discovered in the investigation of 3D genome organization based on High-throughput Chromosome Conformation Capture (Hi-C) data. Continuous considerable efforts have been dedicated to developing methods for detecting TADs from Hi-C data. Different computational methods for TADs identification vary in their assumptions and criteria in calling TADs. As a consequence, the TADs called by these methods differ in their similarities and biological features they are enriched in. In this work, we performed a systematic comparison of twenty-six TAD callers. We first compared the TADs and gaps between adjacent TADs across different methods, resolutions, and sequencing depths. We then assessed the quality of TADs and TAD boundaries according to three criteria: the decay of contact frequencies over the genomic distance, enrichment and depletion of regulatory elements around TAD boundaries, and reproducibility of TADs and TAD boundaries in replicate samples. Last, due to the lack of a gold standard of TADs, we also evaluated the performance of the methods on synthetic datasets. We discussed the key principles of TAD callers, and pinpointed current situation in the detection of TADs. We provide a concise, comprehensive, and systematic framework for evaluating the performance of TAD callers, and expect our work will provide useful guidance in choosing suitable approaches for the detection and evaluation of TADs. Kun Liu 0028, Hong-Dong Li, Yaohang Li, Jun Wang 0153, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | IsoCell: An Approach to Enhance Single Cell Clustering by Integrating Isoform-Level Expression Through Orthogonal ProjectionabstractSingle cell RNA sequencing (scRNA-seq) provides a powerful approach for profiling transcriptomes at single cell resolution. An essential application of scRNA-seq is the discovery of cell types with the aid of clustering analysis. Currently, existing single cell clustering methods are exclusively based on gene-level expression data, without considering alternative splicing information. It has been shown that alternative splicing has an important influence on biological processes such as cell differentiation and cell cycle. We therefore hypothesize that adding information about alternative splicing may help enhance single cell clustering. This motivates us to develop a way to integrate isoform-level expression and gene-level expression. We report an approach to enhance single cell clustering by integrating isoform-level expression through orthogonal projection. First, we construct an orthogonal projection matrix based on gene expression data. Second, isoforms are projected to the gene space to remove the redundant information between them. Third, isoform selection is performed based on the residual of the projected expression and the selected isoforms are combined with gene expression data for subsequent clustering. We applied our method to sixteen scRNA-seq datasets. We find that alternative splicing contains differential information among cell types and can be integrated to enhance single cell clustering. Compared with using only gene-level expression data, the integration of isoform-level expression leads to better clustering performances for most of the datasets. The integration of isoform-level expression also has potential in the detection of novel cell subgroups. Our study shows that integrating isoform and gene-level expression is a promising way to improve single cell clustering. The IsoCell R package is freely available at both Github (https://github.com/genemine/IsoCell) and Zenodo (https://zenodo.org/record/4395707). Yingyi Liu, Hong-Dong Li, Yunpei Xu, Yi-Wei Liu, Xiaoqing Peng, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | KL-RF: Predicting disease-gene associations with model fusionabstractAppropriate disease gene prediction methods are not only critical for the diagnosis of many genotypic diseases, but well-positioned to help scientists discover the links between complex diseases and disease-causing genes. As the development of high-throughput biotechnology, a great quantity of biomolecular interaction networks and omics data emerge rapidly. Because of the capacity of biomolecular networks to indicate complex interactions between elements in disease, they are one of the most frequently used data in the field of disease-gene associations prediction. Fusing these biomolecular networks effectively is hopeful to improve the performance of computational methods. Here, we develop a method to predict the associations between disease and gene by fusing the information of heterogeneous datasets. This method first learns the latent feature representation from disease-gene similarity networks via graph representation learning methods, then utilizes a Beyesian approach to obtain the posterior distribution of the feature matrix, and subsequently fuses the posterior distribution based on KL divergence and predicts the associations by random forest (KL-RF) finally. We compared KL-RF with several classic methods, and the results demonstrate that it performs better than the existing methods. Further analysis of the top-10 unknown disease genes and case study also reveals the ability of the proposed method. The source codes is available on https://github.com/genemine/KL-RF. Hong-Dong Li |
BIBM | 2 |
| 2022 | An Alzheimer's disease gene prediction method based on ensemble of genome-wide association study summary statisticsabstractThe hitherto unknown specific etiology of Alzheimer’s disease (AD) poses a challenge for its prevention, diagnosis and treatment. Although genome-wide association studies (GWAS) are currently making rapid progress in identifying genetic variants associated with AD, the pathogenic mechanisms of the genetic loci identified are largely unknown. Transcriptome-wide association studies (TWAS) are an important class of methods for predicting disease genes. TWAS can explore the association of genes with the disease in relevant tissues by integrating genome-wide genetic regulatory data from specific tissues and disease-associated GWAS summary statistics. We found that TWAS analysis using different GWAS summary statistics may produce inconsistent results. To address this issue, we used ensemble summary statistics for AD-associated gene prediction considering the complementary nature of different datasets and the comparative nature between the results generated from different datasets. The prediction results were compared and analyzed to identify AD associated genes. The predicted genes were validated. In case study of an individual genes, we identified a potential association between AZGP1 and AD disease by this method. Jia-Hao Song, Cui-Xiang Lin, Hong-Dong Li |
BIBM | 3 |
| 2022 | An integrated brain-specific network identifies genes associated with neuropathologic and clinical traits of Alzheimer's diseaseabstractAlzheimer's disease (AD) has a strong genetic predisposition. However, its risk genes remain incompletely identified. We developed an Alzheimer's brain gene network-based approach to predict AD-associated genes by leveraging the functional pattern of known AD-associated genes. Our constructed network outperformed existing networks in predicting AD genes. We then systematically validated the predictions using independent genetic, transcriptomic, proteomic data, neuropathological and clinical data. First, top-ranked genes were enriched in AD-associated pathways. Second, using external gene expression data from the Mount Sinai Brain Bank study, we found that the top-ranked genes were significantly associated with neuropathological and clinical traits, including the Consortium to Establish a Registry for Alzheimer's Disease score, Braak stage score and clinical dementia rating. The analysis of Alzheimer's brain single-cell RNA-seq data revealed cell-type-specific association of predicted genes with early pathology of AD. Third, by interrogating proteomic data in the Religious Orders Study and Memory and Aging Project and Baltimore Longitudinal Study of Aging studies, we observed a significant association of protein expression level with cognitive function and AD clinical severity. The network, method and predictions could become a valuable resource to advance the identification of risk genes for AD. Cui-Xiang Lin, Hong-Dong Li, Weisheng Liu, Shannon Erhardt, Fang-Xiang Wu, Xing-Ming Zhao, Yuanfang Guan, Jun Wang 0153, Daifeng Wang, Bin Hu 0001, Jianxin Wang 0001 |
Briefings Bioinform. | 2 |
| 2022 | GTFtools: a software package for analyzing various features of gene modelsabstractMOTIVATION: Gene-centric bioinformatics studies frequently involve the calculation or the extraction of various features of genes such as splice sites, promoters, independent introns and untranslated regions (UTRs) through manipulation of gene models. Gene models are often annotated in gene transfer format (GTF) files. The features are essential for subsequent analysis such as intron retention detection, DNA-binding site identification and computing splicing strength of splice sites. Some features such as independent introns and splice sites are not provided in existing resources including the commonly used BioMart database. A package that implements and integrates functions to analyze various features of genes will greatly ease routine analysis for related bioinformatics studies. However, to the best of our knowledge, such a package is not available yet. RESULTS: We introduce GTFtools, a stand-alone command-line software that provides a set of functions to calculate various gene features, including splice sites, independent introns, transcription start sites (TSS)-flanking regions, UTRs, isoform coordination and length, different types of gene lengths, etc. It takes the ENSEMBL or GENCODE GTF files as input and can be applied to both human and non-human gene models like the lab mouse. We compare the utilities of GTFtools with those of two related tools: Bedtools and BioMart. GTFtools is implemented in Python and not dependent on any third-party software, making it very easy to install and use. AVAILABILITY AND IMPLEMENTATION: GTFtools is freely available at www.genemine.org/gtftools.php as well as pyPI and Bioconda. Hong-Dong Li, Cui-Xiang Lin, Jiantao Zheng |
Bioinform. | 1 |
| 2022 | AlzCode: a platform for multiview analysis of genes related to Alzheimer's diseaseabstractMOTIVATION: Alzheimer's disease (AD) is a complex brain disorder with risk genes incompletely identified. The candidate genes are dominantly obtained by computational approaches. In order to obtain biological insights of candidate genes or screen genes for experimental testing, it is essential to assess their relevance to AD. A platform that integrates different types of omics data and approaches would facilitate the analysis of candidate genes and is in great need. RESULTS: We report AlzCode, a platform for multiview analysis of genes related to AD. First, this platform integrates a rich collection of functional genomic data, including expression data of AD samples (gene expression, single-cell RNA-seq data and protein expression), AD-specific biological networks (co-expression networks and functional gene networks), neuropathological and clinical traits (CERAD score, Braak staging score, Clinical Dementia Rating, cognitive function and clinical severity) and general data such as protein-protein interaction, regulatory networks, sequence similarity and miRNA-target interactions. These data provide basis for analyzing genes from different views. Second, the platform integrates multiple approaches designed for the various types of data. We implement functions to analyze both individual genes and gene sets. We also compare AlzCode with two existing platforms for AD analysis, which are Agora and AD Atlas. We pinpoint the features of each platform and highlight their differences. This platform would be valuable to the understanding of AD genetics and pathological mechanisms. AVAILABILITY AND IMPLEMENTATION: AlzCode is freely available at: http://www.alzcode.xyz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cui-Xiang Lin, Hong-Dong Li, Shannon Erhardt, Jun Wang 0153, Xiaoqing Peng, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2021 | Improving the Prediction of Disease-associated Genes by Integrating Annotated Gene SetsabstractIdentifying disease-associated genes is key to studying the pathogenic mechanism of complex diseases. Existing methods for predicting disease genes are dominantly based on molecular networks or omics data, including gene expression, protein expression, co-expression networks, protein-protein interaction, etc. The annotated gene sets such as Gene Ontology (GO) terms and Kyoto Encyclopedia of Genes and Genomes pathways represent knowledge about the functional association between genes, which may be of predictive value for disease gene identification. However, the knowledge is barely explored in existing approaches. In this study, we propose a new method for predicting disease-associated genes by integrating annotated gene sets. It first constructs a signal matrix by integrating annotated gene sets in the MSigDB data, a comprehensive database of curated gene sets. Then it uses the signal matrix as input features to build a prediction model with machine learning approaches. We compared our method with existing disease gene prediction methods on five complex diseases. The results showed that our method is superior to other methods. Cui-Xiang Lin, Hong-Dong Li |
BIBM | 3 |
| 2021 | GraphIsoFun: a graph neural network based approach for splice isoform function predictionabstractAlternative splicing has a significant impact in the process of gene regulation and it greatly raises the complexity of protein species and functions. At present, the gene function has been extensively studied, and there are multiple public databases of gene function annotations. Due to the high cost and the large number of isoforms, it is currently impossible to conduct extensive experimental analysis on isoform functions. Functions at the level of splicing isoforms remain largely unknown. Computational prediction provides an alternative to functional annotation of splice isoforms. Here we present GraphIsoFun, a novel computational approach to predict isoform function by integrating both gene and isoform-level expression information with graph neural network. It first constructs a heterogeneous co-expression network with both genes and isoforms being nodes. This network integrates the rich association between isoforms and genes, and provides a bridge for connecting isoform and gene functions. Then GraphIsoFun uses the graph neural network to mine the information of the co-expression network and obtains the final prediction score of isoforms for each function. The experiments on multiple datasets show that GraphIsoFun surpasses existing methods based on different evaluation indicators. Changhuo Yang, Hong-Dong Li, Jianxin Wang 0001 |
BIBM | 3 |
| 2021 | REBET: a method to determine the number of cell clusters based on batch effect removalabstractIn single-cell RNA-seq (scRNA-seq) data analysis, a fundamental problem is to determine the number of cell clusters based on the gene expression profiles. However, the performance of current methods is still far from satisfactory, presumably due to their limitations in capturing the expression variability among cell clusters. Batch effects represent the undesired variability between data measured in different batches. When data are obtained from different labs or protocols batch effects occur. Motivated by the practice of batch effect removal, we considered cell clusters as batches. We hypothesized that the number of cell clusters (i.e. batches) could be correctly determined if the variances among clusters (i.e. batch effects) were removed. We developed a new method, namely, removal of batch effect and testing (REBET), for determining the number of cell clusters. In this method, cells are first partitioned into k clusters. Second, the batch effects among these k clusters are then removed. Third, the quality of batch effect removal is evaluated with the average range of normalized mutual information (ARNMI), which measures how uniformly the cells with batch-effects-removal are mixed. By testing a range of k values, the k value that corresponds to the lowest ARNMI is determined to be the optimal number of clusters. We compared REBET with state-of-the-art methods on 32 simulated datasets and 14 published scRNA-seq datasets. The results show that REBET can accurately and robustly estimate the number of cell clusters and outperform existing methods. Contact: H.D.L. ([email protected]) or Q.S.X. ([email protected]). Zhaoyu Fang, Cui-Xiang Lin, Yun-Pei Xu, Hong-Dong Li |
Briefings Bioinform. | 4 |
| 2021 | Identifying the tissues-of-origin of circulating cell-free DNAs is a promising way in noninvasive diagnosticsabstractAdvances in sequencing technologies facilitate personalized disease-risk profiling and clinical diagnosis. In recent years, some great progress has been made in noninvasive diagnoses based on cell-free DNAs (cfDNAs). It exploits the fact that dead cells release DNA fragments into the circulation, and some DNA fragments carry information that indicates their tissues-of-origin (TOOs). Based on the signals used for identifying the TOOs of cfDNAs, the existing methods can be classified into three categories: cfDNA mutation-based methods, methylation pattern-based methods and cfDNA fragmentation pattern-based methods. In cfDNA mutation-based methods, the SNP information or the detected mutations in driven genes of certain diseases are employed to identify the TOOs of cfDNAs. Methylation pattern-based methods are developed to identify the TOOs of cfDNAs based on the tissue-specific methylation patterns. In cfDNA fragmentation pattern-based methods, cfDNA fragmentation patterns, such as nucleosome positioning or preferred end coordinates of cfDNAs, are used to predict the TOOs of cfDNAs. In this paper, the strategies and challenges in each category are reviewed. Furthermore, the representative applications based on the TOOs of cfDNAs, including noninvasive prenatal testing, noninvasive cancer screening, transplantation rejection monitoring and parasitic infection detection, are also reviewed. Moreover, the challenges and future work in identifying the TOOs of cfDNAs are discussed. Our research provides a comprehensive picture of the development and challenges in identifying the TOOs of cfDNAs, which may benefit bioinformatics researchers to develop new methods to improve the identification of the TOOs of cfDNAs. Xiaoqing Peng, Hong-Dong Li, Fang-Xiang Wu, Jianxin Wang 0001 |
Briefings Bioinform. | 2 |
| 2021 | IsoResolve: predicting splice isoform functions by integrating gene and isoform-level features with domain adaptationabstractMOTIVATION: High resolution annotation of gene functions is a central goal in functional genomics. A single gene may produce multiple isoforms with different functions through alternative splicing. Conventional approaches, however, consider a gene as a single entity without differentiating these functionally different isoforms. Towards understanding gene functions at higher resolution, recent efforts have focused on predicting the functions of isoforms. However, the performance of existing methods is far from satisfactory mainly because of the lack of isoform-level functional annotation. RESULTS: We present IsoResolve, a novel approach for isoform function prediction, which leverages the information from gene function prediction models with domain adaptation (DA). IsoResolve treats gene-level and isoform-level features as source and target domains, respectively. It uses DA to project the two domains into a latent variable space in such a way that the latent variables from the two domains have similar distribution, which enables the gene domain information to be leveraged for isoform function prediction. We systematically evaluated the performance of IsoResolve in predicting functions. Compared with five state-of-the-art methods, IsoResolve achieved significantly better performance. IsoResolve was further validated by case studies of genes with isoform-level functional annotation. AVAILABILITY AND IMPLEMENTATION: IsoResolve is freely available at https://github.com/genemine/IsoResolve. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong-Dong Li, Changhuo Yang, Mengyun Yang, Fang-Xiang Wu, Gilbert S. Omenn, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2021 | A Gene Rank Based Approach for Single Cell Similarity Assessment and ClusteringabstractSingle-cell RNA sequencing (scRNA-seq) technology provides quantitative gene expression profiles at single-cell resolution. As a result, researchers have established new ways to explore cell population heterogeneity and genetic variability of cells. One of the current research directions for scRNA-seq data is to identify different cell types accurately through unsupervised clustering methods. However, scRNA-seq data analysis is challenging because of their high noise level, high dimensionality and sparsity. Moreover, the impact of multiple latent factors on gene expression heterogeneity and on the ability to accurately identify cell types remains unclear. How to overcome these challenges to reveal the biological difference between cell types has become the key to analyze scRNA-seq data. For these reasons, the unsupervised learning for cell population discovery based on scRNA-seq data analysis has become an important research area. A cell similarity assessment method plays a significant role in cell clustering. Here, we present BioRank, a new cell similarity assessment method based on annotated gene sets and gene ranks. To evaluate the performances, we cluster cells by two classical clustering algorithms based on the similarity between cells obtained by BioRank. In addition, BioRank can be used by any clustering algorithm that requires a similarity matrix. Applying BioRank to 12 public scRNA-seq datasets, we show that it is better than or at least as well as several popular similarity assessment methods for single cell clustering. Yunpei Xu, Hong-Dong Li, Yi Pan 0001, Feng Luo 0001, Fang-Xiang Wu, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Joint Learning of Primary and Secondary Labels based on Multi-scale Representation for Alzheimer's Disease DiagnosisabstractThe cause of Alzheimer's disease (AD) is insufficient to understand so far, and its diagnosis is challenging in clinical practice. Recently, the convolutional neural network (CNN) model has shown impressive performance in medical image analysis. Combining CNN with magnetic resonance imaging (MRI) image has excellent potential for AD diagnosis. However, it is still a challenging task. To address the challenge, we propose a joint learning method based on multi-scale representation (JL-MSR). The multi-scale representation is proposed to obtain more feature maps by the multi-scale atrous convolutions. Furthermore, in order to use the intrinsic relationship between diagnostic results and clinical scores, we propose a joint learning strategy using the diagnosis result as the primary label and the Mini-Mental State Examination (MMSE) score as the secondary label to joint training. The proposed method is evaluated on a dataset of 417 subjects (including 188 AD and 229 health controls (HC)) from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The experimental results show that our proposed method achieves an accuracy of 88.1% and an area under the receiver operating characteristic (ROC) curve (AUC) value of 0.942 for AD diagnosis, respectively. Compared with a state-of-the-art method in AD diagnosis, our proposed method performs better, and has potential in clinical diagnosis. Hong-Dong Li, Rui Guo 0009, Junjian Li, Jianxin Wang 0001, Yi Pan 0001, Jin Liu 0012 |
BIBM | 1 |
| 2020 | Identification of Disease-Associated Genes Based on Differential Intron RetentionabstractDifferential analysis of gene-level expression is a commonly used approach for identifying disease-associated genes. Recently, intron retention (IR) has been shown to be associated with complex diseases such as cancers. IR provides value that is complementary to traditional gene-level expression. However, a systematic method to exploit IR for identifying disease-associated genes remains largely unexplored. We developed a pipeline to identify the disease-associated gene based on differential intron retention (IRDAG), which integrates IR events detected by two methods, IRFinder and iREAD. We applied it to Alzheimer's disease (AD). We found that many of the differential genes detected based on IR were not able to be discovered by the traditional gene-level differential expression method, suggesting that our method is complementary to traditional methods. We showed that the differential genes identified with IRDAG were functionally related to AD based on the analysis of protein-protein interaction networks and brain-specific functional gene networks. Being complementary to the existing method, IRDAG provides a new and generic approach for identifying the disease-associated gene. Zhenpeng Wu, Jiantao Zheng, Hong-Dong Li |
BIBM | 3 |
| 2019 | A Global Similarity Learning for Clustering of Single-Cell RNA-Seq DataabstractSingle-cell RNA-seq (scRNA-seq) data analysis is a powerful tool for biological researches. Similarity plays an important role in clustering scRNA-seq data. Existing similarity measurements are mainly based on local distance information that is calculated between directly connected node pairs, or shared nearest neighbours' information, without considering the global information. Therefore, these similarity measurements may be not very accurate based on the insufficient information. Based on multi-kernel indices in a global feature space and path-based similarity, we proposed a new similarity measurement for single-cell clustering, called multi-kernel and path-based global similarity (MPGS). In MPGS, global information was incorporated by a new feature space from Spearman correlation coefficient, and a global similarity matrix calculated by multi-kernel. A path-based similarity metric was designed to expand the relevant node range. Based on this similaritiy, a modified Louvain community detection method was applied to cluster the scRNA-seq data, named MPGS-Louvain. To validate the performance of MPGS, the clustering performances of several clustering methods combined with different similarity measurements were compared. To demonstrate the performance of MPGS-Louvain, we compared MPGS-Louvain and five scRNA-seq clustering methods on twenty scRNA-seq datasets. The experimental results showed that MPGS outperformed other similarity measurements, and MPGS-Louvain achieved better performance on these datasets. It can be observed that MPGS provided a new insight to improve the accuracy of clustering scRNA-seq data by considering the global information in similarity measurement. MPGS-Louvain automatically detected clusters accurately without prior knowledge. Xiaoshu Zhu, Lilu Guo, Yunpei Xu, Hong-Dong Li, Xingyu Liao, Fang-Xiang Wu, Xiaoqing Peng |
BIBM | 4 |
| 2019 | BaiHui: cross-species brain-specific network built with hundreds of hand-curated datasetsabstractMOTIVATION: Functional gene networks, representing how likely two genes work in the same biological process, are important models for studying gene interactions in complex tissues. However, a limitation of the current network-building scheme is the lack of leveraging evidence from multiple model organisms as well as the lack of expert curation and quality control of the input genomic data. RESULTS: Here, we present BaiHui, a brain-specific functional gene network built by probabilistically integrating expertly-hand-curated (by reading original publications) heterogeneous and multi-species genomic data in human, mouse and rat brains. To facilitate the use of this network, we deployed a web server through which users can query their genes of interest, visualize the network, gain functional insight from enrichment analysis and download network data. We also illustrated how this network could be used to generate testable hypotheses on disease gene prioritization of brain disorders. AVAILABILITY AND IMPLEMENTATION: BaiHui is freely available at: http://guanlab.ccmb.med.umich.edu/BaiHui/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong-Dong Li, Tianjian Bai, Erin Sandford, Margit Burmeister, Yuanfang Guan |
Bioinform. | 1 |
| 2018 | BioRank: A Similarity Assessment Method for Single Cell Clustering
Yunpei Xu, Hong-Dong Li, Yi Pan 0001, Feng Luo 0001, Jianxin Wang 0001 |
BIBM | 2 |
| 2017 | An interpretable model for predicting side effects of analgesics for osteoarthritisabstractOsteoarthritis (OA) is the most common type of arthritis. Analgesics are widely used in the process of the treatment of arthritis. Analgesics are particularly used by OA patients which may increase the risk of cardiovascular disease by 20% to 50% overall. In this study, we proposed an interpretable model to predict side effects of analgesics on cardiovascular disease for OA patients. One task of our study is to predict whether OA patients can use analgesics. We weighed accuracy and interpretability among state-of-the-art methods, and constructed a non-linear model by the Gradient Boosting Decision Tree technique. The AUC of the prediction model was 0.96. Another task was to select informative risk features (RFs) by our proposed model. We sought to identify risk features in literature from the biomedical. Most of the selected RFs are validated by the medical literature and some new RFs could attract the interest across the medical research. The performance of the proposed model, showed its superiority compared with well-known machine learning algorithms in terms of AUC. Liangliang Liu 0001, Jianxin Wang 0001, Min Li 0007, Fang-Xiang Wu, Hong-Dong Li, Zhihui Fei |
BIBM | 5 |
| 2011 | Recipe for Uncovering Predictive Genes Using Support Vector Machines Based on Model Population AnalysisabstractSelecting a small number of informative genes for microarray-based tumor classification is central to cancer prediction and treatment. Based on model population analysis, here we present a new approach, called Margin Influence Analysis (MIA), designed to work with support vector machines (SVM) for selecting informative genes. The rationale for performing margin influence analysis lies in the fact that the margin of support vector machines is an important factor which underlies the generalization performance of SVM models. Briefly, MIA could reveal genes which have statistically significant influence on the margin by using Mann-Whitney U test. The reason for using the Mann-Whitney U test rather than two-sample t test is that Mann-Whitney U test is a nonparametric test method without any distribution-related assumptions and is also a robust method. Using two publicly available cancerous microarray data sets, it is demonstrated that MIA could typically select a small number of margin-influencing genes and further achieves comparable classification accuracy compared to those reported in the literature. The distinguished features and outstanding performance may make MIA a good alternative for gene selection of high dimensional microarray data. (The source code in MATLAB with GNU General Public License Version 2.0 is freely available at http://code.google.com/p/mia2009/). Hong-Dong Li, Yi-Zeng Liang, Qingsong Xu 0003, Dong-Sheng Cao 0001, Bin-Bin Tan, Bai-Chuan Deng, Chen-Chen Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |