VLDB 2026 Research / reviewers in the wild / expert
Xingyi Li 0003
dblp:180/6887-3
· DBLP profile ↗
19ranked-venue papers
14as first author
14since 2021 · last 2026
0000-0002-6004-4174ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 14 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModalabstractMOTIVATION: Advances in spatial transcriptomics (ST) technologies have made it possible to jointly acquire gene expression and histological image information while preserving spatial coordinates. This breakthrough presents unprecedented opportunities for the precise dissection of spatial heterogeneity in complex tissues. However, existing computational methods remain limited in their capacity for effective integration and synergistic modelling of multimodal ST data. RESULTS: We propose SpatialModal, a multimodal graph learning framework that learns robust joint representations by combining a hierarchical representation strategy with a dual-level contrastive learning mechanism. We perform extensive validation of SpatialModal across diverse ST datasets spanning human and mouse tissues. The results demonstrate that SpatialModal effectively reveals intricate brain architectures in humans and mice, dissects tumour microenvironment heterogeneity in breast cancer, delineates Alzheimer's disease patterns, and characterizes spatiotemporal developmental trajectories within the embryonic heart, underscoring its capability to decipher the spatial heterogeneity of biological tissues. Furthermore, SpatialModal exhibits remarkable versatility and robustness, maintaining superior efficacy even on unimodal datasets devoid of histological images, thereby ensuring its broad applicability across diverse ST platforms. AVAILABILITY AND IMPLEMENTATION: SpatialModal is implemented in Python and is freely available at https://github.com/xingyili/SpatialModal. The source code used in this study has been archived on Zenodo at DOI: https://doi.org/10.5281/zenodo.21264356. All datasets used in this study are publicly available at https://doi.org/10.5281/zenodo.18220735. Xingyi Li 0003, Dongmin Zhao, Xiangting Jia, Gaoyuan Du, Jialuo Xu, Yingfu Wu, Junnan Zhu, Xuequn Shang 0001 |
Bioinform. | 1 |
| 2026 | WEmarker: Breast Cancer-Specific Prognostic Analysis With Weighted Multiplex Network EmbeddingabstractIn clinical trials, prognostic biomarkers have become essential for guiding treatment decisions after breast cancer surgery. Network-based methods have gained notable attention to reveal marker genes, but many existing methods only focus on a single network, which inevitably neglects the incompleteness of interaction relationships within the network. Even when based upon the multiplex network, most of methods directly integrate the multiplex network into an aggregated network and do not take into account the inherent noise in the biological networks, which can not preserve the topological structure of each original network very well. In this study, we propose a novel method, WEmarker, for breast cancer-specific prognostic analysis. WEmarker reduces the noise level of biological networks and quantifies the probability of interactions between genes, and represents the nodes in the weighted multiplex network as vectors while efficiently retaining the structure information of these networks for identifying prognostic biomarkers. The results show that WEmarker outperforms comparative methods and the case study also demonstrates that biomarkers identified by WEmarker have reliable biological interpretability for breast cancer prognosis. Xingyi Li 0003, Zhelin Zhao, Huihui Kong, Min Li 0007, Xuequn Shang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Multi-Task Learning Framework for Cancer Driver Gene Identification on Multi-Network and Multi-OMICS DataabstractCancer driver genes play an essential role in understanding cancer oncogenesis, tumor progression, and thera-peutic development. The integration of multi-omics data with biological networks has enabled the application of graph deep learning techniques for identifying cancer driver genes. However, most existing methods only use a single biological network as input, inevitably introducing the incompleteness and noise of interactions into models. To address these limitations, we propose MTCDG, a multi-task learning framework for cancer driver gene identification on multi-network and multi-omics data, which can not only enhance the interaction completeness but also enable more comprehensive extraction of graph topological features. The experimental results show the superior predictive performance of MTCDG over other methods. We anticipate that MTCDG will offer new insights for cancer genomic research and can be potentially extended to other areas of biological research in future research. The code of MTCDG is available on github: https://github.com/xingyili/MTCDG. Jialuo Xu, Xuequn Shang 0001, Xingyi Li 0003 |
BIBM | 6 |
| 2025 | Towards simplified graph neural networks for identifying cancer driver genes in heterophilic networksabstractThe identification of cancer driver genes is crucial for understanding the complex processes involved in cancer development, progression, and therapeutic strategies. Multi-omics data and biological networks provided by numerous databases enable the application of graph deep learning techniques that incorporate network structures into the deep learning framework. However, most existing methods do not account for the heterophily in the biological networks, which hinders the improvement of model performance. Meanwhile, feature confusion often arises in models based on graph neural networks in such graphs. To address this, we propose a Simplified Graph neural network for identifying Cancer Driver genes in heterophilic networks (SGCD), which comprises primarily two components: a graph convolutional neural network with representation separation and a bimodal feature extractor. The results demonstrate that SGCD not only performs exceptionally well but also exhibits robust discriminative capabilities compared to state-of-the-art methods across all benchmark datasets. Moreover, subsequent interpretability experiments on both the model and biological aspects provide compelling evidence supporting the reliability of SGCD. Additionally, the model can dissect gene modules, revealing clearer connections between driver genes in cancers. We are confident that SGCD holds potential in the field of precision oncology and may be applied to prognosticate biomarkers for a wide range of complex diseases. Xingyi Li 0003, Jialuo Xu, Xuequn Shang 0001 |
Briefings Bioinform. | 1 |
| 2025 | Deep graph convolutional network-based multi-omics integration for cancer driver gene identificationabstractCancer driver genes play a pivotal role in understanding cancer development, progression, and therapeutic discovery. The plenty of accumulation of multi-omics data and biological networks provides a data foundation for graph neural network (GNN) frameworks. However, most existing methods directly concatenate multi-omics data as features, which may lead to limited performance. To address this limitation, we propose deepCDG, a deep graph convolutional network (GCN)-based multi-omics integration model for cancer driver gene identification. The model first employs shared-parameter GCN encoders to extract representations from three omics perspectives, followed by feature integration through an attention layer, and finally utilizes a residual-connected GCN predictor for cancer driver gene identification. Additionally, deepCDG employs GNNExplainer for cancer driver gene module identification. Experimental results demonstrate the effective predictive performance, model robustness, and computational efficiency of deepCDG. Additionally, biological interpretability analysis further validates the reliability of the identification of cancer driver genes of our framework, and the identified gene modules provide profound insights into complex inter-gene relationships and interactions. We believe our method offers enhanced applicability for cancer driver gene identification and could be extended to other biological research fields in future studies. Yingzhuo Wu, Jialuo Xu, Xuequn Shang 0001, Xingyi Li 0003 |
Briefings Bioinform. | 6 |
| 2025 | PathActMarker: an R package for inferring pathway activity of complex diseases
Xingyi Li 0003, Zhelin Zhao, Xingyu Liao, Min Li 0007, Xuequn Shang 0001 |
Frontiers Comput. Sci. | 1 |
| 2025 | DualMarker: A Multi-Source Fusion Identification Method for Prognostic Biomarkers of Breast Cancer Based on Dual-Layer Heterogeneous NetworkabstractBreast cancer is a complex disease that arises from multiple factors, including genetics, age, and environmental factors. Prognosis prediction for breast cancer is a challenging task that urgently needs to be addressed. Prognostic biomarkers can aid in predicting clinical outcomes for breast cancer patients, and network-based approaches are frequently employed to identify such biomarkers. However, the accuracy of these approaches based on single source biological network is poor due to incomplete interactions of single biological network. Some network-based approaches that integrate multiple biological networks have not considered network denoising, which may lead to the accuracy of these approaches to be improved. We propose a multi-source fusion identification method named DualMarker for prognostic biomarkers of breast cancer. This method constructs a dual-layer heterogeneous network by integrating multiple biological sources. To decrease the negative effects of incomplete interactions in biological networks, we denoise the constructed network. The ranking of features is obtained by the network propagation algorithm and the initial scoring strategy. Compared with six other network-based methods, DualMarker shows the best performance in six breast cancer datasets. Moreover, we have also demonstrated that the biomarkers identified by DualMarker are of interpretability biologically and closely associated with breast cancer patients' prognosis. Xingyi Li 0003, Gaoyuan Du, Zhelin Zhao, Ju Xiang, Jialu Hu, Xuequn Shang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Scbean: a python library for single-cell multi-omics data analysisabstractSUMMARY: Single-cell multi-omics technologies provide a unique platform for characterizing cell states and reconstructing developmental process by simultaneously quantifying and integrating molecular signatures across various modalities, including genome, transcriptome, epigenome, and other omics layers. However, there is still an urgent unmet need for novel computational tools in this nascent field, which are critical for both effective and efficient interrogation of functionality across different omics modalities. Scbean represents a user-friendly Python library, designed to seamlessly incorporate a diverse array of models for the examination of single-cell data, encompassing both paired and unpaired multi-omics data. The library offers uniform and straightforward interfaces for tasks, such as dimensionality reduction, batch effect elimination, cell label transfer from well-annotated scRNA-seq data to scATAC-seq data, and the identification of spatially variable genes. Moreover, Scbean's models are engineered to harness the computational power of GPU acceleration through Tensorflow, rendering them capable of effortlessly handling datasets comprising millions of cells. AVAILABILITY AND IMPLEMENTATION: Scbean is released on the Python Package Index (PyPI) (https://pypi.org/project/scbean/) and GitHub (https://github.com/jhu99/scbean) under the MIT license. The documentation and example code can be found at https://scbean.readthedocs.io/en/latest/. Haohui Zhang, Bin Lian, Xingyi Li 0003, Tao Wang 0082, Xuequn Shang 0001, Ahmad Aziz, Jialu Hu |
Bioinform. | 5 |
| 2023 | A personalized pathway activation inference method based on pathway structure for classification of inflammatory bowel diseaseabstractInflammatory bowel disease (IBD) is a complex disease that mainly consists of two subtypes, ulcerative colitis (UC) and Crohn's disease (CD). These two diseases exhibit similar clinical symptoms, leading to a potential misdiagnosis of patients. Accurately assessing the disease status and identifying specific biomarkers are important for the diagnosis and treatment of IBD. Pathways play a significant role in the occurrence and progression of complex diseases, involving the abnormal functionality or regulatory imbalance. Although many methods integrating pathway information have been proposed to evaluate pathway activity, but these methods rarely take into account the topology of pathway networks. Some algorithms based on pathway structure do not de-noise the pathway networks and characterize the disease-specific state of a single sample from the perspective of pathways. In this study, we present a personalized pathway activation inference method (PPA-PS) based on the pathway structure, which utilizes the topology of pathways to evaluate the importance of nodes and quantify the degree of edge disturbance caused by a single disease sample. The results demonstrate that PPA-PS outperforms the compared approaches in terms of classification performance and robustness, indicating its potential as a valuable tool for the pathway biomarker identification and precise diagnosis of IBD. Xingyi Li 0003, Xuequn Shang 0001, Zhelin Zhao, Chenzhuo Yan |
BIBM | 1 |
| 2022 | A multi-source fusion method to identify biomarkers for breast cancer prognosis based on dual-layer heterogeneous networkabstractThe prognosis of breast cancer is challenging, which is an urgent problem to be solved. The prognostic biomarkers for breast cancer can help us predict the clinical outcomes of patients, and network-based methods are widely introduced to find prognostic biomarkers. According to the difference of input biological data, existing network-based biomarker prediction methods are mainly classified into two types: integrating single-source network or multi-source networks. However, the interactome of single-source network remains incomplete, and biological networks are noisy, which will hamper the network-based identification accuracy of biomarkers. In this study, we introduce a multi-source fusion method, DualMarker, which integrates multiple biological information sources and constructs a dual-layer heterogeneous network by fast network embedding. Next, we introduce a network enhancement method to denoise the constructed dual-layer heterogeneous network, and we implement network propagation algorithm on the constructed dual-layer heterogeneous network to rank the features. After comparing with competitive methods, we find that DualMarker substantially outperforms these methods. In addition, we verify that the biomarkers identified by DualMarker are closely related to the prognosis of breast cancer patients. Xingyi Li 0003, Zhelin Zhao, Ju Xiang, Jialu Hu, Xuequn Shang 0001 |
BIBM | 1 |
| 2022 | SEPA: signaling entropy-based algorithm to evaluate personalized pathway activation for survival analysis on pan-cancer dataabstractMOTIVATION: Biomarkers with prognostic ability and biological interpretability can be used to support decision-making in the survival analysis. Genes usually form functional modules to play synergistic roles, such as pathways. Predicting significant features from the functional level can effectively reduce the adverse effects of heterogeneity and obtain more reproducible and interpretable biomarkers. Personalized pathway activation inference can quantify the dysregulation of essential pathways involved in the initiation and progression of cancers, and can contribute to the development of personalized medical treatments. RESULTS: In this study, we propose a novel method to evaluate personalized pathway activation based on signaling entropy for survival analysis (SEPA), which is a new attempt to introduce the information-theoretic entropy in generating pathway representation for each patient. SEPA effectively integrates pathway-level information into gene expression data, converting the high-dimensional gene expression data into the low-dimensional biological pathway activation scores. SEPA shows its classification power on the prognostic pan-cancer genomic data, and the potential pathway markers identified based on SEPA have statistical significance in the discrimination of high- and low-risk cohorts and are likely to be associated with the initiation and progress of cancers. The results show that SEPA scores can be used as an indicator to precisely distinguish cancer patients with different clinical outcomes, and identify important pathway features with strong discriminative power and biological interpretability. AVAILABILITY AND IMPLEMENTATION: The MATLAB-package for SEPA is freely available from https://github.com/xingyili/SEPA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xingyi Li 0003, Min Li 0007, Ju Xiang, Zhelin Zhao, Xuequn Shang 0001 |
Bioinform. | 1 |
| 2022 | A Dual Ranking Algorithm Based on the Multiplex Network for Heterogeneous Complex Disease AnalysisabstractIdentifying biomarkers of heterogeneous complex diseases has always been one of the focuses in medical research. In previous studies, the powerful network propagation methods have been applied to finding marker genes related to specific diseases, but existing methods are mostly based on a single network, which may be greatly affected by the incompleteness of the network and the ignorance of a large amount of information about physical and functional interactions between biological components. Other methods that directly integrate multiple types of interactions into an aggregate network have the risks that different types of data may conflict with each other and the characteristics and topologies of each individual network are lost. Meanwhile, biomarkers used in clinical trials should have the characteristics of small quantity and strong discriminate ability. In this study, we developed a multiplex network-based dual ranking framework (DualRank) for heterogeneous complex disease analysis. We applied the proposed method to heterogeneous complex diseases for diagnosis, prognosis, and classification. The results showed that DualRank outperformed competing methods and could identify biomarkers with the small quantity, great prediction performance (average AUC = 0.818) and biological interpretability. Xingyi Li 0003, Ju Xiang, Fang-Xiang Wu, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | NIDM: network impulsive dynamics on multiplex biological network for disease-gene predictionabstractThe prediction of genes related to diseases is important to the study of the diseases due to high cost and time consumption of biological experiments. Network propagation is a popular strategy for disease-gene prediction. However, existing methods focus on the stable solution of dynamics while ignoring the useful information hidden in the dynamical process, and it is still a challenge to make use of multiple types of physical/functional relationships between proteins/genes to effectively predict disease-related genes. Therefore, we proposed a framework of network impulsive dynamics on multiplex biological network (NIDM) to predict disease-related genes, along with four variants of NIDM models and four kinds of impulsive dynamical signatures (IDSs). NIDM is to identify disease-related genes by mining the dynamical responses of nodes to impulsive signals being exerted at specific nodes. By a series of experimental evaluations in various types of biological networks, we confirmed the advantage of multiplex network and the important roles of functional associations in disease-gene prediction, demonstrated superior performance of NIDM compared with four types of network-based algorithms and then gave the effective recommendations of NIDM models and IDS signatures. To facilitate the prioritization and analysis of (candidate) genes associated to specific diseases, we developed a user-friendly web server, which provides three kinds of filtering patterns for genes, network visualization, enrichment analysis and a wealth of external links (http://bioinformatics.csu.edu.cn/DGP/NID.jsp). NIDM is a protocol for disease-gene prediction integrating different types of biological networks, which may become a very useful computational tool for the study of disease-related genes. Ju Xiang, Jiashuai Zhang, Ruiqing Zheng, Xingyi Li 0003, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2021 | FUNMarker: Fusion Network-Based Method to Identify Prognostic and Heterogeneous Breast Cancer BiomarkersabstractBreast cancer is a heterogeneous disease with many clinically distinguishable molecular subtypes each corresponding to a cluster of patients. Identification of prognostic and heterogeneous biomarkers for breast cancer is to detect cluster-specific gene biomarkers which can be used for accurate survival prediction of breast cancer outcomes. In this study, we proposed a FUsion Network-based method (FUNMarker) to identify prognostic and heterogeneous breast cancer biomarkers by considering the heterogeneity of patient samples and biological information from multiple sources. To reduce the affect of heterogeneity of patients, samples were first clustered using the K-means algorithm based on the principal components of gene expression. For each cluster, to comprehensively evaluate the influence of genes on breast cancer, genes were weighted from three aspects: biological function, prognostic ability and correlation with known disease genes. Then they were ranked via a label propagation model on a fusion network that combined physical protein interactions from seven types of networks and thus could reduce the impact of incompleteness of interactome. We compared FUNMarker with three state-of-the-art methods and the results showed that biomarkers identified by FUNMarker were biological interpretable and had stronger discriminative power than the existing methods in differentiating patients with different prognostic outcomes. Xingyi Li 0003, Ju Xiang, Jianxin Wang 0001, Jinyan Li 0001, Fang-Xiang Wu, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | A topological AUC-based biomarker ensemble method for the complex disease analysisabstractComplex diseases are affected by many factors, and their pathogenic mechanism is complicated, which brings difficulties to the analysis and treatment of diseases. AUC, the area under the ROC curve, is often used as a gold standard to evaluate the performance of a binary classifier. The existing methods of constructing classifier by optimizing AUC are easy to fall into local optimum, and have high time complexity, which is not suitable for real-time analysis of high-dimensional gene expression data. With the rapid development of high-throughput sequencing technology, feature selection and model estimation become the necessary means to reduce the dimension and complexity of data, and the selected important features have the potential as biomarkers to reveal the pathogenesis of diseases. In this paper, we proposed a topological AUC-based biomarker ensemble method for the complex disease analysis, which uses gene expression data and the topological information derived from the protein-protein interaction network to identify biomarkers. The main contribution is to optimize two objectives simultaneously: maximizing the AUC score and minimizing the number of selected features. We applied the proposed method to analyze two types of problems: 1) prognosis of breast cancer, 2) classification of similar diseases. The results show that our method can effectively identify a small set of biomarkers with the powerful classification ability and the biological interpretability. Xingyi Li 0003, Ju Xiang, Fang-Xiang Wu, Min Li 0007 |
BIBM | 1 |
| 2020 | Network-based methods for predicting essential genes or proteins: a surveyabstractGenes that are thought to be critical for the survival of organisms or cells are called essential genes. The prediction of essential genes and their products (essential proteins) is of great value in exploring the mechanism of complex diseases, the study of the minimal required genome for living cells and the development of new drug targets. As laboratory methods are often complicated, costly and time-consuming, a great many of computational methods have been proposed to identify essential genes/proteins from the perspective of the network level with the in-depth understanding of network biology and the rapid development of biotechnologies. Through analyzing the topological characteristics of essential genes/proteins in protein-protein interaction networks (PINs), integrating biological information and considering the dynamic features of PINs, network-based methods have been proved to be effective in the identification of essential genes/proteins. In this paper, we survey the advanced methods for network-based prediction of essential genes/proteins and present the challenges and directions for future research. Xingyi Li 0003, Min Zeng 0004, Ruiqing Zheng, Min Li 0007 |
Briefings Bioinform. | 1 |
| 2019 | Classification of Schizophrenia by Iterative Random Forest Feature Selection Based on DNA Methylation Array DataabstractChanges in DNA methylation are widely thought to be involved in the evolution of the disease, and most studies suggest that whole genome hypo-methylation levels are widespread in patients with schizophrenia. Since the exploration of DNA methylation changes in the etiology and pathogenesis of schizophrenia will be important for the prevention and early intervention, it is crucial to identify differentially methylated sites with high specificity and sensitivity for disease classification. In this study, we present a comprehensive approach MethIRF for the DNA methylation-based classification of schizophrenia by iterative random forest feature selection. The results show that MethIRF has a powerful discrimination ability compared to other four criteria methods for detecting differentially methylation sites. Moreover, the proposed method can discover significant CpG sites associated with schizophrenia and explore changes in the biological mechanisms of diseases. Min Li 0007, Linconghua Wang, Xingyi Li 0003, Fang-Xiang Wu, Jianxin Wang 0001 |
BIBM | 4 |
| 2019 | DualRank: multiplex network-based dual ranking for heterogeneous complex disease analysisabstractAnalysis of heterogeneous complex diseases based on the expression of biomarkers has always been the focus of medical research. In the past studies, the powerful network propagation has been applied in finding marker genes related to specific diseases. However, the network propagation model largely depends on the reliability and integrity of the network data, current networks may cause some problems due to the incompleteness of the networks. In this study, we developed a multiplex network-based dual ranking framework (DualRank) for heterogeneous complex disease analysis. We applied the proposed method to heterogeneous complex diseases for disease diagnosis, cancer prognosis, and similar disease classification. The results showed that DualRank outperformed current methods and could identify biomarkers with small quantity, strong prediction accuracy and biological interpretability. Xingyi Li 0003, Ju Xiang, Fang-Xiang Wu, Min Li 0007 |
BIBM | 1 |
| 2019 | Identification of Prognostic and Heterogeneous Breast Cancer Biomarkers Based on Fusion Network and Multiple Scoring Strategies
Xingyi Li 0003, Ju Xiang, Jianxin Wang 0001, Fang-Xiang Wu, Min Li 0007 |
ICIC (2) | 1 |