Yusen Zhang 0002

dblp:38/10863-2 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0003-3842-1153ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 MODCAN: driver gene identification based on multi-omics features and differential co-association networks for tumor subtypes
abstract
BACKGROUND: Despite the identification of some pan-cancer driver genes through international collaborative initiatives, the discovery of key cancer driver genes remains a formidable challenge. This limitation continues to hinder progress in critical areas, such as early diagnosis, prognostic evaluation, and precision medicine. Consequently, the accurate identification of key driver genes for specific cancer types has become a central focus of bioinformatics research. RESULTS: We present MODCAN, a novel semi-supervised algorithm based on multi-omics features and differential co-association networks. MODCAN facilitates the meaningful stratification of tumor samples into distinct subtypes, enabling a comprehensive exploration of multi-omics features that reflect the inherent heterogeneity of tumors. By constructing differential co-association networks, MODCAN reveals the unique genetic interactions characteristic of each subtype, thereby facilitating the complementary integration of information. CONCLUSIONS: When applied to ten cancer datasets from TCGA, MODCAN significantly outperforms both existing supervised and unsupervised learning algorithms, exhibiting superior performance in terms of precision, recall, and AUPR. Furthermore, MODCAN demonstrates substantial advantages in predicting potential tumor-specific driver genes. Notably, these genes not only exhibit strong specificity for their respective cancers, but also reveal tumor heterogeneity across distinct subtypes.
Ponian Li, Guodong Xiao, Haihui Wang, Chunrui Xu, Yusen Zhang 0002
BMC Bioinform.5
2025 Multi-scale cancer driver gene prediction by flexible data selection and network topology guidance
abstract
OBJECTIVE: Efficient and comprehensive prioritization of cancer driver genes across individual patients, cancer cohorts, and pan-cancer is crucial for advancing cancer diagnosis and treatment. The existing methods are effective, but they seem to have reached a plateau in accuracy enhancement and lack broad-scale joint analysis, flexibility in adapting to cancer and interpretability. METHODS: Here, we introduce GenMorw, a heterogeneous network framework that discovers a novel association score between patients and their mutated genes, enabling the estimation of the likelihood of the mutated genes acting as drivers in patients. GenMorw flexibly integrates or fully utilize collected mutation, gene/miRNA expression, methylation data and PPI networks to classify patient groups based on data-specific characteristics and identify potential drivers at the individual, cancer and pan-cancer levels. RESULTS: GenMorw outperforms existing algorithms with an average cohort AUC improvement of 17.66% and higher overall accuracy by a cumulative ranking strategy in patient-gene heterogeneous networks. Except for AUC evaluation, other various comparative strategies consistently demonstrate the superior performance of GenMorw across multiple cancers, outperforming other algorithms. Some uniquely predicted genes, such as ANK3, CENPF, and COL7A1, which are absent from standard databases and not identified by other methods, were validated as highly cancer-related through literature review and survival analysis. Based on GenMorw-derived heterogeneous networks, the strongly connected components and cliques, which are extracted from them, capture most of the predicted or known driver genes to help predict driver genes. CONCLUSION: We conclude that GenMorw, with its novel gene-patient score mechanism, offers a significant advance in cancer driver gene discovery by capturing both population-wide and patient-specific network signals, thereby improving predictive power and enabling deeper insights into cancer heterogeneity.
Jian Liu 0039, Yingzan Ren, Guodong Xiao, Ponian Li, Chuanqi Sun, Fubin Ma, Rui Gao 0006, Haiyan Cong, Yusen Zhang 0002
J. Biomed. Informatics12
2025 Diversity and consistency graph learning guided multi-view unsupervised feature selection
Hanxiao Xu, Yusen Zhang 0002
Knowl. Based Syst.3
2024 SCSMD: Single Cell Consistent Clustering based on Spectral Matrix Decomposition
abstract
Cluster analysis, a pivotal step in single-cell sequencing data analysis, presents substantial opportunities to effectively unveil the molecular mechanisms underlying cellular heterogeneity and intercellular phenotypic variations. However, the inherent imperfections arise as different clustering algorithms yield diverse estimates of cluster numbers and cluster assignments. This study introduces Single Cell Consistent Clustering based on Spectral Matrix Decomposition (SCSMD), a comprehensive clustering approach that integrates the strengths of multiple methods to determine the optimal clustering scheme. Testing the performance of SCSMD across different distances and employing the bespoke evaluation metric, the methodological selection undergoes validation to ensure the optimal efficacy of the SCSMD. A consistent clustering test is conducted on 15 authentic scRNA-seq datasets. The application of SCSMD to human embryonic stem cell scRNA-seq data successfully identifies known cell types and delineates their developmental trajectories. Similarly, when applied to glioblastoma cells, SCSMD accurately detects pre-existing cell types and provides finer sub-division within one of the original clusters. The results affirm the robust performance of our SCSMD method in terms of both the number of clusters and cluster assignments. Moreover, we have broadened the application scope of SCSMD to encompass larger datasets, thereby furnishing additional evidence of its superiority. The findings suggest that SCSMD is poised for application to additional scRNA-seq datasets and for further downstream analyses.
Ran Jia, Ying-Zan Ren, Ponian Li, Rui Gao 0006, Yusen Zhang 0002
Briefings Bioinform.5
2024 A novel hypergraph model for identifying and prioritizing personalized drivers in cancer
abstract
Cancer development is driven by an accumulation of a small number of driver genetic mutations that confer the selective growth advantage to the cell, while most passenger mutations do not contribute to tumor progression. The identification of these driver genes responsible for tumorigenesis is a crucial step in designing effective cancer treatments. Although many computational methods have been developed with this purpose, the majority of existing methods solely provided a single driver gene list for the entire cohort of patients, ignoring the high heterogeneity of driver events across patients. It remains challenging to identify the personalized driver genes. Here, we propose a novel method (PDRWH), which aims to prioritize the mutated genes of a single patient based on their impact on the abnormal expression of downstream genes across a group of patients who share the co-mutation genes and similar gene expression profiles. The wide experimental results on 16 cancer datasets from TCGA showed that PDRWH excels in identifying known general driver genes and tumor-specific drivers. In the comparative testing across five cancer types, PDRWH outperformed existing individual-level methods as well as cohort-level methods. Our results also demonstrated that PDRWH could identify both common and rare drivers. The personalized driver profiles could improve tumor stratification, providing new insights into understanding tumor heterogeneity and taking a further step toward personalized treatment. We also validated one of our predicted novel personalized driver genes on tumor cell proliferation by vitro cell-based assays, the promoting effect of the high expression of Low-density lipoprotein receptor-related protein 1 (LRP1) on tumor cell proliferation.
Naiqian Zhang, Fubin Ma, Yuxuan Pang, Chenye Wang, Yusen Zhang 0002, Xiaoqi Zheng
PLoS Comput. Biol.6
2023 MaxCLK: discovery of cancer driver genes via maximal clique and information entropy of modules
abstract
MOTIVATION: Cancer is caused by the accumulation of somatic mutations in multiple pathways, in which driver mutations are typically of the properties of high coverage and high exclusivity in patients. Identifying cancer driver genes has a pivotal role in understanding the mechanisms of oncogenesis and treatment. RESULTS: Here, we introduced MaxCLK, an algorithm for identifying cancer driver genes, which was developed by an integrated analysis of somatic mutation data and protein-protein interaction (PPI) networks and further improved by an information entropy index. Tested on pancancer and single cancers, MaxCLK outperformed other existing methods with higher accuracy. About pancancer, we predicted 154 driver genes and 787 driver modules. The analysis of co-occurrence and exclusivity between modules and pathways reveals the correlation of their combinations. Overall, our study has deepened the understanding of driver mechanism in PPI topology and found novel driver genes. AVAILABILITY AND IMPLEMENTATION: The source codes for MaxCLK are freely available at https://github.com/ShandongUniversityMasterMa/MaxCLK-main.
Jian Liu 0039, Fubin Ma, Yongdi Zhu, Naiqian Zhang, Lingming Kong, Haiyan Cong, Rui Gao 0006, Yusen Zhang 0002
Bioinform.10
2023 A Tensor Method Based on Enhanced Tensor Nuclear Norm and Hypergraph Laplacian Regularization for Pan-Cancer Omics Data Analysis
abstract
As a powerful data representation technique, tensor robust principal component analysis (TRPCA) has been widely used for clustering and feature selection tasks. However, it ignores the significant difference in singular values of tensor data and the manifold information contained in different views, thereby causing serious degradation of conventional TRPCA performance. In this paper, a novel tensor method based on enhanced tensor nuclear norm and hypergraph Laplacian regularization (ETHLR) is developed to address the above problem. ETHLR can jointly learn the prior knowledge of singular values and high-order manifold structures in the unified tensor space and the view-specific feature spaces, respectively. Specifically, the enhanced tensor nuclear norm, namely, the weighted tensor Schatten p-norm, is used to shrink the singular values by fully considering the salient difference information of singular values and the complementary information embedded in the tensor space; the hypergraph Laplacian constraint helps encode high-order geometric structures among multiple samples in the nonlinear view-specific feature space. Furthermore, we employ inexact augmented Lagrange multipliers (ALM) to optimize the ETHLR method. Numerous experiments on pan-cancer omics data show that the superiority of ETHLR over several state-of-the-art competitors.
Na Yu 0004, Yusen Zhang 0002, Rui Gao 0006
IEEE J. Biomed. Health Informatics2
2022 DriverRWH: discovering cancer driver genes by random walk on a gene mutation hypergraph
abstract
BACKGROUND: Recent advances in next-generation sequencing technologies have helped investigators generate massive amounts of cancer genomic data. A critical challenge in cancer genomics is identification of a few cancer driver genes whose mutations cause tumor growth. However, the majority of existing computational approaches underuse the co-occurrence mutation information of the individuals, which are deemed to be important in tumorigenesis and tumor progression, resulting in high rate of false positive. RESULTS: To make full use of co-mutation information, we present a random walk algorithm referred to as DriverRWH on a weighted gene mutation hypergraph model, using somatic mutation data and molecular interaction network data to prioritize candidate driver genes. Applied to tumor samples of different cancer types from The Cancer Genome Atlas, DriverRWH shows significantly better performance than state-of-art prioritization methods in terms of the area under the curve scores and the cumulative number of known driver genes recovered in top-ranked candidate genes. Besides, DriverRWH discovers several potential drivers, which are enriched in cancer-related pathways. DriverRWH recovers approximately 50% known driver genes in the top 30 ranked candidate genes for more than half of the cancer types. In addition, DriverRWH is also highly robust to perturbations in the mutation data and gene functional network data. CONCLUSION: DriverRWH is effective among various cancer types in prioritizes cancer driver genes and provides considerable improvement over other tools with a better balance of precision and sensitivity. It can be a useful tool for detecting potential driver genes and facilitate targeted cancer therapies.
Chenye Wang, Junhan Shi, Jiansheng Cai, Yusen Zhang 0002, Xiaoqi Zheng, Naiqian Zhang
BMC Bioinform.4
2021 FS-GBDT: identification multicancer-risk module via a feature selection algorithm by integrating Fisher score and GBDT
abstract
Cancer is a highly heterogeneous disease caused by dysregulation in different cell types and tissues. However, different cancers may share common mechanisms. It is critical to identify decisive genes involved in the development and progression of cancer, and joint analysis of multiple cancers may help to discover overlapping mechanisms among different cancers. In this study, we proposed a fusion feature selection framework attributed to ensemble method named Fisher score and Gradient Boosting Decision Tree (FS-GBDT) to select robust and decisive feature genes in high-dimensional gene expression datasets. Joint analysis of 11 human cancers types was conducted to explore the key feature genes subset of cancer. To verify the efficacy of FS-GBDT, we compared it with four other common feature selection algorithms by Support Vector Machine (SVM) classifier. The algorithm achieved highest indicators, outperforms other four methods. In addition, we performed gene ontology analysis and literature validation of the key gene subset, and this subset were classified into several functional modules. Functional modules can be used as markers of disease to replace single gene which is difficult to be found repeatedly in applications of gene chip, and to study the core mechanisms of cancer.
Da Xu 0005, Kaijing Hao, Yusen Zhang 0002, Wei Chen 0039, Jiaguo Liu, Rui Gao 0006, Chuanyan Wu, Yang De Marinis
Briefings Bioinform.4
2019 PTPD: predicting therapeutic peptides by deep learning and word2vec
abstract
*: Background In the search for therapeutic peptides for disease treatments, many efforts have been made to identify various functional peptides from large numbers of peptide sequence databases. In this paper, we propose an effective computational model that uses deep learning and word2vec to predict therapeutic peptides (PTPD). *: Results Representation vectors of all k-mers were obtained through word2vec based on k-mer co-existence information. The original peptide sequences were then divided into k-mers using the windowing method. The peptide sequences were mapped to the input layer by the embedding vector obtained by word2vec. Three types of filters in the convolutional layers, as well as dropout and max-pooling operations, were applied to construct feature maps. These feature maps were concatenated into a fully connected dense layer, and rectified linear units (ReLU) and dropout operations were included to avoid over-fitting of PTPD. The classification probabilities were generated by a sigmoid function. PTPD was then validated using two datasets: an independent anticancer peptide dataset and a virulent protein dataset, on which it achieved accuracies of 96% and 94%, respectively. *: Conclusions PTPD identified novel therapeutic peptides efficiently, and it is suitable for application as a useful tool in therapeutic peptide design.
Chuanyan Wu, Rui Gao 0006, Yusen Zhang 0002, Yang De Marinis
BMC Bioinform.3
2019 Towards detecting structural branching and cyclicity in graphs: A polynomial-based approach
Matthias Dehmer, Zengqiang Chen 0001, Frank Emmert-Streib, Abbe Mowshowitz, Yongtang Shi, Shailesh Tripathi, Yusen Zhang 0002
Inf. Sci.7
2018 Feature selection of gene expression data for Cancer classification using double RBF-kernels
abstract
BACKGROUND: Using knowledge-based interpretation to analyze omics data can not only obtain essential information regarding various biological processes, but also reflect the current physiological status of cells and tissue. The major challenge to analyze gene expression data, with a large number of genes and small samples, is to extract disease-related information from a massive amount of redundant data and noise. Gene selection, eliminating redundant and irrelevant genes, has been a key step to address this problem. RESULTS: The modified method was tested on four benchmark datasets with either two-class phenotypes or multiclass phenotypes, outperforming previous methods, with relatively higher accuracy, true positive rate, false positive rate and reduced runtime. CONCLUSIONS: This paper proposes an effective feature selection method, combining double RBF-kernels with weighted analysis, to extract feature genes from gene expression data, by exploring its nonlinear mapping ability.
Shenghui Liu, Chunrui Xu, Yusen Zhang 0002, Jiaguo Liu, Bin Yu 0007, Xiaoping Liu 0002, Matthias Dehmer
BMC Bioinform.3
2017 Computational prediction of therapeutic peptides based on graph index
Chunrui Xu, Li Ge, Yusen Zhang 0002, Matthias Dehmer, Ivan Gutman
J. Biomed. Informatics3
2016 A kernel-based clustering method for gene selection with gene expression data
Huihui Chen, Yusen Zhang 0002, Ivan Gutman
J. Biomed. Informatics2