VLDB 2026 Research / reviewers in the wild / expert
Hao Wu 0062
dblp:72/4250-62
· DBLP profile ↗
19ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0003-2340-9258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MomicPred: A Cell Cycle Prediction Framework Based on Dual-Branch Multi-Modal Feature Fusion for Single-Cell Multi-Omics DataabstractThe cell cycle plays a pivotal role in regulating cell fate and stem cell differentiation. As a rate-limiting step in differentiation, its precise regulation is essential for maintaining cellular diversity and tissue homeostasis. Recent advances in single-cell multi-omics technologies have enabled the integration of gene expression data and chromatin structural regulation, thereby enhancing the prediction of cell cycle using multi-omics approaches. However, current algorithms have yet to effectively integrate transcriptome and three-dimensional (3D) genomic data for cell cycle prediction. We propose MomicPred, an innovative dual-branch multi-modal fusion framework designed to predict cell cycle dynamics. This framework integrates transcriptome-derived gene expression data with global chromatin structural insights from 3D genome data. By leveraging the complementary nature of these multi-omics data, MomicPred extracts three core feature sets that uncover cross-layer associations and synergistic interactions between the two omics modalities, enabling high-precision cell cycle prediction. We further evaluate the framework's performance through various benchmarking strategies, demonstrating its efficiency and robustness. Furthermore, feature importance analysis reveals chromatin structural changes and key biological processes across distinct cell cycle stages, offering new perspectives for future research. Zhenqi Shi, Linxing Cong, Hao Wu 0062 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | EAP-LSTM: A Bi-LSTM-Based Deep Learning Framework for Quantitatively Predicting Enhancer Activity in Drosophila and Human Cell LinesabstractEnhancer activity plays a critical role in gene regulation, influencing various biological processes such as development and disease progression. Accurate prediction of enhancer activity is essential for understanding the mechanisms underlying gene regulation and enhancer function. This study introduces a novel deep learning framework, EAP-LSTM (Enhancer Activity Prediction based on Bi-LSTM), to quantitatively predict enhancer activity across different species and cell lines. The model integrates multiple feature modules, including Word2Vec-based representations of DNA sequences, reverse complement k-mer, mismatch k-mer features, and epigenomic data. Evaluated on six cell lines, including five human cell lines (A549, HCT116, HepG2, K562, and MCF-7) and one Drosophila cell line (S2), EAP-LSTM consistently outperforms state-of-the-art models, such as DeepSTARR and HEAP, in all datasets. For example, on the K562 dataset, EAP-LSTM achieves a Pearson correlation coefficient (PCC) of 0.7944, outperforming DeepSTARR and HEAP by 13.65% and 2.73%, respectively. In addition, EAP-LSTM demonstrates strong performance in small-sample learning scenarios, showing clear improvements compared with baseline models. Furthermore, the study investigates the role of transcription factor binding sites (TFBSs) within enhancer regions, identifying critical motifs associated with enhancer activity. These findings not only improve enhancer prediction accuracy but also provide valuable insights into the molecular mechanisms underlying enhancer function. Lichang Dai, Yu Dou, Xin Li 0137, Chang Lu 0015, Hao Wu 0062 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | A novel deep learning framework with dynamic tokenization for identifying chromatin interactions along with motif importance investigationabstractA comprehensive understanding of chromatin interaction networks is crucial for unraveling the regulatory mechanisms of gene expression. While various computational methods have been developed to predict chromatin interactions and address the limitations and high costs of high-throughput experimental techniques, their performance is often overestimated due to the specificity of chromatin interaction data. In this study, we proposed Inter-Chrom, a novel deep learning model integrating dynamic tokenization, DNABERT's word embedding, and the efficient channel attention mechanism to identify chromatin interactions using sequence and genomic features, leveraging a newly curated dataset. Experimental results demonstrate that Inter-Chrom outperforms existing methods on three cell line datasets. Additionally, we proposed a novel method for calculating motif importance and analyzed the motifs with high importance scores identified through this method, including those that have been extensively studied and others that have received limited attention to date. Inter-Chrom's robustness for input variations and superior ability to leverage sequence features position it as a powerful tool for advancing chromatin interaction research. The source code of Inter-Chrom is freely available at https://github.com/HaoWuLab-Bioinformatics/Inter-Chrom. Liangcan Li, Xin Li 0137, Hao Wu 0062 |
Briefings Bioinform. | 3 |
| 2025 | GRANet: a graph residual attention network for gene regulatory network inferenceabstractThe reconstruction of gene regulatory networks (GRNs) is crucial for uncovering regulatory relationships between genes and understanding the mechanisms of gene expression within cells. With advancements in single-cell RNA sequencing (scRNA-seq) technology, researchers have sought to infer GRNs at the single-cell level. However, existing methods primarily construct global models encompassing entire gene networks. While these approaches aim to capture genome-wide interactions, they frequently suffer from decreased accuracy due to challenges such as network scale, noise interference, and data sparsity. This study proposes GRANet (Graph Residual Attention Network), a novel deep learning framework for inferring GRNs. GRANet leverages residual attention mechanisms to adaptively learn complex gene regulatory relationships while integrating multi-dimensional biological features for a more comprehensive inference process. We evaluated GRANet across multiple datasets, benchmarking its performance against state-of-the-art methods. The experimental results demonstrate that GRANet consistently outperforms existing methods in GRN inference tasks. In addition, in our case study on EGR1, CBFB, and ELF1, GRANet achieved high prediction accuracy, effectively identifying both known and novel regulatory interactions. These findings highlight GRANet's potential to advance research in gene regulation and disease mechanisms. Junliang Zhou, Ningji Gong, Yanjun Hu, Hong Yu 0007, Guoyin Wang 0001, Hao Wu 0062 |
Briefings Bioinform. | 6 |
| 2025 | scHiClassifier: a deep learning framework for cell type prediction by fusing multiple feature sets from single-cell Hi-C dataabstractSingle-cell high-throughput chromosome conformation capture (Hi-C) technology enables capturing chromosomal spatial structure information at the cellular level. However, to effectively investigate changes in chromosomal structure across different cell types, there is a requisite for methods that can identify cell types utilizing single-cell Hi-C data. Current frameworks for cell type prediction based on single-cell Hi-C data are limited, often struggling with features interpretability and biological significance, and lacking convincing and robust classification performance validation. In this study, we propose four new feature sets based on the contact matrix with clear interpretability and biological significance. Furthermore, we develop a novel deep learning framework named scHiClassifier based on multi-head self-attention encoder, 1D convolution and feature fusion, which integrates information from these four feature sets to predict cell types accurately. Through comprehensive comparison experiments with benchmark frameworks on six datasets, we demonstrate the superior classification performance and the universality of the scHiClassifier framework. We further assess the robustness of scHiClassifier through data perturbation experiments and data dropout experiments. Moreover, we demonstrate that using all feature sets in the scHiClassifier framework yields optimal performance, supported by comparisons of different feature set combinations. The effectiveness and the superiority of the multiple feature set extraction are proven by comparison with four unsupervised dimensionality reduction methods. Additionally, we analyze the importance of different feature sets and chromosomes using the "SHapley Additive exPlanations" method. Furthermore, the accuracy and reliability of the scHiClassifier framework in cell classification for single-cell Hi-C data are supported through enrichment analysis. The source code of scHiClassifier is freely available at https://github.com/HaoWuLab-Bioinformatics/scHiClassifier. Xiangfei Zhou, Hao Wu 0062 |
Briefings Bioinform. | 2 |
| 2025 | A Rule-Guided Community Detection Method for Identifying Subpopulations in Medical DataabstractPrecisely identifying and explaining subpopulations in heterogeneous populations is essential to understanding the disease subtype. Using community detection to identify subpopulations is a promising way. However, there remains an issue in the existing community detection: Current methods for identifying subpopulations in medical data rely solely on separate attribute values, ignoring the important association rules between attribute values. Association rules are crucial in medical diagnosis to determine disease subtypes. Thus, We propose a rule-guided community detection (RGCD) method for precisely identifying homogeneous subpopulations. Specifically, the RGCD incorporates association rules into the original network, thereby constructing an augmented network. It proves that decomposing the embedding vectors obtained from biased random walks on the augmented network is equivalent to decomposing the transition probability matrix. Based on this proof, we enhance the transition probability matrix through rule-guided biased random walks, resulting in the rule-augmented matrix. By performing matrix decomposition and clustering on this matrix, we achieve precise identification of subpopulations. To the best of our knowledge, this is the first work that introduces the incorporation of association rules into community detection. Extensive experiments on 10 real-world datasets from medical fields fully show that the RGCD is more competitive than six state-of-the-art community detection methods. The weighted F1 of RGCD increases by up to 22.62%, compared to the best existing community detection methods. Furthermore, We provide a qualitative depiction of the subpopulations obtained through RGCD and acquire medically significant insights. Hong Yu 0007, Hao Wu 0062, Guoyin Wang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Enhancer-MDLF: a novel deep learning framework for identifying cell-specific enhancersabstractEnhancers, noncoding DNA fragments, play a pivotal role in gene regulation, facilitating gene transcription. Identifying enhancers is crucial for understanding genomic regulatory mechanisms, pinpointing key elements and investigating networks governing gene expression and disease-related mechanisms. Existing enhancer identification methods exhibit limitations, prompting the development of our novel multi-input deep learning framework, termed Enhancer-MDLF. Experimental results illustrate that Enhancer-MDLF outperforms the previous method, Enhancer-IF, across eight distinct human cell lines and exhibits superior performance on generic enhancer datasets and enhancer-promoter datasets, affirming the robustness of Enhancer-MDLF. Additionally, we introduce transfer learning to provide an effective and potential solution to address the prediction challenges posed by enhancer specificity. Furthermore, we utilize model interpretation to identify transcription factor binding site motifs that may be associated with enhancer regions, with important implications for facilitating the study of enhancer regulatory mechanisms. The source code is openly accessible at https://github.com/HaoWuLab-Bioinformatics/Enhancer-MDLF. Hao Wu 0062 |
Briefings Bioinform. | 3 |
| 2024 | HHGNN: Hyperbolic Hypergraph Convolutional Neural Network based on variational autoencoder
Zhangyu Mei, Xiao Bi, Yating Wen, Xianchun Kong, Hao Wu 0062 |
Neurocomputing | 5 |
| 2024 | lncLocator-imb: An Imbalance-Tolerant Ensemble Deep Learning Framework for Predicting Long Non-Coding RNA Subcellular LocalizationabstractRecent studies have highlighted the critical roles of long non-coding RNAs (lncRNAs) in various biological processes, including but not limited to dosage compensation, epigenetic regulation, cell cycle regulation, and cell differentiation regulation. Consequently, lncRNAs have emerged as a central focus in genetic studies. The identification of the subcellular localization of lncRNAs is essential for gaining insights into crucial information about lncRNA interaction partners, post- or co-transcriptional regulatory modifications, and external stimuli that directly impact the function of lncRNA. Computational methods have emerged as a promising avenue for predicting the subcellular localization of lncRNAs. However, there is a need for additional enhancement in the performance of current methods when dealing with unbalanced data sets. To address this challenge, we propose a novel ensemble deep learning framework, termed lncLocator-imb, for predicting the subcellular localization of lncRNAs. To fully exploit lncRNA sequence information, lncLocator-imb integrates two base classifiers, including convolutional neural networks (CNN) and gated recurrent units (GRU). Additionally, it incorporates two distinct types of features, including the physicochemical pattern feature and the distributed representation of nucleic acids feature. To address the problem of poor performance exhibited by models when confronted with unbalanced data sets, we utilize the label-distribution-aware margin (LDAM) loss function during the training process. Compared with traditional machine learning models and currently available predictors, lncLocator-imb demonstrates more robust category imbalance tolerance. Our study proposes an ensemble deep learning framework for predicting the subcellular localization of lncRNAs. Additionally, a novel approach is presented for the management of different features and the resolution of unbalanced data sets. The proposed framework exhibits the potential to serve as a significant resource for various sequence-based prediction tasks, providing a versatile tool that can be utilized by professionals in the fields of bioinformatics and genetics. Dianguo Li, Hao Wu 0062 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | IChrom-Deep: An Attention-Based Deep Learning Model for Identifying Chromatin InteractionsabstractIdentification of chromatin interactions is crucial for advancing our knowledge of gene regulation. However, due to the limitations of high-throughput experimental techniques, there is an urgent need to develop computational methods for predicting chromatin interactions. In this study, we propose a novel attention-based deep learning model, termed IChrom-Deep, to identify chromatin interactions using sequence features and genomic features. The experimental results based on the datasets of three cell lines demonstrate that the IChrom-Deep achieves satisfactory performance and is superior to the previous methods. We also investigate the effect of DNA sequence and associated features and genomic features on chromatin interactions, and highlight the applicable scenarios of some features, such as sequence conservation and distance. Moreover, we identify a few genomic features that are extremely important across different cell lines, and IChrom-Deep achieves comparable performance with only these significant genomic features versus using all genomic features. It is believed that IChrom-Deep can serve as a useful tool for future studies that seek to identify chromatin interactions. Hao Wu 0062 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | scHiCSC: A Novel Single-Cell Hi-C Clustering Framework by Contact-Weight-Based Smoothing and Feature FusionabstractSingle-cell Hi-C technology is utilized to obtain chromosome interaction information at the single-cell level and further study the differences in genome structures between different cell types. However, there are few accurate and efficient clustering methods for single-cell Hi-C data, with the following manifestations: The clustering efficacy on the dataset with a large number of cells is not very satisfactory, and it is difficult to cluster these cells with small number in the whole dataset. In this study, we propose a high-performance single-cell Hi-C clustering framework, called scHiCSC. A new smoothing method based on contact number weight is first proposed to generate cell embedding with more accurate cell features. In addition, a novel feature fusion method is proposed to further supplement the feature information of cells by fusing the chromosome structure information within cells and the distance information between cells. The experimental results show that scHiCSC has a strong generalization ability on different sizes of datasets and outperforms the existing single-cell Hi-C clustering frameworks. Moreover, scHiCSC achieves an optimal and stable clustering efficacy in the datasets with large-scale cell numbers and can cluster the cells with small number in the whole dataset. The source code of scHiCSC is freely available at https://github.com/HaoWuLab-Bioinformatics/scHiCSC. Xiangfei Zhou, Zhenqi Shi, Yingfu Wu, Hao Wu 0062 |
BIBM | 5 |
| 2022 | scHiCStackL: a stacking ensemble learning-based method for single-cell Hi-C classification using cell embeddingabstractSingle-cell Hi-C data are a common data source for studying the differences in the three-dimensional structure of cell chromosomes. The development of single-cell Hi-C technology makes it possible to obtain batches of single-cell Hi-C data. How to quickly and effectively discriminate cell types has become one hot research field. However, the existing computational methods to predict cell types based on Hi-C data are found to be low in accuracy. Therefore, we propose a high accuracy cell classification algorithm, called scHiCStackL, based on single-cell Hi-C data. In our work, we first improve the existing data preprocessing method for single-cell Hi-C data, which allows the generated cell embedding better to represent cells. Then, we construct a two-layer stacking ensemble model for classifying cells. Experimental results show that the cell embedding generated by our data preprocessing method increases by 0.23, 1.22, 1.46 and 1.61$\%$ comparing with the cell embedding generated by the previously published method scHiCluster, in terms of the Acc, MCC, F1 and Precision confidence intervals, respectively, on the task of classifying human cells in the ML1 and ML3 datasets. When using the two-layer stacking ensemble framework with the cell embedding, scHiCStackL improves by 13.33, 19, 19.27 and 14.5 over the scHiCluster, in terms of the Acc, ARI, NMI and F1 confidence intervals, respectively. In summary, scHiCStackL achieves superior performance in predicting cell types using the single-cell Hi-C data. The webserver and source code of scHiCStackL are freely available at http://hww.sdu.edu.cn:8002/scHiCStackL/ and https://github.com/HaoWuLab-Bioinformatics/scHiCStackL, respectively. Hao Wu 0062, Yingfu Wu, Haoru Zhou, Zhongli Chen, Yi Xiong 0002, Quanzhong Liu, Hongming Zhang 0002 |
Briefings Bioinform. | 1 |
| 2022 | StackTADB: a stacking-based ensemble learning model for predicting the boundaries of topologically associating domains (TADs) accurately in fruit fliesabstractChromosome is composed of many distinct chromatin domains, referred to variably as topological domains or topologically associating domains (TADs). The domains are stable across different cell types and highly conserved across species, thus these chromatin domains have been considered as the basic units of chromosome folding and regarded as an important secondary structure in chromosome organization. However, the identification of TAD boundaries is still a great challenge due to the high cost and low resolution of Hi-C data or experiments. In this study, we propose a novel ensemble learning framework, termed as StackTADB, for predicting the boundaries of TADs. StackTADB integrates four base classifiers including Random Forest, Logistic Regression, K-NearestNeighbor and Support Vector Machine. From the analysis of a series of examinations on the data set in the previous study, it is concluded that StackTADB has optimal performance in six metrics, AUC, Accuracy, MCC, Precision, Recall and F1 score, and it is superior to the existing methods. In addition, the comparison of the performance of multiple features shows that Kmers-based features play an essential role in predicting TADs boundaries of fruit flies, and we also apply the SHapley Additive exPlanations (SHAP) framework to interpret the predictions of StackTADB to identify the reason why Kmers-based features are vital. The experimental results show that the subsequences matching the BEAF-32 motif play a crucial role in predicting the boundaries of TADs. The source code is freely available at https://github.com/HaoWuLab-Bioinformatics/StackTADB and the webserver of StackTADB is freely available at http://hwtad.sdu.edu.cn:8002/StackTADB. Hao Wu 0062, Zhaoheng Ai, Leyi Wei, Hongming Zhang 0002, Fan Yang 0068, Li-Zhen Cui 0001 |
Briefings Bioinform. | 1 |
| 2022 | Signaling repurposable drug combinations against COVID-19 by developing the heterogeneous deep herb-graph methodabstractBACKGROUND: Coronavirus disease 2019 (COVID-19) has spurred a boom in uncovering repurposable existing drugs. Drug repurposing is a strategy for identifying new uses for approved or investigational drugs that are outside the scope of the original medical indication. MOTIVATION: Current works of drug repurposing for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) are mostly limited to only focusing on chemical medicines, analysis of single drug targeting single SARS-CoV-2 protein, one-size-fits-all strategy using the same treatment (same drug) for different infected stages of SARS-CoV-2. To dilute these issues, we initially set the research focusing on herbal medicines. We then proposed a heterogeneous graph embedding method to signaled candidate repurposing herbs for each SARS-CoV-2 protein, and employed the variational graph convolutional network approach to recommend the precision herb combinations as the potential candidate treatments against the specific infected stage. METHOD: We initially employed the virtual screening method to construct the 'Herb-Compound' and 'Compound-Protein' docking graph based on 480 herbal medicines, 12,735 associated chemical compounds and 24 SARS-CoV-2 proteins. Sequentially, the 'Herb-Compound-Protein' heterogeneous network was constructed by means of the metapath-based embedding approach. We then proposed the heterogeneous-information-network-based graph embedding method to generate the candidate ranking lists of herbs that target structural, nonstructural and accessory SARS-CoV-2 proteins, individually. To obtain precision synthetic effective treatments forvarious COVID-19 infected stages, we employed the variational graph convolutional network method to generate candidate herb combinations as the recommended therapeutic therapies. RESULTS: There were 24 ranking lists, each containing top-10 herbs, targeting 24 SARS-CoV-2 proteins correspondingly, and 20 herb combinations were generated as the candidate-specific treatment to target the four infected stages. The code and supplementary materials are freely available at https://github.com/fanyang-AI/TCM-COVID19. Fan Yang 0068, Shuaijie Zhang, Ruiyuan Yao, Yanchun Zhang, Guoyin Wang 0001, Qianghua Zhang, Yunlong Cheng, Jihua Dong, Chunyang Ruan, Li-Zhen Cui 0001, Hao Wu 0062, Fuzhong Xue |
Briefings Bioinform. | 13 |
| 2022 | CLNN-loop: a deep learning model to predict CTCF-mediated chromatin loops in the different cell lines and CTCF-binding sites (CBS) pair typesabstractMOTIVATION: Three-dimensional (3D) genome organization is of vital importance in gene regulation and disease mechanisms. Previous studies have shown that CTCF-mediated chromatin loops are crucial to studying the 3D structure of cells. Although various experimental techniques have been developed to detect chromatin loops, they have been found to be time-consuming and costly. Nowadays, various sequence-based computational methods can capture significant features of 3D genome organization and help predict chromatin loops. However, these methods have low performance and poor generalization ability in predicting chromatin loops. RESULTS: Here, we propose a novel deep learning model, called CLNN-loop, to predict chromatin loops in different cell lines and CTCF-binding sites (CBS) pair types by fusing multiple sequence-based features. The analysis of a series of examinations based on the datasets in the previous study shows that CLNN-loop has satisfactory performance and is superior to the existing methods in terms of predicting chromatin loops. In addition, we apply the SHAP framework to interpret the predictions of different models, and find that CTCF motif and sequence conservation are important signs of chromatin loops in different cell lines and CBS pair types. AVAILABILITY AND IMPLEMENTATION: The source code of CLNN-loop is freely available at https://github.com/HaoWuLab-Bioinformatics/CLNN-loop and the webserver of CLNN-loop is freely available at http://hwclnn.sdu.edu.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yingfu Wu, Haoru Zhou, Hongming Zhang 0002, Hao Wu 0062 |
Bioinform. | 6 |
| 2021 | Large-scale comparative review and assessment of computational methods for anti-cancer peptide identificationabstractAnti-cancer peptides (ACPs) are known as potential therapeutics for cancer. Due to their unique ability to target cancer cells without affecting healthy cells directly, they have been extensively studied. Many peptide-based drugs are currently evaluated in the preclinical and clinical trials. Accurate identification of ACPs has received considerable attention in recent years; as such, a number of machine learning-based methods for in silico identification of ACPs have been developed. These methods promote the research on the mechanism of ACPs therapeutics against cancer to some extent. There is a vast difference in these methods in terms of their training/testing datasets, machine learning algorithms, feature encoding schemes, feature selection methods and evaluation strategies used. Therefore, it is desirable to summarize the advantages and disadvantages of the existing methods, provide useful insights and suggestions for the development and improvement of novel computational tools to characterize and identify ACPs. With this in mind, we firstly comprehensively investigate 16 state-of-the-art predictors for ACPs in terms of their core algorithms, feature encoding schemes, performance evaluation metrics and webserver/software usability. Then, comprehensive performance assessment is conducted to evaluate the robustness and scalability of the existing predictors using a well-prepared benchmark dataset. We provide potential strategies for the model performance improvement. Moreover, we propose a novel ensemble learning framework, termed ACPredStackL, for the accurate identification of ACPs. ACPredStackL is developed based on the stacking ensemble strategy combined with SVM, Naïve Bayesian, lightGBM and KNN. Empirical benchmarking experiments against the state-of-the-art methods demonstrate that ACPredStackL achieves a comparative performance for predicting ACPs. The webserver and source code of ACPredStackL is freely available at http://bigdata.biocie.cn/ACPredStackL/ and https://github.com/liangxiaoq/ACPredStackL, respectively. Fuyi Li, Hao Wu 0062, Jiangning Song, Quanzhong Liu |
Briefings Bioinform. | 5 |
| 2021 | Strategies of attack-defense game for wireless sensor networks considering the effect of confidence level in fuzzy environment
Yingfu Wu, Bingyi Kang, Hao Wu 0062 |
Eng. Appl. Artif. Intell. | 3 |
| 2016 | Network-Based Method for Inferring Cancer Progression at the Pathway Level from Cross-Sectional Mutation DataabstractLarge-scale cancer genomics projects are providing a wealth of somatic mutation data from a large number of cancer patients. However, it is difficult to obtain several samples with a temporal order from one patient in evaluating the cancer progression. Therefore, one of the most challenging problems arising from the data is to infer the temporal order of mutations across many patients. To solve the problem efficiently, we present a Network-based method (NetInf) to Infer cancer progression at the pathway level from cross-sectional data across many patients, leveraging on the exclusive property of driver mutations within a pathway and the property of linear progression between pathways. To assess the robustness of NetInf, we apply it on simulated data with the addition of different levels of noise. To verify the performance of NetInf, we apply it to analyze somatic mutation data from three real cancer studies with large number of samples. Experimental results reveal that the pathways detected by NetInf show significant enrichment. Our method reduces computational complexity by constructing gene networks without assigning the number of pathways, which also provides new insights on the temporal order of somatic mutations at the pathway level rather than at the gene level. Hao Wu 0062, Lin Gao 0006, Nikola K. Kasabov |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | Identifying overlapping mutated driver pathways by constructing gene networks in cancerabstractBACKGROUND: Large-scale cancer genomic projects are providing lots of data on genomic, epigenomic and gene expression aberrations in many cancer types. One key challenge is to detect functional driver pathways and to filter out nonfunctional passenger genes in cancer genomics. Vandin et al. introduced the Maximum Weight Sub-matrix Problem to find driver pathways and showed that it is an NP-hard problem. METHODS: To find a better solution and solve the problem more efficiently, we present a network-based method (NBM) to detect overlapping driver pathways automatically. This algorithm can directly find driver pathways or gene sets de novo from somatic mutation data utilizing two combinatorial properties, high coverage and high exclusivity, without any prior information. We firstly construct gene networks based on the approximate exclusivity between each pair of genes using somatic mutation data from many cancer patients. Secondly, we present a new greedy strategy to add or remove genes for obtaining overlapping gene sets with driver mutations according to the properties of high exclusivity and high coverage. RESULTS: To assess the efficiency of the proposed NBM, we apply the method on simulated data and compare results obtained from the NBM, RME, Dendrix and Multi-Dendrix. NBM obtains optimal results in less than nine seconds on a conventional computer and the time complexity is much less than the three other methods. To further verify the performance of NBM, we apply the method to analyze somatic mutation data from five real biological data sets such as the mutation profiles of 90 glioblastoma tumor samples and 163 lung carcinoma samples. NBM detects groups of genes which overlap with known pathways, including P53, RB and RTK/RAS/PI(3)K signaling pathways. New gene sets with p-value less than 1e-3 are found from the somatic mutation data. CONCLUSIONS: NBM can detect more biologically relevant gene sets. Results show that NBM outperforms other algorithms for detecting driver pathways or gene sets. Further research will be conducted with the use of novel machine learning techniques. Hao Wu 0062, Lin Gao 0006, Feng Li 0033, Xiaofei Yang 0003, Nikola K. Kasabov |
BMC Bioinform. | 1 |