Zhongming Zhao

dblp:11/4656 · DBLP profile ↗
← Back
89ranked-venue papers
2as first author
20since 2021 · last 2025
0000-0002-3477-0914ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 88 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 iGTP: learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics
abstract
Deep-learning models like Variational AutoEncoder have enabled low dimensional cellular embedding representation for large-scale single-cell transcriptomes and shown great flexibility in downstream tasks. However, biologically meaningful latent space is usually missing if no specific structure is designed. Here, we engineered a novel interpretable generative transcriptional program (iGTP) framework that could model the importance of transcriptional program (TP) space and protein-protein interactions (PPI) between different biological states. We demonstrated the performance of iGTP in a diverse biological context using gene ontology, canonical pathway, and different PPI curation. iGTP not only elucidated the ground truth of cellular responses but also surpassed other deep learning models and traditional bioinformatics methods in functional enrichment tasks. By integrating the latent layer with a graph neural network framework, iGTP could effectively infer cellular responses to perturbations. Lastly, we applied iGTP TP embeddings with a latent diffusion model to accurately generate cell embeddings for specific cell types and states. We anticipate that iGTP will offer insights at both PPI and TP levels and holds promise for predicting responses to novel perturbations.
Kanglin Hsieh, Yan Chu 0005, Lishan Yu, Nuo Hu, Isha Kawosa, Patrick G. Pilié, Pratip K. Bhattacharya, Degui Zhi, Xiaoqian Jiang, Zhongming Zhao, Yulin Dai
Briefings Bioinform.12
2025 BrainGeneBot: a framework for variant prioritization and generative pretrained transformer-informed interpretation across polygenic risk score studies
abstract
Polygenic risk scores (PRS) are widely used to assess genetic susceptibility in Alzheimer's disease (AD) research. However, the rapid expansion of PRS studies has led to dataset-specific biases-stemming from factors like population makeup, genotyping methods, and analysis pipelines-that result in inconsistent variant prioritization and limit generalizability and reproducibility. To address these challenges, we propose a transductive learning framework that integrates multiple PRS datasets for more robust risk variant prioritization, incorporating genome-wide association study (GWAS) priority scores as biologically informed priors. Additionally, we introduce BrainGeneBot, an AI-driven tool leveraging generative pretrained transformers with retrieval-augmented generation technology to streamline genomic analyses in AD, including the STRING for protein interaction analysis, Enrichr for gene set enrichment, ClinVar for genetic variant interpretation, and Biopython for conducting literature searches. We apply our approach to publicly available AD datasets from the PGS Catalog and conduct further analyses to validate its efficacy. In parallel, we perform conventional unsupervised rank aggregation as a baseline. The transductive learning approach not only verifies high-risk variants identified by traditional methods but also reveals unique insights that better correlate with GWAS signals. Our framework streamlines data retrieval and interpretation, effectively prioritizing genetic variants in multiple PRS studies. Moreover, BrainGeneBot facilitates the discovery of biologically meaningful insights to enhance PRS interpretability and applicability in AD research, supporting the development of precise AD interventions and treatments. Our approach provides a robust framework for AD genetic research, improving data accessibility, accelerating discoveries, and refining genetic insights.
Gang Qu 0002, Nitesh Enduru, Xiaoqian Jiang, Zhongming Zhao
Briefings Bioinform.5
2025 Deciphering RNA modification and post-transcriptional regulation with NetRNApan
abstract
RNA modification, which is evolutionarily conserved, is crucial for modulating various biological functions and disease pathogenesis. High resolution transcriptome-wide mapping of RNA modifications has facilitated both data resources and computational prediction of RNA modification. While these prediction algorithms are promising, they are limited in interpretability or generalizability, or the capacity for discovering novel post-transcriptional regulations. Here, we present NetRNApan, a deep learning framework for RNA modification site prediction, motif discovery and trans-regulatory factor identification. Using m5U profiles generated by FICC-seq and miCLIP-seq technologies and single-base resolution m6A sites from multiple experiments as cases, we demonstrated the accuracy of NetRNApan with more efficient and interpretive feature representations. For m5U modification, we uncovered five representative clusters with consensus motifs that may be essential by decoding the informative characteristics detected by NetRNApan. Furthermore, NetRNApan revealed interesting trans-regulatory factors and provided a protein-binding perspective for investigating the function of RNA modifications. Specifically, we discovered 21 potential functional RNA-binding proteins (RBPs) whose binding sites were significantly linked to the extracted top-scoring motifs for m5U modification. Two examples are ANKHD1 and RBM4 with potential regulatory function of m5U modifications. Meanwhile, the analysis of convolution layer parameters within the model offers valuable insights into the regulation of m6A in humans. Collectively, NetRNApan demonstrated high accuracy, interpretability and generalizability for study of RNA modification and mRNA regulation. NetRNApan is freely available at https://github.com/bsml320/NetRNApan.
Haodong Xu, Wankun Deng, Ruifeng Hu 0002, Binfeng Liu, Lujuan Wang, Xiaolei Ren 0007, Chao Tu, Zhongming Zhao
Briefings Bioinform.11
2025 Scupa: single-cell unified polarization assessment of immune cells using the single-cell foundation model
abstract
MOTIVATION: Immune cells undergo cytokine-driven polarization in response to diverse stimuli, altering their transcriptional profiles and functional states. This dynamic process is central to immune responses in health and diseases, yet a systematic approach to assess cytokine-driven polarization in single-cell RNA sequencing data has been lacking. RESULTS: To address this gap, we developed single-cell unified polarization assessment (Scupa), the first computational method for comprehensive immune cell polarization assessment. Scupa leverages data from the Immune Dictionary, which characterizes cytokine-driven polarization states across 14 immune cell types. By integrating cell embeddings from the single-cell foundation model Universal Cell Embeddings, Scupa effectively identifies polarized cells across different species and experimental conditions. Applications of Scupa in independent datasets demonstrated its accuracy in classifying polarized cells and further revealed distinct polarization profiles in tumor-infiltrating myeloid cells across cancers. Scupa complements conventional single-cell data analysis by providing new insights into dynamic immune cell states, and holds potential for advancing therapeutic insights, particularly in cytokine-based therapies. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/bsml320/Scupa.
Wendao Liu, Zhongming Zhao
Bioinform.2
2024 Deep Transfer Learning for Kidney Disease Detection Using CT Scan Images
abstract
Kidney disease is a significant health issue that leads to a high number of deaths worldwide. Accurate and timely detection of kidney diseases, including cysts, stones, and tumours, is critical for effective treatment and patient outcomes. Deep learning methodologies, specifically transfer learning, have recently been widely used in medical image analysis. This study comprehensively evaluates seven transfer learning models—Xception, DenseNet201, MobileNet, InceptionV3, VGG16, ResNet50, and EfficientNetB0—for multiclass kidney disease classification using CT scan images. The models are assessed based on key performance metrics such as accuracy, loss, precision, recall, F1-score, and AUC. Results indicate that ResNet50 delivers the best overall performance, making it a promising model for clinical applications. By comparing these state-of-the-art models, this research provides valuable insights into the most effective transfer learning approaches for kidney disease detection, paving the way for robust computer-aided diagnostic systems.
Shahid Mohammad Ganie, Pijush Kanti Dutta Pramanik, Zhongming Zhao
BIBM3
2024 MetaDegron: multimodal feature-integrated protein language model for predicting E3 ligase targeted degrons
abstract
Protein degradation through the ubiquitin proteasome system at the spatial and temporal regulation is essential for many cellular processes. E3 ligases and degradation signals (degrons), the sequences they recognize in the target proteins, are key parts of the ubiquitin-mediated proteolysis, and their interactions determine the degradation specificity and maintain cellular homeostasis. To date, only a limited number of targeted degron instances have been identified, and their properties are not yet fully characterized. To tackle on this challenge, here we develop a novel deep-learning framework, namely MetaDegron, for predicting E3 ligase targeted degron by integrating the protein language model and comprehensive featurization strategies. Through extensive evaluations using benchmark datasets and comparison with existing method, such as Degpred, we demonstrate the superior performance of MetaDegron. Among functional features, MetaDegron allows batch prediction of targeted degrons of 21 E3 ligases, and provides functional annotations and visualization of multiple degron-related structural and physicochemical features. MetaDegron is freely available at http://modinfor.com/MetaDegron/. We anticipate that MetaDegron will serve as a useful tool for the clinical and translational community to elucidate the mechanisms of regulation of protein homeostasis, cancer research, and drug development.
Mengqiu Zheng, Shaofeng Lin, Kunqi Chen, Ruifeng Hu 0002, Zhongming Zhao, Haodong Xu
Briefings Bioinform.6
2024 GENEVIC: GENetic data Exploration and Visualization via Intelligent interactive Console
abstract
SUMMARY: The vast generation of genetic data poses a significant challenge in efficiently uncovering valuable knowledge. Introducing GENEVIC, an AI-driven chat framework that tackles this challenge by bridging the gap between genetic data generation and biomedical knowledge discovery. Leveraging generative AI, notably ChatGPT, it serves as a biologist's "copilot." It automates the analysis, retrieval, and visualization of customized domain-specific genetic information, and integrates functionalities to generate protein interaction networks, enrich gene sets, and search scientific literature from PubMed, Google Scholar, and arXiv, making it a comprehensive tool for biomedical research. In its pilot phase, GENEVIC is assessed using a curated database that ranks genetic variants associated with Alzheimer's disease, schizophrenia, and cognition, based on their effect weights from the Polygenic Score (PGS) Catalog, thus enabling researchers to prioritize genetic variants in complex diseases. GENEVIC's operation is user-friendly, accessible without any specialized training, secured by Azure OpenAI's HIPAA-compliant infrastructure, and evaluated for its efficacy through real-time query testing. As a prototype, GENEVIC is set to advance genetic research, enabling informed biomedical decisions. AVAILABILITY AND IMPLEMENTATION: GENEVIC is publicly accessible at https://genevicanath2024.streamlit.app. The underlying code is open-source and available via GitHub at https://github.com/bsml320/GENEVIC.git (also at https://github.com/anath2110/GENEVIC.git).
Anindita Nath, Savannah Mwesigwa, Yulin Dai, Xiaoqian Jiang, Zhongming Zhao
Bioinform.5
2024 Enabling the clinical application of artificial intelligence in genomics: a perspective of the AMIA Genomics and Translational Bioinformatics Workgroup
abstract
OBJECTIVE: Given the importance AI in genomics and its potential impact on human health, the American Medical Informatics Association-Genomics and Translational Biomedical Informatics (GenTBI) Workgroup developed this assessment of factors that can further enable the clinical application of AI in this space. PROCESS: A list of relevant factors was developed through GenTBI workgroup discussions in multiple in-person and online meetings, along with review of pertinent publications. This list was then summarized and reviewed to achieve consensus among the group members. CONCLUSIONS: Substantial informatics research and development are needed to fully realize the clinical potential of such technologies. The development of larger datasets is crucial to emulating the success AI is achieving in other domains. It is important that AI methods do not exacerbate existing socio-economic, racial, and ethnic disparities. Genomic data standards are critical to effectively scale such technologies across institutions. With so much uncertainty, complexity and novelty in genomics and medicine, and with an evolving regulatory environment, the current focus should be on using these technologies in an interface with clinicians that emphasizes the value each brings to clinical decision-making.
Nephi Walton, Radhakrishnan Nagarajan, Chen Wang 0001, Murat Sincan, Robert R. Freimuth, David B. Everman, Derek C. Walton, Scott McGrath, Dominick J. Lemas, Panayiotis V. Benos, Alexander V. Alekseyenko, Qianqian Song 0002, Ece D. Gamsiz Uzun, Casey Overby Taylor, Alper Uzun, Thomas N. Person, Nadav Rappoport, Zhongming Zhao, Marc S. Williams
J. Am. Medical Informatics Assoc.18
2023 Benchmark of embedding-based methods for accurate and transferable prediction of drug response
abstract
Prediction of therapy response has been a major challenge in cancer precision medicine due to the extensive tumor heterogeneity. Recently, several deep learning methods have been developed to predict drug response by utilizing various omics data. Most of them train models by using the drug-response screening data generated from cell lines and then use these models to predict response in cancer patient data. In this study, we focus on and evaluate deep learning methods using transcriptome data for the long-standing question of personalized drug-response prediction. We developed an embedding-based approach for drug-response prediction and benchmarked similar methods for their performance. For all methods, we used pretreatment transcriptome data to train models and then conducted a comprehensive evaluation and comparison of the models using cross-panels, cross-datasets and target genes. We further validated the methods using three independent datasets assessing multiple compounds for their predictive capability of drug response, survival outcome and cell line status. As a result, the methods building on gene embeddings had an overall competitive performance with reduced overfitting when we applied evaluation parameters for model fitting as well as the correlation with clinical outcomes in the validation data. We further developed an ensemble model to combine the results from the three most competitive methods for an overall prediction. Finally, we developed DrVAEN (https://bioinfo.uth.edu/drvaen), a user-friendly and easy-accessible web-server that hosts all these methods for drug-response prediction and model comparison for broad use in cancer research, method evaluation and drug development.
Peilin Jia, Ruifeng Hu 0002, Zhongming Zhao
Briefings Bioinform.3
2023 Advances and challenges in Bioinformatics and Biomedical Engineering: IWBBIO 2020
abstract
This Supplement issue, presents five research articles which are distributed, mainly due to the subject they address, from the 8th International Work-Conference on Bioinformatics and Biomedical Engineering (IWBBIO 2020), which was held on line, during September, 30th-2nd October, 2020. These contributions have been chosen because of their quality and the importance of their findings. Those contributions were then invited to participate in this supplement for the following journals of BMC: BMC Bioinformatics and BMC Genomics. In the present Editorial in BMC journal, we summarize the contributions that provide a clear overview of the thematic areas covered by the IWBBIO conference, ranging from theoretical/review aspects to real-world applications of bioinformatic and biomedical engineering.
Olga Valenzuela, Mario Cannataro, Irena Rusur, Jianxin Wang 0001, Zhongming Zhao, Ignacio Rojas
BMC Bioinform.5
2022 Detection of Chronic Kidney Disease Using Neuro-Fuzzy Rule-based Classifier
abstract
Chronic kidney disease (CKD) is as severe as cancer in today s world. It may even lead to the permanent failure of kidney. The initial detection of this disease is needed for timely cure. Our work presents a classifier (named ANFIS) in accordance with the notion of neuro-fuzzy in order to detect the existence of chronic kidney disease. We use blood test results of several patients for our research study. We compare our proposed classifier with some conventional classifiers such as Multi-layer Perceptron, Support Vector Machine, Logistic Regression and Decision Tree. Experimental results indicates that our proposed neuro-fuzzy rule-based classifier performs better than the other classifiers used here. ANFIS has given 3% to 4% better accuracy compared to the other classifiers.
Supantha Das, Arnab Hazra, Soumen Kumar Pati, Soumadip Ghosh, Saurav Mallik, Suharta Banerjee, Ayan Mukherji, Zhongming Zhao
BIBM9
2022 Grid-Search Integrated Optimized Support Vector Machine Model for Breast Cancer Detection
abstract
Breast cancer is a common and highly heterogeneous cancer worldwide. Rapid detection and early diagnosis are essential in its treatment, but it is challenging due to mammogram’s uncertainty. Based on machine learning, operative clinical systems can be implemented to save patients’ lives. This study proposes an optimized support vector machine (SVM) model to predict breast cancer using the grid search method to find the best hyper-parameters. For validation, we offer an in-depth experimental study and we achieve, for Benign|Malignant|Average cases, accuracy of 98.69%|99.72%|99.0%, precision of 96.0%|100.0%|98.0%, recall of 100.0%|98.0%|98.0%, and Fl-score of 98.0%|97.0%|98.0% respectively. Comparison is made between the tuned hyper-parameter and default hyper-parameter performance. SVM performance with default parameters is 96%, while the maximum accuracy is achieved by hyper-parameters tuned. An SVM model is 99%, which can be a better cancer detection system. The results show performance improves when the best hyperparameters are used for SVM training. Thus, the analysis and comparison indicate that our stated system is better than the state-of-the-art ML-based breast cancer detection system.
Partho Ghose, Selina Sharmin, Loveleen Gaur, Zhongming Zhao
BIBM4
2022 KCOSS: an ultra-fast k-mer counter for assembled genome analysis
abstract
MOTIVATION: The k-mer frequency in whole genome sequences provides researchers with an insightful perspective on genomic complexity, comparative genomics, metagenomics and phylogeny. The current k-mer counting tools are typically slow, and they require large memory and hard disk for assembled genome analysis. RESULTS: We propose a novel and ultra-fast k-mer counting algorithm, KCOSS, to fulfill k-mer counting mainly for assembled genomes with segmented Bloom filter, lock-free queue, lock-free thread pool and cuckoo hash table. We optimize running time and memory consumption by recycling memory blocks, merging multiple consecutive first-occurrence k-mers into C-read, and writing a set of C-reads to disk asynchronously. KCOSS was comparatively tested with Jellyfish2, CHTKC and KMC3 on seven assembled genomes and three sequencing datasets in running time, memory consumption, and hard disk occupation. The experimental results show that KCOSS counts k-mer with less memory and disk while having a shorter running time on assembled genomes. KCOSS can be used to calculate the k-mer frequency not only for assembled genomes but also for sequencing data. AVAILABILITYAND IMPLEMENTATION: The KCOSS software is implemented in C++. It is freely available on GitHub: https://github.com/kcoss-2021/KCOSS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Deyou Tang, Daqiang Tan, Juan Fu, Yelei Tang, Jiabin Lin, Hongli Du, Zhongming Zhao
Bioinform.9
2022 Comparison of five supervised feature selection algorithms leading to top features and gene signatures from multi-omics data in cancer
abstract
BACKGROUND: As many complex omics data have been generated during the last two decades, dimensionality reduction problem has been a challenging issue in better mining such data. The omics data typically consists of many features. Accordingly, many feature selection algorithms have been developed. The performance of those feature selection methods often varies by specific data, making the discovery and interpretation of results challenging. METHODS AND RESULTS: In this study, we performed a comprehensive comparative study of five widely used supervised feature selection methods (mRMR, INMIFS, DFS, SVM-RFE-CBR and VWMRmR) for multi-omics datasets. Specifically, we used five representative datasets: gene expression (Exp), exon expression (ExpExon), DNA methylation (hMethyl27), copy number variation (Gistic2), and pathway activity dataset (Paradigm IPLs) from a multi-omics study of acute myeloid leukemia (LAML) from The Cancer Genome Atlas (TCGA). The different feature subsets selected by the aforesaid five different feature selection algorithms are assessed using three evaluation criteria: (1) classification accuracy (Acc), (2) representation entropy (RE) and (3) redundancy rate (RR). Four different classifiers, viz., C4.5, NaiveBayes, KNN, and AdaBoost, were used to measure the classification accuary (Acc) for each selected feature subset. The VWMRmR algorithm obtains the best Acc for three datasets (ExpExon, hMethyl27 and Paradigm IPLs). The VWMRmR algorithm offers the best RR (obtained using normalized mutual information) for three datasets (Exp, Gistic2 and Paradigm IPLs), while it gives the best RR (obtained using Pearson correlation coefficient) for two datasets (Gistic2 and Paradigm IPLs). It also obtains the best RE for three datasets (Exp, Gistic2 and Paradigm IPLs). Overall, the VWMRmR algorithm yields best performance for all three evaluation criteria for majority of the datasets. In addition, we identified signature genes using supervised learning collected from the overlapped top feature set among five feature selection methods. We obtained a 7-gene signature (ZMIZ1, ENG, FGFR1, PAWR, KRT17, MPO and LAT2) for EXP, a 9-gene signature for ExpExon, a 7-gene signature for hMethyl27, one single-gene signature (PIK3CG) for Gistic2 and a 3-gene signature for Paradigm IPLs. CONCLUSION: We performed a comprehensive comparison of the performance evaluation of five well-known feature selection methods for mining features from various high-dimensional datasets. We identified signature genes using supervised learning for the specific omic data for the disease. The study will help incorporate higher order dependencies among features.
Tapas Bhadra, Saurav Mallik, Neaj Hasan, Zhongming Zhao
BMC Bioinform.4
2022 Unsupervised Feature Selection Using an Integrated Strategy of Hierarchical Clustering With Singular Value Decomposition: An Integrative Biomarker Discovery Method With Application to Acute Myeloid Leukemia
abstract
In this article, we propose a novel unsupervised feature selection method by combining hierarchical feature clustering with singular value decomposition (SVD). The proposed algorithm first generates several feature clusters by adopting the hierarchical clustering on the feature space and then applies SVD to each of these feature clusters to find out the feature that contributes most to the SVD-entropy. The proposed feature selection method selects an optimal feature subset that not only minimizes the mutual dependency among the selected features but also maximizes the mutual dependency of the selected features against their nearest neighbor non-selected features to some extent. Each of the selected features also contributes the maximum SVD-entropy among all features of the same feature cluster. The experimental results demonstrate that the proposed algorithm performs well against several state-of-the-art methods of feature selection in terms of various evaluation criteria such as classification accuracy, redundancy rate, and representation entropy. The superiority of the proposed algorithm is demonstrated through analysis of Acute Myeloid Leukemia (AML) multi-omics data that consist of five datasets: gene expression, exon expression, methylation, microRNA, and pathway activity dataset (paradigm IPLs) from The Cancer Genome Atlas (TCGA). Our analysis pinpoints a candidate gene-marker, EREG for AML with an integrative omics evidence. EREG is targeted by two top ranked microRNAs, hsa-miR-1286 and hsa-miR-1976, here in the datasets. The method and results will be useful for biomarker discovery in the era of in precision medicine.
Tapas Bhadra, Saurav Mallik, Amir Sohel, Zhongming Zhao
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 Mining Cancer Cell Line-Based Drugs to Benefit KRAS(G12D) Pancreatic Adenocarcinoma Patients
abstract
Pancreatic adenocarcinoma (PAAD) is one of the most challenging cancers with high morbidity and mortality. KRAS mutations could occur as an early event in PAAD. KRAS genes are highly mutated with recurrent mutations in various cancer types, including PAAD. We aim to depict the omics landscape of KRAS mutations and seek potential novel drugs for pancreatic cancer patients with KRAS mutations. This study used data from The Cancer Genome Atlas (TCGA) and the Cancer Cell Line Encyclopedia (CCLE) for KRAS mutation analysis in a multi-omics manner. We found that the genomics and transcriptomics patterns of KRAS mutations are significantly different compared to the corresponding non-KRAS-mutated PAAD samples. Pancreatic cancer's prognosis is directly associated with a specific KRAS mutation and its protein's structure instability. Our analysis confirmed that Irinotecan could be a potential drug for PAAD patients with KRASG12Dmutation.
Aman Chandra Kaushik, Aamir Mehmood, Ankit Babu, Zhongming Zhao
BIBM6
2021 Negatively-Associated Maximal Frequent Geneset Mining on DNA Methylation Profile
abstract
Association rule mining has been an important approach for feature and biomarker discovery in various omics data. One main challenge is that it generates a large number of itemsets. The effect of this shortcoming increases substantially in the case of negative association rule mining that is useful for detecting significant relationships between genes (items/features) in the form of either presence or absence in disease characterization. In this article, we propose a new algorithm, NegaMax (negatively-associated maximal frequent itemsets) for negative association itemset mining. Our method follows depth-first search rather than breadth-first search used in the other methods. It identifies a much fewer number of non-redundant itemsets than that by the existing methods. Thus, it saves elapsing time for itemset generation which potentially remove false positive intermediate results. We demonstrated NegaMax in a real-world DNA methylation dataset. The proposed method is highly beneficial from a medical perspective.
Saurav Mallik, Souvik Rakshit, Ujjwal Maulik, Zhongming Zhao
BIBM4
2021 Distinct effect of prenatal and postnatal brain expression across 20 brain disorders and anthropometric social traits: a systematic study of spatiotemporal modularity
abstract
Different spatiotemporal abnormalities have been implicated in different neuropsychiatric disorders and anthropometric social traits, yet an investigation in the temporal network modularity with brain tissue transcriptomics has been lacking. We developed a supervised network approach to investigate the genome-wide association study (GWAS) results in the spatial and temporal contexts and demonstrated it in 20 brain disorders and anthropometric social traits. BrainSpan transcriptome profiles were used to discover significant modules enriched with trait susceptibility genes in a developmental stage-stratified manner. We investigated whether, and in which developmental stages, GWAS-implicated genes are coordinately expressed in brain transcriptome. We identified significant network modules for each disorder and trait at different developmental stages, providing a systematic view of network modularity at specific developmental stages for a myriad of brain disorders and traits. Specifically, we observed a strong pattern of the fetal origin for most psychiatric disorders and traits [such as schizophrenia (SCZ), bipolar disorder, obsessive-compulsive disorder and neuroticism], whereas increased co-expression activities of genes were more strongly associated with neurological diseases [such as Alzheimer's disease (AD) and amyotrophic lateral sclerosis] and anthropometric traits (such as college completion, education and subjective well-being) in postnatal brains. Further analyses revealed enriched cell types and functional features that were supported and corroborated prior knowledge in specific brain disorders, such as clathrin-mediated endocytosis in AD, myelin sheath in multiple sclerosis and regulation of synaptic plasticity in both college completion and education. Our study provides a landscape view of the spatiotemporal features in a myriad of brain-related disorders and traits.
Peilin Jia, Astrid Marilyn Manuel, Brisa S. Fernandes, Yulin Dai, Zhongming Zhao
Briefings Bioinform.5
2021 Landscape of drug-resistance mutations in kinase regulatory hotspots
abstract
More than 48 kinase inhibitors (KIs) have been approved by Food and Drug Administration. However, drug-resistance (DR) eventually occurs, and secondary mutations have been found in the previously targeted primary-mutated cancer cells. Cancer and drug research communities recognize the importance of the kinase domain (KD) mutations for kinasopathies. So far, a systematic investigation of kinase mutations on DR hotspots has not been done yet. In this study, we systematically investigated four types of representative mutation hotspots (gatekeeper, G-loop, αC-helix and A-loop) associated with DR in 538 human protein kinases using large-scale cancer data sets (TCGA, ICGC, COSMIC and GDSC). Our results revealed 358 kinases harboring 3318 mutations that covered 702 drug resistance hotspot residues. Among them, 197 kinases had multiple genetic variants on each residue. We further computationally assessed and validated the epidermal growth factor receptor mutations on protein structure and drug-binding efficacy. This is the first study to provide a landscape view of DR-associated mutation hotspots in kinase's secondary structures, and its knowledge will help the development of effective next-generation KIs for better precision medicine.
Pora Kim, Junmei Wang, Zhongming Zhao
Briefings Bioinform.4
2021 Deep4mC: systematic assessment and computational prediction for DNA N4-methylcytosine sites by deep learning
abstract
DNA N4-methylcytosine (4mC) modification represents a novel epigenetic regulation. It involves in various cellular processes, including DNA replication, cell cycle and gene expression, among others. In addition to experimental identification of 4mC sites, in silico prediction of 4mC sites in the genome has emerged as an alternative and promising approach. In this study, we first reviewed the current progress in the computational prediction of 4mC sites and systematically evaluated the predictive capacity of eight conventional machine learning algorithms as well as 12 feature types commonly used in previous studies in six species. Using a representative benchmark dataset, we investigated the contribution of feature selection and stacking approach to the model construction, and found that feature optimization and proper reinforcement learning could improve the performance. We next recollected newly added 4mC sites in the six species' genomes and developed a novel deep learning-based 4mC site predictor, namely Deep4mC. Deep4mC applies convolutional neural networks with four representative features. For species with small numbers of samples, we extended our deep learning framework with a bootstrapping method. Our evaluation indicated that Deep4mC could obtain high accuracy and robust performance with the average area under curve (AUC) values greater than 0.9 in all species (range: 0.9005-0.9722). In comparison, Deep4mC achieved an AUC value improvement from 10.14 to 46.21% when compared to previous tools in these six species. A user-friendly web server (https://bioinfo.uth.edu/Deep4mC) was built for predicting putative 4mC sites in a genome.
Hao-Dong Xu, Peilin Jia, Zhongming Zhao
Briefings Bioinform.3
2020 Critical microRNAs and regulatory motifs in cleft palate identified by a conserved miRNA-TF-gene network approach in humans and mice
abstract
Cleft palate (CP) is the second most common congenital birth defect. The etiology of CP is complicated, with involvement of various genetic and environmental factors. To investigate the gene regulatory mechanisms, we designed a powerful regulatory analytical approach to identify the conserved regulatory networks in humans and mice, from which we identified critical microRNAs (miRNAs), target genes and regulatory motifs (miRNA-TF-gene) related to CP. Using our manually curated genes and miRNAs with evidence in CP in humans and mice, we constructed miRNA and transcription factor (TF) co-regulation networks for both humans and mice. A consensus regulatory loop (miR17/miR20a-FOXE1-PDGFRA) and eight miRNAs (miR-140, miR-17, miR-18a, miR-19a, miR-19b, miR-20a, miR-451a and miR-92a) were discovered in both humans and mice. The role of miR-140, which had the strongest association with CP, was investigated in both human and mouse palate cells. The overexpression of miR-140-5p, but not miR-140-3p, significantly inhibited cell proliferation. We further examined whether miR-140 overexpression could suppress the expression of its predicted target genes (BMP2, FGF9, PAX9 and PDGFRA). Our results indicated that miR-140-5p overexpression suppressed the expression of BMP2 and FGF9 in cultured human palate cells and Fgf9 and Pdgfra in cultured mouse palate cells. In summary, our conserved miRNA-TF-gene regulatory network approach is effective in detecting consensus miRNAs, motifs, and regulatory mechanisms in human and mouse CP.
Peilin Jia, Saurav Mallik, Rong Fei, Hiroki Yoshioka, Akiko Suzuki, Junichi Iwata, Zhongming Zhao
Briefings Bioinform.8
2020 Graph- and rule-based learning algorithms: a comprehensive review of their applications for cancer type classification and prognosis using genomic data
abstract
Cancer is well recognized as a complex disease with dysregulated molecular networks or modules. Graph- and rule-based analytics have been applied extensively for cancer classification as well as prognosis using large genomic and other data over the past decade. This article provides a comprehensive review of various graph- and rule-based machine learning algorithms that have been applied to numerous genomics data to determine the cancer-specific gene modules, identify gene signature-based classifiers and carry out other related objectives of potential therapeutic value. This review focuses mainly on the methodological design and features of these algorithms to facilitate the application of these graph- and rule-based analytical approaches for cancer classification and prognosis. Based on the type of data integration, we divided all the algorithms into three categories: model-based integration, pre-processing integration and post-processing integration. Each category is further divided into four sub-categories (supervised, unsupervised, semi-supervised and survival-driven learning analyses) based on learning style. Therefore, a total of 11 categories of methods are summarized with their inputs, objectives and description, advantages and potential limitations. Next, we briefly demonstrate well-known and most recently developed algorithms for each sub-category along with salient information, such as data profiles, statistical or feature selection methods and outputs. Finally, we summarize the appropriate use and efficiency of all categories of graph- and rule mining-based learning methods when input data and specific objective are given. This review aims to help readers to select and use the appropriate algorithms for cancer classification and prognosis study.
Saurav Mallik, Zhongming Zhao
Briefings Bioinform.2
2020 6mA-Finder: a novel online tool for predicting DNA N6-methyladenine sites in genomes
abstract
MOTIVATION: DNA N6-methyladenine (6 mA) has recently been found as an essential epigenetic modification, playing its roles in a variety of cellular processes. The abnormal status of DNA 6 mA modification has been reported in cancer and other disease. The annotation of 6 mA marks in genome is the first crucial step to explore the underlying molecular mechanisms including its regulatory roles. RESULTS: We present a novel online DNA 6 mA site tool, 6 mA-Finder, by incorporating seven sequence-derived information and three physicochemical-based features through recursive feature elimination strategy. Our multiple cross-validations indicate the promising accuracy and robustness of our model. 6 mA-Finder outperforms its peer tools in general and species-specific 6 mA site prediction, suggesting it can provide a useful resource for further experimental investigation of DNA 6 mA modification. AVAILABILITY AND IMPLEMENTATION: https://bioinfo.uth.edu/6mA_Finder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao-Dong Xu, Ruifeng Hu 0002, Peilin Jia, Zhongming Zhao
Bioinform.4
2020 Accelerating bioinformatics research with International Conference on Intelligent Biology and Medicine 2020
abstract
The International Association for Intelligent Biology and Medicine (IAIBM) is a nonprofit organization that promotes intelligent biology and medical science. It hosts an annual International Conference on Intelligent Biology and Medicine (ICIBM), which was initially established in 2012. Due to the coronavirus (COVID-19) pandemic, the ICIBM 2020 was held for the first time as a virtual online conference on August 9 to 10. The virtual conference had ~ 300 registered participants and featured 41 online real-time presentations. ICIBM 2020 received a total of 75 manuscript submissions, and 12 were selected to be published in this special issue of BMC Bioinformatics. These 12 manuscripts cover a wide range of bioinformatics topics including network analysis, imaging analysis, machine learning, gene expression analysis, and sequence analysis.
Li Shen 0001, Xinghua Shi, Kai Wang 0063, Yulin Dai, Zhongming Zhao
BMC Bioinform.6
2020 In silico ranking of phenolics for therapeutic effectiveness on cancer stem cells
abstract
BACKGROUND: Cancer stem cells (CSCs) have features such as the ability to self-renew, differentiate into defined progenies and initiate the tumor growth. Treatments of cancer include drugs, chemotherapy and radiotherapy or a combination. However, treatment of cancer by various therapeutic strategies often fail. One possible reason is that the nature of CSCs, which has stem-like properties, make it more dynamic and complex and may cause the therapeutic resistance. Another limitation is the side effects associated with the treatment of chemotherapy or radiotherapy. To explore better or alternative treatment options the current study aims to investigate the natural drug-like molecules that can be used as CSC-targeted therapy. Among various natural products, anticancer potential of phenolics is well established. We collected the 21 phytochemicals from phenolic group and their interacting CSC genes from the publicly available databases. Then a bipartite graph is constructed from the collected CSC genes along with their interacting phytochemicals from phenolic group as other. The bipartite graph is then transformed into weighted bipartite graph by considering the interaction strength between the phenolics and the CSC genes. The CSC genes are also weighted by two scores, namely, DSI (Disease Specificity Index) and DPI (Disease Pleiotropy Index). For each gene, its DSI score reflects the specific relationship with the disease and DPI score reflects the association with multiple diseases. Finally, a ranking technique is developed based on PageRank (PR) algorithm for ranking the phenolics. RESULTS: We collected 21 phytochemicals from phenolic group and 1118 CSC genes. The top ranked phenolics were evaluated by their molecular and pharmacokinetics properties and disease association networks. We selected top five ranked phenolics (Resveratrol, Curcumin, Quercetin, Epigallocatechin Gallate, and Genistein) for further examination of their oral bioavailability through molecular properties, drug likeness through pharmacokinetic properties, and associated network with CSC genes. CONCLUSION: Our PR ranking based approach is useful to rank the phenolics that are associated with CSC genes. Our results suggested some phenolics are potential molecules for CSC-related cancer treatment.
Monalisa Mandal, Sanjeeb Kumar Sahoo, Priyadarsan Patra, Saurav Mallik, Zhongming Zhao
BMC Bioinform.5
2020 Correction to: The International Conference on Intelligent Biology and Medicine (ICIBM) 2019: bioinformatics methods and applications for human diseases
abstract
After publication of this supplement article [1], it is requested the grant ID in the Funding section should be corrected from NSF grant IIS-7811367 to NSF grant IIS-1902617. Therefore, the correct 'Funding' section in this article should read: We thank the National Science Foundation (NSF grant IIS-1902617) for the financial support of ICIBM 2019. This article has not received sponsorship for publication.
Zhongming Zhao, Yulin Dai, Ewy A. Mathé, Kai Wang 0063
BMC Bioinform.1
2019 A Multi-classifier Model to Identify Mitochondrial Respiratory Gene Signatures in Human Cancer
abstract
Whether alteration of mitochondrial gene expression can serve as an effective molecular signature for cancer classification currently remains controversy. To tackle this challenge, here we present a multi-classifier model to identify mitochondrial aberrant gene signatures and then assess the effectiveness on cancer classification. Specifically, we first applied a supervised learning model, Empirical Bayes statistics, to detect differentially expressed genes from a liver cancer mitochondrial gene expression dataset (GEO accession: GSE64505). Next, we applied two well-known classifiers, Prediction Analysis of Microarrays (PAM) and Random Forest (RF), with several folds of cross-validation, to find molecular signatures that can classify liver cancer samples from normal samples. We obtained a mitochondrial molecular signature comprising 25 genes. Classification accuracy was 87.5% by PAM classifier (5 fold cross-validation with 30 repeats), while it was 87.5% and 75.0% by Random Forest with 5- and 2-fold cross-validations (30 repeats), respectively. We further performed literature mining and Gene Set Enrichment Analysis (GSEA) to evaluate the biological significance and novelty of genes in this gene signature.
Saurav Mallik, Soumita Seth, Tapas Bhadra, Namrata Tomar, Zhongming Zhao
BIBM5
2019 Distinct telomere length and molecular signatures in seminoma and non-seminoma of testicular germ cell tumor
abstract
Testicular germ cell tumors (TGCTs) are classified into two main subtypes, seminoma (SE) and non-seminoma (NSE), but their molecular distinctions remain largely unexplored. Here, we used expression data for mRNAs and microRNAs (miRNAs) from The Cancer Genome Atlas (TCGA) to perform a systematic investigation to explain the different telomere length (TL) features between NSE (n = 48) and SE (n = 55). We found that TL elongation was dominant in NSE, whereas TL shortening prevailed in SE. We further showed that both mRNA and miRNA expression profiles could clearly distinguish these two subtypes. Notably, four telomere-related genes (TelGenes) showed significantly higher expression and positively correlated with telomere elongation in NSE than SE: three telomerase activity-related genes (TERT, WRAP53 and MYC) and an independent telomerase activity gene (ZSCAN4). We also found that the expression of genes encoding Yamanaka factors was positively correlated with telomere lengthening in NSE. Among them, SOX2 and MYC were highly expressed in NSE versus SE, while POU5F1 and KLF4 had the opposite patterns. These results suggested that enhanced expression of both TelGenes (TERT, WRAP53, MYC and ZSCAN4) and Yamanaka factors might induce telomere elongation in NSE. Conversely, the relative lack of telomerase activation and low expression of independent telomerase activity pathway during cell division may be contributed to telomere shortening in SE. Taken together, our results revealed the potential molecular profiles and regulatory roles involving the TL difference between NSE and SE, and provided a better molecular understanding of this complex disease.
Hua Sun 0003, Pora Kim, Peilin Jia, Aekyung Park, Zhongming Zhao
Briefings Bioinform.6
2019 Translational bioinformatics in mental health: open access data sources and computational biomarker discovery
abstract
Mental illness is increasingly recognized as both a significant cost to society and a significant area of opportunity for biological breakthrough. As -omics and imaging technologies enable researchers to probe molecular and physiological underpinnings of multiple diseases, opportunities arise to explore the biological basis for behavioral health and disease. From individual investigators to large international consortia, researchers have generated rich data sets in the area of mental health, including genomic, transcriptomic, metabolomic, proteomic, clinical and imaging resources. General data repositories such as the Gene Expression Omnibus (GEO) and Database of Genotypes and Phenotypes (dbGaP) and mental health (MH)-specific initiatives, such as the Psychiatric Genomics Consortium, MH Research Network and PsychENCODE represent a wealth of information yet to be gleaned. At the same time, novel approaches to integrate and analyze data sets are enabling important discoveries in the area of mental and behavioral health. This review will discuss and catalog into an organizing framework the increasingly diverse set of MH data resources available, using schizophrenia as a focus area, and will describe novel and integrative approaches to molecular biomarker discovery that make use of mental health data.
Jessica D. Tenenbaum, Krithika Bhuvaneshwar, Jane P. Gagliardi, Kate Fultz Hollis, Peilin Jia, Radhakrishnan Nagarajan, Gopalkumar Rakesh, Vignesh Subbian, Shyam Visweswaran, Zhongming Zhao, Leon Rozenblit
Briefings Bioinform.11
2019 CNet: a multi-omics approach to detecting clinically associated, combinatory genomic signatures
abstract
MOTIVATION: Genome-wide multi-omics profiling of complex diseases provides valuable resources and opportunities to discover associations between various measures of genes and diseases. Currently, a pressing challenge is how to effectively detect functional genes associated with or causing phenotypic outcomes. We developed CNet to identify groups of genomic signatures whose combinatory effect is significantly associated with clinical and phenotypical outcomes. RESULTS: CNet builds on a generalized sequential feedforward method, augmented by a down-sampling bootstrap strategy to reduce random hitchhiking signatures. It further applies a dynamic trimming procedure to remove relatively less informative signatures at every step. CNet can manage heterogeneous genomic signature profiles simultaneously and select the best signature to represent a specific gene. To deal with various forms of clinical and phenotypical measurements, we introduced four models to deal with continuous, categorical and censored data. We tested CNet using drug-response data, multidimensional cancer genomics data and genome-wide association study data for multiple traits. Our results demonstrated that in various scenarios, CNet could effectively identify signatures that are associated with the outcomes. In addition, we applied CNet to identify likely disease-causing chains involving somatic mutations, pathway activities and patient outcomes. With appropriate setting, CNet can be applied in many biological conditions. AVAILABILITY AND IMPLEMENTATION: CNet can be downloaded at https://github.com/bsml320/CNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Peilin Jia, Guangsheng Pei, Zhongming Zhao
Bioinform.3
2019 deTS: tissue-specific enrichment analysis to decode tissue specificity
abstract
MOTIVATION: Diseases and traits are under dynamic tissue-specific regulation. However, heterogeneous tissues are often collected in biomedical studies, which reduce the power in the identification of disease-associated variants and gene expression profiles. RESULTS: We present deTS, an R package, to conduct tissue-specific enrichment analysis with two built-in reference panels. Statistical methods are developed and implemented for detecting tissue-specific genes and for enrichment test of different forms of query data. Our applications using multi-trait genome-wide association studies data and cancer expression data showed that deTS could effectively identify the most relevant tissues for each query trait or sample, providing insights for future studies. AVAILABILITY AND IMPLEMENTATION: https://github.com/bsml320/deTS and CRAN https://cran.r-project.org/web/packages/deTS/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guangsheng Pei, Yulin Dai, Zhongming Zhao, Peilin Jia
Bioinform.3
2019 The International Conference on Intelligent Biology and Medicine (ICIBM) 2019: bioinformatics methods and applications for human diseases
abstract
Between June 9-11, 2019, the International Conference on Intelligent Biology and Medicine (ICIBM 2019) was held in Columbus, Ohio, USA. The conference included 12 scientific sessions, five tutorials or workshops, one poster session, four keynote talks and four eminent scholar talks that covered a wide range of topics in bioinformatics, medical informatics, systems biology and intelligent computing. Here, we describe 13 high quality research articles selected for publishing in BMC Bioinformatics.
Zhongming Zhao, Yulin Dai, Ewy A. Mathé, Kai Wang 0063
BMC Bioinform.1
2018 Kinase impact assessment in the landscape of fusion genes that retain kinase domains: a pan-cancer study
abstract
Assessing the impact of kinase in gene fusion is essential for both identifying driver fusion genes (FGs) and developing molecular targeted therapies. Kinase domain retention is a crucial factor in kinase fusion genes (KFGs), but such a systematic investigation has not been done yet. To this end, we analyzed kinase domain retention (KDR) status in chimeric protein sequences of 914 KFGs covering 312 kinases across 13 major cancer types. Based on 171 kinase domain-retained KFGs including 101 kinases, we studied their recurrence, kinase groups, fusion partners, exon-based expression depth, short DNA motifs around the break points and networks. Our results, such as more KDR than 5'-kinase fusion genes, combinatorial effects between 3'-KDR kinases and their 5'-partners and a signal transduction-specific DNA sequence motif in the break point intronic sequences, supported positive selection on 3'-kinase fusion genes in cancer. We introduced a degree-of-frequency (DoF) score to measure the possible number of KFGs of a kinase. Interestingly, kinases with high DoF scores tended to undergo strong gene expression alteration at the break points. Furthermore, our KDR gene fusion network analysis revealed six of the seven kinases with the highest DoF scores (ALK, BRAF, MET, NTRK1, NTRK3 and RET) were all observed in thyroid carcinoma. Finally, we summarized common features of 'effective' (highly recurrent) kinases in gene fusions such as expression alteration at break point, redundant usage in multiple cancer types and 3'-location tendency. Collectively, our findings are useful for prioritizing driver kinases and FGs and provided insights into KFGs' clinical implications.
Pora Kim, Peilin Jia, Zhongming Zhao
Briefings Bioinform.3
2018 The International Conference on Intelligent Biology and Medicine (ICIBM) 2018: bioinformatics towards translational applications
Xiaoming Liu 0021, Lei Xie 0006, Zhijin Wu, Kai Wang 0063, Zhongming Zhao, Jianhua Ruan, Degui Zhi
BMC Bioinform.5
2017 TrapRM: Transcriptomic and proteomic rule mining using weighted shortest distance based multiple minimum supports for multi-omics dataset
abstract
Association rule mining is an important machine learning tool for unveiling critical biological relations between genes from omics data. Previous approaches typically are designed for one single genomic dataset, and most of them use a single minimum support threshold globally. To overcome the above two general limitations, in this work, we present a novel Transcriptomic and Proteomic Rule Mining (TrapRM) method using Weighted Shortest Distance based Multiple Minimum Supports for Multi-Omics Dataset that integrates gene expression, methylation and protein-protein interaction data. To do so, we initially introduce three new thresholds: Weighted Shortest Distance based Multiple Minimum Supports (WSDMS), Weighted Shortest Distance based Multiple Minimum Confidences (WSDMC), and Weighted Shortest Distance based Multiple Minimum Lifts (WSDML). Our algorithm is superior to the related existing algorithms since it generates substantially fewer number of rules and smaller average weighted shortest distance value than the existing methods. Finally, our TrapRM algorithm is useful for extracting the rules that are critical for translational and clinical applications when being applied to drug or disease related multi-omics data.
Saurav Mallik, Zhongming Zhao
BIBM2
2017 Impacts of somatic mutations on gene expression: an association perspective
abstract
Assessing the functional impacts of somatic mutations in cancer genomes is critical for both identifying driver mutations and developing molecular targeted therapies. Currently, it remains a fundamental challenge to distinguish the patterns through which mutations execute their biological effects and to infer biological mechanisms underlying these patterns. To this end, we systematically studied the association between somatic mutations in protein-coding regions and expression profiles, which represents an indirect measurement of impacts. We defined mutation features (mutation type, cluster and status) and built linear regression models to assess mutation associations with mRNA expression and protein expression. Our results presented a comprehensive landscape of the associations between mutation features and expression profile in multiple cancer types, including 62 genes showing mutation type associated expression changes, 21 genes showing mutation cluster associations and 51 genes showing mutation status associations. We revealed four characteristics of the patterns that mutations impact on expression. First, we showed that mutation type (truncation versus amino acid-altering mutations) was the most important determinant of expression levels. Second, we detected mutation clusters in well-studied oncogenes that were associated with gene expression. Third, we found both similarities and differences in association patterns existed within and across cancer types. Fourth, although many of the observed associations stay stable at both mRNA and protein expression levels, there are also novel associations uniquely observed at the protein level, which warrant future investigation. Taken together, our findings provided implications for cancer driver gene prioritization and insights into the functional consequences of somatic mutations.
Peilin Jia, Zhongming Zhao
Briefings Bioinform.2
2017 The International Conference on Intelligent Biology and Medicine (ICIBM) 2016: from big data to big analytical tools
abstract
The 2016 International Conference on Intelligent Biology and Medicine (ICIBM 2016) was held on December 8-10, 2016 in Houston, Texas, USA. ICIBM included eight scientific sessions, four tutorials, one poster session, four highlighted talks and four keynotes that covered topics on 3D genomics structural analysis, next generation sequencing (NGS) analysis, computational drug discovery, medical informatics, cancer genomics, and systems biology. Here, we present a summary of the nine research articles selected from ICIBM 2016 program for publishing in BMC Bioinformatics.
Zhandong Liu, W. Jim Zheng, Genevera I. Allen, Jianhua Ruan, Zhongming Zhao
BMC Bioinform.6
2017 Investigating MicroRNA and transcription factor co-regulatory networks in colorectal cancer
abstract
BACKGROUND: Colorectal cancer (CRC) is one of the most common malignancies worldwide with poor prognosis. Studies have showed that abnormal microRNA (miRNA) expression can affect CRC pathogenesis and development through targeting critical genes in cellular system. However, it is unclear about which miRNAs play central roles in CRC's pathogenesis and how they interact with transcription factors (TFs) to regulate the cancer-related genes. RESULTS: To address this issue, we systematically explored the major regulation motifs, namely feed-forward loops (FFLs), that consist of miRNAs, TFs and CRC-related genes through the construction of a miRNA-TF regulatory network in CRC. First, we compiled CRC-related miRNAs, CRC-related genes, and human TFs from multiple data sources. Second, we identified 13,123 3-node FFLs including 25 miRNA-FFLs, 13,005 TF-FFLs and 93 composite-FFLs, and merged the 3-node FFLs to construct a CRC-related regulatory network. The network consists of three types of regulatory subnetworks (SNWs): miRNA-SNW, TF-SNW, and composite-SNW. To enhance the accuracy of the network, the results were filtered by using The Cancer Genome Atlas (TCGA) expression data in CRC, whereby we generated a core regulatory network consisting of 58 significant FFLs. We then applied a hub identification strategy to the significant FFLs and found 5 significant components, including two miRNAs (hsa-miR-25 and hsa-miR-31), two genes (ADAMTSL3 and AXIN1) and one TF (BRCA1). The follow up prognosis analysis indicated all of the 5 significant components having good prediction of overall survival of CRC patients. CONCLUSIONS: In summary, we generated a CRC-specific miRNA-TF regulatory network, which is helpful to understand the complex CRC regulatory mechanisms and guide clinical treatment. The discovered 5 regulators might have critical roles in CRC pathogenesis and warrant future investigation.
Jiamao Luo, Huilin Niu, Jing Wang 0026, Qi Liu 0024, Zhongming Zhao, Hua Xu 0001, Yanqing Ding, Jingchun Sun, Qingling Zhang 0003
BMC Bioinform.7
2016 Advances in computational approaches for prioritizing driver mutations and significantly mutated genes in cancer genomes
abstract
Cancer is often driven by the accumulation of genetic alterations, including single nucleotide variants, small insertions or deletions, gene fusions, copy-number variations, and large chromosomal rearrangements. Recent advances in next-generation sequencing technologies have helped investigators generate massive amounts of cancer genomic data and catalog somatic mutations in both common and rare cancer types. So far, the somatic mutation landscapes and signatures of >10 major cancer types have been reported; however, pinpointing driver mutations and cancer genes from millions of available cancer somatic mutations remains a monumental challenge. To tackle this important task, many methods and computational tools have been developed during the past several years and, thus, a review of its advances is urgently needed. Here, we first summarize the main features of these methods and tools for whole-exome, whole-genome and whole-transcriptome sequencing data. Then, we discuss major challenges like tumor intra-heterogeneity, tumor sample saturation and functionality of synonymous mutations in cancer, all of which may result in false-positive discoveries. Finally, we highlight new directions in studying regulatory roles of noncoding somatic mutations and quantitatively measuring circulating tumor DNA in cancer. This review may help investigators find an appropriate tool for detecting potential driver or actionable mutations in rapidly emerging precision cancer medicine.
Feixiong Cheng, Junfei Zhao, Zhongming Zhao
Briefings Bioinform.3
2016 Systematic dissection of dysregulated transcription factor-miRNA feed-forward loops across tumor types
abstract
Transcription factor and microRNA (miRNA) can mutually regulate each other and jointly regulate their shared target genes to form feed-forward loops (FFLs). While there are many studies of dysregulated FFLs in a specific cancer, a systematic investigation of dysregulated FFLs across multiple tumor types (pan-cancer FFLs) has not been performed yet. In this study, using The Cancer Genome Atlas data, we identified 26 pan-cancer FFLs, which were dysregulated in at least five tumor types. These pan-cancer FFLs could communicate with each other and form functionally consistent subnetworks, such as epithelial to mesenchymal transition-related subnetwork. Many proteins and miRNAs in each subnetwork belong to the same protein and miRNA family, respectively. Importantly, cancer-associated genes and drug targets were enriched in these pan-cancer FFLs, in which the genes and miRNAs also tended to be hubs and bottlenecks. Finally, we identified potential anticancer indications for existing drugs with novel mechanism of action. Collectively, this study highlights the potential of pan-cancer FFLs as a novel paradigm in elucidating pathogenesis of cancer and developing anticancer drugs.
Wei Jiang 0023, Ramkrishna Mitra, Chen-Ching Lin, Quan Wang 0004, Feixiong Cheng, Zhongming Zhao
Briefings Bioinform.6
2016 A network-based drug repositioning infrastructure for precision cancer medicine through targeting significantly mutated genes in the human cancer genomes
abstract
OBJECTIVE: Development of computational approaches and tools to effectively integrate multidomain data is urgently needed for the development of newly targeted cancer therapeutics. METHODS: We proposed an integrative network-based infrastructure to identify new druggable targets and anticancer indications for existing drugs through targeting significantly mutated genes (SMGs) discovered in the human cancer genomes. The underlying assumption is that a drug would have a high potential for anticancer indication if its up-/down-regulated genes from the Connectivity Map tended to be SMGs or their neighbors in the human protein interaction network. RESULTS: We assembled and curated 693 SMGs in 29 cancer types and found 121 proteins currently targeted by known anticancer or noncancer (repurposed) drugs. We found that the approved or experimental cancer drugs could potentially target these SMGs in 33.3% of the mutated cancer samples, and this number increased to 68.0% by drug repositioning through surveying exome-sequencing data in approximately 5000 normal-tumor pairs from The Cancer Genome Atlas. Furthermore, we identified 284 potential new indications connecting 28 cancer types and 48 existing drugs (adjusted P < .05), with a 66.7% success rate validated by literature data. Several existing drugs (e.g., niclosamide, valproic acid, captopril, and resveratrol) were predicted to have potential indications for multiple cancer types. Finally, we used integrative analysis to showcase a potential mechanism-of-action for resveratrol in breast and lung cancer treatment whereby it targets several SMGs (ARNTL, ASPM, CTTN, EIF4G1, FOXP1, and STIP1). CONCLUSIONS: In summary, we demonstrated that our integrative network-based infrastructure is a promising strategy to identify potential druggable targets and uncover new indications for existing drugs to speed up molecularly targeted cancer therapeutics.
Feixiong Cheng, Junfei Zhao, Michaela Fooksa, Zhongming Zhao
J. Am. Medical Informatics Assoc.4
2016 An informatics research agenda to support precision medicine: seven key areas
abstract
The recent announcement of the Precision Medicine Initiative by President Obama has brought precision medicine (PM) to the forefront for healthcare providers, researchers, regulators, innovators, and funders alike. As technologies continue to evolve and datasets grow in magnitude, a strong computational infrastructure will be essential to realize PM's vision of improved healthcare derived from personal data. In addition, informatics research and innovation affords a tremendous opportunity to drive the science underlying PM. The informatics community must lead the development of technologies and methodologies that will increase the discovery and application of biomedical knowledge through close collaboration between researchers, clinicians, and patients. This perspective highlights seven key areas that are in need of further informatics research and innovation to support the realization of PM.
Jessica D. Tenenbaum, Paul Avillach, Marge M. Benham-Hutchins, Matthew K. Breitenstein, Erin L. Crowgey, Mark A. Hoffman, Xia Jiang, Subha Madhavan, John E. Mattison, Radhakrishnan Nagarajan, Bisakha Ray, Dmitriy Shin, Shyam Visweswaran, Zhongming Zhao, Robert R. Freimuth
J. Am. Medical Informatics Assoc.14
2016 Systems Biology-Based Investigation of Cellular Antiviral Drug Targets Identified by Gene-Trap Insertional Mutagenesis
abstract
Viruses require host cellular factors for successful replication. A comprehensive systems-level investigation of the virus-host interactome is critical for understanding the roles of host factors with the end goal of discovering new druggable antiviral targets. Gene-trap insertional mutagenesis is a high-throughput forward genetics approach to randomly disrupt (trap) host genes and discover host genes that are essential for viral replication, but not for host cell survival. In this study, we used libraries of randomly mutagenized cells to discover cellular genes that are essential for the replication of 10 distinct cytotoxic mammalian viruses, 1 gram-negative bacterium, and 5 toxins. We herein reported 712 candidate cellular genes, characterizing distinct topological network and evolutionary signatures, and occupying central hubs in the human interactome. Cell cycle phase-specific network analysis showed that host cell cycle programs played critical roles during viral replication (e.g. MYC and TAF4 regulating G0/1 phase). Moreover, the viral perturbation of host cellular networks reflected disease etiology in that host genes (e.g. CTCF, RHOA, and CDKN1B) identified were frequently essential and significantly associated with Mendelian and orphan diseases, or somatic mutations in cancer. Computational drug repositioning framework via incorporating drug-gene signatures from the Connectivity Map into the virus-host interactome identified 110 putative druggable antiviral targets and prioritized several existing drugs (e.g. ajmaline) that may be potential for antiviral indication (e.g. anti-Ebola). In summary, this work provides a powerful methodology with a tight integration of gene-trap insertional mutagenesis testing and systems biology to identify new antiviral targets and drugs for the development of broadly acting and targeted clinical antiviral therapeutics.
Feixiong Cheng, James L. Murray, Junfei Zhao, Jinsong Sheng, Zhongming Zhao, Donald H. Rubin
PLoS Comput. Biol.5
2015 EW_dmGWAS: edge-weighted dense module search for genome-wide association studies and gene expression profiles
abstract
Abstract Summary: We previously developed dmGWAS to search for dense modules in a human protein–protein interaction (PPI) network; it has since become a popular tool for network-assisted analysis of genome-wide association studies (GWAS). dmGWAS weights nodes by using GWAS signals. Here, we introduce an upgraded algorithm, EW_dmGWAS, to boost GWAS signals in a node- and edge-weighted PPI network. In EW_dmGWAS, we utilize condition-specific gene expression profiles for edge weights. Specifically, differential gene co-expression is used to infer the edge weights. We applied EW_dmGWAS to two diseases and compared it with other relevant methods. The results suggest that EW_dmGWAS is more powerful in detecting disease-associated signals. Availability and implementation: The algorithm of EW_dmGWAS is implemented in the R package dmGWAS_3.0 and is available at http://bioinfo.mc.vanderbilt.edu/dmGWAS. Contact: [email protected] or [email protected] Supplementary information: Supplementary materials are available at Bioinformatics online.
Quan Wang 0004, Zhongming Zhao, Peilin Jia
Bioinform.3
2015 Snowball: resampling combined with distance-based regression to discover transcriptional consequences of a driver mutation
abstract
MOTIVATION: Large-scale cancer genomic studies, such as The Cancer Genome Atlas (TCGA), have profiled multidimensional genomic data, including mutation and expression profiles on a variety of cancer cell types, to uncover the molecular mechanism of cancerogenesis. More than a hundred driver mutations have been characterized that confer the advantage of cell growth. However, how driver mutations regulate the transcriptome to affect cellular functions remains largely unexplored. Differential analysis of gene expression relative to a driver mutation on patient samples could provide us with new insights in understanding driver mutation dysregulation in tumor genome and developing personalized treatment strategies. RESULTS: Here, we introduce the Snowball approach as a highly sensitive statistical analysis method to identify transcriptional signatures that are affected by a recurrent driver mutation. Snowball utilizes a resampling-based approach and combines a distance-based regression framework to assign a robust ranking index of genes based on their aggregated association with the presence of the mutation, and further selects the top significant genes for downstream data analyses or experiments. In our application of the Snowball approach to both synthesized and TCGA data, we demonstrated that it outperforms the standard methods and provides more accurate inferences to the functional effects and transcriptional dysregulation of driver mutations. AVAILABILITY AND IMPLEMENTATION: R package and source code are available from CRAN at http://cran.r-project.org/web/packages/DESnowball, and also available at http://bioinfo.mc.vanderbilt.edu/DESnowball/.
Yaomin Xu, Xingyi Guo, Zhongming Zhao
Bioinform.4
2015 SGDriver: a novel structural genomics-based approach to prioritize cancer related and potentially druggable somatic mutations
abstract
Background A huge volume of somatic mutations have been generated through large cancer genome sequencing projects such as The Cancer Genome Atlas (TCGA) and the International Cancer Genome Consortium (ICGC). However, understanding the functional consequences of somatic mutations in cancer and translating the results into clinical use remains a major challenge in cancer genomic studies. Thanks to the rapid development of structural genomic technologies, such as X-ray and NMR, large amounts of protein structure data have been generated during the past decade, which enables us to map somatic mutations to protein functional features (i.e., protein-ligand binding sites) and investigate their potential impacts[1,2].
Junfei Zhao, Feixiong Cheng, Zhongming Zhao
BMC Bioinform.3
2015 A Gene Gravity Model for the Evolution of Cancer Genomes: A Study of 3, 000 Cancer Genomes across 9 Cancer Types
abstract
Cancer development and progression result from somatic evolution by an accumulation of genomic alterations. The effects of those alterations on the fitness of somatic cells lead to evolutionary adaptations such as increased cell proliferation, angiogenesis, and altered anticancer drug responses. However, there are few general mathematical models to quantitatively examine how perturbations of a single gene shape subsequent evolution of the cancer genome. In this study, we proposed the gene gravity model to study the evolution of cancer genomes by incorporating the genome-wide transcription and somatic mutation profiles of ~3,000 tumors across 9 cancer types from The Cancer Genome Atlas into a broad gene network. We found that somatic mutations of a cancer driver gene may drive cancer genome evolution by inducing mutations in other genes. This functional consequence is often generated by the combined effect of genetic and epigenetic (e.g., chromatin regulation) alterations. By quantifying cancer genome evolution using the gene gravity model, we identified six putative cancer genes (AHNAK, COL11A1, DDX3X, FAT4, STAG2, and SYNE1). The tumor genomes harboring the nonsynonymous somatic mutations in these genes had a higher mutation density at the genome level compared to the wild-type groups. Furthermore, we provided statistical evidence that hypermutation of cancer driver genes on inactive X chromosomes is a general feature in female cancer genomes. In summary, this study sheds light on the functional consequences and evolutionary characteristics of somatic mutations during tumorigenesis by propelling adaptive cancer genome evolution, which would provide new perspectives for cancer research and therapeutics.
Feixiong Cheng, Chen-Ching Lin, Junfei Zhao, Peilin Jia, Wen-Hsiung Li, Zhongming Zhao
PLoS Comput. Biol.7
2015 Deciphering Signaling Pathway Networks to Understand the Molecular Mechanisms of Metformin Action
abstract
A drug exerts its effects typically through a signal transduction cascade, which is non-linear and involves intertwined networks of multiple signaling pathways. Construction of such a signaling pathway network (SPNetwork) can enable identification of novel drug targets and deep understanding of drug action. However, it is challenging to synopsize critical components of these interwoven pathways into one network. To tackle this issue, we developed a novel computational framework, the Drug-specific Signaling Pathway Network (DSPathNet). The DSPathNet amalgamates the prior drug knowledge and drug-induced gene expression via random walk algorithms. Using the drug metformin, we illustrated this framework and obtained one metformin-specific SPNetwork containing 477 nodes and 1,366 edges. To evaluate this network, we performed the gene set enrichment analysis using the disease genes of type 2 diabetes (T2D) and cancer, one T2D genome-wide association study (GWAS) dataset, three cancer GWAS datasets, and one GWAS dataset of cancer patients with T2D on metformin. The results showed that the metformin network was significantly enriched with disease genes for both T2D and cancer, and that the network also included genes that may be associated with metformin-associated cancer survival. Furthermore, from the metformin SPNetwork and common genes to T2D and cancer, we generated a subnetwork to highlight the molecule crosstalk between T2D and cancer. The follow-up network analyses and literature mining revealed that seven genes (CDKN1A, ESR1, MAX, MYC, PPARGC1A, SP1, and STK11) and one novel MYC-centered pathway with CDKN1A, SP1, and STK11 might play important roles in metformin's antidiabetic and anticancer effects. Some results are supported by previous studies. In summary, our study 1) develops a novel framework to construct drug-specific signal transduction networks; 2) provides insights into the molecular mode of metformin; 3) serves a model for exploring signaling pathways to facilitate understanding of drug action, disease pathogenesis, and identification of drug targets.
Jingchun Sun, Min Zhao 0006, Peilin Jia, Lily Wang 0001, Yonghui Wu 0001, Carissa Iverson, Yubo Zhou, Erica A. Bowton, Dan M. Roden, Joshua C. Denny, Melinda Aldrich, Hua Xu 0001, Zhongming Zhao
PLoS Comput. Biol.13
2014 Evaluating four major algorithms for identifying differential regulators in condition-specific transcriptional responses
abstract
BackgroundIdentifying molecular regulators underlying conditionspecific transcriptional responses is essential for our understanding of their underlying molecular mechanisms.So far, there have been several computational methods developed for this purpose.Specifically, four major algorithms, TFact [1], RIF [2], CSA [3], and DRrank [4], were released one after another.Each of these algorithms has its own features and evaluation strategies.Thus, these methods should be systematically evaluated so that the users can make the most appropriate method selection for their practical application needs. Materials and methodsIn this work, we evaluated the four algorithms using Escherichia coli transcription network models and synthetic expression datasets that were generated by and GeneNetWeaver [6].Specifically, we developed a simulation-based schema to evaluate each algorithm according to operatively defined, known "differential regulators."In addition, we tested each method's robustness against its key parameter(s) and explored factors that influence algorithm performance in general.
Zhongming Zhao
BMC Bioinform.2
2014 VarWalker: Personalized Mutation Network Analysis of Putative Cancer Genes from Next-Generation Sequencing Data
abstract
A major challenge in interpreting the large volume of mutation data identified by next-generation sequencing (NGS) is to distinguish driver mutations from neutral passenger mutations to facilitate the identification of targetable genes and new drugs. Current approaches are primarily based on mutation frequencies of single-genes, which lack the power to detect infrequently mutated driver genes and ignore functional interconnection and regulation among cancer genes. We propose a novel mutation network method, VarWalker, to prioritize driver genes in large scale cancer mutation data. VarWalker fits generalized additive models for each sample based on sample-specific mutation profiles and builds on the joint frequency of both mutation genes and their close interactors. These interactors are selected and optimized using the Random Walk with Restart algorithm in a protein-protein interaction network. We applied the method in >300 tumor genomes in two large-scale NGS benchmark datasets: 183 lung adenocarcinoma samples and 121 melanoma samples. In each cancer, we derived a consensus mutation subnetwork containing significantly enriched consensus cancer genes and cancer-related functional pathways. These cancer-specific mutation networks were then validated using independent datasets for each cancer. Importantly, VarWalker prioritizes well-known, infrequently mutated genes, which are shown to interact with highly recurrently mutated genes yet have been ignored by conventional single-gene-based approaches. Utilizing VarWalker, we demonstrated that network-assisted approaches can be effectively adapted to facilitate the detection of cancer driver genes in NGS data.
Peilin Jia, Zhongming Zhao
PLoS Comput. Biol.2
2013 Network-based mutation analysis of putative cancer genes from next-generation sequencing data
abstract
Next-generation sequencing (NGS) has enabled fast detection of somatic mutations in cancer genomes. A major challenge in interpreting the large volume of mutation data is to distinguish driver mutations from neutral passenger mutations. Current approaches are primarily single-gene based prioritization according to mutation frequencies, which harbors both high false positive and false negative discoveries. We propose a novel network-based method of mutation data for driver gene prioritization from large scale mutation data for cancer. Our method takes into consideration of the mutation profile of each patient by fitting sample-specific generalized additive models. It builds on joint frequency of both mutation genes and their close interactors, which are optimized by the algorithm Random Walk with Restart in a protein-protein interaction network. We demonstrated our method in two large-scale NGS datasets: a lung adenocarcinoma (LUAD) dataset including 183 patients and a melanoma dataset including 121 samples. In each cancer, we derived a consensus mutation subnetwork with significantly enriched consensus cancer genes and cancer-related functional pathways. The LUAD subnetwork recruited 70 genes of the Cancer Gene Census (CGC) collection (p-value <; 2.2×10-16, Fisher's Exact Test) and the melanoma subnetwork included 65 CGC genes (p-value <; 2.2x10-16). In addition, our results indicate that some well-known, infrequently mutated genes, which have been ignored by conventional single-gene based approaches, are also prioritized and are shown to interact with those highly recurrently mutated genes. In sum, our method is effective in prioritizing candidate driver genes from more than ten thousand mutation genes and provides biological interpretations for future work.
Peilin Jia, Zhongming Zhao
BIBM2
2013 Correlating adverse drug reactions with biological pathways in humans
abstract
It has been well recognized that adverse drug reactions (ADRs) are a significant cause of morbidity and mortality. There is a growing interest in investigating biological pathways involved in cellular response to drugs. Based on examining the co-occurrence of drugs in pathway activity and ADR profiles, in this paper, we propose a new method to explore the relationship between biological pathways and ADRs at a large scale. Using sparse canonical correlation analysis of 495 drugs with two profiles for 173 pathways and 1385 ADRs, a total of 80 correlated sets of pathways and ADRs were extracted. To evaluate the performance of our method, extracted correlated components were used to retrieve known ADR profiles from drug pathway profiles using a 5-fold cross validation. A relatively high prediction performance (AUC: 0.881) was achieved. This work provides a foundation for future investigation of ADRs in the context of biological pathways under different conditions.
Huiru Zheng, Haiying Wang 0001, Hua Xu 0001, Zhongming Zhao, Francisco Azuaje
BIBM4
2013 Application of next generation sequencing to human gene fusion detection: computational tools, features and perspectives
abstract
Gene fusions are important genomic events in human cancer because their fusion gene products can drive the development of cancer and thus are potential prognostic tools or therapeutic targets in anti-cancer treatment. Major advancements have been made in computational approaches for fusion gene discovery over the past 3 years due to improvements and widespread applications of high-throughput next generation sequencing (NGS) technologies. To identify fusions from NGS data, existing methods typically leverage the strengths of both sequencing technologies and computational strategies. In this article, we review the NGS and computational features of existing methods for fusion gene detection and suggest directions for future development.
Qingguo Wang, Junfeng Xia, Peilin Jia, William Pao, Zhongming Zhao
Briefings Bioinform.5
2013 Network analysis of gene fusions in human cancer
abstract
Background Gene fusions are hybrid genes formed when two discrete genes are incorrectly joined together. Gene fusions are found to play roles in tumorigenesis. For example, the fusion gene BCR-ABL translates into an abnormal tyrosine kinase that accelerates development of chronic myelogenous leukemia [1]. A network is a relational representation of nodes (e.g., genes) with edges, and is a useful approach to explore biological interactions among many related nodes. Network analysis of gene fusions in cancer would aid the exploration of gene fusion occurrence and association with tumorigenesis. Hoglund et al [2] performed an initial investigation of gene fusions network after collecting 291 tumorigenesis related gene fusions from the Mitelman database in 2006. Since then, gene fusion data has exponentially increased. There is no current and comprehensive cancer-related gene fusion network to assist in targeting cancer-associated genes.
Morgan Harrell, Junfeng Xia, Zhongming Zhao
BMC Bioinform.3
2013 Identifying transcription factor and microRNA mediated synergetic regulatory networks in lung cancer
abstract
Background It has been demonstrated that, at the network level, the transcriptional regulation by transcription factors (TFs) and post-transcriptional regulation by microRNAs (miRNAs) are tightly coupled. Aberrant expression of these bio-molecules is linked to several diseases, including lung cancer. In this study, we pursued a regulatory networkbased approach mediated by TFs and miRNAs for a comprehensive investigation of gene regulation patterns in lung cancer.
Ramkrishna Mitra, Jingchun Sun, Min Zhao 0006, Zhongming Zhao
BMC Bioinform.4
2013 A short tutorial in analyzing NGS data of cancer genomes for somatic mutation calling
abstract
Background Somatic mutation is the key element of tumorigenesis as these changes in nucleotide sequence of the cancer genome in somatic cells acquired throughout life can lead to protein alteration, cellular damage and thus cause cancer. The advent of next generation sequencing has significantly improved our ability to identify somatic mutations in cancer genomes paving the way for the comprehensive online catalogue for somatic mutation in human cancer (COSMIC) 1 which contains more than 820,000 mutations so far. Nevertheless, there are still many challenges in detecting somatic mutation in cancer especially for low frequency mutation due to either tumor heterogeneity or contamination with normal cells. Here, in this short tutorial, I will present a recent somatic mutation caller tool developed by the Broad Institute called Mutect 2 as part of the GATK (Genome Analysis Toolkit). I will use my own NGS dataset to demonstrate the tools and address some issues of troubleshooting input data and interpreting output.
Huy Vuong, Zhongming Zhao
BMC Bioinform.2
2013 Differential coexpression network modules observed in human hepatocellular carcinoma progression
abstract
Background While an understanding of the human interactome is within attainable reach [1], an impending challenge is to uncover the condition-specific dynamics of the proteinprotein interaction (PPI) network, especially those that coordinate with disease progression [2]. Differential coexpression analysis (DCA) [3,4] has recently emerged as an effective approach to address this issue, but such an effort has yet to be thoroughly tested.
Zhongming Zhao
BMC Bioinform.2
2013 Computational tools for copy number variation (CNV) detection using next-generation sequencing data: features and perspectives
abstract
Copy number variation (CNV) is a prevalent form of critical genetic variation that leads to an abnormal number of copies of large genomic regions in a cell. Microarray-based comparative genome hybridization (arrayCGH) or genotyping arrays have been standard technologies to detect large regions subject to copy number changes in genomes until most recently high-resolution sequence data can be analyzed by next-generation sequencing (NGS). During the last several years, NGS-based analysis has been widely applied to identify CNVs in both healthy and diseased individuals. Correspondingly, the strong demand for NGS-based CNV analyses has fuelled development of numerous computational methods and tools for CNV detection. In this article, we review the recent advances in computational methods pertaining to CNV detection using whole genome and whole exome sequencing data. Additionally, we discuss their strengths and weaknesses and suggest directions for future development.
Min Zhao 0006, Qingguo Wang, Quan Wang 0004, Peilin Jia, Zhongming Zhao
BMC Bioinform.5
2012 DTome: a web-based tool for drug-target interactome construction
abstract
BACKGROUND: Understanding drug bioactivities is crucial for early-stage drug discovery, toxicology studies and clinical trials. Network pharmacology is a promising approach to better understand the molecular mechanisms of drug bioactivities. With a dramatic increase of rich data sources that document drugs' structural, chemical, and biological activities, it is necessary to develop an automated tool to construct a drug-target network for candidate drugs, thus facilitating the drug discovery process. RESULTS: We designed a computational workflow to construct drug-target networks from different knowledge bases including DrugBank, PharmGKB, and the PINA database. To automatically implement the workflow, we created a web-based tool called DTome (Drug-Target interactome tool), which is comprised of a database schema and a user-friendly web interface. The DTome tool utilizes web-based queries to search candidate drugs and then construct a DTome network by extracting and integrating four types of interactions. The four types are adverse drug interactions, drug-target interactions, drug-gene associations, and target-/gene-protein interactions. Additionally, we provided a detailed network analysis and visualization process to illustrate how to analyze and interpret the DTome network. The DTome tool is publicly available at http://bioinfo.mc.vanderbilt.edu/DTome. CONCLUSIONS: As demonstrated with the antipsychotic drug clozapine, the DTome tool was effective and promising for the investigation of relationships among drugs, adverse interaction drugs, drug primary targets, drug-associated genes, and proteins directly interacting with targets or genes. The resultant DTome network provides researchers with direct insights into their interest drug(s), such as the molecular mechanisms of drug actions. We believe such a tool can facilitate identification of drug targets and drug adverse interactions.
Jingchun Sun, Yonghui Wu 0001, Hua Xu 0001, Zhongming Zhao
BMC Bioinform.4
2012 Large-scale prediction of adverse drug reactions using chemical, biological, and phenotypic properties of drugs
abstract
OBJECTIVE: Adverse drug reaction (ADR) is one of the major causes of failure in drug development. Severe ADRs that go undetected until the post-marketing phase of a drug often lead to patient morbidity. Accurate prediction of potential ADRs is required in the entire life cycle of a drug, including early stages of drug design, different phases of clinical trials, and post-marketing surveillance. METHODS: Many studies have utilized either chemical structures or molecular pathways of the drugs to predict ADRs. Here, the authors propose a machine-learning-based approach for ADR prediction by integrating the phenotypic characteristics of a drug, including indications and other known ADRs, with the drug's chemical structures and biological properties, including protein targets and pathway information. A large-scale study was conducted to predict 1385 known ADRs of 832 approved drugs, and five machine-learning algorithms for this task were compared. RESULTS: This evaluation, based on a fivefold cross-validation, showed that the support vector machine algorithm outperformed the others. Of the three types of information, phenotypic data were the most informative for ADR prediction. When biological and phenotypic features were added to the baseline chemical information, the ADR prediction model achieved significant improvements in area under the curve (from 0.9054 to 0.9524), precision (from 43.37% to 66.17%), and recall (from 49.25% to 63.06%). Most importantly, the proposed model successfully predicted the ADRs associated with withdrawal of rofecoxib and cerivastatin. CONCLUSION: The results suggest that phenotypic information on drugs is valuable for ADR prediction. Moreover, they demonstrate that different models that combine chemical, biological, or phenotypic information can be built from approved drugs, and they have the potential to detect clinically important ADRs in both preclinical and post-marketing phases.
Yonghui Wu 0001, Yukun Chen 0001, Jingchun Sun, Zhongming Zhao, Xue-wen Chen 0001, Michael E. Matheny, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2012 Network-Assisted Investigation of Combined Causal Signals from Genome-Wide Association Studies in Schizophrenia
abstract
With the recent success of genome-wide association studies (GWAS), a wealth of association data has been accomplished for more than 200 complex diseases/traits, proposing a strong demand for data integration and interpretation. A combinatory analysis of multiple GWAS datasets, or an integrative analysis of GWAS data and other high-throughput data, has been particularly promising. In this study, we proposed an integrative analysis framework of multiple GWAS datasets by overlaying association signals onto the protein-protein interaction network, and demonstrated it using schizophrenia datasets. Building on a dense module search algorithm, we first searched for significantly enriched subnetworks for schizophrenia in each single GWAS dataset and then implemented a discovery-evaluation strategy to identify module genes with consistent association signals. We validated the module genes in an independent dataset, and also examined them through meta-analysis of the related SNPs using multiple GWAS datasets. As a result, we identified 205 module genes with a joint effect significantly associated with schizophrenia; these module genes included a number of well-studied candidate genes such as DISC1, GNA12, GNA13, GNAI1, GPR17, and GRIN2B. Further functional analysis suggested these genes are involved in neuronal related processes. Additionally, meta-analysis found that 18 SNPs in 9 module genes had P(meta)<1 × 10⁻⁴, including the gene HLA-DQA1 located in the MHC region on chromosome 6, which was reported in previous studies using the largest cohort of schizophrenia patients to date. These results demonstrated our bi-directional network-based strategy is efficient for identifying disease-associated genes with modest signals in GWAS datasets. This approach can be applied to any other complex diseases/traits where multiple GWAS datasets are available.
Peilin Jia, Lily Wang 0001, Ayman H. Fanous, Carlos N. Pato, Todd L. Edwards, Zhongming Zhao
PLoS Comput. Biol.6
2012 Uncovering MicroRNA and Transcription Factor Mediated Regulatory Networks in Glioblastoma
abstract
Glioblastoma multiforme (GBM) is the most common and lethal brain tumor in humans. Recent studies revealed that patterns of microRNA (miRNA) expression in GBM tissue samples are different from those in normal brain tissues, suggesting that a number of miRNAs play critical roles in the pathogenesis of GBM. However, little is yet known about which miRNAs play central roles in the pathology of GBM and their regulatory mechanisms of action. To address this issue, in this study, we systematically explored the main regulation format (feed-forward loops, FFLs) consisting of miRNAs, transcription factors (TFs) and their impacting GBM-related genes, and developed a computational approach to construct a miRNA-TF regulatory network. First, we compiled GBM-related miRNAs, GBM-related genes, and known human TFs. We then identified 1,128 3-node FFLs and 805 4-node FFLs with statistical significance. By merging these FFLs together, we constructed a comprehensive GBM-specific miRNA-TF mediated regulatory network. Then, from the network, we extracted a composite GBM-specific regulatory network. To illustrate the GBM-specific regulatory network is promising for identification of critical miRNA components, we specifically examined a Notch signaling pathway subnetwork. Our follow up topological and functional analyses of the subnetwork revealed that six miRNAs (miR-124, miR-137, miR-219-5p, miR-34a, miR-9, and miR-92b) might play important roles in GBM, including some results that are supported by previous studies. In this study, we have developed a computational framework to construct a miRNA-TF regulatory network and generated the first miRNA-TF regulatory network for GBM, providing a valuable resource for further understanding the complex regulatory mechanisms in GBM. The observation of critical miRNAs in the Notch signaling pathway, with partial verification from previous studies, demonstrates that our network-based approach is promising for the identification of new and important miRNAs in GBM and, potentially, other cancers.
Jingchun Sun, Benjamin Purow, Zhongming Zhao
PLoS Comput. Biol.4
2011 Do MicroRNAs Preferentially Target the Genes with Low DNA Methylation Level at the Promoter Region?
Zhixi Su, Junfeng Xia, Zhongming Zhao
ICIC (3)3
2011 dmGWAS: dense module searching for genome-wide association studies in protein-protein interaction networks
abstract
MOTIVATION: An important question that has emerged from the recent success of genome-wide association studies (GWAS) is how to detect genetic signals beyond single markers/genes in order to explore their combined effects on mediating complex diseases and traits. Integrative testing of GWAS association data with that from prior-knowledge databases and proteome studies has recently gained attention. These methodologies may hold promise for comprehensively examining the interactions between genes underlying the pathogenesis of complex diseases. METHODS: Here, we present a dense module searching (DMS) method to identify candidate subnetworks or genes for complex diseases by integrating the association signal from GWAS datasets into the human protein-protein interaction (PPI) network. The DMS method extensively searches for subnetworks enriched with low P-value genes in GWAS datasets. Compared with pathway-based approaches, this method introduces flexibility in defining a gene set and can effectively utilize local PPI information. RESULTS: We implemented the DMS method in an R package, which can also evaluate and graphically represent the results. We demonstrated DMS in two GWAS datasets for complex diseases, i.e. breast cancer and pancreatic cancer. For each disease, the DMS method successfully identified a set of significant modules and candidate genes, including some well-studied genes not detected in the single-marker analysis of GWA studies. Functional enrichment analysis and comparison with previously published methods showed that the genes we identified by DMS have higher association signal. AVAILABILITY: dmGWAS package and documents are available at http://bioinfo.mc.vanderbilt.edu/dmGWAS.html.
Peilin Jia, Siyuan Zheng, Jirong Long, Zhongming Zhao
Bioinform.5
2011 An efficient hierarchical generalized linear mixed model for pathway analysis of genome-wide association studies
abstract
MOTIVATION: In genome-wide association studies (GWAS) of complex diseases, genetic variants having real but weak associations often fail to be detected at the stringent genome-wide significance level. Pathway analysis, which tests disease association with combined association signals from a group of variants in the same pathway, has become increasingly popular. However, because of the complexities in genetic data and the large sample sizes in typical GWAS, pathway analysis remains to be challenging. We propose a new statistical model for pathway analysis of GWAS. This model includes a fixed effects component that models mean disease association for a group of genes, and a random effects component that models how each gene's association with disease varies about the gene group mean, thus belongs to the class of mixed effects models. RESULTS: The proposed model is computationally efficient and uses only summary statistics. In addition, it corrects for the presence of overlapping genes and linkage disequilibrium (LD). Via simulated and real GWAS data, we showed our model improved power over currently available pathway analysis methods while preserving type I error rate. Furthermore, using the WTCCC Type 1 Diabetes (T1D) dataset, we demonstrated mixed model analysis identified meaningful biological processes that agreed well with previous reports on T1D. Therefore, the proposed methodology provides an efficient statistical modeling framework for systems analysis of GWAS. AVAILABILITY: The software code for mixed models analysis is freely available at http://biostat.mc.vanderbilt.edu/LilyWang.
Lily Wang 0001, Peilin Jia, Russell D. Wolfinger, Britney L. Grayson, Thomas M. Aune, Zhongming Zhao
Bioinform.7
2011 Identifying the key genes and pathways in the progression of hepatitis C virus induced hepatocellular carcinoma using a systems biology approach
abstract
Differentially expressed genes (DEGs) were identified as genes with up or down regulation fold change ≥ 2 and student t test P value ≤ 0.01. †Hub interaction number refers to the total number of interactions involving hub genes. ‡Hub genes were defined to have at least 5 interactions in each network.
Siyuan Zheng, Zhongming Zhao
BMC Bioinform.2
2010 Genome-Wide DNA Methylation Profiling in 40 Breast Cancer Cell Lines
Leng Han, Siyuan Zheng, Shuying Sun, Tim Hui-Ming Huang, Zhongming Zhao
ICIC (1)5
2010 Gene- and evidence-based candidate gene selection for schizophrenia and gene feature analysis
Jingchun Sun, Leng Han, Zhongming Zhao
Artif. Intell. Medicine3
2010 Application of Pearson correlation coefficient (PCC) and Kolmogorov-Smirnov distance (KSD) metrics to identify disease-specific biomarker genes
abstract
Background DNA microarrays have been widely applied in cancer research for better diagnosis and prediction of the disease states. Traditionally, most microarray studies aim to identify differentially expressed genes (DEGs) by comparing the average gene expression levels between two groups (e.g., the treated vs. control or disease vs. non-disease) based on statistical analysis such as t-test and Significance Analysis of Microarrays (SAM) [1,2]. Materials and methods In this study, we defined the gene expression profile (GEP) of a gene as the distribution of the log2 values of its normalized expression signal intensities across the samples in the similarly studied microarrays. We hypothesized that the biomarker genes that distinguish disease samples from normal samples might form distinct GEPs between comparison groups. We applied Pearson Correlation Coefficient (PCC) and Kolmogorov-Smirnov Distance (KSD) metrics to identify disease-specific biomarkers by comparing GEPs between normal and disease states and then applied this technology to disease (e.g., cancer) related studies in order to discover some disease genes as biomarker candidates. These biomarkers’ gene profiles in normal and disease samples might be used to diagnose or monitor patient’s disease state via regular gene expression analysis.
Hung-Chung Huang, Siyuan Zheng, Zhongming Zhao
BMC Bioinform.3
2010 Pathway- and network-based analysis of GWAS data revealed susceptibility gene sets to schizophrenia
abstract
Materials and methods In this study, we uniquely examined GWAS data at the the gene set level (i.e., pathways and protein-protein interaction (PPI) subnetworks) rather than the SNP level. We first collected a comprehensive list of pathways from the KEGG and BioCarta databases. Additionally, we applied a network module searching approach to search for informative subnetworks in a nodeweighted human PPI network by GWAS markers’ P values. Using this module searching method, we identified ~100 candidate genes from top ranked schizophrenia-specific modules and used these genes as a gene set for follow up functional enrichment test. We applied two statistical methods (ge ne set enrichment analysis (GSEA) and hypergeometric test) and also combined them by Fisher’s combined method to analyze the gene sets defined by canonical pathways or our module genes. Results were further validated by permutation analysis. Results and conclusion In GSEA analysis, we found the gene set consisting of the network module genes was the most significant among all gene sets, indicating our module searching being efficient in finding the real effect of multiple genes in am ore flexible way than classically defined pathways. We also identified a few pathways that are consistently associated with schizophrenia by multiple methods; they included glutamate metabolism pathway, TNFR1 pathway, and TGF beta signaling pathway. These results not only improved our understanding of the underlying pathogenesis of schizophrenia, but also suggested that the gene set based approach is powerful to detect common variants conferring risk to complex diseases.
Peilin Jia, Zhongming Zhao
BMC Bioinform.2
2009 A multi-dimensional evidence-based candidate gene prioritization approach for complex diseases-schizophrenia as a case
abstract
MOTIVATION: During the past decade, we have seen an exponential growth of vast amounts of genetic data generated for complex disease studies. Currently, across a variety of complex biological problems, there is a strong trend towards the integration of data from multiple sources. So far, candidate gene prioritization approaches have been designed for specific purposes, by utilizing only some of the available sources of genetic studies, or by using a simple weight scheme. Specifically to psychiatric disorders, there has been no prioritization approach that fully utilizes all major sources of experimental data. RESULTS: Here we present a multi-dimensional evidence-based candidate gene prioritization approach for complex diseases and demonstrate it in schizophrenia. In this approach, we first collect and curate genetic studies for schizophrenia from four major categories: association studies, linkage analyses, gene expression and literature search. Genes in these data sets are initially scored by category-specific scoring methods. Then, an optimal weight matrix is searched by a two-step procedure (core genes and unbiased P-values in independent genome-wide association studies). Finally, genes are prioritized by their combined scores using the optimal weight matrix. Our evaluation suggests this approach generates prioritized candidate genes that are promising for further analysis or replication. The approach can be applied to other complex diseases. AVAILABILITY: The collected data, prioritized candidate genes, and gene prioritization tools are freely available at http://bioinfo.mc.vanderbilt.edu/SZGR/.
Jingchun Sun, Peilin Jia, Ayman H. Fanous, Bradley Todd Webb, Edwin J. C. G. van den Oord, Xiangning Chen, József Bukszár, Kenneth S. Kendler, Zhongming Zhao
Bioinform.9
2009 CpG islands or CpG clusters: how to identify functional GC-rich regions in a genome?
abstract
BACKGROUND: CpG islands (CGIs), clusters of CpG dinucleotides in GC-rich regions, are often located in the 5' end of genes and considered gene markers. Hackenberg et al. (2006) recently developed a new algorithm, CpGcluster, which uses a completely different mathematical approach from previous traditional algorithms. Their evaluation suggests that CpGcluster provides a much more efficient approach to detecting functional clusters or islands of CpGs. RESULTS: We systematically compared CpGcluster with the traditional algorithm by Takai and Jones (2002). Our comparisons of (1) the number of islands versus the number of genes in a genome, (2) the distribution of islands in different genomic regions, (3) island length, (4) the distance between two neighboring islands, and (5) methylation status suggest that Takai and Jones' algorithm is overall more appropriate for identifying promoter-associated islands of CpGs in vertebrate genomes. CONCLUSION: The generation of genome sequence and DNA methylation data is expected to accelerate greatly. The information in this study is important for its extensive utility in gene feature analysis and epigenomics including gene prediction and methylation chip design in different genomes.
Leng Han, Zhongming Zhao
BMC Bioinform.2
2008 FindSUMO: A PSSM-Based Method for Sumoylation Site Prediction
Christopher J. Friedline, Xueping Zhang, Zendra E. Zehner, Zhongming Zhao
ICIC (2)4
2008 An SVM-Based Algorithm for Classifying Promoter-Associated CpG Islands in the Human and Mouse Genomes
Leng Han, Ruolin Yang 0003, Zhongming Zhao
ICIC (2)4
2008 A fast two-step deblurring method for satellite images
abstract
In satellite images processing, image restoration is an important pre-processing step. As some traditional deblurring methods like Wiener filtering often obtain an either smooth or noisy solution, whereas some edge-preserving methods are iterative and time-consuming; a modified Fourier-wavelet deconvolution method (ForWaRD) is proposed in this paper. This method could be implemented forward directly with two steps. From the view of practical engineering, this approach could obtain pretty good results in a fast speed, so it is a suitable method in satellite image restoration area.
Zhensong Wang, Zhongming Zhao
SMC3
2007 InPrePPI: an integrated evaluation method based on genomic context for predicting protein-protein interactions in prokaryotic genomes
abstract
BACKGROUND: Although many genomic features have been used in the prediction of protein-protein interactions (PPIs), frequently only one is used in a computational method. After realizing the limited power in the prediction using only one genomic feature, investigators are now moving toward integration. So far, there have been few integration studies for PPI prediction; one failed to yield appreciable improvement of prediction and the others did not conduct performance comparison. It remains unclear whether an integration of multiple genomic features can improve the PPI prediction and, if it can, how to integrate these features. RESULTS: In this study, we first performed a systematic evaluation on the PPI prediction in Escherichia coli (E. coli) by four genomic context based methods: the phylogenetic profile method, the gene cluster method, the gene fusion method, and the gene neighbor method. The number of predicted PPIs and the average degree in the predicted PPI networks varied greatly among the four methods. Further, no method outperformed the others when we tested using three well-defined positive datasets from the KEGG, EcoCyc, and DIP databases. Based on these comparisons, we developed a novel integrated method, named InPrePPI. InPrePPI first normalizes the AC value (an integrated value of the accuracy and coverage) of each method using three positive datasets, then calculates a weight for each method, and finally uses the weight to calculate an integrated score for each protein pair predicted by the four genomic context based methods. We demonstrate that InPrePPI outperforms each of the four individual methods and, in general, the other two existing integrated methods: the joint observation method and the integrated prediction method in STRING. These four methods and InPrePPI are implemented in a user-friendly web interface. CONCLUSION: This study evaluated the PPI prediction by four genomic context based methods, and presents an integrated evaluation method that shows better performance in E. coli.
Jingchun Sun, Guohui Ding 0001, Qi Liu 0024, Youyu He, Tieliu Shi, Zhongming Zhao
BMC Bioinform.9
2005 Image compression based on unrestrained sized wavelet transform
abstract
In this paper, we adopt a wavelet decomposition scheme that doesn't require 2/sup N/ size image and avoid data padding. Instead of tree-structure coding algorithms, we just apply scalar quantization and arithmetic coding to every sub-band separately after wavelet transformation. The optimized bit-rates for every sub-band are confirmed by the BFOS algorithm. The experiment shows that this scheme can perform as well as the sophisticated algorithm SPIHT and it has the new property of resolution progressive.
Renxi Chen, Zhongming Zhao
IGARSS2
2005 Clouds and cloud shadows removal from high-resolution remote sensing images
Zhongming Zhao, Dongmei Yan
IGARSS2
2005 Multi-scale segmentation of the high resolution remote sensing image
abstract
The high spatial resolution remote sensing image provide more details such as color, size, shape, context and texture. The traditional pixel-based classifier cannot provide satisfying results and may reduce the classification accuracy. So the object-oriented processing for extraction of information from remote sensing data become of interest. The first step of the object-oriented processing is the segmentation. The object is derived by means of multi-scale segmentation in this paper. The hierarchical image segmentation and region-merging are implemented. The procedure of the region-merging does not stop until the average size of all object regions exceeds the scale threshold. Lastly, in order to provide an appropriate link between remote sensing image and GIS data, the region boundary will be obtained from the segmented remote sensing image and initialize the regions’ boundary to a series of polygon by Douglas-Peucker (DP) algorithm.
Zhongming Zhao, Dongmei Yan, Renxi Chen
IGARSS2
2005 The indigenous remote sensing image processing system and its exertion strategy
Zhongming Zhao, Jianglin Ma
IGARSS2
2005 A fused road detection approach in high resolution multi-spectrum remote sensing imagery
Dongmei Yan, Zhongming Zhao
IGARSS2
2005 SNPNB: analyzing neighboring-nucleotide biases on single nucleotide polymorphisms (SNPs)
abstract
UNLABELLED: SNPNB is a user-friendly and platform-independent application for analyzing Single Nucleotide Polymorphism NeighBoring sequence context and nucleotide bias patterns, and subsequently evaluating the effective SNP size for the bias patterns observed from the whole data. It was implemented by Java and Perl. SNPNB can efficiently handle genome-wide or chromosome-wide SNP data analysis in a PC or a workstation. It provides visualizations of the bias patterns for SNPs or each type of SNPs. AVAILABILITY: SNPNB and its full description are freely available at http://bioinfo.vipbg.vcu.edu/SNPNB/
Fengkai Zhang, Zhongming Zhao
Bioinform.2
2004 An integrated classification strategy of hyperspectral imaging spectrometer data
abstract
It is one of the hotspots to apply the advanced remote sensing data and processing techniques to monitor the desertification. Some monitor factors, such as vegetation, sand and soil moisture, were identified by use of the OMIS-I hyperspectral data individually and its integration with the 7/sub th/ band of ETM data in this study. The results indicate that the former has a high identification precision in vegetation and sand, but in soil moisture it is not well because of the influence of upper vegetation; this can been greatly improved in the latter and the overall identification precision is higher than the former.
Linli Cui, Zhongming Zhao
IGARSS3
2004 Remote sensing study based on IRSA Remote Sensing Image Processing System
abstract
The IRSA Remote Sensing Image Processing System is multi-functional software used for satellite image processing. It consists of over ten parts of the routine and typical used modules in Remote Sensing Image Processing project, such as viewer & file import/export, basic processing, image restoration. As an indigenous developed software, IRSA combines the advantages and kernels of many import famous systems, such as ERDAS imagine, PCI, ENVI and ER-mapper, and avoids some infrequently used functions or details. Hence, it appears concisely, refinedly and practically, acceptable and understandable. Based on this characteristic, we develop an additional set of interrelated data together with the system to face the college students and people who are not familiar with the Remote Sensing Image Processing work. Our experiences prove that we are successful. With the detailed help documents and instruction as well as our elaborately chose, arranged data, the students can study the system step by step. From these data, they get very intuitionistic and sensible cognition to the remote sensing study
Zhongming Zhao, Linli Cui
IGARSS2
2004 Texture image segmentation based on wavelet-domain hidden Markov models
abstract
Many approaches have been used to segment texture-based image, but can't have one's wish fulfilled. People unremittingly try to find high quality, more effective method. The 2D discrete wavelet transform, as a powerful and effective approach, have got preferable production in image analyzing and image processing. However routine method focuses on the assumption that the wavelet coefficients are independent and jointly Gaussian. In fact, most real-world images are not always Gaussian distributed, there are underlying relationships and rules among these wavelet coefficients on both the same scale and the inter-scale. We first reveal the dependencies among these coefficients through the wavelet-domain HMM's, then estimate the model parameters using the expectation maximization (EM) algorithms. The classification can be first realized through maximum likely method in each band, and then combine with the classification results of the three sub-bands from the 2D wavelet transform and also integrate the classification results of different scale. These approaches offer improved segmentation accuracy.
Zhongming Zhao, Jianglin Ma
IGARSS2
2004 An algorithm for eliminating the isolated regions based on connected area in image classification
abstract
It is well known that the results of remote sensing image classification contain many small isolated regions. These isolated regions need to be eliminated and merged into the existing classes in the post-processing. The most critical issue among all is how we could recognize the isolated regions, and what strategy we should adopt in automatic elimination. We present An algorithm for isolated regions' recognition based on connected area, in isolated region generation, we adopt an improved area filling algorithm in computer graphics. In automatic elimination, we propose a merge strategy which is based on the number of immediate pixels. Experimental results show that our recognition algorithm and merge strategy can effectively and automatically eliminate the isolated regions in the results of remote sensing image classification.
Jianghong Song, Zhongming Zhao, Qingye Zeng, Yanfeng Wei
IGARSS2
2004 Urban building extraction from high-resolution satellite panchromatic image using clustering and edge detection
abstract
For decades, large-scale aerial photos have been employed to extract building for mapping application. With the successively launching of high-resolution commercial satellites (e.g. IKONOS and QuickBird), high-resolution satellite imagery has been shown to be a cost-effective alternative to aerial photography in many applications. Drawing on the traditional building extraction approach, this paper proposes an algorithm to extract urban building from high-resolution panchromatic QuickBird image using clustering and edge detection. In the first step, an unsupervised clustering by histogram peak selection is used to split the image into a number of classes. The shadows of building are extracted from the lowest gray class. In the second step, the shadows are used as one of the evidences to verify the presence of buildings. Thus, the candidate building objects are extracted from the clustering classes except for the shadow class. Finally, to refine building boundary and further exclude some false building objects, the Canny operator is applied to detect edge of the candidate building objects in the PAN image. From the Hough transform of the detected edges, the main lines, which compose the polyhedral description of the building, can be found. The building extraction results are compared with manually delineated results. The comparison illustrates the efficiency of the proposed algorithm.
Yanfeng Wei, Zhongming Zhao, Jianghong Song
IGARSS2
2003 Automatic change detection of artificial objects in multitemporal high spatial resolution remotely sensed imagery
abstract
Change detection is one of the most important processes in various monitoring applications in multi-temporal remote sensed imagery. We focus on changes of artificial objects, including whether new artifical objects occur or existing artificial objects have changes. This paper proposes a new method to discriminate such changes in multi-temporal images using optimal quantization and block-based linear regression techniques. In the method, multi-temporal images are represented by less quantization level through optimal quantization method respectively; consequently, a block-based linear regression model is used to establish the relationship between multi-temporal images getting the changes effectively and automatically. The method is successfully applied to detect the changes of artificial objects without being affected by various vegetation covers for panchromatic high spatial resolution images such as IRS satellite images.
Zhongming Zhao
IGARSS2
2003 Road detection from Quickbird fused image using IHS transform and morphology
abstract
With the development of sensor technique, the commercial high resolution remote image would be directly served in digital city or monitoring urban changing. The topology feature of urban road is changed form line feature to fixed feature of line and segment. Especially to asphalt road, the shadow of urban building and high tree increases the difficult on describing the fixed feature from high resolution panchromatic image. In this paper, we present a region segmentation algorithm based on IHS transform of multispectral image, develop the orientation projection of candidate segment to get road segments and the intersections, and apply the morphology filter to improve the quality of the road detection. At the end of paper, some result image is presented to show the properties of the road network detection from the Quickbird fused multispectral image.
Dongmei Yan, Zhongming Zhao
IGARSS2