EDBT 2026 Demo / reviewers in the wild / expert
Robert Clarke
dblp:23/3883
· DBLP profile ↗
58ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 8 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bayesian identification of differentially expressed isoforms using a novel joint model of RNA-seq dataabstractWe develop a Bayesian approach, BayesIso, to identify differentially expressed isoforms from RNA-seq data. The approach features a novel joint model of the sample variability and the deferential state of isoforms. Specifically, the within-sample variability and the between-sample variability of each isoform are modeled by a Poisson-Lognormal model and a Gamma-Gamma model, respectively. Using a Bayesian framework, the differential state of each isoform and the model parameters are jointly estimated by a Markov Chain Monte Carlo (MCMC) method. Extensive studies using simulation and real data demonstrate that BayesIso can effectively detect isoforms of less differentially expressed and differential transcripts for genes with multiple isoforms. We applied the approach to breast cancer RNA-seq data and uncovered a unique set of isoforms that form key pathways associated with breast cancer recurrence. First, PI3K/AKT/mTOR signaling and PTEN signaling pathways are identified as being involved in breast cancer development. Further integrated with protein-protein interaction data, pathways of Jak-STAT, mTOR, MAPK and Wnt signaling are revealed in association with breast cancer recurrence. Finally, several pathways are activated in the early recurrence of breast cancer. In tumors that occur early, members of pathways of cellular metabolism and cell cycle (such as CD36 and TOP2A) are upregulated, while immune response genes such as NFATC1 are downregulated. Xu Shi 0003, Xiao Wang 0031, Leena Halakivi-Clarke, Robert Clarke, Andrew F. Neuwald, Jianhua Xuan |
PLoS Comput. Biol. | 5 |
| 2024 | DDN3.0: determining significant rewiring of biological network structure with differential dependency networksabstractMOTIVATION: Complex diseases are often caused and characterized by misregulation of multiple biological pathways. Differential network analysis aims to detect significant rewiring of biological network structures under different conditions and has become an important tool for understanding the molecular etiology of disease progression and therapeutic response. With few exceptions, most existing differential network analysis tools perform differential tests on separately learned network structures that are computationally expensive and prone to collapse when grouped samples are limited or less consistent. RESULTS: We previously developed an accurate differential network analysis method-differential dependency networks (DDN), that enables joint learning of common and rewired network structures under different conditions. We now introduce the DDN3.0 tool that improves this framework with three new and highly efficient algorithms, namely, unbiased model estimation with a weighted error measure applicable to imbalance sample groups, multiple acceleration strategies to improve learning efficiency, and data-driven determination of proper hyperparameters. The comparative experimental results obtained from both realistic simulations and case studies show that DDN3.0 can help biologists more accurately identify, in a study-specific and often unknown conserved regulatory circuitry, a network of significantly rewired molecular players potentially responsible for phenotypic transitions. AVAILABILITY AND IMPLEMENTATION: The Python package of DDN3.0 is freely available at https://github.com/cbil-vt/DDN3. A user's guide and a vignette are provided at https://ddn-30.readthedocs.io/. Yingzhou Lu, Yizhi Wang 0009, Bai Zhang, Guoqiang Yu, Chunyu Liu 0001, Robert Clarke, David M. Herrington, Yue Joseph Wang |
Bioinform. | 8 |
| 2024 | CAM3.0: determining cell type composition and expression from bulk tissues with fully unsupervised deconvolutionabstractMOTIVATION: Complex tissues are dynamic ecosystems consisting of molecularly distinct yet interacting cell types. Computational deconvolution aims to dissect bulk tissue data into cell type compositions and cell-specific expressions. With few exceptions, most existing deconvolution tools exploit supervised approaches requiring various types of references that may be unreliable or even unavailable for specific tissue microenvironments. RESULTS: We previously developed a fully unsupervised deconvolution method-Convex Analysis of Mixtures (CAM), that enables estimation of cell type composition and expression from bulk tissues. We now introduce CAM3.0 tool that improves this framework with three new and highly efficient algorithms, namely, radius-fixed clustering to identify reliable markers, linear programming to detect an initial scatter simplex, and a smart floating search for the optimum latent variable model. The comparative experimental results obtained from both realistic simulations and case studies show that the CAM3.0 tool can help biologists more accurately identify known or novel cell markers, determine cell proportions, and estimate cell-specific expressions, complementing the existing tools particularly when study- or datatype-specific references are unreliable or unavailable. AVAILABILITY AND IMPLEMENTATION: The open-source R Scripts of CAM3.0 is freely available at https://github.com/ChiungTingWu/CAM3/(https://github.com/Bioconductor/Contributions/issues/3205). A user's guide and a vignette are provided. Chiung-Ting Wu, Dongping Du, Lulu Chen, Rujia Dai, Chunyu Liu 0001, Guoqiang Yu, Saurabh Bhardwaj, Sarah J. Parker, Robert Clarke, David M. Herrington, Yue Joseph Wang |
Bioinform. | 10 |
| 2024 | AutoNet-Generated Deep Layer-Wise Convex Networks for ECG ClassificationabstractThe design of neural networks typically involves trial-and-error, a time-consuming process for obtaining an optimal architecture, even for experienced researchers. Additionally, it is widely accepted that loss functions of deep neural networks are generally non-convex with respect to the parameters to be optimised. We propose the Layer-wise Convex Theorem to ensure that the loss is convex with respect to the parameters of a given layer, achieved by constraining each layer to be an overdetermined system of non-linear equations. Based on this theorem, we developed an end-to-end algorithm (the AutoNet) to automatically generate layer-wise convex networks (LCNs) for any given training set. We then demonstrate the performance of the AutoNet-generated LCNs (AutoNet-LCNs) compared to state-of-the-art models on three electrocardiogram (ECG) classification benchmark datasets, with further validation on two non-ECG benchmark datasets for more general tasks. The AutoNet-LCN was able to find networks customised for each dataset without manual fine-tuning under 2 GPU-hours, and the resulting networks outperformed the state-of-the-art models with fewer than 5% parameters on all the above five benchmark datasets. The efficiency and robustness of the AutoNet-LCN markedly reduce model discovery costs and enable efficient training of deep learning models in resource-constrained settings. Yanting Shen, Tingting Zhu 0001, Xinshao Wang, Lei A. Clifton, Zhengming Chen 0003, Robert Clarke, David A. Clifton |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | swCAM: estimation of subtype-specific expressions in individual samples with unsupervised sample-wise deconvolutionabstractMOTIVATION: Complex biological tissues are often a heterogeneous mixture of several molecularly distinct cell subtypes. Both subtype compositions and subtype-specific (STS) expressions can vary across biological conditions. Computational deconvolution aims to dissect patterns of bulk tissue data into subtype compositions and STS expressions. Existing deconvolution methods can only estimate averaged STS expressions in a population, while many downstream analyses such as inferring co-expression networks in particular subtypes require subtype expression estimates in individual samples. However, individual-level deconvolution is a mathematically underdetermined problem because there are more variables than observations. RESULTS: We report a sample-wise Convex Analysis of Mixtures (swCAM) method that can estimate subtype proportions and STS expressions in individual samples from bulk tissue transcriptomes. We extend our previous CAM framework to include a new term accounting for between-sample variations and formulate swCAM as a nuclear-norm and ℓ2,1-norm regularized matrix factorization problem. We determine hyperparameter values using cross-validation with random entry exclusion and obtain a swCAM solution using an efficient alternating direction method of multipliers. Experimental results on realistic simulation data show that swCAM can accurately estimate STS expressions in individual samples and successfully extract co-expression networks in particular subtypes that are otherwise unobtainable using bulk data. In two real-world applications, swCAM analysis of bulk RNASeq data from brain tissue of cases and controls with bipolar disorder or Alzheimer's disease identified significant changes in cell proportion, expression pattern and co-expression module in patient neurons. Comparative evaluation of swCAM versus peer methods is also provided. AVAILABILITY AND IMPLEMENTATION: The R Scripts of swCAM are freely available at https://github.com/Lululuella/swCAM. A user's guide and a vignette are provided. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lulu Chen, Chiung-Ting Wu, Chia-Hsiang Lin, Rujia Dai, Chunyu Liu 0001, Robert Clarke, Guoqiang Yu, Jennifer E. Van Eyk, David M. Herrington, Yue Joseph Wang |
Bioinform. | 6 |
| 2021 | IntAPT: integrated assembly of phenotype-specific transcripts from multiple RNA-seq profilesabstractMOTIVATION: High-throughput RNA sequencing has revolutionized the scope and depth of transcriptome analysis. Accurate reconstruction of a phenotype-specific transcriptome is challenging due to the noise and variability of RNA-seq data. This requires computational identification of transcripts from multiple samples of the same phenotype, given the underlying consensus transcript structure. RESULTS: We present a Bayesian method, integrated assembly of phenotype-specific transcripts (IntAPT), that identifies phenotype-specific isoforms from multiple RNA-seq profiles. IntAPT features a novel two-layer Bayesian model to capture the presence of isoforms at the group layer and to quantify the abundance of isoforms at the sample layer. A spike-and-slab prior is used to model the isoform expression and to enforce the sparsity of expressed isoforms. Dependencies between the existence of isoforms and their expression are modeled explicitly to facilitate parameter estimation. Model parameters are estimated iteratively using Gibbs sampling to infer the joint posterior distribution, from which the presence and abundance of isoforms can reliably be determined. Studies using both simulations and real datasets show that IntAPT consistently outperforms existing methods for the IntAPT. Experimental results demonstrate that, despite sequencing errors, IntAPT exhibits a robust performance among multiple samples, resulting in notably improved identification of expressed isoforms of low abundance. AVAILABILITY AND IMPLEMENTATION: The IntAPT package is available at http://github.com/henryxushi/IntAPT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xu Shi 0003, Andrew F. Neuwald, Xiao Wang 0031, Tian-Li Wang, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 6 |
| 2021 | ChIP-BIT2: a software tool to detect weak binding events using a Bayesian integration approachabstractBACKGROUND: ChIP-seq combines chromatin immunoprecipitation assays with sequencing and identifies genome-wide binding sites for DNA binding proteins. While many binding sites have strong ChIP-seq 'peak' observations and are well captured, there are still regions bound by proteins weakly, with a relatively low ChIP-seq signal enrichment. These weak binding sites, especially those at promoters and enhancers, are functionally important because they also regulate nearby gene expression. Yet, it remains a challenge to accurately identify weak binding sites in ChIP-seq data due to the ambiguity in differentiating these weak binding sites from the amplified background DNAs. RESULTS: ChIP-BIT2 ( http://sourceforge.net/projects/chipbitc/ ) is a software package for ChIP-seq peak detection. ChIP-BIT2 employs a mixture model integrating protein and control ChIP-seq data and predicts strong or weak protein binding sites at promoters, enhancers, or other genomic locations. For binding sites at gene promoters, ChIP-BIT2 simultaneously predicts their target genes. ChIP-BIT2 has been validated on benchmark regions and tested using large-scale ENCODE ChIP-seq data, demonstrating its high accuracy and wide applicability. CONCLUSION: ChIP-BIT2 is an efficient ChIP-seq peak caller. It provides a better lens to examine weak binding sites and can refine or extend the existing binding site collection, providing additional regulatory regions for decoding the mechanism of gene expression regulation. Xi Chen 0056, Xu Shi 0003, Andrew F. Neuwald, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
BMC Bioinform. | 5 |
| 2021 | Stroke risk prediction using machine learning: a prospective cohort study of 0.5 million Chinese adultsabstractOBJECTIVE: To compare Cox models, machine learning (ML), and ensemble models combining both approaches, for prediction of stroke risk in a prospective study of Chinese adults. MATERIALS AND METHODS: We evaluated models for stroke risk at varying intervals of follow-up (<9 years, 0-3 years, 3-6 years, 6-9 years) in 503 842 adults without prior history of stroke recruited from 10 areas in China in 2004-2008. Inputs included sociodemographic factors, diet, medical history, physical activity, and physical measurements. We compared discrimination and calibration of Cox regression, logistic regression, support vector machines, random survival forests, gradient boosted trees (GBT), and multilayer perceptrons, benchmarking performance against the 2017 Framingham Stroke Risk Profile. We then developed an ensemble approach to identify individuals at high risk of stroke (>10% predicted 9-yr stroke risk) by selectively applying either a GBT or Cox model based on individual-level characteristics. RESULTS: For 9-yr stroke risk prediction, GBT provided the best discrimination (AUROC: 0.833 in men, 0.836 in women) and calibration, with consistent results in each interval of follow-up. The ensemble approach yielded incrementally higher accuracy (men: 76%, women: 80%), specificity (men: 76%, women: 81%), and positive predictive value (men: 26%, women: 24%) compared to any of the single-model approaches. DISCUSSION AND CONCLUSION: Among several approaches, an ensemble model combining both GBT and Cox models achieved the best performance for identifying individuals at high risk of stroke in a contemporary study of Chinese adults. The results highlight the potential value of expanding the use of ML in clinical practice. Matthew Chun, Robert Clarke, Benjamin J. Cairns, David A. Clifton, Derrick Bennett, Pei Pei, Canqing Yu, Zhengming Chen 0003, Tingting Zhu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | ChIP-GSM: Inferring active transcription factor modules to predict functional regulatory elementsabstractTranscription factors (TFs) often function as a module including both master factors and mediators binding at cis-regulatory regions to modulate nearby gene transcription. ChIP-seq profiling of multiple TFs makes it feasible to infer functional TF modules. However, when inferring TF modules based on co-localization of ChIP-seq peaks, often many weak binding events are missed, especially for mediators, resulting in incomplete identification of modules. To address this problem, we develop a ChIP-seq data-driven Gibbs Sampler to infer Modules (ChIP-GSM) using a Bayesian framework that integrates ChIP-seq profiles of multiple TFs. ChIP-GSM samples read counts of module TFs iteratively to estimate the binding potential of a module to each region and, across all regions, estimates the module abundance. Using inferred module-region probabilistic bindings as feature units, ChIP-GSM then employs logistic regression to predict active regulatory elements. Validation of ChIP-GSM predicted regulatory regions on multiple independent datasets sharing the same context confirms the advantage of using TF modules for predicting regulatory activity. In a case study of K562 cells, we demonstrate that the ChIP-GSM inferred modules form as groups, activate gene expression at different time points, and mediate diverse functional cellular processes. Hence, ChIP-GSM infers biologically meaningful TF modules and improves the prediction accuracy of regulatory region activities. Xi Chen 0056, Andrew F. Neuwald, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
PLoS Comput. Biol. | 4 |
| 2020 | debCAM: a bioconductor R package for fully unsupervised deconvolution of complex tissuesabstractSUMMARY: We develop a fully unsupervised deconvolution method to dissect complex tissues into molecularly distinctive tissue or cell subtypes based on bulk expression profiles. We implement an R package, deconvolution by Convex Analysis of Mixtures (debCAM) that can automatically detect tissue/cell-specific markers, determine the number of constituent subtypes, calculate subtype proportions in individual samples and estimate tissue/cell-specific expression profiles. We demonstrate the performance and biomedical utility of debCAM on gene expression, methylation, proteomics and imaging data. With enhanced data preprocessing and prior knowledge incorporation, debCAM software tool will allow biologists to perform a more comprehensive and unbiased characterization of tissue remodeling in many biomedical contexts. AVAILABILITY AND IMPLEMENTATION: http://bioconductor.org/packages/debCAM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lulu Chen, Chiung-Ting Wu, Niya Wang, David M. Herrington, Robert Clarke, Yue Joseph Wang |
Bioinform. | 5 |
| 2018 | CRNET: an efficient sampling approach to infer functional regulatory networks by integrating large-scale ChIP-seq and time-course RNA-seq dataabstractMotivation: NGS techniques have been widely applied in genetic and epigenetic studies. Multiple ChIP-seq and RNA-seq profiles can now be jointly used to infer functional regulatory networks (FRNs). However, existing methods suffer from either oversimplified assumption on transcription factor (TF) regulation or slow convergence of sampling for FRN inference from large-scale ChIP-seq and time-course RNA-seq data. Results: We developed an efficient Bayesian integration method (CRNET) for FRN inference using a two-stage Gibbs sampler to estimate iteratively hidden TF activities and the posterior probabilities of binding events. A novel statistic measure that jointly considers regulation strength and regression error enables the sampling process of CRNET to converge quickly, thus making CRNET very efficient for large-scale FRN inference. Experiments on synthetic and benchmark data showed a significantly improved performance of CRNET when compared with existing methods. CRNET was applied to breast cancer data to identify FRNs functional at promoter or enhancer regions in breast cancer MCF-7 cells. Transcription factor MYC is predicted as a key functional factor in both promoter and enhancer FRNs. We experimentally validated the regulation effects of MYC on CRNET-predicted target genes using appropriate RNAi approaches in MCF-7 cells. Availability and implementation: R scripts of CRNET are available at http://www.cbil.ece.vt.edu/software.htm. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xi Chen 0056, Jinghua Gu, Xiao Wang 0031, Jin-Gyoung Jung, Tian-Li Wang, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 7 |
| 2018 | SparseIso: a novel Bayesian approach to identify alternatively spliced isoforms from RNA-seq dataabstractMotivation: Recent advances in high-throughput RNA sequencing (RNA-seq) technologies have made it possible to reconstruct the full transcriptome of various types of cells. It is important to accurately assemble transcripts or identify isoforms for an improved understanding of molecular mechanisms in biological systems. Results: We have developed a novel Bayesian method, SparseIso, to reliably identify spliced isoforms from RNA-seq data. A spike-and-slab prior is incorporated into the Bayesian model to enforce the sparsity for isoform identification, effectively alleviating the problem of overfitting. A Gibbs sampling procedure is further developed to simultaneously identify and quantify transcripts from RNA-seq data. With the sampling approach, SparseIso estimates the joint distribution of all candidate transcripts, resulting in a significantly improved performance in detecting lowly expressed transcripts and multiple expressed isoforms of genes. Both simulation study and real data analysis have demonstrated that the proposed SparseIso method significantly outperforms existing methods for improved transcript assembly and isoform identification. Availability and implementation: The SparseIso package is available at http://github.com/henryxushi/SparseIso. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xu Shi 0003, Xiao Wang 0031, Tian-Li Wang, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 5 |
| 2017 | PSSV: a novel pattern-based probabilistic approach for somatic structural variation identificationabstractMOTIVATION: Whole genome DNA-sequencing (WGS) of paired tumor and normal samples has enabled the identification of somatic DNA changes in an unprecedented detail. Large-scale identification of somatic structural variations (SVs) for a specific cancer type will deepen our understanding of driver mechanisms in cancer progression. However, the limited number of WGS samples, insufficient read coverage, and the impurity of tumor samples that contain normal and neoplastic cells, limit reliable and accurate detection of somatic SVs. RESULTS: We present a novel pattern-based probabilistic approach, PSSV, to identify somatic structural variations from WGS data. PSSV features a mixture model with hidden states representing different mutation patterns; PSSV can thus differentiate heterozygous and homozygous SVs in each sample, enabling the identification of those somatic SVs with heterozygous mutations in normal samples and homozygous mutations in tumor samples. Simulation studies demonstrate that PSSV outperforms existing tools. PSSV has been successfully applied to breast cancer data to identify somatic SVs of key factors associated with breast cancer development. AVAILABILITY AND IMPLEMENTATION: An R package of PSSV is available at http://www.cbil.ece.vt.edu/software.htm CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Xi Chen 0056, Xu Shi 0003, Leena Hilakivi-Clarke, Ayesha N. Shajahan, Robert Clarke, Jianhua Xuan |
Bioinform. | 5 |
| 2017 | DM-BLD: differential methylation detection using a hierarchical Bayesian model exploiting local dependencyabstractMOTIVATION: The advent of high-throughput DNA methylation profiling techniques has enabled the possibility of accurate identification of differentially methylated genes for cancer research. The large number of measured loci facilitates whole genome methylation study, yet posing great challenges for differential methylation detection due to the high variability in tumor samples. RESULTS: We have developed a novel probabilistic approach, D: ifferential M: ethylation detection using a hierarchical B: ayesian model exploiting L: ocal D: ependency (DM-BLD), to detect differentially methylated genes based on a Bayesian framework. The DM-BLD approach features a joint model to capture both the local dependency of measured loci and the dependency of methylation change in samples. Specifically, the local dependency is modeled by Leroux conditional autoregressive structure; the dependency of methylation changes is modeled by a discrete Markov random field. A hierarchical Bayesian model is developed to fully take into account the local dependency for differential analysis, in which differential states are embedded as hidden variables. Simulation studies demonstrate that DM-BLD outperforms existing methods for differential methylation detection, particularly when the methylation change is moderate and the variability of methylation in samples is high. DM-BLD has been applied to breast cancer data to identify important methylated genes (such as polycomb target genes and genes involved in transcription factor activity) associated with breast cancer recurrence. AVAILABILITY AND IMPLEMENTATION: A Matlab package of DM-BLD is available at http://www.cbil.ece.vt.edu/software.htm CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Xiao Wang 0031, Jinghua Gu, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 4 |
| 2016 | G-DOC Plus - an integrative bioinformatics platform for precision medicineabstractBACKGROUND: G-DOC Plus is a data integration and bioinformatics platform that uses cloud computing and other advanced computational tools to handle a variety of biomedical BIG DATA including gene expression arrays, NGS and medical images so that they can be analyzed in the full context of other omics and clinical information. RESULTS: G-DOC Plus currently holds data from over 10,000 patients selected from private and public resources including Gene Expression Omnibus (GEO), The Cancer Genome Atlas (TCGA) and the recently added datasets from REpository for Molecular BRAin Neoplasia DaTa (REMBRANDT), caArray studies of lung and colon cancer, ImmPort and the 1000 genomes data sets. The system allows researchers to explore clinical-omic data one sample at a time, as a cohort of samples; or at the level of population, providing the user with a comprehensive view of the data. G-DOC Plus tools have been leveraged in cancer and non-cancer studies for hypothesis generation and validation; biomarker discovery and multi-omics analysis, to explore somatic mutations and cancer MRI images; as well as for training and graduate education in bioinformatics, data and computational sciences. Several of these use cases are described in this paper to demonstrate its multifaceted usability. CONCLUSION: G-DOC Plus can be used to support a variety of user groups in multiple domains to enable hypothesis generation for precision medicine research. The long-term vision of G-DOC Plus is to extend this translational bioinformatics platform to stay current with emerging omics technologies and analysis methods to continue supporting novel hypothesis generation, analysis and validation for integrative biomedical research. By integrating several aspects of the disease and exposing various data elements, such as outpatient lab workup, pathology, radiology, current treatments, molecular signatures and expected outcomes over a web interface, G-DOC Plus will continue to strengthen precision medicine research. G-DOC Plus is available at: https://gdoc.georgetown.edu . Krithika Bhuvaneshwar, Anas Belouali, Varun Singh, Robert M. Johnson, Adil Alaoui, Michael Harris, Robert Clarke, Louis M. Weiner, Yuriy Gusev, Subha Madhavan |
BMC Bioinform. | 8 |
| 2015 | BMRF-Net: a software tool for identification of protein interaction subnetworks by a bagging Markov random field-based methodabstractUNLABELLED: Identification of protein interaction subnetworks is an important step to help us understand complex molecular mechanisms in cancer. In this paper, we develop a BMRF-Net package, implemented in Java and C++, to identify protein interaction subnetworks based on a bagging Markov random field (BMRF) framework. By integrating gene expression data and protein-protein interaction data, this software tool can be used to identify biologically meaningful subnetworks. A user friendly graphic user interface is developed as a Cytoscape plugin for the BMRF-Net software to deal with the input/output interface. The detailed structure of the identified networks can be visualized in Cytoscape conveniently. The BMRF-Net package has been applied to breast cancer data to identify significant subnetworks related to breast cancer recurrence. AVAILABILITY AND IMPLEMENTATION: The BMRF-Net package is available at http://sourceforge.net/projects/bmrfcjava/. The package is tested under Ubuntu 12.04 (64-bit), Java 7, glibc 2.15 and Cytoscape 3.1.0. Xu Shi 0003, Robert O. Barnes, Li Chen 0018, Ayesha N. Shajahan, Leena Hilakivi-Clarke, Robert Clarke, Yue Joseph Wang, Jianhua Xuan |
Bioinform. | 6 |
| 2015 | KDDN: an open-source Cytoscape app for constructing differential dependency networks with significant rewiringabstractUNLABELLED: We have developed an integrated molecular network learning method, within a well-grounded mathematical framework, to construct differential dependency networks with significant rewiring. This knowledge-fused differential dependency networks (KDDN) method, implemented as a Java Cytoscape app, can be used to optimally integrate prior biological knowledge with measured data to simultaneously construct both common and differential networks, to quantitatively assign model parameters and significant rewiring p-values and to provide user-friendly graphical results. The KDDN algorithm is computationally efficient and provides users with parallel computing capability using ubiquitous multi-core machines. We demonstrate the performance of KDDN on various simulations and real gene expression datasets, and further compare the results with those obtained by the most relevant peer methods. The acquired biologically plausible results provide new insights into network rewiring as a mechanistic principle and illustrate KDDN's ability to detect them efficiently and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY: Source code and compiled package are freely available for download at http://apps.cytoscape.org/apps/kddn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bai Zhang, Eric P. Hoffman, Robert Clarke, Ie-Ming Shih, Jianhua Xuan, David M. Herrington, Yue Joseph Wang |
Bioinform. | 4 |
| 2015 | UNDO: a Bioconductor R package for unsupervised deconvolution of mixed gene expressions in tumor samplesabstractSUMMARY: We develop a novel unsupervised deconvolution method, within a well-grounded mathematical framework, to dissect mixed gene expressions in heterogeneous tumor samples. We implement an R package, UNsupervised DecOnvolution (UNDO), that can be used to automatically detect cell-specific marker genes (MGs) located on the scatter radii of mixed gene expressions, estimate cellular proportions in each sample and deconvolute mixed expressions into cell-specific expression profiles. We demonstrate the performance of UNDO over a wide range of tumor-stroma mixing proportions, validate UNDO on various biologically mixed benchmark gene expression datasets and further estimate tumor purity in TCGA/CPTAC datasets. The highly accurate deconvolution results obtained suggest not only the existence of cell-specific MGs but also UNDO's ability to detect them blindly and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY AND IMPLEMENTATION: UNDO is available at http://bioconductor.org/packages. Niya Wang, Robert Clarke, Lulu Chen, Ie-Ming Shih, Douglas A. Levine, Jianhua Xuan, Yue Joseph Wang |
Bioinform. | 3 |
| 2014 | A Markov random field-based Bayesian model to identify genes with differential methylationabstractThe rapid development of biotechnology makes it possible to explore genome-wide DNA methylation mapping which has been demonstrated to be related to diseases including cancer. However, it also posts substantial challenges in identifying biologically meaningful methylation pattern changes. Several algorithms have been proposed to detect differential methylation events, such as differentially methylated CpG sites and differentially methylated regions. However, the intrinsic dependency of the CpG sites in a neighboring area has not yet been fully considered. In this paper, we propose a novel method for the identification of differentially methylated genes in a Markov random field-based Bayesian framework. Specifically, we use Markov random field to model the dependency of the neighboring CpG sites, and then estimate the differential methylation score of the CpG sites in a Bayesian framework through a sampling scheme. Finally, the differential methylation statuses of the genes are determined by the estimated scores of the involved CpG sites. In addition, significance test is conducted to assess the significance of the identified differentially methylated genes. Experimental results on both synthetic data and real data demonstrate the effectiveness of the proposed method in identifying genes with differential methylation patterns under different conditions. Xiao Wang 0031, Jinghua Gu, Jianhua Xuan, Robert Clarke, Leena Hilakivi-Clarke |
CIBCB | 4 |
| 2014 | An Observational Analysis of the Range and Extent of Contract Cheating from Online Courses Found on Agency WebsitesabstractAlthough online courses can provide access to higher education through e-learning systems which would not otherwise be available for students, they also pose challenges for academic integrity. Paramount to this is contract cheating, where students have been observed paying other people to complete work for them to complete their online courses. This paper analyses attempts by students at contract cheating using Transtutors.com, which is a billed as a site for homework support. A sample of 174 online assignments found on Transtutors.com are analysed and traced back to 17 online universities. Assignments from online institutions are demonstrated to be a particular problem for contract cheating detectives, since notifying staff at those institutions of attempts by their students to cheat has proved to be difficult or impossible. The paper concludes by looking at the wider issues posed by online contract cheating and the opportunities for automated detection within this field. Thomas Lancaster, Robert Clarke |
CISIS | 2 |
| 2014 | Robust identification of transcriptional regulatory networks using a Gibbs sampler on outlier sum statisticabstractContact: [email protected] Bioinformatics (2012) 28 (15), 1990–1997 doi:10.1093/bioinformatics/bts296 The authors wish to add one citation to a relevant conference report, Gu, J., Xuan, J., Wang, Y., Riggins, R.B. and Clarke R. (2010) Identification of transcriptional regulatory networks by learning the marginal function of outlier sum statistic. Proceedings of International Conference on Machine Learning and Applications , 281–286. The formatted reference is given below and should read in the sentence: In particular, a novel statistic for testing the confidence of target genes, namely, outlier sum of regression t -statistic ( Gu et al. , 2010 ), is specifically designed to pin-down confident target genes; based on this statistic, a Gibbs sampling strategy is used to sample target genes in a high probability as governed by the underlying distribution. The authors apologize for this oversight. Jinghua Gu, Jianhua Xuan, Rebecca B. Riggins, Li Chen 0018, Yue Joseph Wang, Robert Clarke |
Bioinform. | 6 |
| 2014 | AISAIC: a software suite for accurate identification of significant aberrations in cancersabstractUNLABELLED: Accurate identification of significant aberrations in cancers (AISAIC) is a systematic effort to discover potential cancer-driving genes such as oncogenes and tumor suppressors. Two major confounding factors against this goal are the normal cell contamination and random background aberrations in tumor samples. We describe a Java AISAIC package that provides comprehensive analytic functions and graphic user interface for integrating two statistically principled in silico approaches to address the aforementioned challenges in DNA copy number analyses. In addition, the package provides a command-line interface for users with scripting and programming needs to incorporate or extend AISAIC to their customized analysis pipelines. This open-source multiplatform software offers several attractive features: (i) it implements a user friendly complete pipeline from processing raw data to reporting analytic results; (ii) it detects deletion types directly from copy number signals using a Bayes hypothesis test; (iii) it estimates the fraction of normal contamination for each sample; (iv) it produces unbiased null distribution of random background alterations by iterative aberration-exclusive permutations; and (v) it identifies significant consensus regions and the percentage of homozygous/hemizygous deletions across multiple samples. AISAIC also provides users with a parallel computing option to leverage ubiquitous multicore machines. AVAILABILITY AND IMPLEMENTATION: AISAIC is available as a Java application, with a user's guide and source code, at https://code.google.com/p/aisaic/. Bai Zhang, Xuchu Hou, Xiguo Yuan, Ie-Ming Shih, Robert Clarke, Roger R. Wang, Subha Madhavan, Yue Joseph Wang, Guoqiang Yu |
Bioinform. | 6 |
| 2014 | BADGE: A novel Bayesian model for accurate abundance quantification and differential analysis of RNA-Seq dataabstractBACKGROUND: Recent advances in RNA sequencing (RNA-Seq) technology have offered unprecedented scope and resolution for transcriptome analysis. However, precise quantification of mRNA abundance and identification of differentially expressed genes are complicated due to biological and technical variations in RNA-Seq data. RESULTS: We systematically study the variation in count data and dissect the sources of variation into between-sample variation and within-sample variation. A novel Bayesian framework is developed for joint estimate of gene level mRNA abundance and differential state, which models the intrinsic variability in RNA-Seq to improve the estimation. Specifically, a Poisson-Lognormal model is incorporated into the Bayesian framework to model within-sample variation; a Gamma-Gamma model is then used to model between-sample variation, which accounts for over-dispersion of read counts among multiple samples. Simulation studies, where sequencing counts are synthesized based on parameters learned from real datasets, have demonstrated the advantage of the proposed method in both quantification of mRNA abundance and identification of differentially expressed genes. Moreover, performance comparison on data from the Sequencing Quality Control (SEQC) Project with ERCC spike-in controls has shown that the proposed method outperforms existing RNA-Seq methods in differential analysis. Application on breast cancer dataset has further illustrated that the proposed Bayesian model can 'blindly' estimate sources of variation caused by sequencing biases. CONCLUSIONS: We have developed a novel Bayesian hierarchical approach to investigate within-sample and between-sample variations in RNA-Seq data. Simulation and real data applications have validated desirable performance of the proposed method. The software package is available at http://www.cbil.ece.vt.edu/software.htm. Jinghua Gu, Xiao Wang 0031, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
BMC Bioinform. | 4 |
| 2014 | Integration of Network Biology and Imaging to Study Cancer Phenotypes and ResponsesabstractEver growing "omics" data and continuously accumulated biological knowledge provide an unprecedented opportunity to identify molecular biomarkers and their interactions that are responsible for cancer phenotypes that can be accurately defined by clinical measurements such as in vivo imaging. Since signaling or regulatory networks are dynamic and context-specific, systematic efforts to characterize such structural alterations must effectively distinguish significant network rewiring from random background fluctuations. Here we introduced a novel integration of network biology and imaging to study cancer phenotypes and responses to treatments at the molecular systems level. Specifically, Differential Dependence Network (DDN) analysis was used to detect statistically significant topological rewiring in molecular networks between two phenotypic conditions, and in vivo Magnetic Resonance Imaging (MRI) was used to more accurately define phenotypic sample groups for such differential analysis. We applied DDN to analyze two distinct phenotypic groups of breast cancer and study how genomic instability affects the molecular network topologies in high-grade ovarian cancer. Further, FDA-approved arsenic trioxide (ATO) and the ND2-SmoA1 mouse model of Medulloblastoma (MB) were used to extend our analyses of combined MRI and Reverse Phase Protein Microarray (RPMA) data to assess tumor responses to ATO and to uncover the complexity of therapeutic molecular biology. Sean S. Wang, Olga C. Rodriguez, Emanuel Petricoin III, Ie-Ming Shih, Daniel Chan, Maria Avantaggiati, Guoqiang Yu, Shaozhen Ye, Robert Clarke, Chao Wang 0005, Bai Zhang, Yue Joseph Wang, Chris Albanese |
IEEE ACM Trans. Comput. Biol. Bioinform. | 11 |
| 2013 | A novel statistical approach to identify co-regulatory gene modulesabstractChlP-chip experiments are performed to determine binding sites for transcription factors (TFs). Conventional TF-gene regulation is generated based on p-value cutoff of the binding sites as well as their distance to nearest genes. Taking into account that binding sites of one ChlP-chip experiment should follow the same specific location distribution, we proposed a statistical model using both location and significance information to weigh target genes. With multiple ChlP-chip experiments and gene expression data, we identified co-regulatory and differentially expressed gene modules with a joint clustering and Metropolis sampling approach. We demonstrated the efficiency of our method on a ChlP-chip data set with 38 breast cancer related TFs. Xi Chen 0056, Jianhua Xuan, Xu Shi 0003, Ayesha N. Shajahan, Leena Hilakivi-Clarke, Robert Clarke |
BIBM | 6 |
| 2013 | Commercial aspects of contract cheatingabstractThe process of contract cheating, the form of academic dishonesty where students outsource the creation of work on their behalf, has been recognised as a serious threat to the quality of academic awards. Unlike student plagiarism, this cheating behaviour is not currently detectable using automated tools. Robert Clarke, Thomas Lancaster |
ITiCSE | 1 |
| 2013 | The CAM software for nonnegative blind source separation in R-Java
Niya Wang, Li Chen 0018, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Yue Joseph Wang |
J. Mach. Learn. Res. | 5 |
| 2013 | Reconstruction of Transcriptional Regulatory Networks by Stability-Based Network Component AnalysisabstractReliable inference of transcription regulatory networks is a challenging task in computational biology. Network component analysis (NCA) has become a powerful scheme to uncover regulatory networks behind complex biological processes. However, the performance of NCA is impaired by the high rate of false connections in binding information. In this paper, we integrate stability analysis with NCA to form a novel scheme, namely stability-based NCA (sNCA), for regulatory network identification. The method mainly addresses the inconsistency between gene expression data and binding motif information. Small perturbations are introduced to prior regulatory network, and the distance among multiple estimated transcript factor (TF) activities is computed to reflect the stability for each TF's binding network. For target gene identification, multivariate regression and t-statistic are used to calculate the significance for each TF-gene connection. Simulation studies are conducted and the experimental results show that sNCA can achieve an improved and robust performance in TF identification as compared to NCA. The approach for target gene identification is also demonstrated to be suitable for identifying true connections between TFs and their target genes. Furthermore, we have successfully applied sNCA to breast cancer data to uncover the role of TFs in regulating endocrine resistance in breast cancer. Xi Chen 0056, Jianhua Xuan, Chen Wang 0001, Ayesha N. Shajahan, Rebecca B. Riggins, Robert Clarke |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2012 | Sampling-Based Subnetwork Identification from Microarray Data and Protein-Protein Interaction NetworkabstractIdentification of condition-specific protein interaction subnetworks has emerged as an attractive research field to reveal molecular mechanisms of diseases and provide reliable network biomarkers for disease diagnosis. Several methods have been proposed, which integrate gene expression and protein-protein interaction (PPI) data to identify subnetworks. However, existing methods treat differential expression of genes and network topology independently, which is an oversimplified assumption to model real biological systems. In this paper, we propose a sampling-based subnetwork identification approach to take into account the dependency between gene expression and network topology. Specifically, we apply Markov random field (MRF) theory to model the dependency of genes in PPI network using a Bayesian framework, followed by a Markov Chain Monte Carlo (MCMC) approach to identify significant subnetworks. The MCMC approach estimates the posterior distribution of genes' significant scores and network structure iteratively. Experimental results on both synthetic data and real breast cancer data demonstrated the effectiveness of the proposed method in identifying subnetworks, especially several functionally important, aberrant subnetworks associated with pathways involved in the development and recurrence of breast cancer. Xiao Wang 0031, Jinghua Gu, Jianhua Xuan, Ayesha N. Shajahan, Robert Clarke, Li Chen 0018 |
ICMLA (2) | 5 |
| 2012 | Reconstruction of Transcription Regulatory Networks by Stability-Based Network Component Analysis
Xi Chen 0056, Chen Wang 0001, Ayesha N. Shajahan, Rebecca B. Riggins, Robert Clarke, Jianhua Xuan |
ISBRA | 5 |
| 2012 | Robust identification of transcriptional regulatory networks using a Gibbs sampler on outlier sum statisticabstractMOTIVATION: Identification of transcriptional regulatory networks (TRNs) is of significant importance in computational biology for cancer research, providing a critical building block to unravel disease pathways. However, existing methods for TRN identification suffer from the inclusion of excessive 'noise' in microarray data and false-positives in binding data, especially when applied to human tumor-derived cell line studies. More robust methods that can counteract the imperfection of data sources are therefore needed for reliable identification of TRNs in this context. RESULTS: In this article, we propose to establish a link between the quality of one target gene to represent its regulator and the uncertainty of its expression to represent other target genes. Specifically, an outlier sum statistic was used to measure the aggregated evidence for regulation events between target genes and their corresponding transcription factors. A Gibbs sampling method was then developed to estimate the marginal distribution of the outlier sum statistic, hence, to uncover underlying regulatory relationships. To evaluate the effectiveness of our proposed method, we compared its performance with that of an existing sampling-based method using both simulation data and yeast cell cycle data. The experimental results show that our method consistently outperforms the competing method in different settings of signal-to-noise ratio and network topology, indicating its robustness for biological applications. Finally, we applied our method to breast cancer cell line data and demonstrated its ability to extract biologically meaningful regulatory modules related to estrogen signaling and action in breast cancer. AVAILABILITY AND IMPLEMENTATION: The Gibbs sampler MATLAB package is freely available at http://www.cbil.ece.vt.edu/software.htm. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jinghua Gu, Jianhua Xuan, Rebecca B. Riggins, Li Chen 0018, Yue Joseph Wang, Robert Clarke |
Bioinform. | 6 |
| 2012 | Regulatory component analysis: A semi-blind extraction approach to infer gene regulatory networks with imperfect biological knowledge
Chen Wang 0001, Jianhua Xuan, Ie-Ming Shih, Robert Clarke, Yue Joseph Wang |
Signal Process. | 4 |
| 2011 | PUGSVM: a caBIGTM analytical tool for multiclass gene selection and predictive classificationabstractUNLABELLED: Phenotypic Up-regulated Gene Support Vector Machine (PUGSVM) is a cancer Biomedical Informatics Grid (caBIG™) analytical tool for multiclass gene selection and classification. PUGSVM addresses the problem of imbalanced class separability, small sample size and high gene space dimensionality, where multiclass gene markers are defined by the union of one-versus-everyone phenotypic upregulated genes, and used by a well-matched one-versus-rest support vector machine. PUGSVM provides a simple yet more accurate strategy to identify statistically reproducible mechanistic marker genes for characterization of heterogeneous diseases. AVAILABILITY: http://www.cbil.ece.vt.edu/caBIG-PUGSVM.htm. Guoqiang Yu, Huai Li, Sook Shin Ha, Ie-Ming Shih, Robert Clarke, Eric P. Hoffman, Subha Madhavan, Jianhua Xuan, Yue Joseph Wang |
Bioinform. | 5 |
| 2011 | DDN: a caBIG® analytical tool for differential network analysisabstractUNLABELLED: Differential dependency network (DDN) is a caBIG® (cancer Biomedical Informatics Grid) analytical tool for detecting and visualizing statistically significant topological changes in transcriptional networks representing two biological conditions. Developed under caBIG®'s In Silico Research Centers of Excellence (ISRCE) Program, DDN enables differential network analysis and provides an alternative way for defining network biomarkers predictive of phenotypes. DDN also serves as a useful systems biology tool for users across biomedical research communities to infer how genetic, epigenetic or environment variables may affect biological networks and clinical phenotypes. Besides the standalone Java application, we have also developed a Cytoscape plug-in, CytoDDN, to integrate network analysis and visualization seamlessly. AVAILABILITY: The Java and MATLAB source code can be downloaded at the authors' web site http://www.cbil.ece.vt.edu/software.htm. Bai Zhang, Huai Li, Ie-Ming Shih, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Leena Hilakivi-Clarke, Yue Joseph Wang |
Bioinform. | 7 |
| 2011 | Motif-guided sparse decomposition of gene expression data for regulatory module identificationabstractBACKGROUND: Genes work coordinately as gene modules or gene networks. Various computational approaches have been proposed to find gene modules based on gene expression data; for example, gene clustering is a popular method for grouping genes with similar gene expression patterns. However, traditional gene clustering often yields unsatisfactory results for regulatory module identification because the resulting gene clusters are co-expressed but not necessarily co-regulated. RESULTS: We propose a novel approach, motif-guided sparse decomposition (mSD), to identify gene regulatory modules by integrating gene expression data and DNA sequence motif information. The mSD approach is implemented as a two-step algorithm comprising estimates of (1) transcription factor activity and (2) the strength of the predicted gene regulation event(s). Specifically, a motif-guided clustering method is first developed to estimate the transcription factor activity of a gene module; sparse component analysis is then applied to estimate the regulation strength, and so predict the target genes of the transcription factors. The mSD approach was first tested for its improved performance in finding regulatory modules using simulated and real yeast data, revealing functionally distinct gene modules enriched with biologically validated transcription factors. We then demonstrated the efficacy of the mSD approach on breast cancer cell line data and uncovered several important gene regulatory modules related to endocrine therapy of breast cancer. CONCLUSION: We have developed a new integrated strategy, namely motif-guided sparse decomposition (mSD) of gene expression data, for regulatory module identification. The mSD method features a novel motif-guided clustering method for transcription factor activity estimation by finding a balance between co-regulation and co-expression. The mSD method further utilizes a sparse decomposition method for regulation strength estimation. The experimental results show that such a motif-guided strategy can provide context-specific regulatory modules in both yeast and breast cancer studies. Jianhua Xuan, Li Chen 0018, Rebecca B. Riggins, Huai Li, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
BMC Bioinform. | 7 |
| 2010 | Module-based biomarker discovery in breast cancerabstractThe availability of genome-wide biological network data opens up new possibilities to discover novel biomarkers and elucidate cancer-related complex mechanisms at network level. In this paper, we propose a novel module-based feature selection framework, which integrates biological network information and gene expression data to identify biomarkers, not as individual genes but as functional modules. Also, a large-scale analysis of ensemble feature selection concept is presented. The method allows combining features selected from multiple runs with various data subsampling to increase the reliability and classification accuracy of the final set of selected features. The results from four breast cancer studies demonstrate that the identified module biomarkers achieve: i) higher classification accuracy in independent validation datasets; ii) better reproducibility than individual gene biomarkers; iii) improved biological interpretability; and iv) enhanced enrichment in cancer-related “disease drivers”. Yuji Zhang 0001, Jianhua Xuan, Robert Clarke, Habtom W. Ressom |
BIBM | 3 |
| 2010 | Identification of Transcriptional Regulatory Networks by Learning the Marginal Function of Outlier Sum StatisticabstractNetwork component analysis (NCA) and other methods based on the NCA model have become powerful bioinformatics tools to reconstruct underlying regulatory networks and recover hidden biological processes. However, due to the existence of experimental noises in micro array data and false information in network connectivity data (e.g., ChIP-on-chip binding data, motif information, etc.), it still remains challenging to reconstruct gene regulatory networks for real biomedical applications such as human cancer studies. In this paper, we model the relationship between the genes that share the same transcription factors (TF) from the angle of regression. We propose a statistic called outlier sum testing the conditional significance of the target genes. A Gibbs strategy is utilized in order to estimate the marginal value of outlier sum from its conditional function. Based on the outlier sum statistic we are able to extract the true target genes that carry information about transcription factor activities (TFAs) from the whole population. As a proof-of-concept, we demonstrated the efficiency and robustness of the proposed method on both simulation data and yeast cell cycle data. Jinghua Gu, Jianhua Xuan, Yue Joseph Wang, Rebecca B. Riggins, Robert Clarke |
ICMLA | 5 |
| 2010 | Multilevel support vector regression analysis to identify condition-specific regulatory networksabstractMOTIVATION: The identification of gene regulatory modules is an important yet challenging problem in computational biology. While many computational methods have been proposed to identify regulatory modules, their initial success is largely compromised by a high rate of false positives, especially when applied to human cancer studies. New strategies are needed for reliable regulatory module identification. RESULTS: We present a new approach, namely multilevel support vector regression (ml-SVR), to systematically identify condition-specific regulatory modules. The approach is built upon a multilevel analysis strategy designed for suppressing false positive predictions. With this strategy, a regulatory module becomes ever more significant as more relevant gene sets are formed at finer levels. At each level, a two-stage support vector regression (SVR) method is utilized to help reduce false positive predictions by integrating binding motif information and gene expression data; a significant analysis procedure is followed to assess the significance of each regulatory module. To evaluate the effectiveness of the proposed strategy, we first compared the ml-SVR approach with other existing methods on simulation data and yeast cell cycle data. The resulting performance shows that the ml-SVR approach outperforms other methods in the identification of both regulators and their target genes. We then applied our method to breast cancer cell line data to identify condition-specific regulatory modules associated with estrogen treatment. Experimental results show that our method can identify biologically meaningful regulatory modules related to estrogen signaling and action in breast cancer. AVAILABILITY AND IMPLEMENTATION: The ml-SVR MATLAB package can be downloaded at http://www.cbil.ece.vt.edu/software.htm. Li Chen 0018, Jianhua Xuan, Rebecca B. Riggins, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
Bioinform. | 6 |
| 2010 | Knowledge-guided gene ranking by coordinative component analysisabstractBACKGROUND: In cancer, gene networks and pathways often exhibit dynamic behavior, particularly during the process of carcinogenesis. Thus, it is important to prioritize those genes that are strongly associated with the functionality of a network. Traditional statistical methods are often inept to identify biologically relevant member genes, motivating researchers to incorporate biological knowledge into gene ranking methods. However, current integration strategies are often heuristic and fail to incorporate fully the true interplay between biological knowledge and gene expression data. RESULTS: To improve knowledge-guided gene ranking, we propose a novel method called coordinative component analysis (COCA) in this paper. COCA explicitly captures those genes within a specific biological context that are likely to be expressed in a coordinative manner. Formulated as an optimization problem to maximize the coordinative effort, COCA is designed to first extract the coordinative components based on a partial guidance from knowledge genes and then rank the genes according to their participation strengths. An embedded bootstrapping procedure is implemented to improve statistical robustness of the solutions. COCA was initially tested on simulation data and then on published gene expression microarray data to demonstrate its improved performance as compared to traditional statistical methods. Finally, the COCA approach has been applied to stem cell data to identify biologically relevant genes in signaling pathways. As a result, the COCA approach uncovers novel pathway members that may shed light into the pathway deregulation in cancers. CONCLUSION: We have developed a new integrative strategy to combine biological knowledge and microarray data for gene ranking. The method utilizes knowledge genes for a guidance to first extract coordinative components, and then rank the genes according to their contribution related to a network or pathway. The experimental results show that such a knowledge-guided strategy can provide context-specific gene ranking with an improved performance in pathway member identification. Chen Wang 0001, Jianhua Xuan, Huai Li, Yue Joseph Wang, Ming Zhan, Eric P. Hoffman, Robert Clarke |
BMC Bioinform. | 7 |
| 2010 | Matched Gene Selection and Committee Classifier for Molecular Classification of Heterogeneous Diseases
Guoqiang Yu, Yuanjian Feng, David J. Miller 0001, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Ben Davidson, Ie-Ming Shih, Yue Joseph Wang |
J. Mach. Learn. Res. | 6 |
| 2009 | Differential dependency network analysis to identify condition-specific topological changes in biological networksabstractMOTIVATION: Significant efforts have been made to acquire data under different conditions and to construct static networks that can explain various gene regulation mechanisms. However, gene regulatory networks are dynamic and condition-specific; under different conditions, networks exhibit different regulation patterns accompanied by different transcriptional network topologies. Thus, an investigation on the topological changes in transcriptional networks can facilitate the understanding of cell development or provide novel insights into the pathophysiology of certain diseases, and help identify the key genetic players that could serve as biomarkers or drug targets. RESULTS: Here, we report a differential dependency network (DDN) analysis to detect statistically significant topological changes in the transcriptional networks between two biological conditions. We propose a local dependency model to represent the local structures of a network by a set of conditional probabilities. We develop an efficient learning algorithm to learn the local dependency model using the Lasso technique. A permutation test is subsequently performed to estimate the statistical significance of each learned local structure. In testing on a simulation dataset, the proposed algorithm accurately detected all the genes with network topological changes. The method was then applied to the estrogen-dependent T-47D estrogen receptor-positive (ER+) breast cancer cell line datasets and human and mouse embryonic stem cell datasets. In both experiments using real microarray datasets, the proposed method produced biologically meaningful results. We expect DDN to emerge as an important bioinformatics tool in transcriptional network analyses. While we focus specifically on transcriptional networks, the DDN method we introduce here is generally applicable to other biological networks with similar characteristics. AVAILABILITY: The DDN MATLAB toolbox and experiment data are available at http://www.cbil.ece.vt.edu/software.htm. Bai Zhang, Huai Li, Rebecca B. Riggins, Ming Zhan, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
Bioinform. | 8 |
| 2008 | Annotating breast cancer microarray samples using ontologies
Victoria Y. Yoon, Robert Clarke |
AMIA | 4 |
| 2008 | Network-Constrained Support Vector Machine for ClassificationabstractOne of the major goals in microarray data analysis is to identify biomarkers and build a classification model for future prediction. Many traditional statistical models, based on microarray data alone, often fail in identifying biologically meaningful genes, which should have synergistic effect on determine the clinical outcomes through some interactions rather than work individually. In this paper, we proposed a network-constrained support vector machine (nSVM) for classification by incorporating prior knowledge, which could be protein-protein interactions, protein-gene regulation relationships or pathways information. Specifically, we use Laplacian matrix to represent gene-gene interaction network to regularize the objective function of SVM, which imposes the smoothness of coefficients over the network. The experimental results on simulation and real microarray datasets demonstrate that our method could not only improve classification performance compared to conventional SVM, but more importantly, it could identify significant sub-networks belonging to several pathways which might be related to underlying mechanism associated with clinical outcomes. Li Chen 0018, Jianhua Xuan, Yue Joseph Wang, Rebecca B. Riggins, Robert Clarke |
ICMLA | 5 |
| 2008 | Sparse Decomposition of Gene Expression Data to Infer Transcriptional Modules Guided by Motif Information
Jianhua Xuan, Li Chen 0018, Rebecca B. Riggins, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
ISBRA | 7 |
| 2008 | Integrative Network Component Analysis for Regulatory Network Reconstruction
Chen Wang 0001, Jianhua Xuan, Li Chen 0018, Po Zhao, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman |
ISBRA | 6 |
| 2008 | Knowledge-guided multi-scale independent component analysis for biomarker identificationabstractBACKGROUND: Many statistical methods have been proposed to identify disease biomarkers from gene expression profiles. However, from gene expression profile data alone, statistical methods often fail to identify biologically meaningful biomarkers related to a specific disease under study. In this paper, we develop a novel strategy, namely knowledge-guided multi-scale independent component analysis (ICA), to first infer regulatory signals and then identify biologically relevant biomarkers from microarray data. RESULTS: Since gene expression levels reflect the joint effect of several underlying biological functions, disease-specific biomarkers may be involved in several distinct biological functions. To identify disease-specific biomarkers that provide unique mechanistic insights, a meta-data "knowledge gene pool" (KGP) is first constructed from multiple data sources to provide important information on the likely functions (such as gene ontology information) and regulatory events (such as promoter responsive elements) associated with potential genes of interest. The gene expression and biological meta data associated with the members of the KGP can then be used to guide subsequent analysis. ICA is then applied to multi-scale gene clusters to reveal regulatory modes reflecting the underlying biological mechanisms. Finally disease-specific biomarkers are extracted by their weighted connectivity scores associated with the extracted regulatory modes. A statistical significance test is used to evaluate the significance of transcription factor enrichment for the extracted gene set based on motif information. We applied the proposed method to yeast cell cycle microarray data and Rsf-1-induced ovarian cancer microarray data. The results show that our knowledge-guided ICA approach can extract biologically meaningful regulatory modes and outperform several baseline methods for biomarker identification. CONCLUSION: We have proposed a novel method, namely knowledge-guided multi-scale ICA, to identify disease-specific biomarkers. The goal is to infer knowledge-relevant regulatory signals and then identify corresponding biomarkers through a multi-scale strategy. The approach has been successfully applied to two expression profiling experiments to demonstrate its improved performance in extracting biologically meaningful and disease-related biomarkers. More importantly, the proposed approach shows promising results to infer novel biomarkers for ovarian cancer and extend current knowledge. Li Chen 0018, Jianhua Xuan, Chen Wang 0001, Ie-Ming Shih, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
BMC Bioinform. | 8 |
| 2008 | Motif-directed network component analysis for regulatory network inferenceabstractBACKGROUND: Network Component Analysis (NCA) has shown its effectiveness in discovering regulators and inferring transcription factor activities (TFAs) when both microarray data and ChIP-on-chip data are available. However, a NCA scheme is not applicable to many biological studies due to limited topology information available, such as lack of ChIP-on-chip data. We propose a new approach, motif-directed NCA (mNCA), to integrate motif information and gene expression data to infer regulatory networks. RESULTS: We develop motif-directed NCA (mNCA) to incorporate motif information into NCA for regulatory network inference. While motif information is readily available from knowledge databases, it is a "noisy" source of network topology information consisting of many false positives. To overcome this problem, we develop a stability analysis procedure embedded in mNCA to resolve the inconsistency between motif information and gene expression data, and to enable the identification of stable TFAs. The mNCA approach has been applied to a time course microarray data set of muscle regeneration. The experimental results show that the inferred TFAs are not only numerically stable but also biologically relevant to muscle differentiation process. In particular, several inferred TFAs like those of MyoD, myogenin and YY1 are well supported by biological experiments. CONCLUSION: A novel computational approach, mNCA, has been developed to integrate motif information and gene expression data for regulatory network reconstruction. Specifically, motif analysis is used to obtain initial network topology, and stability analysis is developed and applied with mNCA to extract stable TFAs. Experimental results on muscle regeneration microarray data have demonstrated that mNCA is a practical and reliable computational method for regulatory network inference and pathway discovery. Chen Wang 0001, Jianhua Xuan, Li Chen 0018, Po Zhao, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman |
BMC Bioinform. | 6 |
| 2008 | Network motif-based identification of transcription factor-target gene relationships by integrating multi-source biological dataabstractBACKGROUND: Integrating data from multiple global assays and curated databases is essential to understand the spatio-temporal interactions within cells. Different experiments measure cellular processes at various widths and depths, while databases contain biological information based on established facts or published data. Integrating these complementary datasets helps infer a mutually consistent transcriptional regulatory network (TRN) with strong similarity to the structure of the underlying genetic regulatory modules. Decomposing the TRN into a small set of recurring regulatory patterns, called network motifs (NM), facilitates the inference. Identifying NMs defined by specific transcription factors (TF) establishes the framework structure of a TRN and allows the inference of TF-target gene relationship. This paper introduces a computational framework for utilizing data from multiple sources to infer TF-target gene relationships on the basis of NMs. The data include time course gene expression profiles, genome-wide location analysis data, binding sequence data, and gene ontology (GO) information. RESULTS: The proposed computational framework was tested using gene expression data associated with cell cycle progression in yeast. Among 800 cell cycle related genes, 85 were identified as candidate TFs and classified into four previously defined NMs. The NMs for a subset of TFs are obtained from literature. Support vector machine (SVM) classifiers were used to estimate NMs for the remaining TFs. The potential downstream target genes for the TFs were clustered into 34 biologically significant groups. The relationships between TFs and potential target gene clusters were examined by training recurrent neural networks whose topologies mimic the NMs to which the TFs are classified. The identified relationships between TFs and gene clusters were evaluated using the following biological validation and statistical analyses: (1) Gene set enrichment analysis (GSEA) to evaluate the clustering results; (2) Leave-one-out cross-validation (LOOCV) to ensure that the SVM classifiers assign TFs to NM categories with high confidence; (3) Binding site enrichment analysis (BSEA) to determine enrichment of the gene clusters for the cognate binding sites of their predicted TFs; (4) Comparison with previously reported results in the literatures to confirm the inferred regulations. CONCLUSION: The major contribution of this study is the development of a computational framework to assist the inference of TRN by integrating heterogeneous data from multiple sources and by decomposing a TRN into NM-based modules. The inference capability of the proposed framework is verified statistically (e.g., LOOCV) and biologically (e.g., GSEA, BSEA, and literature validation). The proposed framework is useful for inferring small NM-based modules of TF-target gene relationships that can serve as a basis for generating new testable hypotheses. Yuji Zhang 0001, Jianhua Xuan, Benildo de los Reyes, Robert Clarke, Habtom W. Ressom |
BMC Bioinform. | 4 |
| 2008 | caBIGTM VISDA: Modeling, visualization, and discovery for cluster analysis of genomic dataabstractBACKGROUND: The main limitations of most existing clustering methods used in genomic data analysis include heuristic or random algorithm initialization, the potential of finding poor local optima, the lack of cluster number detection, an inability to incorporate prior/expert knowledge, black-box and non-adaptive designs, in addition to the curse of dimensionality and the discernment of uninformative, uninteresting cluster structure associated with confounding variables. RESULTS: In an effort to partially address these limitations, we develop the VIsual Statistical Data Analyzer (VISDA) for cluster modeling, visualization, and discovery in genomic data. VISDA performs progressive, coarse-to-fine (divisive) hierarchical clustering and visualization, supported by hierarchical mixture modeling, supervised/unsupervised informative gene selection, supervised/unsupervised data visualization, and user/prior knowledge guidance, to discover hidden clusters within complex, high-dimensional genomic data. The hierarchical visualization and clustering scheme of VISDA uses multiple local visualization subspaces (one at each node of the hierarchy) and consequent subspace data modeling to reveal both global and local cluster structures in a "divide and conquer" scenario. Multiple projection methods, each sensitive to a distinct type of clustering tendency, are used for data visualization, which increases the likelihood that cluster structures of interest are revealed. Initialization of the full dimensional model is based on first learning models with user/prior knowledge guidance on data projected into the low-dimensional visualization spaces. Model order selection for the high dimensional data is accomplished by Bayesian theoretic criteria and user justification applied via the hierarchy of low-dimensional visualization subspaces. Based on its complementary building blocks and flexible functionality, VISDA is generally applicable for gene clustering, sample clustering, and phenotype clustering (wherein phenotype labels for samples are known), albeit with minor algorithm modifications customized to each of these tasks. CONCLUSION: VISDA achieved robust and superior clustering accuracy, compared with several benchmark clustering schemes. The model order selection scheme in VISDA was shown to be effective for high dimensional genomic data clustering. On muscular dystrophy data and muscle regeneration data, VISDA identified biologically relevant co-expressed gene clusters. VISDA also captured the pathological relationships among different phenotypes revealed at the molecular level, through phenotype clustering on muscular dystrophy data and multi-category cancer data. Yitan Zhu, Huai Li, David J. Miller 0001, Zuyi Wang, Jianhua Xuan, Robert Clarke, Eric P. Hoffman, Yue Joseph Wang |
BMC Bioinform. | 6 |
| 2007 | Biomarker Identification by Knowledge-Driven Multi-Level ICA and Motif AnalysisabstractMany statistical methods often fail to identify biologically meaningful biomarkers related to a specific disease under study from expression data alone. In this paper, we develop a novel strategy, namely knowledge-driven multi-level independent component analysis (ICA), to infer regulatory signals and identify biologically relevant biomarkers from microarray data. Specifically, based on multi-level clustering results and partial prior knowledge, we apply ICA to find stable disease specific linear regulatory modes and then extract associated biomarker genes. A statistical test is designed to evaluate the significance of transcription factor enrichment for extracted gene set based on motif information. The experimental results on an Rsf-1 induced microarray data set show that our knowledge-driven method can extract more biologically meaningful biomarkers with significant enrichment of transcription factors related to ovarian cancer compared to other gene selection methods with/without prior knowledge. Li Chen 0018, Chen Wang 0001, Ie-Ming Shih, Tian-Li Wang, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman, Jianhua Xuan |
ICMLA | 7 |
| 2007 | VISDA: an open-source caBIGTM analytical tool for data clustering and beyondabstractSUMMARY: VISDA (Visual Statistical Data Analyzer) is a caBIG analytical tool for cluster modeling, visualization and discovery that has met silver-level compatibility under the caBIG initiative. Being statistically principled and visually interfaced, VISDA exploits both hierarchical statistics modeling and human gift for pattern recognition to allow a progressive yet interactive discovery of hidden clusters within high dimensional and complex biomedical datasets. The distinctive features of VISDA are particularly useful for users across the cancer research and broader research communities to analyze complex biological data. AVAILABILITY: http://gforge.nci.nih.gov/projects/visda/ Jiajing Wang, Huai Li, Yitan Zhu, Malik Yousef, Michael Nebozhyn, Michael M. Showe, Louise C. Showe, Jianhua Xuan, Robert Clarke, Yue Joseph Wang |
Bioinform. | 9 |
| 2006 | ModVis: An information visualization tool for gene module discovery
Justin Molineaux, Jianhua Xuan, Yitan Zhu, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
CAINE | 6 |
| 2006 | Inference of Gene Regulatory Networks from Time Course Gene Expression Data Using Neural Networks and Swarm IntelligenceabstractWe present a novel algorithm that combines a recurrent neural network (RNN) and two swarm intelligence (SI) methods to infer a gene regulatory network (GRN) from time course gene expression data. The algorithm uses ant colony optimization (ACO) to identify the optimal architecture of an RNN, while the weights of the RNN are optimized using particle swarm optimization (PSO). Our goal is to construct an RNN whose response mimics gene expression data generated by time course DNA microarray experiments. We observed promising results in applying the proposed hybrid SI-RNN algorithm to infer networks of interaction from simulated and real-world gene expression data Habtom W. Ressom, Yuji Zhang 0001, Jianhua Xuan, Yue Joseph Wang, Robert Clarke |
CIBCB | 5 |
| 2006 | Optimized multilayer perceptrons for molecular classification and diagnosis using genomic dataabstractMOTIVATION: Multilayer perceptrons (MLP) represent one of the widely used and effective machine learning methods currently applied to diagnostic classification based on high-dimensional genomic data. Since the dimensionalities of the existing genomic data often exceed the available sample sizes by orders of magnitude, the MLP performance may degrade owing to the curse of dimensionality and over-fitting, and may not provide acceptable prediction accuracy. RESULTS: Based on Fisher linear discriminant analysis, we designed and implemented an MLP optimization scheme for a two-layer MLP that effectively optimizes the initialization of MLP parameters and MLP architecture. The optimized MLP consistently demonstrated its ability in easing the curse of dimensionality in large microarray datasets. In comparison with a conventional MLP using random initialization, we obtained significant improvements in major performance measures including Bayes classification accuracy, convergence properties and area under the receiver operating characteristic curve (A(z)). SUPPLEMENTARY INFORMATION: The Supplementary information is available on http://www.cbil.ece.vt.edu/publications.htm Zuyi Wang, Yue Joseph Wang, Jianhua Xuan, Yibin Dong, Marina Bakay, Yuanjian Feng, Robert Clarke, Eric P. Hoffman |
Bioinform. | 7 |
| 2005 | Normalization of Microarray Data by Iterative Nonlinear RegressionabstractNormalization is an important prerequisite for almost all follow-up microarray data analysis steps. Accurate normalization assures a common base for comparative biomedical studies using gene expression profiles across different experiments and phenotypes. In this paper, we present a novel normalization approach - iterative nonlinear regression (INR) method - that exploits concurrent identification of invariantly expressed genes (IEGs) and implementation of nonlinear regression normalization. We demonstrate the principle and performance of the INR approach on two real microarray data sets. As compared to major peer methods (e.g., linear regression method, Loess method and iterative ranking method), INR method shows a superior performance in achieving low expression variance across replicates and excellent fold change preservation. Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
BIBE | 3 |
| 2005 | Secure Mutual Distrust Transaction Tracking Using Cryptographic Elements
Angela S. L. Wong, Matthew Sorell, Robert Clarke |
IWDW | 3 |
| 2004 | Gene selection in class space for molecular classification of cancer
Yue Joseph Wang, Javed I. Khan, Robert Clarke |
Sci. China Ser. F Inf. Sci. | 4 |
| 2002 | Iterative normalization of cDNA microarray dataabstractThis paper describes a new approach to normalizing microarray expression data. The novel feature is to unify the tasks of estimating normalization coefficients and identifying control gene set. Unification is realized by constructing a window function over the scatter plot defining the subset of constantly expressed genes and by affecting optimization using an iterative procedure. The structure of window function gates contributions to the control gene set used to estimate normalization coefficients. This window measures the consistency of the matched neighborhoods in the scatter plot and provides a means of rejecting control gene outliers. The recovery of normalizational regression and control gene selection are interleaved and are realized by applying coupled operations to the mean square error function. In this way, the two processes bootstrap one another. We evaluate the technique on real microarray data from breast cancer cell lines and complement the experiment with a data cluster visualization study. Yue Joseph Wang, Jianping Lu, Richard Lee 0002, Zhiping Gu, Robert Clarke |
IEEE Trans. Inf. Technol. Biomed. | 5 |