VLDB 2026 Research / reviewers in the wild / expert
Jianhua Xuan
dblp:77/6937 · also Jianhua Jason Xuan
· DBLP profile ↗
59ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0001-7256-1374ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 43 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-authorArtificial intelligence and machine learning · 9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bayesian identification of differentially expressed isoforms using a novel joint model of RNA-seq dataabstractWe develop a Bayesian approach, BayesIso, to identify differentially expressed isoforms from RNA-seq data. The approach features a novel joint model of the sample variability and the deferential state of isoforms. Specifically, the within-sample variability and the between-sample variability of each isoform are modeled by a Poisson-Lognormal model and a Gamma-Gamma model, respectively. Using a Bayesian framework, the differential state of each isoform and the model parameters are jointly estimated by a Markov Chain Monte Carlo (MCMC) method. Extensive studies using simulation and real data demonstrate that BayesIso can effectively detect isoforms of less differentially expressed and differential transcripts for genes with multiple isoforms. We applied the approach to breast cancer RNA-seq data and uncovered a unique set of isoforms that form key pathways associated with breast cancer recurrence. First, PI3K/AKT/mTOR signaling and PTEN signaling pathways are identified as being involved in breast cancer development. Further integrated with protein-protein interaction data, pathways of Jak-STAT, mTOR, MAPK and Wnt signaling are revealed in association with breast cancer recurrence. Finally, several pathways are activated in the early recurrence of breast cancer. In tumors that occur early, members of pathways of cellular metabolism and cell cycle (such as CD36 and TOP2A) are upregulated, while immune response genes such as NFATC1 are downregulated. Xu Shi 0003, Xiao Wang 0031, Leena Halakivi-Clarke, Robert Clarke, Andrew F. Neuwald, Jianhua Xuan |
PLoS Comput. Biol. | 7 |
| 2021 | IntAPT: integrated assembly of phenotype-specific transcripts from multiple RNA-seq profilesabstractMOTIVATION: High-throughput RNA sequencing has revolutionized the scope and depth of transcriptome analysis. Accurate reconstruction of a phenotype-specific transcriptome is challenging due to the noise and variability of RNA-seq data. This requires computational identification of transcripts from multiple samples of the same phenotype, given the underlying consensus transcript structure. RESULTS: We present a Bayesian method, integrated assembly of phenotype-specific transcripts (IntAPT), that identifies phenotype-specific isoforms from multiple RNA-seq profiles. IntAPT features a novel two-layer Bayesian model to capture the presence of isoforms at the group layer and to quantify the abundance of isoforms at the sample layer. A spike-and-slab prior is used to model the isoform expression and to enforce the sparsity of expressed isoforms. Dependencies between the existence of isoforms and their expression are modeled explicitly to facilitate parameter estimation. Model parameters are estimated iteratively using Gibbs sampling to infer the joint posterior distribution, from which the presence and abundance of isoforms can reliably be determined. Studies using both simulations and real datasets show that IntAPT consistently outperforms existing methods for the IntAPT. Experimental results demonstrate that, despite sequencing errors, IntAPT exhibits a robust performance among multiple samples, resulting in notably improved identification of expressed isoforms of low abundance. AVAILABILITY AND IMPLEMENTATION: The IntAPT package is available at http://github.com/henryxushi/IntAPT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xu Shi 0003, Andrew F. Neuwald, Xiao Wang 0031, Tian-Li Wang, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 7 |
| 2021 | ChIP-BIT2: a software tool to detect weak binding events using a Bayesian integration approachabstractBACKGROUND: ChIP-seq combines chromatin immunoprecipitation assays with sequencing and identifies genome-wide binding sites for DNA binding proteins. While many binding sites have strong ChIP-seq 'peak' observations and are well captured, there are still regions bound by proteins weakly, with a relatively low ChIP-seq signal enrichment. These weak binding sites, especially those at promoters and enhancers, are functionally important because they also regulate nearby gene expression. Yet, it remains a challenge to accurately identify weak binding sites in ChIP-seq data due to the ambiguity in differentiating these weak binding sites from the amplified background DNAs. RESULTS: ChIP-BIT2 ( http://sourceforge.net/projects/chipbitc/ ) is a software package for ChIP-seq peak detection. ChIP-BIT2 employs a mixture model integrating protein and control ChIP-seq data and predicts strong or weak protein binding sites at promoters, enhancers, or other genomic locations. For binding sites at gene promoters, ChIP-BIT2 simultaneously predicts their target genes. ChIP-BIT2 has been validated on benchmark regions and tested using large-scale ENCODE ChIP-seq data, demonstrating its high accuracy and wide applicability. CONCLUSION: ChIP-BIT2 is an efficient ChIP-seq peak caller. It provides a better lens to examine weak binding sites and can refine or extend the existing binding site collection, providing additional regulatory regions for decoding the mechanism of gene expression regulation. Xi Chen 0056, Xu Shi 0003, Andrew F. Neuwald, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
BMC Bioinform. | 6 |
| 2021 | ChIP-GSM: Inferring active transcription factor modules to predict functional regulatory elementsabstractTranscription factors (TFs) often function as a module including both master factors and mediators binding at cis-regulatory regions to modulate nearby gene transcription. ChIP-seq profiling of multiple TFs makes it feasible to infer functional TF modules. However, when inferring TF modules based on co-localization of ChIP-seq peaks, often many weak binding events are missed, especially for mediators, resulting in incomplete identification of modules. To address this problem, we develop a ChIP-seq data-driven Gibbs Sampler to infer Modules (ChIP-GSM) using a Bayesian framework that integrates ChIP-seq profiles of multiple TFs. ChIP-GSM samples read counts of module TFs iteratively to estimate the binding potential of a module to each region and, across all regions, estimates the module abundance. Using inferred module-region probabilistic bindings as feature units, ChIP-GSM then employs logistic regression to predict active regulatory elements. Validation of ChIP-GSM predicted regulatory regions on multiple independent datasets sharing the same context confirms the advantage of using TF modules for predicting regulatory activity. In a case study of K562 cells, we demonstrate that the ChIP-GSM inferred modules form as groups, activate gene expression at different time points, and mediate diverse functional cellular processes. Hence, ChIP-GSM infers biologically meaningful TF modules and improves the prediction accuracy of regulatory region activities. Xi Chen 0056, Andrew F. Neuwald, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
PLoS Comput. Biol. | 5 |
| 2018 | CRNET: an efficient sampling approach to infer functional regulatory networks by integrating large-scale ChIP-seq and time-course RNA-seq dataabstractMotivation: NGS techniques have been widely applied in genetic and epigenetic studies. Multiple ChIP-seq and RNA-seq profiles can now be jointly used to infer functional regulatory networks (FRNs). However, existing methods suffer from either oversimplified assumption on transcription factor (TF) regulation or slow convergence of sampling for FRN inference from large-scale ChIP-seq and time-course RNA-seq data. Results: We developed an efficient Bayesian integration method (CRNET) for FRN inference using a two-stage Gibbs sampler to estimate iteratively hidden TF activities and the posterior probabilities of binding events. A novel statistic measure that jointly considers regulation strength and regression error enables the sampling process of CRNET to converge quickly, thus making CRNET very efficient for large-scale FRN inference. Experiments on synthetic and benchmark data showed a significantly improved performance of CRNET when compared with existing methods. CRNET was applied to breast cancer data to identify FRNs functional at promoter or enhancer regions in breast cancer MCF-7 cells. Transcription factor MYC is predicted as a key functional factor in both promoter and enhancer FRNs. We experimentally validated the regulation effects of MYC on CRNET-predicted target genes using appropriate RNAi approaches in MCF-7 cells. Availability and implementation: R scripts of CRNET are available at http://www.cbil.ece.vt.edu/software.htm. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xi Chen 0056, Jinghua Gu, Xiao Wang 0031, Jin-Gyoung Jung, Tian-Li Wang, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 8 |
| 2018 | SparseIso: a novel Bayesian approach to identify alternatively spliced isoforms from RNA-seq dataabstractMotivation: Recent advances in high-throughput RNA sequencing (RNA-seq) technologies have made it possible to reconstruct the full transcriptome of various types of cells. It is important to accurately assemble transcripts or identify isoforms for an improved understanding of molecular mechanisms in biological systems. Results: We have developed a novel Bayesian method, SparseIso, to reliably identify spliced isoforms from RNA-seq data. A spike-and-slab prior is incorporated into the Bayesian model to enforce the sparsity for isoform identification, effectively alleviating the problem of overfitting. A Gibbs sampling procedure is further developed to simultaneously identify and quantify transcripts from RNA-seq data. With the sampling approach, SparseIso estimates the joint distribution of all candidate transcripts, resulting in a significantly improved performance in detecting lowly expressed transcripts and multiple expressed isoforms of genes. Both simulation study and real data analysis have demonstrated that the proposed SparseIso method significantly outperforms existing methods for improved transcript assembly and isoform identification. Availability and implementation: The SparseIso package is available at http://github.com/henryxushi/SparseIso. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xu Shi 0003, Xiao Wang 0031, Tian-Li Wang, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 6 |
| 2017 | PSSV: a novel pattern-based probabilistic approach for somatic structural variation identificationabstractMOTIVATION: Whole genome DNA-sequencing (WGS) of paired tumor and normal samples has enabled the identification of somatic DNA changes in an unprecedented detail. Large-scale identification of somatic structural variations (SVs) for a specific cancer type will deepen our understanding of driver mechanisms in cancer progression. However, the limited number of WGS samples, insufficient read coverage, and the impurity of tumor samples that contain normal and neoplastic cells, limit reliable and accurate detection of somatic SVs. RESULTS: We present a novel pattern-based probabilistic approach, PSSV, to identify somatic structural variations from WGS data. PSSV features a mixture model with hidden states representing different mutation patterns; PSSV can thus differentiate heterozygous and homozygous SVs in each sample, enabling the identification of those somatic SVs with heterozygous mutations in normal samples and homozygous mutations in tumor samples. Simulation studies demonstrate that PSSV outperforms existing tools. PSSV has been successfully applied to breast cancer data to identify somatic SVs of key factors associated with breast cancer development. AVAILABILITY AND IMPLEMENTATION: An R package of PSSV is available at http://www.cbil.ece.vt.edu/software.htm CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Xi Chen 0056, Xu Shi 0003, Leena Hilakivi-Clarke, Ayesha N. Shajahan, Robert Clarke, Jianhua Xuan |
Bioinform. | 6 |
| 2017 | DM-BLD: differential methylation detection using a hierarchical Bayesian model exploiting local dependencyabstractMOTIVATION: The advent of high-throughput DNA methylation profiling techniques has enabled the possibility of accurate identification of differentially methylated genes for cancer research. The large number of measured loci facilitates whole genome methylation study, yet posing great challenges for differential methylation detection due to the high variability in tumor samples. RESULTS: We have developed a novel probabilistic approach, D: ifferential M: ethylation detection using a hierarchical B: ayesian model exploiting L: ocal D: ependency (DM-BLD), to detect differentially methylated genes based on a Bayesian framework. The DM-BLD approach features a joint model to capture both the local dependency of measured loci and the dependency of methylation change in samples. Specifically, the local dependency is modeled by Leroux conditional autoregressive structure; the dependency of methylation changes is modeled by a discrete Markov random field. A hierarchical Bayesian model is developed to fully take into account the local dependency for differential analysis, in which differential states are embedded as hidden variables. Simulation studies demonstrate that DM-BLD outperforms existing methods for differential methylation detection, particularly when the methylation change is moderate and the variability of methylation in samples is high. DM-BLD has been applied to breast cancer data to identify important methylated genes (such as polycomb target genes and genes involved in transcription factor activity) associated with breast cancer recurrence. AVAILABILITY AND IMPLEMENTATION: A Matlab package of DM-BLD is available at http://www.cbil.ece.vt.edu/software.htm CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Xiao Wang 0031, Jinghua Gu, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
Bioinform. | 5 |
| 2015 | BMRF-Net: a software tool for identification of protein interaction subnetworks by a bagging Markov random field-based methodabstractUNLABELLED: Identification of protein interaction subnetworks is an important step to help us understand complex molecular mechanisms in cancer. In this paper, we develop a BMRF-Net package, implemented in Java and C++, to identify protein interaction subnetworks based on a bagging Markov random field (BMRF) framework. By integrating gene expression data and protein-protein interaction data, this software tool can be used to identify biologically meaningful subnetworks. A user friendly graphic user interface is developed as a Cytoscape plugin for the BMRF-Net software to deal with the input/output interface. The detailed structure of the identified networks can be visualized in Cytoscape conveniently. The BMRF-Net package has been applied to breast cancer data to identify significant subnetworks related to breast cancer recurrence. AVAILABILITY AND IMPLEMENTATION: The BMRF-Net package is available at http://sourceforge.net/projects/bmrfcjava/. The package is tested under Ubuntu 12.04 (64-bit), Java 7, glibc 2.15 and Cytoscape 3.1.0. Xu Shi 0003, Robert O. Barnes, Li Chen 0018, Ayesha N. Shajahan, Leena Hilakivi-Clarke, Robert Clarke, Yue Joseph Wang, Jianhua Xuan |
Bioinform. | 8 |
| 2015 | KDDN: an open-source Cytoscape app for constructing differential dependency networks with significant rewiringabstractUNLABELLED: We have developed an integrated molecular network learning method, within a well-grounded mathematical framework, to construct differential dependency networks with significant rewiring. This knowledge-fused differential dependency networks (KDDN) method, implemented as a Java Cytoscape app, can be used to optimally integrate prior biological knowledge with measured data to simultaneously construct both common and differential networks, to quantitatively assign model parameters and significant rewiring p-values and to provide user-friendly graphical results. The KDDN algorithm is computationally efficient and provides users with parallel computing capability using ubiquitous multi-core machines. We demonstrate the performance of KDDN on various simulations and real gene expression datasets, and further compare the results with those obtained by the most relevant peer methods. The acquired biologically plausible results provide new insights into network rewiring as a mechanistic principle and illustrate KDDN's ability to detect them efficiently and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY: Source code and compiled package are freely available for download at http://apps.cytoscape.org/apps/kddn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bai Zhang, Eric P. Hoffman, Robert Clarke, Ie-Ming Shih, Jianhua Xuan, David M. Herrington, Yue Joseph Wang |
Bioinform. | 7 |
| 2015 | UNDO: a Bioconductor R package for unsupervised deconvolution of mixed gene expressions in tumor samplesabstractSUMMARY: We develop a novel unsupervised deconvolution method, within a well-grounded mathematical framework, to dissect mixed gene expressions in heterogeneous tumor samples. We implement an R package, UNsupervised DecOnvolution (UNDO), that can be used to automatically detect cell-specific marker genes (MGs) located on the scatter radii of mixed gene expressions, estimate cellular proportions in each sample and deconvolute mixed expressions into cell-specific expression profiles. We demonstrate the performance of UNDO over a wide range of tumor-stroma mixing proportions, validate UNDO on various biologically mixed benchmark gene expression datasets and further estimate tumor purity in TCGA/CPTAC datasets. The highly accurate deconvolution results obtained suggest not only the existence of cell-specific MGs but also UNDO's ability to detect them blindly and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY AND IMPLEMENTATION: UNDO is available at http://bioconductor.org/packages. Niya Wang, Robert Clarke, Lulu Chen, Ie-Ming Shih, Douglas A. Levine, Jianhua Xuan, Yue Joseph Wang |
Bioinform. | 8 |
| 2014 | A Markov random field-based Bayesian model to identify genes with differential methylationabstractThe rapid development of biotechnology makes it possible to explore genome-wide DNA methylation mapping which has been demonstrated to be related to diseases including cancer. However, it also posts substantial challenges in identifying biologically meaningful methylation pattern changes. Several algorithms have been proposed to detect differential methylation events, such as differentially methylated CpG sites and differentially methylated regions. However, the intrinsic dependency of the CpG sites in a neighboring area has not yet been fully considered. In this paper, we propose a novel method for the identification of differentially methylated genes in a Markov random field-based Bayesian framework. Specifically, we use Markov random field to model the dependency of the neighboring CpG sites, and then estimate the differential methylation score of the CpG sites in a Bayesian framework through a sampling scheme. Finally, the differential methylation statuses of the genes are determined by the estimated scores of the involved CpG sites. In addition, significance test is conducted to assess the significance of the identified differentially methylated genes. Experimental results on both synthetic data and real data demonstrate the effectiveness of the proposed method in identifying genes with differential methylation patterns under different conditions. Xiao Wang 0031, Jinghua Gu, Jianhua Xuan, Robert Clarke, Leena Hilakivi-Clarke |
CIBCB | 3 |
| 2014 | Robust identification of transcriptional regulatory networks using a Gibbs sampler on outlier sum statisticabstractContact: [email protected] Bioinformatics (2012) 28 (15), 1990–1997 doi:10.1093/bioinformatics/bts296 The authors wish to add one citation to a relevant conference report, Gu, J., Xuan, J., Wang, Y., Riggins, R.B. and Clarke R. (2010) Identification of transcriptional regulatory networks by learning the marginal function of outlier sum statistic. Proceedings of International Conference on Machine Learning and Applications , 281–286. The formatted reference is given below and should read in the sentence: In particular, a novel statistic for testing the confidence of target genes, namely, outlier sum of regression t -statistic ( Gu et al. , 2010 ), is specifically designed to pin-down confident target genes; based on this statistic, a Gibbs sampling strategy is used to sample target genes in a high probability as governed by the underlying distribution. The authors apologize for this oversight. Jinghua Gu, Jianhua Xuan, Rebecca B. Riggins, Li Chen 0018, Yue Joseph Wang, Robert Clarke |
Bioinform. | 2 |
| 2014 | BADGE: A novel Bayesian model for accurate abundance quantification and differential analysis of RNA-Seq dataabstractBACKGROUND: Recent advances in RNA sequencing (RNA-Seq) technology have offered unprecedented scope and resolution for transcriptome analysis. However, precise quantification of mRNA abundance and identification of differentially expressed genes are complicated due to biological and technical variations in RNA-Seq data. RESULTS: We systematically study the variation in count data and dissect the sources of variation into between-sample variation and within-sample variation. A novel Bayesian framework is developed for joint estimate of gene level mRNA abundance and differential state, which models the intrinsic variability in RNA-Seq to improve the estimation. Specifically, a Poisson-Lognormal model is incorporated into the Bayesian framework to model within-sample variation; a Gamma-Gamma model is then used to model between-sample variation, which accounts for over-dispersion of read counts among multiple samples. Simulation studies, where sequencing counts are synthesized based on parameters learned from real datasets, have demonstrated the advantage of the proposed method in both quantification of mRNA abundance and identification of differentially expressed genes. Moreover, performance comparison on data from the Sequencing Quality Control (SEQC) Project with ERCC spike-in controls has shown that the proposed method outperforms existing RNA-Seq methods in differential analysis. Application on breast cancer dataset has further illustrated that the proposed Bayesian model can 'blindly' estimate sources of variation caused by sequencing biases. CONCLUSIONS: We have developed a novel Bayesian hierarchical approach to investigate within-sample and between-sample variations in RNA-Seq data. Simulation and real data applications have validated desirable performance of the proposed method. The software package is available at http://www.cbil.ece.vt.edu/software.htm. Jinghua Gu, Xiao Wang 0031, Leena Hilakivi-Clarke, Robert Clarke, Jianhua Xuan |
BMC Bioinform. | 5 |
| 2013 | A novel statistical approach to identify co-regulatory gene modulesabstractChlP-chip experiments are performed to determine binding sites for transcription factors (TFs). Conventional TF-gene regulation is generated based on p-value cutoff of the binding sites as well as their distance to nearest genes. Taking into account that binding sites of one ChlP-chip experiment should follow the same specific location distribution, we proposed a statistical model using both location and significance information to weigh target genes. With multiple ChlP-chip experiments and gene expression data, we identified co-regulatory and differentially expressed gene modules with a joint clustering and Metropolis sampling approach. We demonstrated the efficiency of our method on a ChlP-chip data set with 38 breast cancer related TFs. Xi Chen 0056, Jianhua Xuan, Xu Shi 0003, Ayesha N. Shajahan, Leena Hilakivi-Clarke, Robert Clarke |
BIBM | 2 |
| 2013 | The CAM software for nonnegative blind source separation in R-Java
Niya Wang, Li Chen 0018, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Yue Joseph Wang |
J. Mach. Learn. Res. | 7 |
| 2013 | Reconstruction of Transcriptional Regulatory Networks by Stability-Based Network Component AnalysisabstractReliable inference of transcription regulatory networks is a challenging task in computational biology. Network component analysis (NCA) has become a powerful scheme to uncover regulatory networks behind complex biological processes. However, the performance of NCA is impaired by the high rate of false connections in binding information. In this paper, we integrate stability analysis with NCA to form a novel scheme, namely stability-based NCA (sNCA), for regulatory network identification. The method mainly addresses the inconsistency between gene expression data and binding motif information. Small perturbations are introduced to prior regulatory network, and the distance among multiple estimated transcript factor (TF) activities is computed to reflect the stability for each TF's binding network. For target gene identification, multivariate regression and t-statistic are used to calculate the significance for each TF-gene connection. Simulation studies are conducted and the experimental results show that sNCA can achieve an improved and robust performance in TF identification as compared to NCA. The approach for target gene identification is also demonstrated to be suitable for identifying true connections between TFs and their target genes. Furthermore, we have successfully applied sNCA to breast cancer data to uncover the role of TFs in regulating endocrine resistance in breast cancer. Xi Chen 0056, Jianhua Xuan, Chen Wang 0001, Ayesha N. Shajahan, Rebecca B. Riggins, Robert Clarke |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Detecting aberrant signal transduction pathways from high-throughput data using GIST algorithmabstractIt is biologically important to integrate high-throughput data to identify aberrant signal transduction pathways in cancer research. The high-throughput data acquired from The Cancer Genome Atlas (TCGA) Project offer a comprehensive picture of the genomic and transcriptional changes across hundreds of tumor samples. In this paper we propose a novel method, namely Gibbs sampler to Infer Signal Transduction pathways (GIST), to detect aberrant pathways that are highly associated with biological phenotypes or clinical information. GIST endeavors to estimate the edge probability by using a Markov Chain Monte Carlo (MCMC) method (i.e., a Gibbs sampling strategy). Through the sampling process, GIST is able to infer the correct signal transduction direction because the sampled edge probabilities are jointly determined by gene expression data and network topology. We first tested the efficacy of the GIST algorithm on yeast data and successfully uncovered several biologically meaningful signaling pathways. A case study on TCGA ovarian cancer data was further designed, aiming to unravel diverse signaling pathways associated with the development of ovarian cancer. The experimental results demonstrated the feasibility of applying GIST to identify and prioritize important signaling pathways in ovarian cancer for further biological validation. Jinghua Gu, Jianhua Xuan, Chen Wang 0001, Li Chen 0018, Tian-Li Wang, Ie-Ming Shih |
CIBCB | 2 |
| 2012 | Sampling-Based Subnetwork Identification from Microarray Data and Protein-Protein Interaction NetworkabstractIdentification of condition-specific protein interaction subnetworks has emerged as an attractive research field to reveal molecular mechanisms of diseases and provide reliable network biomarkers for disease diagnosis. Several methods have been proposed, which integrate gene expression and protein-protein interaction (PPI) data to identify subnetworks. However, existing methods treat differential expression of genes and network topology independently, which is an oversimplified assumption to model real biological systems. In this paper, we propose a sampling-based subnetwork identification approach to take into account the dependency between gene expression and network topology. Specifically, we apply Markov random field (MRF) theory to model the dependency of genes in PPI network using a Bayesian framework, followed by a Markov Chain Monte Carlo (MCMC) approach to identify significant subnetworks. The MCMC approach estimates the posterior distribution of genes' significant scores and network structure iteratively. Experimental results on both synthetic data and real breast cancer data demonstrated the effectiveness of the proposed method in identifying subnetworks, especially several functionally important, aberrant subnetworks associated with pathways involved in the development and recurrence of breast cancer. Xiao Wang 0031, Jinghua Gu, Jianhua Xuan, Ayesha N. Shajahan, Robert Clarke, Li Chen 0018 |
ICMLA (2) | 3 |
| 2012 | Reconstruction of Transcription Regulatory Networks by Stability-Based Network Component Analysis
Xi Chen 0056, Chen Wang 0001, Ayesha N. Shajahan, Rebecca B. Riggins, Robert Clarke, Jianhua Xuan |
ISBRA | 6 |
| 2012 | Robust identification of transcriptional regulatory networks using a Gibbs sampler on outlier sum statisticabstractMOTIVATION: Identification of transcriptional regulatory networks (TRNs) is of significant importance in computational biology for cancer research, providing a critical building block to unravel disease pathways. However, existing methods for TRN identification suffer from the inclusion of excessive 'noise' in microarray data and false-positives in binding data, especially when applied to human tumor-derived cell line studies. More robust methods that can counteract the imperfection of data sources are therefore needed for reliable identification of TRNs in this context. RESULTS: In this article, we propose to establish a link between the quality of one target gene to represent its regulator and the uncertainty of its expression to represent other target genes. Specifically, an outlier sum statistic was used to measure the aggregated evidence for regulation events between target genes and their corresponding transcription factors. A Gibbs sampling method was then developed to estimate the marginal distribution of the outlier sum statistic, hence, to uncover underlying regulatory relationships. To evaluate the effectiveness of our proposed method, we compared its performance with that of an existing sampling-based method using both simulation data and yeast cell cycle data. The experimental results show that our method consistently outperforms the competing method in different settings of signal-to-noise ratio and network topology, indicating its robustness for biological applications. Finally, we applied our method to breast cancer cell line data and demonstrated its ability to extract biologically meaningful regulatory modules related to estrogen signaling and action in breast cancer. AVAILABILITY AND IMPLEMENTATION: The Gibbs sampler MATLAB package is freely available at http://www.cbil.ece.vt.edu/software.htm. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jinghua Gu, Jianhua Xuan, Rebecca B. Riggins, Li Chen 0018, Yue Joseph Wang, Robert Clarke |
Bioinform. | 2 |
| 2012 | Computational analysis of muscular dystrophy sub-types using a novel integrative scheme
Chen Wang 0001, Sook Shin Ha, Jianhua Xuan, Yue Joseph Wang, Eric P. Hoffman |
Neurocomputing | 3 |
| 2012 | Regulatory component analysis: A semi-blind extraction approach to infer gene regulatory networks with imperfect biological knowledge
Chen Wang 0001, Jianhua Xuan, Ie-Ming Shih, Robert Clarke, Yue Joseph Wang |
Signal Process. | 2 |
| 2011 | PUGSVM: a caBIGTM analytical tool for multiclass gene selection and predictive classificationabstractUNLABELLED: Phenotypic Up-regulated Gene Support Vector Machine (PUGSVM) is a cancer Biomedical Informatics Grid (caBIG™) analytical tool for multiclass gene selection and classification. PUGSVM addresses the problem of imbalanced class separability, small sample size and high gene space dimensionality, where multiclass gene markers are defined by the union of one-versus-everyone phenotypic upregulated genes, and used by a well-matched one-versus-rest support vector machine. PUGSVM provides a simple yet more accurate strategy to identify statistically reproducible mechanistic marker genes for characterization of heterogeneous diseases. AVAILABILITY: http://www.cbil.ece.vt.edu/caBIG-PUGSVM.htm. Guoqiang Yu, Huai Li, Sook Shin Ha, Ie-Ming Shih, Robert Clarke, Eric P. Hoffman, Subha Madhavan, Jianhua Xuan, Yue Joseph Wang |
Bioinform. | 8 |
| 2011 | DDN: a caBIG® analytical tool for differential network analysisabstractUNLABELLED: Differential dependency network (DDN) is a caBIG® (cancer Biomedical Informatics Grid) analytical tool for detecting and visualizing statistically significant topological changes in transcriptional networks representing two biological conditions. Developed under caBIG®'s In Silico Research Centers of Excellence (ISRCE) Program, DDN enables differential network analysis and provides an alternative way for defining network biomarkers predictive of phenotypes. DDN also serves as a useful systems biology tool for users across biomedical research communities to infer how genetic, epigenetic or environment variables may affect biological networks and clinical phenotypes. Besides the standalone Java application, we have also developed a Cytoscape plug-in, CytoDDN, to integrate network analysis and visualization seamlessly. AVAILABILITY: The Java and MATLAB source code can be downloaded at the authors' web site http://www.cbil.ece.vt.edu/software.htm. Bai Zhang, Huai Li, Ie-Ming Shih, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Leena Hilakivi-Clarke, Yue Joseph Wang |
Bioinform. | 9 |
| 2011 | Motif-guided sparse decomposition of gene expression data for regulatory module identificationabstractBACKGROUND: Genes work coordinately as gene modules or gene networks. Various computational approaches have been proposed to find gene modules based on gene expression data; for example, gene clustering is a popular method for grouping genes with similar gene expression patterns. However, traditional gene clustering often yields unsatisfactory results for regulatory module identification because the resulting gene clusters are co-expressed but not necessarily co-regulated. RESULTS: We propose a novel approach, motif-guided sparse decomposition (mSD), to identify gene regulatory modules by integrating gene expression data and DNA sequence motif information. The mSD approach is implemented as a two-step algorithm comprising estimates of (1) transcription factor activity and (2) the strength of the predicted gene regulation event(s). Specifically, a motif-guided clustering method is first developed to estimate the transcription factor activity of a gene module; sparse component analysis is then applied to estimate the regulation strength, and so predict the target genes of the transcription factors. The mSD approach was first tested for its improved performance in finding regulatory modules using simulated and real yeast data, revealing functionally distinct gene modules enriched with biologically validated transcription factors. We then demonstrated the efficacy of the mSD approach on breast cancer cell line data and uncovered several important gene regulatory modules related to endocrine therapy of breast cancer. CONCLUSION: We have developed a new integrated strategy, namely motif-guided sparse decomposition (mSD) of gene expression data, for regulatory module identification. The mSD method features a novel motif-guided clustering method for transcription factor activity estimation by finding a balance between co-regulation and co-expression. The mSD method further utilizes a sparse decomposition method for regulation strength estimation. The experimental results show that such a motif-guided strategy can provide context-specific regulatory modules in both yeast and breast cancer studies. Jianhua Xuan, Li Chen 0018, Rebecca B. Riggins, Huai Li, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
BMC Bioinform. | 2 |
| 2010 | Module-based biomarker discovery in breast cancerabstractThe availability of genome-wide biological network data opens up new possibilities to discover novel biomarkers and elucidate cancer-related complex mechanisms at network level. In this paper, we propose a novel module-based feature selection framework, which integrates biological network information and gene expression data to identify biomarkers, not as individual genes but as functional modules. Also, a large-scale analysis of ensemble feature selection concept is presented. The method allows combining features selected from multiple runs with various data subsampling to increase the reliability and classification accuracy of the final set of selected features. The results from four breast cancer studies demonstrate that the identified module biomarkers achieve: i) higher classification accuracy in independent validation datasets; ii) better reproducibility than individual gene biomarkers; iii) improved biological interpretability; and iv) enhanced enrichment in cancer-related “disease drivers”. Yuji Zhang 0001, Jianhua Xuan, Robert Clarke, Habtom W. Ressom |
BIBM | 2 |
| 2010 | Identification of Transcriptional Regulatory Networks by Learning the Marginal Function of Outlier Sum StatisticabstractNetwork component analysis (NCA) and other methods based on the NCA model have become powerful bioinformatics tools to reconstruct underlying regulatory networks and recover hidden biological processes. However, due to the existence of experimental noises in micro array data and false information in network connectivity data (e.g., ChIP-on-chip binding data, motif information, etc.), it still remains challenging to reconstruct gene regulatory networks for real biomedical applications such as human cancer studies. In this paper, we model the relationship between the genes that share the same transcription factors (TF) from the angle of regression. We propose a statistic called outlier sum testing the conditional significance of the target genes. A Gibbs strategy is utilized in order to estimate the marginal value of outlier sum from its conditional function. Based on the outlier sum statistic we are able to extract the true target genes that carry information about transcription factor activities (TFAs) from the whole population. As a proof-of-concept, we demonstrated the efficiency and robustness of the proposed method on both simulation data and yeast cell cycle data. Jinghua Gu, Jianhua Xuan, Yue Joseph Wang, Rebecca B. Riggins, Robert Clarke |
ICMLA | 2 |
| 2010 | Computational Analysis of Muscular Dystrophy Sub-types Using a Novel Integrative SchemeabstractTo construct biologically interpretable features and facilitate Muscular Dystrophy (MD) sub-types classification, we propose a novel integrative scheme utilizing PPI network, functional gene sets information, and mRNA profiling. The workflow of the proposed scheme includes three major steps: First, by combining protein-protein interaction network structure and gene co-expression relationship into new distance metric, we apply affinity propagation clustering to build gene sub-networks. Secondly, we further incorporate functional gene sets knowledge to complement the physical interaction information. Finally, based on constructed sub-network and gene set features, we apply multi-class support vector machine (MSVM) for MD sub-type classification, and highlight the biomarkers contributing to the sub-type prediction. The experimental results show that our scheme could construct sub-networks that are more relevant to MD than those constructed by conventional approach. Furthermore, our integrative strategy substantially improved the prediction accuracy, especially for those hard-to-classify sub-types. Chen Wang 0001, Sook Shin Ha, Yue Joseph Wang, Jianhua Xuan, Eric P. Hoffman |
ICMLA | 4 |
| 2010 | Multilevel support vector regression analysis to identify condition-specific regulatory networksabstractMOTIVATION: The identification of gene regulatory modules is an important yet challenging problem in computational biology. While many computational methods have been proposed to identify regulatory modules, their initial success is largely compromised by a high rate of false positives, especially when applied to human cancer studies. New strategies are needed for reliable regulatory module identification. RESULTS: We present a new approach, namely multilevel support vector regression (ml-SVR), to systematically identify condition-specific regulatory modules. The approach is built upon a multilevel analysis strategy designed for suppressing false positive predictions. With this strategy, a regulatory module becomes ever more significant as more relevant gene sets are formed at finer levels. At each level, a two-stage support vector regression (SVR) method is utilized to help reduce false positive predictions by integrating binding motif information and gene expression data; a significant analysis procedure is followed to assess the significance of each regulatory module. To evaluate the effectiveness of the proposed strategy, we first compared the ml-SVR approach with other existing methods on simulation data and yeast cell cycle data. The resulting performance shows that the ml-SVR approach outperforms other methods in the identification of both regulators and their target genes. We then applied our method to breast cancer cell line data to identify condition-specific regulatory modules associated with estrogen treatment. Experimental results show that our method can identify biologically meaningful regulatory modules related to estrogen signaling and action in breast cancer. AVAILABILITY AND IMPLEMENTATION: The ml-SVR MATLAB package can be downloaded at http://www.cbil.ece.vt.edu/software.htm. Li Chen 0018, Jianhua Xuan, Rebecca B. Riggins, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
Bioinform. | 2 |
| 2010 | Knowledge-guided gene ranking by coordinative component analysisabstractBACKGROUND: In cancer, gene networks and pathways often exhibit dynamic behavior, particularly during the process of carcinogenesis. Thus, it is important to prioritize those genes that are strongly associated with the functionality of a network. Traditional statistical methods are often inept to identify biologically relevant member genes, motivating researchers to incorporate biological knowledge into gene ranking methods. However, current integration strategies are often heuristic and fail to incorporate fully the true interplay between biological knowledge and gene expression data. RESULTS: To improve knowledge-guided gene ranking, we propose a novel method called coordinative component analysis (COCA) in this paper. COCA explicitly captures those genes within a specific biological context that are likely to be expressed in a coordinative manner. Formulated as an optimization problem to maximize the coordinative effort, COCA is designed to first extract the coordinative components based on a partial guidance from knowledge genes and then rank the genes according to their participation strengths. An embedded bootstrapping procedure is implemented to improve statistical robustness of the solutions. COCA was initially tested on simulation data and then on published gene expression microarray data to demonstrate its improved performance as compared to traditional statistical methods. Finally, the COCA approach has been applied to stem cell data to identify biologically relevant genes in signaling pathways. As a result, the COCA approach uncovers novel pathway members that may shed light into the pathway deregulation in cancers. CONCLUSION: We have developed a new integrative strategy to combine biological knowledge and microarray data for gene ranking. The method utilizes knowledge genes for a guidance to first extract coordinative components, and then rank the genes according to their contribution related to a network or pathway. The experimental results show that such a knowledge-guided strategy can provide context-specific gene ranking with an improved performance in pathway member identification. Chen Wang 0001, Jianhua Xuan, Huai Li, Yue Joseph Wang, Ming Zhan, Eric P. Hoffman, Robert Clarke |
BMC Bioinform. | 2 |
| 2010 | Matched Gene Selection and Committee Classifier for Molecular Classification of Heterogeneous Diseases
Guoqiang Yu, Yuanjian Feng, David J. Miller 0001, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Ben Davidson, Ie-Ming Shih, Yue Joseph Wang |
J. Mach. Learn. Res. | 4 |
| 2009 | Enhancing Pathway Based Analysis Using Different Weighting SchemesabstractIn this paper, we propose to apply non-uniform weighs to the genes in the pathway based analysis and present two weighting schemes for the genes. Specifically, we incorporate our weighting schemes into the global test pathway based analysis approach and investigate the effects of our weighting schemes. We observe that when non-uniform weights are applied, some originally lower ranked pathways are elevated to the high ranks, prediction performances of the selected genes are improved, and some genes associated with the related phenotype are identified, which are missed by the uniform weighting approach. Sook Shin Ha, Inyoung Kim, Jianhua Xuan |
BIBM | 3 |
| 2009 | Multi-level Ground Glass Nodule Detection and Segmentation in CT Lung Images
Yimo Tao, Le Lu 0001, Maneesh Dewan, Albert Y. Chen, Jason J. Corso, Jianhua Xuan, Marcos Salganicoff, Arun Krishnan |
MICCAI (1) | 6 |
| 2009 | Differential dependency network analysis to identify condition-specific topological changes in biological networksabstractMOTIVATION: Significant efforts have been made to acquire data under different conditions and to construct static networks that can explain various gene regulation mechanisms. However, gene regulatory networks are dynamic and condition-specific; under different conditions, networks exhibit different regulation patterns accompanied by different transcriptional network topologies. Thus, an investigation on the topological changes in transcriptional networks can facilitate the understanding of cell development or provide novel insights into the pathophysiology of certain diseases, and help identify the key genetic players that could serve as biomarkers or drug targets. RESULTS: Here, we report a differential dependency network (DDN) analysis to detect statistically significant topological changes in the transcriptional networks between two biological conditions. We propose a local dependency model to represent the local structures of a network by a set of conditional probabilities. We develop an efficient learning algorithm to learn the local dependency model using the Lasso technique. A permutation test is subsequently performed to estimate the statistical significance of each learned local structure. In testing on a simulation dataset, the proposed algorithm accurately detected all the genes with network topological changes. The method was then applied to the estrogen-dependent T-47D estrogen receptor-positive (ER+) breast cancer cell line datasets and human and mouse embryonic stem cell datasets. In both experiments using real microarray datasets, the proposed method produced biologically meaningful results. We expect DDN to emerge as an important bioinformatics tool in transcriptional network analyses. While we focus specifically on transcriptional networks, the DDN method we introduce here is generally applicable to other biological networks with similar characteristics. AVAILABILITY: The DDN MATLAB toolbox and experiment data are available at http://www.cbil.ece.vt.edu/software.htm. Bai Zhang, Huai Li, Rebecca B. Riggins, Ming Zhan, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
Bioinform. | 5 |
| 2008 | Network-Constrained Support Vector Machine for ClassificationabstractOne of the major goals in microarray data analysis is to identify biomarkers and build a classification model for future prediction. Many traditional statistical models, based on microarray data alone, often fail in identifying biologically meaningful genes, which should have synergistic effect on determine the clinical outcomes through some interactions rather than work individually. In this paper, we proposed a network-constrained support vector machine (nSVM) for classification by incorporating prior knowledge, which could be protein-protein interactions, protein-gene regulation relationships or pathways information. Specifically, we use Laplacian matrix to represent gene-gene interaction network to regularize the objective function of SVM, which imposes the smoothness of coefficients over the network. The experimental results on simulation and real microarray datasets demonstrate that our method could not only improve classification performance compared to conventional SVM, but more importantly, it could identify significant sub-networks belonging to several pathways which might be related to underlying mechanism associated with clinical outcomes. Li Chen 0018, Jianhua Xuan, Yue Joseph Wang, Rebecca B. Riggins, Robert Clarke |
ICMLA | 2 |
| 2008 | Imaging biomarker analysis of rat mammary fat pads and glandular tissues in MRI imagesabstractIn studying the relationship between risk factors and breast cancer, the growth patterns of fat pads and glandular tissues are considered as important biomarkers. The aim of this study is to measure the growth pattern statistics of rat mammary pads and glandular tissues with magnetic resonance (MR) time sequence images. In this paper, we proposed methods containing sequential steps to extract and analyze imaging biomarkers of rat mammary pad and glandular tissues. Firstly, to accurately segment out pads in MR images with noisy bias filed, we proposed a level set method combining local binary fitting (LBF) and geodesic active contour (GAC). The salient glandular tissue regions within the fat pads are further extracted by a scale-space analysis procedure. Then, the volume data of a single rat at different time points are aligned through profile correlation analysis. Finally, the growth rates are calculated and compared to show the changing patterns of fat pads and glandular tissues within separate groups. The experimental results showed the great utility of this approach in providing accurate measurements for novel risk factors of breast cancer. Yimo Tao, Jianhua Xuan, Matthew T. Freedman, Gloria Chepko, Peter G. Shields, Yue Joseph Wang |
ICPR | 2 |
| 2008 | Sparse Decomposition of Gene Expression Data to Infer Transcriptional Modules Guided by Motif Information
Jianhua Xuan, Li Chen 0018, Rebecca B. Riggins, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
ISBRA | 2 |
| 2008 | Integrative Network Component Analysis for Regulatory Network Reconstruction
Chen Wang 0001, Jianhua Xuan, Li Chen 0018, Po Zhao, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman |
ISBRA | 2 |
| 2008 | Knowledge-guided multi-scale independent component analysis for biomarker identificationabstractBACKGROUND: Many statistical methods have been proposed to identify disease biomarkers from gene expression profiles. However, from gene expression profile data alone, statistical methods often fail to identify biologically meaningful biomarkers related to a specific disease under study. In this paper, we develop a novel strategy, namely knowledge-guided multi-scale independent component analysis (ICA), to first infer regulatory signals and then identify biologically relevant biomarkers from microarray data. RESULTS: Since gene expression levels reflect the joint effect of several underlying biological functions, disease-specific biomarkers may be involved in several distinct biological functions. To identify disease-specific biomarkers that provide unique mechanistic insights, a meta-data "knowledge gene pool" (KGP) is first constructed from multiple data sources to provide important information on the likely functions (such as gene ontology information) and regulatory events (such as promoter responsive elements) associated with potential genes of interest. The gene expression and biological meta data associated with the members of the KGP can then be used to guide subsequent analysis. ICA is then applied to multi-scale gene clusters to reveal regulatory modes reflecting the underlying biological mechanisms. Finally disease-specific biomarkers are extracted by their weighted connectivity scores associated with the extracted regulatory modes. A statistical significance test is used to evaluate the significance of transcription factor enrichment for the extracted gene set based on motif information. We applied the proposed method to yeast cell cycle microarray data and Rsf-1-induced ovarian cancer microarray data. The results show that our knowledge-guided ICA approach can extract biologically meaningful regulatory modes and outperform several baseline methods for biomarker identification. CONCLUSION: We have proposed a novel method, namely knowledge-guided multi-scale ICA, to identify disease-specific biomarkers. The goal is to infer knowledge-relevant regulatory signals and then identify corresponding biomarkers through a multi-scale strategy. The approach has been successfully applied to two expression profiling experiments to demonstrate its improved performance in extracting biologically meaningful and disease-related biomarkers. More importantly, the proposed approach shows promising results to infer novel biomarkers for ovarian cancer and extend current knowledge. Li Chen 0018, Jianhua Xuan, Chen Wang 0001, Ie-Ming Shih, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
BMC Bioinform. | 2 |
| 2008 | Motif-directed network component analysis for regulatory network inferenceabstractBACKGROUND: Network Component Analysis (NCA) has shown its effectiveness in discovering regulators and inferring transcription factor activities (TFAs) when both microarray data and ChIP-on-chip data are available. However, a NCA scheme is not applicable to many biological studies due to limited topology information available, such as lack of ChIP-on-chip data. We propose a new approach, motif-directed NCA (mNCA), to integrate motif information and gene expression data to infer regulatory networks. RESULTS: We develop motif-directed NCA (mNCA) to incorporate motif information into NCA for regulatory network inference. While motif information is readily available from knowledge databases, it is a "noisy" source of network topology information consisting of many false positives. To overcome this problem, we develop a stability analysis procedure embedded in mNCA to resolve the inconsistency between motif information and gene expression data, and to enable the identification of stable TFAs. The mNCA approach has been applied to a time course microarray data set of muscle regeneration. The experimental results show that the inferred TFAs are not only numerically stable but also biologically relevant to muscle differentiation process. In particular, several inferred TFAs like those of MyoD, myogenin and YY1 are well supported by biological experiments. CONCLUSION: A novel computational approach, mNCA, has been developed to integrate motif information and gene expression data for regulatory network reconstruction. Specifically, motif analysis is used to obtain initial network topology, and stability analysis is developed and applied with mNCA to extract stable TFAs. Experimental results on muscle regeneration microarray data have demonstrated that mNCA is a practical and reliable computational method for regulatory network inference and pathway discovery. Chen Wang 0001, Jianhua Xuan, Li Chen 0018, Po Zhao, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman |
BMC Bioinform. | 2 |
| 2008 | Network motif-based identification of transcription factor-target gene relationships by integrating multi-source biological dataabstractBACKGROUND: Integrating data from multiple global assays and curated databases is essential to understand the spatio-temporal interactions within cells. Different experiments measure cellular processes at various widths and depths, while databases contain biological information based on established facts or published data. Integrating these complementary datasets helps infer a mutually consistent transcriptional regulatory network (TRN) with strong similarity to the structure of the underlying genetic regulatory modules. Decomposing the TRN into a small set of recurring regulatory patterns, called network motifs (NM), facilitates the inference. Identifying NMs defined by specific transcription factors (TF) establishes the framework structure of a TRN and allows the inference of TF-target gene relationship. This paper introduces a computational framework for utilizing data from multiple sources to infer TF-target gene relationships on the basis of NMs. The data include time course gene expression profiles, genome-wide location analysis data, binding sequence data, and gene ontology (GO) information. RESULTS: The proposed computational framework was tested using gene expression data associated with cell cycle progression in yeast. Among 800 cell cycle related genes, 85 were identified as candidate TFs and classified into four previously defined NMs. The NMs for a subset of TFs are obtained from literature. Support vector machine (SVM) classifiers were used to estimate NMs for the remaining TFs. The potential downstream target genes for the TFs were clustered into 34 biologically significant groups. The relationships between TFs and potential target gene clusters were examined by training recurrent neural networks whose topologies mimic the NMs to which the TFs are classified. The identified relationships between TFs and gene clusters were evaluated using the following biological validation and statistical analyses: (1) Gene set enrichment analysis (GSEA) to evaluate the clustering results; (2) Leave-one-out cross-validation (LOOCV) to ensure that the SVM classifiers assign TFs to NM categories with high confidence; (3) Binding site enrichment analysis (BSEA) to determine enrichment of the gene clusters for the cognate binding sites of their predicted TFs; (4) Comparison with previously reported results in the literatures to confirm the inferred regulations. CONCLUSION: The major contribution of this study is the development of a computational framework to assist the inference of TRN by integrating heterogeneous data from multiple sources and by decomposing a TRN into NM-based modules. The inference capability of the proposed framework is verified statistically (e.g., LOOCV) and biologically (e.g., GSEA, BSEA, and literature validation). The proposed framework is useful for inferring small NM-based modules of TF-target gene relationships that can serve as a basis for generating new testable hypotheses. Yuji Zhang 0001, Jianhua Xuan, Benildo de los Reyes, Robert Clarke, Habtom W. Ressom |
BMC Bioinform. | 2 |
| 2008 | caBIGTM VISDA: Modeling, visualization, and discovery for cluster analysis of genomic dataabstractBACKGROUND: The main limitations of most existing clustering methods used in genomic data analysis include heuristic or random algorithm initialization, the potential of finding poor local optima, the lack of cluster number detection, an inability to incorporate prior/expert knowledge, black-box and non-adaptive designs, in addition to the curse of dimensionality and the discernment of uninformative, uninteresting cluster structure associated with confounding variables. RESULTS: In an effort to partially address these limitations, we develop the VIsual Statistical Data Analyzer (VISDA) for cluster modeling, visualization, and discovery in genomic data. VISDA performs progressive, coarse-to-fine (divisive) hierarchical clustering and visualization, supported by hierarchical mixture modeling, supervised/unsupervised informative gene selection, supervised/unsupervised data visualization, and user/prior knowledge guidance, to discover hidden clusters within complex, high-dimensional genomic data. The hierarchical visualization and clustering scheme of VISDA uses multiple local visualization subspaces (one at each node of the hierarchy) and consequent subspace data modeling to reveal both global and local cluster structures in a "divide and conquer" scenario. Multiple projection methods, each sensitive to a distinct type of clustering tendency, are used for data visualization, which increases the likelihood that cluster structures of interest are revealed. Initialization of the full dimensional model is based on first learning models with user/prior knowledge guidance on data projected into the low-dimensional visualization spaces. Model order selection for the high dimensional data is accomplished by Bayesian theoretic criteria and user justification applied via the hierarchy of low-dimensional visualization subspaces. Based on its complementary building blocks and flexible functionality, VISDA is generally applicable for gene clustering, sample clustering, and phenotype clustering (wherein phenotype labels for samples are known), albeit with minor algorithm modifications customized to each of these tasks. CONCLUSION: VISDA achieved robust and superior clustering accuracy, compared with several benchmark clustering schemes. The model order selection scheme in VISDA was shown to be effective for high dimensional genomic data clustering. On muscular dystrophy data and muscle regeneration data, VISDA identified biologically relevant co-expressed gene clusters. VISDA also captured the pathological relationships among different phenotypes revealed at the molecular level, through phenotype clustering on muscular dystrophy data and multi-category cancer data. Yitan Zhu, Huai Li, David J. Miller 0001, Zuyi Wang, Jianhua Xuan, Robert Clarke, Eric P. Hoffman, Yue Joseph Wang |
BMC Bioinform. | 5 |
| 2007 | Rat Mammary Fat Pad Segmentation and Growth Rate Evaluation in T1 Weighted MR ImagesabstractIn studying the relationship between risk factors and breast cancer, growth patterns of the fat pads and glandular tissues are important features. The goal of this small animal study is to measure the size of mammary pads over the time. To achieve this goal, we propose a hierarchical approach to segmenting out rat body, mammary fat pads and evaluating their development in Tl weighted magnetic resonance (TlW-MR) images. Particularly, we have developed a new approach combining watershed transform and region competition for improved fat pad segmentation. An efficient strategy, termed as competition propagation, is developed to propagate the region competition result from one slice to next slice, resulting in a fast convergence in region competition algorithm otherwise computationally costly. To evaluate the development of the fat pads, the volume data of the scans for a single rat to compare are aligned and the common valid range is acquired through correlation analysis. The method has been applied to 18 volumetric sets of Tl W-MR images acquired from this study. The experimental results showed the great utility of this approach as it can provide accurate measurements to assess novel risk factors for breast cancer. Bin Wang 0064, Jianhua Xuan, Matthew T. Freedman, Peter G. Shields, Yue Joseph Wang |
BIBE | 2 |
| 2007 | Biomarker Identification by Knowledge-Driven Multi-Level ICA and Motif AnalysisabstractMany statistical methods often fail to identify biologically meaningful biomarkers related to a specific disease under study from expression data alone. In this paper, we develop a novel strategy, namely knowledge-driven multi-level independent component analysis (ICA), to infer regulatory signals and identify biologically relevant biomarkers from microarray data. Specifically, based on multi-level clustering results and partial prior knowledge, we apply ICA to find stable disease specific linear regulatory modes and then extract associated biomarker genes. A statistical test is designed to evaluate the significance of transcription factor enrichment for extracted gene set based on motif information. The experimental results on an Rsf-1 induced microarray data set show that our knowledge-driven method can extract more biologically meaningful biomarkers with significant enrichment of transcription factors related to ovarian cancer compared to other gene selection methods with/without prior knowledge. Li Chen 0018, Chen Wang 0001, Ie-Ming Shih, Tian-Li Wang, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman, Jianhua Xuan |
ICMLA | 9 |
| 2007 | VISDA: an open-source caBIGTM analytical tool for data clustering and beyondabstractSUMMARY: VISDA (Visual Statistical Data Analyzer) is a caBIG analytical tool for cluster modeling, visualization and discovery that has met silver-level compatibility under the caBIG initiative. Being statistically principled and visually interfaced, VISDA exploits both hierarchical statistics modeling and human gift for pattern recognition to allow a progressive yet interactive discovery of hidden clusters within high dimensional and complex biomedical datasets. The distinctive features of VISDA are particularly useful for users across the cancer research and broader research communities to analyze complex biological data. AVAILABILITY: http://gforge.nci.nih.gov/projects/visda/ Jiajing Wang, Huai Li, Yitan Zhu, Malik Yousef, Michael Nebozhyn, Michael M. Showe, Louise C. Showe, Jianhua Xuan, Robert Clarke, Yue Joseph Wang |
Bioinform. | 8 |
| 2006 | Learning the Tree of Phenotypes Using Genomic Data and VISDAabstractThough supervised and unsupervised analyses of genomic data have been intensively studied in recent years, little effort has been made to discover the structural information contained in the data. In this work, we propose a stability analysis guided supervised clustering and visualization method aiming to discover the hierarchical structure in gene expression data, which we call the "tree of phenotypes". We applied the method on two multiclass gene expression microarray data sets and presented the biological plausibility of the learned trees. We also tested the multiclass classifiers built on the learned trees and demonstrated their good classification performance Yuanjian Feng, Zuyi Wang, Yitan Zhu, Jianhua Xuan, David J. Miller 0001 |
BIBE | 4 |
| 2006 | ModVis: An information visualization tool for gene module discovery
Justin Molineaux, Jianhua Xuan, Yitan Zhu, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
CAINE | 2 |
| 2006 | Inference of Gene Regulatory Networks from Time Course Gene Expression Data Using Neural Networks and Swarm IntelligenceabstractWe present a novel algorithm that combines a recurrent neural network (RNN) and two swarm intelligence (SI) methods to infer a gene regulatory network (GRN) from time course gene expression data. The algorithm uses ant colony optimization (ACO) to identify the optimal architecture of an RNN, while the weights of the RNN are optimized using particle swarm optimization (PSO). Our goal is to construct an RNN whose response mimics gene expression data generated by time course DNA microarray experiments. We observed promising results in applying the proposed hybrid SI-RNN algorithm to infer networks of interaction from simulated and real-world gene expression data Habtom W. Ressom, Yuji Zhang 0001, Jianhua Xuan, Yue Joseph Wang, Robert Clarke |
CIBCB | 3 |
| 2006 | Optimized multilayer perceptrons for molecular classification and diagnosis using genomic dataabstractMOTIVATION: Multilayer perceptrons (MLP) represent one of the widely used and effective machine learning methods currently applied to diagnostic classification based on high-dimensional genomic data. Since the dimensionalities of the existing genomic data often exceed the available sample sizes by orders of magnitude, the MLP performance may degrade owing to the curse of dimensionality and over-fitting, and may not provide acceptable prediction accuracy. RESULTS: Based on Fisher linear discriminant analysis, we designed and implemented an MLP optimization scheme for a two-layer MLP that effectively optimizes the initialization of MLP parameters and MLP architecture. The optimized MLP consistently demonstrated its ability in easing the curse of dimensionality in large microarray datasets. In comparison with a conventional MLP using random initialization, we obtained significant improvements in major performance measures including Bayes classification accuracy, convergence properties and area under the receiver operating characteristic curve (A(z)). SUPPLEMENTARY INFORMATION: The Supplementary information is available on http://www.cbil.ece.vt.edu/publications.htm Zuyi Wang, Yue Joseph Wang, Jianhua Xuan, Yibin Dong, Marina Bakay, Yuanjian Feng, Robert Clarke, Eric P. Hoffman |
Bioinform. | 3 |
| 2005 | Normalization of Microarray Data by Iterative Nonlinear RegressionabstractNormalization is an important prerequisite for almost all follow-up microarray data analysis steps. Accurate normalization assures a common base for comparative biomedical studies using gene expression profiles across different experiments and phenotypes. In this paper, we present a novel normalization approach - iterative nonlinear regression (INR) method - that exploits concurrent identification of invariantly expressed genes (IEGs) and implementation of nonlinear regression normalization. We demonstrate the principle and performance of the INR approach on two real microarray data sets. As compared to major peer methods (e.g., linear regression method, Loess method and iterative ranking method), INR method shows a superior performance in achieving low expression variance across replicates and excellent fold change preservation. Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
BIBE | 1 |
| 2005 | Discontinuity-embedded deformable models for surface reconstruction from range imagesabstractSurface reconstruction is a critical step in three-dimensional image processing and understanding. In this letter, a discontinuity-embedded deformable model has been developed to model surfaces with discontinuities. Governed by the Lagrange motion equation, a finite-element representation of the model can dynamically fit the data in both continuous and discontinuous components, reaching its equilibrium in response to induced forces. Experimental results on synthetic and range images demonstrate a significant improvement in preserving depth discontinuities over conventional approaches. Jianhua Xuan, Yue Joseph Wang, Qinfen Zheng, Tülay Adali |
IEEE Signal Process. Lett. | 1 |
| 2003 | Computational intelligence approach for gene expression data mining and classificationabstractThe exploration of high dimensional gene expression microarray data demands powerful analytical tools. Our data mining software, visual data analyzer (VISDA) for cluster discovery, reveals many distinguishing patterns among gene expression profiles. The model-supported hierarchical data exploration tool has two complementary schemes: discriminatory dimensionality reduction for structure-focused data visualization, and cluster decomposition by probabilistic clustering. Reducing dimensionality generates the visualization of the complete data set at the top level. This data set is then partitioned into subclusters that can consequently be visualized at lower levels and if necessary partitioned again. These approaches produce different visualizations that are compared against known phenotypes from the microarray experiments. For class prediction on cancers using miroarray data, multilayer perceptrons (MLPs) are trained and optimized, whose architecture and parameters are regularized and initialized by weighted Fisher criterion (wFC)-based discriminatory component analysis (DCA). The prediction performance is compared and evaluated via multifold cross-validation. Zuyi Wang, Sun-Yuan Kung, Javed I. Khan, Jianhua Xuan, Yue Joseph Wang |
ICME | 5 |
| 2001 | Magnetic resonance image analysis by information theoretic criteria and stochastic site modelsabstractQuantitative analysis of magnetic resonance (MR) images is a powerful tool for image-guided diagnosis, monitoring, and intervention. The major tasks involve tissue quantification and image segmentation where both the pixel and context images are considered. To extract clinically useful information from images that might be lacking in prior knowledge, we introduce an unsupervised tissue characterization algorithm that is both statistically principled and patient specific. The method uses adaptive standard finite normal mixture and inhomogeneous Markov random field models, whose parameters are estimated using expectation-maximization and relaxation labeling algorithms under information theoretic criteria. We demonstrate the successful applications of the approach with synthetic data sets and then with real MR brain images. Yue Joseph Wang, Tülay Adali, Jianhua Xuan, Zsolt Szabo |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 1998 | 3-D Model Supported Prostrate Biopsy Simulation and Evaluation
Jianhua Xuan, Yue Joseph Wang, Isabell A. Sesterhenn, Judd W. Moul, Seong Ki Mun |
MICCAI | 1 |
| 1997 | Modeling of Wavelet Coefficients in Medical Image CompressionabstractThe discrete wavelet transform provides a new framework of multiresolution space-frequency representation. Its preliminary applications in medical image compression are promising. An accurate modeling of the spatial and frequency characteristics of the wavelet coefficients is a key to designing efficient and accurate quantization for wavelet-based source coding. In this study, we investigate the modeling of a finite mixture distribution of the wavelet coefficients, within the context of information theory and statistical model identification. Using a finite generalized Gaussian mixture to model the overall distribution of the coefficients, an unsupervised learning procedure is developed to quantify the histogram through a tripled adaptive algorithm including detection of the number of local kernels, approximation of the shape of local kernels, and estimation of model parameter values. Our preliminary experimental results indicate that the unsupervised and adaptive histogram quantification can efficiently and accurately fit to the overall mixture distribution of the coefficients for any given frequency subband with unknown characteristics. Yue Joseph Wang, Huao Li, Jianhua Xuan, Shih-Chung Ben Lo, Seong Ki Mun |
ICIP (1) | 3 |
| 1997 | A Deformable Surface-Spine Model for 3-D Surface RegistrationabstractA finite-element deformable surface-spine model is developed in this paper to register two surfaces by recovering the nonlinear deformation with respect to each other. The deformable surface-spine model is a dynamic model governed by Lagrangian motion equations. A 9 degree-of-freedom (dof) finite-element surface element and a 4-dof spine element are developed to iteratively solve Lagrangian equations for computing the deformation between two surfaces. The method has been applied to registration of computerized surgical prostate models. Experimental results have demonstrated that the new registration method can successfully match complex-structured surfaces by recovering the nonlinear deformation. Jianhua Xuan, Yue Joseph Wang, Tülay Adali, Qinfen Zheng |
ICIP (3) | 1 |
| 1996 | Information geometry of maximum partial likelihood estimation for channel equalizationabstractInformation geometry of partial likelihood is constructed and is used to derive the em-algorithm for learning parameters of a conditional distribution model through information-theoretic projections. To construct the coordinates of the information geometry, an expectation maximization (EM) framework is described for the distribution learning problem using the Gaussian mixture probability model. It is shown that the information-geometric em-algorithm is equivalent to EM to establish its convergence. The algorithm is applied to channel equalization by distribution learning and its rapid convergence characteristics are demonstrated through simulation studies. Jianhua Xuan, Tülay Adali, Xiao Liu 0024 |
ICASSP | 1 |
| 1995 | Segmentation of magnetic resonance brain image: integrating region growing and edge detectionabstractThe authors present a method that combines region growing and edge detection for magnetic resonance (MR) brain image segmentation. Starting with a simple region growing algorithm which produces an over segmented image, the authors apply a sophisticated region merging method which is capable of handling complex image structures. Edge information is then integrated to verify and, where necessary, to correct region boundaries. The results show that this method is reliable and efficient for MR brain image segmentation. Jianhua Xuan, Tülay Adali, Yue Joseph Wang |
ICIP (3) | 1 |