VLDB 2026 Research / reviewers in the wild / expert
Yue Joseph Wang
dblp:89/2341 · also Yue Wang 0021
· DBLP profile ↗
96ranked-venue papers
9as first author
4since 2021 · last 2025
0000-0002-1788-1102ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 58 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 22 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-authorSystems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prediction of gene regulatory connections with joint single-cell foundation models and graph-based learningabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) data offers unprecedented opportunities to infer gene regulatory networks (GRNs) at a fine-grained resolution, shedding light on cellular phenotypes at the molecular level. However, the high sparsity, noise, and dropout events inherent in scRNA-seq data pose significant challenges for accurate and reliable GRN inference. The rapid growth in experimentally validated transcription factor-DNA binding data has enabled supervised machine learning methods, which rely on known regulatory interactions to learn patterns, and achieve high accuracy in GRN inference by framing it as a gene regulatory link prediction task. This study addresses the gene regulatory link prediction problem by learning vectorized representations at the gene level to predict missing regulatory interactions. However, a higher performance of supervised learning methods requires a large amount of known TF-DNA binding data, which is often experimentally expensive and therefore limited in amount. Advances in large-scale pre-training and transfer learning provide a transformative opportunity to address this challenge. In this study, we leverage large-scale pre-trained models, trained on extensive scRNA-seq datasets and known as single-cell foundation models (scFMs). These models are combined with joint graph-based learning to establish a robust foundation for gene regulatory link prediction. RESULTS: We propose scRegNet, a novel and effective framework that leverages scFMs with joint graph-based learning for gene regulatory link prediction. scRegNet achieves state-of-the-art results in comparison with nine baseline methods on seven scRNA-seq benchmark datasets. Additionally, scRegNet is more robust than the baseline methods on noisy training data. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/sindhura-cs/scRegNet. Sindhura Kommu, Yizhi Wang 0009, Yue Joseph Wang, Xuan Wang 0008 |
Bioinform. | 3 |
| 2024 | DDN3.0: determining significant rewiring of biological network structure with differential dependency networksabstractMOTIVATION: Complex diseases are often caused and characterized by misregulation of multiple biological pathways. Differential network analysis aims to detect significant rewiring of biological network structures under different conditions and has become an important tool for understanding the molecular etiology of disease progression and therapeutic response. With few exceptions, most existing differential network analysis tools perform differential tests on separately learned network structures that are computationally expensive and prone to collapse when grouped samples are limited or less consistent. RESULTS: We previously developed an accurate differential network analysis method-differential dependency networks (DDN), that enables joint learning of common and rewired network structures under different conditions. We now introduce the DDN3.0 tool that improves this framework with three new and highly efficient algorithms, namely, unbiased model estimation with a weighted error measure applicable to imbalance sample groups, multiple acceleration strategies to improve learning efficiency, and data-driven determination of proper hyperparameters. The comparative experimental results obtained from both realistic simulations and case studies show that DDN3.0 can help biologists more accurately identify, in a study-specific and often unknown conserved regulatory circuitry, a network of significantly rewired molecular players potentially responsible for phenotypic transitions. AVAILABILITY AND IMPLEMENTATION: The Python package of DDN3.0 is freely available at https://github.com/cbil-vt/DDN3. A user's guide and a vignette are provided at https://ddn-30.readthedocs.io/. Yingzhou Lu, Yizhi Wang 0009, Bai Zhang, Guoqiang Yu, Chunyu Liu 0001, Robert Clarke, David M. Herrington, Yue Joseph Wang |
Bioinform. | 10 |
| 2024 | CAM3.0: determining cell type composition and expression from bulk tissues with fully unsupervised deconvolutionabstractMOTIVATION: Complex tissues are dynamic ecosystems consisting of molecularly distinct yet interacting cell types. Computational deconvolution aims to dissect bulk tissue data into cell type compositions and cell-specific expressions. With few exceptions, most existing deconvolution tools exploit supervised approaches requiring various types of references that may be unreliable or even unavailable for specific tissue microenvironments. RESULTS: We previously developed a fully unsupervised deconvolution method-Convex Analysis of Mixtures (CAM), that enables estimation of cell type composition and expression from bulk tissues. We now introduce CAM3.0 tool that improves this framework with three new and highly efficient algorithms, namely, radius-fixed clustering to identify reliable markers, linear programming to detect an initial scatter simplex, and a smart floating search for the optimum latent variable model. The comparative experimental results obtained from both realistic simulations and case studies show that the CAM3.0 tool can help biologists more accurately identify known or novel cell markers, determine cell proportions, and estimate cell-specific expressions, complementing the existing tools particularly when study- or datatype-specific references are unreliable or unavailable. AVAILABILITY AND IMPLEMENTATION: The open-source R Scripts of CAM3.0 is freely available at https://github.com/ChiungTingWu/CAM3/(https://github.com/Bioconductor/Contributions/issues/3205). A user's guide and a vignette are provided. Chiung-Ting Wu, Dongping Du, Lulu Chen, Rujia Dai, Chunyu Liu 0001, Guoqiang Yu, Saurabh Bhardwaj, Sarah J. Parker, Robert Clarke, David M. Herrington, Yue Joseph Wang |
Bioinform. | 12 |
| 2022 | swCAM: estimation of subtype-specific expressions in individual samples with unsupervised sample-wise deconvolutionabstractMOTIVATION: Complex biological tissues are often a heterogeneous mixture of several molecularly distinct cell subtypes. Both subtype compositions and subtype-specific (STS) expressions can vary across biological conditions. Computational deconvolution aims to dissect patterns of bulk tissue data into subtype compositions and STS expressions. Existing deconvolution methods can only estimate averaged STS expressions in a population, while many downstream analyses such as inferring co-expression networks in particular subtypes require subtype expression estimates in individual samples. However, individual-level deconvolution is a mathematically underdetermined problem because there are more variables than observations. RESULTS: We report a sample-wise Convex Analysis of Mixtures (swCAM) method that can estimate subtype proportions and STS expressions in individual samples from bulk tissue transcriptomes. We extend our previous CAM framework to include a new term accounting for between-sample variations and formulate swCAM as a nuclear-norm and ℓ2,1-norm regularized matrix factorization problem. We determine hyperparameter values using cross-validation with random entry exclusion and obtain a swCAM solution using an efficient alternating direction method of multipliers. Experimental results on realistic simulation data show that swCAM can accurately estimate STS expressions in individual samples and successfully extract co-expression networks in particular subtypes that are otherwise unobtainable using bulk data. In two real-world applications, swCAM analysis of bulk RNASeq data from brain tissue of cases and controls with bipolar disorder or Alzheimer's disease identified significant changes in cell proportion, expression pattern and co-expression module in patient neurons. Comparative evaluation of swCAM versus peer methods is also provided. AVAILABILITY AND IMPLEMENTATION: The R Scripts of swCAM are freely available at https://github.com/Lululuella/swCAM. A user's guide and a vignette are provided. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lulu Chen, Chiung-Ting Wu, Chia-Hsiang Lin, Rujia Dai, Chunyu Liu 0001, Robert Clarke, Guoqiang Yu, Jennifer E. Van Eyk, David M. Herrington, Yue Joseph Wang |
Bioinform. | 10 |
| 2020 | debCAM: a bioconductor R package for fully unsupervised deconvolution of complex tissuesabstractSUMMARY: We develop a fully unsupervised deconvolution method to dissect complex tissues into molecularly distinctive tissue or cell subtypes based on bulk expression profiles. We implement an R package, deconvolution by Convex Analysis of Mixtures (debCAM) that can automatically detect tissue/cell-specific markers, determine the number of constituent subtypes, calculate subtype proportions in individual samples and estimate tissue/cell-specific expression profiles. We demonstrate the performance and biomedical utility of debCAM on gene expression, methylation, proteomics and imaging data. With enhanced data preprocessing and prior knowledge incorporation, debCAM software tool will allow biologists to perform a more comprehensive and unbiased characterization of tissue remodeling in many biomedical contexts. AVAILABILITY AND IMPLEMENTATION: http://bioconductor.org/packages/debCAM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lulu Chen, Chiung-Ting Wu, Niya Wang, David M. Herrington, Robert Clarke, Yue Joseph Wang |
Bioinform. | 6 |
| 2020 | SynQuant: an automatic tool to quantify synapses from microscopy imagesabstractMOTIVATION: Synapses are essential to neural signal transmission. Therefore, quantification of synapses and related neurites from images is vital to gain insights into the underlying pathways of brain functionality and diseases. Despite the wide availability of synaptic punctum imaging data, several issues are impeding satisfactory quantification of these structures by current tools. First, the antibodies used for labeling synapses are not perfectly specific to synapses. These antibodies may exist in neurites or other cell compartments. Second, the brightness of different neurites and synaptic puncta is heterogeneous due to the variation of antibody concentration and synapse-intrinsic differences. Third, images often have low signal to noise ratio due to constraints of experiment facilities and availability of sensitive antibodies. These issues make the detection of synapses challenging and necessitates developing a new tool to easily and accurately quantify synapses. RESULTS: We present an automatic probability-principled synapse detection algorithm and integrate it into our synapse quantification tool SynQuant. Derived from the theory of order statistics, our method controls the false discovery rate and improves the power of detecting synapses. SynQuant is unsupervised, works for both 2D and 3D data, and can handle multiple staining channels. Through extensive experiments on one synthetic and three real datasets with ground truth annotation or manually labeling, SynQuant was demonstrated to outperform peer specialized unsupervised synapse detection tools as well as generic spot detection methods. AVAILABILITY AND IMPLEMENTATION: Java source code, Fiji plug-in, and test data are available at https://github.com/yu-lab-vt/SynQuant. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yizhi Wang 0009, Congchao Wang, Petter Ranefall, Gerard Broussard, Yinxue Wang, Guilai Shi, Boyu Lyu, Chiung-Ting Wu, Yue Joseph Wang, Guoqiang Yu |
Bioinform. | 9 |
| 2020 | Targeted realignment of LC-MS profiles by neighbor-wise compound-specific graphical time warping with misalignment detectionabstractMOTIVATION: Liquid chromatography-mass spectrometry (LC-MS) is a standard method for proteomics and metabolomics analysis of biological samples. Unfortunately, it suffers from various changes in the retention times (RT) of the same compound in different samples, and these must be subsequently corrected (aligned) during data processing. Classic alignment methods such as in the popular XCMS package often assume a single time-warping function for each sample. Thus, the potentially varying RT drift for compounds with different masses in a sample is neglected in these methods. Moreover, the systematic change in RT drift across run order is often not considered by alignment algorithms. Therefore, these methods cannot effectively correct all misalignments. For a large-scale experiment involving many samples, the existence of misalignment becomes inevitable and concerning. RESULTS: Here, we describe an integrated reference-free profile alignment method, neighbor-wise compound-specific Graphical Time Warping (ncGTW), that can detect misaligned features and align profiles by leveraging expected RT drift structures and compound-specific warping functions. Specifically, ncGTW uses individualized warping functions for different compounds and assigns constraint edges on warping functions of neighboring samples. Validated with both realistic synthetic data and internal quality control samples, ncGTW applied to two large-scale metabolomics LC-MS datasets identifies many misaligned features and successfully realigns them. These features would otherwise be discarded or uncorrected using existing methods. The ncGTW software tool is developed currently as a plug-in to detect and realign misaligned features present in standard XCMS output. AVAILABILITY AND IMPLEMENTATION: An R package of ncGTW is freely available at Bioconductor and https://github.com/ChiungTingWu/ncGTW. A detailed user's manual and a vignette are provided within the package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chiung-Ting Wu, Yizhi Wang 0009, Yinxue Wang, Timothy M. D. Ebbels, Ibrahim Karaman, Gonçalo Graça, David M. Herrington, Yue Joseph Wang, Guoqiang Yu |
Bioinform. | 9 |
| 2019 | DBS: a fast and informative segmentation algorithm for DNA copy number analysisabstractBACKGROUND: Genome-wide DNA copy number changes are the hallmark events in the initiation and progression of cancers. Quantitative analysis of somatic copy number alterations (CNAs) has broad applications in cancer research. With the increasing capacity of high-throughput sequencing technologies, fast and efficient segmentation algorithms are required when characterizing high density CNAs data. RESULTS: A fast and informative segmentation algorithm, DBS (Deviation Binary Segmentation), is developed and discussed. The DBS method is based on the least absolute error principles and is inspired by the segmentation method rooted in the circular binary segmentation procedure. DBS uses point-by-point model calculation to ensure the accuracy of segmentation and combines a binary search algorithm with heuristics derived from the Central Limit Theorem. The DBS algorithm is very efficient requiring a computational complexity of O(n*log n), and is faster than its predecessors. Moreover, DBS measures the change-point amplitude of mean values of two adjacent segments at a breakpoint, where the significant degree of change-point amplitude is determined by the weighted average deviation at breakpoints. Accordingly, using the constructed binary tree of significant degree, DBS informs whether the results of segmentation are over- or under-segmented. CONCLUSION: DBS is implemented in a platform-independent and open-source Java application (ToolSeg), including a graphical user interface and simulation data generation, as well as various segmentation methods in the native Java language. Jun Ruan, Yue Joseph Wang, Junqiu Yue, Guoqiang Yu |
BMC Bioinform. | 4 |
| 2018 | Maximum Volume Inscribed Ellipsoid: A New Simplex-Structured Matrix Factorization Framework via Facet Enumeration and Convex OptimizationabstractConsider a structured matrix factorization model where one factor is restricted to have its columns lying in the unit simplex. This simplex-structured matrix factorization (SSMF) model and the associated factorization techniques have spurred much interest in research topics over different areas, such as hyperspectral unmixing in remote sensing and topic discovery in machine learning, to name a few. In this paper we develop a new theoretical SSMF framework whose idea is to study a maximum volume ellipsoid inscribed in the convex hull of the data points. This maximum volume inscribed ellipsoid (MVIE) idea has not been attempted in prior literature, and we show a sufficient condition under which the MVIE framework guarantees exact recovery of the factors. The sufficient recovery condition we show for MVIE is much more relaxed than that of separable nonnegative matrix factorization (or pure-pixel search); coincidentally, it is also identical to that of minimum volume enclosing simplex, which is known to be a powerful SSMF framework for nonseparable problem instances. We also show that MVIE can be practically implemented by performing facet enumeration and then by solving a convex optimization problem. The potential of the MVIE framework is illustrated by numerical results. Chia-Hsiang Lin, Ruiyuan Wu, Wing-Kin Ma, Chong-Yung Chi, Yue Joseph Wang |
SIAM J. Imaging Sci. | 5 |
| 2018 | Multi-Pass Fast Watershed for Accurate Segmentation of Overlapping Cervical CellsabstractThe task of segmenting cell nuclei and cytoplasm in pap smear images is one of the most challenging tasks in automated cervix cytological analysis due to specifically the presence of overlapping cells. This paper introduces a multi-pass fast watershed-based method (MPFW) to segment both nucleus and cytoplasm from large cell masses of overlapping cervical cells in three watershed passes. The first pass locates the nuclei with barrier-based watershed on the gradient-based edge map of a pre-processed image. The next pass segments the isolated, touching, and partially overlapping cells with a watershed transform adapted to the cell shape and location. The final pass introduces mutual iterative watersheds separately applied to each nucleus in the largely overlapping clusters to estimate the cell shape. In MPFW, the line-shaped contours of the watershed cells are deformed with ellipse fitting and contour adjustment to give a better representation of cell shapes. The performance of the proposed method has been evaluated using synthetic, real extended depth-of-field, and multi-layers cervical cytology images provided by the first and second overlapping cervical cytology image segmentation challenges in ISBI 2014 and ISBI 2015. The experimental results demonstrate superior performance of the proposed MPFW in terms of segmentation accuracy, detection rate, and time complexity, compared with recent peer methods. Afaf Tareef, Yang Song 0001, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Detection of Sources in Non-Negative Blind Source Separation by Minimum Description Length CriterionabstractWhile non-negative blind source separation (nBSS) has found many successful applications in science and engineering, model order selection, determining the number of sources, remains a critical yet unresolved problem. Various model order selection methods have been proposed and applied to real-world data sets but with limited success, with both order over- and under-estimation reported. By studying existing schemes, we have found that the unsatisfactory results are mainly due to invalid assumptions, model oversimplification, subjective thresholding, and/or to assumptions made solely for mathematical convenience. Building on our earlier work that reformulated model order selection for nBSS with more realistic assumptions and models, we report a newly and formally revised model order selection criterion rooted in the minimum description length (MDL) principle. Adopting widely invoked assumptions for achieving a unique nBSS solution, we consider the mixing matrix as consisting of deterministic unknowns, with the source signals following a multivariate Dirichlet distribution. We derive a computationally efficient, stochastic algorithm to obtain approximate maximum-likelihood estimates of model parameters and apply Monte Carlo integration to determine the description length. Our modeling and estimation strategy exploits the characteristic geometry of the data simplex in nBSS. We validate our nBSS-MDL criterion through extensive simulation studies and on four real-world data sets, demonstrating its strong performance and general applicability to nBSS. The proposed nBSS-MDL criterion consistently detects the true number of sources, in all of our case studies. Chia-Hsiang Lin, Chong-Yung Chi, Lulu Chen, David J. Miller 0001, Yue Joseph Wang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Automatic segmentation of overlapping cervical smear cells based on local distinctive features and guided shape deformation
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Hang Chang, Yue Joseph Wang, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 6 |
| 2017 | Optimizing the cervix cytological examination based on deep learning and dynamic shape modeling
Afaf Tareef, Yang Song 0001, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng, Tom Weidong Cai |
Neurocomputing | 4 |
| 2017 | Dual discriminative local coding for tissue aging analysis
Yang Song 0001, Qing Li 0012, Fan Zhang 0013, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai |
Medical Image Anal. | 6 |
| 2016 | Graphical Time Warping for Joint Alignment of Multiple CurvesabstractDynamic time warping (DTW) is a fundamental technique in time series analysis for comparing one curve to another using a flexible time-warping function. However, it was designed to compare a single pair of curves. In many applications, such as in metabolomics and image series analysis, alignment is simultaneously needed for multiple pairs. Because the underlying warping functions are often related, independent application of DTW to each pair is a sub-optimal solution. Yet, it is largely unknown how to efficiently conduct a joint alignment with all warping functions simultaneously considered, since any given warping function is constrained by the others and dynamic programming cannot be applied. In this paper, we show that the joint alignment problem can be transformed into a network flow problem and thus can be exactly and efficiently solved by the max flow algorithm, with a guarantee of global optimality. We name the proposed approach graphical time warping (GTW), emphasizing the graphical nature of the solution and that the dependency structure of the warping functions can be represented by a graph. Modifications of DTW, such as windowing and weighting, are readily derivable within GTW. We also discuss optimal tuning of parameters and hyperparameters in GTW. We illustrate the power of GTW using both synthetic data and a real case study of an astrocyte calcium movie. Yizhi Wang 0009, David J. Miller 0001, Kira Poskanzer, Yue Joseph Wang, Guoqiang Yu |
NIPS | 4 |
| 2016 | Bioimage classification with subcategory discriminant transform of high dimensional visual descriptorsabstractBACKGROUND: Bioimage classification is a fundamental problem for many important biological studies that require accurate cell phenotype recognition, subcellular localization, and histopathological classification. In this paper, we present a new bioimage classification method that can be generally applicable to a wide variety of classification problems. We propose to use a high-dimensional multi-modal descriptor that combines multiple texture features. We also design a novel subcategory discriminant transform (SDT) algorithm to further enhance the discriminative power of descriptors by learning convolution kernels to reduce the within-class variation and increase the between-class difference. RESULTS: We evaluate our method on eight different bioimage classification tasks using the publicly available IICBU 2008 database. Each task comprises a separate dataset, and the collection represents typical subcellular, cellular, and tissue level classification problems. Our method demonstrates improved classification accuracy (0.9 to 9%) on six tasks when compared to state-of-the-art approaches. We also find that SDT outperforms the well-known dimension reduction techniques, with for example 0.2 to 13% improvement over linear discriminant analysis. CONCLUSIONS: We present a general bioimage classification method, which comprises a highly descriptive visual feature representation and a learning-based discriminative feature transformation algorithm. Our evaluation on the IICBU 2008 database demonstrates improved performance over the state-of-the-art for six different classification tasks. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang |
BMC Bioinform. | 5 |
| 2015 | Learning Shape-Driven Segmentation Based on Neural Network and Sparse Reconstruction Toward Automated Cell Analysis of Cervical Smears
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng |
ICONIP (1) | 5 |
| 2015 | BMRF-Net: a software tool for identification of protein interaction subnetworks by a bagging Markov random field-based methodabstractUNLABELLED: Identification of protein interaction subnetworks is an important step to help us understand complex molecular mechanisms in cancer. In this paper, we develop a BMRF-Net package, implemented in Java and C++, to identify protein interaction subnetworks based on a bagging Markov random field (BMRF) framework. By integrating gene expression data and protein-protein interaction data, this software tool can be used to identify biologically meaningful subnetworks. A user friendly graphic user interface is developed as a Cytoscape plugin for the BMRF-Net software to deal with the input/output interface. The detailed structure of the identified networks can be visualized in Cytoscape conveniently. The BMRF-Net package has been applied to breast cancer data to identify significant subnetworks related to breast cancer recurrence. AVAILABILITY AND IMPLEMENTATION: The BMRF-Net package is available at http://sourceforge.net/projects/bmrfcjava/. The package is tested under Ubuntu 12.04 (64-bit), Java 7, glibc 2.15 and Cytoscape 3.1.0. Xu Shi 0003, Robert O. Barnes, Li Chen 0018, Ayesha N. Shajahan, Leena Hilakivi-Clarke, Robert Clarke, Yue Joseph Wang, Jianhua Xuan |
Bioinform. | 7 |
| 2015 | KDDN: an open-source Cytoscape app for constructing differential dependency networks with significant rewiringabstractUNLABELLED: We have developed an integrated molecular network learning method, within a well-grounded mathematical framework, to construct differential dependency networks with significant rewiring. This knowledge-fused differential dependency networks (KDDN) method, implemented as a Java Cytoscape app, can be used to optimally integrate prior biological knowledge with measured data to simultaneously construct both common and differential networks, to quantitatively assign model parameters and significant rewiring p-values and to provide user-friendly graphical results. The KDDN algorithm is computationally efficient and provides users with parallel computing capability using ubiquitous multi-core machines. We demonstrate the performance of KDDN on various simulations and real gene expression datasets, and further compare the results with those obtained by the most relevant peer methods. The acquired biologically plausible results provide new insights into network rewiring as a mechanistic principle and illustrate KDDN's ability to detect them efficiently and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY: Source code and compiled package are freely available for download at http://apps.cytoscape.org/apps/kddn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bai Zhang, Eric P. Hoffman, Robert Clarke, Ie-Ming Shih, Jianhua Xuan, David M. Herrington, Yue Joseph Wang |
Bioinform. | 9 |
| 2015 | UNDO: a Bioconductor R package for unsupervised deconvolution of mixed gene expressions in tumor samplesabstractSUMMARY: We develop a novel unsupervised deconvolution method, within a well-grounded mathematical framework, to dissect mixed gene expressions in heterogeneous tumor samples. We implement an R package, UNsupervised DecOnvolution (UNDO), that can be used to automatically detect cell-specific marker genes (MGs) located on the scatter radii of mixed gene expressions, estimate cellular proportions in each sample and deconvolute mixed expressions into cell-specific expression profiles. We demonstrate the performance of UNDO over a wide range of tumor-stroma mixing proportions, validate UNDO on various biologically mixed benchmark gene expression datasets and further estimate tumor purity in TCGA/CPTAC datasets. The highly accurate deconvolution results obtained suggest not only the existence of cell-specific MGs but also UNDO's ability to detect them blindly and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY AND IMPLEMENTATION: UNDO is available at http://bioconductor.org/packages. Niya Wang, Robert Clarke, Lulu Chen, Ie-Ming Shih, Douglas A. Levine, Jianhua Xuan, Yue Joseph Wang |
Bioinform. | 9 |
| 2015 | SIMAT: GC-SIM-MS data analysis toolabstractBACKGROUND: Gas chromatography coupled with mass spectrometry (GC-MS) is one of the technologies widely used for qualitative and quantitative analysis of small molecules. In particular, GC coupled to single quadrupole MS can be utilized for targeted analysis by selected ion monitoring (SIM). However, to our knowledge, there are no software tools specifically designed for analysis of GC-SIM-MS data. In this paper, we introduce a new R/Bioconductor package called SIMAT for quantitative analysis of the levels of targeted analytes. SIMAT provides guidance in choosing fragments for a list of targets. This is accomplished through an optimization algorithm that has the capability to select the most appropriate fragments from overlapping chromatographic peaks based on a pre-specified library of background analytes. The tool also allows visualization of the total ion chromatograms (TIC) of runs and extracted ion chromatograms (EIC) of analytes of interest. Moreover, retention index (RI) calibration can be performed and raw GC-SIM-MS data can be imported in netCDF or NIST mass spectral library (MSL) formats. RESULTS: We evaluated the performance of SIMAT using two GC-SIM-MS datasets obtained by targeted analysis of: (1) plasma samples from 86 patients in a targeted metabolomic experiment; and (2) mixtures of internal standards spiked in plasma samples at varying concentrations in a method development study. Our results demonstrate that SIMAT offers alternative solutions to AMDIS and MetaboliteDetector to achieve accurate detection of targets and estimation of their relative intensities by analysis of GC-SIM-MS data. CONCLUSIONS: We introduce a new R package called SIMAT that allows the selection of the optimal set of fragments and retention time windows for target analytes in GC-SIM-MS based analysis. Also, various functions and algorithms are implemented in the tool to: (1) read and import raw data and spectral libraries; (2) perform GC-SIM-MS data preprocessing; and (3) plot and visualize EICs and TICs. Mohammad R. Nezami Ranjbar, Cristina Di Poto, Yue Joseph Wang, Habtom W. Ressom |
BMC Bioinform. | 3 |
| 2015 | Locality-constrained Subcluster Representation Ensemble for lung image classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, Yue Joseph Wang, David Dagan Feng |
Medical Image Anal. | 5 |
| 2015 | Bayesian Normalization Model for Label-Free Quantitative Analysis by LC-MSabstractWe introduce a new method for normalization of data acquired by liquid chromatography coupled with mass spectrometry (LC-MS) in label-free differential expression analysis. Normalization of LC-MS data is desired prior to subsequent statistical analysis to adjust variabilities in ion intensities that are not caused by biological differences but experimental bias. There are different sources of bias including variabilities during sample collection and sample storage, poor experimental design, noise, etc. In addition, instrument variability in experiments involving a large number of LC-MS runs leads to a significant drift in intensity measurements. Although various methods have been proposed for normalization of LC-MS data, there is no universally applicable approach. In this paper, we propose a Bayesian normalization model (BNM) that utilizes scan-level information from LC-MS data. Specifically, the proposed method uses peak shapes to model the scan-level data acquired from extracted ion chromatograms (EIC) with parameters considered as a linear mixed effects model. We extended the model into BNM with drift (BNMD) to compensate for the variability in intensity measurements due to long LC-MS runs. We evaluated the performance of our method using synthetic and experimental data. In comparison with several existing methods, the proposed BNM and BNMD yielded significant improvement. Mohammad R. Nezami Ranjbar, Mahlet G. Tadesse, Yue Joseph Wang, Habtom W. Ressom |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2015 | Large Margin Local Estimate With Applications to Medical Image ClassificationabstractMedical images usually exhibit large intra-class variation and inter-class ambiguity in the feature space, which could affect classification accuracy. To tackle this issue, we propose a new Large Margin Local Estimate (LMLE) classification model with sub-categorization based sparse representation. We first sub-categorize the reference sets of different classes into multiple clusters, to reduce feature variation within each subcategory compared to the entire reference set. Local estimates are generated for the test image using sparse representation with reference subcategories as the dictionaries. The similarity between the test image and each class is then computed by fusing the distances with the local estimates in a learning-based large margin aggregation construct to alleviate the problem of inter-class ambiguity. The derived similarities are finally used to determine the class label. We demonstrate that our LMLE model is generally applicable to different imaging modalities, and applied it to three tasks: interstitial lung disease (ILD) classification on high-resolution computed tomography (HRCT) images, phenotype binary classification and continuous regression on brain magnetic resonance (MR) imaging. Our experimental results show statistically significant performance improvements over existing popular classifiers. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng, Yue Joseph Wang, Michael J. Fulham |
IEEE Trans. Medical Imaging | 6 |
| 2014 | Ensemble random projection for multi-label classification with application to protein subcellular localizationabstractThe curse of dimensionality severely restricts the predictive power of multi-label classification systems. High-dimensional feature vectors may contain redundant or irrelevant information, causing the classification systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble of multilabel classifiers. The merits of the proposed method are demonstrated through a multi-label protein classification task. Specifically, high-dimensional feature vectors are extracted from protein sequences using the gene ontology (GO) and Swiss-Prot databases. The feature vectors are then projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for predicting the subcellular localization of proteins. Experimental results suggest that the proposed method can reduce the dimensions by seven folds and impressively improve the classification performance. Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung |
ICASSP | 4 |
| 2014 | Robust identification of transcriptional regulatory networks using a Gibbs sampler on outlier sum statisticabstractContact: [email protected] Bioinformatics (2012) 28 (15), 1990–1997 doi:10.1093/bioinformatics/bts296 The authors wish to add one citation to a relevant conference report, Gu, J., Xuan, J., Wang, Y., Riggins, R.B. and Clarke R. (2010) Identification of transcriptional regulatory networks by learning the marginal function of outlier sum statistic. Proceedings of International Conference on Machine Learning and Applications , 281–286. The formatted reference is given below and should read in the sentence: In particular, a novel statistic for testing the confidence of target genes, namely, outlier sum of regression t -statistic ( Gu et al. , 2010 ), is specifically designed to pin-down confident target genes; based on this statistic, a Gibbs sampling strategy is used to sample target genes in a high probability as governed by the underlying distribution. The authors apologize for this oversight. Jinghua Gu, Jianhua Xuan, Rebecca B. Riggins, Li Chen 0018, Yue Joseph Wang, Robert Clarke |
Bioinform. | 5 |
| 2014 | AISAIC: a software suite for accurate identification of significant aberrations in cancersabstractUNLABELLED: Accurate identification of significant aberrations in cancers (AISAIC) is a systematic effort to discover potential cancer-driving genes such as oncogenes and tumor suppressors. Two major confounding factors against this goal are the normal cell contamination and random background aberrations in tumor samples. We describe a Java AISAIC package that provides comprehensive analytic functions and graphic user interface for integrating two statistically principled in silico approaches to address the aforementioned challenges in DNA copy number analyses. In addition, the package provides a command-line interface for users with scripting and programming needs to incorporate or extend AISAIC to their customized analysis pipelines. This open-source multiplatform software offers several attractive features: (i) it implements a user friendly complete pipeline from processing raw data to reporting analytic results; (ii) it detects deletion types directly from copy number signals using a Bayes hypothesis test; (iii) it estimates the fraction of normal contamination for each sample; (iv) it produces unbiased null distribution of random background alterations by iterative aberration-exclusive permutations; and (v) it identifies significant consensus regions and the percentage of homozygous/hemizygous deletions across multiple samples. AISAIC also provides users with a parallel computing option to leverage ubiquitous multicore machines. AVAILABILITY AND IMPLEMENTATION: AISAIC is available as a Java application, with a user's guide and source code, at https://code.google.com/p/aisaic/. Bai Zhang, Xuchu Hou, Xiguo Yuan, Ie-Ming Shih, Robert Clarke, Roger R. Wang, Subha Madhavan, Yue Joseph Wang, Guoqiang Yu |
Bioinform. | 10 |
| 2014 | Integration of Network Biology and Imaging to Study Cancer Phenotypes and ResponsesabstractEver growing "omics" data and continuously accumulated biological knowledge provide an unprecedented opportunity to identify molecular biomarkers and their interactions that are responsible for cancer phenotypes that can be accurately defined by clinical measurements such as in vivo imaging. Since signaling or regulatory networks are dynamic and context-specific, systematic efforts to characterize such structural alterations must effectively distinguish significant network rewiring from random background fluctuations. Here we introduced a novel integration of network biology and imaging to study cancer phenotypes and responses to treatments at the molecular systems level. Specifically, Differential Dependence Network (DDN) analysis was used to detect statistically significant topological rewiring in molecular networks between two phenotypic conditions, and in vivo Magnetic Resonance Imaging (MRI) was used to more accurately define phenotypic sample groups for such differential analysis. We applied DDN to analyze two distinct phenotypic groups of breast cancer and study how genomic instability affects the molecular network topologies in high-grade ovarian cancer. Further, FDA-approved arsenic trioxide (ATO) and the ND2-SmoA1 mouse model of Medulloblastoma (MB) were used to extend our analyses of combined MRI and Reverse Phase Protein Microarray (RPMA) data to assess tumor responses to ATO and to uncover the complexity of therapeutic molecular biology. Sean S. Wang, Olga C. Rodriguez, Emanuel Petricoin III, Ie-Ming Shih, Daniel Chan, Maria Avantaggiati, Guoqiang Yu, Shaozhen Ye, Robert Clarke, Chao Wang 0005, Bai Zhang, Yue Joseph Wang, Chris Albanese |
IEEE ACM Trans. Comput. Biol. Bioinform. | 14 |
| 2013 | An ensemble classifier with random projection for predicting multi-label protein subcellular localizationabstractIn protein subcellular localization prediction, a predominant scenario is that the number of available features is much larger than the number of data samples. Among the large number of features, many of them may contain redundant or irrelevant information, causing the prediction systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble multi-label classifier for predicting protein subcellular localization. Specifically, the frequencies of occurrences of gene-ontology terms are used as feature vectors, which are projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for making the final decision. Experimental results on two recent datasets suggest that the proposed method can reduce the dimensions by six folds and remarkably improve the classification performance. Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung |
BIBM | 4 |
| 2013 | Multi-profile Bayesian alignment model for LC-MS data analysis with integration of internal standardsabstractMOTIVATION: Liquid chromatography-mass spectrometry (LC-MS) has been widely used for profiling expression levels of biomolecules in various '-omic' studies including proteomics, metabolomics and glycomics. Appropriate LC-MS data preprocessing steps are needed to detect true differences between biological groups. Retention time (RT) alignment, which is required to ensure that ion intensity measurements among multiple LC-MS runs are comparable, is one of the most important yet challenging preprocessing steps. Current alignment approaches estimate RT variability using either single chromatograms or detected peaks, but do not simultaneously take into account the complementary information embedded in the entire LC-MS data. RESULTS: We propose a Bayesian alignment model for LC-MS data analysis. The alignment model provides estimates of the RT variability along with uncertainty measures. The model enables integration of multiple sources of information including internal standards and clustered chromatograms in a mathematically rigorous framework. We apply the model to LC-MS metabolomic, proteomic and glycomic data. The performance of the model is evaluated based on ground-truth data, by measuring correlation of variation, RT difference across runs and peak-matching performance. We demonstrate that Bayesian alignment model improves significantly the RT alignment performance through appropriate integration of relevant information. AVAILABILITY AND IMPLEMENTATION: MATLAB code, raw and preprocessed LC-MS data are available at http://omics.georgetown.edu/alignLCMS.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tsung-Heng Tsai, Mahlet G. Tadesse, Cristina Di Poto, Lewis K. Pannell, Yehia Mechref, Yue Joseph Wang, Habtom W. Ressom |
Bioinform. | 6 |
| 2013 | Region-based progressive localization of cell nuclei in microscopic images with data adaptive modelingabstractBACKGROUND: Segmenting cell nuclei in microscopic images has become one of the most important routines in modern biological applications. With the vast amount of data, automatic localization, i.e. detection and segmentation, of cell nuclei is highly desirable compared to time-consuming manual processes. However, automated segmentation is challenging due to large intensity inhomogeneities in the cell nuclei and the background. RESULTS: We present a new method for automated progressive localization of cell nuclei using data-adaptive models that can better handle the inhomogeneity problem. We perform localization in a three-stage approach: first identify all interest regions with contrast-enhanced salient region detection, then process the clusters to identify true cell nuclei with probability estimation via feature-distance profiles of reference regions, and finally refine the contours of detected regions with regional contrast-based graphical model. The proposed region-based progressive localization (RPL) method is evaluated on three different datasets, with the first two containing grayscale images, and the third one comprising of color images with cytoplasm in addition to cell nuclei. We demonstrate performance improvement over the state-of-the-art. For example, compared to the second best approach, on the first dataset, our method achieves 2.8 and 3.7 reduction in Hausdorff distance and false negatives; on the second dataset that has larger intensity inhomogeneity, our method achieves 5% increase in Dice coefficient and Rand index; on the third dataset, our method achieves 4% increase in object-level accuracy. CONCLUSIONS: To tackle the intensity inhomogeneities in cell nuclei and background, a region-based progressive localization method is proposed for cell nuclei localization in fluorescence microscopy images. The RPL method is demonstrated highly effective on three different public datasets, with on average 3.5% and 7% improvement of region- and contour-based segmentation performance over the state-of-the-art. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng |
BMC Bioinform. | 4 |
| 2013 | The CAM software for nonnegative blind source separation in R-Java
Niya Wang, Li Chen 0018, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Yue Joseph Wang |
J. Mach. Learn. Res. | 8 |
| 2013 | Profile-Based LC-MS Data Alignment-A Bayesian ApproachabstractA Bayesian alignment model (BAM) is proposed for alignment of liquid chromatography-mass spectrometry (LC-MS) data. BAM belongs to the category of profile-based approaches, which are composed of two major components: a prototype function and a set of mapping functions. Appropriate estimation of these functions is crucial for good alignment results. BAM uses Markov chain Monte Carlo (MCMC) methods to draw inference on the model parameters and improves on existing MCMC-based alignment methods through 1) the implementation of an efficient MCMC sampler and 2) an adaptive selection of knots. A block Metropolis-Hastings algorithm that mitigates the problem of the MCMC sampler getting stuck at local modes of the posterior distribution is used for the update of the mapping function coefficients. In addition, a stochastic search variable selection (SSVS) methodology is used to determine the number and positions of knots. We applied BAM to a simulated data set, an LC-MS proteomic data set, and two LC-MS metabolomic data sets, and compared its performance with the Bayesian hierarchical curve registration (BHCR) model, the dynamic time-warping (DTW) model, and the continuous profile model (CPM). The advantage of applying appropriate profile-based retention time correction prior to performing a feature-based approach is also demonstrated through the metabolomic data sets. Tsung-Heng Tsai, Mahlet G. Tadesse, Yue Joseph Wang, Habtom W. Ressom |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2012 | Robust identification of transcriptional regulatory networks using a Gibbs sampler on outlier sum statisticabstractMOTIVATION: Identification of transcriptional regulatory networks (TRNs) is of significant importance in computational biology for cancer research, providing a critical building block to unravel disease pathways. However, existing methods for TRN identification suffer from the inclusion of excessive 'noise' in microarray data and false-positives in binding data, especially when applied to human tumor-derived cell line studies. More robust methods that can counteract the imperfection of data sources are therefore needed for reliable identification of TRNs in this context. RESULTS: In this article, we propose to establish a link between the quality of one target gene to represent its regulator and the uncertainty of its expression to represent other target genes. Specifically, an outlier sum statistic was used to measure the aggregated evidence for regulation events between target genes and their corresponding transcription factors. A Gibbs sampling method was then developed to estimate the marginal distribution of the outlier sum statistic, hence, to uncover underlying regulatory relationships. To evaluate the effectiveness of our proposed method, we compared its performance with that of an existing sampling-based method using both simulation data and yeast cell cycle data. The experimental results show that our method consistently outperforms the competing method in different settings of signal-to-noise ratio and network topology, indicating its robustness for biological applications. Finally, we applied our method to breast cancer cell line data and demonstrated its ability to extract biologically meaningful regulatory modules related to estrogen signaling and action in breast cancer. AVAILABILITY AND IMPLEMENTATION: The Gibbs sampler MATLAB package is freely available at http://www.cbil.ece.vt.edu/software.htm. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jinghua Gu, Jianhua Xuan, Rebecca B. Riggins, Li Chen 0018, Yue Joseph Wang, Robert Clarke |
Bioinform. | 5 |
| 2012 | Computational analysis of muscular dystrophy sub-types using a novel integrative scheme
Chen Wang 0001, Sook Shin Ha, Jianhua Xuan, Yue Joseph Wang, Eric P. Hoffman |
Neurocomputing | 4 |
| 2012 | Regulatory component analysis: A semi-blind extraction approach to infer gene regulatory networks with imperfect biological knowledge
Chen Wang 0001, Jianhua Xuan, Ie-Ming Shih, Robert Clarke, Yue Joseph Wang |
Signal Process. | 5 |
| 2012 | Nonlinear System Modeling With Random Matrices: Echo State Networks RevisitedabstractEcho state networks (ESNs) are a novel form of recurrent neural networks (RNNs) that provide an efficient and powerful computational model approximating nonlinear dynamical systems. A unique feature of an ESN is that a large number of neurons (the "reservoir") are used, whose synaptic connections are generated randomly, with only the connections from the reservoir to the output modified by learning. Why a large randomly generated fixed RNN gives such excellent performance in approximating nonlinear systems is still not well understood. In this brief, we apply random matrix theory to examine the properties of random reservoirs in ESNs under different topologies (sparse or fully connected) and connection weights (Bernoulli or Gaussian). We quantify the asymptotic gap between the scaling factor bounds for the necessary and sufficient conditions previously proposed for the echo state property. We then show that the state transition mapping is contractive with high probability when only the necessary condition is satisfied, which corroborates and thus analytically explains the observation that in practice one obtains echo states when the spectral radius of the reservoir weight matrix is smaller than 1. Bai Zhang, David J. Miller 0001, Yue Joseph Wang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2011 | Bayesian Alignment Model for LC-MS DataabstractA Bayesian alignment model (BAM) is proposed for alignment of liquid chromatography-mass spectrometry (LC-MS) data. BAM is composed of two important components: prototype function and mapping function. Estimation of both functions is crucial for the alignment result. We use Markov chain Monte Carlo (MCMC) methods for inference of model parameters. To address the trapping effect in local modes, we propose a block Metropolis-Hastings algorithm that leads to better mixing behavior in updating the mapping function coefficients. We applied BAM to both simulated and real LC-MS datasets, and compared its performance with the Bayesian hierarchical curve registration model (BHCR). Performance evaluation on both simulated and real datasets shows satisfactory results in terms of correlation coefficients and ratio of overlapping peak areas. Tsung-Heng Tsai, Mahlet G. Tadesse, Yue Joseph Wang, Habtom W. Ressom |
BIBM | 3 |
| 2011 | CAM-CM: a signal deconvolution tool for in vivo dynamic contrast-enhanced imaging of complex tissuesabstractSUMMARY: In vivo dynamic contrast-enhanced imaging tools provide non-invasive methods for analyzing various functional changes associated with disease initiation, progression and responses to therapy. The quantitative application of these tools has been hindered by its inability to accurately resolve and characterize targeted tissues due to spatially mixed tissue heterogeneity. Convex Analysis of Mixtures - Compartment Modeling (CAM-CM) signal deconvolution tool has been developed to automatically identify pure-volume pixels located at the corners of the clustered pixel time series scatter simplex and subsequently estimate tissue-specific pharmacokinetic parameters. CAM-CM can dissect complex tissues into regions with differential tracer kinetics at pixel-wise resolution and provide a systems biology tool for defining imaging signatures predictive of phenotypes. AVAILABILITY: The MATLAB source code can be downloaded at the authors' website www.cbil.ece.vt.edu/software.htm CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li Chen 0018, Tsung-Han Chan, Peter L. Choyke, Elizabeth M. C. Hillman, Chong-Yung Chi, Zaver M. Bhujwalla, Ge Wang 0001, Sean S. Wang, Zsolt Szabo, Yue Joseph Wang |
Bioinform. | 10 |
| 2011 | PUGSVM: a caBIGTM analytical tool for multiclass gene selection and predictive classificationabstractUNLABELLED: Phenotypic Up-regulated Gene Support Vector Machine (PUGSVM) is a cancer Biomedical Informatics Grid (caBIG™) analytical tool for multiclass gene selection and classification. PUGSVM addresses the problem of imbalanced class separability, small sample size and high gene space dimensionality, where multiclass gene markers are defined by the union of one-versus-everyone phenotypic upregulated genes, and used by a well-matched one-versus-rest support vector machine. PUGSVM provides a simple yet more accurate strategy to identify statistically reproducible mechanistic marker genes for characterization of heterogeneous diseases. AVAILABILITY: http://www.cbil.ece.vt.edu/caBIG-PUGSVM.htm. Guoqiang Yu, Huai Li, Sook Shin Ha, Ie-Ming Shih, Robert Clarke, Eric P. Hoffman, Subha Madhavan, Jianhua Xuan, Yue Joseph Wang |
Bioinform. | 9 |
| 2011 | BACOM: in silico detection of genomic deletion types and correction of normal cell contamination in copy number dataabstractMOTIVATION: Identification of somatic DNA copy number alterations (CNAs) and significant consensus events (SCEs) in cancer genomes is a main task in discovering potential cancer-driving genes such as oncogenes and tumor suppressors. The recent development of SNP array technology has facilitated studies on copy number changes at a genome-wide scale with high resolution. However, existing copy number analysis methods are oblivious to normal cell contamination and cannot distinguish between contributions of cancerous and normal cells to the measured copy number signals. This contamination could significantly confound downstream analysis of CNAs and affect the power to detect SCEs in clinical samples. RESULTS: We report here a statistically principled in silico approach, Bayesian Analysis of COpy number Mixtures (BACOM), to accurately estimate genomic deletion type and normal tissue contamination, and accordingly recover the true copy number profile in cancer cells. We tested the proposed method on two simulated datasets, two prostate cancer datasets and The Cancer Genome Atlas high-grade ovarian dataset, and obtained very promising results supported by the ground truth and biological plausibility. Moreover, based on a large number of comparative simulation studies, the proposed method gives significantly improved power to detect SCEs after in silico correction of normal tissue contamination. We develop a cross-platform open-source Java application that implements the whole pipeline of copy number analysis of heterogeneous cancer tissues including relevant processing steps. We also provide an R interface, bacomR, for running BACOM within the R environment, making it straightforward to include in existing data pipelines. AVAILABILITY: The cross-platform, stand-alone Java application, BACOM, the R interface, bacomR, all source code and the simulation data used in this article are freely available at authors' web site: http://www.cbil.ece.vt.edu/software.htm. Guoqiang Yu, Bai Zhang, G. Steven Bova, Ie-Ming Shih, Yue Joseph Wang |
Bioinform. | 6 |
| 2011 | DDN: a caBIG® analytical tool for differential network analysisabstractUNLABELLED: Differential dependency network (DDN) is a caBIG® (cancer Biomedical Informatics Grid) analytical tool for detecting and visualizing statistically significant topological changes in transcriptional networks representing two biological conditions. Developed under caBIG®'s In Silico Research Centers of Excellence (ISRCE) Program, DDN enables differential network analysis and provides an alternative way for defining network biomarkers predictive of phenotypes. DDN also serves as a useful systems biology tool for users across biomedical research communities to infer how genetic, epigenetic or environment variables may affect biological networks and clinical phenotypes. Besides the standalone Java application, we have also developed a Cytoscape plug-in, CytoDDN, to integrate network analysis and visualization seamlessly. AVAILABILITY: The Java and MATLAB source code can be downloaded at the authors' web site http://www.cbil.ece.vt.edu/software.htm. Bai Zhang, Huai Li, Ie-Ming Shih, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Leena Hilakivi-Clarke, Yue Joseph Wang |
Bioinform. | 11 |
| 2011 | Motif-guided sparse decomposition of gene expression data for regulatory module identificationabstractBACKGROUND: Genes work coordinately as gene modules or gene networks. Various computational approaches have been proposed to find gene modules based on gene expression data; for example, gene clustering is a popular method for grouping genes with similar gene expression patterns. However, traditional gene clustering often yields unsatisfactory results for regulatory module identification because the resulting gene clusters are co-expressed but not necessarily co-regulated. RESULTS: We propose a novel approach, motif-guided sparse decomposition (mSD), to identify gene regulatory modules by integrating gene expression data and DNA sequence motif information. The mSD approach is implemented as a two-step algorithm comprising estimates of (1) transcription factor activity and (2) the strength of the predicted gene regulation event(s). Specifically, a motif-guided clustering method is first developed to estimate the transcription factor activity of a gene module; sparse component analysis is then applied to estimate the regulation strength, and so predict the target genes of the transcription factors. The mSD approach was first tested for its improved performance in finding regulatory modules using simulated and real yeast data, revealing functionally distinct gene modules enriched with biologically validated transcription factors. We then demonstrated the efficacy of the mSD approach on breast cancer cell line data and uncovered several important gene regulatory modules related to endocrine therapy of breast cancer. CONCLUSION: We have developed a new integrated strategy, namely motif-guided sparse decomposition (mSD) of gene expression data, for regulatory module identification. The mSD method features a novel motif-guided clustering method for transcription factor activity estimation by finding a balance between co-regulation and co-expression. The mSD method further utilizes a sparse decomposition method for regulation strength estimation. The experimental results show that such a motif-guided strategy can provide context-specific regulatory modules in both yeast and breast cancer studies. Jianhua Xuan, Li Chen 0018, Rebecca B. Riggins, Huai Li, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
BMC Bioinform. | 8 |
| 2011 | Tissue-Specific Compartmental Analysis for Dynamic Contrast-Enhanced MR Imaging of Complex TumorsabstractDynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) provides a noninvasive method for evaluating tumor vasculature patterns based on contrast accumulation and washout. However, due to limited imaging resolution and tumor tissue heterogeneity, tracer concentrations at many pixels often represent a mixture of more than one distinct compartment. This pixel-wise partial volume effect (PVE) would have profound impact on the accuracy of pharmacokinetics studies using existing compartmental modeling (CM) methods. We, therefore, propose a convex analysis of mixtures (CAM) algorithm to explicitly mitigate PVE by expressing the kinetics in each pixel as a nonnegative combination of underlying compartments and subsequently identifying pure volume pixels at the corners of the clustered pixel time series scatter plot simplex. The algorithm is supported theoretically by a well-grounded mathematical framework and practically by plug-in noise filtering and normalization preprocessing. We demonstrate the principle and feasibility of the CAM-CM approach on realistic synthetic data involving two functional tissue compartments, and compare the accuracy of parameter estimates obtained with and without PVE elimination using CAM or other relevant techniques. Experimental results show that CAM-CM achieves a significant improvement in the accuracy of kinetic parameter estimation. We apply the algorithm to real DCE-MRI breast cancer data and observe improved pharmacokinetic parameter estimation, separating tumor tissue into regions with differential tracer kinetics on a pixel-by-pixel basis and revealing biologically plausible tumor tissue heterogeneity patterns. This method combines the advantages of multivariate clustering, convex geometry analysis, and compartmental modeling approaches. The open-source MATLAB software of CAM-CM is publicly available from the Web. Li Chen 0018, Peter L. Choyke, Tsung-Han Chan, Chong-Yung Chi, Ge Wang 0001, Yue Joseph Wang |
IEEE Trans. Medical Imaging | 6 |
| 2010 | Identification of Transcriptional Regulatory Networks by Learning the Marginal Function of Outlier Sum StatisticabstractNetwork component analysis (NCA) and other methods based on the NCA model have become powerful bioinformatics tools to reconstruct underlying regulatory networks and recover hidden biological processes. However, due to the existence of experimental noises in micro array data and false information in network connectivity data (e.g., ChIP-on-chip binding data, motif information, etc.), it still remains challenging to reconstruct gene regulatory networks for real biomedical applications such as human cancer studies. In this paper, we model the relationship between the genes that share the same transcription factors (TF) from the angle of regression. We propose a statistic called outlier sum testing the conditional significance of the target genes. A Gibbs strategy is utilized in order to estimate the marginal value of outlier sum from its conditional function. Based on the outlier sum statistic we are able to extract the true target genes that carry information about transcription factor activities (TFAs) from the whole population. As a proof-of-concept, we demonstrated the efficiency and robustness of the proposed method on both simulation data and yeast cell cycle data. Jinghua Gu, Jianhua Xuan, Yue Joseph Wang, Rebecca B. Riggins, Robert Clarke |
ICMLA | 3 |
| 2010 | Computational Analysis of Muscular Dystrophy Sub-types Using a Novel Integrative SchemeabstractTo construct biologically interpretable features and facilitate Muscular Dystrophy (MD) sub-types classification, we propose a novel integrative scheme utilizing PPI network, functional gene sets information, and mRNA profiling. The workflow of the proposed scheme includes three major steps: First, by combining protein-protein interaction network structure and gene co-expression relationship into new distance metric, we apply affinity propagation clustering to build gene sub-networks. Secondly, we further incorporate functional gene sets knowledge to complement the physical interaction information. Finally, based on constructed sub-network and gene set features, we apply multi-class support vector machine (MSVM) for MD sub-type classification, and highlight the biomarkers contributing to the sub-type prediction. The experimental results show that our scheme could construct sub-networks that are more relevant to MD than those constructed by conventional approach. Furthermore, our integrative strategy substantially improved the prediction accuracy, especially for those hard-to-classify sub-types. Chen Wang 0001, Sook Shin Ha, Yue Joseph Wang, Jianhua Xuan, Eric P. Hoffman |
ICMLA | 3 |
| 2010 | Learning Structural Changes of Gaussian Graphical Models in Controlled Experiments
Bai Zhang, Yue Joseph Wang |
UAI | 2 |
| 2010 | Multilevel support vector regression analysis to identify condition-specific regulatory networksabstractMOTIVATION: The identification of gene regulatory modules is an important yet challenging problem in computational biology. While many computational methods have been proposed to identify regulatory modules, their initial success is largely compromised by a high rate of false positives, especially when applied to human cancer studies. New strategies are needed for reliable regulatory module identification. RESULTS: We present a new approach, namely multilevel support vector regression (ml-SVR), to systematically identify condition-specific regulatory modules. The approach is built upon a multilevel analysis strategy designed for suppressing false positive predictions. With this strategy, a regulatory module becomes ever more significant as more relevant gene sets are formed at finer levels. At each level, a two-stage support vector regression (SVR) method is utilized to help reduce false positive predictions by integrating binding motif information and gene expression data; a significant analysis procedure is followed to assess the significance of each regulatory module. To evaluate the effectiveness of the proposed strategy, we first compared the ml-SVR approach with other existing methods on simulation data and yeast cell cycle data. The resulting performance shows that the ml-SVR approach outperforms other methods in the identification of both regulators and their target genes. We then applied our method to breast cancer cell line data to identify condition-specific regulatory modules associated with estrogen treatment. Experimental results show that our method can identify biologically meaningful regulatory modules related to estrogen signaling and action in breast cancer. AVAILABILITY AND IMPLEMENTATION: The ml-SVR MATLAB package can be downloaded at http://www.cbil.ece.vt.edu/software.htm. Li Chen 0018, Jianhua Xuan, Rebecca B. Riggins, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
Bioinform. | 4 |
| 2010 | Knowledge-guided gene ranking by coordinative component analysisabstractBACKGROUND: In cancer, gene networks and pathways often exhibit dynamic behavior, particularly during the process of carcinogenesis. Thus, it is important to prioritize those genes that are strongly associated with the functionality of a network. Traditional statistical methods are often inept to identify biologically relevant member genes, motivating researchers to incorporate biological knowledge into gene ranking methods. However, current integration strategies are often heuristic and fail to incorporate fully the true interplay between biological knowledge and gene expression data. RESULTS: To improve knowledge-guided gene ranking, we propose a novel method called coordinative component analysis (COCA) in this paper. COCA explicitly captures those genes within a specific biological context that are likely to be expressed in a coordinative manner. Formulated as an optimization problem to maximize the coordinative effort, COCA is designed to first extract the coordinative components based on a partial guidance from knowledge genes and then rank the genes according to their participation strengths. An embedded bootstrapping procedure is implemented to improve statistical robustness of the solutions. COCA was initially tested on simulation data and then on published gene expression microarray data to demonstrate its improved performance as compared to traditional statistical methods. Finally, the COCA approach has been applied to stem cell data to identify biologically relevant genes in signaling pathways. As a result, the COCA approach uncovers novel pathway members that may shed light into the pathway deregulation in cancers. CONCLUSION: We have developed a new integrative strategy to combine biological knowledge and microarray data for gene ranking. The method utilizes knowledge genes for a guidance to first extract coordinative components, and then rank the genes according to their contribution related to a network or pathway. The experimental results show that such a knowledge-guided strategy can provide context-specific gene ranking with an improved performance in pathway member identification. Chen Wang 0001, Jianhua Xuan, Huai Li, Yue Joseph Wang, Ming Zhan, Eric P. Hoffman, Robert Clarke |
BMC Bioinform. | 4 |
| 2010 | Matched Gene Selection and Committee Classifier for Molecular Classification of Heterogeneous Diseases
Guoqiang Yu, Yuanjian Feng, David J. Miller 0001, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Ben Davidson, Ie-Ming Shih, Yue Joseph Wang |
J. Mach. Learn. Res. | 9 |
| 2010 | Nonnegative Least-Correlated Component Analysis for Separation of Dependent Sources by Volume MaximizationabstractAlthough significant efforts have been made in developing nonnegative blind source separation techniques, accurate separation of positive yet dependent sources remains a challenging task. In this paper, a joint correlation function of multiple signals is proposed to reveal and confirm that the observations after nonnegative mixing would have higher joint correlation than the original unknown sources. Accordingly, a new nonnegative least-correlated component analysis (n/LCA) method is proposed to design the unmixing matrix by minimizing the joint correlation function among the estimated nonnegative sources. In addition to a closed-form solution for unmixing two mixtures of two sources, the general algorithm of n/LCA for the multisource case is developed based on an iterative volume maximization (IVM) principle and linear programming. The source identifiability and required conditions are discussed and proven. The proposed n/LCA algorithm, denoted by n/LCA-IVM, is evaluated with both simulation data and real biomedical data to demonstrate its superior performance over several existing benchmark methods. Chong-Yung Chi, Tsung-Han Chan, Yue Joseph Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2009 | Analyzing DNA Copy Number Changes Using Fused Margin RegressionabstractDNA copy number change is an important form of structural variations in human genomes. Detecting copy number changes using DNA array data is a challenging task due to high density genomic loci, low signal to noise ratios, and normal tissue contamination. We propose fused margin regression (FMR) method that combines a variable fusion rule and robust epsilon-insensitive loss criterion to approximate piecewise constant segments of the underlying copy number profile. We tested FMR on both simulation and real CGH and SNP array datasets, and observed competitively improved performance as compared to several widely-adopted existing methods. Yuanjian Feng, Guoqiang Yu, Tian-Li Wang, Ie-Ming Shih, Yue Joseph Wang |
BIBM | 5 |
| 2009 | Accurate Estimation of Genomic Deletions and Normal Cell Contamination by Bayesian Analysis of MixturesabstractCopy number change is an important form of structural variation in human genomes. Somatic copy number alterations can cause the acquisition of oncogenes and loss of tumor suppressor genes in tumorigenesis. Recent development of SNP array technology facilitates studies on copy number changes in a genome-wide scale with high resolution. However, tumor samples often consist of mixed cancer and normal cells. Such tissue heterogeneity poses as a serious hurdle to analyzing copy number changes and could confound subsequent marker identification and diagnostic classification rooted in specific cells. We report here a statistically-principled in silico approach to accurately estimate genomic deletions and normal tissue contamination, and accordingly recover the true copy number profile in cancer cells. We tested the proposed method on three simulation and one real datasets and obtained highly promising results validated by the ground truth and figure of merit. We expect this newly developed method to be a useful tool in routine copy number analysis of heterogeneous tissues. Guoqiang Yu, Bai Zhang, Ie-Ming Shih, Yue Joseph Wang |
BIBM | 5 |
| 2009 | An algorithm for learning maximum entropy probability models of disease risk that efficiently searches and sparingly encodes multilocus genomic interactionsabstractMOTIVATION: In both genome-wide association studies (GWAS) and pathway analysis, the modest sample size relative to the number of genetic markers presents formidable computational, statistical and methodological challenges for accurately identifying markers/interactions and for building phenotype-predictive models. RESULTS: We address these objectives via maximum entropy conditional probability modeling (MECPM), coupled with a novel model structure search. Unlike neural networks and support vector machines (SVMs), MECPM makes explicit and is determined by the interactions that confer phenotype-predictive power. Our method identifies both a marker subset and the multiple k-way interactions between these markers. Additional key aspects are: (i) evaluation of a select subset of up to five-way interactions while retaining relatively low complexity; (ii) flexible single nucleotide polymorphism (SNP) coding (dominant, recessive) within each interaction; (iii) no mathematical interaction form assumed; (iv) model structure and order selection based on the Bayesian Information Criterion, which fairly compares interactions at different orders and automatically sets the experiment-wide significance level; (v) MECPM directly yields a phenotype-predictive model. MECPM was compared with a panel of methods on datasets with up to 1000 SNPs and up to eight embedded penetrance function (i.e. ground-truth) interactions, including a five-way, involving less than 20 SNPs. MECPM achieved improved sensitivity and specificity for detecting both ground-truth markers and interactions, compared with previous methods. AVAILABILITY: http://www.cbil.ece.vt.edu/ResearchOngoingSNP.htm David J. Miller 0001, Guoqiang Yu, Yongmei Liu 0003, Li Chen 0018, Carl D. Langefeld, David M. Herrington, Yue Joseph Wang |
Bioinform. | 8 |
| 2009 | Differential dependency network analysis to identify condition-specific topological changes in biological networksabstractMOTIVATION: Significant efforts have been made to acquire data under different conditions and to construct static networks that can explain various gene regulation mechanisms. However, gene regulatory networks are dynamic and condition-specific; under different conditions, networks exhibit different regulation patterns accompanied by different transcriptional network topologies. Thus, an investigation on the topological changes in transcriptional networks can facilitate the understanding of cell development or provide novel insights into the pathophysiology of certain diseases, and help identify the key genetic players that could serve as biomarkers or drug targets. RESULTS: Here, we report a differential dependency network (DDN) analysis to detect statistically significant topological changes in the transcriptional networks between two biological conditions. We propose a local dependency model to represent the local structures of a network by a set of conditional probabilities. We develop an efficient learning algorithm to learn the local dependency model using the Lasso technique. A permutation test is subsequently performed to estimate the statistical significance of each learned local structure. In testing on a simulation dataset, the proposed algorithm accurately detected all the genes with network topological changes. The method was then applied to the estrogen-dependent T-47D estrogen receptor-positive (ER+) breast cancer cell line datasets and human and mouse embryonic stem cell datasets. In both experiments using real microarray datasets, the proposed method produced biologically meaningful results. We expect DDN to emerge as an important bioinformatics tool in transcriptional network analyses. While we focus specifically on transcriptional networks, the DDN method we introduce here is generally applicable to other biological networks with similar characteristics. AVAILABILITY: The DDN MATLAB toolbox and experiment data are available at http://www.cbil.ece.vt.edu/software.htm. Bai Zhang, Huai Li, Rebecca B. Riggins, Ming Zhan, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
Bioinform. | 9 |
| 2008 | Blind separation of non-negative sources by convex analysis: Effective method using linear programmingabstractWe recently reported a criterion for blind separation of non-negative sources, using a new concept called convex analysis for mixtures of non-negative sources (CAMNS). Under some assumptions that are considered realistic for sparse or high-contrast signals, the criterion is that the true source signals can be perfectly recovered by finding the extreme points of some observation-constructed convex set. In our last work we also developed methods for fulfilling the CAMNS criterion, but only for two to three sources. In this paper we propose a systematic linear programming (LP) based method that is applicable to any number of sources. The proposed method has two advantages. First, its dependence on LP means that the method does not suffer from local minima. Second, the maturity of LP solvers enables efficient implementation of the proposed method in practice. Simulation results are provided to demonstrate the efficacy of the proposed method. Tsung-Han Chan, Wing-Kin Ma, Chong-Yung Chi, Yue Joseph Wang |
ICASSP | 4 |
| 2008 | Network-Constrained Support Vector Machine for ClassificationabstractOne of the major goals in microarray data analysis is to identify biomarkers and build a classification model for future prediction. Many traditional statistical models, based on microarray data alone, often fail in identifying biologically meaningful genes, which should have synergistic effect on determine the clinical outcomes through some interactions rather than work individually. In this paper, we proposed a network-constrained support vector machine (nSVM) for classification by incorporating prior knowledge, which could be protein-protein interactions, protein-gene regulation relationships or pathways information. Specifically, we use Laplacian matrix to represent gene-gene interaction network to regularize the objective function of SVM, which imposes the smoothness of coefficients over the network. The experimental results on simulation and real microarray datasets demonstrate that our method could not only improve classification performance compared to conventional SVM, but more importantly, it could identify significant sub-networks belonging to several pathways which might be related to underlying mechanism associated with clinical outcomes. Li Chen 0018, Jianhua Xuan, Yue Joseph Wang, Rebecca B. Riggins, Robert Clarke |
ICMLA | 3 |
| 2008 | Imaging biomarker analysis of rat mammary fat pads and glandular tissues in MRI imagesabstractIn studying the relationship between risk factors and breast cancer, the growth patterns of fat pads and glandular tissues are considered as important biomarkers. The aim of this study is to measure the growth pattern statistics of rat mammary pads and glandular tissues with magnetic resonance (MR) time sequence images. In this paper, we proposed methods containing sequential steps to extract and analyze imaging biomarkers of rat mammary pad and glandular tissues. Firstly, to accurately segment out pads in MR images with noisy bias filed, we proposed a level set method combining local binary fitting (LBF) and geodesic active contour (GAC). The salient glandular tissue regions within the fat pads are further extracted by a scale-space analysis procedure. Then, the volume data of a single rat at different time points are aligned through profile correlation analysis. Finally, the growth rates are calculated and compared to show the changing patterns of fat pads and glandular tissues within separate groups. The experimental results showed the great utility of this approach in providing accurate measurements for novel risk factors of breast cancer. Yimo Tao, Jianhua Xuan, Matthew T. Freedman, Gloria Chepko, Peter G. Shields, Yue Joseph Wang |
ICPR | 6 |
| 2008 | Sparse Decomposition of Gene Expression Data to Infer Transcriptional Modules Guided by Motif Information
Jianhua Xuan, Li Chen 0018, Rebecca B. Riggins, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
ISBRA | 5 |
| 2008 | Integrative Network Component Analysis for Regulatory Network Reconstruction
Chen Wang 0001, Jianhua Xuan, Li Chen 0018, Po Zhao, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman |
ISBRA | 5 |
| 2008 | Knowledge-guided multi-scale independent component analysis for biomarker identificationabstractBACKGROUND: Many statistical methods have been proposed to identify disease biomarkers from gene expression profiles. However, from gene expression profile data alone, statistical methods often fail to identify biologically meaningful biomarkers related to a specific disease under study. In this paper, we develop a novel strategy, namely knowledge-guided multi-scale independent component analysis (ICA), to first infer regulatory signals and then identify biologically relevant biomarkers from microarray data. RESULTS: Since gene expression levels reflect the joint effect of several underlying biological functions, disease-specific biomarkers may be involved in several distinct biological functions. To identify disease-specific biomarkers that provide unique mechanistic insights, a meta-data "knowledge gene pool" (KGP) is first constructed from multiple data sources to provide important information on the likely functions (such as gene ontology information) and regulatory events (such as promoter responsive elements) associated with potential genes of interest. The gene expression and biological meta data associated with the members of the KGP can then be used to guide subsequent analysis. ICA is then applied to multi-scale gene clusters to reveal regulatory modes reflecting the underlying biological mechanisms. Finally disease-specific biomarkers are extracted by their weighted connectivity scores associated with the extracted regulatory modes. A statistical significance test is used to evaluate the significance of transcription factor enrichment for the extracted gene set based on motif information. We applied the proposed method to yeast cell cycle microarray data and Rsf-1-induced ovarian cancer microarray data. The results show that our knowledge-guided ICA approach can extract biologically meaningful regulatory modes and outperform several baseline methods for biomarker identification. CONCLUSION: We have proposed a novel method, namely knowledge-guided multi-scale ICA, to identify disease-specific biomarkers. The goal is to infer knowledge-relevant regulatory signals and then identify corresponding biomarkers through a multi-scale strategy. The approach has been successfully applied to two expression profiling experiments to demonstrate its improved performance in extracting biologically meaningful and disease-related biomarkers. More importantly, the proposed approach shows promising results to infer novel biomarkers for ovarian cancer and extend current knowledge. Li Chen 0018, Jianhua Xuan, Chen Wang 0001, Ie-Ming Shih, Yue Joseph Wang, Eric P. Hoffman, Robert Clarke |
BMC Bioinform. | 5 |
| 2008 | Motif-directed network component analysis for regulatory network inferenceabstractBACKGROUND: Network Component Analysis (NCA) has shown its effectiveness in discovering regulators and inferring transcription factor activities (TFAs) when both microarray data and ChIP-on-chip data are available. However, a NCA scheme is not applicable to many biological studies due to limited topology information available, such as lack of ChIP-on-chip data. We propose a new approach, motif-directed NCA (mNCA), to integrate motif information and gene expression data to infer regulatory networks. RESULTS: We develop motif-directed NCA (mNCA) to incorporate motif information into NCA for regulatory network inference. While motif information is readily available from knowledge databases, it is a "noisy" source of network topology information consisting of many false positives. To overcome this problem, we develop a stability analysis procedure embedded in mNCA to resolve the inconsistency between motif information and gene expression data, and to enable the identification of stable TFAs. The mNCA approach has been applied to a time course microarray data set of muscle regeneration. The experimental results show that the inferred TFAs are not only numerically stable but also biologically relevant to muscle differentiation process. In particular, several inferred TFAs like those of MyoD, myogenin and YY1 are well supported by biological experiments. CONCLUSION: A novel computational approach, mNCA, has been developed to integrate motif information and gene expression data for regulatory network reconstruction. Specifically, motif analysis is used to obtain initial network topology, and stability analysis is developed and applied with mNCA to extract stable TFAs. Experimental results on muscle regeneration microarray data have demonstrated that mNCA is a practical and reliable computational method for regulatory network inference and pathway discovery. Chen Wang 0001, Jianhua Xuan, Li Chen 0018, Po Zhao, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman |
BMC Bioinform. | 5 |
| 2008 | caBIGTM VISDA: Modeling, visualization, and discovery for cluster analysis of genomic dataabstractBACKGROUND: The main limitations of most existing clustering methods used in genomic data analysis include heuristic or random algorithm initialization, the potential of finding poor local optima, the lack of cluster number detection, an inability to incorporate prior/expert knowledge, black-box and non-adaptive designs, in addition to the curse of dimensionality and the discernment of uninformative, uninteresting cluster structure associated with confounding variables. RESULTS: In an effort to partially address these limitations, we develop the VIsual Statistical Data Analyzer (VISDA) for cluster modeling, visualization, and discovery in genomic data. VISDA performs progressive, coarse-to-fine (divisive) hierarchical clustering and visualization, supported by hierarchical mixture modeling, supervised/unsupervised informative gene selection, supervised/unsupervised data visualization, and user/prior knowledge guidance, to discover hidden clusters within complex, high-dimensional genomic data. The hierarchical visualization and clustering scheme of VISDA uses multiple local visualization subspaces (one at each node of the hierarchy) and consequent subspace data modeling to reveal both global and local cluster structures in a "divide and conquer" scenario. Multiple projection methods, each sensitive to a distinct type of clustering tendency, are used for data visualization, which increases the likelihood that cluster structures of interest are revealed. Initialization of the full dimensional model is based on first learning models with user/prior knowledge guidance on data projected into the low-dimensional visualization spaces. Model order selection for the high dimensional data is accomplished by Bayesian theoretic criteria and user justification applied via the hierarchy of low-dimensional visualization subspaces. Based on its complementary building blocks and flexible functionality, VISDA is generally applicable for gene clustering, sample clustering, and phenotype clustering (wherein phenotype labels for samples are known), albeit with minor algorithm modifications customized to each of these tasks. CONCLUSION: VISDA achieved robust and superior clustering accuracy, compared with several benchmark clustering schemes. The model order selection scheme in VISDA was shown to be effective for high dimensional genomic data clustering. On muscular dystrophy data and muscle regeneration data, VISDA identified biologically relevant co-expressed gene clusters. VISDA also captured the pathological relationships among different phenotypes revealed at the molecular level, through phenotype clustering on muscular dystrophy data and multi-category cancer data. Yitan Zhu, Huai Li, David J. Miller 0001, Zuyi Wang, Jianhua Xuan, Robert Clarke, Eric P. Hoffman, Yue Joseph Wang |
BMC Bioinform. | 8 |
| 2008 | Extensions of transductive learning for distributed ensemble classification and application to biometric authentication
David J. Miller 0001, Siddharth Pal, Yue Joseph Wang |
Neurocomputing | 3 |
| 2007 | Rat Mammary Fat Pad Segmentation and Growth Rate Evaluation in T1 Weighted MR ImagesabstractIn studying the relationship between risk factors and breast cancer, growth patterns of the fat pads and glandular tissues are important features. The goal of this small animal study is to measure the size of mammary pads over the time. To achieve this goal, we propose a hierarchical approach to segmenting out rat body, mammary fat pads and evaluating their development in Tl weighted magnetic resonance (TlW-MR) images. Particularly, we have developed a new approach combining watershed transform and region competition for improved fat pad segmentation. An efficient strategy, termed as competition propagation, is developed to propagate the region competition result from one slice to next slice, resulting in a fast convergence in region competition algorithm otherwise computationally costly. To evaluate the development of the fat pads, the volume data of the scans for a single rat to compare are aligned and the common valid range is acquired through correlation analysis. The method has been applied to 18 volumetric sets of Tl W-MR images acquired from this study. The experimental results showed the great utility of this approach as it can provide accurate measurements to assess novel risk factors for breast cancer. Bin Wang 0064, Jianhua Xuan, Matthew T. Freedman, Peter G. Shields, Yue Joseph Wang |
BIBE | 5 |
| 2007 | A Convex Analysis Based Criterion for Blind Separation of Non-Negative SourcesabstractIn this paper, we apply convex analysis to the problem of blind source separation (BSS) of non-negative signals. Under realistic assumptions applicable to many real-world problems such as multichannel biomedical imaging, we formulate a new BSS criterion that does not require statistical source independence, a fundamental assumption to many existing BSS approaches. The new criterion guarantees perfect separation (in the absence of noise), by constructing a convex set from the observations and then finding the extreme points of the convex set. Some experimental results are provided to demonstrate the efficacy of the proposed method. Tsung-Han Chan, Wing-Kin Ma, Chong-Yung Chi, Yue Joseph Wang |
ICASSP (3) | 4 |
| 2007 | Biomarker Identification by Knowledge-Driven Multi-Level ICA and Motif AnalysisabstractMany statistical methods often fail to identify biologically meaningful biomarkers related to a specific disease under study from expression data alone. In this paper, we develop a novel strategy, namely knowledge-driven multi-level independent component analysis (ICA), to infer regulatory signals and identify biologically relevant biomarkers from microarray data. Specifically, based on multi-level clustering results and partial prior knowledge, we apply ICA to find stable disease specific linear regulatory modes and then extract associated biomarker genes. A statistical test is designed to evaluate the significance of transcription factor enrichment for extracted gene set based on motif information. The experimental results on an Rsf-1 induced microarray data set show that our knowledge-driven method can extract more biologically meaningful biomarkers with significant enrichment of transcription factors related to ovarian cancer compared to other gene selection methods with/without prior knowledge. Li Chen 0018, Chen Wang 0001, Ie-Ming Shih, Tian-Li Wang, Yue Joseph Wang, Robert Clarke, Eric P. Hoffman, Jianhua Xuan |
ICMLA | 6 |
| 2007 | VISDA: an open-source caBIGTM analytical tool for data clustering and beyondabstractSUMMARY: VISDA (Visual Statistical Data Analyzer) is a caBIG analytical tool for cluster modeling, visualization and discovery that has met silver-level compatibility under the caBIG initiative. Being statistically principled and visually interfaced, VISDA exploits both hierarchical statistics modeling and human gift for pattern recognition to allow a progressive yet interactive discovery of hidden clusters within high dimensional and complex biomedical datasets. The distinctive features of VISDA are particularly useful for users across the cancer research and broader research communities to analyze complex biological data. AVAILABILITY: http://gforge.nci.nih.gov/projects/visda/ Jiajing Wang, Huai Li, Yitan Zhu, Malik Yousef, Michael Nebozhyn, Michael M. Showe, Louise C. Showe, Jianhua Xuan, Robert Clarke, Yue Joseph Wang |
Bioinform. | 10 |
| 2006 | ModVis: An information visualization tool for gene module discovery
Justin Molineaux, Jianhua Xuan, Yitan Zhu, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
CAINE | 7 |
| 2006 | Inference of Gene Regulatory Networks from Time Course Gene Expression Data Using Neural Networks and Swarm IntelligenceabstractWe present a novel algorithm that combines a recurrent neural network (RNN) and two swarm intelligence (SI) methods to infer a gene regulatory network (GRN) from time course gene expression data. The algorithm uses ant colony optimization (ACO) to identify the optimal architecture of an RNN, while the weights of the RNN are optimized using particle swarm optimization (PSO). Our goal is to construct an RNN whose response mimics gene expression data generated by time course DNA microarray experiments. We observed promising results in applying the proposed hybrid SI-RNN algorithm to infer networks of interaction from simulated and real-world gene expression data Habtom W. Ressom, Yuji Zhang 0001, Jianhua Xuan, Yue Joseph Wang, Robert Clarke |
CIBCB | 4 |
| 2006 | Blind Source Separation with Pattern Expression NMF
Le Wei, Yue Joseph Wang |
ISNN (1) | 4 |
| 2006 | Optimized multilayer perceptrons for molecular classification and diagnosis using genomic dataabstractMOTIVATION: Multilayer perceptrons (MLP) represent one of the widely used and effective machine learning methods currently applied to diagnostic classification based on high-dimensional genomic data. Since the dimensionalities of the existing genomic data often exceed the available sample sizes by orders of magnitude, the MLP performance may degrade owing to the curse of dimensionality and over-fitting, and may not provide acceptable prediction accuracy. RESULTS: Based on Fisher linear discriminant analysis, we designed and implemented an MLP optimization scheme for a two-layer MLP that effectively optimizes the initialization of MLP parameters and MLP architecture. The optimized MLP consistently demonstrated its ability in easing the curse of dimensionality in large microarray datasets. In comparison with a conventional MLP using random initialization, we obtained significant improvements in major performance measures including Bayes classification accuracy, convergence properties and area under the receiver operating characteristic curve (A(z)). SUPPLEMENTARY INFORMATION: The Supplementary information is available on http://www.cbil.ece.vt.edu/publications.htm Zuyi Wang, Yue Joseph Wang, Jianhua Xuan, Yibin Dong, Marina Bakay, Yuanjian Feng, Robert Clarke, Eric P. Hoffman |
Bioinform. | 2 |
| 2005 | Normalization of Microarray Data by Iterative Nonlinear RegressionabstractNormalization is an important prerequisite for almost all follow-up microarray data analysis steps. Accurate normalization assures a common base for comparative biomedical studies using gene expression profiles across different experiments and phenotypes. In this paper, we present a novel normalization approach - iterative nonlinear regression (INR) method - that exploits concurrent identification of invariantly expressed genes (IEGs) and implementation of nonlinear regression normalization. We demonstrate the principle and performance of the INR approach on two real microarray data sets. As compared to major peer methods (e.g., linear regression method, Loess method and iterative ranking method), INR method shows a superior performance in achieving low expression variance across replicates and excellent fold change preservation. Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang |
BIBE | 4 |
| 2005 | Discontinuity-embedded deformable models for surface reconstruction from range imagesabstractSurface reconstruction is a critical step in three-dimensional image processing and understanding. In this letter, a discontinuity-embedded deformable model has been developed to model surfaces with discontinuities. Governed by the Lagrange motion equation, a finite-element representation of the model can dynamically fit the data in both continuous and discontinuous components, reaching its equilibrium in response to induced forces. Experimental results on synthetic and range images demonstrate a significant improvement in preserving depth discontinuities over conventional approaches. Jianhua Xuan, Yue Joseph Wang, Qinfen Zheng, Tülay Adali |
IEEE Signal Process. Lett. | 2 |
| 2004 | Image fusion based on non-negative matrix factorizationabstractNonnegative Matrix Factorization technique (NMF) has been shown to have various applications to image processing, because of its power of local or part-based representation of objects and/or images. In this paper, we present an image fusion method based on NMF, not by the part-based representation feature of NMF, but by its wholly representation of the images needed to be fused: the images are fused by NMF with the parameter r of the NMF to be set to 1. Our experimental results show that the proposed method is efficient and effective for image fusion compared with many other image fusion methods. Le Wei, Qiguang Miao, Yue Joseph Wang |
ICIP | 4 |
| 2004 | Gene selection in class space for molecular classification of cancer
Yue Joseph Wang, Javed I. Khan, Robert Clarke |
Sci. China Ser. F Inf. Sci. | 2 |
| 2004 | Output-threshold coupled neural network for solving the shortest path problems
Defeng Wang, Meihong Shi, Yue Joseph Wang |
Sci. China Ser. F Inf. Sci. | 4 |
| 2003 | Computational intelligence approach for gene expression data mining and classificationabstractThe exploration of high dimensional gene expression microarray data demands powerful analytical tools. Our data mining software, visual data analyzer (VISDA) for cluster discovery, reveals many distinguishing patterns among gene expression profiles. The model-supported hierarchical data exploration tool has two complementary schemes: discriminatory dimensionality reduction for structure-focused data visualization, and cluster decomposition by probabilistic clustering. Reducing dimensionality generates the visualization of the complete data set at the top level. This data set is then partitioned into subclusters that can consequently be visualized at lower levels and if necessary partitioned again. These approaches produce different visualizations that are compared against known phenotypes from the microarray experiments. For class prediction on cancers using miroarray data, multilayer perceptrons (MLPs) are trained and optimized, whose architecture and parameters are regularized and initialized by weighted Fisher criterion (wFC)-based discriminatory component analysis (DCA). The prediction performance is compared and evaluated via multifold cross-validation. Zuyi Wang, Sun-Yuan Kung, Javed I. Khan, Jianhua Xuan, Yue Joseph Wang |
ICME | 6 |
| 2002 | Information-theoretic matching of two point setsabstractThis paper describes the theoretic roadmap of least relative entropy matching of two point sets. The novel feature is to align two point sets without needing to establish explicit point correspondences. The recovery of transformational geometry is achieved using a mixture of principal axes registrations, whose parameters are estimated by minimizing the relative entropy between the two point distributions and using the expectation-maximization algorithm. We give evidence of the optimality of the method and we then evaluate the algorithm's performance in both rigid and nonrigid image registration cases. Yue Joseph Wang, Kelvin Woods, Maxine A. McClain |
IEEE Trans. Image Process. | 1 |
| 2002 | Iterative normalization of cDNA microarray dataabstractThis paper describes a new approach to normalizing microarray expression data. The novel feature is to unify the tasks of estimating normalization coefficients and identifying control gene set. Unification is realized by constructing a window function over the scatter plot defining the subset of constantly expressed genes and by affecting optimization using an iterative procedure. The structure of window function gates contributions to the control gene set used to estimate normalization coefficients. This window measures the consistency of the matched neighborhoods in the scatter plot and provides a means of rejecting control gene outliers. The recovery of normalizational regression and control gene selection are interleaved and are realized by applying coupled operations to the mean square error function. In this way, the two processes bootstrap one another. We evaluate the technique on real microarray data from breast cancer cell lines and complement the experiment with a data cluster visualization study. Yue Joseph Wang, Jianping Lu, Richard Lee 0002, Zhiping Gu, Robert Clarke |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2002 | A Multiple Circular Paths Convolution Neural Network System for Detection of Mammographic MassesabstractA multiple circular path convolution neural network (MCPCNN) architecture specifically designed for the analysis of tumor and tumor-like structures has been constructed. We first divided each suspected tumor area into sectors and computed the defined mass features for each sector independently. These sector features were used on the input layer and were coordinated by convolution kernels of different sizes that propagated signals to the second layer in the neural network system. The convolution kernels were trained, as required, by presenting the training cases to the neural network. In this study, randomly selected mammograms were processed by a dual morphological enhancement technique. Radiodense areas were isolated and were delineated using a region growing algorithm. The boundary of each region of interest was then divided into 36 sectors using 36 equi-angular dividers radiated from the center of the region. A total of 144 Breast Imaging-Reporting and Data System-based features (i.e., four features per sector for 36 sectors) were computed as input values for the evaluation of this newly invented neural network system. The overall performance was 0.78-0.80 for the areas (Az) under the receiver operating characteristic curves using the conventional feed-forward neural network in the detection of mammographic masses. The performance was markedly improved with Az values ranging from 0.84 to 0.89 using the MCPCNN. This paper does not intend to claim the best mass detection system. Instead it reports a potentially better neural network structure for analyzing a set of the mass features defined by an investigator. Shih-Chung Ben Lo, Huai Li, Yue Joseph Wang, Lisa Kinnard, Matthew T. Freedman |
IEEE Trans. Medical Imaging | 3 |
| 2001 | Magnetic resonance image analysis by information theoretic criteria and stochastic site modelsabstractQuantitative analysis of magnetic resonance (MR) images is a powerful tool for image-guided diagnosis, monitoring, and intervention. The major tasks involve tissue quantification and image segmentation where both the pixel and context images are considered. To extract clinically useful information from images that might be lacking in prior knowledge, we introduce an unsupervised tissue characterization algorithm that is both statistically principled and patient specific. The method uses adaptive standard finite normal mixture and inhomogeneous Markov random field models, whose parameters are estimated using expectation-maximization and relaxation labeling algorithms under information theoretic criteria. We demonstrate the successful applications of the approach with synthetic data sets and then with real MR brain images. Yue Joseph Wang, Tülay Adali, Jianhua Xuan, Zsolt Szabo |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2001 | Computerized Radiographic Mass Detection - Part I: Lesion Site Selection by Morphological Enhancement and Contextual SegmentationabstractThis paper presents a statistical model supported approach for enhanced segmentation and extraction of suspicious mass areas from mammographic images. With an appropriate statistical description of various discriminate characteristics of both true and false candidates from the localized areas, an improved mass detection may be achieved in computer-assisted diagnosis (CAD). In this study, one type of morphological operation is derived to enhance disease patterns of suspected masses by cleaning up unrelated background clutters, and a model-based image segmentation is performed to localize the suspected mass areas using stochastic relaxation labeling scheme. We discuss the importance of model selection when a finite generalized Gaussian mixture is employed, and use the information theoretic criteria to determine the optimal model structure and parameters. Examples are presented to show the effectiveness of the proposed methods on mass lesion enhancement and segmentation when applied to mammographical images. Experimental results demonstrate that the proposed method achieves a very satisfactory performance as a preprocessing procedure for mass detection in CAD. Huai Li, Yue Joseph Wang, K. J. Ray Liu, Shih-Chung Ben Lo, Matthew T. Freedman |
IEEE Trans. Medical Imaging | 2 |
| 2001 | Computerized Radiographic Mass Detection - Part II: Decision Support by Featured Database Visualization and Modular Neural NetworksabstractBased on the enhanced segmentation of suspicious mass areas, further development of computer-assisted mass detection may be decomposed into three distinctive machine learning tasks: 1) construction of the featured knowledge database; 2) mapping of the classified and/or unclassified data points in the database; and 3) development of an intelligent user interface. A decision support system may then be constructed as a complementary machine observer that should enhance the radiologists performance in mass detection. We adopt a mathematical feature extraction procedure to construct the featured knowledge database from all the suspicious mass sites localized by the enhanced segmentation. The optimal mapping of the data points is then obtained by learning the generalized normal mixtures and decision boundaries, where a is developed to carry out both soft and hard clustering. A visual explanation of the decision making is further invented as a decision support, based on an interactive visualization hierarchy through the probabilistic principal component projections of the knowledge database and the localized optimal displays of the retrieved raw data. A prototype system is developed and pilot tested to demonstrate the applicability of this framework to mammographic mass detection. Huai Li, Yue Joseph Wang, K. J. Ray Liu, Shih-Chung Ben Lo, Matthew T. Freedman |
IEEE Trans. Medical Imaging | 2 |
| 2000 | Probabilistic principal component subspaces: a hierarchical finite mixture model for data visualizationabstractVisual exploration has proven to be a powerful tool for multivariate data mining and knowledge discovery. Most visualization algorithms aim to find a projection from the data space down to a visually perceivable rendering space. To reveal all of the interesting aspects of multimodal data sets living in a high-dimensional space, a hierarchical visualization algorithm is introduced which allows the complete data set to be visualized at the top level, with clusters and subclusters of data points visualized at deeper levels. The methods involve hierarchical use of standard finite normal mixtures and probabilistic principal component projections, whose parameters are estimated using the expectation-maximization and principal component neural networks under the information theoretic criteria.We demonstrate the principle of the approach on several multimodal numerical data sets, and we then apply the method to the visual explanation in computer-aided diagnosis for breast cancer detection from digital mammograms. Yue Joseph Wang, Matthew T. Freedman, Sun-Yuan Kung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 1999 | Tactile Mapping of Palpable Abnormalities for Breast Cancer DiagnosisabstractThis paper presents the development of a prototype tactile mapping device (TMD) system comprised mainly of a tactile sensor array probe, a 3D camera and a force/torque sensor, which can provide the means to produce tactile maps of the breast lumps during a breast palpation. Focusing on the key tactile topology features from breast palpation such as spatial location, size and shape of the detected lesion, and the force levels used to demonstrate the palpable abnormalities, these maps can record the results of clinical breast examination with a set of pressure distribution profiles and force sensor measurements due to detected lesion. These maps will serve as an objective documentation of palpable lesions for future comparative examinations. Preliminary results of simulated experiments and limited pre-clinical evaluations of the TMD prototype have pilot-tested our hypothesis and provided solid promising data showing the feasibility of the TMD in real clinical applications. Yue Joseph Wang, Charles C. Nguyen, Rujirutana Srikanchana, Z. Geng, Matthew T. Freedman |
ICRA | 1 |
| 1999 | Hierarchical probabilistic principal component subspaces for data visualizationabstractVisual exploration has proven to be a powerful tool for multivariate data mining. Most visualization algorithms aim to find a projection from the data space down to a visually perceivable rendering space. To reveal all of the interesting aspects of complex data sets existing in a high-dimensional space, a hierarchical visualization algorithm is introduced, which allows the complete data set to be visualized at the top level, with clusters and subclusters of data points visualized at deeper levels. The methods involve multiple use of standard finite normal mixture models and probabilistic principal component projections, whose parameters are estimated using the expectation-maximization and principal component neural networks under the information theoretic criteria. We demonstrate the principle of the approach on two 3D synthetic data sets. Yue Joseph Wang, Matthew T. Freedman, Sun-Yuan Kung |
IJCNN | 1 |
| 1998 | 3-D Model Supported Prostrate Biopsy Simulation and Evaluation
Jianhua Xuan, Yue Joseph Wang, Isabell A. Sesterhenn, Judd W. Moul, Seong Ki Mun |
MICCAI | 2 |
| 1998 | Quantification and segmentation of brain tissues from MR images: a probabilistic neural network approachabstractThis paper presents a probabilistic neural network based technique for unsupervised quantification and segmentation of brain tissues from magnetic resonance images. It is shown that this problem can be solved by distribution learning and relaxation labeling, resulting in an efficient method that may be particularly useful in quantifying and segmenting abnormal brain tissues where the number of tissue types is unknown and the distributions of tissue types heavily overlap. The new technique uses suitable statistical models for both the pixel and context images and formulates the problem in terms of model-histogram fitting and global consistency labeling. The quantification is achieved by probabilistic self-organizing mixtures and the segmentation by a probabilistic constraint relaxation network. The experimental results show the efficient and robust performance of the new algorithm and that it outperforms the conventional classification based approaches. Yue Joseph Wang, Tülay Adali, Sun-Yuan Kung, Zsolt Szabo |
IEEE Trans. Image Process. | 1 |
| 1997 | Stochastic Model and Probabilistic Decision-Based Calssifier for Mass Detection in Digital MammographyabstractWe have developed a combined method utilizing morphological operations, a finite generalized Gaussian mixture (FGGM) modeling, and a contextual Bayesian relaxation labeling technique (CBRL) to enhance and extract suspicious masses. A feature space is constructed based on multiple feature extraction from the regions of interest (ROIs). Finally, a multi-modular probabilistic decision-based classifier is employed to distinguish true masses from non-masses. Huai Li, K. J. Ray Liu, Shih-Chung Ben Lo, Yue Joseph Wang |
ICIP (3) | 4 |
| 1997 | Modeling of Wavelet Coefficients in Medical Image CompressionabstractThe discrete wavelet transform provides a new framework of multiresolution space-frequency representation. Its preliminary applications in medical image compression are promising. An accurate modeling of the spatial and frequency characteristics of the wavelet coefficients is a key to designing efficient and accurate quantization for wavelet-based source coding. In this study, we investigate the modeling of a finite mixture distribution of the wavelet coefficients, within the context of information theory and statistical model identification. Using a finite generalized Gaussian mixture to model the overall distribution of the coefficients, an unsupervised learning procedure is developed to quantify the histogram through a tripled adaptive algorithm including detection of the number of local kernels, approximation of the shape of local kernels, and estimation of model parameter values. Our preliminary experimental results indicate that the unsupervised and adaptive histogram quantification can efficiently and accurately fit to the overall mixture distribution of the coefficients for any given frequency subband with unknown characteristics. Yue Joseph Wang, Huao Li, Jianhua Xuan, Shih-Chung Ben Lo, Seong Ki Mun |
ICIP (1) | 1 |
| 1997 | A Deformable Surface-Spine Model for 3-D Surface RegistrationabstractA finite-element deformable surface-spine model is developed in this paper to register two surfaces by recovering the nonlinear deformation with respect to each other. The deformable surface-spine model is a dynamic model governed by Lagrangian motion equations. A 9 degree-of-freedom (dof) finite-element surface element and a 4-dof spine element are developed to iteratively solve Lagrangian equations for computing the deformation between two surfaces. The method has been applied to registration of computerized surgical prostate models. Experimental results have demonstrated that the new registration method can successfully match complex-structured surfaces by recovering the nonlinear deformation. Jianhua Xuan, Yue Joseph Wang, Tülay Adali, Qinfen Zheng |
ICIP (3) | 2 |
| 1997 | Color-feature-based finger tracking for breast palpation quantificationabstractWe have developed a system using vision-based motion tracking technology to gather quantitative data about the breast palpation process for analysis of breast self-examination (BSE) technique. By tracking the position of the fingers, the system is able to objectively quantify the procedure of the BSE process, thus can improve our knowledge on the effectiveness of BSE so as to optimize the search strategy and assure full coverage for breast cancer detection. By visually displaying all the touched position information to the patient as the BSE is being conducted, the system can provide an interactive feedback to the patient and create a prototype for a computer-based BSE training system. In this paper, we describe the vision-based finger tracking technique used in the system, which uses color information in feature extraction and can track 3D positions of the fingers in real time. Experimental results are shown to confirm the performance and effectiveness of the technique. Several issues on implementation of the technique are also discussed. Jianchao Zeng 0002, Yue Joseph Wang, Matthew T. Freedman, Seong Ki Mun |
ICRA | 2 |
| 1996 | Efficient learning of standard finite normal mixtures for image quantificationabstractThis paper presents an efficient on-line distribution learning procedure of standard finite normal mixtures for image quantification. Based on the standard finite normal mixture (SFNM) model, we formulate image quantification as a distribution learning problem, and derive the probabilistic self-organizing map (PSOM) algorithm by minimizing the relative entropy between the SFNM distribution and the image histogram. We justify our formulation and hence provide a basis for the use of SFNM, in pixel image modeling in terms of large sample properties of the maximum likelihood estimator. We then establish convergence properties of the PSOM which simulates a Bayesian rule network structure with Gaussian activation functions forming soft splits of the data, and thus providing unbiased estimates. It is shown that by incorporating learning rate adaptation in a sequential mode, PSOM achieves fast convergence and has efficient learning capabilities which make it very attractive for many practical image quantification applications; such as unsupervised image segmentation and diagnosis by medical images. Yue Joseph Wang, Tülay Adali |
ICASSP | 1 |
| 1995 | Segmentation of magnetic resonance brain image: integrating region growing and edge detectionabstractThe authors present a method that combines region growing and edge detection for magnetic resonance (MR) brain image segmentation. Starting with a simple region growing algorithm which produces an over segmented image, the authors apply a sophisticated region merging method which is capable of handling complex image structures. Edge information is then integrated to verify and, where necessary, to correct region boundaries. The results show that this method is reliable and efficient for MR brain image segmentation. Jianhua Xuan, Tülay Adali, Yue Joseph Wang |
ICIP (3) | 3 |
| 1994 | Probabilistic Neural Networks for Medical Image QuantificationabstractA probabilistic neural network structure is designed for estimating the parameters of a standard finite normal mixture (SFNM) model in medical image analysis. This neural network employs an unsupervised learning scheme based on the unification of Bayesian and least relative entropy principles, and has Bayes and maximum likelihood neurons which adaptively update the local fuzzy variables in the classification space with the capability of achieving flexible boundary shapes. The optimal network size and hence the number of regions for the SFNM model are determined by various information theoretic criteria, and their performances are compared for images with different stochastic characterizations. A Lloyd-Max quantizer is used to improve the initialization of this self-learning procedure. The performance of this learning technique is tested with both simulated and real medical images, and is shown to be an efficient learning scheme.> Tülay Adali, Yue Joseph Wang |
ICIP (3) | 2 |