Edward R. Dougherty

dblp:90/3909 · also Edward Russell Dougherty · DBLP profile ↗
← Back
182ranked-venue papers
30as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 78 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 64 · 16 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 16 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2024 Multi-fidelity Bayesian Optimization with Multiple Information Sources of Input-dependent Fidelity
abstract
By querying approximate surrogate models of different fidelity as available information sources, Multi-Fidelity Bayesian Optimization (MFBO) aims at optimizing unknown functions that are costly if not infeasible to evaluate. Existing MFBO methods often assume that approximate surrogates have consistently high/low fidelity across the input domain. However, approximate evaluations from the same surrogate can have different fidelity at different input regions due to data availability and model constraints, especially when considering machine learning surrogates. In this work, we investigate MFBO when multi-fidelity approximations have input-dependent fidelity. By explicitly capturing input dependency for multi-fidelity queries in Gaussian Process (GP), our new input-dependent MFBO (iMFBO) with learnable noise models better captures the fidelity of each information source in an intuitive way. We further design a new acquisition function for iMFBO and prove that the queries selected by iMFBO have higher quality than those by naive MFBO methods, with the derived sub-linear regret bound. Experiments on both synthetic and real-world data demonstrate its superior empirical performance.
Mingzhou Fan, Byung-Jun Yoon, Edward R. Dougherty, Nathan M. Urban, Francis J. Alexander, Raymundo Arróyave, Xiaoning Qian
UAI3
2022 Comprehensive analysis of gene expression profiles to radiation exposure reveals molecular signatures of low-dose radiation response
abstract
There are various sources of ionizing radiation exposure, where medical exposure for radiation therapy or diagnosis is the most common human-made source. Understanding how gene expression is modulated after ionizing radiation exposure and investigating the presence of any dose-dependent gene expression patterns have broad implications for health risks from radiotherapy, medical radiation diagnostic procedures, as well as other environmental exposure. In this paper, we perform a comprehensive pathway-based analysis of gene expression profiles in response to low-dose radiation exposure, in order to examine the potential mechanism of gene regulation underlying such responses. To accomplish this goal, we employ a statistical framework to determine whether a specific group of genes belonging to a known pathway display coordinated expression patterns that are modulated in a manner consistent with the radiation level. Findings in our study suggest that there exist complex yet consistent signatures that reflect the molecular response to radiation exposure, which differ between low-dose and high-dose radiation.
Xihaier Luo, Sean McCorkle, Gilchan Park, Vanessa López-Marrero, Shinjae Yoo, Edward R. Dougherty, Xiaoning Qian, Francis J. Alexander, Byung-Jun Yoon
BIBM6
2022 Adaptive Group Testing with Mismatched Models
abstract
Accurate detection of infected individuals is one of the critical steps in stopping any pandemic. When the underlying infection rate of the disease is low, testing people in groups, instead of testing each individual in the population, can be more efficient. In this work, we consider noisy adaptive group testing design with specific test sensitivity and specificity that select the optimal group given previous test results based on pre-selected utility function. As in prior studies on group testing, we model this problem as a sequential Bayesian Optimal Experimental Design (BOED) to adaptively design the groups for each test. We analyze the required number of group tests when using the updated posterior on the infection status and the corresponding Mutual Information (MI) as our utility function for selecting new groups. More importantly, we study how the potential bias on the ground-truth noise of group tests may affect the group testing sample complexity.
Mingzhou Fan, Byung-Jun Yoon, Francis J. Alexander, Edward R. Dougherty, Xiaoning Qian
ICASSP4
2022 Network Classification Based on Reducibility With Respect to the Stability of Canalizing Power of Genes in a Gene Regulatory Network - A Boolean Network Modeling Perspective
abstract
A key objective of studying biological systems is to design therapeutic intervention strategies for beneficially altering cell dynamics. Derivation of control policies is hindered by the high-dimensional state spaces associated with gene regulatory networks. Hence, it is critical to reduce the network complexity and the paper aims to address this issue by focusing on the distribution of the canalizing power (CP) of the genes in the model. Canalizing genes enforce broad corrective actions on cellular processes and play a crucial role in producing optimal reactions to external stimuli. Therefore, it is critical to reduce the network while preserving the canalizing power of genes. We reduce Boolean networks with perturbation by removing genes with the smallest canalizing power consecutively, and evaluate the stability of canalizing power. A systematic empirical study demonstrates that there are two classes of networks, reducible and irreducible with respect to the preservation of canalizing power of the genes. Based on these observations, we introduce the definition of reducible networks and proceed with the problem of selecting the relevant network features that allow for discriminating networks from the two different classes. We demonstrate the efficacy of the selected features on synthetic and real gene regulatory networks.
Ivan Ivanov 0001, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Bayesian Active Learning by Soft Mean Objective Cost of Uncertainty
abstract
To achieve label efficiency for training supervised learning models, pool-based active learning sequentially selects samples from a set of candidates as queries to label by optimizing an acquisition function. One category of existing methods adopts one-step-look-ahead strategies based on acquisition functions tailored with the learning objectives, for example based on the expected loss reduction (ELR) or the mean objective cost of uncertainty (MOCU) proposed recently. These active learning methods are optimal with the maximum classification error reduction when one considers a single query. However, it is well-known that there is no performance guarantee in the long run for these myopic methods. In this paper, we show that these methods are not guaranteed to converge to the optimal classifier of the true model because MOCU is not strictly concave. Moreover, we suggest a strictly concave approximation of MOCU—Soft MOCU—that can be used to define an acquisition function to guide Bayesian active learning with theoretical convergence guarantee. For training Bayesian classifiers with both synthetic and real-world data, our experiments demonstrate the superior performance of active learning by Soft MOCU compared to other existing methods.
Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander, Xiaoning Qian
AISTATS2
2021 Uncertainty-aware Active Learning for Optimal Bayesian Classifier
Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander, Xiaoning Qian
ICLR2
2021 Efficient Active Learning for Gaussian Process Classification by Error Reduction
abstract
Active learning sequentially selects the best instance for labeling by optimizing an acquisition function to enhance data/label efficiency. The selection can be either from a discrete instance set (pool-based scenario) or a continuous instance space (query synthesis scenario). In this work, we study both active learning scenarios for Gaussian Process Classification (GPC). The existing active learning strategies that maximize the Estimated Error Reduction (EER) aim at reducing the classification error after training with the new acquired instance in a one-step-look-ahead manner. The computation of EER-based acquisition functions is typically prohibitive as it requires retraining the GPC with every new query. Moreover, as the EER is not smooth, it can not be combined with gradient-based optimization techniques to efficiently explore the continuous instance space for query synthesis. To overcome these critical limitations, we develop computationally efficient algorithms for EER-based active learning with GPC. We derive the joint predictive distribution of label pairs as a one-dimensional integral, as a result of which the computation of the acquisition function avoids retraining the GPC for each query, remarkably reducing the computational overhead. We also derive the gradient chain rule to efficiently calculate the gradient of the acquisition function, which leads to the first query synthesis active learning algorithm implementing EER-based strategies. Our experiments clearly demonstrate the computational efficiency of the proposed algorithms. We also benchmark our algorithms on both synthetic and real-world datasets, which show superior performance in terms of sampling efficiency compared to the existing state-of-the-art algorithms.
Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander, Xiaoning Qian
NeurIPS2
2021 Optimal Bayesian supervised domain adaptation for RNA sequencing data
abstract
MOTIVATION: When learning to subtype complex disease based on next-generation sequencing data, the amount of available data is often limited. Recent works have tried to leverage data from other domains to design better predictors in the target domain of interest with varying degrees of success. But they are either limited to the cases requiring the outcome label correspondence across domains or cannot leverage the label information at all. Moreover, the existing methods cannot usually benefit from other information available a priori such as gene interaction networks. RESULTS: In this article, we develop a generative optimal Bayesian supervised domain adaptation (OBSDA) model that can integrate RNA sequencing (RNA-Seq) data from different domains along with their labels for improving prediction accuracy in the target domain. Our model can be applied in cases where different domains share the same labels or have different ones. OBSDA is based on a hierarchical Bayesian negative binomial model with parameter factorization, for which the optimal predictor can be derived by marginalization of likelihood over the posterior of the parameters. We first provide an efficient Gibbs sampler for parameter inference in OBSDA. Then, we leverage the gene-gene network prior information and construct an informed and flexible variational family to infer the posterior distributions of model parameters. Comprehensive experiments on real-world RNA-Seq data demonstrate the superior performance of OBSDA, in terms of accuracy in identifying cancer subtypes by utilizing data from different domains. Moreover, we show that by taking advantage of the prior network information we can further improve the performance. AVAILABILITY AND IMPLEMENTATION: The source code for implementations of OBSDA and SI-OBSDA are available at the following link. https://github.com/SHBLK/BSDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shahin Boluki, Xiaoning Qian, Edward R. Dougherty
Bioinform.3
2021 Optimal Bayesian Transfer Learning for Count Data
abstract
There is often a limited amount of omics data to design predictive models in biomedicine. Knowing that these omics data come from underlying processes that may share common pathways and disease mechanisms, it may be beneficial for designing a more accurate and reliable predictor in a target domain of interest, where there is a lack of labeled data to leverage available data in relevant source domains. Here, we focus on developing Bayesian transfer learning methods for analyzing next-generation sequencing (NGS) data to help improve predictions in the target domain. We formulate transfer learning in a fully Bayesian framework and define the relatedness by a joint prior distribution of the model parameters of the source and target domains. Defining joint priors acts as a bridge across domains, through which the related knowledge of source data is transferred to the target domain. We focus on RNA-seq discrete count data, which are often overdispersed. To appropriately model them, we consider the Negative Binomial model and propose an Optimal Bayesian Transfer Learning (OBTL) classifier that minimizes the expected classification error in the target domain. We evaluate the performance of the OBTL classifier via both synthetic and cancer data from The Cancer Genome Atlas (TCGA).
Alireza Karbalayghareh, Xiaoning Qian, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 Quantifying the notions of canalizing and master genes in a gene regulatory network - a Boolean network modeling perspective
abstract
MOTIVATION: Canalizing genes enforce broad corrective actions on cellular processes for the purpose of biological robustness maintaining a constant phenotype to remain unchanged in spite of genetic mutations or environmental perturbations. Despite their central role in biological systems, the observation/detection of canalizing genes is often impeded because the behavior of affected genes is highly varied relative to the inactive canalizer. Therefore, the activity of canalizing genes is difficult to predict to any significant degree by their subject genes under normal cell conditions. RESULTS: We investigate this question and present a quantitative framework that allows for the estimation of the power of canalizing genes in the context of Boolean Networks (BNs) with perturbation. This framework borrows tools from the Pattern Recognition theory and uses the coefficient of determination (CoD) to capture the capacity of the canalizing genes. The canalizing power (CP) of a gene is quantitatively characterized by two terms: regulation power (RP) and incapacitating power (IP). We base this assumption on the idea that canalizing power of a gene should be quantified by the extent of its regulation on the overall network and the extent of control that the gene takes over from other master genes when it is activated, which is equivalent to reduction of the control of other master genes upon its activation. Following this, the CP concept is illustrated with examples in which the goal is to provide preliminary evidence that CP can be used to characterize the ability of canalizing genes. AVAILABILITY AND IMPLEMENTATION: A library of functions written in MATLAB for computing CP is available at http://github.com/eunjikim-angie/CanalizingPower.
Ivan Ivanov 0001, Edward R. Dougherty
Bioinform.3
2019 Optimal clustering with missing values
abstract
BACKGROUND: Missing values frequently arise in modern biomedical studies due to various reasons, including missing tests or complex profiling technologies for different omics measurements. Missing values can complicate the application of clustering algorithms, whose goals are to group points based on some similarity criterion. A common practice for dealing with missing values in the context of clustering is to first impute the missing values, and then apply the clustering algorithm on the completed data. RESULTS: We consider missing values in the context of optimal clustering, which finds an optimal clustering operator with reference to an underlying random labeled point process (RLPP). We show how the missing-value problem fits neatly into the overall framework of optimal clustering by incorporating the missing value mechanism into the random labeled point process and then marginalizing out the missing-value process. In particular, we demonstrate the proposed framework for the Gaussian model with arbitrary covariance structures. Comprehensive experimental studies on both synthetic and real-world RNA-seq data show the superior performance of the proposed optimal clustering with missing values when compared to various clustering approaches. CONCLUSION: Optimal clustering with missing values obviates the need for imputation-based pre-processing of the data, while at the same time possessing smaller clustering errors.
Shahin Boluki, Siamak Zamani Dadaneh, Xiaoning Qian, Edward R. Dougherty
BMC Bioinform.4
2019 Constructing Pathway-Based Priors within a Gaussian Mixture Model for Bayesian Regression and Classification
abstract
Gene-expression-based classification and regression are major concerns in translational genomics. If the feature-label distribution is known, then an optimal classifier can be derived. If the predictor-target distribution is known, then an optimal regression function can be derived. In practice, neither is known, data must be employed, and, for small samples, prior knowledge concerning the feature-label or predictor-target distribution can be used in the learning process. Optimal Bayesian classification and optimal Bayesian regression provide optimality under uncertainty. With optimal Bayesian classification (or regression), uncertainty is treated directly on the feature-label (or predictor-target) distribution. The fundamental engineering problem is prior construction. The Regularized Expected Mean Log-Likelihood Prior (REMLP) utilizes pathway information and provides viable priors for the feature-label distribution, assuming that the training data contain labels. In practice, the labels may not be observed. This paper extends the REMLP methodology to a Gaussian mixture model (GMM) when the labels are unknown. Prior construction bundled with prior update via Bayesian sampling results in Monte Carlo approximations to the optimal Bayesian regression function and optimal Bayesian classifier. Simulations demonstrate that the GMM REMLP prior yields better performance than the EM algorithm for small data sets. We apply it to phenotype classification when the prior knowledge consists of colon cancer pathways.
Shahin Boluki, Mohammad Shahrokh Esfahani, Xiaoning Qian, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 Classification of Single-Cell Gene Expression Trajectories from Incomplete and Noisy Data
abstract
This paper studies classification of gene-expression trajectories coming from two classes, healthy and mutated (cancerous) using Boolean networks with perturbation (BNps) to model the dynamics of each class at the state level. Each class has its own BNp, which is partially known based on gene pathways. We employ a Gaussian model at the observation level to show the expression values of the genes given the hidden binary states at each time point. We use expectation maximization (EM) to learn the BNps and the unknown model parameters, derive closed-form updates for the parameters, and propose a learning algorithm. After learning, a plug-in Bayes classifier is used to classify unlabeled trajectories, which can have missing data. Measuring gene expressions at different times yields trajectories only when measurements come from a single cell. In multiple-cell scenarios, the expression values are averages over many cells with possibly different states. Via the central-limit theorem, we propose another model for expression data in multiple-cell scenarios. Simulations demonstrate that single-cell trajectory data can outperform multiple-cell average expression data relative to classification error, especially in high-noise situations. We also consider data generated via a mammalian cell-cycle network, both the wild-type and with a common mutation affecting p27.
Alireza Karbalayghareh, Ulisses Braga-Neto, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 A fast Branch-and-Bound algorithm for U-curve feature selection
Esmaeil Atashpaz-Gargari, Marcelo da Silva Reis, Ulisses Braga-Neto, Junior Barrera, Edward R. Dougherty
Pattern Recognit.5
2018 Intrinsically Bayesian robust Karhunen-Loève compression
Roozbeh Dehghannasiri, Xiaoning Qian, Edward R. Dougherty
Signal Process.3
2018 Optimal Bayesian Transfer Regression
abstract
Transfer learning studies effective ways to derive better predictors for a system of interest in atargetdomain, where there is lack of data, by utilizing data from other related systems assourcedomain(s). We define a Bayesian transfer learning framework for regression to integrate data between the domains through a joint prior distribution for the source and target parameters. We derive closed-form posteriors of the target parameters integrating both the source and target data, from which closed-form effective joint distributions in the target domain can be derived in terms of generalized hypergeometric functions of matrix argument to define the optimal Bayesian transfer regression (OBTR) operator. We show that the OBTR improves the mean squared error when the source and target domains are related on both synthetic and real-world data.
Alireza Karbalayghareh, Xiaoning Qian, Edward R. Dougherty
IEEE Signal Process. Lett.3
2018 Classification of State Trajectories in Gene Regulatory Networks
abstract
Gene-expression-based phenotype classification is used for disease diagnosis and prognosis relating to treatment strategies. The present paper considers classification based on sequential measurements of multiple genes using gene regulatory network (GRN) modeling. There are two networks, original and mutated, and observations consist of trajectories of network states. The problem is to classify an observation trajectory as coming from either the original or mutated network. GRNs are modeled via probabilistic Boolean networks, which incorporate stochasticity at both the gene and network levels. Mutation affects the regulatory logic. Classification is based upon observing a trajectory of states of some given length. We characterize the Bayes classifier and find the Bayes error for a general PBN and the special case of a single Boolean network affected by random perturbations (BNp). The Bayes error is related to network sensitivity, meaning the extent of alteration in the steady-state distribution of the original network owing to mutation. Using standard methods to calculate steady-state distributions is cumbersome and sometimes impossible, so we provide an efficient algorithm and approximations. Extensive simulations are performed to study the effects of various factors, including approximation accuracy. We apply the classification procedure to a p53 BNp and a mammalian cell cycle PBN.
Alireza Karbalayghareh, Ulisses Braga-Neto, Jianping Hua, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 Detecting Multivariate Gene Interactions in RNA-Seq Data Using Optimal Bayesian Classification
abstract
Differential gene expression testing is an analysis commonly applied to RNA-Seq data. These statistical tests identify genes that are significantly different across phenotypes. We extend this testing paradigm to multivariate gene interactions from a classification perspective with the goal to detect novel gene interactions for the phenotypes of interest. This is achieved through our novel computational framework comprised of a hierarchical statistical model of the RNA-Seq processing pipeline and the corresponding optimal Bayesian classifier. Through Markov Chain Monte Carlo sampling and Monte Carlo integration, we compute quantities where no analytical formulation exists. The performance is then illustrated on an expression dataset from a dietary intervention study where we identify gene pairs that have low classification error yet were not identified as differentially expressed. Additionally, we have released the software package to perform OBC classification on RNA-Seq data under an open source license and is available at http://bit.ly/obc_package.
Jason M. Knight, Ivan Ivanov 0001, Karen Triff, Robert S. Chapkin, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.5
2018 Optimal Objective-Based Experimental Design for Uncertain Dynamical Gene Networks with Experimental Error
abstract
In systems biology, network models are often used to study interactions among cellular components, a salient aim being to develop drugs and therapeutic mechanisms to change the dynamical behavior of the network to avoid undesirable phenotypes. Owing to limited knowledge, model uncertainty is commonplace and network dynamics can be updated in different ways, thereby giving multiple dynamic trajectories, that is, dynamics uncertainty. In this manuscript, we propose an experimental design method that can effectively reduce the dynamics uncertainty and improve performance in an interaction-based network. Both dynamics uncertainty and experimental error are quantified with respect to the modeling objective, herein, therapeutic intervention. The aim of experimental design is to select among a set of candidate experiments the experiment whose outcome, when applied to the network model, maximally reduces the dynamics uncertainty pertinent to the intervention objective.
Daniel N. Mohsenizadeh, Roozbeh Dehghannasiri, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Classification of Gaussian trajectories with missing data in Boolean gene regulatory networks
abstract
This paper studies the classification of gene regulatory networks (GRNs) modeled by probabilistic Boolean networks (PBNs). After observing Gaussian expression values of n genes at m consecutive time points, with consideration of missing data, an algorithm based on expectation maximization (EM) is proposed to estimate the parameters and infer the unknown parts of the networks in the maximum likelihood (ML) sense. Then the estimated values are plugged in to the Bayes classifier, which is optimal, and the performance of the classifier is investigated through various simulations.
Alireza Karbalayghareh, Ulisses Braga-Neto, Edward R. Dougherty
ICASSP3
2017 Incorporating biological prior knowledge for Bayesian learning via maximal knowledge-driven information priors
abstract
BACKGROUND: Phenotypic classification is problematic because small samples are ubiquitous; and, for these, use of prior knowledge is critical. If knowledge concerning the feature-label distribution - for instance, genetic pathways - is available, then it can be used in learning. Optimal Bayesian classification provides optimal classification under model uncertainty. It differs from classical Bayesian methods in which a classification model is assumed and prior distributions are placed on model parameters. With optimal Bayesian classification, uncertainty is treated directly on the feature-label distribution, which assures full utilization of prior knowledge and is guaranteed to outperform classical methods. RESULTS: The salient problem confronting optimal Bayesian classification is prior construction. In this paper, we propose a new prior construction methodology based on a general framework of constraints in the form of conditional probability statements. We call this prior the maximal knowledge-driven information prior (MKDIP). The new constraint framework is more flexible than our previous methods as it naturally handles the potential inconsistency in archived regulatory relationships and conditioning can be augmented by other knowledge, such as population statistics. We also extend the application of prior construction to a multinomial mixture model when labels are unknown, which often occurs in practice. The performance of the proposed methods is examined on two important pathway families, the mammalian cell-cycle and a set of p53-related pathways, and also on a publicly available gene expression dataset of non-small cell lung cancer when combined with the existing prior knowledge on relevant signaling pathways. CONCLUSION: The new proposed general prior construction framework extends the prior construction methodology to a more flexible framework that results in better inference when proper prior knowledge exists. Moreover, the extension of optimal Bayesian classification to multinomial mixtures where data sets are both small and unlabeled, enables superior classifier design using small, unstructured data sets. We have demonstrated the effectiveness of our approach using pathway information and available knowledge of gene regulating functions; however, the underlying theory can be applied to a wide variety of knowledge types, and other applications when there are small samples.
Shahin Boluki, Mohammad Shahrokh Esfahani, Xiaoning Qian, Edward R. Dougherty
BMC Bioinform.4
2017 Optimal experimental design in the context of canonical expansions
abstract
In a wide variety of engineering applications, the mathematical model cannot be fully identified. Therefore, one would like to construct robust operators (filters, classifiers, controllers etc.) that perform optimally relative to incomplete knowledge. Improving model identification through determining unknown parameters can enhance the performance of robust operators. One would like to perform the experiment that provides the most information relative to the engineering objective. The authors present an experimental design framework for parameter estimation in signal processing when the random process model is in the form of canonical expansions. The proposed experimental design is based on the concept of the mean objective cost of uncertainty, which quantifies model uncertainty by taking into account the performance degradation of the designed operator owing to the presence of uncertainty. They provide the general framework for experimental design in the context of canonical expansions and solve it for two major signal processing problems: optimal linear filtering and signal detection.
Roozbeh Dehghannasiri, Xiaoning Qian, Edward R. Dougherty
IET Signal Process.3
2016 Optimal Wiener and homomorphic filtration: Review
Artyom M. Grigoryan, Edward R. Dougherty, Sos S. Agaian
Signal Process.2
2015 Erratum to: Efficient experimental design for uncertainty reduction in gene regulatory networks
abstract
Erratum to: Efficient experimental design for uncertainty reduction in gene regulatory networks https://dx.doi.org/10.1186/1471-2105-16-s13-s2, published online 09 September 2018. During the production of this article [1], errors occurred in equations and algorithms. The Editorial Department of BMC Bioinformatics would like to apologise and inform its readers that an updated version is now available on the BMC Bioinformatics website. Other Information Published in: BMC Bioinformatics License: https://creativecommons.org/licenses/by/4.0/ See article on publisher's website: https://dx.doi.org/10.1186/s12859-015-0839-y
Roozbeh Dehghannasiri, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.3
2015 Efficient experimental design for uncertainty reduction in gene regulatory networks
abstract
BACKGROUND: An accurate understanding of interactions among genes plays a major role in developing therapeutic intervention methods. Gene regulatory networks often contain a significant amount of uncertainty. The process of prioritizing biological experiments to reduce the uncertainty of gene regulatory networks is called experimental design. Under such a strategy, the experiments with high priority are suggested to be conducted first. RESULTS: The authors have already proposed an optimal experimental design method based upon the objective for modeling gene regulatory networks, such as deriving therapeutic interventions. The experimental design method utilizes the concept of mean objective cost of uncertainty (MOCU). MOCU quantifies the expected increase of cost resulting from uncertainty. The optimal experiment to be conducted first is the one which leads to the minimum expected remaining MOCU subsequent to the experiment. In the process, one must find the optimal intervention for every gene regulatory network compatible with the prior knowledge, which can be prohibitively expensive when the size of the network is large. In this paper, we propose a computationally efficient experimental design method. This method incorporates a network reduction scheme by introducing a novel cost function that takes into account the disruption in the ranking of potential experiments. We then estimate the approximate expected remaining MOCU at a lower computational cost using the reduced networks. CONCLUSIONS: Simulation results based on synthetic and real gene regulatory networks show that the proposed approximate method has close performance to that of the optimal method but at lower computational cost. The proposed approximate method also outperforms the random selection policy significantly. A MATLAB software implementing the proposed experimental design method is available at http://gsp.tamu.edu/Publications/supplementary/roozbeh15a/.
Roozbeh Dehghannasiri, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.3
2015 Dynamical modeling of uncertain interaction-based genomic networks
abstract
BACKGROUND: Most dynamical models for genomic networks are built upon two current methodologies, one process-based and the other based on Boolean-type networks. Both are problematic when it comes to experimental design purposes in the laboratory. The first approach requires a comprehensive knowledge of the parameters involved in all biological processes a priori, whereas the results from the second method may not have a biological correspondence and thus cannot be tested in the laboratory. Moreover, the current methods cannot readily utilize existing curated knowledge databases and do not consider uncertainty in the knowledge. Therefore, a new methodology is needed that can generate a dynamical model based on available biological data, assuming uncertainty, while the results from experimental design can be examined in the laboratory. RESULTS: We propose a new methodology for dynamical modeling of genomic networks that can utilize the interaction knowledge provided in public databases. The model assigns discrete states for physical entities, sets priorities among interactions based on information provided in the database, and updates each interaction based on associated node states. Whenever uncertainty in dynamics arises, it explores all possible outcomes. By using the proposed model, biologists can study regulation networks that are too complex for manual analysis. CONCLUSIONS: The proposed approach can be effectively used for constructing dynamical models of interaction-based genomic networks without requiring a complete knowledge of all parameters affecting the network dynamics, and thus based on a small set of available data.
Daniel N. Mohsenizadeh, Jianping Hua, Michael L. Bittner, Edward R. Dougherty
BMC Bioinform.4
2015 Discrete optimal Bayesian classification with error-conditioned sequential sampling
Ariana Broumand, Mohammad Shahrokh Esfahani, Byung-Jun Yoon, Edward R. Dougherty
Pattern Recognit.4
2015 Optimal Experimental Design for Gene Regulatory Networks in the Presence of Uncertainty
abstract
Of major interest to translational genomics is the intervention in gene regulatory networks (GRNs) to affect cell behavior; in particular, to alter pathological phenotypes. Owing to the complexity of GRNs, accurate network inference is practically challenging and GRN models often contain considerable amounts of uncertainty. Considering the cost and time required for conducting biological experiments, it is desirable to have a systematic method for prioritizing potential experiments so that an experiment can be chosen to optimally reduce network uncertainty. Moreover, from a translational perspective it is crucial that GRN uncertainty be quantified and reduced in a manner that pertains to the operational cost that it induces, such as the cost of network intervention. In this work, we utilize the concept of mean objective cost of uncertainty (MOCU) to propose a novel framework for optimal experimental design. In the proposed framework, potential experiments are prioritized based on the MOCU expected to remain after conducting the experiment. Based on this prioritization, one can select an optimal experiment with the largest potential to reduce the pertinent uncertainty present in the current network model. We demonstrate the effectiveness of the proposed method via extensive simulations based on synthetic and real gene regulatory networks.
Roozbeh Dehghannasiri, Byung-Jun Yoon, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 An Optimization-Based Framework for the Transformation of Incomplete Biological Knowledge into a Probabilistic Structure and Its Application to the Utilization of Gene/Protein Signaling Pathways in Discrete Phenotype Classification
abstract
Phenotype classification via genomic data is hampered by small sample sizes that negatively impact classifier design. Utilization of prior biological knowledge in conjunction with training data can improve both classifier design and error estimation via the construction of the optimal Bayesian classifier. In the genomic setting, gene/protein signaling pathways provide a key source of biological knowledge. Although these pathways are neither complete, nor regulatory, with no timing associated with them, they are capable of constraining the set of possible models representing the underlying interaction between molecules. The aim of this paper is to provide a framework and the mathematical tools to transform signaling pathways to prior probabilities governing uncertainty classes of feature-label distributions used in classifier design. Structural motifs extracted from the signaling pathways are mapped to a set of constraints on a prior probability on a Multinomial distribution. Being the conjugate prior for the Multinomial distribution, we propose optimization paradigms to estimate the parameters of a Dirichlet distribution in the Bayesian setting. The performance of the proposed methods is tested on two widely studied pathways: mammalian cell cycle and a p53 pathway model.
Mohammad Shahrokh Esfahani, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 Modeling mass protest adoption in social network communities using geometric brownian motion
abstract
Modeling the movement of information within social media outlets, like Twitter, is key to understanding to how ideas spread but quantifying such movement runs into several difficulties. Two specific areas that elude a clear characterization are (i) the intrinsic random nature of individuals to potentially adopt and subsequently broadcast a Twitter topic, and (ii) the dissemination of information via non-Twitter sources, such as news outlets and word of mouth, and its impact on Twitter propagation. These distinct yet inter-connected areas must be incorporated to generate a comprehensive model of information diffusion. We propose a bispace model to capture propagation in the union of (exclusively) Twitter and non-Twitter environments. To quantify the stochastic nature of Twitter topic propagation, we combine principles of geometric Brownian motion and traditional network graph theory. We apply Poisson process functions to model information diffusion outside of the Twitter mentions network. We discuss techniques to unify the two sub-models to accurately model information dissemination. We demonstrate the novel application of these techniques on real Twitter datasets related to mass protest adoption in social communities.
Fang Jin, Rupinder Paul Khandpur, Nathan Self, Edward R. Dougherty, Sheng Guo 0002, Feng Chen 0001, B. Aditya Prakash, Naren Ramakrishnan
KDD4
2014 Cross-validation under separate sampling: strong bias and how to correct it
abstract
MOTIVATION: It is commonly assumed in pattern recognition that cross-validation error estimation is 'almost unbiased' as long as the number of folds is not too small. While this is true for random sampling, it is not true with separate sampling, where the populations are independently sampled, which is a common situation in bioinformatics. RESULTS: We demonstrate, via analytical and numerical methods, that classical cross-validation can have strong bias under separate sampling, depending on the difference between the sampling ratios and the true population probabilities. We propose a new separate-sampling cross-validation error estimator, and prove that it satisfies an 'almost unbiased' theorem similar to that of random-sampling cross-validation. We present two case studies with previously published data, which show that the results can change drastically if the correct form of cross-validation is used. AVAILABILITY AND IMPLEMENTATION: The source code in C++, along with the Supplementary Materials, is available at: http://gsp.tamu.edu/Publications/supplementary/zollanvari13/.
Ulisses Braga-Neto, Amin Zollanvari, Edward R. Dougherty
Bioinform.3
2014 Effect of separate sampling on classification accuracy
abstract
MOTIVATION: Measurements are commonly taken from two phenotypes to build a classifier, where the number of data points from each class is predetermined, not random. In this 'separate sampling' scenario, the data cannot be used to estimate the class prior probabilities. Moreover, predetermined class sizes can severely degrade classifier performance, even for large samples. RESULTS: We employ simulations using both synthetic and real data to show the detrimental effect of separate sampling on a variety of classification rules. We establish propositions related to the effect on the expected classifier error owing to a sampling ratio different from the population class ratio. From these we derive a sample-based minimax sampling ratio and provide an algorithm for approximating it from the data. We also extend to arbitrary distributions the classical population-based Anderson linear discriminant analysis minimax sampling ratio derived from the discriminant form of the Bayes classifier. AVAILABILITY: All the codes for synthetic data and real data examples are written in MATLAB. A function called mmratio, whose output is an approximation of the minimax sampling ratio of a given dataset, is also written in MATLAB. All the codes are available at: http://gsp.tamu.edu/Publications/supplementary/shahrokh13b.
Mohammad Shahrokh Esfahani, Edward R. Dougherty
Bioinform.2
2014 MCMC implementation of the optimal Bayesian classifier for non-Gaussian models: model-based RNA-Seq classification
abstract
BACKGROUND: Sequencing datasets consist of a finite number of reads which map to specific regions of a reference genome. Most effort in modeling these datasets focuses on the detection of univariate differentially expressed genes. However, for classification, we must consider multiple genes and their interactions. RESULTS: Thus, we introduce a hierarchical multivariate Poisson model (MP) and the associated optimal Bayesian classifier (OBC) for classifying samples using sequencing data. Lacking closed-form solutions, we employ a Monte Carlo Markov Chain (MCMC) approach to perform classification. We demonstrate superior or equivalent classification performance compared to typical classifiers for two synthetic datasets and over a range of classification problem difficulties. We also introduce the Bayesian minimum mean squared error (MMSE) conditional error estimator and demonstrate its computation over the feature space. In addition, we demonstrate superior or leading class performance over an RNA-Seq dataset containing two lung cancer tumor types from The Cancer Genome Atlas (TCGA). CONCLUSIONS: Through model-based, optimal Bayesian classification, we demonstrate superior classification performance for both synthetic and real RNA-Seq datasets. A tutorial video and Python source code is available under an open source license at http://bit.ly/1gimnss .
Jason M. Knight, Ivan Ivanov 0001, Edward R. Dougherty
BMC Bioinform.3
2014 Moments and root-mean-square error of the Bayesian MMSE estimator of classification error in the Gaussian model
Amin Zollanvari, Edward R. Dougherty
Pattern Recognit.2
2014 Incorporation of Biological PathwayKnowledge in the Construction of Priorsfor Optimal Bayesian Classification
abstract
Small samples are commonplace in genomic/proteomic classification, the result being inadequate classifier design and poor error estimation. The problem has recently been addressed by utilizing prior knowledge in the form of a prior distribution on an uncertainty class of feature-label distributions. A critical issue remains: how to incorporate biological knowledge into the prior distribution. For genomics/proteomics, the most common kind of knowledge is in the form of signaling pathways. Thus, it behooves us to find methods of transforming pathway knowledge into knowledge of the feature-label distribution governing the classification problem. In this paper, we address the problem of prior probability construction by proposing a series of optimization paradigms that utilize the incomplete prior information contained in pathways (both topological and regulatory). The optimization paradigms employ the marginal log-likelihood, established using a small number of feature-label realizations (sample points) regularized with the prior pathway information about the variables. In the special case of a Normal-Wishart prior distribution on the mean and inverse covariance matrix (precision matrix) of a Gaussian distribution, these optimization problems become convex. Companion website: gsp.tamu.edu/Publications/supplementary/shahrokh13a.
Mohammad Shahrokh Esfahani, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.2
2013 Performance of linear discriminant analysis in stochastic settings
abstract
This paper provides, for the first time, exact analytical expressions for the first moment of the true error of linear discriminant analysis (LDA) when the data are univariate and taken from two stochastic Gaussian processes. We assume a general setting in which the sample data from each class do not need to be identically distributed or independent within or between classes. As an application of this framework, we characterize the performance of LDA in situations that the data are generated from autoregressive models of the first order.
Amin Zollanvari, Jianping Hua, Edward R. Dougherty
ICASSP3
2013 Intervention in gene regulatory networks with maximal phenotype alteration
abstract
MOTIVATION: A basic issue for translational genomics is to model gene interaction via gene regulatory networks (GRNs) and thereby provide an informatics environment to study the effects of intervention (say, via drugs) and to derive effective intervention strategies. Taking the view that the phenotype is characterized by the long-run behavior (steady-state distribution) of the network, we desire interventions to optimally move the probability mass from undesirable to desirable states Heretofore, two external control approaches have been taken to shift the steady-state mass of a GRN: (i) use a user-defined cost function for which desirable shift of the steady-state mass is a by-product and (ii) use heuristics to design a greedy algorithm. Neither approach provides an optimal control policy relative to long-run behavior. RESULTS: We use a linear programming approach to optimally shift the steady-state mass from undesirable to desirable states, i.e. optimization is directly based on the amount of shift and therefore must outperform previously proposed methods. Moreover, the same basic linear programming structure is used for both unconstrained and constrained optimization, where in the latter case, constraints on the optimization limit the amount of mass that may be shifted to 'ambiguous' states, these being states that are not directly undesirable relative to the pathology of interest but which bear some perceived risk. We apply the method to probabilistic Boolean networks, but the theory applies to any Markovian GRN. AVAILABILITY: Supplementary materials, including the simulation results, MATLAB source code and description of suboptimal methods are available at http://gsp.tamu.edu/Publications/supplementary/yousefi13b. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mohammadmahdi R. Yousefi, Edward R. Dougherty
Bioinform.2
2013 Modeling the Next Generation Sequencing sample processing pipeline for the purposes of classification
abstract
BACKGROUND: A key goal of systems biology and translational genomics is to utilize high-throughput measurements of cellular states to develop expression-based classifiers for discriminating among different phenotypes. Recent developments of Next Generation Sequencing (NGS) technologies can facilitate classifier design by providing expression measurements for tens of thousands of genes simultaneously via the abundance of their mRNA transcripts. Because NGS technologies result in a nonlinear transformation of the actual expression distributions, their application can result in data that are less discriminative than would be the actual expression levels themselves, were they directly observable. RESULTS: Using state-of-the-art distributional modeling for the NGS processing pipeline, this paper studies how that pipeline, via the resulting nonlinear transformation, affects classification and feature selection. The effects of different factors are considered and NGS-based classification is compared to SAGE-based classification and classification directly on the raw expression data, which is represented by a very high-dimensional model previously developed for gene expression. As expected, the nonlinear transformation resulting from NGS processing diminishes classification accuracy; however, owing to a larger number of reads, NGS-based classification outperforms SAGE-based classification. CONCLUSIONS: Having high numbers of reads can mitigate the degradation in classification performance resulting from the effects of NGS technologies. Hence, when performing a RNA-Seq analysis, using the highest possible coverage of the genome is recommended for the purposes of classification.
Noushin Ghaffari, Mohammadmahdi R. Yousefi, Charles D. Johnson, Ivan Ivanov 0001, Edward R. Dougherty
BMC Bioinform.5
2013 Relationship between the accuracy of classifier error estimation and complexity of decision boundary
Esmaeil Atashpaz-Gargari, Chao Sima, Ulisses Braga-Neto, Edward R. Dougherty
Pattern Recognit.4
2013 Optimal classifiers with minimum expected error within a Bayesian framework - Part II: Properties and performance analysis
Lori A. Dalton, Edward R. Dougherty
Pattern Recognit.2
2013 Optimal classifiers with minimum expected error within a Bayesian framework - Part I: Discrete and Gaussian models
Lori A. Dalton, Edward R. Dougherty
Pattern Recognit.2
2013 Classifier design given an uncertainty class of feature distributions via regularized maximum likelihood and the incorporation of biological pathway knowledge in steady-state phenotype classification
Mohammad Shahrokh Esfahani, Jason M. Knight, Amin Zollanvari, Byung-Jun Yoon, Edward R. Dougherty
Pattern Recognit.5
2013 The reliability of estimated confidence intervals for classification error rates when only a single sample is available
Blaise Hanczar, Edward R. Dougherty
Pattern Recognit.2
2013 Analytical study of performance of linear discriminant analysis in stochastic settings
Amin Zollanvari, Jianping Hua, Edward R. Dougherty
Pattern Recognit.3
2012 Structural intervention of gene regulatory networks by general rank-k matrix perturbation
abstract
One of the ultimate objectives of studying gene regulatory networks is to derive potential intervention strategies to avoid aberrant cellular behavior. Boolean networks (BNs) and their stochastic extension, probabilistic Boolean networks (PBNs), provide a convenient framework to design different types of intervention strategies. In this paper, we focus on studying structural intervention, in which we perturb regulatory Boolean functions to alter the long-term network dynamics to obtain desirable behavior. Specifically, we extend our previous work that derives optimal structural intervention for rank-1 function perturbations to more general solutions for arbitrary rank-k function perturbations. The analytic solution is derived using the Sherman-Morrison-Woodbury (SMW) formula. We apply the derived structural intervention to a mutated mammalian cell cycle network. Our results show that our intervention strategy correctly identifies the main targets to stop uncontrolled cell growth in the mutated cell cycle network.
Xiaoning Qian, Byung-Jun Yoon, Edward R. Dougherty
ICASSP3
2012 BPDA2d - a 2D global optimization-based Bayesian peptide detection algorithm for liquid chromatograph-mass spectrometry
abstract
MOTIVATION: Peptide detection is a crucial step in mass spectrometry (MS) based proteomics. Most existing algorithms are based upon greedy isotope template matching and thus may be prone to error propagation and ineffective to detect overlapping peptides. In addition, existing algorithms usually work at different charge states separately, isolating useful information that can be drawn from other charge states, which may lead to poor detection of low abundance peptides. RESULTS: BPDA2d models spectra as a mixture of candidate peptide signals and systematically evaluates all possible combinations of possible peptide candidates to interpret the given spectra. For each candidate, BPDA2d takes into account its elution profile, charge state distribution and isotope pattern, and it combines all evidence to infer the candidate's signal and existence probability. By piecing all evidence together--especially by deriving information across charge states--low abundance peptides can be better identified and peptide detection rates can be improved. Instead of local template matching, BPDA2d performs global optimization for all candidates and systematically optimizes their signals. Since BPDA2d looks for the optimal among all possible interpretations of the given spectra, it has the capability in handling complex spectra where features overlap. BPDA2d estimates the posterior existence probability of detected peptides, which can be directly used for probability-based evaluation in subsequent processing steps. Our experiments indicate that BPDA2d outperforms state-of-the-art detection methods on both simulated data and real liquid chromatography-mass spectrometry data, according to sensitivity and detection accuracy. AVAILABILITY: The BPDA2d software package is available at http://gsp.tamu.edu/Publications/supplementary/sun11a/.
Youting Sun, Jianqiu Zhang 0002, Ulisses Braga-Neto, Edward R. Dougherty
Bioinform.4
2012 Performance reproducibility index for classification
abstract
MOTIVATION: A common practice in biomarker discovery is to decide whether a large laboratory experiment should be carried out based on the results of a preliminary study on a small set of specimens. Consideration of the efficacy of this approach motivates the introduction of a probabilistic measure, for whether a classifier showing promising results in a small-sample preliminary study will perform similarly on a large independent sample. Given the error estimate from the preliminary study, if the probability of reproducible error is low, then there is really no purpose in substantially allocating more resources to a large follow-on study. Indeed, if the probability of the preliminary study providing likely reproducible results is small, then why even perform the preliminary study? RESULTS: This article introduces a reproducibility index for classification, measuring the probability that a sufficiently small error estimate on a small sample will motivate a large follow-on study. We provide a simulation study based on synthetic distribution models that possess known intrinsic classification difficulties and emulate real-world scenarios. We also set up similar simulations on four real datasets to show the consistency of results. The reproducibility indices for different distributional models, real datasets and classification schemes are empirically calculated. The effects of reporting and multiple-rule biases on the reproducibility index are also analyzed. AVAILABILITY: We have implemented in C code the synthetic data distribution model, classification rules, feature selection routine and error estimation methods. The source code is available at http://gsp.tamu.edu/Publications/supplementary/yousefi12a/.
Mohammadmahdi R. Yousefi, Edward R. Dougherty
Bioinform.2
2012 Identifying mechanistic similarities in drug responses
abstract
MOTIVATION: In early drug development, it would be beneficial to be able to identify those dynamic patterns of gene response that indicate that drugs targeting a particular gene will be likely or not to elicit the desired response. One approach would be to quantitate the degree of similarity between the responses that cells show when exposed to drugs, so that consistencies in the regulation of cellular response processes that produce success or failure can be more readily identified. RESULTS: We track drug response using fluorescent proteins as transcription activity reporters. Our basic assumption is that drugs inducing very similar alteration in transcriptional regulation will produce similar temporal trajectories on many of the reporter proteins and hence be identified as having similarities in their mechanisms of action (MOA). The main body of this work is devoted to characterizing similarity in temporal trajectories/signals. To do so, we must first identify the key points that determine mechanistic similarity between two drug responses. Directly comparing points on the two signals is unrealistic, as it cannot handle delays and speed variations on the time axis. Hence, to capture the similarities between reporter responses, we develop an alignment algorithm that is robust to noise, time delays and is able to find all the contiguous parts of signals centered about a core alignment (reflecting a core mechanism in drug response). Applying the proposed algorithm to a range of real drug experiments shows that the result agrees well with the prior drug MOA knowledge. AVAILABILITY: The R code for the RLCSS algorithm is available at http://gsp.tamu.edu/Publications/supplementary/zhao12a.
Jianping Hua, Michael L. Bittner, Ivan Ivanov 0001, Edward R. Dougherty
Bioinform.5
2012 Optimal mean-square-error calibration of classifier error estimators under Bayesian models
Lori A. Dalton, Edward R. Dougherty
Pattern Recognit.2
2012 Exact representation of the second-order moments for resubstitution and leave-one-out error estimation for linear discriminant analysis in the univariate heteroskedastic Gaussian model
Amin Zollanvari, Ulisses Braga-Neto, Edward R. Dougherty
Pattern Recognit.3
2012 Multiscale Denoising of Biological Data: A Comparative Analysis
abstract
Measured microarray genomic and metabolic data are a rich source of information about the biological systems they represent. For example, time-series biological data can be used to construct dynamic genetic regulatory network models, which can be used to design intervention strategies to cure or manage major diseases. Also, copy number data can be used to determine the locations and extent of aberrations in chromosome sequences. Unfortunately, measured biological data are usually contaminated with errors that mask the important features in the data. Therefore, these noisy measurements need to be filtered to enhance their usefulness in practice. Wavelet-based multiscale filtering has been shown to be a powerful denoising tool. In this work, different batch as well as online multiscale filtering techniques are used to denoise biological data contaminated with white or colored noise. The performances of these techniques are demonstrated and compared to those of some conventional low-pass filters using two case studies. The first case study uses simulated dynamic metabolic data, while the second case study uses real copy number data. Simulation results show that significant improvement can be achieved using multiscale filtering over conventional filtering techniques.
Mohamed N. Nounou, Hazem N. Nounou, Nader Meskin, Aniruddha Datta, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.5
2012 Fuzzy Intervention in Biological Phenomena
abstract
An important objective of modeling biological phenomena is to develop therapeutic intervention strategies to move an undesirable state of a diseased network toward a more desirable one. Such transitions can be achieved by the use of drugs to act on some genes/metabolites that affect the undesirable behavior. Due to the fact that biological phenomena are complex processes with nonlinear dynamics that are impossible to perfectly represent with a mathematical model, the need for model-free nonlinear intervention strategies that are capable of guiding the target variables to their desired values often arises. In many applications, fuzzy systems have been found to be very useful for parameter estimation, model development and control design of nonlinear processes. In this paper, a model-free fuzzy intervention strategy (that does not require a mathematical model of the biological phenomenon) is proposed to guide the target variables of biological systems to their desired values. The proposed fuzzy intervention strategy is applied to three different biological models: a glycolytic-glycogenolytic pathway model, a purine metabolism pathway model, and a generic pathway model. The simulation results for all models demonstrate the effectiveness of the proposed scheme.
Hazem N. Nounou, Mohamed N. Nounou, Nader Meskin, Aniruddha Datta, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.5
2012 Intervention in Gene Regulatory Networks via Phenotypically Constrained Control Policies Based on Long-Run Behavior
abstract
A salient purpose for studying gene regulatory networks is to derive intervention strategies to identify potential drug targets and design gene-based therapeutic intervention. Optimal and approximate intervention strategies based on the transition probability matrix of the underlying Markov chain have been studied extensively for probabilistic Boolean networks. While the key goal of control is to reduce the steady-state probability mass of undesirable network states, in practice it is important to limit collateral damage and this constraint should be taken into account when designing intervention strategies with network models. In this paper, we propose two new phenotypically constrained stationary control policies by directly investigating the effects on the network long-run behavior. They are derived to reduce the risk of visiting undesirable states in conjunction with constraints on the shift of undesirable steady-state mass so that only limited collateral damage can be introduced. We have studied the performance of the new constrained control policies together with the previous greedy control policies to randomly generated probabilistic Boolean networks. A preliminary example for intervening in a metastatic melanoma network is also given to show their potential application in designing genetic therapeutics to reduce the risk of entering both aberrant phenotypes and other ambiguous states corresponding to complications or collateral damage. Experiments on both random network ensembles and the melanoma network demonstrate that, in general, the new proposed control policies exhibit the desired performance. As shown by intervening in the melanoma network, these control policies can potentially serve as future practical gene therapeutic intervention strategies.
Xiaoning Qian, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 Validation of gene regulatory networks: scientific and inferential
abstract
Gene regulatory network models are a major area of study in systems and computational biology and the construction of network models is among the most important problems in these disciplines. The critical epistemological issue concerns validation. Validity can be approached from two different perspectives (i) given a hypothesized network model, its scientific validity relates to the ability to make predictions from the model that can be checked against experimental observations; and (ii) the validity of a network inference procedure must be evaluated relative to its ability to infer a network from sample points generated by the network. This article examines both perspectives in the framework of a distance function between two networks. It considers some of the obstacles to validation and provides examples of both validation paradigms.
Edward R. Dougherty
Briefings Bioinform.1
2011 Application of the Bayesian MMSE estimator for classification error to gene expression microarray data
abstract
MOTIVATION: With the development of high-throughput genomic and proteomic technologies, coupled with the inherent difficulties in obtaining large samples, biomedicine faces difficult small-sample classification issues, in particular, error estimation. Most popular error estimation methods are motivated by intuition rather than mathematical inference. A recently proposed error estimator based on Bayesian minimum mean square error estimation places error estimation in an optimal filtering framework. In this work, we examine the application of this error estimator to gene expression microarray data, including the suitability of the Gaussian model with normal-inverse-Wishart priors and how to find prior probabilities. RESULTS: We provide an implementation for non-linear classification, where closed form solutions are not available. We propose a methodology for calibrating normal-inverse-Wishart priors based on discarded microarray data and examine the performance on synthetic high-dimensional data and a real dataset from a breast cancer study. The calibrated Bayesian error estimator has superior root mean square performance, especially with moderate to high expected true errors and small feature sizes. AVAILABILITY: We have implemented in C code the Bayesian error estimator for Gaussian distributions and normal-inverse-Wishart priors for both linear classifiers, with exact closed-form representations, and arbitrary classifiers, where we use a Monte Carlo approximation. Our code for the Bayesian error estimator and a toolbox of related utilities are available at http://gsp.tamu.edu/Publications/supplementary/dalton11a. Several supporting simulations are also included. CONTACT: [email protected]
Lori A. Dalton, Edward R. Dougherty
Bioinform.2
2011 Cancer therapy design based on pathway logic
abstract
MOTIVATION: Cancer encompasses various diseases associated with loss of cell cycle control, leading to uncontrolled cell proliferation and/or reduced apoptosis. Cancer is usually caused by malfunction(s) in the cellular signaling pathways. Malfunctions occur in different ways and at different locations in a pathway. Consequently, therapy design should first identify the location and type of malfunction to arrive at a suitable drug combination. RESULTS: We consider the growth factor (GF) signaling pathways, widely studied in the context of cancer. Interactions between different pathway components are modeled using Boolean logic gates. All possible single malfunctions in the resulting circuit are enumerated and responses of the different malfunctioning circuits to a 'test' input are used to group the malfunctions into classes. Effects of different drugs, targeting different parts of the Boolean circuit, are taken into account in deciding drug efficacy, thereby mapping each malfunction to an appropriate set of drugs.
Ritwik Layek, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.4
2011 High-dimensional bolstered error estimation
abstract
MOTIVATION: In small-sample settings, bolstered error estimation has been shown to perform better than cross-validation and competitively with bootstrap with regard to various criteria. The key issue for bolstering performance is the variance setting for the bolstering kernel. Heretofore, this variance has been determined in a non-parametric manner from the data. Although bolstering based on this variance setting works well for small feature sets, results can deteriorate for high-dimensional feature spaces. RESULTS: This article computes an optimal kernel variance depending on the classification rule, sample size, model and feature space, both the original number and the number remaining after feature selection. A key point is that the optimal variance is robust relative to the model. This allows us to develop a method for selecting a suitable variance to use in real-world applications where the model is not known, but the other factors in determining the optimal kernel are known. AVAILABILITY: Companion website at http://compbio.tgen.org/paper_supp/high_dim_bolstering. CONTACT: [email protected].
Chao Sima, Ulisses Braga-Neto, Edward R. Dougherty
Bioinform.3
2011 Multiple-rule bias in the comparison of classification rules
abstract
MOTIVATION: There is growing discussion in the bioinformatics community concerning overoptimism of reported results. Two approaches contributing to overoptimism in classification are (i) the reporting of results on datasets for which a proposed classification rule performs well and (ii) the comparison of multiple classification rules on a single dataset that purports to show the advantage of a certain rule. RESULTS: This article provides a careful probabilistic analysis of the second issue and the 'multiple-rule bias', resulting from choosing a classification rule having minimum estimated error on the dataset. It quantifies this bias corresponding to estimating the expected true error of the classification rule possessing minimum estimated error and it characterizes the bias from estimating the true comparative advantage of the chosen classification rule relative to the others by the estimated comparative advantage on the dataset. The analysis is applied to both synthetic and real data using a number of classification rules and error estimators. AVAILABILITY: We have implemented in C code the synthetic data distribution model, classification rules, feature selection routines and error estimation methods. The code for multiple-rule analysis is implemented in MATLAB. The source code is available at http://gsp.tamu.edu/Publications/supplementary/yousefi11a/. Supplementary simulation results are also included.
Mohammadmahdi R. Yousefi, Jianping Hua, Edward R. Dougherty
Bioinform.3
2011 Probabilistic reconstruction of the tumor progression process in gene regulatory networks in the presence of uncertainty
abstract
BACKGROUND: Accumulation of gene mutations in cells is known to be responsible for tumor progression, driving it from benign states to malignant states. However, previous studies have shown that the detailed sequence of gene mutations, or the steps in tumor progression, may vary from tumor to tumor, making it difficult to infer the exact path that a given type of tumor may have taken. RESULTS: In this paper, we propose an effective probabilistic algorithm for reconstructing the tumor progression process based on partial knowledge of the underlying gene regulatory network and the steady state distribution of the gene expression values in a given tumor. We take the BNp (Boolean networks with pertubation) framework to model the gene regulatory networks. We assume that the true network is not exactly known but we are given an uncertainty class of networks that contains the true network. This network uncertainty class arises from our partial knowledge of the true network, typically represented as a set of local pathways that are embedded in the global network. Given the SSD of the cancerous network, we aim to simultaneously identify the true normal (healthy) network and the set of gene mutations that drove the network into the cancerous state. This is achieved by analyzing the effect of gene mutation on the SSD of a gene regulatory network. At each step, the proposed algorithm reduces the uncertainty class by keeping only those networks whose SSDs get close enough to the cancerous SSD as a result of additional gene mutation. These steps are repeated until we can find the best candidate for the true network and the most probable path of tumor progression. CONCLUSIONS: Simulation results based on both synthetic networks and networks constructed from actual pathway knowledge show that the proposed algorithm can identify the normal network and the actual path of tumor progression with high probability. The algorithm is also robust to model mismatch and allows us to control the trade-off between efficiency and accuracy.
Mohammad Shahrokh Esfahani, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.3
2011 A CoD-based stationary control policy for intervening in large gene regulatory networks
abstract
BACKGROUND: One of the most important goals of the mathematical modeling of gene regulatory networks is to alter their behavior toward desirable phenotypes. Therapeutic techniques are derived for intervention in terms of stationary control policies. In large networks, it becomes computationally burdensome to derive an optimal control policy. To overcome this problem, greedy intervention approaches based on the concept of the Mean First Passage Time or the steady-state probability mass of the network states were previously proposed. Another possible approach is to use reduction mappings to compress the network and develop control policies on its reduced version. However, such mappings lead to loss of information and require an induction step when designing the control policy for the original network. RESULTS: In this paper, we propose a novel solution, CoD-CP, for designing intervention policies for large Boolean networks. The new method utilizes the Coefficient of Determination (CoD) and the Steady-State Distribution (SSD) of the model. The main advantage of CoD-CP in comparison with the previously proposed methods is that it does not require any compression of the original model, and thus can be directly designed on large networks. The simulation studies on small synthetic networks shows that CoD-CP performs comparable to previously proposed greedy policies that were induced from the compressed versions of the networks. Furthermore, on a large 17-gene gastrointestinal cancer network, CoD-CP outperforms other two available greedy techniques, which is precisely the kind of case for which CoD-CP has been developed. Finally, our experiments show that CoD-CP is robust with respect to the attractor structure of the model. CONCLUSIONS: The newly proposed CoD-CP provides an attractive alternative for intervening large networks where other available greedy methods require size reduction on the network and an extra induction step before designing a control policy.
Noushin Ghaffari, Ivan Ivanov 0001, Xiaoning Qian, Edward R. Dougherty
BMC Bioinform.4
2010 Modeling treatment and drug effects at the molecular level using hybrid system theory
abstract
In this paper, we propose to study the treatment and drug effects at the molecular level using a hybrid system model. Specifically, we propose a generic piecewise linear model to analyze drug effects on the state of the genes in a genetic regulatory network. We intend to answer the following question: given an initial state, would a treatment or drug (control input) drive the target gene to a new desired state that are not reachable without the treatment or drug? assuming that the concentration level of the drug remains constant. In other words, we try to identify whether there is a chance that the treatment or drug will be effective for changing gene expressions at all. We provide detailed analysis for two cases. In the first case, there is only one target gene; while in the second case, there is also another gene interacting with the target gene. The relationships between various parameters (of the genetic regulatory network and the design of the drug) and the convergence and the steady state of the controlled genes are derived analytically and discussed in detail. Simulations are performed using MATLAB/SIMULINK and the results confirmed our analytical findings.
Xiangfang Li, Lijun Qian, Edward R. Dougherty
CIBCB3
2010 A CoD-based reduction algorithm for designing stationary control policies on Boolean networks
abstract
MOTIVATION: Gene regulatory networks serve as models from which to derive therapeutic intervention strategies, in particular, stationary control policies over time that shift the probability mass of the steady state distribution (SSD) away from states associated with undesirable phenotypes. Derivation of control policies is hindered by the high-dimensional state spaces associated with gene regulatory networks. Hence, network reduction is a fundamental issue for intervention. RESULTS: The network model that has been most used for the study of intervention in gene regulatory networks is the probabilistic Boolean network (PBN), which is a collection of constituent Boolean networks (BNs) with perturbation. In this article, we propose an algorithm that reduces a BN with perturbation, designs a control policy on the reduced network and then induces that policy to the original network. The coefficient of determination (CoD) is used to choose a gene for deletion, and a reduction mapping is used to rewire the remaining genes. This CoD-reduction procedure is used to construct a reduced network, then either the previously proposed mean first-passage time (MFPT) or SSD stationary control policy is designed on the reduced network, and these policies are induced to the original network. The efficacy of the overall algorithm is demonstrated on networks of 10 genes or less, where it is possible to compare the steady state shifts of the induced and original policies (because the latter can be derived), and by applying it to a 17-gene gastrointestinal network where it is shown that there is substantial beneficial steady state shift. AVAILABILITY: The code for the algorithms is available at: http://gsp.tamu.edu/Publications/supplementary/ghaffari10a/ Please Contact Noushin Ghaffari at [email protected] for further questions. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Noushin Ghaffari, Ivan Ivanov 0001, Xiaoning Qian, Edward R. Dougherty
Bioinform.4
2010 Small-sample precision of ROC-related estimates
abstract
MOTIVATION: The receiver operator characteristic (ROC) curves are commonly used in biomedical applications to judge the performance of a discriminant across varying decision thresholds. The estimated ROC curve depends on the true positive rate (TPR) and false positive rate (FPR), with the key metric being the area under the curve (AUC). With small samples these rates need to be estimated from the training data, so a natural question arises: How well do the estimates of the AUC, TPR and FPR compare with the true metrics? RESULTS: Through a simulation study using data models and analysis of real microarray data, we show that (i) for small samples the root mean square differences of the estimated and true metrics are considerable; (ii) even for large samples, there is only weak correlation between the true and estimated metrics; and (iii) generally, there is weak regression of the true metric on the estimated metric. For classification rules, we consider linear discriminant analysis, linear support vector machine (SVM) and radial basis function SVM. For error estimation, we consider resubstitution, three kinds of cross-validation and bootstrap. Using resampling, we show the unreliability of some published ROC results. AVAILABILITY: Companion web site at http://compbio.tgen.org/paper_supp/ROC/roc.html CONTACT: [email protected].
Blaise Hanczar, Jianping Hua, Chao Sima, John N. Weinstein, Michael L. Bittner, Edward R. Dougherty
Bioinform.6
2010 State reduction for network intervention in probabilistic Boolean networks
abstract
MOTIVATION: A key goal of studying biological systems is to design therapeutic intervention strategies. Probabilistic Boolean networks (PBNs) constitute a mathematical model which enables modeling, predicting and intervening in their long-run behavior using Markov chain theory. The long-run dynamics of a PBN, as represented by its steady-state distribution (SSD), can guide the design of effective intervention strategies for the modeled systems. A major obstacle for its application is the large state space of the underlying Markov chain, which poses a serious computational challenge. Hence, it is critical to reduce the model complexity of PBNs for practical applications. RESULTS: We propose a strategy to reduce the state space of the underlying Markov chain of a PBN based on a criterion that the reduction least distorts the proportional change of stationary masses for critical states, for instance, the network attractors. In comparison to previous reduction methods, we reduce the state space directly, without deleting genes. We then derive stationary control policies on the reduced network that can be naturally induced back to the original network. Computational experiments study the effects of the reduction on model complexity and the performance of designed control policies which is measured by the shift of stationary mass away from undesirable states, those associated with undesirable phenotypes. We consider randomly generated networks as well as a 17-gene gastrointestinal cancer network, which, if not reduced, has a 2(17) × 2(17) transition probability matrix. Such a dimension is too large for direct application of many previously proposed PBN intervention strategies.
Xiaoning Qian, Noushin Ghaffari, Ivan Ivanov 0001, Edward R. Dougherty
Bioinform.4
2010 Reporting bias when using real data sets to analyze classification performance
abstract
MOTIVATION: It is commonplace for authors to propose a new classification rule, either the operator construction part or feature selection, and demonstrate its performance on real data sets, which often come from high-dimensional studies, such as from gene-expression microarrays, with small samples. Owing to the variability in feature selection and error estimation, individual reported performances are highly imprecise. Hence, if only the best test results are reported, then these will be biased relative to the overall performance of the proposed procedure. RESULTS: This article characterizes reporting bias with several statistics and computes these statistics in a large simulation study using both modeled and real data. The results appear as curves giving the different reporting biases as functions of the number of samples tested when reporting only the best or second best performance. It does this for two classification rules, linear discriminant analysis (LDA) and 3-nearest-neighbor (3NN), and for filter and wrapper feature selection, t-test and sequential forward search. These were chosen on account of their well-studied properties and because they were amenable to the extremely large amount of processing required for the simulations. The results across all the experiments are consistent: there is generally large bias overriding what would be considered a significant performance differential, when reporting the best or second best performing data set. We conclude that there needs to be a database of data sets and that, for those studies depending on real data, results should be reported for all data sets in the database. AVAILABILITY: Companion web site at http://gsp.tamu.edu/Publications/supplementary/yousefi09a/
Mohammadmahdi R. Yousefi, Jianping Hua, Chao Sima, Edward R. Dougherty
Bioinform.4
2010 Identification of diagnostic subnetwork markers for cancer in human protein-protein interaction network
abstract
BACKGROUND: Finding reliable gene markers for accurate disease classification is very challenging due to a number of reasons, including the small sample size of typical clinical data, high noise in gene expression measurements, and the heterogeneity across patients. In fact, gene markers identified in independent studies often do not coincide with each other, suggesting that many of the predicted markers may have no biological significance and may be simply artifacts of the analyzed dataset. To find more reliable and reproducible diagnostic markers, several studies proposed to analyze the gene expression data at the level of groups of functionally related genes, such as pathways. Studies have shown that pathway markers tend to be more robust and yield more accurate classification results. One practical problem of the pathway-based approach is the limited coverage of genes by currently known pathways. As a result, potentially important genes that play critical roles in cancer development may be excluded. To overcome this problem, we propose a novel method for identifying reliable subnetwork markers in a human protein-protein interaction (PPI) network. RESULTS: In this method, we overlay the gene expression data with the PPI network and look for the most discriminative linear paths that consist of discriminative genes that are highly correlated to each other. The overlapping linear paths are then optimally combined into subnetworks that can potentially serve as effective diagnostic markers. We tested our method on two independent large-scale breast cancer datasets and compared the effectiveness and reproducibility of the identified subnetwork markers with gene-based and pathway-based markers. We also compared the proposed method with an existing subnetwork-based method. CONCLUSIONS: The proposed method can efficiently find reliable subnetwork markers that outperform the gene-based and pathway-based markers in terms of discriminative power, reproducibility and classification performance. Subnetwork markers found by our method are highly enriched in common GO terms, and they can more accurately classify breast cancer metastasis compared to markers found by a previous method.
Junjie Su, Byung-Jun Yoon, Edward R. Dougherty
BMC Bioinform.3
2010 BPDA - A Bayesian peptide detection algorithm for mass spectrometry
abstract
BACKGROUND: Mass spectrometry (MS) is an essential analytical tool in proteomics. Many existing algorithms for peptide detection are based on isotope template matching and usually work at different charge states separately, making them ineffective to detect overlapping peptides and low abundance peptides. RESULTS: We present BPDA, a Bayesian approach for peptide detection in data produced by MS instruments with high enough resolution to baseline-resolve isotopic peaks, such as MALDI-TOF and LC-MS. We model the spectra as a mixture of candidate peptide signals, and the model is parameterized by MS physical properties. BPDA is based on a rigorous statistical framework and avoids problems, such as voting and ad-hoc thresholding, generally encountered in algorithms based on template matching. It systematically evaluates all possible combinations of possible peptide candidates to interpret a given spectrum, and iteratively finds the best fitting peptide signal in order to minimize the mean squared error of the inferred spectrum to the observed spectrum. In contrast to previous detection methods, BPDA performs deisotoping and deconvolution of mass spectra simultaneously, which enables better identification of weak peptide signals and produces higher sensitivities and more robust results. Unlike template-matching algorithms, BPDA can handle complex data where features overlap. Our experimental results indicate that BPDA performs well on simulated data and real MS data sets, for various resolutions and signal to noise ratios, and compares very favorably with commonly used commercial and open-source software, such as flexAnalysis, OpenMS, and Decon2LS, according to sensitivity and detection accuracy. CONCLUSION: Unlike previous detection methods, which only employ isotopic distributions and work at each single charge state alone, BPDA takes into account the charge state distribution as well, thus lending information to better identify weak peptide signals and produce more robust results. The proposed approach is based on a rigorous statistical framework, which avoids problems generally encountered in algorithms based on template matching. Our experiments indicate that BPDA performs well on both simulated data and real data, and compares very favorably with commonly used commercial and open-source software. The BPDA software can be downloaded from http://gsp.tamu.edu/Publications/supplementary/sun10a/bpda.
Youting Sun, Jianqiu Zhang 0002, Ulisses Braga-Neto, Edward R. Dougherty
BMC Bioinform.4
2010 Exact correlation between actual and estimated errors in discrete classification
Ulisses Braga-Neto, Edward R. Dougherty
Pattern Recognit. Lett.2
2010 Joint sampling distribution between actual and estimated classification errors for linear discriminant analysis
abstract
Error estimation must be used to find the accuracy of a designed classifier, an issue that is critical in biomarker discovery for disease diagnosis and prognosis in genomics and proteomics. This paper presents, for what is believed to be the first time, the analytical formulation for the joint sampling distribution of the actual and estimated errors of a classification rule. The analysis presented here concerns the linear discriminant analysis (LDA) classification rule and the resubstitution and leave-one-out error estimators, under a general parametric Gaussian assumption. Exact results are provided in the univariate case, and a simple method is suggested to obtain an accurate approximation in the multivariate case. It is also shown how these results can be applied in the computation of condition bounds and the regression of the actual error, given the observed error estimate. In contrast to asymptotic results, the analysis presented here is applicable to finite training data. In particular, it applies in the small-sample settings commonly found in genomics and proteomics applications. Numerical examples, which include parameters estimated from actual microarray data, illustrate the analysis throughout.
Amin Zollanvari, Ulisses Braga-Neto, Edward R. Dougherty
IEEE Trans. Inf. Theory3
2009 Steady-state analysis of genetic regulatory networks modeled by nonlinear ordinary differential equations
abstract
Although Ordinary Differential Equations (ODEs) have been used to model Genetic Regulatory Networks (GRNs) in many previous works, their steady-state behaviors are not well studied. However, a phenotype corresponds to a steady-state gene expression pattern and steady-state analysis of GRNs can provide valuable information on the stability of the GRNs, insights into cellular regulatory mechanisms underlying disease development as well as possible interventions for disease control. In this study, the steady-state behaviors of the nonlinear GRN models are analyzed based on time series data. The steady-state solutions and stability of nonlinear GRNs including polynomial model, sigmoidal model and S-system model are discussed in details.
Haixin Wang 0004, Lijun Qian, Edward R. Dougherty
CIBCB3
2009 Adaptive intervention in probabilistic boolean networks
abstract
MOTIVATION: A basic problem of translational systems biology is to utilize gene regulatory networks as a vehicle to design therapeutic intervention strategies to beneficially alter network and, therefore, cellular dynamics. One strain of research has this problem from the perspective of control theory via the design of optimal Markov chain decision processes, mainly in the framework of probabilistic Boolean networks (PBNs). Full optimization assumes that the network is accurately modeled and, to the extent that model inference is inaccurate, which can be expected for gene regulatory networks owing to the combination of model complexity and a paucity of time-course data, the designed intervention strategy may perform poorly. We desire intervention strategies that do not assume accurate full-model inference. RESULTS: This article demonstrates the feasibility of applying on-line adaptive control to improve intervention performance in genetic regulatory networks modeled by PBNs. It shows via simulations that when the network is modeled by a member of a known family of PBNs, an adaptive design can yield improved performance in terms of the average cost. Two algorithms are presented, one better suited for instantaneously random PBNs and the other better suited for context-sensitive PBNs with low switching probability between the constituent BNs.
Ritwik Layek, Aniruddha Datta, Ranadip Pal, Edward R. Dougherty
Bioinform.4
2009 Analysis and modeling of time-course gene-expression profiles from nanomaterial-exposed primary human epidermal keratinocytes
abstract
BACKGROUND: Nanomaterials are being manufactured on a commercial scale for use in medical, diagnostic, energy, component and communications industries. However, concerns over the safety of engineered nanomaterials have surfaced. Humans can be exposed to nanomaterials in different ways such as inhalation or exposure through the integumentary system. RESULTS: The interactions of engineered nanomaterials with primary human cells was investigated, using a systems biology approach combining gene expression microarray profiling with dynamic experimental parameters. In this experiment, primary human epidermal keratinocytes cells were exposed to several low-micron to nano-scale materials, and gene expression was profiled over both time and dose to compile a comprehensive picture of nanomaterial-cellular interactions. Very few gene-expression studies so far have dealt with both time and dose response simultaneously. Here, we propose different approaches to this kind of analysis. First, we used heat maps and multi-dimensional scaling (MDS) plots to visualize the dose response of nanomaterials over time. Then, in order to find out the most common patterns in gene-expression profiles, we used self-organizing maps (SOM) combined with two different criteria to determine the number of clusters. The consistency of SOM results is discussed in context of the information derived from the MDS plots. Finally, in order to identify the genes that have significantly different responses among different levels of dose of each treatment while accounting for the effect of time at the same time, we used a two-way ANOVA model, in connection with Tukey's additivity test and the Box-Cox transformation. The results are discussed in the context of the cellular responses of engineered nanomaterials. CONCLUSION: The analysis presented here lead to interesting and complementary conclusions about the response across time of human epidermal keratinocytes after exposure to nanomaterials. For example, we observed that gene expression for most treatments become closer to the expression of the baseline cultures as time proceeds. The genes found to be differentially-expressed are involved in a number of cellular processes, including regulation of transcription and translation, protein localization, transport, cell cycle progression, cell migration, cytoskeletal reorganization, signal transduction, and development.
Amin Zollanvari, Mary Jane Cunningham, Ulisses Braga-Neto, Edward R. Dougherty
BMC Bioinform.4
2009 Performance of feature-selection methods in the classification of high-dimension data
Jianping Hua, Waibhav Tembe, Edward R. Dougherty
Pattern Recognit.3
2009 On the sampling distribution of resubstitution and leave-one-out error estimators for linear classifiers
Amin Zollanvari, Ulisses Braga-Neto, Edward R. Dougherty
Pattern Recognit.3
2009 Conditioning-Based Modeling of Contextual Genomic Regulation
abstract
A more complete understanding of the alterations in cellular regulatory and control mechanisms that occur in the various forms of cancer has been one of the central targets of the genomic and proteomic methods that allow surveys of the abundance and/or state of cellular macromolecules. This preference is driven both by the intractability of cancer to generic therapies, assumed to be due to the highly varied molecular etiologies observed in cancer, and by the opportunity to discern and dissect the regulatory and control interactions presented by the highly diverse assortment of perturbations of regulation and control that arise in cancer. Exploiting the opportunities for inference on the regulatory and control connections offered by these revealing system perturbations is fraught with the practical problems that arise from the way biological systems operate. Two classes of regulatory action in biological systems are particularly inimical to inference, convergent regulation, where a variety of regulatory actions result in a common set of control responses (crosstalk), and divergent regulation, where a single regulatory action produces entirely different sets of control responses, depending on cellular context (conditioning). We have constructed a coarse mathematical model of the propagation of regulatory influence in such distributed, context-sensitive regulatory networks that allows a quantitative estimation of the amount of crosstalk and conditioning associated with a candidate regulatory gene taken from a set of genes that have been profiled over a series of samples where the candidate's activity varies.
Edward R. Dougherty, Marcel Brun, Jeffrey M. Trent, Michael L. Bittner
IEEE ACM Trans. Comput. Biol. Bioinform.1
2008 Identifying drosophila cell-cycle regulated genes from irregular microarray data
abstract
Due to experimental constraints, most microarray observations are obtained through irregular sampling. In this paper three popular spectral analyzing schemes, i.e., Lomb-Scargle, Capon and the missing data amplitude and phase estimation (MAPES), are compared in terms of their ability and efficiency to recover the periodically expressed genes. The in silico experiments based on microarray measurements of drosophila melanogaster not only verify half of the published cell-cycle genes, but also corroborate genes that behave periodically in human Hela time series experiments.
Kwadwo Agyepong, Erchin Serpedin, Edward R. Dougherty
ICASSP4
2008 Classification with reject option in gene expression data
abstract
MOTIVATION: The classification methods typically used in bioinformatics classify all examples, even if the classification is ambiguous, for instance, when the example is close to the separating hyperplane in linear classification. For medical applications, it may be better to classify an example only when there is a sufficiently high degree of accuracy, rather than classify all examples with decent accuracy. Moreover, when all examples are classified, the classification rule has no control over the accuracy of the classifier; the algorithm just aims to produce a classifier with the smallest error rate possible. In our approach, we fix the accuracy of the classifier and thereby choose a desired risk of error. RESULTS: Our method consists of defining a rejection region in the feature space. This region contains the examples for which classification is ambiguous. These are rejected by the classifier. The accuracy of the classifier becomes a user-defined parameter of the classification rule. The task of the classification rule is to minimize the rejection region with the constraint that the error rate of the classifier be bounded by the chosen target error. This approach is also used in the feature-selection step. The results computed on both synthetic and real data show that classifier accuracy is significantly improved. AVAILABILITY: Companion Website. http://gsp.tamu.edu/Publications/rejectoption/
Blaise Hanczar, Edward R. Dougherty
Bioinform.2
2008 The peaking phenomenon in the presence of feature-selection
Chao Sima, Edward R. Dougherty
Pattern Recognit. Lett.2
2008 Inferring Connectivity of Genetic Regulatory Networks Using Information-Theoretic Criteria
abstract
Recently, the concept of mutual information has been proposed for inferring the structure of genetic regulatory networks from gene expression profiling. After analyzing the limitations of mutual information in inferring the gene-to-gene interactions, this paper introduces the concept of conditional mutual information and based on it proposes two novel algorithms to infer the connectivity structure of genetic regulatory networks. One of the proposed algorithms exhibits a better accuracy while the other algorithm excels in simplicity and flexibility. By exploiting the mutual information and conditional mutual information, a practical metric is also proposed to assess the likeliness of direct connectivity between genes. This novel metric resolves a common limitation associated with the current inference algorithms, namely the situations where the gene connectivity is established in terms of the dichotomy of being either connected or disconnected. Based on the data sets generated by synthetic networks, the performance of the proposed algorithms is compared favorably relative to existing state-of-the-art schemes. The proposed algorithms are also applied on realistic biological measurements, such as the cutaneous melanoma data set, and biological meaningful results are inferred.
Erchin Serpedin, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.3
2007 Inference of Gene Regulatory Networks using S-System: A Unified Approach
abstract
In this paper, a unified approach to infer gene regulatory networks using the S-system model is proposed. In order to discover the structure of large-scale gene regulatory networks, a simplified S-system model is proposed that enables fast parameter estimation to determine the major gene interactions. If a detailed S-system model is desirable for a subset of genes, a two-step method is proposed where the range of the parameters will be determined first using genetic programming and recursive least square estimation. Then the exact values of the parameters will be calculated using a multi-dimensional optimization algorithm. Both downhill simplex algorithm and modified Powell algorithm are tested for multi-dimensional optimization. Simulation results using both synthetic data and real microarray measurements demonstrate the effectiveness of the proposed methods
Haixin Wang 0004, Lijun Qian, Edward R. Dougherty
CIBCB3
2007 Reconstruction of Genetic Regulatory Networks Based on the Posterior Probabilities of Gene Regulations
abstract
Recent advances in high throughput microarray data have enabled the learning of the structure and operation of gene regulatory networks. This paper proposes a novel approach for reconstruction of gene regulatory networks based on the posterior probabilities of gene regulations. Built within the framework of Bayesian statistics and exploiting efficient computational Monte Carlo techniques, the proposed approach prevents the dichotomy of classifying gene interactions as either being connected or disconnected, and thereby it reduces significantly the inference errors. Simulation results corroborate the superior performance of the proposed approach relative to the existing state-of-the-art algorithms.
Kwadwo Agyepong, Erchin Serpedin, Edward R. Dougherty
ICASSP (1)4
2007 SNiPer-HD: improved genotype calling accuracy by an expectation-maximization algorithm for high-density SNP arrays
abstract
MOTIVATION: The technology to genotype single nucleotide polymorphisms (SNPs) at extremely high densities provides for hypothesis-free genome-wide scans for common polymorphisms associated with complex disease. However, we find that some errors introduced by commonly employed genotyping algorithms may lead to inflation of false associations between markers and phenotype. RESULTS: We have developed a novel SNP genotype calling program, SNiPer-High Density (SNiPer-HD), for highly accurate genotype calling across hundreds of thousands of SNPs. The program employs an expectation-maximization (EM) algorithm with parameters based on a training sample set. The algorithm choice allows for highly accurate genotyping for most SNPs. Also, we introduce a quality control metric for each assayed SNP, such that poor-behaving SNPs can be filtered using a metric correlating to genotype class separation in the calling algorithm. SNiPer-HD is superior to the standard dynamic modeling algorithm and is complementary and non-redundant to other algorithms, such as BRLMM. Implementing multiple algorithms together may provide highly accurate genotyping calls, without inflation of false positives due to systematically miss-called SNPs. A reliable and accurate set of SNP genotypes for increasingly dense panels will eliminate some false association signals and false negative signals, allowing for rapid identification of disease susceptibility loci for complex traits. AVAILABILITY: SNiPer-HD is available at TGen's website: http://www.tgen.org/neurogenomics/data.
Jianping Hua, David W. Craig, Marcel Brun, Jennifer Webster, Victoria Zismann, Waibhav Tembe, Keta Joshipura, Matthew J. Huentelman, Edward R. Dougherty, Dietrich A. Stephan
Bioinform.9
2007 The impact of function perturbations in Boolean networks
abstract
MOTIVATION: A network is said to be robust relative to a certain network characteristic if a small change in network structure does not significantly affect the characteristic. From the perspective of network stability, robustness is desirable; however, from the perspective of intervention to exert influence on network behavior, it is undesirable. For Boolean networks, there are two fundamental types of robustness. One type pertains to perturbing the state of the network and the other to perturbing the rule-based structure. RESULTS: This article explores the impact of function perturbations in Boolean networks from two aspects: (1) analysis: predict the impact on network state transitions and attractors via analytical approaches or identify a perturbation by observing its consequences; (2) synthesis: preserve or modify the network characteristics, especially attractors, by introducing a judicious change to the functions. The results are applied to achieve intervention that structurally alters the network to achieve a more favorable steady-state distribution and to identify the function perturbation that has led to altered observed behavior. The intervention procedure is applied to a WNT5A network to reduce the risk of metastasis in melanoma, and the identification procedure is applied to a Drosophila melanogaster segmentation polarity gene network to identify regulatory function perturbation.
Yufei Xiao, Edward R. Dougherty
Bioinform.2
2007 Model-based evaluation of clustering validation measures
Marcel Brun, Chao Sima, Jianping Hua, James Lowey, Brent Carroll, Edward Suh, Edward R. Dougherty
Pattern Recognit.7
2006 Genetic test bed for feature selection
abstract
MOTIVATION: Given a large set of potential features, such as the set of all gene-expression values from a microarray, it is necessary to find a small subset with which to classify. The task of finding an optimal feature set of a given size is inherently combinatoric because to assure optimality all feature sets of a given size must be checked. Thus, numerous suboptimal feature-selection algorithms have been proposed. There are strong impediments to evaluate feature-selection algorithms using real data when data are limited, a common situation in genetic classification. The difficulty is compound. First, there are no class-conditional distributions from which to draw data points, only a single small labeled sample. Second, there are no test data with which to estimate the feature-set errors, and one must depend on a training-data-based error estimator. Finally, there is no optimal feature set with which to compare the feature sets found by the algorithms. RESULTS: This paper describes a genetic test bed for the evaluation of feature-selection algorithms. It begins with a large biological feature-label dataset that is used as an empirical distribution and, using massively parallel computation, finds the top feature sets of various sizes based on a given sample size and classification rule. The user can draw random samples from the data, apply a proposed algorithm, and evaluate the proficiency of the proposed algorithm via three different measures (code provided). A key feature of the test bed is that, once a dataset is input, a single command creates the entire test bed relative to the dataset. The particular dataset used for the first version of the test bed comes from a microarray-based classification study that analyzes a large number of microarrays, prepared with RNA from breast tumor samples from each of 295 patients. AVAILABILITY: The software and supplementary material are available at http://public.tgen.org/tgen-cb/support/testbed/ CONTACT: [email protected].
Ashish Choudhury, Marcel Brun, Jianping Hua, James Lowey, Edward Suh, Edward R. Dougherty
Bioinform.6
2006 Intervention in a family of Boolean networks
abstract
MOTIVATION: Intervention in a gene regulatory network is used to avoid undesirable states, such as those associated with a disease. Several types of intervention have been studied in the framework of a probabilistic Boolean network (PBN), which is a collection of Boolean networks in which the gene state vector transitions according to the rules of one of the constituent networks and where network choice is governed by a selection distribution. The theory of automatic control has been applied to find optimal strategies for manipulating external control variables that affect the transition probabilities to desirably affect dynamic evolution over a finite time horizon. In this paper we treat a case in which we lack the governing probability structure for Boolean network selection, so we simply have a family of Boolean networks, but where these networks possess a common attractor structure. This corresponds to the situation in which network construction is treated as an ill-posed inverse problem in which there are many Boolean networks created from the data under the constraint that they all possess attractor structures matching the data states, which are assumed to arise from sampling the steady state of the real biological network. RESULTS: Given a family of Boolean networks possessing a common attractor structure composed of singleton attractors, a control algorithm is derived by minimizing a composite finite-horizon cost function that is a weighted average over all the individual networks, the idea being that we desire a control policy that on average suits the networks because these are viewed as equivalent relative to the data. The weighting for each network at any time point is taken to be proportional to the instantaneous estimated probability of that network being the underlying network governing the state transition. The results are applied to a family of Boolean networks derived from gene-expression data collected in a study of metastatic melanoma, the intent being to devise a control strategy that reduces the WNT5A gene's action in affecting biological regulation. AVAILABILITY: The software is available on request. SUPPLEMENTARY INFORMATION: The supplementary Information is available at http://ee.tamu.edu/~edward/tree
Ashish Choudhury, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.4
2006 What should be expected from feature selection in small-sample settings
abstract
MOTIVATION: High-throughput technologies for rapid measurement of vast numbers of biological variables offer the potential for highly discriminatory diagnosis and prognosis; however, high dimensionality together with small samples creates the need for feature selection, while at the same time making feature-selection algorithms less reliable. Feature selection must typically be carried out from among thousands of gene-expression features and in the context of a small sample (small number of microarrays). Two basic questions arise: (1) Can one expect feature selection to yield a feature set whose error is close to that of an optimal feature set? (2) If a good feature set is not found, should it be expected that good feature sets do not exist? RESULTS: The two questions translate quantitatively into questions concerning conditional expectation. (1) Given the error of an optimal feature set, what is the conditionally expected error of the selected feature set? (2) Given the error of the selected feature set, what is the conditionally expected error of the optimal feature set? We address these questions using three classification rules (linear discriminant analysis, linear support vector machine and k-nearest-neighbor classification) and feature selection via sequential floating forward search and the t-test. We consider three feature-label models and patient data from a study concerning survival prognosis for breast cancer. With regard to the two focus questions, there is similarity across all experiments: (1) One cannot expect to find a feature set whose error is close to optimal, and (2) the inability to find a good feature set should not lead to the conclusion that good feature sets do not exist. In practice, the latter conclusion may be more immediately relevant, since when faced with the common occurrence that a feature set discovered from the data does not give satisfactory results, the experimenter can draw no conclusions regarding the existence or nonexistence of suitable feature sets. AVAILABILITY: http://ee.tamu.edu/~edward/feature_regression/
Chao Sima, Edward R. Dougherty
Bioinform.2
2006 Inferring gene regulatory networks from time series data using the minimum description length principle
abstract
MOTIVATION: A central question in reverse engineering of genetic networks consists in determining the dependencies and regulating relationships among genes. This paper addresses the problem of inferring genetic regulatory networks from time-series gene-expression profiles. By adopting a probabilistic modeling framework compatible with the family of models represented by dynamic Bayesian networks and probabilistic Boolean networks, this paper proposes a network inference algorithm to recover not only the direct gene connectivity but also the regulating orientations. RESULTS: Based on the minimum description length principle, a novel network inference algorithm is proposed that greatly shrinks the search space for graphical solutions and achieves a good trade-off between modeling complexity and data fitting. Simulation results show that the algorithm achieves good performance in the case of synthetic networks. Compared with existing state-of-the-art results in the literature, the proposed algorithm exceptionally excels in efficiency, accuracy, robustness and scalability. Given a time-series dataset for Drosophila melanogaster, the paper proposes a genetic regulatory network involved in Drosophila's muscle development. AVAILABILITY: Available from the authors upon request.
Erchin Serpedin, Edward R. Dougherty
Bioinform.3
2006 Noise-injected neural networks show promise for use on small-sample expression data
abstract
BACKGROUND: Overfitting the data is a salient issue for classifier design in small-sample settings. This is why selecting a classifier from a constrained family of classifiers, ones that do not possess the potential to too finely partition the feature space, is typically preferable. But overfitting is not merely a consequence of the classifier family; it is highly dependent on the classification rule used to design a classifier from the sample data. Thus, it is possible to consider families that are rather complex but for which there are classification rules that perform well for small samples. Such classification rules can be advantageous because they facilitate satisfactory classification when the class-conditional distributions are not easily separated and the sample is not large. Here we consider neural networks, from the perspectives of classical design based solely on the sample data and from noise-injection-based design. RESULTS: This paper provides an extensive simulation-based comparative study of noise-injected neural-network design. It considers a number of different feature-label models across various small sample sizes using varying amounts of noise injection. Besides comparing noise-injected neural-network design to classical neural-network design, the paper compares it to a number of other classification rules. Our particular interest is with the use of microarray data for expression-based classification for diagnosis and prognosis. To that end, we consider noise-injected neural-network design as it relates to a study of survivability of breast cancer patients. CONCLUSION: The conclusion is that in many instances noise-injected neural network design is superior to the other tested methods, and in almost all cases it does not perform substantially worse than the best of the other methods. Since the amount of noise injected is consequential, the effect of differing amounts of injected noise must be considered.
Jianping Hua, James Lowey, Zixiang Xiong, Edward R. Dougherty
BMC Bioinform.4
2006 Optimal convex error estimators for classification
Chao Sima, Edward R. Dougherty
Pattern Recognit.2
2005 How many samples are needed to build a classifier: a general sequential approach
abstract
MOTIVATION: The standard paradigm for a classifier design is to obtain a sample of feature-label pairs and then to apply a classification rule to derive a classifier from the sample data. Typically in laboratory situations the sample size is limited by cost, time or availability of sample material. Thus, an investigator may wish to consider a sequential approach in which there is a sufficient number of patients to train a classifier in order to make a sound decision for diagnosis while at the same time keeping the number of patients as small as possible to make the studies affordable. RESULTS: A sequential classification procedure is studied via the martingale central limit theorem. It updates the classification rule at each step and provides stopping criteria to ensure with a certain confidence that at stopping a future subject will have misclassification probability smaller than a predetermined threshold. Simulation studies and applications to microarray data analysis are provided. The procedure possesses several attractive properties: (1) it updates the classification rule sequentially and thus does not rely on distributions of primary measurements from other studies; (2) it assesses the stopping criteria at each sequential step and thus can substantially reduce cost via early stopping; and (3) it is not restricted to any particular classification rule and therefore applies to any parametric or non-parametric method, including feature selection or extraction. AVAILABILITY: R-code for the sequential stopping rule is available at http://stat.tamu.edu/~wfu/microarray/sequential/R-code.html
Wenjiang J. Fu, Edward R. Dougherty, Bani K. Mallick, Raymond J. Carroll
Bioinform.2
2005 Optimal number of features as a function of sample size for various classification rules
abstract
MOTIVATION: Given the joint feature-label distribution, increasing the number of features always results in decreased classification error; however, this is not the case when a classifier is designed via a classification rule from sample data. Typically (but not always), for fixed sample size, the error of a designed classifier decreases and then increases as the number of features grows. The potential downside of using too many features is most critical for small samples, which are commonplace for gene-expression-based classifiers for phenotype discrimination. For fixed sample size and feature-label distribution, the issue is to find an optimal number of features. RESULTS: Since only in rare cases is there a known distribution of the error as a function of the number of features and sample size, this study employs simulation for various feature-label distributions and classification rules, and across a wide range of sample and feature-set sizes. To achieve the desired end, finding the optimal number of features as a function of sample size, it employs massively parallel computation. Seven classifiers are treated: 3-nearest-neighbor, Gaussian kernel, linear support vector machine, polynomial support vector machine, perceptron, regular histogram and linear discriminant analysis. Three Gaussian-based models are considered: linear, nonlinear and bimodal. In addition, real patient data from a large breast-cancer study is considered. To mitigate the combinatorial search for finding optimal feature sets, and to model the situation in which subsets of genes are co-regulated and correlation is internal to these subsets, we assume that the covariance matrix of the features is blocked, with each block corresponding to a group of correlated features. Altogether there are a large number of error surfaces for the many cases. These are provided in full on a companion website, which is meant to serve as resource for those working with small-sample classification. AVAILABILITY: For the companion website, please visit http://public.tgen.org/tamu/ofs/ CONTACT: [email protected].
Jianping Hua, Zixiang Xiong, James Lowey, Edward Suh, Edward R. Dougherty
Bioinform.5
2005 Intervention in context-sensitive probabilistic Boolean networks
abstract
MOTIVATION: Intervention in a gene regulatory network is used to help it avoid undesirable states, such as those associated with a disease. Several types of intervention have been studied in the framework of a probabilistic Boolean network (PBN), which is essentially a finite collection of Boolean networks in which at any discrete time point the gene state vector transitions according to the rules of one of the constituent networks. For an instantaneously random PBN, the governing Boolean network is randomly chosen at each time point. For a context-sensitive PBN, the governing Boolean network remains fixed for an interval of time until a binary random variable determines a switch. The theory of automatic control has been previously applied to find optimal strategies for manipulating external (control) variables that affect the transition probabilities of an instantaneously random PBN to desirably affect its dynamic evolution over a finite time horizon. This paper extends the methods of external control to context-sensitive PBNs. RESULTS: This paper treats intervention via external control variables in context-sensitive PBNs by extending the results for instantaneously random PBNs in several directions. First, and most importantly, whereas an instantaneously random PBN yields a Markov chain whose state space is composed of gene vectors, each state of the Markov chain corresponding to a context-sensitive PBN is composed of a pair, the current gene vector occupied by the network and the current constituent Boolean network. Second, the analysis is applied to PBNs with perturbation, meaning that random gene perturbation is permitted at each instant with some probability. Third, the (mathematical) influence of genes within the network is used to choose the particular gene with which to intervene. Lastly, PBNs are designed from data using a recently proposed inference procedure that takes steady-state considerations into account. The results are applied to a context-sensitive PBN derived from gene-expression data collected in a study of metastatic melanoma, the intent being to devise a control strategy that reduces the WNT5A gene's action in affecting biological regulation, since the available data suggest that disruption of this influence could reduce the chance of a melanoma metastasizing.
Ranadip Pal, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.4
2005 Boolean relationships among genes responsive to ionizing radiation in the NCI 60 ACDS
abstract
MOTIVATION: An early use of gene-expression data coming from microarrays was to discover non-linear multivariate intergene relationships. Pursuing this direction, the motivation for this paper is 2-fold: (1) to discover and elucidate multivariate logical predictive relations among gene expressions in a dataset arising from radiation studies using the NCI 60 Anti-Cancer Drug Screen (ACDS) cell lines; and (2) to demonstrate how these logical relations based on coarse quantization reflect corresponding relations in the continuous data. RESULTS: Using the coefficient of determination, a large number of logical relationships have been discovered among genes in the NCI 60 ACDS cell lines. Moreover, these relationships can be seen directly in the original continuous data, and many are robust relative to the thresholds used to obtain the logical data from the continuous data. A key observation is that a number of intergene relationships appear to be considerably stronger when p53 is functional as compared to when it is not, which is consistent with earlier findings in the literature. AVAILABILITY: The appendix is available at http://gsp.tamu.edu/Publications/supplement.htm CONTACT: [email protected].
Ranadip Pal, Aniruddha Datta, Albert J. Fornace Jr., Michael L. Bittner, Edward R. Dougherty
Bioinform.5
2005 Generating Boolean networks with a prescribed attractor structure
abstract
MOTIVATION: Dynamical modeling of gene regulation via network models constitutes a key problem for genomics. The long-run characteristics of a dynamical system are critical and their determination is a primary aspect of system analysis. In the other direction, system synthesis involves constructing a network possessing a given set of properties. This constitutes the inverse problem. Generally, the inverse problem is ill-posed, meaning there will be many networks, or perhaps none, possessing the desired properties. Relative to long-run behavior, we may wish to construct networks possessing a desirable steady-state distribution. This paper addresses the long-run inverse problem pertaining to Boolean networks (BNs). RESULTS: The long-run behavior of a BN is characterized by its attractors. The rest of the state transition diagram is partitioned into level sets, the j-th level set being composed of all states that transition to one of the attractor states in exactly j transitions. We present two algorithms for the attractor inverse problem. The attractors are specified, and the sizes of the predictor sets and the number of levels are constrained. Algorithm complexity and performance are analyzed. The algorithmic solutions have immediate application. Under the assumption that sampling is from the steady state, a basic criterion for checking the validity of a designed network is that there should be concordance between the attractor states of the model and the data states. This criterion can be used to test a design algorithm: randomly select a set of states to be used as data states; generate a BN possessing the selected states as attractors, perhaps with some added requirements such as constraints on the number of predictors and the level structure; apply the design algorithm; and check the concordance between the attractor states of the designed network and the data states. AVAILABILITY: The software and supplementary material is available at http://gsp.tamu.edu/Publications/BNs/bn.htm
Ranadip Pal, Ivan Ivanov 0001, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.5
2005 Superior feature-set ranking for small samples using bolstered error estimation
abstract
Abstract Motivation: Ranking feature sets is a key issue for classification, for instance, phenotype classification based on gene expression. Since ranking is often based on error estimation, and error estimators suffer to differing degrees of imprecision in small-sample settings, it is important to choose a computationally feasible error estimator that yields good feature-set ranking. Results: This paper examines the feature-ranking performance of several kinds of error estimators: resubstitution, cross-validation, bootstrap and bolstered error estimation. It does so for three classification rules: linear discriminant analysis, three-nearest-neighbor classification and classification trees. Two measures of performance are considered. One counts the number of the truly best feature sets appearing among the best feature sets discovered by the error estimator and the other computes the mean absolute error between the top ranks of the truly best feature sets and their ranks as given by the error estimator. Our results indicate that bolstering is superior to bootstrap, and bootstrap is better than cross-validation, for discovering top-performing feature sets for classification when using small samples. A key issue is that bolstered error estimation is tens of times faster than bootstrap, and faster than cross-validation, and is therefore feasible for feature-set ranking when the number of feature sets is extremely large. Availability: We provide a companion website, which contains the complete set of tables and plots regarding the simulation study, and a compilation of references on feature-set ranking with applications in Genomics. The companion website can be accessed at the URL http://ee.tamu.edu/~edward/bolster_ranking Contact: [email protected]
Chao Sima, Ulisses Braga-Neto, Edward R. Dougherty
Bioinform.3
2005 Exact performance of error estimators for discrete classifiers
Ulisses Braga-Neto, Edward R. Dougherty
Pattern Recognit.2
2005 The fundamental role of pattern recognition for gene-expression/microarray data in bioinformatics
Edward R. Dougherty
Pattern Recognit.1
2005 Optimal robust classifiers
Edward R. Dougherty, Jianping Hua, Zixiang Xiong, Yidong Chen 0002
Pattern Recognit.1
2005 The coefficient of intrinsic dependence (feature selection using el CID)
Tailen Hsing, Li-Yu Liu, Marcel Brun, Edward R. Dougherty
Pattern Recognit.4
2005 Determination of the optimal number of features for quadratic discriminant analysis via the normal approximation to the discriminant distribution
Jianping Hua, Zixiang Xiong, Edward R. Dougherty
Pattern Recognit.3
2005 Impact of error estimation on feature selection
Chao Sima, Sanju Attoor, Ulisses Braga-Neto, James Lowey, Edward Suh, Edward R. Dougherty
Pattern Recognit.6
2005 Feature selection algorithms to find strong genes
Paulo J. S. Silva, Ronaldo Fumio Hashimoto, Seungchan Kim, Junior Barrera, Leônidas de Oliveira Brandão, Edward Suh, Edward R. Dougherty
Pattern Recognit. Lett.7
2005 Steady-state probabilities for attractors in probabilistic Boolean networks
Marcel Brun, Edward R. Dougherty, Ilya Shmulevich
Signal Process.2
2004 Which is better for cDNA-microarray-based classification: ratios or direct intensities
abstract
MOTIVATION: There are two general methods for making gene-expression microarrays: one is to hybridize a single test set of labeled targets to the probe, and measure the background-subtracted intensity at each probe site; the other is to hybridize both a test and a reference set of differentially labeled targets to a single detector array, and measure the ratio of the background-subtracted intensities at each probe site. Which method is better depends on the variability in the cell system and the random factors resulting from the microarray technology. It also depends on the purpose for which the microarray is being used. Classification is a fundamental application and it is the one considered here. RESULTS: This paper describes a model-based simulation paradigm that compares the classification accuracy provided by these methods over a variety of noise types and presents the results of a study modeled on noise typical of cDNA microarray data. The model consists of four parts: (1) the measurement equation for genes in the reference state; (2) the measurement equation for genes in the test state; (3) the ratio and normalization procedure for a dual-channel system; and (4) the intensity and normalization procedure for a single-channel system. In the reference state, the mean intensities are modeled as a shifted exponential distribution, and the intensity for a particular gene is modeled via a normal distribution, Normal(I, alphaI), about its mean intensity I, with alpha being the coefficient of variation of the cell system. In the test state, some genes have their intensities up-regulated by a random factor. The model includes a number of random factors affecting intensity measurement: deposition gain d, labeling gain, and post-image-processing residual noise. The key conclusion resulting from the study is that the coefficient of variation governing the randomness of the intensities and the deposition gain are the most important factors for determining whether a single-channel or dual-channel system provides superior classification, and the decision region in the alpha-d plane is approximately linear.
Sanju Attoor, Edward R. Dougherty, Yidong Chen 0002, Michael L. Bittner, Jeffrey M. Trent
Bioinform.2
2004 Is cross-validation valid for small-sample microarray classification?
abstract
Abstract Motivation: Microarray classification typically possesses two striking attributes: (1) classifier design and error estimation are based on remarkably small samples and (2) cross-validation error estimation is employed in the majority of the papers. Thus, it is necessary to have a quantifiable understanding of the behavior of cross-validation in the context of very small samples. Results: An extensive simulation study has been performed comparing cross-validation, resubstitution and bootstrap estimation for three popular classification rules—linear discriminant analysis, 3-nearest-neighbor and decision trees (CART)—using both synthetic and real breast-cancer patient data. Comparison is via the distribution of differences between the estimated and true errors. Various statistics for the deviation distribution have been computed: mean (for estimator bias), variance (for estimator precision), root-mean square error (for composition of bias and variance) and quartile ranges, including outlier behavior. In general, while cross-validation error estimation is much less biased than resubstitution, it displays excessive variance, which makes individual estimates unreliable for small samples. Bootstrap methods provide improved performance relative to variance, but at a high computational cost and often with increased bias (albeit, much less than with resubstitution). Availability and Supplementary information: A companion web site can be accessed at the URL http://ee.tamu.edu/~edward/cv_paper. The companion web site contains: (1) the complete set of tables and plots regarding the simulation study; (2) additional figures; (3) a compilation of references for microarray classification studies and (4) the source code used, with full documentation and examples.
Ulisses Braga-Neto, Edward R. Dougherty
Bioinform.2
2004 Is cross-validation better than resubstitution for ranking genes?
abstract
MOTIVATION: Ranking gene feature sets is a key issue for both phenotype classification, for instance, tumor classification in a DNA microarray experiment, and prediction in the context of genetic regulatory networks. Two broad methods are available to estimate the error (misclassification rate) of a classifier. Resubstitution fits a single classifier to the data, and applies this classifier in turn to each data observation. Cross-validation (in leave-one-out form) removes each observation in turn, constructs the classifier, and then computes whether this leave-one-out classifier correctly classifies the deleted observation. Resubstitution typically underestimates classifier error, severely so in many cases. Cross-validation has the advantage of producing an effectively unbiased error estimate, but the estimate is highly variable. In many applications it is not the misclassification rate per se that is of interest, but rather the construction of gene sets that have the potential to classify or predict. Hence, one needs to rank feature sets based on their performance. RESULTS: A model-based approach is used to compare the ranking performances of resubstitution and cross-validation for classification based on real-valued feature sets and for prediction in the context of probabilistic Boolean networks (PBNs). For classification, a Gaussian model is considered, along with classification via linear discriminant analysis and the 3-nearest-neighbor classification rule. Prediction is examined in the steady-distribution of a PBN. Three metrics are proposed to compare feature-set ranking based on error estimation with ranking based on the true error, which is known owing to the model-based approach. In all cases, resubstitution is competitive with cross-validation relative to ranking accuracy. This is in addition to the enormous savings in computation time afforded by resubstitution.
Ulisses Braga-Neto, Ronaldo Fumio Hashimoto, Edward R. Dougherty, Danh V. Nguyen, Raymond J. Carroll
Bioinform.3
2004 External control in Markovian genetic regulatory networks: the imperfect information case
abstract
Probabilistic Boolean Networks, which form a subclass of Markovian Genetic Regulatory Networks, have been recently introduced as a rule-based paradigm for modeling gene regulatory networks. In an earlier paper, we introduced external control into Markovian Genetic Regulatory networks. More precisely, given a Markovian genetic regulatory network whose state transition probabilities depend on an external (control) variable, a Dynamic Programming-based procedure was developed by which one could choose the sequence of control actions that minimized a given performance index over a finite number of steps. The control algorithm of that paper, however, could be implemented only when one had perfect knowledge of the states of the Markov Chain. This paper presents a control strategy that can be implemented in the imperfect information case, and makes use of the available measurements which are assumed to be probabilistically related to the states of the underlying Markov Chain.
Aniruddha Datta, Ashish Choudhury, Michael L. Bittner, Edward R. Dougherty
Bioinform.4
2004 Growing genetic regulatory networks from seed genes
abstract
MOTIVATION: A number of models have been proposed for genetic regulatory networks. In principle, a network may contain any number of genes, so long as data are available to make inferences about their relationships. Nevertheless, there are two important reasons why the size of a constructed network should be limited. Computationally and mathematically, it is more feasible to model and simulate a network with a small number of genes. In addition, it is more likely that a small set of genes maintains a specific core regulatory mechanism. RESULTS: Subnetworks are constructed in the context of a directed graph by beginning with a seed consisting of one or more genes believed to participate in a viable subnetwork. Functionalities and regulatory relationships among seed genes may be partially known or they may simply be of interest. Given the seed, we iteratively adjoin new genes in a manner that enhances subnetwork autonomy. The algorithm is applied using both the coefficient of determination and the Boolean-function influence among genes, and it is illustrated using a glioma gene-expression dataset. AVAILABILITY: Software for the seed-growing algorithm will be available at the website for Probabilistic Boolean Networks: http://www2.mdanderson.org/app/ilya/PBN/PBN.htm
Ronaldo Fumio Hashimoto, Seungchan Kim, Ilya Shmulevich, Wei Zhang 0011, Michael L. Bittner, Edward R. Dougherty
Bioinform.6
2004 A Bayesian connectivity-based approach to constructing probabilistic gene regulatory networks
abstract
MOTIVATION: We have hypothesized that the construction of transcriptional regulatory networks using a method that optimizes connectivity would lead to regulation consistent with biological expectations. A key expectation is that the hypothetical networks should produce a few, very strong attractors, highly similar to the original observations, mimicking biological state stability and determinism. Another central expectation is that, since it is expected that the biological control is distributed and mutually reinforcing, interpretation of the observations should lead to a very small number of connection schemes. RESULTS: We propose a fully Bayesian approach to constructing probabilistic gene regulatory networks (PGRNs) that emphasizes network topology. The method computes the possible parent sets of each gene, the corresponding predictors and the associated probabilities based on a nonlinear perceptron model, using a reversible jump Markov chain Monte Carlo (MCMC) technique, and an MCMC method is employed to search the network configurations to find those with the highest Bayesian scores to construct the PGRN. The Bayesian method has been used to construct a PGRN based on the observed behavior of a set of genes whose expression patterns vary across a set of melanoma samples exhibiting two very different phenotypes with respect to cell motility and invasiveness. Key biological features have been faithfully reflected in the model. Its steady-state distribution contains attractors that are either identical or very similar to the states observed in the data, and many of the attractors are singletons, which mimics the biological propensity to stably occupy a given state. Most interestingly, the connectivity rules for the most optimal generated networks constituting the PGRN are remarkably similar, as would be expected for a network operating on a distributed basis, with strong interactions between the components.
Xiaobo Zhou 0001, Xiaodong Wang 0001, Ranadip Pal, Ivan Ivanov 0001, Michael L. Bittner, Edward R. Dougherty
Bioinform.6
2004 Classifier performance as a function of distributional complexity
Sanju Attoor, Edward R. Dougherty
Pattern Recognit.2
2004 Bolstered error estimation
Ulisses Braga-Neto, Edward R. Dougherty
Pattern Recognit.2
2004 A probabilistic theory of clustering
Edward R. Dougherty, Marcel Brun
Pattern Recognit.1
2003 Corrected Small-sample Estimation of the Bayes Error
abstract
MOTIVATION: A major problem of pattern classification is estimation of the Bayes error when only small samples are available. One way to estimate the Bayes error is to design a classifier based on some classification rule applied to sample data, estimate the error of the designed classifier, and then use this estimate as an estimate of the Bayes error. Relative to the Bayes error, the expected error of the designed classifier is biased high, and this bias can be severe with small samples. RESULTS: This paper provides a correction for the bias by subtracting a term derived from the representation of the estimation error. It does so for Boolean classifiers, these being defined on binary features. Although the general theory applies to any Boolean classifier, a model is introduced to reduce the number of parameters. A key point is that the expected correction is conservative. Properties of the corrected estimate are studied via simulation. The correction applies to binary predictors because they are mathematically identical to Boolean classifiers. In this context the correction is adapted to the coefficient of determination, which has been used to measure nonlinear multivariate relations between genes and design genetic regulatory networks. An application using gene-expression data from a microarray experiment is provided on the website http://gspsnap.tamu.edu/smallsample/ (user:'smallsample', password:'smallsample)').
Marcel Brun, David Sabbagh, Seungchan Kim, Edward R. Dougherty
Bioinform.4
2003 Gene selection: a Bayesian variable selection approach
abstract
UNLABELLED: Selection of significant genes via expression patterns is an important problem in microarray experiments. Owing to small sample size and the large number of variables (genes), the selection process can be unstable. This paper proposes a hierarchical Bayesian model for gene (variable) selection. We employ latent variables to specialize the model to a regression setting and uses a Bayesian mixture prior to perform the variable selection. We control the size of the model by assigning a prior distribution over the dimension (number of significant genes) of the model. The posterior distributions of the parameters are not in explicit form and we need to use a combination of truncated sampling and Markov Chain Monte Carlo (MCMC) based computation techniques to simulate the parameters from the posteriors. The Bayesian model is flexible enough to identify significant genes as well as to perform future predictions. The method is applied to cancer classification via cDNA microarrays where the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the method is used to identify a set of significant genes. The method is also applied successfully to the leukemia data. SUPPLEMENTARY INFORMATION: http://stat.tamu.edu/people/faculty/bmallick.html.
Kyeong Eun Lee, Naijun Sha, Edward R. Dougherty, Marina Vannucci, Bani K. Mallick
Bioinform.3
2003 Missing-value estimation using linear and non-linear regression with Bayesian gene selection
abstract
MOTIVATION: Data from microarray experiments are usually in the form of large matrices of expression levels of genes under different experimental conditions. Owing to various reasons, there are frequently missing values. Estimating these missing values is important because they affect downstream analysis, such as clustering, classification and network design. Several methods of missing-value estimation are in use. The problem has two parts: (1) selection of genes for estimation and (2) design of an estimation rule. RESULTS: We propose Bayesian variable selection to obtain genes to be used for estimation, and employ both linear and nonlinear regression for the estimation rule itself. Fast implementation issues for these methods are discussed, including the use of QR decomposition for parameter estimation. The proposed methods are tested on data sets arising from hereditary breast cancer and small round blue-cell tumors. The results compare very favorably with currently used methods based on the normalized root-mean-square error. AVAILABILITY: The appendix is available from http://gspsnap.tamu.edu/gspweb/zxb/missing_zxb/ (user: gspweb; passwd: gsplab).
Xiaobo Zhou 0001, Xiaodong Wang 0001, Edward R. Dougherty
Bioinform.3
2003 Morphological Texture Analysis Using the Texture Evolution Function
abstract
This paper develops a new technique for modeling and classifying a growing texture using its evolution function over time. It encompasses morphological texture classification and parameter estimation with the objective of assessing the state of growth achieved by the texture using only a small sample set to train on, consistent with many real world situations for quality control. It is assumed that the texture model evolves over time according to the way in which its evolution function determines the parameters of its defining random process. This paper considers the random Boolean model for both binary and gray-scale images. A multiple linear regression model is used to estimate the Boolean model parameters as functions of the granulometric moments of the textures. Once the texture-model parameters are estimated, the time of the process can be found via the manner in which the parameters are determined by the dynamic evolutionary model.
J. McKenzie, Stephen Marshall, Alison J. Gray, Edward R. Dougherty
Int. J. Pattern Recognit. Artif. Intell.4
2003 External Control in Markovian Genetic Regulatory Networks
Aniruddha Datta, Ashish Choudhury, Michael L. Bittner, Edward R. Dougherty
Mach. Learn.4
2003 Relation Between Permutation-Test P Values and Classifier Error Estimates
Tailen Hsing, Sanju Attoor, Edward R. Dougherty
Mach. Learn.3
2003 Granulometric parametric estimation for the random Boolean model using optimal linear filters and optimal structuring elements
Yoganand Balagurunathan, Edward R. Dougherty
Pattern Recognit. Lett.2
2003 Design of optimal binary filters under joint multiresolution-envelope constraint
Marcel Brun, Edward R. Dougherty, Roberto Hirata Jr., Junior Barrera
Pattern Recognit. Lett.2
2003 Genomic signal processing
Jaakko Astola, Edward R. Dougherty, Ilya Shmulevich, Ioan Tabus
Signal Process.2
2003 Mappings between probabilistic Boolean networks
Edward R. Dougherty, Ilya Shmulevich
Signal Process.1
2003 Design of multi-mask aperture filters
Alan C. Green, Stephen Marshall, David Greenhalgh, Edward R. Dougherty
Signal Process.4
2003 Efficient selection of feature sets possessing high coefficients of determination based on incremental determinations
Ronaldo Fumio Hashimoto, Edward R. Dougherty, Marcel Brun, Zheng-Zheng Zhou, Michael L. Bittner, Jeffrey M. Trent
Signal Process.2
2003 Construction of genomic networks using mutual-information clustering and reversible-jump Markov-chain-Monte-Carlo predictor design
Xiaobo Zhou 0001, Xiaodong Wang 0001, Edward R. Dougherty
Signal Process.3
2002 Ratio statistics of gene expression levels and applications to microarray data analysis
abstract
MOTIVATION: Expression-based analysis for large families of genes has recently become possible owing to the development of cDNA microarrays, which allow simultaneous measurement of transcript levels for thousands of genes. For each spot on a microarray, signals in two channels must be extracted from their backgrounds. This requires algorithms to extract signals arising from tagged mRNA hybridized to arrayed cDNA locations and algorithms to determine the significance of signal ratios. RESULTS: This paper focuses on estimation of signal ratios from the two channels, and the significance of those ratios. The key issue is the determination of whether a ratio is significantly high or low in order to conclude whether the gene is upregulated or downregulated. The paper builds on an earlier study that involved a hypothesis test based on a ratio statistic under the supposition that the measured fluorescent intensities subsequent to image processing can be assumed to reflect the signal intensities. Here, a refined hypothesis test is considered in which the measured intensities forming the ratio are assumed to be combinations of signal and background. The new method involves a signal-to-noise ratio, and for a high signal-to-noise ratio the new test reduces (with close approximation) to the original test. The effect of low signal-to-noise ratio on the ratio statistics constitutes the main theme of the paper. Finally, and in this vein, a quality metric is formulated for spots. This measure can be used to decide whether or not a spot ratio should be deleted, or to adjust various measurements to reflect confidence in the quality of the measurement. CONTACT: [email protected]
Yidong Chen 0002, Vishnu G. Kamat, Edward R. Dougherty, Michael L. Bittner, Paul S. Meltzer, Jeffrey M. Trent
Bioinform.3
2002 Gene perturbation and intervention in probabilistic Boolean networks
abstract
MOTIVATION: A major objective of gene regulatory network modeling, in addition to gaining a deeper understanding of genetic regulation and control, is the development of computational tools for the identification and discovery of potential targets for therapeutic intervention in diseases such as cancer. We consider the general question of the potential effect of individual genes on the global dynamical network behavior, both from the view of random gene perturbation as well as intervention in order to elicit desired network behavior. RESULTS: Using a recently introduced class of models, called Probabilistic Boolean Networks (PBNs), this paper develops a model for random gene perturbations and derives an explicit formula for the transition probabilities in the new PBN model. This result provides a building block for performing simulations and deriving other results concerning network dynamics. An example is provided to show how the gene perturbation model can be used to compute long-term influences of genes on other genes. Following this, the problem of intervention is addressed via the development of several computational tools based on first-passage times in Markov chains. The consequence is a methodology for finding the best gene with which to intervene in order to most likely achieve desirable network behavior. The ideas are illustrated with several examples in which the goal is to induce the network to transition into a desired state, or set of states. The corresponding issue of avoiding undesirable states is also addressed. Finally, the paper turns to the important problem of assessing the effect of gene perturbations on long-run network behavior. A bound on the steady-state probabilities is derived in terms of the perturbation probability. The result demonstrates that states of the network that are more 'easily reachable' from other states are more stable in the presence of gene perturbations. Consequently, these are hypothesized to correspond to cellular functional states. AVAILABILITY: A library of functions written in MATLAB for simulating PBNs, constructing state-transition matrices, computing steady-state distributions, computing influences, modeling random gene perturbations, and finding optimal intervention targets, as described in this paper, is available on request from [email protected].
Ilya Shmulevich, Edward R. Dougherty, Wei Zhang 0011
Bioinform.2
2002 Probabilistic Boolean networks: a rule-based uncertainty model for gene regulatory networks
abstract
MOTIVATION: Our goal is to construct a model for genetic regulatory networks such that the model class: (i) incorporates rule-based dependencies between genes; (ii) allows the systematic study of global network dynamics; (iii) is able to cope with uncertainty, both in the data and the model selection; and (iv) permits the quantification of the relative influence and sensitivity of genes in their interactions with other genes. RESULTS: We introduce Probabilistic Boolean Networks (PBN) that share the appealing rule-based properties of Boolean networks, but are robust in the face of uncertainty. We show how the dynamics of these networks can be studied in the probabilistic context of Markov chains, with standard Boolean networks being special cases. Then, we discuss the relationship between PBNs and Bayesian networks--a family of graphical models that explicitly represent probabilistic relationships between variables. We show how probabilistic dependencies between a gene and its parent genes, constituting the basic building blocks of Bayesian networks, can be obtained from PBNs. Finally, we present methods for quantifying the influence of genes on other genes, within the context of PBNs. Examples illustrating the above concepts are presented throughout the paper.
Ilya Shmulevich, Edward R. Dougherty, Seungchan Kim, Wei Zhang 0011
Bioinform.2
2002 ayesian automatic relevance determination algorithms for classifying gene expression data
abstract
Abstract Motivation: We investigate two new Bayesian classification algorithms incorporating feature selection. These algorithms are applied to the classification of gene expression data derived from cDNA microarrays. Results: We demonstrate the effectiveness of the algorithms on three gene expression datasets for cancer, showing they compare well with alternative kernel-based techniques. By automatically incorporating feature selection, accurate classifiers can be constructed utilizing very few features and with minimal hand-tuning. We argue that the feature selection is meaningful and some of the highlighted genes appear to be medically important. Contact: [email protected] * To whom correspondence should be addressed. Present address: Information and Mathematical Sciences, Genome Institute of Singapore, 1 Science Park Road, The Capricorn #05-01, Singapore 117528, Republic of Singapore
Ilya Shmulevich, Edward R. Dougherty, Wei Zhang 0011
Bioinform.2
2002 From Boolean to probabilistic Boolean networks as models of genetic regulatory networks
abstract
Mathematical and computational modeling of genetic regulatory networks promises to uncover the fundamental principles governing biological systems in an integrative and holistic manner. It also paves the way toward the development of systematic approaches for effective therapeutic intervention in disease. The central theme in this paper is the Boolean formalism as a building block for modeling complex, large-scale, and dynamical networks of genetic interactions. We discuss the goals of modeling genetic networks as well as the data requirements. The Boolean formalism is justified from several points of view. We then introduce Boolean networks and discuss their relationships to nonlinear digital filters. The role of Boolean networks in understanding cell differentiation and cellular functional states is discussed. The inference of Boolean networks from real gene expression data is considered from the viewpoints of computational learning theory and nonlinear signal processing, touching on computational complexity of learning and robustness. Then, a discussion of the need to handle uncertainty in a probabilistic framework is presented, leading to an introduction of probabilistic Boolean networks and their relationships to Markov chains. Methods for quantifying the influence of genes on other genes are presented. The general question of the potential effect of individual genes on the global dynamical network behavior is considered using stochastic perturbation analysis. This discussion then leads into the problem of target identification for therapeutic intervention via the development of several computational tools based on first-passage times in Markov chains. Examples from biology are presented throughout the paper.
Ilya Shmulevich, Edward R. Dougherty, Wei Zhang 0011
Proc. IEEE2
2002 Optimal linear granulometric estimation for random sets
Yoganand Balagurunathan, Edward R. Dougherty
Pattern Recognit.2
2001 Morphological granulometric estimation of random patterns in the context of parameterized random sets
Sinan Batman, Edward R. Dougherty
Pattern Recognit.2
2001 Non-homothetic granulometric mixing theory with application to blood cell counting
Nipon Theera-Umpon, Edward R. Dougherty, Paul D. Gader
Pattern Recognit.2
2001 Asymptotic joint normality of the granulometric moments
Krishnamoorthy Sivakumar, Yoganand Balagurunathan, Edward R. Dougherty
Pattern Recognit. Lett.3
2001 Robust optimal granulometric bandpass filters
Edward R. Dougherty, Yidong Chen 0002
Signal Process.1
2001 Bayesian robust optimal linear filters
Artyom M. Grigoryan, Edward R. Dougherty
Signal Process.2
2000 Heterogeneous morphological granulometries
Sinan Batman, Edward R. Dougherty, Francis Sand
Pattern Recognit.2
2000 A switching algorithm for design of optimal increasing binary filters over large windows
Nina Sumiko Tomita Hirata, Edward R. Dougherty, Junior Barrera
Pattern Recognit.2
2000 Hybrid human-machine binary morphological operator design. An independent constraint approach
Junior Barrera, Edward R. Dougherty, Marcel Brun
Signal Process.2
2000 Coefficient of determination in nonlinear signal processing
Edward R. Dougherty, Seungchan Kim, Yidong Chen 0002
Signal Process.1
2000 Aperture filters
Roberto Hirata Jr., Edward R. Dougherty, Junior Barrera
Signal Process.2
1999 Maximum-likelihood estimation and optimal filtering in the nondirectional, one-dimensional binomial germ-grain model
John C. Handley, Edward R. Dougherty
Pattern Recognit.2
1999 Robustness of granulometric moments
Francis Sand, Edward R. Dougherty
Pattern Recognit.2
1999 Probability distributions for discrete one-dimensional coverage processes
John C. Handley, Edward R. Dougherty
Signal Process.2
1998 Segmentation of Mammograms into Distinct Morphological Texture Regions
abstract
Presents a comprehensive discussion on the segmentation of mammograms using morphological texture features. These features are derived from morphological granulometries with various structuring elements. Each structuring element captures a specific texture content. The segmentation is carried out in an unsupervised manner by applying the KL (Karhunen-Loeve) transform feature reduction and Voronoi clustering on the extracted morphological texture features. The evaluation of the segmentation outcome by a trained radiologist is provided.
Sooncheol Baeg, Anthony T. Popov, Vishnu G. Kamat, Sinan Batman, Krishnamoorthy Sivakumar, Nasser Kehtarnavaz, Edward R. Dougherty, Robert B. Shah
CBMS7
1998 Asymptotic granulometric mixing theorem: Morphological estimation of sizing parameters and mixture proportions
Francis Sand, Edward R. Dougherty
Pattern Recognit.2
1998 Secondarily constrained Boolean filters
Octavian V. Sarca, Edward R. Dougherty, Jaakko Astola
Signal Process.2
1997 Maximum-Likelihood Estimation for the Two-Dimensional Discrete Boolean Random Set and Function Models Using Multidimensional Linear Samples
John C. Handley, Edward R. Dougherty
CVGIP Graph. Model. Image Process.2
1997 Optimal and adaptive reconstructive granulometric bandpass filters
Yidong Chen 0002, Edward R. Dougherty
Signal Process.2
1997 Optimal reconstructive τ-openings for disjoint and statistically modeled nondisjoint grains
Edward R. Dougherty, Clara Cuciurean-Zapan
Signal Process.1
1997 Design and analysis of fuzzy morphological algorithms for image processing
abstract
A general paradigm for lifting binary morphological algorithms to fuzzy algorithms is employed to construct fuzzy versions of classical binary morphological operations. The lifting procedure is based upon an epistemological interpretation of both image and filter fuzzification. Algorithms are designed via the paradigm for various fuzzifications and their performances are analyzed to provide insight into the kind of liftings that produce suitable results. Algorithms are discussed for three image processing tasks: shape detection, edge detection, and clutter removal. Detailed analyses are given for the effect of noise and its mitigation owing to fuzzy approaches. It is demonstrated how the fuzzy hit-or-miss transform can be used in conjunction with a decision procedure to achieve word recognition.
Divyendu Sinha, Purnendu Sinha, Edward R. Dougherty, Sinan Batman
IEEE Trans. Fuzzy Syst.3
1996 Bayesian morphological peak estimation and its application to chromosome counting via fluorescence In situ hybridization
Edward R. Dougherty, Yidong Chen 0002, Amir Waks
Pattern Recognit.1
1996 Optimal nonlinear filter for signal-union-noise and runlength analysis in the directional one-dimensional discrete Boolean random set model
John C. Handley, Edward R. Dougherty
Signal Process.2
1995 Representation of Linear Granulometric Moments for Deterministic and Random Binary Euclidean Images
Edward R. Dougherty, Francis Sand
J. Vis. Commun. Image Represent.1
1995 Morphological pattern-spectrum classification of noisy shapes: Exterior granulometries
Edward R. Dougherty, Yingchong Cheng
Pattern Recognit.1
1995 Recursive maximum-likelihood estimation in the one-dimensional discrete Boolean random set model
Edward R. Dougherty, John C. Handley
Signal Process.1
1995 A general axiomatic theory of intrinsically fuzzy mathematical morphologies
abstract
Intrinsic fuzzification of mathematical morphology is grounded on an axiomatic characterization of subset fuzzification. The result is an axiomatic formulation of fuzzy Minkowski algebra. Part of the Minkowski algebra results solely from the axioms themselves and part results from a specific postulated form of a subsethood indicator function. There exists an infinite number of fuzzy morphologies satisfying the axioms; in particular, there are uncountably many indicators satisfying the postulated form. This paper develops fuzzy Minkowski algebra, with special emphasis on fitting characterizations of fuzzy erosion and opening, examines key properties of the indicator function, and provides fuzzy extensions of the basic binary Matheron representations for openings and increasing, translation-invariant operators.
Divyendu Sinha, Edward R. Dougherty
IEEE Trans. Fuzzy Syst.2
1994 Minimal representation of -openings via pattern bases
Edward R. Dougherty
Pattern Recognit. Lett.1
1994 Precision of morphological-representation estimators for translation-invariant binary filters: Increasing and nonincreasing
Edward R. Dougherty, Robert P. Loce
Signal Process.1
1994 Computational mathematical morphology
Edward R. Dougherty, Divyendu Sinha
Signal Process.1
1993 Computational morphology and representation of operators between complete lattices
abstract
The representations of translation-invariant mappings in the context of computational morphology are in terms of elementary binary erosions and dilations. The general representations in the context of complete lattices are in terms of abstract dilations and antidilations. The present paper considers the relationship between the specialized computational and general lattice representations. The computational and lattice-based representations are directly demonstrated to be equivalent in the computational setting.
Clara Cuciurean-Zapan, Edward R. Dougherty
VCIP2
1993 Gray-scale granulometries compatible with spatial scalings
Eugene J. Kraus, Henk J. A. M. Heijmans, Edward R. Dougherty
Signal Process.3
1992 The use of first-order structuring-element libraries to design morphological filters
abstract
Statistically optimized morphological filters are preferable to those traditionally selected by humans. Nevertheless, full optimization has been shown to be computationally intractable. By applying first-order knowledge to select a predetermined structuring-element library upon which to apply optimization, one can greatly reduce design computation, while at the same time producing good filters. The paper sets down a paradigm for library optimization and presents a methodology for first-order-library construction. Experimental results depicted herein illustrate the goodness of the estimations.>
Robert P. Loce, Edward R. Dougherty
ICPR (3)2
1992 Optimal mean-square N-observation digital morphological filters : I. Optimal binary filters
Edward R. Dougherty
CVGIP Image Underst.1
1992 Optimal mean-square N-observation digital morphological filters : II. Optimal gray-scale filters
Edward R. Dougherty
CVGIP Image Underst.1
1992 Robust Morphologically Continuous Fourier Descriptors I: Projection-Generated Descriptors
abstract
By generating Fourier descriptors based upon the waveform induced by a pattern's geometric projection, a number of classic difficulties with the Fourier-descriptor methodology are mitigated. Not only are the descriptors invariant with respect to scale, translation, and rotation (as is usually the case), they are also continuous in the Hausdorff metric and robust with respect to both point noise and occlusion. An additional advantage is that they can be computed relative to a thresholded image without first finding an edge, thereby avoiding the difficulties typically present in thinning and orientation determination. The present paper discusses the method of projection-generated Fourier descriptors, as well as a study of the sensitivity to point noise. A companion paper will present the morphological properties and the effect of pattern occlusion.
Edward R. Dougherty, Robert P. Loce
Int. J. Pattern Recognit. Artif. Intell.1
1992 Robust Morphologically Continuous Fourier Descriptors II: Continuity and Occlusion Analysis
abstract
Fourier descriptors based upon the waveform induced by a pattern's projection overcome a number of classic difficulties with Fourier-descriptor methodology. Not only are the descriptors invariant with respect to scale, translation, and rotation (as is usually the case), they are also continuous in the Hausdorff metric and robust with respect to both point noise and occlusion; insensitivity with respect to minimum occlusions is perhaps their most significant advantage. Continuity in the Hausdorff metric allows prediction of the effect on the descriptors when morphologically filtering a pattern. The effect of occlusion is also predictable.
Edward R. Dougherty, Robert P. Loce
Int. J. Pattern Recognit. Artif. Intell.1
1992 Model-based characterization of statistically optimal design for morphological shape recognition algorithms via the hit-or-miss transform
Edward R. Dougherty, Dongming Zhao 0001
J. Vis. Commun. Image Represent.1
1992 Optimal morphological restoration: The morphological filter mean-absolute-error theorem
Robert P. Loce, Edward R. Dougherty
J. Vis. Commun. Image Represent.2
1992 Asymptotic normality of the morphological pattern-spectrum moments and orthogonal granulometric generators
Francis Sand, Edward R. Dougherty
J. Vis. Commun. Image Represent.2
1992 Preface
Jean Paul Frédéric Serra, Edward R. Dougherty
J. Vis. Commun. Image Represent.2
1992 Fuzzy mathematical morphology
Divyendu Sinha, Edward R. Dougherty
J. Vis. Commun. Image Represent.2
1992 Morphological texture-based maximum-likelihood pixel classification based on local granulometric moments
Edward R. Dougherty, John T. Newell, Jeff B. Pelz
Pattern Recognit.1
1992 Estimation of optimal morphological τ-opening parameters based on independent observation of signal and noise pattern spectra
Edward R. Dougherty, Robert M. Haralick, Yidong Chen 0002, Carsten Agerskov, Ulrik Jacobi, Poul Henrik Sloth
Signal Process.1
1991 Application of the Hausdorff metric in gray-scale mathematical morphology via truncated umbrae
Edward R. Dougherty
J. Vis. Commun. Image Represent.1
1990 Hausdorf-metric interpretation of convergence in the Matheron topology for binary mathematical morphology
abstract
The basic convergence properties of mathematical morphology are characterized in terms of the topology of G. Matheron (1975). That topology is grounded on a particular subbase that can often mask the important metric properties that are consequential to Euclidean morphology. The author presents a development of some of the key Matheron theory in terms of the Hausdorff metric, thereby bypassing the Matheron subbase and giving both theorems and proofs in a metric framework.>
Edward R. Dougherty
ICPR (1)1
1990 Minimal search for the optimal mean-square digital gray-scale morphological filter
abstract
To characterize optimal mean-square morphological filters, it is first necessary to interpret morphological operations in a functional manner appropriate to the theory of statistical estimation. The present paper takes such an approach in the case of digital N-observation grayscale filters, these being defined via the Matheron representation. Having obtained the optimality criterion, we are lead to the characterization of a minimal search space, the nodes of the space being potential erosion structuring elements. More precisely, there exists a set of structuring elements which will always contain elements forming the basis for an optimal MS filter. Moreover, the set, called the fundamental set, is minimal, in the sense that no element can be deleted from it without possibly yielding a set not containing the optimal structuring element for a single-erosion filter.
Edward R. Dougherty
VCIP1
1989 The dual representation of gray-scale morphological filters
abstract
One of the classic results of mathematical morphology is the filter-representation theorem of G. Matheron (1975) for black-and-white images. The theorem states that any morphological filter can be represented as a union of erosions by elements in the filter's kernel. In its dual form, it states that the erosion representation can be replaced by an intersection of dilations by elements of the dual filter's kernel. Here, the dual-form of the gray-scale representation is derived in terms of a minimum of dilations by elements in the dual filter's kernel.>
Edward R. Dougherty
CVPR1
1988 Closed-form representation of convolution, dilation, and erosion in the context of image algebra
abstract
Using fundamental operators from image algebra, the authors present simple closed-form expressions for dilation, erosion, and convolution. Algebraically, these expressions appear as terms within the algebra. Moreover, the methodology for obtaining the expressions reveals a universal operational structure within image algebra, of which the three aforementioned operations are particular instances. The result is a natural parallel mechanism for computation and a representation of convolution that naturally overcomes the difficulties arising from the variability of image domains in the defining relation.>
Edward R. Dougherty, Charles R. Giardina
CVPR1
1988 A robust image processing language in the context of image algebra
abstract
A high level image processing language is discussed that provides a robust environment for the universal specification and implementation of image processing algorithms. In essence, an image processing program consists of a collection of procedure calls. To implement an algorithm, one needs only to construct a block-diagram representation of the algorithm and the proceed with one-to-one translation of the block diagram into the host Pascal language. The result is that an investigator can proceed with rapid variations at the algorithm design stage without having to be involved with writing code. The development of the language rests on three pillars: (1) the existence of a rigorous and robust image algebra, (2) the block diagram specification of image-processing algorithms, and (3) the bound matrix image representation.>
Edward R. Dougherty, Paramjit S. Sehdev
CVPR1
1988 Morphology on umbra Matrices
abstract
The umbra transform serves as a connection between gray-scale morphology and the classical two-valued morphology of G. Matheron and H. Hadwiger. From a general set-theoretic perspective, the umbra transform of an image (or signal) results in an infinite set, even in the discrete case. By employing bound matrix image representation it is possible to represent the umbra by a finite data structure, the result being an approach that is both intuitive and computational. Moreover, the method is essentially dimensionally independent and thus applies to both morphological image and signal processing.
Edward R. Dougherty, Charles R. Giardina
Int. J. Pattern Recognit. Artif. Intell.1