VLDB 2026 Research / reviewers in the wild / expert
William S. Hlavacek
dblp:02/3329
· DBLP profile ↗
17ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0003-4383-8711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 1 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 50% Cloud and datacenter computing · 50% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
markov chain monte carlo |
0.6 | 1 | 2022 | Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFit · Bioinform. 2022 |
Bioinformatics and computational biology › systems biology
rule-based modeling |
0.4 | 3 | 2017 | SPATKIN: a simulator for rule-based modeling of biomolecular site dynamics on surfaces · Bioinform. 2017 GetBonNie for building, analyzing and sharing rule-based models · Bioinform. 2009 BioNetGen: software for rule-based modeling of signal transduction based on the interactions of molecular domains · Bioinform. 2004 |
Bioinformatics and computational biology › systems biology
parameter estimation |
0.2 | 1 | 2016 | BioNetFit: a fitting tool compatible with BioNetGen, NFsim and distributed computing environments · Bioinform. 2016 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling |
0.2 | 1 | 2016 | BioNetFit: a fitting tool compatible with BioNetGen, NFsim and distributed computing environments · Bioinform. 2016 |
High-performance computing › scientific computing
distributed scientific computing |
0.2 | 1 | 2016 | BioNetFit: a fitting tool compatible with BioNetGen, NFsim and distributed computing environments · Bioinform. 2016 |
Bioinformatics and computational biology
systems biology |
0.2 | 2 | 2009 | GetBonNie for building, analyzing and sharing rule-based models · Bioinform. 2009 Simulation of large-scale rule-based models · Bioinform. 2009 |
Bioinformatics and computational biology › systems biology › metabolic network reconstruction
metabolic reaction prediction |
0.2 | 2 | 2011 | Prediction of metabolic reactions based on atomic and molecular properties of small-molecule compounds · Bioinform. 2011 Prediction of oxidoreductase-catalyzed reactions based on atomic properties of metabolites · Bioinform. 2006 |
Bioinformatics and computational biology › systems biology
metabolic network analysis |
0.1 | 2 | 2007 | Carbon-fate maps for metabolic reactions · Bioinform. 2007 Prediction of oxidoreductase-catalyzed reactions based on atomic properties of metabolites · Bioinform. 2006 |
Bioinformatics and computational biology › metabolomics
isotope labeling analysis |
0.1 | 1 | 2007 | Carbon-fate maps for metabolic reactions · Bioinform. 2007 |
Bioinformatics and computational biology › systems biology › cellular signaling
signal transduction modeling |
0.0 | 1 | 2004 | BioNetGen: software for rule-based modeling of signal transduction based on the interactions of molecular domains · Bioinform. 2004 |
Bioinformatics and computational biology › systems biology
model reuse |
0.0 | 1 | 2009 | GetBonNie for building, analyzing and sharing rule-based models · Bioinform. 2009 |
Methods — techniques the papers use, named apart from their topics
bayesian inference · 1.0markov chain monte carlo · 0.6adaptive metropolis · 0.6rule-based simulation · 0.5parameter optimization · 0.5likelihood function formulation · 0.4stochastic simulation · 0.3lattice-based brownian dynamics · 0.3support vector machine · 0.2bionetgen language · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFitabstractSUMMARY: Bayesian inference in biological modeling commonly relies on Markov chain Monte Carlo (MCMC) sampling of a multidimensional and non-Gaussian posterior distribution that is not analytically tractable. Here, we present the implementation of a practical MCMC method in the open-source software package PyBioNetFit (PyBNF), which is designed to support parameterization of mathematical models for biological systems. The new MCMC method, am, incorporates an adaptive move proposal distribution. For warm starts, sampling can be initiated at a specified location in parameter space and with a multivariate Gaussian proposal distribution defined initially by a specified covariance matrix. Multiple chains can be generated in parallel using a computer cluster. We demonstrate that am can be used to successfully solve real-world Bayesian inference problems, including forecasting of new Coronavirus Disease 2019 case detection with Bayesian quantification of forecast uncertainty. AVAILABILITY AND IMPLEMENTATION: PyBNF version 1.1.9, the first stable release with am, is available at PyPI and can be installed using the pip package-management system on platforms that have a working installation of Python 3. PyBNF relies on libRoadRunner and BioNetGen for simulations (e.g. numerical integration of ordinary differential equations defined in SBML or BNGL files) and Dask.Distributed for task scheduling on Linux computer clusters. The Python source code can be freely downloaded/cloned from GitHub and used and modified under terms of the BSD-3 license (https://github.com/lanl/pybnf). Online documentation covering installation/usage is available (https://pybnf.readthedocs.io/en/latest/). A tutorial video is available on YouTube (https://www.youtube.com/watch?v=2aRqpqFOiS4&t=63s). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jacob Neumann, Abhishek Mallela, Ely F. Miller, Joshua Colvin, Abell T. Duprat, William S. Hlavacek, Richard G. Posner |
Bioinform. | 8 |
| 2020 | Bayesian inference using qualitative observations of underlying continuous variablesabstractMOTIVATION: Recent work has demonstrated the feasibility of using non-numerical, qualitative data to parameterize mathematical models. However, uncertainty quantification (UQ) of such parameterized models has remained challenging because of a lack of a statistical interpretation of the objective functions used in optimization. RESULTS: We formulated likelihood functions suitable for performing Bayesian UQ using qualitative observations of underlying continuous variables or a combination of qualitative and quantitative data. To demonstrate the resulting UQ capabilities, we analyzed a published model for immunoglobulin E (IgE) receptor signaling using synthetic qualitative and quantitative datasets. Remarkably, estimates of parameter values derived from the qualitative data were nearly as consistent with the assumed ground-truth parameter values as estimates derived from the lower throughput quantitative data. These results provide further motivation for leveraging qualitative data in biological modeling. AVAILABILITY AND IMPLEMENTATION: The likelihood functions presented here are implemented in a new release of PyBioNetFit, an open-source application for analyzing Systems Biology Markup Language- and BioNetGen Language-formatted models, available online at www.github.com/lanl/PyBNF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eshan D. Mitra, William S. Hlavacek |
Bioinform. | 2 |
| 2019 | Modeling cell line-specific recruitment of signaling proteins to the insulin-like growth factor 1 receptorabstractReceptor tyrosine kinases (RTKs) typically contain multiple autophosphorylation sites in their cytoplasmic domains. Once activated, these autophosphorylation sites can recruit downstream signaling proteins containing Src homology 2 (SH2) and phosphotyrosine-binding (PTB) domains, which recognize phosphotyrosine-containing short linear motifs (SLiMs). These domains and SLiMs have polyspecific or promiscuous binding activities. Thus, multiple signaling proteins may compete for binding to a common SLiM and vice versa. To investigate the effects of competition on RTK signaling, we used a rule-based modeling approach to develop and analyze models for ligand-induced recruitment of SH2/PTB domain-containing proteins to autophosphorylation sites in the insulin-like growth factor 1 (IGF1) receptor (IGF1R). Models were parameterized using published datasets reporting protein copy numbers and site-specific binding affinities. Simulations were facilitated by a novel application of model restructuration, to reduce redundancy in rule-derived equations. We compare predictions obtained via numerical simulation of the model to those obtained through simple prediction methods, such as through an analytical approximation, or ranking by copy number and/or KD value, and find that the simple methods are unable to recapitulate the predictions of numerical simulations. We created 45 cell line-specific models that demonstrate how early events in IGF1R signaling depend on the protein abundance profile of a cell. Simulations, facilitated by model restructuration, identified pairs of IGF1R binding partners that are recruited in anti-correlated and correlated fashions, despite no inclusion of cooperativity in our models. This work shows that the outcome of competition depends on the physicochemical parameters that characterize pairwise interactions, as well as network properties, including network connectivity and the relative abundances of competitors. Keesha E. Erickson, Oleksii S. Rukhlenko, MD Shahinuzzaman, Kalina P. Slavkova, Ryan Suderman, Edward C. Stites, Marian Anghel, Richard G. Posner, Dipak Barua, Boris N. Kholodenko, William S. Hlavacek |
PLoS Comput. Biol. | 12 |
| 2017 | SPATKIN: a simulator for rule-based modeling of biomolecular site dynamics on surfacesabstractSUMMARY: Rule-based modeling is a powerful approach for studying biomolecular site dynamics. Here, we present SPATKIN, a general-purpose simulator for rule-based modeling in two spatial dimensions. The simulation algorithm is a lattice-based method that tracks Brownian motion of individual molecules and the stochastic firing of rule-defined reaction events. Because rules are used as event generators, the algorithm is network-free, meaning that it does not require to generate the complete reaction network implied by rules prior to simulation. In a simulation, each molecule (or complex of molecules) is taken to occupy a single lattice site that cannot be shared with another molecule (or complex). SPATKIN is capable of simulating a wide array of membrane-associated processes, including adsorption, desorption and crowding. Models are specified using an extension of the BioNetGen language, which allows to account for spatial features of the simulated process. AVAILABILITY AND IMPLEMENTATION: The C ++ source code for SPATKIN is distributed freely under the terms of the GNU GPLv3 license. The source code can be compiled for execution on popular platforms (Windows, Mac and Linux). An installer for 64-bit Windows and a macOS app are available. The source code and precompiled binaries are available at the SPATKIN Web site (http://pmbm.ippt.pan.pl/software/spatkin). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marek Kochanczyk, William S. Hlavacek, Tomasz Lipniacki |
Bioinform. | 2 |
| 2016 | BioNetFit: a fitting tool compatible with BioNetGen, NFsim and distributed computing environmentsabstractUNLABELLED: Rule-based models are analyzed with specialized simulators, such as those provided by the BioNetGen and NFsim open-source software packages. Here, we present BioNetFit, a general-purpose fitting tool that is compatible with BioNetGen and NFsim. BioNetFit is designed to take advantage of distributed computing resources. This feature facilitates fitting (i.e. optimization of parameter values for consistency with data) when simulations are computationally expensive. AVAILABILITY AND IMPLEMENTATION: BioNetFit can be used on stand-alone Mac, Windows/Cygwin, and Linux platforms and on Linux-based clusters running SLURM, Torque/PBS, or SGE. The BioNetFit source code (Perl) is freely available (http://bionetfit.nau.edu). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Brandon R. Thomas, Lily A. Chylek, Joshua Colvin, Suman Sirimulla, Andrew H. A. Clayton, William S. Hlavacek, Richard G. Posner |
Bioinform. | 6 |
| 2013 | Modeling the Effect of APC Truncation on Destruction Complex Function in Colorectal Cancer CellsabstractIn colorectal cancer cells, APC, a tumor suppressor protein, is commonly expressed in truncated form. Truncation of APC is believed to disrupt degradation of β-catenin, which is regulated by a multiprotein complex called the destruction complex. The destruction complex comprises APC, Axin, β-catenin, serine/threonine kinases, and other proteins. The kinases CK1α and GSK -3β, which are recruited by Axin, mediate phosphorylation of β-catenin, which initiates its ubiquitination and proteosomal degradation. The mechanism of regulation of β-catenin degradation by the destruction complex and the role of truncation of APC in colorectal cancer are not entirely understood. Through formulation and analysis of a rule-based computational model, we investigated the regulation of β-catenin phosphorylation and degradation by APC and the effect of APC truncation on function of the destruction complex. The model integrates available mechanistic knowledge about site-specific interactions and phosphorylation of destruction complex components and is consistent with an array of published data. We find that the phosphorylated truncated form of APC can outcompete Axin for binding to β-catenin, provided that Axin is limiting, and thereby sequester β-catenin away from Axin and the Axin-recruited kinases CK1α and GSK -3β. Full-length APC also competes with Axin for binding to β-catenin; however, full-length APC is able, through its SAMP repeats, which bind Axin and which are missing in truncated oncogenic forms of APC, to bring β-catenin into indirect association with Axin and Axin-recruited kinases. Because our model indicates that the positive effects of truncated APC on β-catenin levels depend on phosphorylation of APC, at the first 20-amino acid repeat, and because phosphorylation of this site is mediated by CK1ε, we suggest that CK1ε is a potential target for therapeutic intervention in colorectal cancer. Specific inhibition of CK1ε is predicted to limit binding of β-catenin to truncated APC and thereby to reverse the effect of APC truncation. Dipak Barua, William S. Hlavacek |
PLoS Comput. Biol. | 2 |
| 2013 | Binding of Nucleoid-Associated Protein Fis to DNA Is Regulated by DNA Breathing DynamicsabstractPhysicochemical properties of DNA, such as shape, affect protein-DNA recognition. However, the properties of DNA that are most relevant for predicting the binding sites of particular transcription factors (TFs) or classes of TFs have yet to be fully understood. Here, using a model that accurately captures the melting behavior and breathing dynamics (spontaneous local openings of the double helix) of double-stranded DNA, we simulated the dynamics of known binding sites of the TF and nucleoid-associated protein Fis in Escherichia coli. Our study involves simulations of breathing dynamics, analysis of large published in vitro and genomic datasets, and targeted experimental tests of our predictions. Our simulation results and available in vitro binding data indicate a strong correlation between DNA breathing dynamics and Fis binding. Indeed, we can define an average DNA breathing profile that is characteristic of Fis binding sites. This profile is significantly enriched among the identified in vivo E. coli Fis binding sites. To test our understanding of how Fis binding is influenced by DNA breathing dynamics, we designed base-pair substitutions, mismatch, and methylation modifications of DNA regions that are known to interact (or not interact) with Fis. The goal in each case was to make the local DNA breathing dynamics either closer to or farther from the breathing profile characteristic of a strong Fis binding site. For the modified DNA segments, we found that Fis-DNA binding, as assessed by gel-shift assay, changed in accordance with our expectations. We conclude that Fis binding is associated with DNA breathing dynamics, which in turn may be regulated by various nucleotide modifications. Kristy Nowak-Lovato, Ludmil B. Alexandrov, Afsheen Banisadr, Amy L. Bauer, Alan R. Bishop, Anny Usheva, Fangping Mu, Elizabeth Hong-Geller, Kim Ø. Rasmussen, William S. Hlavacek, Boian S. Alexandrov |
PLoS Comput. Biol. | 10 |
| 2011 | Prediction of metabolic reactions based on atomic and molecular properties of small-molecule compoundsabstractMOTIVATION: Our knowledge of the metabolites in cells and their reactions is far from complete as revealed by metabolomic measurements that detect many more small molecules than are documented in metabolic databases. Here, we develop an approach for predicting the reactivity of small-molecule metabolites in enzyme-catalyzed reactions that combines expert knowledge, computational chemistry and machine learning. RESULTS: We classified 4843 reactions documented in the KEGG database, from all six Enzyme Commission classes (EC 1-6), into 80 reaction classes, each of which is marked by a characteristic functional group transformation. Reaction centers and surrounding local structures in substrates and products of these reactions were represented using SMARTS. We found that each of the SMARTS-defined chemical substructures is widely distributed among metabolites, but only a fraction of the functional groups in these substructures are reactive. Using atomic properties of atoms in a putative reaction center and molecular properties as features, we trained support vector machine (SVM) classifiers to discriminate between functional groups that are reactive and non-reactive. Classifier accuracy was assessed by cross-validation analysis. A typical sensitivity [TP/(TP+FN)] or specificity [TN/(TN+FP)] is ≈0.8. Our results suggest that metabolic reactivity of small-molecule compounds can be predicted with reasonable accuracy based on the presence of a potentially reactive functional group and the chemical features of its local environment. AVAILABILITY: The classifiers presented here can be used to predict reactions via a web site (http://cellsignaling.lanl.gov/Reactivity/). The web site is freely available. Fangping Mu, Clifford J. Unkefer, Pat J. Unkefer, William S. Hlavacek |
Bioinform. | 4 |
| 2011 | Hierarchical graphs for rule-based modeling of biochemical systemsabstractBACKGROUND: In rule-based modeling, graphs are used to represent molecules: a colored vertex represents a component of a molecule, a vertex attribute represents the internal state of a component, and an edge represents a bond between components. Components of a molecule share the same color. Furthermore, graph-rewriting rules are used to represent molecular interactions. A rule that specifies addition (removal) of an edge represents a class of association (dissociation) reactions, and a rule that specifies a change of a vertex attribute represents a class of reactions that affect the internal state of a molecular component. A set of rules comprises an executable model that can be used to determine, through various means, the system-level dynamics of molecular interactions in a biochemical system. RESULTS: For purposes of model annotation, we propose the use of hierarchical graphs to represent structural relationships among components and subcomponents of molecules. We illustrate how hierarchical graphs can be used to naturally document the structural organization of the functional components and subcomponents of two proteins: the protein tyrosine kinase Lck and the T cell receptor (TCR) complex. We also show that computational methods developed for regular graphs can be applied to hierarchical graphs. In particular, we describe a generalization of Nauty, a graph isomorphism and canonical labeling algorithm. The generalized version of the Nauty procedure, which we call HNauty, can be used to assign canonical labels to hierarchical graphs or more generally to graphs with multiple edge types. The difference between the Nauty and HNauty procedures is minor, but for completeness, we provide an explanation of the entire HNauty algorithm. CONCLUSIONS: Hierarchical graphs provide more intuitive formal representations of proteins and other structured molecules with multiple functional components than do the regular graphs of current languages for specifying rule-based models, such as the BioNetGen language (BNGL). Thus, the proposed use of hierarchical graphs should promote clarity and better understanding of rule-based models. Nathan Lemons, William S. Hlavacek |
BMC Bioinform. | 3 |
| 2010 | Systems Biology Approaches for Developing Artificial Pathogen Detection Networks
Anu Chaudhary, Geoffrey Waldo, William S. Hlavacek, Chang-Shung Tung |
ALIFE | 3 |
| 2010 | RuleMonkey: software for stochastic simulation of rule-based modelsabstractBACKGROUND: The system-level dynamics of many molecular interactions, particularly protein-protein interactions, can be conveniently represented using reaction rules, which can be specified using model-specification languages, such as the BioNetGen language (BNGL). A set of rules implicitly defines a (bio)chemical reaction network. The reaction network implied by a set of rules is often very large, and as a result, generation of the network implied by rules tends to be computationally expensive. Moreover, the cost of many commonly used methods for simulating network dynamics is a function of network size. Together these factors have limited application of the rule-based modeling approach. Recently, several methods for simulating rule-based models have been developed that avoid the expensive step of network generation. The cost of these "network-free" simulation methods is independent of the number of reactions implied by rules. Software implementing such methods is now needed for the simulation and analysis of rule-based models of biochemical systems. RESULTS: Here, we present a software tool called RuleMonkey, which implements a network-free method for simulation of rule-based models that is similar to Gillespie's method. The method is suitable for rule-based models that can be encoded in BNGL, including models with rules that have global application conditions, such as rules for intramolecular association reactions. In addition, the method is rejection free, unlike other network-free methods that introduce null events, i.e., steps in the simulation procedure that do not change the state of the reaction system being simulated. We verify that RuleMonkey produces correct simulation results, and we compare its performance against DYNSTOC, another BNGL-compliant tool for network-free simulation of rule-based models. We also compare RuleMonkey against problem-specific codes implementing network-free simulation methods. CONCLUSIONS: RuleMonkey enables the simulation of rule-based models for which the underlying reaction networks are large. It is typically faster than DYNSTOC for benchmark problems that we have examined. RuleMonkey is freely available as a stand-alone application http://public.tgen.org/rulemonkey. It is also available as a simulation engine within GetBonNie, a web-based environment for building, analyzing and sharing rule-based models. Joshua Colvin, Michael I. Monine, Ryan N. Gutenkunst, William S. Hlavacek, Daniel D. Von Hoff, Richard G. Posner |
BMC Bioinform. | 4 |
| 2010 | Using Sequence-Specific Chemical and Structural Properties of DNA to Predict Transcription Factor Binding SitesabstractAn important step in understanding gene regulation is to identify the DNA binding sites recognized by each transcription factor (TF). Conventional approaches to prediction of TF binding sites involve the definition of consensus sequences or position-specific weight matrices and rely on statistical analysis of DNA sequences of known binding sites. Here, we present a method called SiteSleuth in which DNA structure prediction, computational chemistry, and machine learning are applied to develop models for TF binding sites. In this approach, binary classifiers are trained to discriminate between true and false binding sites based on the sequence-specific chemical and structural features of DNA. These features are determined via molecular dynamics calculations in which we consider each base in different local neighborhoods. For each of 54 TFs in Escherichia coli, for which at least five DNA binding sites are documented in RegulonDB, the TF binding sites and portions of the non-coding genome sequence are mapped to feature vectors and used in training. According to cross-validation analysis and a comparison of computational predictions against ChIP-chip data available for the TF Fis, SiteSleuth outperforms three conventional approaches: Match, MATRIX SEARCH, and the method of Berg and von Hippel. SiteSleuth also outperforms QPMEME, a method similar to SiteSleuth in that it involves a learning algorithm. The main advantage of SiteSleuth is a lower false positive rate. Amy L. Bauer, William S. Hlavacek, Pat J. Unkefer, Fangping Mu |
PLoS Comput. Biol. | 2 |
| 2009 | Simulation of large-scale rule-based modelsabstractMOTIVATION: Interactions of molecules, such as signaling proteins, with multiple binding sites and/or multiple sites of post-translational covalent modification can be modeled using reaction rules. Rules comprehensively, but implicitly, define the individual chemical species and reactions that molecular interactions can potentially generate. Although rules can be automatically processed to define a biochemical reaction network, the network implied by a set of rules is often too large to generate completely or to simulate using conventional procedures. To address this problem, we present DYNSTOC, a general-purpose tool for simulating rule-based models. RESULTS: DYNSTOC implements a null-event algorithm for simulating chemical reactions in a homogenous reaction compartment. The simulation method does not require that a reaction network be specified explicitly in advance, but rather takes advantage of the availability of the reaction rules in a rule-based specification of a network to determine if a randomly selected set of molecular components participates in a reaction during a time step. DYNSTOC reads reaction rules written in the BioNetGen language which is useful for modeling protein-protein interactions involved in signal transduction. The method of DYNSTOC is closely related to that of StochSim. DYNSTOC differs from StochSim by allowing for model specification in terms of BNGL, which extends the range of protein complexes that can be considered in a model. DYNSTOC enables the simulation of rule-based models that cannot be simulated by conventional methods. We demonstrate the ability of DYNSTOC to simulate models accounting for multisite phosphorylation and multivalent binding processes that are characterized by large numbers of reactions. AVAILABILITY: DYNSTOC is free for non-commercial use. The C source code, supporting documentation and example input files are available at http://public.tgen.org/dynstoc/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joshua Colvin, Michael I. Monine, James R. Faeder, William S. Hlavacek, Daniel D. Von Hoff, Richard G. Posner |
Bioinform. | 4 |
| 2009 | GetBonNie for building, analyzing and sharing rule-based modelsabstractSUMMARY: GetBonNie is a web-based application for building, analyzing and sharing rule-based models encoded in the BioNetGen language (BNGL). Tools accessible within the GetBonNie environment include (i) an applet for drawing graphs that correspond to BNGL code; (ii) a network-generation engine for translating a set of rules into a chemical reaction network; (iii) simulation engines that implement generate-first, on-the-fly and network-free methods for simulating rule-based models; and (iv) a database for sharing models, parameter values, annotations, simulation tasks and results. AVAILABILITY: GetBonNie is free at (http://getbonnie.org). Bin Hu 0025, G. Matthew Fricke, James R. Faeder, Richard G. Posner, William S. Hlavacek |
Bioinform. | 5 |
| 2007 | Carbon-fate maps for metabolic reactionsabstractMOTIVATION: Stable isotope labeling of small-molecule metabolites (e.g. (13)C-labeling of glucose) is a powerful tool for characterizing pathways and reaction fluxes in a metabolic network. Analysis of isotope labeling patterns requires knowledge of the fates of individual atoms and moieties in reactions, which can be difficult to collect in a useful form when considering a large number of enzymatic reactions. RESULTS: We report carbon-fate maps for 4605 enzyme-catalyzed reactions documented in the KEGG database. Every fate map has been manually checked for consistency with known reaction mechanisms. A map includes a standardized structure-based identifier for each reactant (namely, an InChI string); indices for carbon atoms that are uniquely derived from the metabolite identifiers; structural data, including an identification of homotopic and prochiral carbon atoms; and a bijective map relating the corresponding carbon atoms in substrates and products. Fate maps are defined using the BioNetGen language (BNGL), a formal model-specification language, which allows a set of maps to be automatically translated into isotopomer mass-balance equations. AVAILABILITY: The carbon-fate maps and software for visualizing the maps are freely available (http://cellsignaling.lanl.gov/FateMaps/). Fangping Mu, Robert F. Williams, Clifford J. Unkefer, Pat J. Unkefer, James R. Faeder, William S. Hlavacek |
Bioinform. | 6 |
| 2006 | Prediction of oxidoreductase-catalyzed reactions based on atomic properties of metabolitesabstractMOTIVATION: Our knowledge of metabolism is far from complete, and the gaps in our knowledge are being revealed by metabolomic detection of small-molecules not previously known to exist in cells. An important challenge is to determine the reactions in which these compounds participate, which can lead to the identification of gene products responsible for novel metabolic pathways. To address this challenge, we investigate how machine learning can be used to predict potential substrates and products of oxidoreductase-catalyzed reactions. RESULTS: We examined 1956 oxidation/reduction reactions in the KEGG database. The vast majority of these reactions (1626) can be divided into 12 subclasses, each of which is marked by a particular type of functional group transformation. For a given transformation, the local structures of reaction centers in substrates and products can be characterized by patterns. These patterns are not unique to reactants but are widely distributed among KEGG metabolites. To distinguish reactants from non-reactants, we trained classifiers (linear-kernel Support Vector Machines) using negative and positive examples. The input to a classifier is a set of atomic features that can be determined from the 2D chemical structure of a compound. Depending on the subclass of reaction, the accuracy of prediction for positives (negatives) is 64 to 93% (44 to 92%) when asking if a compound is a substrate and 71 to 98% (50 to 92%) when asking if a compound is a product. Sensitivity analysis reveals that this performance is robust to variations of the training data. Our results suggest that metabolic connectivity can be predicted with reasonable accuracy from the presence or absence of local structural motifs in compounds and their readily calculated atomic features. AVAILABILITY: Classifiers reported here can be used freely for noncommercial purposes via a Java program available upon request. Fangping Mu, Pat J. Unkefer, Clifford J. Unkefer, William S. Hlavacek |
Bioinform. | 4 |
| 2004 | BioNetGen: software for rule-based modeling of signal transduction based on the interactions of molecular domainsabstractBioNetGen allows a user to create a computational model that characterizes the dynamics of a signal transduction system, and that accounts comprehensively and precisely for specified enzymatic activities, potential post-translational modifications and interactions of the domains of signaling molecules. The output defines and parameterizes the network of molecular species that can arise during signaling and provides functions that relate model variables to experimental readouts of interest. Models that can be generated are relevant for rational drug discovery, analysis of proteomic data and mechanistic studies of signal transduction. Michael L. Blinov, James R. Faeder, Byron Goldstein, William S. Hlavacek |
Bioinform. | 4 |