EDBT 2026 Demo / reviewers in the wild / expert
Bernhard O. Palsson
dblp:55/3828 · also Bernhard Ø. Palsson
· DBLP profile ↗
52ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0003-2357-6785ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 51 · 12 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RBC-GEM: A genome-scale metabolic model for systems biology of the human red blood cellabstractAdvancements with cost-effective, high-throughput omics technologies have had a transformative effect on both fundamental and translational research in the medical sciences. These advancements have facilitated a departure from the traditional view of human red blood cells (RBCs) as mere carriers of hemoglobin, devoid of significant biological complexity. Over the past decade, proteomic analyses have identified a growing number of different proteins present within RBCs, enabling systems biology analysis of their physiological functions. Here, we introduce RBC-GEM, one of the most comprehensive, curated genome-scale metabolic reconstructions of a specific human cell type to-date. It was developed through meta-analysis of proteomic data from 29 studies published over the past two decades resulting in an RBC proteome composed of more than 4,600 distinct proteins. Through workflow-guided manual curation, we have compiled the metabolic reactions carried out by this proteome to form a genome-scale metabolic model (GEM) of the RBC. RBC-GEM is hosted on a version-controlled GitHub repository, ensuring adherence to the standardized protocols for metabolic reconstruction quality control and data stewardship principles. RBC-GEM represents a metabolic network is a consisting of 820 genes encoding proteins acting on 1,685 unique metabolites through 2,723 biochemical reactions: a 740% size expansion over its predecessor. We demonstrated the utility of RBC-GEM by creating context-specific proteome-constrained models derived from proteomic data of stored RBCs for 616 blood donors, and classified reactions based on their simulated abundance dependence. This reconstruction as an up-to-date curated GEM can be used for contextualization of data and for the construction of a computational whole-cell models of the human RBC. Zachary B. Haiman, Alicia Key, Angelo D'alessandro, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2024 | Inferred regulons are consistent with regulator binding sequences in E. coliabstractThe transcriptional regulatory network (TRN) of E. coli consists of thousands of interactions between regulators and DNA sequences. Regulons are typically determined either from resource-intensive experimental measurement of functional binding sites, or inferred from analysis of high-throughput gene expression datasets. Recently, independent component analysis (ICA) of RNA-seq compendia has shown to be a powerful method for inferring bacterial regulons. However, it remains unclear to what extent regulons predicted by ICA structure have a biochemical basis in promoter sequences. Here, we address this question by developing machine learning models that predict inferred regulon structures in E. coli based on promoter sequence features. Models were constructed successfully (cross-validation AUROC > = 0.8) for 85% (40/47) of ICA-inferred E. coli regulons. We found that: 1) The presence of a high scoring regulator motif in the promoter region was sufficient to specify regulatory activity in 40% (19/47) of the regulons, 2) Additional features, such as DNA shape and extended motifs that can account for regulator multimeric binding, helped to specify regulon structure for the remaining 60% of regulons (28/47); 3) investigating regulons where initial machine learning models failed revealed new regulator-specific sequence features that improved model accuracy. Finally, we found that strong regulatory binding sequences underlie both the genes shared between ICA-inferred and experimental regulons as well as genes in the E. coli core pan-regulon of Fur. This work demonstrates that the structure of ICA-inferred regulons largely can be understood through the strength of regulator binding sites in promoter regions, reinforcing the utility of top-down inference for regulon discovery. Sizhe Qiu, Xinlong Wan, Yueshan Liang, Cameron R. Lamoureux, Amir Akbari, Bernhard O. Palsson, Daniel C. Zielinski |
PLoS Comput. Biol. | 6 |
| 2024 | iModulonMiner and PyModulon: Software for unsupervised mining of gene expression compendiaabstractPublic gene expression databases are a rapidly expanding resource of organism responses to diverse perturbations, presenting both an opportunity and a challenge for bioinformatics workflows to extract actionable knowledge of transcription regulatory network function. Here, we introduce a five-step computational pipeline, called iModulonMiner, to compile, process, curate, analyze, and characterize the totality of RNA-seq data for a given organism or cell type. This workflow is centered around the data-driven computation of co-regulated gene sets using Independent Component Analysis, called iModulons, which have been shown to have broad applications. As a demonstration, we applied this workflow to generate the iModulon structure of Bacillus subtilis using all high-quality, publicly-available RNA-seq data. Using this structure, we predicted regulatory interactions for multiple transcription factors, identified groups of co-expressed genes that are putatively regulated by undiscovered transcription factors, and predicted properties of a recently discovered single-subunit phage RNA polymerase. We also present a Python package, PyModulon, with functions to characterize, visualize, and explore computed iModulons. The pipeline, available at https://github.com/SBRG/iModulonMiner, can be readily applied to diverse organisms to gain a rapid understanding of their transcriptional regulatory network structure and condition-specific activity. Anand Sastry 0002, Saugat Poudel, Kevin Rychel, Reo Yoo, Cameron R. Lamoureux, Gaoyuan Li, Joshua T. Burrows, Siddharth Chauhan, Zachary B. Haiman, Tahani Al Bulushi, Yara Seif, Bernhard O. Palsson, Daniel C. Zielinski |
PLoS Comput. Biol. | 13 |
| 2024 | StressME: Unified computing framework of Escherichia coli metabolism, gene expression, and stress responsesabstractGeneralist microbes have adapted to a multitude of environmental stresses through their integrated stress response system. Individual stress responses have been quantified by E. coli metabolism and expression (ME) models under thermal, oxidative and acid stress, respectively. However, the systematic quantification of cross-stress & cross-talk among these stress responses remains lacking. Here, we present StressME: the unified stress response model of E. coli combining thermal (FoldME), oxidative (OxidizeME) and acid (AcidifyME) stress responses. StressME is the most up to date ME model for E. coli and it reproduces all published single-stress ME models. Additionally, it includes refined rate constants to improve prediction accuracy for wild-type and stress-evolved strains. StressME revealed certain optimal proteome allocation strategies associated with cross-stress and cross-talk responses. These stress-optimal proteomes were shaped by trade-offs between protective vs. metabolic enzymes; cytoplasmic vs. periplasmic chaperones; and expression of stress-specific proteins. As StressME is tuned to compute metabolic and gene expression responses under mild acid, oxidative, and thermal stresses, it is useful for engineering and health applications. The modular design of our open-source package also facilitates model expansion (e.g., to new stress mechanisms) by the computational biology community. Jiao Zhao, Bernhard O. Palsson, Laurence Yang 0001 |
PLoS Comput. Biol. | 3 |
| 2023 | Deep-learning optimized DEOCSU suite provides an iterable pipeline for accurate ChIP-exo peak callingabstractRecognizing binding sites of DNA-binding proteins is a key factor for elucidating transcriptional regulation in organisms. ChIP-exo enables researchers to delineate genome-wide binding landscapes of DNA-binding proteins with near single base-pair resolution. However, the peak calling step hinders ChIP-exo application since the published algorithms tend to generate false-positive and false-negative predictions. Here, we report the development of DEOCSU (DEep-learning Optimized ChIP-exo peak calling SUite), a novel machine learning-based ChIP-exo peak calling suite. DEOCSU entails the deep convolutional neural network model which was trained with curated ChIP-exo peak data to distinguish the visualized data of bona fide peaks from false ones. Performance validation of the trained deep-learning model indicated its high accuracy, high precision and high recall of over 95%. Applying the new suite to both in-house and publicly available ChIP-exo datasets obtained from bacteria, eukaryotes and archaea revealed an accurate prediction of peaks containing canonical motifs, highlighting the versatility and efficiency of DEOCSU. Furthermore, DEOCSU can be executed on a cloud computing platform or the local environment. With visualization software included in the suite, adjustable options such as the threshold of peak probability, and iterable updating of the pre-trained model, DEOCSU can be optimized for users' specific needs. Ina Bang, Sang-Mok Lee, Seojoung Park, Joon Young Park, Linh Khanh Nong, Ye Gao 0005, Bernhard O. Palsson |
Briefings Bioinform. | 7 |
| 2022 | High-quality genome-scale metabolic network reconstruction of probiotic bacterium Escherichia coli Nissle 1917abstractAbstract Background Escherichia coli Nissle 1917 (EcN) is a probiotic bacterium used to treat various gastrointestinal diseases. EcN is increasingly being used as a chassis for the engineering of advanced microbiome therapeutics. To aid in future engineering efforts, our aim was to construct an updated metabolic model of EcN with extended secondary metabolite representation. Results An updated high-quality genome-scale metabolic model of EcN, iHM1533, was developed based on comparison with 55 E. coli/Shigella reference GEMs and manual curation, including expanded secondary metabolite pathways (enterobactin, salmochelins, aerobactin, yersiniabactin, and colibactin). The model was validated and improved using phenotype microarray data, resulting in an 82.3% accuracy in predicting growth phenotypes on various nutrition sources. Flux variability analysis with previously published 13C fluxomics data validated prediction of the internal central carbon fluxes. A standardised test suite called Memote assessed the quality of iHM1533 to have an overall score of 89%. The model was applied by using constraint-based flux analysis to predict targets for optimisation of secondary metabolite production. Modelling predicted design targets from across amino acid metabolism, carbon metabolism, and other subsystems that are common or unique for influencing the production of various secondary metabolites. Conclusion iHM1533 represents a well-annotated metabolic model of EcN with extended secondary metabolite representation. Phenotype characterisation and the iHM1533 model provide a better understanding of the metabolic capabilities of EcN and will help future metabolic engineering efforts. Max van 't Hof, Omkar S. Mohite, Jonathan Monk, Tilmann Weber, Bernhard O. Palsson, Morten O. A. Sommer |
BMC Bioinform. | 5 |
| 2022 | Positively charged mineral surfaces promoted the accumulation of organic intermediates at the origin of metabolismabstractIdentifying plausible mechanisms for compartmentalization and accumulation of the organic intermediates of early metabolic cycles in primitive cells has been a major challenge in theories of life's origins. Here, we propose a mechanism, where positive membrane potentials elevate the concentration of the organic intermediates. Positive membrane potentials are generated by positively charged surfaces of protocell membranes due to accumulation of transition metals. We find that (i) positive membrane potentials comparable in magnitude to those of modern cells can increase the concentration of the organic intermediates by several orders of magnitude; (ii) generation of large membrane potentials destabilize ion distributions; (iii) violation of electroneutrality is necessary to induce nonzero membrane potentials; and (iv) violation of electroneutrality enhances osmotic pressure and diminishes reaction efficiency, resulting in an evolutionary driving force for the formation of lipid membranes, specialized ion channels, and active transport systems. Amir Akbari, Bernhard O. Palsson |
PLoS Comput. Biol. | 2 |
| 2021 | Optimal dimensionality selection for independent component analysis of transcriptomic dataabstractBACKGROUND: Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. METHODS: We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. RESULTS: We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. CONCLUSIONS: OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism's underlying transcriptional regulatory network. John Luke McConn, Cameron R. Lamoureux, Saugat Poudel, Bernhard O. Palsson, Anand Sastry 0002 |
BMC Bioinform. | 4 |
| 2021 | Bacterial fitness landscapes stratify based on proteome allocation associated with discrete aero-typesabstractThe fitness landscape is a concept commonly used to describe evolution towards optimal phenotypes. It can be reduced to mechanistic detail using genome-scale models (GEMs) from systems biology. We use recently developed GEMs of Metabolism and protein Expression (ME-models) to study the distribution of Escherichia coli phenotypes on the rate-yield plane. We found that the measured phenotypes distribute non-uniformly to form a highly stratified fitness landscape. Systems analysis of the ME-model simulations suggest that this stratification results from discrete ATP generation strategies. Accordingly, we define "aero-types", a phenotypic trait that characterizes how a balanced proteome can achieve a given growth rate by modulating 1) the relative utilization of oxidative phosphorylation, glycolysis, and fermentation pathways; and 2) the differential employment of electron-transport-chain enzymes. This global, quantitative, and mechanistic systems biology interpretation of fitness landscape formed upon proteome allocation offers a fundamental understanding of bacterial physiology and evolution dynamics. Amitesh Anand, Connor A. Olson, Troy E. Sandberg, Ye Gao 0005, Nathan Mih, Bernhard O. Palsson |
PLoS Comput. Biol. | 7 |
| 2021 | MASSpy: Building, simulating, and visualizing dynamic biological models in Python using mass action kineticsabstractMathematical models of metabolic networks utilize simulation to study system-level mechanisms and functions. Various approaches have been used to model the steady state behavior of metabolic networks using genome-scale reconstructions, but formulating dynamic models from such reconstructions continues to be a key challenge. Here, we present the Mass Action Stoichiometric Simulation Python (MASSpy) package, an open-source computational framework for dynamic modeling of metabolism. MASSpy utilizes mass action kinetics and detailed chemical mechanisms to build dynamic models of complex biological processes. MASSpy adds dynamic modeling tools to the COnstraint-Based Reconstruction and Analysis Python (COBRApy) package to provide an unified framework for constraint-based and kinetic modeling of metabolic networks. MASSpy supports high-performance dynamic simulation through its implementation of libRoadRunner: the Systems Biology Markup Language (SBML) simulation engine. Three examples are provided to demonstrate how to use MASSpy: (1) a validation of the MASSpy modeling tool through dynamic simulation of detailed mechanisms of enzyme regulation; (2) a feature demonstration using a workflow for generating ensemble of kinetic models using Monte Carlo sampling to approximate missing numerical values of parameters and to quantify biological uncertainty, and (3) a case study in which MASSpy is utilized to overcome issues that arise when integrating experimental data with the computation of functional states of detailed biological mechanisms. MASSpy represents a powerful tool to address challenges that arise in dynamic modeling of metabolic networks, both at small and large scales. Zachary B. Haiman, Daniel C. Zielinski, Yuko Koike, James T. Yurkovich, Bernhard O. Palsson |
PLoS Comput. Biol. | 5 |
| 2021 | Computation of condition-dependent proteome allocation reveals variability in the macro and micro nutrient requirements for growthabstractSustaining a robust metabolic network requires a balanced and fully functioning proteome. In addition to amino acids, many enzymes require cofactors (coenzymes and engrafted prosthetic groups) to function properly. Extensively validated resource allocation models, such as genome-scale models of metabolism and gene expression (ME-models), have the ability to compute an optimal proteome composition underlying a metabolic phenotype, including the provision of all required cofactors. Here we apply the ME-model for Escherichia coli K-12 MG1655 to computationally examine how environmental conditions change the proteome and its accompanying cofactor usage. We found that: (1) The cofactor requirements computed by the ME-model mostly agree with the standard biomass objective function used in models of metabolism alone (M-models); (2) ME-model computations reveal non-intuitive variability in cofactor use under different growth conditions; (3) An analysis of ME-model predicted protein use in aerobic and anaerobic conditions suggests an enrichment in the use of peroxyl scavenging acids in the proteins used to sustain aerobic growth; (4) The ME-model could describe how limitation in key protein components affect the metabolic state of E. coli. Genome-scale models have thus reached a level of sophistication where they reveal intricate properties of functional proteomes and how they support different E. coli lifestyles. Colton J. Lloyd, Jonathan Monk, Laurence Yang 0001, Ali Ebrahim, Bernhard O. Palsson |
PLoS Comput. Biol. | 5 |
| 2021 | Independent component analysis recovers consistent regulatory signals from disparate datasetsabstractThe availability of bacterial transcriptomes has dramatically increased in recent years. This data deluge could result in detailed inference of underlying regulatory networks, but the diversity of experimental platforms and protocols introduces critical biases that could hinder scalable analysis of existing data. Here, we show that the underlying structure of the E. coli transcriptome, as determined by Independent Component Analysis (ICA), is conserved across multiple independent datasets, including both RNA-seq and microarray datasets. We subsequently combined five transcriptomics datasets into a large compendium containing over 800 expression profiles and discovered that its underlying ICA-based structure was still comparable to that of the individual datasets. With this understanding, we expanded our analysis to over 3,000 E. coli expression profiles and predicted three high-impact regulons that respond to oxidative stress, anaerobiosis, and antibiotic treatment. ICA thus enables deep analysis of disparate data to uncover new insights that were not visible in the individual datasets. Anand Sastry 0002, Alyssa Hu, David Heckmann, Saugat Poudel, Erol S. Kavvas, Bernhard O. Palsson |
PLoS Comput. Biol. | 6 |
| 2020 | Adaptations of Escherichia coli strains to oxidative stress are reflected in properties of their structural proteomesabstractBACKGROUND: The reconstruction of metabolic networks and the three-dimensional coverage of protein structures have reached the genome-scale in the widely studied Escherichia coli K-12 MG1655 strain. The combination of the two leads to the formation of a structural systems biology framework, which we have used to analyze differences between the reactive oxygen species (ROS) sensitivity of the proteomes of sequenced strains of E. coli. As proteins are one of the main targets of oxidative damage, understanding how the genetic changes of different strains of a species relates to its oxidative environment can reveal hypotheses as to why these variations arise and suggest directions of future experimental work. RESULTS: Creating a reference structural proteome for E. coli allows us to comprehensively map genetic changes in 1764 different strains to their locations on 4118 3D protein structures. We use metabolic modeling to predict basal ROS production levels (ROStype) for 695 of these strains, finding that strains with both higher and lower basal levels tend to enrich their proteomes with antioxidative properties, and speculate as to why that is. We computationally assess a strain's sensitivity to an oxidative environment, based on known chemical mechanisms of oxidative damage to protein groups, defined by their localization and functionality. Two general groups - metalloproteins and periplasmic proteins - show enrichment of their antioxidative properties between the 695 strains with a predicted ROStype as well as 116 strains with an assigned pathotype. Specifically, proteins that a) utilize a molybdenum ion as a cofactor and b) are involved in the biogenesis of fimbriae show intriguing protective properties to resist oxidative damage. Overall, these findings indicate that a strain's sensitivity to oxidative damage can be elucidated from the structural proteome, though future experimental work is needed to validate our model assumptions and findings. CONCLUSION: We thus demonstrate that structural systems biology enables a proteome-wide, computational assessment of changes to atomic-level physicochemical properties and of oxidative damage mechanisms for multiple strains in a species. This integrative approach opens new avenues to study adaptation to a particular environment based on physiological properties predicted from sequence alone. Nathan Mih, Jonathan Monk, Edward Catoiu, David Heckmann, Laurence Yang 0001, Bernhard O. Palsson |
BMC Bioinform. | 7 |
| 2020 | Machine learning with random subspace ensembles identifies antimicrobial resistance determinants from pan-genomes of three pathogensabstractThe evolution of antimicrobial resistance (AMR) poses a persistent threat to global public health.Sequencing efforts have already yielded genome sequences for thousands of resistant microbial isolates and require robust computational tools to systematically elucidate the genetic basis for AMR.Here, we present a generalizable machine learning workflow for identifying genetic features driving AMR based on constructing reference strain-agnostic pan-genomes and training random subspace ensembles (RSEs).This workflow was applied to the resistance profiles of 14 antimicrobials across three urgent threat pathogens encompassing 288 Staphylococcus aureus, 456 Pseudomonas aeruginosa, and 1588 Escherichia coli genomes.We find that feature selection by RSE detects known AMR associations more reliably than common statistical tests and previous ensemble approaches, identifying a total of 45 known AMR-conferring genes and alleles across the three organisms, as well as 25 candidate associations backed by domain-level annotations.Furthermore, we find that results from the RSE approach are consistent with existing understanding of fluoroquinolone (FQ) resistance due to mutations in the main drug targets, gyrA and parC, in all three organisms, and suggest the mutational landscape of those genes with respect to FQ resistance is simple.As larger datasets become available, we expect this approach to more reliably predict AMR determinants for a wider range of microbial pathogens. Author summaryAntimicrobial resistance remains a persistent threat to global public health, with 700,000 deaths each year attributable to resistant bacterial infections.The falling cost of genome sequencing offers an avenue for rapidly predicting and elucidating the resistance profiles of infectious isolates, which is necessary for the design of more effective antimicrobial therapies from existing drugs.As such, clinical surveillance programs have already yielded sequences for thousands of distinct, resistant strains of most major pathogens.Here, we have developed a workflow for training machine learning models capable of not just predicting resistance profiles from genome sequences, but also Jason C. Hyun, Erol S. Kavvas, Jonathan Monk, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2019 | Estimating Cellular Goals from High-Dimensional Biological DataabstractOptimization-based models have been used to predict cellular behavior for over 25 years. The constraints in these models are derived from genome annotations, measured macromolecular composition of cells, and by measuring the cell's growth rate and metabolism in different conditions. The cellular goal (the optimization problem that the cell is trying to solve) can be challenging to derive experimentally for many organisms, including human or mammalian cells, which have complex metabolic capabilities and are not well understood. Existing approaches to learning goals from data include (a) estimating a linear objective function, or (b) estimating linear constraints that model complex biochemical reactions and constrain the cell's operation. The latter approach is important because often the known reactions are not enough to explain observations; therefore, there is a need to extend automatically the model complexity by learning new reactions. However, this leads to nonconvex optimization problems, and existing tools cannot scale to realistically large metabolic models. Hence, constraint estimation is still used sparingly despite its benefits for modeling cell metabolism, which is important for developing novel antimicrobials against pathogens, discovering cancer drug targets, and producing value-added chemicals. Here, we develop the first approach to estimating constraint reactions from data that can scale to realistically large metabolic models. Previous tools were used on problems having less than 75 reactions and 60 metabolites, which limits real-life-size applications. We perform extensive experiments using 75 large-scale metabolic network models for different organisms (including bacteria, yeasts, and mammals) and show that our algorithm can recover cellular constraint reactions. The recovered constraints enable accurate prediction of metabolic states in hundreds of growth environments not seen in training data, and we recover useful cellular goals even when some measurements are missing. Laurence Yang 0001, Michael A. Saunders, Jean-Christophe Lachance, Bernhard O. Palsson, José Bento 0001 |
KDD | 4 |
| 2019 | Laboratory evolution reveals a two-dimensional rate-yield tradeoff in microbial metabolismabstractGrowth rate and yield are fundamental features of microbial growth. However, we lack a mechanistic and quantitative understanding of the rate-yield relationship. Studies pairing computational predictions with experiments have shown the importance of maintenance energy and proteome allocation in explaining rate-yield tradeoffs and overflow metabolism. Recently, adaptive evolution experiments of Escherichia coli reveal a phenotypic diversity beyond what has been explained using simple models of growth rate versus yield. Here, we identify a two-dimensional rate-yield tradeoff in adapted E. coli strains where the dimensions are (A) a tradeoff between growth rate and yield and (B) a tradeoff between substrate (glucose) uptake rate and growth yield. We employ a multi-scale modeling approach, combining a previously reported coarse-grained small-scale proteome allocation model with a fine-grained genome-scale model of metabolism and gene expression (ME-model), to develop a quantitative description of the full rate-yield relationship for E. coli K-12 MG1655. The multi-scale analysis resolves the complexity of ME-model which hindered its practical use in proteome complexity analysis, and provides a mechanistic explanation of the two-dimensional tradeoff. Further, the analysis identifies modifications to the P/O ratio and the flux allocation between glycolysis and pentose phosphate pathway (PPP) as potential mechanisms that enable the tradeoff between glucose uptake rate and growth yield. Thus, the rate-yield tradeoffs that govern microbial adaptation to new environments are more complex than previously reported, and they can be understood in mechanistic detail using a multi-scale modeling approach. Chuankai Cheng, Edward J. O'Brien, Douglas McCloskey, Jose Utrilla, Connor A. Olson, Ryan A. LaCroix, Troy E. Sandberg, Adam M. Feist, Bernhard O. Palsson, Zachary A. King |
PLoS Comput. Biol. | 9 |
| 2019 | Genome-scale model of metabolism and gene expression provides a multi-scale description of acid stress responses in Escherichia coliabstractResponse to acid stress is critical for Escherichia coli to successfully complete its life-cycle by passing through the stomach to colonize the digestive tract. To develop a fundamental understanding of this response, we established a molecular mechanistic description of acid stress mitigation responses in E. coli and integrated them with a genome-scale model of its metabolism and macromolecular expression (ME-model). We considered three known mechanisms of acid stress mitigation: 1) change in membrane lipid fatty acid composition, 2) change in periplasmic protein stability over external pH and periplasmic chaperone protection mechanisms, and 3) change in the activities of membrane proteins. After integrating these mechanisms into an established ME-model, we could simulate their responses in the context of other cellular processes. We validated these simulations using RNA sequencing data obtained from five E. coli strains grown under external pH ranging from 5.5 to 7.0. We found: i) that for the differentially expressed genes accounted for in the ME-model, 80% of the upregulated genes were correctly predicted by the ME-model, and ii) that these genes are mainly involved in translation processes (45% of genes), membrane proteins and related processes (18% of genes), amino acid metabolism (12% of genes), and cofactor and prosthetic group biosynthesis (8% of genes). We also demonstrated several intervention strategies on acid tolerance that can be simulated by the ME-model. We thus established a quantitative framework that describes, on a genome-scale, the acid stress mitigation response of E. coli that has both scientific and practical uses. Bin Du 0008, Laurence Yang 0001, Colton J. Lloyd, Bernhard O. Palsson |
PLoS Comput. Biol. | 5 |
| 2019 | BOFdat: Generating biomass objective functions for genome-scale metabolic models from experimental dataabstractGenome-scale metabolic models (GEMs) are mathematically structured knowledge bases of metabolism that provide phenotypic predictions from genomic information. GEM-guided predictions of growth phenotypes rely on the accurate definition of a biomass objective function (BOF) that is designed to include key cellular biomass components such as the major macromolecules (DNA, RNA, proteins), lipids, coenzymes, inorganic ions and species-specific components. Despite its importance, no standardized computational platform is currently available to generate species-specific biomass objective functions in a data-driven, unbiased fashion. To fill this gap in the metabolic modeling software ecosystem, we implemented BOFdat, a Python package for the definition of a Biomass Objective Function from experimental data. BOFdat has a modular implementation that divides the BOF definition process into three independent modules defined here as steps: 1) the coefficients for major macromolecules are calculated, 2) coenzymes and inorganic ions are identified and their stoichiometric coefficients estimated, 3) the remaining species-specific metabolic biomass precursors are algorithmically extracted in an unbiased way from experimental data. We used BOFdat to reconstruct the BOF of the Escherichia coli model iML1515, a gold standard in the field. The BOF generated by BOFdat resulted in the most concordant biomass composition, growth rate, and gene essentiality prediction accuracy when compared to other methods. Installation instructions for BOFdat are available in the documentation and the source code is available on GitHub (https://github.com/jclachance/BOFdat). Jean-Christophe Lachance, Colton J. Lloyd, Jonathan Monk, Laurence Yang 0001, Anand Sastry 0002, Yara Seif, Bernhard O. Palsson, Sébastien Rodrigue, Adam M. Feist, Zachary A. King, Pierre-Étienne Jacques |
PLoS Comput. Biol. | 7 |
| 2019 | A computational knowledge-base elucidates the response of Staphylococcus aureus to different media typesabstractS. aureus is classified as a serious threat pathogen and is a priority that guides the discovery and development of new antibiotics. Despite growing knowledge of S. aureus metabolic capabilities, our understanding of its systems-level responses to different media types remains incomplete. Here, we develop a manually reconstructed genome-scale model (GEM-PRO) of metabolism with 3D protein structures for S. aureus USA300 str. JE2 containing 854 genes, 1,440 reactions, 1,327 metabolites and 673 3-dimensional protein structures. Computations were in 85% agreement with gene essentiality data from random barcode transposon site sequencing (RB-TnSeq) and 68% agreement with experimental physiological data. Comparisons of computational predictions with experimental observations highlight: 1) cases of non-essential biomass precursors; 2) metabolic genes subject to transcriptional regulation involved in Staphyloxanthin biosynthesis; 3) the essentiality of purine and amino acid biosynthesis in synthetic physiological media; and 4) a switch to aerobic fermentation upon exposure to extracellular glucose elucidated as a result of integrating time-course of quantitative exo-metabolomics data. An up-to-date GEM-PRO thus serves as a knowledge-based platform to elucidate S. aureus' metabolic response to its environment. Yara Seif, Jonathan Monk, Nathan Mih, Hannah Tsunemoto, Saugat Poudel, Cristal Zuñiga, Jared Broddrick, Karsten Zengler, Bernhard O. Palsson |
PLoS Comput. Biol. | 9 |
| 2019 | Systems-level analysis of NalD mutation, a recurrent driver of rapid drug resistance in acute Pseudomonas aeruginosa infectionabstractPseudomonas aeruginosa, a main cause of human infection, can gain resistance to the antibiotic aztreonam through a mutation in NalD, a transcriptional repressor of cellular efflux. Here we combine computational analysis of clinical isolates, transcriptomics, metabolic modeling and experimental validation to find a strong association between NalD mutations and resistance to aztreonam-as well as resistance to other antibiotics-across P. aeruginosa isolated from different patients. A detailed analysis of one patient's timeline shows how this mutation can emerge in vivo and drive rapid evolution of resistance while the patient received cancer treatment, a bone marrow transplantation, and antibiotics up to the point of causing the patient's death. Transcriptomics analysis confirmed the primary mechanism of NalD action-a loss-of-function mutation that caused constitutive overexpression of the MexAB-OprM efflux system-which lead to aztreonam resistance but, surprisingly, had no fitness cost in the absence of the antibiotic. We constrained a genome-scale metabolic model using the transcriptomics data to investigate changes beyond the primary mechanism of resistance, including adaptations in major metabolic pathways and membrane transport concurrent with aztreonam resistance, which may explain the lack of a fitness cost. We propose that metabolic adaptations may allow resistance mutations to endure in the absence of antibiotics and could be targeted by future therapies against antibiotic resistant pathogens. Jinyuan Yan, Henri Estanbouli, Chen Liao, Wook Kim, Jonathan Monk, Rayees Rahman, Mini Kamboj, Bernhard O. Palsson, Wei-Gang Qiu, João B. Xavier |
PLoS Comput. Biol. | 8 |
| 2018 | ssbio: a Python framework for structural systems biologyabstractSummary: Working with protein structures at the genome-scale has been challenging in a variety of ways. Here, we present ssbio, a Python package that provides a framework to easily work with structural information in the context of genome-scale network reconstructions, which can contain thousands of individual proteins. The ssbio package provides an automated pipeline to construct high quality genome-scale models with protein structures (GEM-PROs), wrappers to popular third-party programs to compute associated protein properties, and methods to visualize and annotate structures directly in Jupyter notebooks, thus lowering the barrier of linking 3D structural data with established systems workflows. Availability and implementation: ssbio is implemented in Python and available to download under the MIT license at http://github.com/SBRG/ssbio. Documentation and Jupyter notebook tutorials are available at http://ssbio.readthedocs.io/en/latest/. Interactive notebooks can be launched using Binder at https://mybinder.org/v2/gh/SBRG/ssbio/master?filepath=Binder.ipynb. Supplementary information: Supplementary data are available at Bioinformatics online. Nathan Mih, Elizabeth Brunk, Edward Catoiu, Anand Sastry 0002, Erol S. Kavvas, Jonathan Monk, Bernhard O. Palsson |
Bioinform. | 9 |
| 2018 | Functional interrogation of Plasmodium genus metabolism identifies species- and stage-specific differences in nutrient essentiality and drug targetingabstractSeveral antimalarial drugs exist, but differences between life cycle stages among malaria species pose challenges for developing more effective therapies.To understand the diversity among stages and species, we reconstructed genome-scale metabolic models (GeMMs) of metabolism for five life cycle stages and five species of Plasmodium spanning the blood, transmission, and mosquito stages.The stage-specific models of Plasmodium falciparum uncovered stage-dependent changes in central carbon metabolism and predicted potential targets that could affect several life cycle stages.The species-specific models further highlight differences between experimental animal models and the humaninfecting species.Comparisons between human-and rodent-infecting species revealed differences in thiamine (vitamin B1), choline, and pantothenate (vitamin B5) metabolism.Thus, we show that genome-scale analysis of multiple stages and species of Plasmodium can prioritize potential drug targets that could be both anti-malarials and transmission blocking agents, in addition to guiding translation from non-human experimental disease models. Author summaryMalaria kills nearly one-half million people a year and over 1 billion people are at risk of becoming infected by the parasite.Plasmodial infections are difficult to treat for a myriad of reasons, but the ability of the organism to remain latent in hosts and the complex life cycles greatly contributed to the difficulty in treat malaria.Genome-scale metabolic models (GeMMs) enable hierarchical integration of disparate data types into a framework Alyaa M. Abdel-Haleem, Hooman Hefzi, Katsuhiko Mineta, Xin Gao 0001, Takashi Gojobori, Bernhard O. Palsson, Nathan E. Lewis, Neema Jamshidi |
PLoS Comput. Biol. | 6 |
| 2018 | COBRAme: A computational framework for genome-scale models of metabolism and gene expressionabstractGenome-scale models of metabolism and macromolecular expression (ME-models) explicitly compute the optimal proteome composition of a growing cell. ME-models expand upon the well-established genome-scale models of metabolism (M-models), and they enable a new fundamental understanding of cellular growth. ME-models have increased predictive capabilities and accuracy due to their inclusion of the biosynthetic costs for the machinery of life, but they come with a significant increase in model size and complexity. This challenge results in models which are both difficult to compute and challenging to understand conceptually. As a result, ME-models exist for only two organisms (Escherichia coli and Thermotoga maritima) and are still used by relatively few researchers. To address these challenges, we have developed a new software framework called COBRAme for building and simulating ME-models. It is coded in Python and built on COBRApy, a popular platform for using M-models. COBRAme streamlines computation and analysis of ME-models. It provides tools to simplify constructing and editing ME-models to enable ME-model reconstructions for new organisms. We used COBRAme to reconstruct a condensed E. coli ME-model called iJL1678b-ME. This reformulated model gives functionally identical solutions to previous E. coli ME-models while using 1/6 the number of free variables and solving in less than 10 minutes, a marked improvement over the 6 hour solve time of previous ME-model formulations. Errors in previous ME-models were also corrected leading to 52 additional genes that must be expressed in iJL1678b-ME to grow aerobically in glucose minimal in silico media. This manuscript outlines the architecture of COBRAme and demonstrates how ME-models can be created, modified, and shared most efficiently using the new software framework. Colton J. Lloyd, Ali Ebrahim, Laurence Yang 0001, Zachary A. King, Edward Catoiu, Edward J. O'Brien, Joanne K. Liu, Bernhard O. Palsson |
PLoS Comput. Biol. | 8 |
| 2018 | Network-level allosteric effects are elucidated by detailing how ligand-binding events modulate utilization of catalytic potentialsabstractAllosteric regulation has traditionally been described by mathematically-complex allosteric rate laws in the form of ratios of polynomials derived from the application of simplifying kinetic assumptions. Alternatively, an approach that explicitly describes all known ligand-binding events requires no simplifying assumptions while allowing for the computation of enzymatic states. Here, we employ such a modeling approach to examine the "catalytic potential" of an enzyme-an enzyme's capacity to catalyze a biochemical reaction. The catalytic potential is the fundamental result of multiple ligand-binding events that represents a "tug of war" among the various regulators and substrates within the network. This formalism allows for the assessment of interacting allosteric enzymes and development of a network-level understanding of regulation. We first define the catalytic potential and use it to characterize the response of three key kinases (hexokinase, phosphofructokinase, and pyruvate kinase) in human red blood cell glycolysis to perturbations in ATP utilization. Next, we examine the sensitivity of the catalytic potential by using existing personalized models, finding that the catalytic potential allows for the identification of subtle but important differences in how individuals respond to such perturbations. Finally, we explore how the catalytic potential can help to elucidate how enzymes work in tandem to maintain a homeostatic state. Taken together, this work provides an interpretation and visualization of the dynamic interactions and network-level effects of interacting allosteric enzymes. James T. Yurkovich, Miguel A. Alcantar, Zachary B. Haiman, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2017 | Machine learning in computational biology to accelerate high-throughput protein expressionabstractMOTIVATION: The Human Protein Atlas (HPA) enables the simultaneous characterization of thousands of proteins across various tissues to pinpoint their spatial location in the human body. This has been achieved through transcriptomics and high-throughput immunohistochemistry-based approaches, where over 40 000 unique human protein fragments have been expressed in E. coli. These datasets enable quantitative tracking of entire cellular proteomes and present new avenues for understanding molecular-level properties influencing expression and solubility. RESULTS: Combining computational biology and machine learning identifies protein properties that hinder the HPA high-throughput antibody production pipeline. We predict protein expression and solubility with accuracies of 70% and 80%, respectively, based on a subset of key properties (aromaticity, hydropathy and isoelectric point). We guide the selection of protein fragments based on these characteristics to optimize high-throughput experimentation. AVAILABILITY AND IMPLEMENTATION: We present the machine learning workflow as a series of IPython notebooks hosted on GitHub (https://github.com/SBRG/Protein_ML). The workflow can be used as a template for analysis of further expression and solubility datasets. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anand Sastry 0002, Jonathan Monk, Hanna Tegel, Mathias Uhlen, Bernhard O. Palsson, Johan Rockberg, Elizabeth Brunk |
Bioinform. | 5 |
| 2017 | Biomarkers are used to predict quantitative metabolite concentration profiles in human red blood cellsabstractDeep-coverage metabolomic profiling has revealed a well-defined development of metabolic decay in human red blood cells (RBCs) under cold storage conditions.A set of extracellular biomarkers has been recently identified that reliably defines the qualitative state of the metabolic network throughout this metabolic decay process.Here, we extend the utility of these biomarkers by using them to quantitatively predict the concentrations of other metabolites in the red blood cell.We are able to accurately predict the concentration profile of 84 of the 91 (92%) measured metabolites (p < 0.05) in RBC metabolism using only measurements of these five biomarkers.The median of prediction errors (symmetric mean absolute percent error) across all metabolites was 13%.The ability to predict numerous metabolite concentrations from a simple set of biomarkers offers the potential for the development of a powerful workflow that could be used to evaluate the metabolic state of a biological system using a minimal set of measurements. Author summaryWhile deep-coverage omics data sets are allowing for more complete characterization of biological systems, there has been a concerted effort to identify a subset of measurements that are representative of qualitative network-level behavior.For some systems-like the human red blood cell (RBC)-such biomarkers have already been identified.Using the concentration profiles of these biomarkers as input to a statistical model, we predict quantitative concentration profiles of other metabolites in the RBC network.These results demonstrate that if good biomarkers are available for a biological system, it is possible to use these measurements to gain insight into the quantitative state of the rest of the network. James T. Yurkovich, Laurence Yang 0001, Bernhard O. Palsson |
PLoS Comput. Biol. | 3 |
| 2016 | solveME: fast and reliable solution of nonlinear ME modelsabstractBACKGROUND: Genome-scale models of metabolism and macromolecular expression (ME) significantly expand the scope and predictive capabilities of constraint-based modeling. ME models present considerable computational challenges: they are much (>30 times) larger than corresponding metabolic reconstructions (M models), are multiscale, and growth maximization is a nonlinear programming (NLP) problem, mainly due to macromolecule dilution constraints. RESULTS: Here, we address these computational challenges. We develop a fast and numerically reliable solution method for growth maximization in ME models using a quad-precision NLP solver (Quad MINOS). Our method was up to 45 % faster than binary search for six significant digits in growth rate. We also develop a fast, quad-precision flux variability analysis that is accelerated (up to 60× speedup) via solver warm-starts. Finally, we employ the tools developed to investigate growth-coupled succinate overproduction, accounting for proteome constraints. CONCLUSIONS: Just as genome-scale metabolic reconstructions have become an invaluable tool for computational and systems biologists, we anticipate that these fast and numerically reliable ME solution methods will accelerate the wide-spread adoption of ME models for researchers in these fields. Laurence Yang 0001, Ding Ma 0003, Ali Ebrahim, Colton J. Lloyd, Michael A. Saunders, Bernhard O. Palsson |
BMC Bioinform. | 6 |
| 2016 | Solving Puzzles With Missing Pieces: The Power of Systems Biology [Point of View]abstractLife is a program written in DNA. Starting in 1995, genome sequences detailing this program have ushered in a new point of view in biology: a true systems-level, or genome-scale, perspective. The genome sequence for an organism is analogous to having a component list for a circuit, except that many connections between components, and even some of the component functions themselves, are unknown. So how do we solve a puzzle when there are pieces missing? Enter systems biology. The combination of full genome sequences with over half a century of research in genetics, molecular biology, and biochemistry has enabled the genome-scale reconstruction of networks underlying well-studied cellular functions, such as metabolism. A quality-controlled reconstruction process effectively produces a circuit diagram of the metabolic network encoded in an organism's genome that can be modeled mathematically. Thus, a first principles ``bottom-up'' approach to systems biology rooted in fundamental mechanisms has arisen, and the quest to reveal the program that DNA encodes is underway. This article will familiarize you with some of the engineering concepts, methods, and applications in systems biology. James T. Yurkovich, Bernhard O. Palsson |
Proc. IEEE | 2 |
| 2016 | A Multi-scale Computational Platform to Mechanistically Assess the Effect of Genetic Variation on Drug Responses in Human Erythrocyte MetabolismabstractProgress in systems medicine brings promise to addressing patient heterogeneity and individualized therapies. Recently, genome-scale models of metabolism have been shown to provide insight into the mechanistic link between drug therapies and systems-level off-target effects while being expanded to explicitly include the three-dimensional structure of proteins. The integration of these molecular-level details, such as the physical, structural, and dynamical properties of proteins, notably expands the computational description of biochemical network-level properties and the possibility of understanding and predicting whole cell phenotypes. In this study, we present a multi-scale modeling framework that describes biological processes which range in scale from atomistic details to an entire metabolic network. Using this approach, we can understand how genetic variation, which impacts the structure and reactivity of a protein, influences both native and drug-induced metabolic states. As a proof-of-concept, we study three enzymes (catechol-O-methyltransferase, glucose-6-phosphate dehydrogenase, and glyceraldehyde-3-phosphate dehydrogenase) and their respective genetic variants which have clinically relevant associations. Using all-atom molecular dynamic simulations enables the sampling of long timescale conformational dynamics of the proteins (and their mutant variants) in complex with their respective native metabolites or drug molecules. We find that changes in a protein's structure due to a mutation influences protein binding affinity to metabolites and/or drug molecules, and inflicts large-scale changes in metabolism. Nathan Mih, Elizabeth Brunk, Aarash Bordbar, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2016 | Quantification and Classification of E. coli Proteome Utilization and Unused Protein Costs across EnvironmentsabstractThe costs and benefits of protein expression are balanced through evolution. Expression of un-utilized protein (that have no benefits in the current environment) incurs a quantifiable fitness costs on cellular growth rates; however, the magnitude and variability of un-utilized protein expression in natural settings is unknown, largely due to the challenge in determining environment-specific proteome utilization. We address this challenge using absolute and global proteomics data combined with a recently developed genome-scale model of Escherichia coli that computes the environment-specific cost and utility of the proteome on a per gene basis. We show that nearly half of the proteome mass is unused in certain environments and accounting for the cost of this unused protein expression explains >95% of the variance in growth rates of Escherichia coli across 16 distinct environments. Furthermore, reduction in unused protein expression is shown to be a common mechanism to increase cellular growth rates in adaptive evolution experiments. Classification of the unused protein reveals that the unused protein encodes several nutrient- and stress- preparedness functions, which may convey fitness benefits in varying environments. Thus, unused protein expression is the source of large and pervasive fitness costs that may provide the benefit of hedging against environmental change. Edward J. O'Brien, Jose Utrilla, Bernhard O. Palsson |
PLoS Comput. Biol. | 3 |
| 2015 | JSBML 1.0: providing a smorgasbord of options to encode systems biology modelsabstractUNLABELLED: JSBML, the official pure Java programming library for the Systems Biology Markup Language (SBML) format, has evolved with the advent of different modeling formalisms in systems biology and their ability to be exchanged and represented via extensions of SBML. JSBML has matured into a major, active open-source project with contributions from a growing, international team of developers who not only maintain compatibility with SBML, but also drive steady improvements to the Java interface and promote ease-of-use with end users. AVAILABILITY AND IMPLEMENTATION: Source code, binaries and documentation for JSBML can be freely obtained under the terms of the LGPL 2.1 from the website http://sbml.org/Software/JSBML. More information about JSBML can be found in the user guide at http://sbml.org/Software/JSBML/docs/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nicolas Rodriguez 0001, Alex Thomas, Leandro H. Watanabe, Ibrahim Y. Vazirabad, Victor Kofia, Harold F. Gómez, Florian Mittag, Jakob Matthes, Jan Rudolph, Finja Wrzodek, Eugen Netz, Alexander Diamantikos, Johannes Eichner, Roland Keller, Clemens Wrzodek, Sebastian Fröhlich, Nathan E. Lewis, Chris J. Myers, Nicolas Le Novère, Bernhard O. Palsson, Michael Hucka, Andreas Dräger |
Bioinform. | 20 |
| 2015 | Escher: A Web Application for Building, Sharing, and Embedding Data-Rich Visualizations of Biological PathwaysabstractEscher is a web application for visualizing data on biological pathways. Three key features make Escher a uniquely effective tool for pathway visualization. First, users can rapidly design new pathway maps. Escher provides pathway suggestions based on user data and genome-scale models, so users can draw pathways in a semi-automated way. Second, users can visualize data related to genes or proteins on the associated reactions and pathways, using rules that define which enzymes catalyze each reaction. Thus, users can identify trends in common genomic data types (e.g. RNA-Seq, proteomics, ChIP)--in conjunction with metabolite- and reaction-oriented data types (e.g. metabolomics, fluxomics). Third, Escher harnesses the strengths of web technologies (SVG, D3, developer tools) so that visualizations can be rapidly adapted, extended, shared, and embedded. This paper provides examples of each of these features and explains how the development approach used for Escher can be used to guide the development of future visualization tools. Zachary A. King, Andreas Dräger, Ali Ebrahim, Nikolaus Sonnenschein, Nathan E. Lewis, Bernhard O. Palsson |
PLoS Comput. Biol. | 6 |
| 2014 | A Systems Approach to Predict Oncometabolites via Context-Specific Genome-Scale Metabolic NetworksabstractAltered metabolism in cancer cells has been viewed as a passive response required for a malignant transformation. However, this view has changed through the recently described metabolic oncogenic factors: mutated isocitrate dehydrogenases (IDH), succinate dehydrogenase (SDH), and fumarate hydratase (FH) that produce oncometabolites that competitively inhibit epigenetic regulation. In this study, we demonstrate in silico predictions of oncometabolites that have the potential to dysregulate epigenetic controls in nine types of cancer by incorporating massive scale genetic mutation information (collected from more than 1,700 cancer genomes), expression profiling data, and deploying Recon 2 to reconstruct context-specific genome-scale metabolic models. Our analysis predicted 15 compounds and 24 substructures of potential oncometabolites that could result from the loss-of-function and gain-of-function mutations of metabolic enzymes, respectively. These results suggest a substantial potential for discovering unidentified oncometabolites in various forms of cancers. Hojung Nam, Miguel Campodonico, Aarash Bordbar, Daniel R. Hyduke, Sangwoo Kim, Daniel C. Zielinski, Bernhard O. Palsson |
PLoS Comput. Biol. | 7 |
| 2013 | GIM3E: condition-specific models of cellular metabolism developed from metabolomics and expression dataabstractMOTIVATION: Genome-scale metabolic models have been used extensively to investigate alterations in cellular metabolism. The accuracy of these models to represent cellular metabolism in specific conditions has been improved by constraining the model with omics data sources. However, few practical methods for integrating metabolomics data with other omics data sources into genome-scale models of metabolism have been developed. RESULTS: GIM(3)E (Gene Inactivation Moderated by Metabolism, Metabolomics and Expression) is an algorithm that enables the development of condition-specific models based on an objective function, transcriptomics and cellular metabolomics data. GIM(3)E establishes metabolite use requirements with metabolomics data, uses model-paired transcriptomics data to find experimentally supported solutions and provides calculations of the turnover (production/consumption) flux of metabolites. GIM(3)E was used to investigate the effects of integrating additional omics datasets to create increasingly constrained solution spaces of Salmonella Typhimurium metabolism during growth in both rich and virulence media. This integration proved to be informative and resulted in a requirement of additional active reactions (12 in each case) or metabolites (26 or 29, respectively). The addition of constraints from transcriptomics also impacted the allowed solution space, and the cellular metabolites with turnover fluxes that were necessarily altered by the change in conditions increased from 118 to 271 of 1397. AVAILABILITY: GIM(3)E has been implemented in Python and requires a COBRApy 0.2.x. The algorithm and sample data described here are freely available at: http://opencobra.sourceforge.net/ CONTACTS: [email protected] Brian J. Schmidt, Ali Ebrahim, Thomas O. Metz, Joshua N. Adkins, Bernhard O. Palsson, Daniel R. Hyduke |
Bioinform. | 5 |
| 2010 | BiGG: a Biochemical Genetic and Genomic knowledgebase of large scale metabolic reconstructionsabstractBACKGROUND: Genome-scale metabolic reconstructions under the Constraint Based Reconstruction and Analysis (COBRA) framework are valuable tools for analyzing the metabolic capabilities of organisms and interpreting experimental data. As the number of such reconstructions and analysis methods increases, there is a greater need for data uniformity and ease of distribution and use. DESCRIPTION: We describe BiGG, a knowledgebase of Biochemically, Genetically and Genomically structured genome-scale metabolic network reconstructions. BiGG integrates several published genome-scale metabolic networks into one resource with standard nomenclature which allows components to be compared across different organisms. BiGG can be used to browse model content, visualize metabolic pathway maps, and export SBML files of the models for further analysis by external software packages. Users may follow links from BiGG to several external databases to obtain additional information on genes, proteins, reactions, metabolites and citations of interest. CONCLUSIONS: BiGG addresses a need in the systems biology community to have access to high quality curated metabolic models and reconstructions. It is freely available for academic use at http://bigg.ucsd.edu. Jan Schellenberger, Junyoung O. Park, Tom M. Conrad, Bernhard O. Palsson |
BMC Bioinform. | 4 |
| 2010 | Drug Off-Target Effects Predicted Using Structural Analysis in the Context of a Metabolic Network ModelabstractRecent advances in structural bioinformatics have enabled the prediction of protein-drug off-targets based on their ligand binding sites. Concurrent developments in systems biology allow for prediction of the functional effects of system perturbations using large-scale network models. Integration of these two capabilities provides a framework for evaluating metabolic drug response phenotypes in silico. This combined approach was applied to investigate the hypertensive side effect of the cholesteryl ester transfer protein inhibitor torcetrapib in the context of human renal function. A metabolic kidney model was generated in which to simulate drug treatment. Causal drug off-targets were predicted that have previously been observed to impact renal function in gene-deficient patients and may play a role in the adverse side effects observed in clinical trials. Genetic risk factors for drug treatment were also predicted that correspond to both characterized and unknown renal metabolic disorders as well as cryptic genetic deficiencies that are not expected to exhibit a renal disorder phenotype except under drug treatment. This study represents a novel integration of structural and systems biology and a first step towards computational systems medicine. The methodology introduced herein has important implications for drug development and personalized medicine. Roger L. Chang, Li Xie 0002, Lei Xie 0006, Philip E. Bourne, Bernhard O. Palsson |
PLoS Comput. Biol. | 5 |
| 2009 | Functional States of the Genome-Scale Escherichia Coli Transcriptional Regulatory SystemabstractA transcriptional regulatory network (TRN) constitutes the collection of regulatory rules that link environmental cues to the transcription state of a cell's genome. We recently proposed a matrix formalism that quantitatively represents a system of such rules (a transcriptional regulatory system [TRS]) and allows systemic characterization of TRS properties. The matrix formalism not only allows the computation of the transcription state of the genome but also the fundamental characterization of the input-output mapping that it represents. Furthermore, a key advantage of this "pseudo-stoichiometric" matrix formalism is its ability to easily integrate with existing stoichiometric matrix representations of signaling and metabolic networks. Here we demonstrate for the first time how this matrix formalism is extendable to large-scale systems by applying it to the genome-scale Escherichia coli TRS. We analyze the fundamental subspaces of the regulatory network matrix (R) to describe intrinsic properties of the TRS. We further use Monte Carlo sampling to evaluate the E. coli transcription state across a subset of all possible environments, comparing our results to published gene expression data as validation. Finally, we present novel in silico findings for the E. coli TRS, including (1) a gene expression correlation matrix delineating functional motifs; (2) sets of gene ontologies for which regulatory rules governing gene transcription are poorly understood and which may direct further experimental characterization; and (3) the appearance of a distributed TRN structure, which is in stark contrast to the more hierarchical organization of metabolic networks. Erwin P. Gianchandani, Andrew R. Joyce, Bernhard O. Palsson, Jason A. Papin |
PLoS Comput. Biol. | 3 |
| 2009 | Identification of Potential Pathway Mediation Targets in Toll-like Receptor SignalingabstractRecent advances in reconstruction and analytical methods for signaling networks have spurred the development of large-scale models that incorporate fully functional and biologically relevant features. An extended reconstruction of the human Toll-like receptor signaling network is presented herein. This reconstruction contains an extensive complement of kinases, phosphatases, and other associated proteins that mediate the signaling cascade along with a delineation of their associated chemical reactions. A computational framework based on the methods of large-scale convex analysis was developed and applied to this network to characterize input-output relationships. The input-output relationships enabled significant modularization of the network into ten pathways. The analysis identified potential candidates for inhibitory mediation of TLR signaling with respect to their specificity and potency. Subsequently, we were able to identify eight novel inhibition targets through constraint-based modeling methods. The results of this study are expected to yield meaningful avenues for further research in the task of mediating the Toll-like receptor signaling network and its effects. Fan Li 0002, Ines Thiele, Neema Jamshidi, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2009 | Genome-Scale Reconstruction of Escherichia coli's Transcriptional and Translational Machinery: A Knowledge Base, Its Mathematical Formulation, and Its Functional CharacterizationabstractMetabolic network reconstructions represent valuable scaffolds for '-omics' data integration and are used to computationally interrogate network properties. However, they do not explicitly account for the synthesis of macromolecules (i.e., proteins and RNA). Here, we present the first genome-scale, fine-grained reconstruction of Escherichia coli's transcriptional and translational machinery, which produces 423 functional gene products in a sequence-specific manner and accounts for all necessary chemical transformations. Legacy data from over 500 publications and three databases were reviewed, and many pathways were considered, including stable RNA maturation and modification, protein complex formation, and iron-sulfur cluster biogenesis. This reconstruction represents the most comprehensive knowledge base for these important cellular functions in E. coli and is unique in its scope. Furthermore, it was converted into a mathematical model and used to: (1) quantitatively integrate gene expression data as reaction constraints and (2) compute functional network states, which were compared to reported experimental data. For example, the model predicted accurately the ribosome production, without any parameterization. Also, in silico rRNA operon deletion suggested that a high RNA polymerase density on the remaining rRNA operons is needed to reproduce the reported experimental ribosome numbers. Moreover, functional protein modules were determined, and many were found to contain gene products from multiple subsystems, highlighting the functional interaction of these proteins. This genome-scale reconstruction of E. coli's transcriptional and translational machinery presents a milestone in systems biology because it will enable quantitative integration of '-omics' datasets and thus the study of the mechanistic principles underlying the genotype-phenotype relationship. Ines Thiele, Neema Jamshidi, Ronan M. T. Fleming, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2008 | Context-Specific Metabolic Networks Are Consistent with ExperimentsabstractReconstructions of cellular metabolism are publicly available for a variety of different microorganisms and some mammalian genomes. To date, these reconstructions are "genome-scale" and strive to include all reactions implied by the genome annotation, as well as those with direct experimental evidence. Clearly, many of the reactions in a genome-scale reconstruction will not be active under particular conditions or in a particular cell type. Methods to tailor these comprehensive genome-scale reconstructions into context-specific networks will aid predictive in silico modeling for a particular situation. We present a method called Gene Inactivity Moderated by Metabolism and Expression (GIMME) to achieve this goal. The GIMME algorithm uses quantitative gene expression data and one or more presupposed metabolic objectives to produce the context-specific reconstruction that is most consistent with the available data. Furthermore, the algorithm provides a quantitative inconsistency score indicating how consistent a set of gene expression data is with a particular metabolic objective. We show that this algorithm produces results consistent with biological experiments and intuition for adaptive evolution of bacteria, rational design of metabolic engineering strains, and human skeletal muscle cells. This work represents progress towards producing constraint-based models of metabolism that are specific to the conditions where the expression profiling data is available. Scott A. Becker, Bernhard O. Palsson |
PLoS Comput. Biol. | 2 |
| 2008 | Top-Down Analysis of Temporal Hierarchy in Biochemical Reaction NetworksabstractThe study of dynamic functions of large-scale biological networks has intensified in recent years. A critical component in developing an understanding of such dynamics involves the study of their hierarchical organization. We investigate the temporal hierarchy in biochemical reaction networks focusing on: (1) the elucidation of the existence of "pools" (i.e., aggregate variables) formed from component concentrations and (2) the determination of their composition and interactions over different time scales. To date the identification of such pools without prior knowledge of their composition has been a challenge. A new approach is developed for the algorithmic identification of pool formation using correlations between elements of the modal matrix that correspond to a pair of concentrations and how such correlations form over the hierarchy of time scales. The analysis elucidates a temporal hierarchy of events that range from chemical equilibration events to the formation of physiologically meaningful pools, culminating in a network-scale (dynamic) structure-(physiological) function relationship. This method is validated on a model of human red blood cell metabolism and further applied to kinetic models of yeast glycolysis and human folate metabolism, enabling the simplification of these models. The understanding of temporal hierarchy and the formation of dynamic aggregates on different time scales is foundational to the study of network dynamics and has relevance in multiple areas ranging from bacterial strain design and metabolic engineering to the understanding of disease processes in humans. Neema Jamshidi, Bernhard O. Palsson |
PLoS Comput. Biol. | 2 |
| 2007 | Estimation of the number of extreme pathways for metabolic networksabstractABSTRACT: BACKGROUND: The set of extreme pathways (ExPa), {pi}, defines the convex basis vectors used for the mathematical characterization of the null space of the stoichiometric matrix for biochemical reaction networks. ExPa analysis has been used for a number of studies to determine properties of metabolic networks as well as to obtain insight into their physiological and functional states in silico. However, the number of ExPas, p = |{pi}|, grows with the size and complexity of the network being studied, and this poses a computational challenge. For this study, we investigated the relationship between the number of extreme pathways and simple network properties. RESULTS: We established an estimating function for the number of ExPas using these easily obtainable network measurements. In particular, it was found that log [p] had an exponential relationship with log[ summation operatori=1Rd-id+ici] MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacyGGSbaBcqGGVbWBcqGGNbWzdaWadaqaamaaqadabaGaemizaq2aaSbaaSqaaiabgkHiTmaaBaaameaacqWGPbqAaeqaaaWcbeaakiabdsgaKnaaBaaaleaacqGHRaWkdaWgaaadbaGaemyAaKgabeaaaSqabaGccqWGJbWydaWgaaWcbaGaemyAaKgabeaaaeaacqWGPbqAcqGH9aqpcqaIXaqmaeaacqWGsbGua0GaeyyeIuoaaOGaay5waiaaw2faaaaa@4414@, where R = |Reff| is the number of active reactions in a network, d-i MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacqWGKbazdaWgaaWcbaGaeyOeI0YaaSbaaWqaaiabdMgaPbqabaaaleqaaaaa@30A9@ and d+i MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacqWGKbazdaWgaaWcbaGaey4kaSYaaSbaaWqaaiabdMgaPbqabaaaleqaaaaa@309E@ the incoming and outgoing degrees of the reactions ri in Reff, and ci the clustering coefficient for each active reaction. CONCLUSION: This relationship typically gave an estimate of the number of extreme pathways to within a factor of 10 of the true number. Such a function providing an estimate for the total number of ExPas for a given system will enable researchers to decide whether ExPas analysis is an appropriate investigative tool. Matthew Yeung, Ines Thiele, Bernhard O. Palsson |
BMC Bioinform. | 3 |
| 2007 | Metabolic Reconstruction and Modeling of Nitrogen Fixation in Rhizobium etliabstractRhizobiaceas are bacteria that fix nitrogen during symbiosis with plants. This symbiotic relationship is crucial for the nitrogen cycle, and understanding symbiotic mechanisms is a scientific challenge with direct applications in agronomy and plant development. Rhizobium etli is a bacteria which provides legumes with ammonia (among other chemical compounds), thereby stimulating plant growth. A genome-scale approach, integrating the biochemical information available for R. etli, constitutes an important step toward understanding the symbiotic relationship and its possible improvement. In this work we present a genome-scale metabolic reconstruction (iOR363) for R. etli CFN42, which includes 387 metabolic and transport reactions across 26 metabolic pathways. This model was used to analyze the physiological capabilities of R. etli during stages of nitrogen fixation. To study the physiological capacities in silico, an objective function was formulated to simulate symbiotic nitrogen fixation. Flux balance analysis (FBA) was performed, and the predicted active metabolic pathways agreed qualitatively with experimental observations. In addition, predictions for the effects of gene deletions during nitrogen fixation in Rhizobia in silico also agreed with reported experimental data. Overall, we present some evidence supporting that FBA of the reconstructed metabolic network for R. etli provides results that are in agreement with physiological observations. Thus, as for other organisms, the reconstructed genome-scale metabolic network provides an important framework which allows us to compare model predictions with experimental measurements and eventually generate hypotheses on ways to improve nitrogen fixation. Osbaldo Resendis-Antonio, Jennifer L. Reed, Sergio Encarnación-Guevara, Julio Collado-Vides, Bernhard O. Palsson |
PLoS Comput. Biol. | 5 |
| 2006 | Network-level analysis of metabolic regulation in the human red blood cell using random sampling and singular value decompositionabstractBACKGROUND: Extreme pathways (ExPas) have been shown to be valuable for studying the functions and capabilities of metabolic networks through characterization of the null space of the stoichiometric matrix (S). Singular value decomposition (SVD) of the ExPa matrix P has previously been used to characterize the metabolic regulatory problem in the human red blood cell (hRBC) from a network perspective. The calculation of ExPas is NP-hard, and for genome-scale networks the computation of ExPas has proven to be infeasible. Therefore an alternative approach is needed to reveal regulatory properties of steady state solution spaces of genome-scale stoichiometric matrices. RESULTS: We show that the SVD of a matrix (W) formed of random samples from the steady-state solution space of the hRBC metabolic network gives similar insights into the regulatory properties of the network as was obtained with SVD of P. This new approach has two main advantages. First, it works with a direct representation of the shape of the metabolic solution space without the confounding factor of a non-uniform distribution of the extreme pathways and second, the SVD procedure can be applied to a very large number of samples, such as will be produced from genome-scale networks. CONCLUSION: These results show that we are now in a position to study the network aspects of the regulatory problem in genome-scale metabolic networks through the use of random sampling. Christian L. Barrett, Nathan D. Price 0001, Bernhard O. Palsson |
BMC Bioinform. | 3 |
| 2006 | Metabolite coupling in genome-scale metabolic networksabstractBACKGROUND: Biochemically detailed stoichiometric matrices have now been reconstructed for various bacteria, yeast, and for the human cardiac mitochondrion based on genomic and proteomic data. These networks have been manually curated based on legacy data and elementally and charge balanced. Comparative analysis of these well curated networks is now possible. Pairs of metabolites often appear together in several network reactions, linking them topologically. This co-occurrence of pairs of metabolites in metabolic reactions is termed herein "metabolite coupling." These metabolite pairs can be directly computed from the stoichiometric matrix, S. Metabolite coupling is derived from the matrix ŝŝT, whose off-diagonal elements indicate the number of reactions in which any two metabolites participate together, where ŝ is the binary form of S. RESULTS: Metabolite coupling in the studied networks was found to be dominated by a relatively small group of highly interacting pairs of metabolites. As would be expected, metabolites with high individual metabolite connectivity also tended to be those with the highest metabolite coupling, as the most connected metabolites couple more often. For metabolite pairs that are not highly coupled, we show that the number of reactions a pair of metabolites shares across a metabolic network closely approximates a line on a log-log scale. We also show that the preferential coupling of two metabolites with each other is spread across the spectrum of metabolites and is not unique to the most connected metabolites. We provide a measure for determining which metabolite pairs couple more often than would be expected based on their individual connectivity in the network and show that these metabolites often derive their principal biological functions from existing in pairs. Thus, analysis of metabolite coupling provides information beyond that which is found from studying the individual connectivity of individual metabolites. CONCLUSION: The coupling of metabolites is an important topological property of metabolic networks. By computing coupling quantitatively for the first time in genome-scale metabolic networks, we provide insight into the basic structure of these networks. Scott A. Becker, Nathan D. Price 0001, Bernhard O. Palsson |
BMC Bioinform. | 3 |
| 2006 | Long-Range Periodic Patterns in Microbial Genomes Indicate Significant Multi-Scale Chromosomal OrganizationabstractGenome organization can be studied through analysis of chromosome position-dependent patterns in sequence-derived parameters. A comprehensive analysis of such patterns in prokaryotic sequences and genome-scale functional data has yet to be performed. We detected spatial patterns in sequence-derived parameters for 163 chromosomes occurring in 135 bacterial and 16 archaeal organisms using wavelet analysis. Pattern strength was found to correlate with organism-specific features such as genome size, overall GC content, and the occurrence of known motility and chromosomal binding proteins. Given additional functional data for Escherichia coli, we found significant correlations among chromosome position dependent patterns in numerous properties, some of which are consistent with previously experimentally identified chromosome macrodomains. These results demonstrate that the large-scale organization of most sequenced genomes is significantly nonrandom, and, moreover, that this organization is likely linked to genome size, nucleotide composition, and information transfer processes. Constraints on genome evolution and design are thus not solely dependent upon information content, but also upon an intricate multi-parameter, multi-length-scale organization of the chromosome. Timothy E. Allen, Nathan D. Price 0001, Andrew R. Joyce, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2006 | Iterative Reconstruction of Transcriptional Regulatory Networks: An Algorithmic ApproachabstractThe number of complete, publicly available genome sequences is now greater than 200, and this number is expected to rapidly grow in the near future as metagenomic and environmental sequencing efforts escalate and the cost of sequencing drops. In order to make use of this data for understanding particular organisms and for discerning general principles about how organisms function, it will be necessary to reconstruct their various biochemical reaction networks. Principal among these will be transcriptional regulatory networks. Given the physical and logical complexity of these networks, the various sources of (often noisy) data that can be utilized for their elucidation, the monetary costs involved, and the huge number of potential experiments approximately 10(12)) that can be performed, experiment design algorithms will be necessary for synthesizing the various computational and experimental data to maximize the efficiency of regulatory network reconstruction. This paper presents an algorithm for experimental design to systematically and efficiently reconstruct transcriptional regulatory networks. It is meant to be applied iteratively in conjunction with an experimental laboratory component. The algorithm is presented here in the context of reconstructing transcriptional regulation for metabolism in Escherichia coli, and, through a retrospective analysis with previously performed experiments, we show that the produced experiment designs conform to how a human would design experiments. The algorithm is able to utilize probability estimates based on a wide range of computational and experimental sources to suggest experiments with the highest potential of discovering the greatest amount of new regulatory knowledge. Christian L. Barrett, Bernhard O. Palsson |
PLoS Comput. Biol. | 2 |
| 2006 | Matrix Formalism to Describe Functional States of Transcriptional Regulatory SystemsabstractComplex regulatory networks control the transcription state of a genome. These transcriptional regulatory networks (TRNs) have been mathematically described using a Boolean formalism, in which the state of a gene is represented as either transcribed or not transcribed in response to regulatory signals. The Boolean formalism results in a series of regulatory rules for the individual genes of a TRN that in turn can be used to link environmental cues to the transcription state of a genome, thereby forming a complete transcriptional regulatory system (TRS). Herein, we develop a formalism that represents such a set of regulatory rules in a matrix form. Matrix formalism allows for the systemic characterization of the properties of a TRS and facilitates the computation of the transcriptional state of the genome under any given set of environmental conditions. Additionally, it provides a means to incorporate mechanistic detail of a TRS as it becomes available. In this study, the regulatory network matrix, R, for a prototypic TRS is characterized and the fundamental subspaces of this matrix are described. We illustrate how the matrix representation of a TRS coupled with its environment (R*) allows for a sampling of all possible expression states of a given network, and furthermore, how the fundamental subspaces of the matrix provide a way to study key TRS features and may assist in experimental design. Erwin P. Gianchandani, Jason A. Papin, Nathan D. Price 0001, Andrew R. Joyce, Bernhard O. Palsson |
PLoS Comput. Biol. | 5 |
| 2006 | Identification of Genome-Scale Metabolic Network Models Using Experimentally Measured Flux ProfilesabstractGenome-scale metabolic network models can be reconstructed for well-characterized organisms using genomic annotation and literature information. However, there are many instances in which model predictions of metabolic fluxes are not entirely consistent with experimental data, indicating that the reactions in the model do not match the active reactions in the in vivo system. We introduce a method for determining the active reactions in a genome-scale metabolic network based on a limited number of experimentally measured fluxes. This method, called optimal metabolic network identification (OMNI), allows efficient identification of the set of reactions that results in the best agreement between in silico predicted and experimentally measured flux distributions. We applied the method to intracellular flux data for evolved Escherichia coli mutant strains with lower than predicted growth rates in order to identify reactions that act as flux bottlenecks in these strains. The expression of the genes corresponding to these bottleneck reactions was often found to be downregulated in the evolved strains relative to the wild-type strain. We also demonstrate the ability of the OMNI method to diagnose problems in E. coli strains engineered for metabolite overproduction that have not reached their predicted production potential. The OMNI method applied to flux data for evolved strains can be used to provide insights into mechanisms that limit the ability of microbial strains to evolve towards their predicted optimal growth phenotypes. When applied to industrial production strains, the OMNI method can also be used to suggest metabolic engineering strategies to improve byproduct secretion. In addition to these applications, the method should prove to be useful in general for reconstructing metabolic networks of ill-characterized microbial organisms based on limited amounts of experimental data. Markus J. Herrgård, Stephen S. Fong, Bernhard O. Palsson |
PLoS Comput. Biol. | 3 |
| 2005 | expa: a program for calculating extreme pathways in biochemical reaction networksabstractUNLABELLED: The set of extreme pathways, a generating set for all possible steady-state flux maps in a biochemical reaction network, can be computed from the stoichiometric matrix, an incidence-like matrix reflecting the network topology. Here, we describe the implementation of a well-known algorithm to compute these pathways and give a summary of the features of the available software. AVAILABILITY: The C-code, along with a Windows executable and sample network reaction files, are available at http://systemsbiology.ucsd.edu CONTACT: [email protected]. Steven L. Bell, Bernhard O. Palsson |
Bioinform. | 2 |
| 2001 | Dynamic simulation of the human red blood cell metabolic networkabstractAbstract Summary: We have developed a Mathematica®application package to perform dynamic simulations of the red blood cell (RBC) metabolic network. The package relies on, and integrates, many years of mathematical modeling and biochemical work on red blood cell metabolism. The extensive data regarding the red blood cell metabolic network and the previous kinetic analysis of all the individual components makes the human RBC an ideal ‘model’ system for mathematical metabolic models. The Mathematica package can be used to understand the dynamics and regulatory characteristics of the red blood cell. Availability: The Mathematica package and an example file can be downloaded from http://gcrg.ucsd.edu Neema Jamshidi, Jeremy S. Edwards, Tom Fahland, George M. Church, Bernhard O. Palsson |
Bioinform. | 5 |
| 2000 | Metabolic flux balance analysis and the in silico analysis of Escherichia coli K-12 gene deletionsabstractBACKGROUND: Genome sequencing and bioinformatics are producing detailed lists of the molecular components contained in many prokaryotic organisms. From this 'parts catalogue' of a microbial cell, in silico representations of integrated metabolic functions can be constructed and analyzed using flux balance analysis (FBA). FBA is particularly well-suited to study metabolic networks based on genomic, biochemical, and strain specific information. RESULTS: Herein, we have utilized FBA to interpret and analyze the metabolic capabilities of Escherichia coli. We have computationally mapped the metabolic capabilities of E. coli using FBA and examined the optimal utilization of the E. coli metabolic pathways as a function of environmental variables. We have used an in silico analysis to identify seven gene products of central metabolism (glycolysis, pentose phosphate pathway, TCA cycle, electron transport system) essential for aerobic growth of E. coli on glucose minimal media, and 15 gene products essential for anaerobic growth on glucose minimal media. The in silico tpi-, zwf, and pta- mutant strains were examined in more detail by mapping the capabilities of these in silico isogenic strains. CONCLUSIONS: We found that computational models of E. coli metabolism based on physicochemical constraints can be used to interpret mutant behavior. These in silica results lead to a further understanding of the complex genotype-phenotype relation. Jeremy S. Edwards, Bernhard O. Palsson |
BMC Bioinform. | 2 |