VLDB 2026 Research / reviewers in the wild / expert
Jason A. Papin
dblp:41/3031
· DBLP profile ↗
41ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-2769-5805ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 41 · 4 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ten simple rules for developing a training program
Leanna B. Blevins, Amy M. Harrigan, Kevin A. Janes, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2025 | Ten simple rules for good model-sharing practicesabstractComputational models are complex scientific constructs that have become essential for us to better understand the world. Many models are valuable for peers within and beyond disciplinary boundaries. However, there are no widely agreed-upon standards for sharing models. This paper suggests 10 simple rules for you to both (i) ensure you share models in a way that is at least "good enough," and (ii) enable others to lead the change towards better model-sharing practices. Ismael Kherroubi Garcia, Christopher Erdmann, Sandra Gesing, C. Michael Barton, Lauren Cadwallader, Geerten M. Hengeveld, Christine R. Kirkpatrick, Kathryn Knight, Carsten Lemmen, Rebecca Ringuette, Qing Zhan, Melissa Harrison, Feilim Mac Gabhann, Natalie Meyers, Cailean Osborne, Charlotte Till, Paul R. Brenner, Matt Buys, Min Chen 0008, Allen Lee, Jason A. Papin, Yuhan Rao |
PLoS Comput. Biol. | 21 |
| 2024 | Identifying metabolic adaptations characteristic of cardiotoxicity using paired transcriptomics and metabolomics data integrated with a computational model of heart metabolismabstractImprovements in the diagnosis and treatment of cancer have revealed long-term side effects of chemotherapeutics, particularly cardiotoxicity. Here, we present paired transcriptomics and metabolomics data characterizing in vitro cardiotoxicity to three compounds: 5-fluorouracil, acetaminophen, and doxorubicin. Standard gene enrichment and metabolomics approaches identify some commonly affected pathways and metabolites but are not able to readily identify metabolic adaptations in response to cardiotoxicity. The paired data was integrated with a genome-scale metabolic network reconstruction of the heart to identify shifted metabolic functions, unique metabolic reactions, and changes in flux in metabolic reactions in response to these compounds. Using this approach, we confirm previously seen changes in the p53 pathway by doxorubicin and RNA synthesis by 5-fluorouracil, we find evidence for an increase in phospholipid metabolism in response to acetaminophen, and we see a shift in central carbon metabolism suggesting an increase in metabolic demand after treatment with doxorubicin and 5-fluorouracil. Bonnie V. Dougherty, Connor J. Moore, Kristopher D. Rawls, Matthew L. Jenior, Bryan Chun, Sarbajeet Nagdas, Jeffrey J. Saucerman, Glynis L. Kolling, Anders Wallqvist, Jason A. Papin |
PLoS Comput. Biol. | 10 |
| 2024 | Spatial transcriptome-guided multi-scale framework connects P. aeruginosa metabolic states to oxidative stress biofilm microenvironmentabstractWith the generation of spatially resolved transcriptomics of microbial biofilms, computational tools can be used to integrate this data to elucidate the multi-scale mechanisms controlling heterogeneous biofilm metabolism. This work presents a Multi-scale model of Metabolism In Cellular Systems (MiMICS) which is a computational framework that couples a genome-scale metabolic network reconstruction (GENRE) with Hybrid Automata Library (HAL), an existing agent-based model and reaction-diffusion model platform. A key feature of MiMICS is the ability to incorporate multiple -omics-guided metabolic models, which can represent unique metabolic states that yield different metabolic parameter values passed to the extracellular models. We used MiMICS to simulate Pseudomonas aeruginosa regulation of denitrification and oxidative stress metabolism in hypoxic and nitric oxide (NO) biofilm microenvironments. Integration of P. aeruginosa PA14 biofilm spatial transcriptomic data into a P. aeruginosa PA14 GENRE generated four PA14 metabolic model states that were input into MiMICS. Characteristic of aerobic, denitrification, and oxidative stress metabolism, the four metabolic model states predicted different oxygen, nitrate, and NO exchange fluxes that were passed as inputs to update the agent's local metabolite concentrations in the extracellular reaction-diffusion model. Individual bacterial agents chose a PA14 metabolic model state based on a combination of stochastic rules, and agents sensing local oxygen and NO. Transcriptome-guided MiMICS predictions suggested microscale denitrification and oxidative stress metabolic heterogeneity emerged due to local variability in the NO biofilm microenvironment. MiMICS accurately predicted the biofilm's spatial relationships between denitrification, oxidative stress, and central carbon metabolism. As simulated cells responded to extracellular NO, MiMICS revealed dynamics of cell populations heterogeneously upregulating reactions in the denitrification pathway, which may function to maintain NO levels within non-toxic ranges. We demonstrated that MiMICS is a valuable computational tool to incorporate multiple -omics-guided metabolic models to mechanistically map heterogeneous microbial metabolic states to the biofilm microenvironment. Tracy J. Kuper, Mohammad Mazharul Islam, Shayn M. Peirce, Jason A. Papin, Roseanne M. Ford |
PLoS Comput. Biol. | 4 |
| 2024 | Celebrating a body of workabstractMy first exposure to visibly fluorescent proteins (FPs) was near the end of my time as a faculty member at the University of California, Berkeley.Prof. Alexander Glazer, a friend and colleague there, was the world's expert on phycobiliproteins, the brilliantly colored and intensely fluorescent proteins that serve as light-harvesting antennae for the photosynthetic apparatus of blue-green algae or cyanobacteria.One day, probably around 1987-88, Glazer told me that his lab had cloned the gene for one of the phycobiliproteins.Furthermore, he said, the apoprotein produced from this gene became fluorescent when mixed with its chromophore, a small molecule cofactor that could be extracted from dried cyanobacteria under conditions that cleaved its bond to the phycobiliprotein.I remember becoming very excited about the prospect that an arbitrary protein could be fluorescently tagged in situ by genetically fusing it to the phycobiliprotein, then administering the chromophore, which I hoped would be able to cross membranes and get inside cells.Unfortunately, Glazer's lab then found out that the spontaneous reaction between the apoprotein and the chromophore produced the "wrong" product, whose fluorescence was red-shifted and five-fold lower than that of the native phycobiliprotein [1][2][3] .An enzyme from the cyanobacteria was required to insert the chromophore correctly into the apoprotein.This enzyme was a heterodimer of two gene products, so at least three cyanobacterial genes would have to be introduced into any other organism, not counting any gene products needed to synthesize the chromophore 4 .Meanwhile fluorescence imaging of the second messenger cAMP (cyclic adenosine 3',5'-monophosphate) had become one of my main research goals by 1988.I reasoned that the best way to create a fluorescent sensor to detect cAMP with the necessary affinity and selectivity inside cells would be to hijack a natural cAMP-binding protein.After much consideration of the various candidates known at the time, I chose cAMP-dependent protein kinase, now more commonly abbreviated PKA.PKA contains two types of Jason A. Papin, Feilim Mac Gabhann, Virginia E. Pitzer |
PLoS Comput. Biol. | 1 |
| 2023 | Reconstructor: a COBRApy compatible tool for automated genome-scale metabolic network reconstruction with parsimonious flux-based gap-fillingabstractMOTIVATION: Genome-scale metabolic network reconstructions (GENREs) are valuable for understanding cellular metabolism in silico. Several tools exist for automatic GENRE generation. However, these tools frequently (i) do not readily integrate with some of the widely-used suites of packaged methods available for network analysis, (ii) lack effective network curation tools, (iii) are not sufficiently user-friendly, and (iv) often produce low-quality draft reconstructions. RESULTS: Here, we present Reconstructor, a user-friendly, COBRApy-compatible tool that produces high-quality draft reconstructions with reaction and metabolite naming conventions that are consistent with the ModelSEED biochemistry database and includes a gap-filling technique based on the principles of parsimony. Reconstructor can generate SBML GENREs from three input types: annotated protein .fasta sequences (Type 1 input), a BLASTp output (Type 2), or an existing SBML GENRE that can be further gap-filled (Type 3). While Reconstructor can be used to create GENREs of any species, we demonstrate the utility of Reconstructor with bacterial reconstructions. We demonstrate how Reconstructor readily generates high-quality GENRES that capture strain, species, and higher taxonomic differences in functional metabolism of bacteria and are useful for further biological discovery. AVAILABILITY AND IMPLEMENTATION: The Reconstructor Python package is freely available for download. Complete installation and usage instructions and benchmarking data are available at http://github.com/emmamglass/reconstructor. Matthew L. Jenior, Emma M. Glass, Jason A. Papin |
Bioinform. | 3 |
| 2023 | The blossoming of methods and software in computational biologyabstractAs we wrote previously [1], science benefits when we share not only our insights and discoveries but also the tools and approaches that we develop.This sharing improves reproducibility and reuse, and it enables others to build on our work.These tools and approaches are research accelerants and are the focus of 2 key sections in our journal: Methods and Software.Since its founding in 2005, PLOS Computational Biology has been the home of exceptional computational research and cutting-edge methodological advances.We publish papers that use computational techniques to generate new biological insight, and we publish papers that describe software or methods that many other researchers in the field can use independently to generate new biological insight.There is great merit in using existing techniques to advance biological understanding, and there is great merit in developing and sharing new techniques that will result in further advances.The field of Computational Biology supports both.In 2013, the journal introduced the Methods section and the Software section to encourage more researchers to publish methodological advancements.From the inception of these sections, there were dedicated Methods Editors and Software Editors on the Editorial Board responsible for the papers submitted to each section.Through the end of 2022, those editors have guided the publication of 647 Methods papers and 286 Software papers, from among the 1,790 and 646 manuscripts submitted, respectively.The standard of those papers has been high, with many excellent and impactful papers published, covering methods for microbiome data analysis [2], deep learning to predict molecular interactions [3], and quantitative analysis of live-cell imaging [4], and software tools ranging from Bayesian Evolutionary Analysis [5] to multiomics integration and feature selection [6], and genome assembly [7].At the outset, the concept of having separate Methods-specific and Software-specific editors was seen as beneficial to being able to set appropriate criteria for the review and publication of these papers, for example, making clear that new work that facilitated original research was publishable, even if the paper itself did not include original research findings.By creating and stewarding a clear standard for these papers, it has been possible to maintain and even increase the high standard of methods and software papers published in our journal.By all metrics, the Methods and Software sections have been a massive success and continue to grow.More than half of all submissions received since 2013 in these sections have come in the last 3 years (Fig 1).Our editors made a Herculean effort to meet this demand, but it became harder to keep up.Clearly, we needed to spread the handling of these papers across more people. Feilim Mac Gabhann, Virginia E. Pitzer, Jason A. Papin |
PLoS Comput. Biol. | 3 |
| 2023 | Metabolic modeling of sex-specific liver tissue suggests mechanism of differences in toxicological responsesabstractMale subjects in animal and human studies are disproportionately used for toxicological testing. This discrepancy is evidenced in clinical medicine where females are more likely than males to experience liver-related adverse events in response to xenobiotics. While previous work has shown gene expression differences between the sexes, there is a lack of systems-level approaches to understand the direct clinical impact of these differences. Here, we integrate gene expression data with metabolic network models to characterize the impact of transcriptional changes of metabolic genes in the context of sex differences and drug treatment. We used Tasks Inferred from Differential Expression (TIDEs), a reaction-centric approach to analyzing differences in gene expression, to discover that several metabolic pathways exhibit sex differences including glycolysis, fatty acid metabolism, nucleotide metabolism, and xenobiotics metabolism. When TIDEs is used to compare expression differences in treated and untreated hepatocytes, we find several subsystems with differential expression overlap with the sex-altered pathways such as fatty acid metabolism, purine and pyrimidine metabolism, and xenobiotics metabolism. Finally, using sex-specific transcriptomic data, we create individual and averaged male and female liver models and find differences in the pentose phosphate pathway and other metabolic pathways. These results suggest potential sex differences in the contribution of the pentose phosphate pathway to oxidative stress, and we recommend further research into how these reactions respond to hepatotoxic pharmaceuticals. Connor J. Moore, Christopher P. Holstege, Jason A. Papin |
PLoS Comput. Biol. | 3 |
| 2023 | Network analysis of toxin production in Clostridioides difficile identifies key metabolic dependenciesabstractClostridioides difficile pathogenesis is mediated through its two toxin proteins, TcdA and TcdB, which induce intestinal epithelial cell death and inflammation. It is possible to alter C. difficile toxin production by changing various metabolite concentrations within the extracellular environment. However, it is unknown which intracellular metabolic pathways are involved and how they regulate toxin production. To investigate the response of intracellular metabolic pathways to diverse nutritional environments and toxin production states, we use previously published genome-scale metabolic models of C. difficile strains CD630 and CDR20291 (iCdG709 and iCdR703). We integrated publicly available transcriptomic data with the models using the RIPTiDe algorithm to create 16 unique contextualized C. difficile models representing a range of nutritional environments and toxin states. We used Random Forest with flux sampling and shadow pricing analyses to identify metabolic patterns correlated with toxin states and environment. Specifically, we found that arginine and ornithine uptake is particularly active in low toxin states. Additionally, uptake of arginine and ornithine is highly dependent on intracellular fatty acid and large polymer metabolite pools. We also applied the metabolic transformation algorithm (MTA) to identify model perturbations that shift metabolism from a high toxin state to a low toxin state. This analysis expands our understanding of toxin production in C. difficile and identifies metabolic dependencies that could be leveraged to mitigate disease severity. Deborah A. Powers, Matthew L. Jenior, Glynis L. Kolling, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2023 | A computational model of Pseudomonas syringae metabolism unveils a role for branched-chain amino acids in Arabidopsis leaf colonizationabstractBacterial pathogens adapt their metabolism to the plant environment to successfully colonize their hosts. In our efforts to uncover the metabolic pathways that contribute to the colonization of Arabidopsis thaliana leaves by Pseudomonas syringae pv tomato DC3000 (Pst DC3000), we created iPst19, an ensemble of 100 genome-scale network reconstructions of Pst DC3000 metabolism. We developed a novel approach for gene essentiality screens, leveraging the predictive power of iPst19 to identify core and ancillary condition-specific essential genes. Constraining the metabolic flux of iPst19 with Pst DC3000 gene expression data obtained from naïve-infected or pre-immunized-infected plants, revealed changes in bacterial metabolism imposed by plant immunity. Machine learning analysis revealed that among other amino acids, branched-chain amino acids (BCAAs) metabolism significantly contributed to the overall metabolic status of each gene-expression-contextualized iPst19 simulation. These predictions were tested and confirmed experimentally. Pst DC3000 growth and gene expression analysis showed that BCAAs suppress virulence gene expression in vitro without affecting bacterial growth. In planta, however, an excess of BCAAs suppress the expression of virulence genes at the early stages of infection and significantly impair the colonization of Arabidopsis leaves. Our findings suggesting that BCAAs catabolism is necessary to express virulence and colonize the host. Overall, this study provides valuable insights into how plant immunity impacts Pst DC3000 metabolism, and how bacterial metabolism impacts the expression of virulence. Philip J. Tubergen, Gregory L. Medlock, Anni Moore, Xiaomu Zhang, Jason A. Papin, Cristian H. Danna |
PLoS Comput. Biol. | 5 |
| 2022 | Advancing code sharing in the computational biology communityabstractOn March 30, 2021, a new code sharing policy was introduced at PLOS Computational Biology [1].This policy requires any code supporting a publication to be shared unless there are ethical or legal restrictions that prevent sharing.The policy was introduced in response to community desire for a stronger position on code sharing to reflect the fact that the majority of the community already voluntarily share code [2,3].This community-driven support for open science practices aligns well with the PLOS mission, and, therefore, the implementation of the new policy was a logical progression for the journal.The policy focuses on increasing code sAU : Pleasenotet haring as its primary aim, which, in turn, will support reproducibility, and so is not prescriptive to authors about how or where to share their code.The policy (https://journals.plos.org/ ploscompbiol/s/code-availability) allows authors to comply in ways which work for them.By the end of the first year of the policy, we expected to see an increase in code sharing rates (the percentage of published research articles that share code) without any negative impact on the publishing demographics or the author, editor, and journal staff experiences.This Editorial reports on the impact of policy over the first 12 months, provides a longitudinal view of code sharing in the journal since 2019, and articulates how this effort can move forward to enhance further sharing, reproducibility, and openness. Lauren Cadwallader, Feilim Mac Gabhann, Jason A. Papin, Virginia E. Pitzer |
PLoS Comput. Biol. | 3 |
| 2022 | Comparative analyses of parasites with a comprehensive database of genome-scale metabolic modelsabstractProtozoan parasites cause diverse diseases with large global impacts. Research on the pathogenesis and biology of these organisms is limited by economic and experimental constraints. Accordingly, studies of one parasite are frequently extrapolated to infer knowledge about another parasite, across and within genera. Model in vitro or in vivo systems are frequently used to enhance experimental manipulability, but these systems generally use species related to, yet distinct from, the clinically relevant causal pathogen. Characterization of functional differences among parasite species is confined to post hoc or single target studies, limiting the utility of this extrapolation approach. To address this challenge and to accelerate parasitology research broadly, we present a functional comparative analysis of 192 genomes, representing every high-quality, publicly-available protozoan parasite genome including Plasmodium, Toxoplasma, Cryptosporidium, Entamoeba, Trypanosoma, Leishmania, Giardia, and other species. We generated an automated metabolic network reconstruction pipeline optimized for eukaryotic organisms. These metabolic network reconstructions serve as biochemical knowledgebases for each parasite, enabling qualitative and quantitative comparisons of metabolic behavior across parasites. We identified putative differences in gene essentiality and pathway utilization to facilitate the comparison of experimental findings and discovered that phylogeny is not the sole predictor of metabolic similarity. This knowledgebase represents the largest collection of genome-scale metabolic models for both pathogens and eukaryotes; with this resource, we can predict species-specific functions, contextualize experimental results, and optimize selection of experimental systems for fastidious species. Maureen A. Carey, Gregory L. Medlock, Michal Stolarczyk, William A. Petri Jr., Jennifer L. Guler, Jason A. Papin |
PLoS Comput. Biol. | 6 |
| 2022 | Quantifying cumulative phenotypic and genomic evidence for procedural generation of metabolic network reconstructionsabstractGenome-scale metabolic network reconstructions (GENREs) are valuable tools for understanding microbial metabolism. The process of automatically generating GENREs includes identifying metabolic reactions supported by sufficient genomic evidence to generate a draft metabolic network. The draft GENRE is then gapfilled with additional reactions in order to recapitulate specific growth phenotypes as indicated with associated experimental data. Previous methods have implemented absolute mapping thresholds for the reactions automatically included in draft GENREs; however, there is growing evidence that integrating annotation evidence in a continuous form can improve model accuracy. There is a need for flexibility in the structure of GENREs to better account for uncertainty in biological data, unknown regulatory mechanisms, and context-specificity associated with data inputs. To address this issue, we present a novel method that provides a framework for quantifying combined genomic, biochemical, and phenotypic evidence for each biochemical reaction during automated GENRE construction. Our method, Constraint-based Analysis Yielding reaction Usage across metabolic Networks (CANYUNs), generates accurate GENREs with a quantitative metric for the cumulative evidence for each reaction included in the network. The structuring of CANYUNs allows for the simultaneous integration of three data inputs while maintaining all supporting evidence for biochemical reactions that may be active in an organism. CANYUNs is designed to maximize the utility of experimental and annotation datasets and to ultimately assist in the curation of the reference datasets used for the automatic construction of metabolic networks. We validated CANYUNs by generating an E. coli K-12 model and compared it to the manually curated reconstruction iML1515. Finally, we demonstrated the use of CANYUNs to build a model by generating an E. coli Nissle CANYUNs model using novel phenotypic data that we collected. This method may address key challenges for the procedural construction of metabolic networks by leveraging uncertainty and redundancy in biological data. Thomas J. Moutinho Jr., Benjamin C. Neubert, Matthew L. Jenior, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2022 | Ten simple rules for launching an academic research career
Jason A. Papin, Jessica Keim-Malpass, Sana Syed |
PLoS Comput. Biol. | 1 |
| 2021 | Collaborating with our community to increase code sharingabstractBiology was launched in 2005 as a journal driven by the computational biology research community and with the principle that "open access ensures not only that everything we publish is immediately freely available to anyone, anywhere in the world, but also that the contents of this journal can be redistributed and reused in ways that increase their value."[1] Delivering this important open vision has relied on the enthusiasm and integrity of the broader computational biology community as editors, as reviewers, and as dedicated authors submitting excellent work to the journal.This vision of openness, redistribution, and reuse is so essential because most scientific progress builds on prior efforts.This progress is possible because as scientists we share not just our research results, but also our methods and protocols and key research tools.Open sharing allows others to check and reproduce our observations, and to build on our work, giving it even more impact over time.This value of sharing is not just true in experimental work, where a key antibody or plasmid must be shared to enable scrutiny and further research; it is also true in computational biology, where computational code is a key reagent.Given how essential newly developed code can be to computational biology research we have been collaborating with the Editorial Board of PLOS Computational Biology and consulting with computational biology researchers to develop a new more-rigorous code policy that is intended to increase code sharing on publication of articles.Code sharing is not new to many of our authors, and in 2019 over 40% of research articles published in the journal reported sharing some code.[2] Assuming this percentage to be a baseline of code sharing behaviour, and after consulting with individual researchers about their code sharing practices, we surveyed our authors and others working in the computational biology field.The objective of this survey was to determine what proportion of articles have code associated with them, and what proportion of this code has not been openly shared in the past.We also aimed to better understand challenges researchers have to overcome to share their code, and how a stronger code sharing policy might affect their opinion of the journal.[3] The results indicated that around 70% of articles have code associated with them based on the cohort who completed the survey.However, for about a third of these articles the authors have not shared the code in the past.The reasons given for this discrepancy range from practical issues, such as not having enough time, to legal and ethical reasons, which is supported by previous research into code sharing.[4] Analysis of the survey data suggested around 5% of articles could not share code for legal and ethical reasons, and therefore more articles published in PLOS Computational Biology could share code than have to date.While there are legitimate restrictions on the sharing of some code, we know that availability of code assists with the reproducibility of research, and journal policies are effective Lauren Cadwallader, Jason A. Papin, Feilim Mac Gabhann, Rebecca Kirk |
PLoS Comput. Biol. | 2 |
| 2020 | Transcriptome-guided parsimonious flux analysis improves predictions with metabolic networks in complex environmentsabstractThe metabolic responses of bacteria to dynamic extracellular conditions drives not only the behavior of single species, but also entire communities of microbes. Over the last decade, genome-scale metabolic network reconstructions have assisted in our appreciation of important metabolic determinants of bacterial physiology. These network models have been a powerful force in understanding the metabolic capacity that species may utilize in order to succeed in an environment. Increasingly, an understanding of context-specific metabolism is critical for elucidating metabolic drivers of larger phenotypes and disease. However, previous approaches to use network models in concert with omics data to better characterize experimental systems have met challenges due to assumptions necessary by the various integration platforms or due to large input data requirements. With these challenges in mind, we developed RIPTiDe (Reaction Inclusion by Parsimony and Transcript Distribution) which uses both transcriptomic abundances and parsimony of overall flux to identify the most cost-effective usage of metabolism that also best reflects the cell's investments into transcription. Additionally, in biological samples where it is difficult to quantify specific growth conditions, it becomes critical to develop methods that require lower amounts of user intervention in order to generate accurate metabolic predictions. Utilizing a metabolic network reconstruction for the model organism Escherichia coli str. K-12 substr. MG1655 (iJO1366), we found that RIPTiDe correctly identifies context-specific metabolic pathway activity without supervision or knowledge of specific media conditions. We also assessed the application of RIPTiDe to in vivo metatranscriptomic data where E. coli was present at high abundances, and found that our approach also effectively predicts metabolic behaviors of host-associated bacteria. In the setting of human health, understanding metabolic changes within bacteria in environments where growth substrate availability is difficult to quantify can have large downstream impacts on our ability to elucidate molecular drivers of disease-associated dysbiosis across the microbiota. Our results indicate that RIPTiDe may have potential to provide understanding of context-specific metabolism of bacteria within complex communities. Matthew L. Jenior, Thomas J. Moutinho Jr., Bonnie V. Dougherty, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2020 | Medusa: Software to build and analyze ensembles of genome-scale metabolic network reconstructionsabstractUncertainty in the structure and parameters of networks is ubiquitous across computational biology. In constraint-based reconstruction and analysis of metabolic networks, this uncertainty is present both during the reconstruction of networks and in simulations performed with them. Here, we present Medusa, a Python package for the generation and analysis of ensembles of genome-scale metabolic network reconstructions. Medusa builds on the COBRApy package for constraint-based reconstruction and analysis by compressing a set of models into a compact ensemble object, providing functions for the generation of ensembles using experimental data, and extending constraint-based analyses to ensemble scale. We demonstrate how Medusa can be used to generate ensembles and perform ensemble simulations, and how machine learning can be used in conjunction with Medusa to guide the curation of genome-scale metabolic network reconstructions. Medusa is available under the permissive MIT license from the Python Packaging Index (https://pypi.org) and from github (https://github.com/opencobra/Medusa), and comprehensive documentation is available at https://medusa.readthedocs.io/en/latest. Gregory L. Medlock, Thomas J. Moutinho Jr., Jason A. Papin |
PLoS Comput. Biol. | 3 |
| 2020 | Improving reproducibility in computational biology researchabstractThere has been much discussion in the scientific literature on a crisis of reproducibility in science [1,2].It has been reported that the percentage of studies that are reproducible is as low as 10% or less, depending on the discipline [3].This inability to reproduce scientific findings from a given paper has been attributed to a lack of clarity in the methods and inherent variability in the biological system being studied [4].Reproducibility in computational biology research is certainly a problem, yet perhaps a challenge that our field can uniquely tackle.A lack of reproducibility in computational biology research can be attributed to many factors, but incomplete or erroneous descriptions of the simulations (e.g., which software version was used), incomplete documentation on how to run simulations, or simply failing to post the relevant computer code needed to run a given simulation are common issues that occur.Many tools have emerged that we can leverage to make computational biology research more reproducible (e.g., http://co.mbine.org/and https://normsys.h-its.org/)and there exist articles that propose best practices, such as Ten Simple Rules for Reproducible Computational Research [5] or Ten Simple Rules for Writing and Sharing Computational Analyses in Jupyter Notebooks [6]. Jason A. Papin, Feilim Mac Gabhann, Herbert M. Sauro, David P. Nickerson, Anand K. Rampadarath |
PLoS Comput. Biol. | 1 |
| 2019 | Leveraging the effects of chloroquine on resistant malaria parasites for combination therapiesabstractBACKGROUND: Malaria is a major global health problem, with the Plasmodium falciparum protozoan parasite causing the most severe form of the disease. Prevalence of drug-resistant P. falciparum highlights the need to understand the biology of resistance and to identify novel combination therapies that are effective against resistant parasites. Resistance has compromised the therapeutic use of many antimalarial drugs, including chloroquine, and limited our ability to treat malaria across the world. Fortunately, chloroquine resistance comes at a fitness cost to the parasite; this can be leveraged in developing combination therapies or to reinstate use of chloroquine. RESULTS: To understand biological changes induced by chloroquine treatment, we compared transcriptomics data from chloroquine-resistant parasites in the presence or absence of the drug. Using both linear models and a genome-scale metabolic network reconstruction of the parasite to interpret the expression data, we identified targetable pathways in resistant parasites. This study identified an increased importance of lipid synthesis, glutathione production/cycling, isoprenoids biosynthesis, and folate metabolism in response to chloroquine. CONCLUSIONS: We identified potential drug targets for chloroquine combination therapies. Significantly, our analysis predicts that the combination of chloroquine and sulfadoxine-pyrimethamine or fosmidomycin may be more effective against chloroquine-resistant parasites than either drug alone; further studies will explore the use of these drugs as chloroquine resistance blockers. Additional metabolic weaknesses were found in glutathione generation and lipid synthesis during chloroquine treatment. These processes could be targeted with novel inhibitors to reduce parasite growth and reduce the burden of malaria infections. Thus, we identified metabolic weaknesses of chloroquine-resistant parasites and propose targeted chloroquine combination therapies. Ana M. Untaroiu, Maureen A. Carey, Jennifer L. Guler, Jason A. Papin |
BMC Bioinform. | 4 |
| 2019 | Reconciling high-throughput gene essentiality data with metabolic network reconstructionsabstractThe identification of genes essential for bacterial growth and survival represents a promising strategy for the discovery of antimicrobial targets. Essential genes can be identified on a genome-scale using transposon mutagenesis approaches; however, variability between screens and challenges with interpretation of essentiality data hinder the identification of both condition-independent and condition-dependent essential genes. To illustrate the scope of these challenges, we perform a large-scale comparison of multiple published Pseudomonas aeruginosa gene essentiality datasets, revealing substantial differences between the screens. We then contextualize essentiality using genome-scale metabolic network reconstructions and demonstrate the utility of this approach in providing functional explanations for essentiality and reconciling differences between screens. Genome-scale metabolic network reconstructions also enable a high-throughput, quantitative analysis to assess the impact of media conditions on the identification of condition-independent essential genes. Our computational model-driven analysis provides mechanistic insight into essentiality and contributes novel insights for design of future gene essentiality screens and the identification of core metabolic processes. Anna S. Blazier, Jason A. Papin |
PLoS Comput. Biol. | 2 |
| 2019 | Wisdom of crowds in computational biologyabstractScientific advances are frequently catalyzed by exploring the intersection of disciplines.From its inception, PLOS Computational Biology has published key insights that advance our understanding of biology and medicine-advances enabled by developments in computation and quantitative analyses.As an example, a recent initiative across the journals PLOS Medicine, PLOS ONE, and PLOS Computational Biology [1] resulted in a fantastic collection of research in Machine Learning in Health and Biomedicine.The breadth of research published in PLOS Computational Biology at this intersection of machine learning and health and biology was incredible, from the use of machine learning analysis to delineate biomarkers for soft tissue sarcomas [2] to the prediction of antibiotic resistance in Escherichia coli from pan-genome data [3].We're only beginning to see the power of machine learning applied to health and biology, with the hope of identifying patterns in the biological and clinical data that will lead to biomarkers of disease and the development of new clinical intervention strategies.As these data-driven strategies evolve and mature, they may also lead to a richer understanding of biological mechanisms, enabling models to predict the outcomes of scenarios and perturbations beyond the bounds of previous studies and data.Many more cross-journal initiatives are in the works, exploring how disparate disciplines can be brought together to tackle seemingly intransigent problems with unique perspectives.Just launched is a Targeted Anticancer Therapies and Precision Medicine call for papers jointly with PLOS ONE and PLOS Computational Biology [4].With nearly 10 million people dying from cancer in 2018 [5] and an increasing appreciation of the heterogeneity of the disease [6], there is an urgent need to develop targeted therapies that can be dosed, scheduled, delivered, and combined in ways that are specific to the patient.Computational approaches to this precision medicine challenge can serve as the common framework that integrates the disparate fields of expertise needed to understand the biological, pharmacological, and physiological complexity of the system.Computation can link, leverage, and amplify expertise in all the areas that are needed to understand biology and medicine: molecular dynamics, biochemistry, cell biology, human physiology, pharmacometrics, clinical practice, and more.These interdisciplinary links enable us to predict protein structure changes from gene variant data, to predict pharmacodynamics of associated drug compounds, to identify correlations in data, and simply to handle and process the vast amounts of data.As we do in these scientific ventures, so we do in managing the activity of PLOS Computational Biology.This journal, like the other PLOS community journals, is led by more than 160 practicing scientists with a variety of disciplinary backgrounds.All papers submitted to the journal are evaluated by multiple scientists through the peer review process and by teams of scientists at the editorial stage.It is with this integration of these different opinions of scientists from different subdisciplines of computational biology that we try to identify and support the Jason A. Papin, Feilim Mac Gabhann |
PLoS Comput. Biol. | 1 |
| 2018 | One thousand simple rulesabstractWhat began as a one-off in 2005 as Ten Simple Rules for Getting Published [1] has, in thirteen years, now multiplied a hundredfold to become One Thousand Simple Rules for many aspects of one's professional development and led to Quick Tips in the journal's Education section.This milestone of a thousand rules has been reached thanks to the unselfish work of all stakeholders-authors, editors, reviewers, and readers.Let's face it, writing, editing or reviewing a Ten Simple Rules (TSR) article is not the same as a publication that advances a scientific field.What it is going to get you is the satisfaction of knowing you have passed on a part of your experience in a form that is easily understood and acted upon by those following in your footsteps and hence have a very different kind of positive impact on science.Thank you.On the other hand, as a reader the TSRs might contribute in some small way to you getting tenure, or whatever else you care about in your professional life."How to write a review, by M Pautasso.It changed my approach to writing reviews."Anonymous, Academic faculty "Ten Simple Rules for Reproducible Computational Research, and other articles related to computational biology, programming, etc.They really helped in organizing my projects and code more efficiently." Philip E. Bourne, Fran Lewitter, Scott Markel, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2018 | Ten simple rules for biologists learning to program
Maureen A. Carey, Jason A. Papin |
PLoS Comput. Biol. | 2 |
| 2017 | Managing uncertainty in metabolic network structure and improving predictions using EnsembleFBAabstractGenome-scale metabolic network reconstructions (GENREs) are repositories of knowledge about the metabolic processes that occur in an organism. GENREs have been used to discover and interpret metabolic functions, and to engineer novel network structures. A major barrier preventing more widespread use of GENREs, particularly to study non-model organisms, is the extensive time required to produce a high-quality GENRE. Many automated approaches have been developed which reduce this time requirement, but automatically-reconstructed draft GENREs still require curation before useful predictions can be made. We present a novel approach to the analysis of GENREs which improves the predictive capabilities of draft GENREs by representing many alternative network structures, all equally consistent with available data, and generating predictions from this ensemble. This ensemble approach is compatible with many reconstruction methods. We refer to this new approach as Ensemble Flux Balance Analysis (EnsembleFBA). We validate EnsembleFBA by predicting growth and gene essentiality in the model organism Pseudomonas aeruginosa UCBPP-PA14. We demonstrate how EnsembleFBA can be included in a systems biology workflow by predicting essential genes in six Streptococcus species and mapping the essential genes to small molecule ligands from DrugBank. We found that some metabolic subsystems contributed disproportionately to the set of predicted essential reactions in a way that was unique to each Streptococcus species, leading to species-specific outcomes from small molecule interactions. Through our analyses of P. aeruginosa and six Streptococci, we show that ensembles increase the quality of predictions without drastically increasing reconstruction time, thus making GENRE approaches more practical for applications which require predictions for many non-model organisms. All of our functions and accompanying example code are available in an open online repository. Matthew B. Biggs, Jason A. Papin |
PLoS Comput. Biol. | 2 |
| 2017 | How can computation advance microbiome research?abstractMicrobiome science is already in the fast lane, and computations share the credit.How can computations further speed up the pace of microbiome research as it dashes farther, scaling higher summits?What can computations accomplish?What questions can computations help address and what types of algorithms and software should computational biologists aim to develop?These are tantalizing questions that many of us are facing, probably in particular new groups aiming to venture into a relatively nascent and fast-developing area that promises rapid accumulation of data, emergence of alluring new concepts, and significant discoveries.It almost seems like every day, inspiring observations come to light.Consider, for example, the possible link that was observed between gut bacteria of mice and Parkinson disease, in which changes in the bacteria populating the gut apparently appear to be associated with a decline in motor skills.How can this be?It turns out that 70% of all neurons in the peripheral nervous system are located in the gut, and these are directly connected to the central nervous system through the vagus nerve [1].A remarkable earlier study suggested that Parkinson disease may start in the stomach, because people who had their vagus nerve cut to treat gastric ulcers exhibited a lower risk of Parkinson disease than those whose treatment involved only a partial dissection [2].In another striking discovery, it was found that gut bacteria may have a role in autism and, curiously, there is evidence that a single species of gut bacteria can reverse autism-related social behavior in mice [3].In fact, there is emerging evidence for relationships between disruptions in the human microbiome and cancer [4], cardiovascular disease [5], obesity ([6] and more, e.g., [7]), food allergies [8], and asthma [9], among many other diseases.Beyond these emerging examples, there are well-established links (yet still quite recent!) between the diversity of a healthy gut microbiome and protection against Clostridium difficile infection, which has led to therapeutic interventions with remarkable success [10].The microbiota also has important roles in cancer therapy [11].Tumor growth can be suppressed by biofilm-producing bacteria [12].So how can computation accelerate research in the examples above?Broadly, computations can make headway in problems ranging from characterizing taxonomic diversity, classification of microbial species, and tracing their evolution.Computational methods are emerging to facilitate the detection and quantification of diverse patterns among these data, as well as the construction of microbial networks and cross interactions between members of microbial communities.Computation can tackle complex data (e.g., genomic, transcriptomic, proteomic, and metabolomic) on the interactions between microbial communities and their hosts, towards the most challenging question of the quantification of the impact of the human microbiome on our health, as in the examples above. Ruth Nussinov, Jason A. Papin |
PLoS Comput. Biol. | 2 |
| 2017 | Computing the Dynamic Supramolecular Structural ProteomeabstractCells execute their functions through protein interactions.The pathways they link are neither discrete nor spatially separated, as typically depicted in cellular diagrams.Cellular diagrams are useful; however, they neglect the physical structure of cell signaling.In reality, functions are shaped by molecular transitions between small-and large-supramolecular assemblies.Even though they are preorganized, they consist of clusters that are loose and dynamic.Importantly too, they are often anchored in the membrane and interact with scaffolding proteins and the cytoskeleton.Their continuum may physically span the cell.Indeed, efficient, productive, and reliable cell signaling can only take place through transient and cooperative protein-protein interactions, not through stochastic, diffusion-controlled processes.Despite this, current computational approaches to the modeling of the structural proteome still do not fully account for the in vivo, real physical cell organization.The enigmas of the assembly sizes and dynamic conformational distributions-and the diverse cellular environments that influence thempresent daunting challenges, which we have only begun to address.How will we then compute the realistic structural proteome in the next decade?How will we overcome the challenges that we confront, and which methods should we develop to meet them?Clearly, our views of protein structure and function have undergone a revolution.We no longer believe that a protein exists in only two distinct (active and inactive) states.We now recognize that even though a specific function is executed by a distinct active state, proteins (and other bio-macromolecules) exist in ensembles of states.The structure-function paradigm that now dominates molecular biology was inspired by physics and chemistry, which stipulate that even living things must abide by the laws of quantum mechanics and structural chemistry.This paradigm argues that biomolecules should be viewed-and described-statistically, not statically.Though challenging, eventually, to realistically capture the functional versatility and model the working proteome, we must consider conformational ensembles and their allosteric shifts, which result in changes to the populations of the conformations.Moreover, within this framework, the heterogeneous cellular environments, as well as allosteric covalent post-translational modifications, cannot be overlooked.Have we indeed treated the structural proteome as such in our computations?Determining the structures of protein assemblies has long been a vastly important aim of structural biology.The problem is challenging: a pair of protein structures can interact by complementary patches of surfaces.The patch size and identity are unknown, and it is difficult to assess which patches on one protein interact with which patches on the other.In principle, Ruth Nussinov, Jason A. Papin, Ilya A. Vakser |
PLoS Comput. Biol. | 2 |
| 2016 | Metabolic network-guided binning of metagenomic sequence fragmentsabstractMOTIVATION: Most microbes on Earth have never been grown in a laboratory, and can only be studied through DNA sequences. Environmental DNA sequence samples are complex mixtures of fragments from many different species, often unknown. There is a pressing need for methods that can reliably reconstruct genomes from complex metagenomic samples in order to address questions in ecology, bioremediation, and human health. RESULTS: We present the SOrting by NEtwork Completion (SONEC) approach for assigning reactions to incomplete metabolic networks based on a metabolite connectivity score. We successfully demonstrate proof of concept in a set of 100 genome-scale metabolic network reconstructions, and delineate the variables that impact reaction assignment accuracy. We further demonstrate the integration of SONEC with existing approaches (such as cross-sample scaffold abundance profile clustering) on a set of 94 metagenomic samples from the Human Microbiome Project. We show that not only does SONEC aid in reconstructing species-level genomes, but it also improves functional predictions made with the resulting metabolic networks. AVAILABILITY AND IMPLEMENTATION: The datasets and code presented in this work are available at: https://bitbucket.org/mattbiggs/sorting_by_network_completion/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthew B. Biggs, Jason A. Papin |
Bioinform. | 2 |
| 2016 | Computing BiologyabstractIs Computational Biology increasingly-and steadily-progressing toward addressing the mammoth challenge of actually computing biology?That is, have we reached the stage where we do not support biological research but drive it?This question is vitally important for allyoung and established computational biologists.Even though forecasting future research can be risky, we still venture to predict that the future will see considerably more research projects drifting toward this ambitious aspiration.Computational Biology is powerful for abstracting signatures of disease, for predicting it, and for proposing medications.It is effective in figuring out disease mechanisms and forceful in bridging experimental disciplines to obtain testable predictions.However, perhaps its biggest challenges lie in putting together the available broad and disparate information, devising tools to efficiently and effectively carry out these tasks while sifting through noise and recognizing cell specificity, and most importantly coming up with sound, coherent, and testable schemes.Among the examples of the complexity and the type of questions that we will increasingly face are the vast potential implications of findings that underscore the role of commensal microbiota in disease treatments.Further, the mechanisms which are involved-on the molecular level-are not understood and neither do we fully understand in detail how pathogens can modulate the host immune response and subvert it to their own advantage.We would like to identify-and understand-signatures of chronic inflammatory diseases; recurrent viral sequences in multiple patients and multiple cancers; we would like to map disease risks and to figure out what are the mechanisms for the distinct signaling of specific isoforms of oncogenic proteins in specific cancers.Analyses of cancer genomes points to driver mutations; however, we are baffled by the higher frequencies of specific mutations in certain tissues.We are also mystified by the complexity of the cellular network and its apparent redundancy which results in drug resistance.These are mere examples demonstrating the enormity of the questions which are facing us.Perhaps most tellingly is the fact that we are often even struggling to articulate specific objectives, to delineate the available knowledge and to evaluate apparent successes.Within the grand challenge to compute biology, efforts to compute human health fosters multidirectional and multidisciplinary integration of basic, patient-oriented, and populationbased research at different levels and across scales.Diverse data accumulate at increasingly high rates and quantities.We do not have a crystal ball for the future of research; we expect, however, that human health is going to be at the center.Human health is complex and encompasses multiple areas; thus, focusing on human health does not necessarily imply a narrower direction.Indeed, we expect that it will drive expansion, algorithm and tool development, and integration, all merging with experiments.Tools involving pattern recognition in images may Ruth Nussinov, Jason A. Papin |
PLoS Comput. Biol. | 2 |
| 2015 | From "What Is?" to "What Isn't?" Computational BiologyabstractISSN:1553-734X Ruth Nussinov, Sebastian Bonhoeffer, Jason A. Papin, Olaf Sporns |
PLoS Comput. Biol. | 3 |
| 2015 | Inference of Network Dynamics and Metabolic Interactions in the Gut MicrobiomeabstractWe present a novel methodology to construct a Boolean dynamic model from time series metagenomic information and integrate this modeling with genome-scale metabolic network reconstructions to identify metabolic underpinnings for microbial interactions. We apply this in the context of a critical health issue: clindamycin antibiotic treatment and opportunistic Clostridium difficile infection. Our model recapitulates known dynamics of clindamycin antibiotic treatment and C. difficile infection and predicts therapeutic probiotic interventions to suppress C. difficile infection. Genome-scale metabolic network reconstructions reveal metabolic differences between community members and are used to explore the role of metabolism in the observed microbial interactions. In vitro experimental data validate a key result of our computational model, that B. intestinihominis can in fact slow C. difficile growth. Steven N. Steinway, Matthew B. Biggs, Thomas P. Loughran Jr., Jason A. Papin, Réka Albert |
PLoS Comput. Biol. | 4 |
| 2014 | MetDraw: automated visualization of genome-scale metabolic network reconstructions and high-throughput dataabstractAbstract Motivation: Metabolic reaction maps allow visualization of genome-scale models and high-throughput data in a format familiar to many biologists. However, creating a map of a large metabolic model is a difficult and time-consuming process. MetDraw fully automates the map-drawing process for metabolic models containing hundreds to thousands of reactions. MetDraw can also overlay high-throughput ‘omics’ data directly on the generated maps. Availability and implementation: Web interface and source code are freely available at http://www.metdraw.com. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Paul A. Jensen, Jason A. Papin |
Bioinform. | 2 |
| 2011 | Functional integration of a metabolic network model and expression data without arbitrary thresholdingabstractMOTIVATION: Flux balance analysis (FBA) has been used extensively to analyze genome-scale, constraint-based models of metabolism in a variety of organisms. The predictive accuracy of such models has recently been improved through the integration of high-throughput expression profiles of metabolic genes and proteins. However, extensions of FBA often require that such data be discretized a priori into sets of genes or proteins that are either 'on' or 'off'. This procedure requires selecting relatively subjective expression thresholds, often requiring several iterations and refinements to capture the expression dynamics and retain model functionality. RESULTS: We present a method for mapping expression data from a set of environmental, genetic or temporal conditions onto a metabolic network model without the need for arbitrary expression thresholds. Metabolic Adjustment by Differential Expression (MADE) uses the statistical significance of changes in gene or protein expression to create a functional metabolic model that most accurately recapitulates the expression dynamics. MADE was used to generate a series of models that reflect the metabolic adjustments seen in the transition from fermentative- to glycerol-based respiration in Saccharomyces cerevisiae. The calculated gene states match 98.7% of possible changes in expression, and the resulting models capture functional characteristics of the metabolic shift. AVAILABILITY: MADE is implemented in Matlab and requires a mixed-integer linear program solver. Source code is freely available at http://www.bme.virginia.edu/csbl/downloads/. Paul A. Jensen, Jason A. Papin |
Bioinform. | 2 |
| 2011 | Reconciliation of Genome-Scale Metabolic Reconstructions for Comparative Systems AnalysisabstractIn the past decade, over 50 genome-scale metabolic reconstructions have been built for a variety of single- and multi- cellular organisms. These reconstructions have enabled a host of computational methods to be leveraged for systems-analysis of metabolism, leading to greater understanding of observed phenotypes. These methods have been sparsely applied to comparisons between multiple organisms, however, due mainly to the existence of differences between reconstructions that are inherited from the respective reconstruction processes of the organisms to be compared. To circumvent this obstacle, we developed a novel process, termed metabolic network reconciliation, whereby non-biological differences are removed from genome-scale reconstructions while keeping the reconstructions as true as possible to the underlying biological data on which they are based. This process was applied to two organisms of great importance to disease and biotechnological applications, Pseudomonas aeruginosa and Pseudomonas putida, respectively. The result is a pair of revised genome-scale reconstructions for these organisms that can be analyzed at a systems level with confidence that differences are indicative of true biological differences (to the degree that is currently known), rather than artifacts of the reconstruction process. The reconstructions were re-validated with various experimental data after reconciliation. With the reconciled and validated reconstructions, we performed a genome-wide comparison of metabolic flexibility between P. aeruginosa and P. putida that generated significant new insight into the underlying biology of these important organisms. Through this work, we provide a novel methodology for reconciling models, present new genome-scale reconstructions of P. aeruginosa and P. putida that can be directly compared at a network level, and perform a network-wide comparison of the two species. These reconstructions provide fresh insights into the metabolic similarities and differences between these important Pseudomonads, and pave the way towards full comparative analysis of genome-scale metabolic reconstructions of multiple species. Matthew A. Oberhardt, Jacek Puchalka, Vítor A. P. Martins dos Santos, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2009 | Functional States of the Genome-Scale Escherichia Coli Transcriptional Regulatory SystemabstractA transcriptional regulatory network (TRN) constitutes the collection of regulatory rules that link environmental cues to the transcription state of a cell's genome. We recently proposed a matrix formalism that quantitatively represents a system of such rules (a transcriptional regulatory system [TRS]) and allows systemic characterization of TRS properties. The matrix formalism not only allows the computation of the transcription state of the genome but also the fundamental characterization of the input-output mapping that it represents. Furthermore, a key advantage of this "pseudo-stoichiometric" matrix formalism is its ability to easily integrate with existing stoichiometric matrix representations of signaling and metabolic networks. Here we demonstrate for the first time how this matrix formalism is extendable to large-scale systems by applying it to the genome-scale Escherichia coli TRS. We analyze the fundamental subspaces of the regulatory network matrix (R) to describe intrinsic properties of the TRS. We further use Monte Carlo sampling to evaluate the E. coli transcription state across a subset of all possible environments, comparing our results to published gene expression data as validation. Finally, we present novel in silico findings for the E. coli TRS, including (1) a gene expression correlation matrix delineating functional motifs; (2) sets of gene ontologies for which regulatory rules governing gene transcription are poorly understood and which may direct further experimental characterization; and (3) the appearance of a distributed TRN structure, which is in stark contrast to the more hierarchical organization of metabolic networks. Erwin P. Gianchandani, Andrew R. Joyce, Bernhard O. Palsson, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2009 | Nano-motion Dynamics are Determined by Surface-Tethered Selectin Mechanokinetics and Bond FormationabstractThe interaction of proteins at cellular interfaces is critical for many biological processes, from intercellular signaling to cell adhesion. For example, the selectin family of adhesion receptors plays a critical role in trafficking during inflammation and immunosurveillance. Quantitative measurements of binding rates between surface-constrained proteins elicit insight into how molecular structural details and post-translational modifications contribute to function. However, nano-scale transport effects can obfuscate measurements in experimental assays. We constructed a biophysical simulation of the motion of a rigid microsphere coated with biomolecular adhesion receptors in shearing flow undergoing thermal motion. The simulation enabled in silico investigation of the effects of kinetic force dependence, molecular deformation, grouping adhesion receptors into clusters, surface-constrained bond formation, and nano-scale vertical transport on outputs that directly map to observable motions. Simulations recreated the jerky, discrete stop-and-go motions observed in P-selectin/PSGL-1 microbead assays with physiologic ligand densities. Motion statistics tied detailed simulated motion data to experimentally reported quantities. New deductions about biomolecular function for P-selectin/PSGL-1 interactions were made. Distributing adhesive forces among P-selectin/PSGL-1 molecules closely grouped in clusters was necessary to achieve bond lifetimes observed in microbead assays. Initial, capturing bond formation effectively occurred across the entire molecular contour length. However, subsequent rebinding events were enhanced by the reduced separation distance following the initial capture. The result demonstrates that vertical transport can contribute to an enhancement in the apparent bond formation rate. A detailed analysis of in silico motions prompted the proposition of wobble autocorrelation as an indicator of two-dimensional function. Insight into two-dimensional bond formation gained from flow cell assays might therefore be important to understand processes involving extended cellular interactions, such as immunological synapse formation. A biologically informative in silico system was created with minimal, high-confidence inputs. Incorporating random effects in surface separation through thermal motion enabled new deductions of the effects of surface-constrained biomolecular function. Important molecular information is embedded in the patterns and statistics of motion. Brian J. Schmidt, Jason A. Papin, Michael B. Lawrence |
PLoS Comput. Biol. | 2 |
| 2008 | Novel pathway compendium analysis elucidates mechanism of pro-angiogenic synthetic small moleculeabstractAbstract Motivation: Computational techniques have been applied to experimental datasets to identify drug mode-of-action. A shortcoming of existing approaches is the requirement of large reference databases of compound expression profiles. Here, we developed a new pathway-based compendium analysis that couples multi-timepoint, controlled microarray data for a single compound with systems-based network analysis to elucidate drug mechanism more efficiently. Results: We applied this approach to a transcriptional regulatory footprint of phthalimide neovascular factor 1 (PNF1)—a novel synthetic small molecule that exhibits significant in vitro endothelial potency—spanning 1–48 h post-supplementation in human micro-vascular endothelial cells (HMVEC) to comprehensively interrogate PNF1 effects. We concluded that PNF1 first induces tumor necrosis factor-alpha (TNF-α) signaling pathway function which in turn affects transforming growth factor-beta (TGF-β) signaling. These results are consistent with our previous observations of PNF1-directed TGF-β signaling at 24 h, including differential regulation of TGF-β-induced matrix metalloproteinase 14 (MMP14/MT1-MMP) which is implicated in angiogenesis. Ultimately, we illustrate how our pathway-based compendium analysis more efficiently generates hypotheses for compound mechanism than existing techniques. Availability: The microarray data generated as part of this study are available in the Gene Expression Omnibus (http://www.ncbi.nlm.nih.gov/geo/). Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Kristen A. Wieghaus, Erwin P. Gianchandani, Mikell A. Paige, Milton L. Brown, Edward A. Botchwey, Jason A. Papin |
Bioinform. | 6 |
| 2008 | Predicting biological system objectives de novo from internal state measurementsabstractBACKGROUND: Optimization theory has been applied to complex biological systems to interrogate network properties and develop and refine metabolic engineering strategies. For example, methods are emerging to engineer cells to optimally produce byproducts of commercial value, such as bioethanol, as well as molecular compounds for disease therapy. Flux balance analysis (FBA) is an optimization framework that aids in this interrogation by generating predictions of optimal flux distributions in cellular networks. Critical features of FBA are the definition of a biologically relevant objective function (e.g., maximizing the rate of synthesis of biomass, a unit of measurement of cellular growth) and the subsequent application of linear programming (LP) to identify fluxes through a reaction network. Despite the success of FBA, a central remaining challenge is the definition of a network objective with biological meaning. RESULTS: We present a novel method called Biological Objective Solution Search (BOSS) for the inference of an objective function of a biological system from its underlying network stoichiometry as well as experimentally-measured state variables. Specifically, BOSS identifies a system objective by defining a putative stoichiometric "objective reaction," adding this reaction to the existing set of stoichiometric constraints arising from known interactions within a network, and maximizing the putative objective reaction via LP, all the while minimizing the difference between the resultant in silico flux distribution and available experimental (e.g., isotopomer) flux data. This new approach allows for discovery of objectives with previously unknown stoichiometry, thus extending the biological relevance from earlier methods. We verify our approach on the well-characterized central metabolic network of Saccharomyces cerevisiae. CONCLUSION: We illustrate how BOSS offers insight into the functional organization of biochemical networks, facilitating the interrogation of cellular design principles and development of cellular engineering applications. Furthermore, we describe how growth is the best-fit objective function for the yeast metabolic network given experimentally-measured fluxes. Erwin P. Gianchandani, Matthew A. Oberhardt, Anthony P. Burgard, Costas D. Maranas, Jason A. Papin |
BMC Bioinform. | 5 |
| 2008 | Dynamic Analysis of Integrated Signaling, Metabolic, and Regulatory NetworksabstractExtracellular cues affect signaling, metabolic, and regulatory processes to elicit cellular responses. Although intracellular signaling, metabolic, and regulatory networks are highly integrated, previous analyses have largely focused on independent processes (e.g., metabolism) without considering the interplay that exists among them. However, there is evidence that many diseases arise from multifunctional components with roles throughout signaling, metabolic, and regulatory networks. Therefore, in this study, we propose a flux balance analysis (FBA)-based strategy, referred to as integrated dynamic FBA (idFBA), that dynamically simulates cellular phenotypes arising from integrated networks. The idFBA framework requires an integrated stoichiometric reconstruction of signaling, metabolic, and regulatory processes. It assumes quasi-steady-state conditions for "fast" reactions and incorporates "slow" reactions into the stoichiometric formalism in a time-delayed manner. To assess the efficacy of idFBA, we developed a prototypic integrated system comprising signaling, metabolic, and regulatory processes with network features characteristic of actual systems and incorporated kinetic parameters based on typical time scales observed in literature. idFBA was applied to the prototypic system, which was evaluated for different environments and gene regulatory rules. In addition, we applied the idFBA framework in a similar manner to a representative module of the single-cell eukaryotic organism Saccharomyces cerevisiae. Ultimately, idFBA facilitated quantitative, dynamic analysis of systemic effects of extracellular cues on cellular phenotypes and generated comparable time-course predictions when contrasted with an equivalent kinetic model. Since idFBA solves a linear programming problem and does not require an exhaustive list of detailed kinetic parameters, it may be efficiently scaled to integrated intracellular systems that incorporate signaling, metabolic, and regulatory processes at the genome scale, such as the S. cerevisiae system presented here. Jong Min Lee 0002, Erwin P. Gianchandani, James A. Eddy, Jason A. Papin |
PLoS Comput. Biol. | 4 |
| 2008 | Genome-Scale Reconstruction and Analysis of the Pseudomonas putida KT2440 Metabolic Network Facilitates Applications in BiotechnologyabstractA cornerstone of biotechnology is the use of microorganisms for the efficient production of chemicals and the elimination of harmful waste. Pseudomonas putida is an archetype of such microbes due to its metabolic versatility, stress resistance, amenability to genetic modifications, and vast potential for environmental and industrial applications. To address both the elucidation of the metabolic wiring in P. putida and its uses in biocatalysis, in particular for the production of non-growth-related biochemicals, we developed and present here a genome-scale constraint-based model of the metabolism of P. putida KT2440. Network reconstruction and flux balance analysis (FBA) enabled definition of the structure of the metabolic network, identification of knowledge gaps, and pin-pointing of essential metabolic functions, facilitating thereby the refinement of gene annotations. FBA and flux variability analysis were used to analyze the properties, potential, and limits of the model. These analyses allowed identification, under various conditions, of key features of metabolism such as growth yield, resource distribution, network robustness, and gene essentiality. The model was validated with data from continuous cell cultures, high-throughput phenotyping data, (13)C-measurement of internal flux distributions, and specifically generated knock-out mutants. Auxotrophy was correctly predicted in 75% of the cases. These systematic analyses revealed that the metabolic network structure is the main factor determining the accuracy of predictions, whereas biomass composition has negligible influence. Finally, we drew on the model to devise metabolic engineering strategies to improve production of polyhydroxyalkanoates, a class of biotechnologically useful compounds whose synthesis is not coupled to cell survival. The solidly validated model yields valuable insights into genotype-phenotype relationships and provides a sound framework to explore this versatile bacterium and to capitalize on its vast biotechnological potential. Jacek Puchalka, Matthew A. Oberhardt, Miguel Godinho, Agata Bielecka, Daniela Regenhardt, Kenneth N. Timmis, Jason A. Papin, Vítor A. P. Martins dos Santos |
PLoS Comput. Biol. | 7 |
| 2006 | Flux balance analysis in the era of metabolomicsabstractFlux balance analysis (FBA) has emerged as an effective means to analyse biological networks in a quantitative manner. Much progress has been made on the extension of FBA to incorporate a priori biological knowledge, provide more practical descriptions of observed cell behaviours, and predict the outcome of network perturbations. Metabolomics is independently advancing as a set of high-throughput data acquisition tools providing dynamic profiles of metabolites in an unbiased manner. These data sets are neither yet sufficiently comprehensive nor accurate enough for generating large-scale kinetic models. Thus, there is a pressing need to develop quantitative techniques that can make use of the emerging data and embrace the associated uncertainties. This article reviews recent advances in FBA to meet this need and discusses the utility of FBA as a complement to metabolomics and the expected synergy as a result of combining these two techniques. Jong Min Lee 0002, Erwin P. Gianchandani, Jason A. Papin |
Briefings Bioinform. | 3 |
| 2006 | Matrix Formalism to Describe Functional States of Transcriptional Regulatory SystemsabstractComplex regulatory networks control the transcription state of a genome. These transcriptional regulatory networks (TRNs) have been mathematically described using a Boolean formalism, in which the state of a gene is represented as either transcribed or not transcribed in response to regulatory signals. The Boolean formalism results in a series of regulatory rules for the individual genes of a TRN that in turn can be used to link environmental cues to the transcription state of a genome, thereby forming a complete transcriptional regulatory system (TRS). Herein, we develop a formalism that represents such a set of regulatory rules in a matrix form. Matrix formalism allows for the systemic characterization of the properties of a TRS and facilitates the computation of the transcriptional state of the genome under any given set of environmental conditions. Additionally, it provides a means to incorporate mechanistic detail of a TRS as it becomes available. In this study, the regulatory network matrix, R, for a prototypic TRS is characterized and the fundamental subspaces of this matrix are described. We illustrate how the matrix representation of a TRS coupled with its environment (R*) allows for a sampling of all possible expression states of a given network, and furthermore, how the fundamental subspaces of the matrix provide a way to study key TRS features and may assist in experimental design. Erwin P. Gianchandani, Jason A. Papin, Nathan D. Price 0001, Andrew R. Joyce, Bernhard O. Palsson |
PLoS Comput. Biol. | 2 |