VLDB 2026 Research / reviewers in the wild / expert
Costas D. Maranas
dblp:06/5713
· DBLP profile ↗
25ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-1508-1398ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 5 since 2021Theory of computation · 4 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BindPred: a framework for predicting protein-protein binding affinity from language model embeddingsabstractMOTIVATION: Reliable predictions of protein-protein binding affinities are essential for molecular biology and therapeutic discovery. However, most computational methods rely on three-dimensional structural models, which are often unavailable for many complexes. RESULTS: We introduce BindPred, a structure-agnostic input framework that predicts affinities directly from amino acid sequences by combining embeddings from large protein language models with gradient boosting trees. On the protein-protein binding (PPB)-Affinity benchmark, which comprises 11 919 diverse complexes, BindPred achieves a Pearson correlation coefficient of 0.86 in random split five-fold cross-validation. Ablation analysis indicates that evolutionary embeddings alone capture most of the predictive signals, while augmenting with physics-based energy terms from PyRosetta and BindCraft increases the correlation only by 0.01. A more stringent protein-level split that places entire protein families (wild-type and all mutants) exclusively in either training or testing sets, resulting in only a modest decline in performance, demonstrating robust generalization to novel interaction pairs. Because BindPred operates exclusively on sequence input, it enables rapid inference [approximately 3 million complexes per GPU (T4) hour], making proteome-scale screening computationally feasible. AVAILABILITY: The pretrained model and inference pipeline are available in a Google Colab notebook: BindPred Colab notebook. The training dataset, code, and model weights are available on the hugging face: https://huggingface.co/hbp5181/BindPred. Haixing Piao, Veda Sheersh Boorla, Somtirtha Santra, Costas D. Maranas |
Bioinform. | 4 |
| 2025 | Parameterization of cell-free systems with time-series data using KETCHUPabstractKinetic models mechanistically link enzyme levels, metabolite concentrations, and allosteric regulation to metabolic reaction fluxes. This coupling allows for the quantitative elucidation of the dynamics of the evolution of metabolite concentrations and metabolic fluxes as a function of time. So far, most large-scale kinetic model parameterizations are carried out using mostly steady-state flux measurements supplemented with metabolomics and/or proteomics data when available. Even though the parameterized kinetic model can trace a temporal evolution of the system, lack of anchoring to temporal data reduces confidence in the dynamics predictions. Notably, the simulation of enzymatic cascade reactions requires a full description of the dynamics of the system as a steady-state is not applicable given that all measured metabolite concentrations vary with time. Here we describe how kinetic parameters fitted to the dynamics of single-enzyme assays remain accurate for the simulation of multi-enzyme cell-free systems. Herein, we demonstrate two extensions for the Kinetic Estimation Tool Capturing Heterogeneous datasets Using Pyomo (KETCHUP) software tool for parameterizing a kinetic model of the cell-free kinetics of formate dehydrogenase (FDH) and 2,3-butanediol dehydrogenase (BDH) through the use of time-course data across various initial conditions. An implemented extension of KETCHUP allowing for the reconciliation of measurement time-lag errors present in datasets was used to parameterize kinetic models using multiple datasets. By combining the kinetic parameters identified by the FDH and BDH assays, accurate simulation of the binary FDH-BDH system was achieved. Syed Bilal Jilani, Daniel G. Olson, Costas D. Maranas |
PLoS Comput. Biol. | 4 |
| 2025 | novoStoic2.0: An integrated framework for pathway synthesis, thermodynamic evaluation, and enzyme selectionabstractComputational pathway design and retro-biosynthetic approaches can facilitate the development of innovative biochemical production routes, biodegradation strategies, and the funneling of multiple precursors into a single bioproduct. However, effective pathway design necessitates a comprehensive understanding of biochemistries, enzyme activities, and thermodynamic feasibility. Herein, we introduce novoStoic2.0, an integrated platform that combines tools for estimating overall stoichiometry, designing de novo synthesis pathways, assessing thermodynamic feasibility, and selecting enzymes. novoStoic2.0 offers a unified web-based interface as a part of the AlphaSynthesis platform (http://novostoic.platform.moleculemaker.org/) tailored for the synthesis of thermodynamically viable pathways as well as the selection of enzymes for re-engineering required for novel reaction steps. We exemplify the utility of the platform to identify novel pathways for hydroxytyrosol synthesis, which are shorter than the known pathways and require reduced cofactor usage. In summary, novoStoic2.0 aims to streamline the process of pathway design contributing to the development of sustainable biotechnological solutions. Vikas Upadhyay, Mohit Anand, Costas D. Maranas |
PLoS Comput. Biol. | 3 |
| 2021 | Elucidation of trophic interactions in an unusual single-cell nitrogen-fixing symbiosis using metabolic modelingabstractMarine nitrogen-fixing microorganisms are an important source of fixed nitrogen in oceanic ecosystems. The colonial cyanobacterium Trichodesmium and diatom symbionts were thought to be the primary contributors to oceanic N2 fixation until the discovery of the unusual uncultivated symbiotic cyanobacterium UCYN-A (Candidatus Atelocyanobacterium thalassa). UCYN-A has atypical metabolic characteristics lacking the oxygen-evolving photosystem II, the tricarboxylic acid cycle, the carbon-fixation enzyme RuBisCo and de novo biosynthetic pathways for a number of amino acids and nucleotides. Therefore, it is obligately symbiotic with its single-celled haptophyte algal host. UCYN-A receives fixed carbon from its host and returns fixed nitrogen, but further insights into this symbiosis are precluded by both UCYN-A and its host being uncultured. In order to investigate how this syntrophy is coordinated, we reconstructed bottom-up genome-scale metabolic models of UCYN-A and its algal partner to explore possible trophic scenarios, focusing on nitrogen fixation and biomass synthesis. Since both partners are uncultivated and only the genome sequence of UCYN-A is available, we used the phylogenetically related Chrysochromulina tobin as a proxy for the host. Through the use of flux balance analysis (FBA), we determined the minimal set of metabolites and biochemical functions that must be shared between the two organisms to ensure viability and growth. We quantitatively investigated the metabolic characteristics that facilitate daytime N2 fixation in UCYN-A and possible oxygen-scavenging mechanisms needed to create an anaerobic environment to allow nitrogenase to function. This is the first application of an FBA framework to examine the tight metabolic coupling between uncultivated microbes in marine symbiotic communities and provides a roadmap for future efforts focusing on such specialized systems. Debolina Sarkar, Marine Landa, Anindita Bandyopadhyay, Himadri B. Pakrasi, Jonathan P. Zehr, Costas D. Maranas |
PLoS Comput. Biol. | 6 |
| 2021 | dGPredictor: Automated fragmentation method for metabolic reaction free energy prediction and de novo pathway designabstractGroup contribution (GC) methods are conventionally used in thermodynamics analysis of metabolic pathways to estimate the standard Gibbs energy change (ΔrG'o) of enzymatic reactions from limited experimental measurements. However, these methods are limited by their dependence on manually curated groups and inability to capture stereochemical information, leading to low reaction coverage. Herein, we introduce an automated molecular fingerprint-based thermodynamic analysis tool called dGPredictor that enables the consideration of stereochemistry within metabolite structures and thus increases reaction coverage. dGPredictor has comparable prediction accuracy compared to existing GC methods and can capture Gibbs energy changes for isomerase and transferase reactions, which exhibit no overall group changes. We also demonstrate dGPredictor's ability to predict the Gibbs energy change for novel reactions and seamless integration within de novo metabolic pathway design tools such as novoStoic for safeguarding against the inclusion of reaction steps with infeasible directionalities. To facilitate easy access to dGPredictor, we developed a graphical user interface to predict the standard Gibbs energy change for reactions at various pH and ionic strengths. The tool allows customized user input of known metabolites as KEGG IDs and novel metabolites as InChI strings (https://github.com/maranasgroup/dGPredictor). Lin Wang 0031, Vikas Upadhyay, Costas D. Maranas |
PLoS Comput. Biol. | 3 |
| 2019 | From Escherichia coli mutant 13C labeling data to a core kinetic model: A kinetic model parameterization pipelineabstractKinetic models of metabolic networks offer the promise of quantitative phenotype prediction. The mechanistic characterization of enzyme catalyzed reactions allows for tracing the effect of perturbations in metabolite concentrations and reaction fluxes in response to genetic and environmental perturbation that are beyond the scope of stoichiometric models. In this study, we develop a two-step computational pipeline for the rapid parameterization of kinetic models of metabolic networks using a curated metabolic model and available 13C-labeling distributions under multiple genetic and environmental perturbations. The first step involves the elucidation of all intracellular fluxes in a core model of E. coli containing 74 reactions and 61 metabolites using 13C-Metabolic Flux Analysis (13C-MFA). Here, fluxes corresponding to the mid-exponential growth phase are elucidated for seven single gene deletion mutants from upper glycolysis, pentose phosphate pathway and the Entner-Doudoroff pathway. The computed flux ranges are then used to parameterize the same (i.e., k-ecoli74) core kinetic model for E. coli with 55 substrate-level regulations using the newly developed K-FIT parameterization algorithm. The K-FIT algorithm employs a combination of equation decomposition and iterative solution techniques to evaluate steady-state fluxes in response to genetic perturbations. k-ecoli74 predicted 86% of flux values for strains used during fitting within a single standard deviation of 13C-MFA estimated values. By performing both tasks using the same network, errors associated with lack of congruity between the two networks are avoided, allowing for seamless integration of data with model building. Product yield predictions and comparison with previously developed kinetic models indicate shifts in flux ranges and the presence or absence of mutant strains delivering flux towards pathways of interest from training data significantly impact predictive capabilities. Using this workflow, the impact of completeness of fluxomic datasets and the importance of specific genetic perturbations on uncertainties in kinetic parameter estimation are evaluated. Charles J. Foster, Saratram Gopalakrishnan, Maciek R. Antoniewicz, Costas D. Maranas |
PLoS Comput. Biol. | 4 |
| 2019 | A diurnal flux balance model of Synechocystis sp. PCC 6803 metabolismabstractPhototrophic organisms such as cyanobacteria utilize the sun's energy to convert atmospheric carbon dioxide into organic carbon, resulting in diurnal variations in the cell's metabolism. Flux balance analysis is a widely accepted constraint-based optimization tool for analyzing growth and metabolism, but it is generally used in a time-invariant manner with no provisions for sequestering different biomass components at different time periods. Here we present CycleSyn, a periodic model of Synechocystis sp. PCC 6803 metabolism that spans a 12-hr light/12-hr dark cycle by segmenting it into 12 Time Point Models (TPMs) with a uniform duration of two hours. The developed framework allows for the flow of metabolites across TPMs while inventorying metabolite levels and only allowing for the utilization of currently or previously produced compounds. The 12 TPMs allow for the incorporation of time-dependent constraints that capture the cyclic nature of cellular processes. Imposing bounds on reactions informed by temporally-segmented transcriptomic data enables simulation of phototrophic growth as a single linear programming (LP) problem. The solution provides the time varying reaction fluxes over a 24-hour cycle and the accumulation/consumption of metabolites. The diurnal rhythm of metabolic gene expression driven by the circadian clock and its metabolic consequences is explored. Predicted flux and metabolite pools are in line with published studies regarding the temporal organization of phototrophic growth in Synechocystis PCC 6803 paving the way for constructing time-resolved genome-scale models (GSMs) for organisms with a circadian clock. In addition, the metabolic reorganization that would be required to enable Synechocystis PCC 6803 to temporally separate photosynthesis from oxygen-sensitive nitrogen fixation is also explored using the developed model formalism. Debolina Sarkar, Thomas J. Mueller, Deng Liu, Himadri B. Pakrasi, Costas D. Maranas |
PLoS Comput. Biol. | 5 |
| 2018 | Accelerating flux balance calculations in genome-scale metabolic models by localizing the application of loopless constraintsabstractBackground: Genome-scale metabolic network models and constraint-based modeling techniques have become important tools for analyzing cellular metabolism. Thermodynamically infeasible cycles (TICs) causing unbounded metabolic flux ranges are often encountered. TICs satisfy the mass balance and directionality constraints but violate the second law of thermodynamics. Current practices involve implementing additional constraints to ensure not only optimal but also loopless flux distributions. However, the mixed integer linear programming problems required to solve become computationally intractable for genome-scale metabolic models. Results: We aimed to identify the fewest needed constraints sufficient for optimality under the loopless requirement. We found that loopless constraints are required only for the reactions that share elementary flux modes representing TICs with reactions that are part of the objective function. We put forth the concept of localized loopless constraints (LLCs) to enforce this minimal required set of loopless constraints. By combining with a novel procedure for minimal null-space calculation, the computational time for loopless flux variability analysis (ll-FVA) is reduced by a factor of 10-150 compared to the original loopless constraints and by 4-20 times compared to the current fastest method Fast-SNP with the percent improvement increasing with model size. Importantly, LLCs offer a scalable strategy for loopless flux calculations for multi-compartment/multi-organism models of large sizes, for example, shortening the CPU time for ll-FVA from 35 h to less than 2 h for a model with more than104 reactions. Availability and implementation: Matlab functions are available in the Supplementary Material or at https://github.com/maranasgroup/lll-FVA. Supplementary information: Supplementary data are available at Bioinformatics online. Siu Hung Joshua Chan, Lin Wang 0031, Satyakam Dash, Costas D. Maranas |
Bioinform. | 4 |
| 2017 | Standardizing biomass reactions and ensuring complete mass balance in genome-scale metabolic modelsabstractMOTIVATION: In a genome-scale metabolic model, the biomass produced is defined to have a molecular weight (MW) of 1 g mmol-1. This is critical for correctly predicting growth yields, contrasting multiple models and more importantly modeling microbial communities. However, the standard is rarely verified in the current practice and the chemical formulae of biomass components such as proteins, nucleic acids and lipids are often represented by undefined side groups (e.g. X, R). RESULTS: We introduced a systematic procedure for checking the biomass weight and ensuring complete mass balance of a model. We identified significant departures after examining 64 published models. The biomass weights of 34 models differed by 5-50%, while 8 models have discrepancies >50%. In total 20 models were manually curated. By maximizing the original versus corrected biomass reactions, flux balance analysis revealed >10% differences in growth yields for 12 of the curated models. Biomass MW discrepancies are accentuated in microbial community simulations as they can cause significant and systematic errors in the community composition. Microbes with underestimated biomass MWs are overpredicted in the community whereas microbes with overestimated biomass weights are underpredicted. The observed departures in community composition are disproportionately larger than the discrepancies in the biomass weight estimate. We propose the presented procedure as a standard practice for metabolic reconstructions. AVAILABILITY AND IMPLEMENTATION: The MALTAB and Python scripts are available in the Supplementary Material. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Siu Hung Joshua Chan, Jingyi Cai, Lin Wang 0031, Margaret N. Simons-Senftle, Costas D. Maranas |
Bioinform. | 5 |
| 2017 | SteadyCom: Predicting microbial abundances while ensuring community stabilityabstractGenome-scale metabolic modeling has become widespread for analyzing microbial metabolism. Extending this established paradigm to more complex microbial communities is emerging as a promising way to unravel the interactions and biochemical repertoire of these omnipresent systems. While several modeling techniques have been developed for microbial communities, little emphasis has been placed on the need to impose a time-averaged constant growth rate across all members for a community to ensure co-existence and stability. In the absence of this constraint, the faster growing organism will ultimately displace all other microbes in the community. This is particularly important for predicting steady-state microbiota composition as it imposes significant restrictions on the allowable community membership, composition and phenotypes. In this study, we introduce the SteadyCom optimization framework for predicting metabolic flux distributions consistent with the steady-state requirement. SteadyCom can be rapidly converged by iteratively solving linear programming (LP) problem and the number of iterations is independent of the number of organisms. A significant advantage of SteadyCom is compatibility with flux variability analysis. SteadyCom is first demonstrated for a community of four E. coli double auxotrophic mutants and is then applied to a gut microbiota model consisting of nine species, with representatives from the phyla Bacteroidetes, Firmicutes, Actinobacteria and Proteobacteria. In contrast to the direct use of FBA, SteadyCom is able to predict the change in species abundance in response to changes in diets with minimal additional imposed constraints on the model. By randomizing the uptake rates of microbes, an abundance profile with a good agreement to experimental gut microbiota is inferred. SteadyCom provides an important step towards the cross-cutting task of predicting the composition of a microbial community in a given environment. Siu Hung Joshua Chan, Margaret N. Simons-Senftle, Costas D. Maranas |
PLoS Comput. Biol. | 3 |
| 2014 | k-OptForce: Integrating Kinetics with Flux Balance Analysis for Strain DesignabstractComputational strain design protocols aim at the system-wide identification of intervention strategies for the enhanced production of biochemicals in microorganisms. Existing approaches relying solely on stoichiometry and rudimentary constraint-based regulation overlook the effects of metabolite concentrations and substrate-level enzyme regulation while identifying metabolic interventions. In this paper, we introduce k-OptForce, which integrates the available kinetic descriptions of metabolic steps with stoichiometric models to sharpen the prediction of intervention strategies for improving the bio-production of a chemical of interest. It enables identification of a minimal set of interventions comprised of both enzymatic parameter changes (for reactions with available kinetics) and reaction flux changes (for reactions with only stoichiometric information). Application of k-OptForce to the overproduction of L-serine in E. coli and triacetic acid lactone (TAL) in S. cerevisiae revealed that the identified interventions tend to cause less dramatic rearrangements of the flux distribution so as not to violate concentration bounds. In some cases the incorporation of kinetic information leads to the need for additional interventions as kinetic expressions render stoichiometry-only derived interventions infeasible by violating concentration bounds, whereas in other cases the kinetic expressions impart flux changes that favor the overproduction of the target product thereby requiring fewer direct interventions. A sensitivity analysis on metabolite concentrations shows that the required number of interventions can be significantly affected by changing the imposed bounds on metabolite concentrations. Furthermore, k-OptForce was capable of finding non-intuitive interventions aiming at alleviating the substrate-level inhibition of key enzymes in order to enhance the flux towards the product of interest, which cannot be captured by stoichiometry-alone analysis. This study paves the way for the integrated analysis of kinetic and stoichiometric models and enables elucidating system-wide metabolic interventions while capturing regulatory and kinetic effects. Anupam Chowdhury, Ali R. Zomorrodi, Costas D. Maranas |
PLoS Comput. Biol. | 3 |
| 2013 | MAPs: a database of modular antibody parts for predicting tertiary structures and designing affinity matured antibodiesabstractBACKGROUND: The de novo design of a novel protein with a particular function remains a formidable challenge with only isolated and hard-to-repeat successes to date. Due to their many structurally conserved features, antibodies are a family of proteins amenable to predictable rational design. Design algorithms must consider the structural diversity of possible naturally occurring antibodies. The human immune system samples this design space (2 1012) by randomly combining variable, diversity, and joining genes in a process known as V-(D)-J recombination. DESCRIPTION: By analyzing structural features found in affinity matured antibodies, a database of Modular Antibody Parts (MAPs) analogous to the variable, diversity, and joining genes has been constructed for the prediction of antibody tertiary structures. The database contains 929 parts constructed from an analysis of 1168 human, humanized, chimeric, and mouse antibody structures and encompasses all currently observed structural diversity of antibodies. CONCLUSIONS: The generation of 260 antibody structures shows that the MAPs database can be used to reliably predict antibody tertiary structures with an average all-atom RMSD of 1.9 Å. Using the broadly neutralizing anti-influenza antibody CH65 and anti-HIV antibody 4E10 as examples, promising starting antibodies for affinity maturation are identified and amino acid changes are traced as antibody affinity maturation occurs. Robert J. Pantazes, Costas D. Maranas |
BMC Bioinform. | 2 |
| 2012 | MetRxn: a knowledgebase of metabolites and reactions spanning metabolic models and databasesabstractBACKGROUND: Increasingly, metabolite and reaction information is organized in the form of genome-scale metabolic reconstructions that describe the reaction stoichiometry, directionality, and gene to protein to reaction associations. A key bottleneck in the pace of reconstruction of new, high-quality metabolic models is the inability to directly make use of metabolite/reaction information from biological databases or other models due to incompatibilities in content representation (i.e., metabolites with multiple names across databases and models), stoichiometric errors such as elemental or charge imbalances, and incomplete atomistic detail (e.g., use of generic R-group or non-explicit specification of stereo-specificity). DESCRIPTION: MetRxn is a knowledgebase that includes standardized metabolite and reaction descriptions by integrating information from BRENDA, KEGG, MetaCyc, Reactome.org and 44 metabolic models into a single unified data set. All metabolite entries have matched synonyms, resolved protonation states, and are linked to unique structures. All reaction entries are elementally and charge balanced. This is accomplished through the use of a workflow of lexicographic, phonetic, and structural comparison algorithms. MetRxn allows for the download of standardized versions of existing genome-scale metabolic models and the use of metabolic information for the rapid reconstruction of new ones. CONCLUSIONS: The standardization in description allows for the direct comparison of the metabolite and reaction content between metabolic models and databases and the exhaustive prospecting of pathways for biotechnological production. This ever-growing dataset currently consists of over 76,000 metabolites participating in more than 72,000 reactions (including unresolved entries). MetRxn is hosted on a web-based platform that uses relational database models (MySQL). Akhil Kumar 0003, Patrick F. Suthers, Costas D. Maranas |
BMC Bioinform. | 3 |
| 2012 | Impact of Stoichiometry Representation on Simulation of Genotype-Phenotype Relationships in Metabolic NetworksabstractGenome-scale metabolic networks provide a comprehensive structural framework for modeling genotype-phenotype relationships through flux simulations. The solution space for the metabolic flux state of the cell is typically very large and optimization-based approaches are often necessary for predicting the active metabolic state under specific environmental conditions. The objective function to be used in such optimization algorithms is directly linked with the biological hypothesis underlying the model and therefore it is one of the most relevant parameters for successful modeling. Although linear combination of selected fluxes is widely used for formulating metabolic objective functions, we show that the resulting optimization problem is sensitive towards stoichiometry representation of the metabolic network. This undesirable sensitivity leads to different simulation results when using numerically different but biochemically equivalent stoichiometry representations and thereby makes biological interpretation intrinsically subjective and ambiguous. We hereby propose a new method, Minimization of Metabolites Balance (MiMBl), which decouples the artifacts of stoichiometry representation from the formulation of the desired objective functions, by casting objective functions using metabolite turnovers rather than fluxes. By simulating perturbed metabolic networks, we demonstrate that the use of stoichiometry representation independent algorithms is fundamental for unambiguously linking modeling results with biological interpretation. For example, MiMBl allowed us to expand the scope of metabolic modeling in elucidating the mechanistic basis of several genetic interactions in Saccharomyces cerevisiae. Ana Rita Brochado, Sergej Andrejev, Costas D. Maranas, Kiran Raosaheb Patil |
PLoS Comput. Biol. | 3 |
| 2012 | OptCom: A Multi-Level Optimization Framework for the Metabolic Modeling and Analysis of Microbial CommunitiesabstractMicroorganisms rarely live isolated in their natural environments but rather function in consolidated and socializing communities. Despite the growing availability of high-throughput sequencing and metagenomic data, we still know very little about the metabolic contributions of individual microbial players within an ecological niche and the extent and directionality of interactions among them. This calls for development of efficient modeling frameworks to shed light on less understood aspects of metabolism in microbial communities. Here, we introduce OptCom, a comprehensive flux balance analysis framework for microbial communities, which relies on a multi-level and multi-objective optimization formulation to properly describe trade-offs between individual vs. community level fitness criteria. In contrast to earlier approaches that rely on a single objective function, here, we consider species-level fitness criteria for the inner problems while relying on community-level objective maximization for the outer problem. OptCom is general enough to capture any type of interactions (positive, negative or combinations thereof) and is capable of accommodating any number of microbial species (or guilds) involved. We applied OptCom to quantify the syntrophic association in a well-characterized two-species microbial system, assess the level of sub-optimal growth in phototrophic microbial mats, and elucidate the extent and direction of inter-species metabolite and electron transfer in a model microbial community. We also used OptCom to examine addition of a new member to an existing community. Our study demonstrates the importance of trade-offs between species- and community-level fitness driving forces and lays the foundation for metabolic-driven analysis of various types of interactions in multi-species microbial systems using genome-scale metabolic models. Ali R. Zomorrodi, Costas D. Maranas |
PLoS Comput. Biol. | 2 |
| 2010 | OptForce: An Optimization Procedure for Identifying All Genetic Manipulations Leading to Targeted OverproductionsabstractComputational procedures for predicting metabolic interventions leading to the overproduction of biochemicals in microbial strains are widely in use. However, these methods rely on surrogate biological objectives (e.g., maximize growth rate or minimize metabolic adjustments) and do not make use of flux measurements often available for the wild-type strain. In this work, we introduce the OptForce procedure that identifies all possible engineering interventions by classifying reactions in the metabolic model depending upon whether their flux values must increase, decrease or become equal to zero to meet a pre-specified overproduction target. We hierarchically apply this classification rule for pairs, triples, quadruples, etc. of reactions. This leads to the identification of a sufficient and non-redundant set of fluxes that must change (i.e., MUST set) to meet a pre-specified overproduction target. Starting with this set we subsequently extract a minimal set of fluxes that must actively be forced through genetic manipulations (i.e., FORCE set) to ensure that all fluxes in the network are consistent with the overproduction objective. We demonstrate our OptForce framework for succinate production in Escherichia coli using the most recent in silico E. coli model, iAF1260. The method not only recapitulates existing engineering strategies but also reveals non-intuitive ones that boost succinate production by performing coordinated changes on pathways distant from the last steps of succinate synthesis. Sridhar Ranganathan, Patrick F. Suthers, Costas D. Maranas |
PLoS Comput. Biol. | 3 |
| 2009 | GrowMatch: An Automated Method for Reconciling In Silico/In Vivo Growth PredictionsabstractGenome-scale metabolic reconstructions are typically validated by comparing in silico growth predictions across different mutants utilizing different carbon sources with in vivo growth data. This comparison results in two types of model-prediction inconsistencies; either the model predicts growth when no growth is observed in the experiment (GNG inconsistencies) or the model predicts no growth when the experiment reveals growth (NGG inconsistencies). Here we propose an optimization-based framework, GrowMatch, to automatically reconcile GNG predictions (by suppressing functionalities in the model) and NGG predictions (by adding functionalities to the model). We use GrowMatch to resolve inconsistencies between the predictions of the latest in silico Escherichia coli (iAF1260) model and the in vivo data available in the Keio collection and improved the consistency of in silico with in vivo predictions from 90.6% to 96.7%. Specifically, we were able to suggest consistency-restoring hypotheses for 56/72 GNG mutants and 13/38 NGG mutants. GrowMatch resolved 18 GNG inconsistencies by suggesting suppressions in the mutant metabolic networks. Fifteen inconsistencies were resolved by suppressing isozymes in the metabolic network, and the remaining 23 GNG mutants corresponding to blocked genes were resolved by suitably modifying the biomass equation of iAF1260. GrowMatch suggested consistency-restoring hypotheses for five NGG mutants by adding functionalities to the model whereas the remaining eight inconsistencies were resolved by pinpointing possible alternate genes that carry out the function of the deleted gene. For many cases, GrowMatch identified fairly nonintuitive model modification hypotheses that would have been difficult to pinpoint through inspection alone. In addition, GrowMatch can be used during the construction phase of new, as opposed to existing, genome-scale metabolic models, leading to more expedient and accurate reconstructions. Vinay Satish Kumar, Costas D. Maranas |
PLoS Comput. Biol. | 2 |
| 2009 | A Genome-Scale Metabolic Reconstruction of Mycoplasma genitalium, iPS189abstractWith a genome size of approximately 580 kb and approximately 480 protein coding regions, Mycoplasma genitalium is one of the smallest known self-replicating organisms and, additionally, has extremely fastidious nutrient requirements. The reduced genomic content of M. genitalium has led researchers to suggest that the molecular assembly contained in this organism may be a close approximation to the minimal set of genes required for bacterial growth. Here, we introduce a systematic approach for the construction and curation of a genome-scale in silico metabolic model for M. genitalium. Key challenges included estimation of biomass composition, handling of enzymes with broad specificities, and the lack of a defined medium. Computational tools were subsequently employed to identify and resolve connectivity gaps in the model as well as growth prediction inconsistencies with gene essentiality experimental data. The curated model, M. genitalium iPS189 (262 reactions, 274 metabolites), is 87% accurate in recapitulating in vivo gene essentiality results for M. genitalium. Approaches and tools described herein provide a roadmap for the automated construction of in silico metabolic models of other organisms. Patrick F. Suthers, Madhukar S. Dasika, Vinay Satish Kumar, Gennady Denisov, John I. Glass, Costas D. Maranas |
PLoS Comput. Biol. | 6 |
| 2008 | Predicting biological system objectives de novo from internal state measurementsabstractBACKGROUND: Optimization theory has been applied to complex biological systems to interrogate network properties and develop and refine metabolic engineering strategies. For example, methods are emerging to engineer cells to optimally produce byproducts of commercial value, such as bioethanol, as well as molecular compounds for disease therapy. Flux balance analysis (FBA) is an optimization framework that aids in this interrogation by generating predictions of optimal flux distributions in cellular networks. Critical features of FBA are the definition of a biologically relevant objective function (e.g., maximizing the rate of synthesis of biomass, a unit of measurement of cellular growth) and the subsequent application of linear programming (LP) to identify fluxes through a reaction network. Despite the success of FBA, a central remaining challenge is the definition of a network objective with biological meaning. RESULTS: We present a novel method called Biological Objective Solution Search (BOSS) for the inference of an objective function of a biological system from its underlying network stoichiometry as well as experimentally-measured state variables. Specifically, BOSS identifies a system objective by defining a putative stoichiometric "objective reaction," adding this reaction to the existing set of stoichiometric constraints arising from known interactions within a network, and maximizing the putative objective reaction via LP, all the while minimizing the difference between the resultant in silico flux distribution and available experimental (e.g., isotopomer) flux data. This new approach allows for discovery of objectives with previously unknown stoichiometry, thus extending the biological relevance from earlier methods. We verify our approach on the well-characterized central metabolic network of Saccharomyces cerevisiae. CONCLUSION: We illustrate how BOSS offers insight into the functional organization of biochemical networks, facilitating the interrogation of cellular design principles and development of cellular engineering applications. Furthermore, we describe how growth is the best-fit objective function for the yeast metabolic network given experimentally-measured fluxes. Erwin P. Gianchandani, Matthew A. Oberhardt, Anthony P. Burgard, Costas D. Maranas, Jason A. Papin |
BMC Bioinform. | 4 |
| 2007 | Optimization based automated curation of metabolic reconstructionsabstractBACKGROUND: Currently, there exists tens of different microbial and eukaryotic metabolic reconstructions (e.g., Escherichia coli, Saccharomyces cerevisiae, Bacillus subtilis) with many more under development. All of these reconstructions are inherently incomplete with some functionalities missing due to the lack of experimental and/or homology information. A key challenge in the automated generation of genome-scale reconstructions is the elucidation of these gaps and the subsequent generation of hypotheses to bridge them. RESULTS: In this work, an optimization based procedure is proposed to identify and eliminate network gaps in these reconstructions. First we identify the metabolites in the metabolic network reconstruction which cannot be produced under any uptake conditions and subsequently we identify the reactions from a customized multi-organism database that restores the connectivity of these metabolites to the parent network using four mechanisms. This connectivity restoration is hypothesized to take place through four mechanisms: a) reversing the directionality of one or more reactions in the existing model, b) adding reaction from another organism to provide functionality absent in the existing model, c) adding external transport mechanisms to allow for importation of metabolites in the existing model and d) restore flow by adding intracellular transport reactions in multi-compartment models. We demonstrate this procedure for the genome- scale reconstruction of Escherichia coli and also Saccharomyces cerevisiae wherein compartmentalization of intra-cellular reactions results in a more complex topology of the metabolic network. We determine that about 10% of metabolites in E. coli and 30% of metabolites in S. cerevisiae cannot carry any flux. Interestingly, the dominant flow restoration mechanism is directionality reversals of existing reactions in the respective models. CONCLUSION: We have proposed systematic methods to identify and fill gaps in genome-scale metabolic reconstructions. The identified gaps can be filled both by making modifications in the existing model and by adding missing reactions by reconciling multi-organism databases of reactions with existing genome-scale models. Computational results provide a list of hypotheses to be queried further and tested experimentally. Vinay Satish Kumar, Madhukar S. Dasika, Costas D. Maranas |
BMC Bioinform. | 3 |
| 2006 | Elucidation of directionality for co-expressed genes: predicting intra-operon termination sitesabstractMOTIVATION: In this paper, we present a novel framework for inferring regulatory and sequence-level information from gene co-expression networks. The key idea of our methodology is the systematic integration of network inference and network topological analysis approaches for uncovering biological insights. RESULTS: We determine the gene co-expression network of Bacillus subtilis using Affymetrix GeneChip time-series data and show how the inferred network topology can be linked to sequence-level information hard-wired in the organism's genome. We propose a systematic way for determining the correlation threshold at which two genes are assessed to be co-expressed using the clustering coefficient and we expand the scope of the gene co-expression network by proposing the slope ratio metric as a means for incorporating directionality on the edges. We show through specific examples for B. subtilis that by incorporating expression level information in addition to the temporal expression patterns, we can uncover sequence-level biological insights. In particular, we are able to identify a number of cases where (1) the co-expressed genes are part of a single transcriptional unit or operon and (2) the inferred directionality arises due to the presence of intra-operon transcription termination sites. AVAILABILITY: The software will be provided on request. SUPPLEMENTARY INFORMATION: http://www.phys.psu.edu/~ralbert/pdf/gma_bioinf_supp.pdf Anshuman Gupta, Costas D. Maranas, Réka Albert |
Bioinform. | 2 |
| 1997 | Prediction of Oligopeptide Conformations via Deterministic Global Optimization
Ioannis P. Androulakis, Costas D. Maranas, Christodoulos A. Floudas |
J. Glob. Optim. | 2 |
| 1995 | αBB: A global optimization method for general constrained nonconvex problems
Ioannis P. Androulakis, Costas D. Maranas, Christodoulos A. Floudas |
J. Glob. Optim. | 2 |
| 1995 | Finding all solutions of nonlinearly constrained systems of equations
Costas D. Maranas, Christodoulos A. Floudas |
J. Glob. Optim. | 1 |
| 1994 | Global minimum potential energy conformations of small molecules
Costas D. Maranas, Christodoulos A. Floudas |
J. Glob. Optim. | 1 |