Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Peter D. Karp

dblp:68/4051 · DBLP profile ↗
← Back
51ranked-venue papers
21as first author
2since 2021 · last 2021
0000-0002-5876-6418ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 44 · 15 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 5 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
22 papers
Bioinformatics and computational biology · 99% Computational science and engineering · 1%
Artificial intelligence
3 papers
Knowledge representation and reasoning · 100%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › systems biology
metabolic network reconstruction
0.422020
Taxonomic weighting improves the accuracy of a gap-filling algorithm for metabolic models · Bioinform. 2020
Evaluation of computational metabolic-pathway predictions for Helicobacter pylori · Bioinform. 2002
Bioinformatics and computational biology › systems biology › metabolic network reconstruction
gap filling
0.412020
Taxonomic weighting improves the accuracy of a gap-filling algorithm for metabolic models · Bioinform. 2020
Bioinformatics and computational biology › systems biology
metabolic network analysis
0.432014
Optimal metabolic route search based on atom mappings · Bioinform. 2014
Construction and completion of flux balance models from pathway databases · Bioinform. 2012
Annotation-based inference of transporter function · ISMB 2008
Bioinformatics and computational biology › systems biology
metabolic modeling
0.212016
Representation and inference of cellular architecture for metabolic reconstruction and modeling · Bioinform. 2016
Bioinformatics and computational biology › molecular informatics › cheminformatics
atom mapping
0.212014
Optimal metabolic route search based on atom mappings · Bioinform. 2014
Bioinformatics and computational biology › systems bioinformatics
metabolic route search
0.212014
Optimal metabolic route search based on atom mappings · Bioinform. 2014
Bioinformatics and computational biology › systems biology › constraint-based modeling
flux balance analysis
0.112012
Construction and completion of flux balance models from pathway databases · Bioinform. 2012
Bioinformatics and computational biology
genome annotation
0.122008
Annotation-based inference of transporter function · ISMB 2008
Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases · Bioinform. 2004
Bioinformatics and computational biology
comparative genomics
0.112011
Discovering novel subsystems using comparative genomics · Bioinform. 2011
Bioinformatics and computational biology
functional genomics
0.112011
Discovering novel subsystems using comparative genomics · Bioinform. 2011
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway discovery
0.112011
Discovering novel subsystems using comparative genomics · Bioinform. 2011
Bioinformatics and computational biology › biological database
pathway/genome database
0.122004
Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases · Bioinform. 2004
The Pathway Tools software · ISMB 2002
Bioinformatics and computational biology › ontology
ontology development
0.112016
Representation and inference of cellular architecture for metabolic reconstruction and modeling · Bioinform. 2016
Bioinformatics and computational biology
metabolic pathway
0.142002
The Pathway Tools software · ISMB 2002
HinCyc: A Knowledge Base of the Complete Genome and Metabolic Pathways of H. influenzae · ISMB 1996
Representations of Metabolic Knowledge: Pathways · ISMB 1994
Bioinformatics and computational biology › biological database
pathway database
0.122005
Querying and computing with BioCyc databases · Bioinform. 2005
HinCyc: A Knowledge Base of the Complete Genome and Metabolic Pathways of H. influenzae · ISMB 1996
Bioinformatics and computational biology
biological database
0.112005
Querying and computing with BioCyc databases · Bioinform. 2005
Bioinformatics and computational biology
genomics
0.012004
Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases · Bioinform. 2004
Bioinformatics and computational biology › systems biology
metabolic network
0.012004
Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases · Bioinform. 2004
Bioinformatics and computational biology › genome annotation
transcriptional unit prediction
0.012004
Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases · Bioinform. 2004
Bioinformatics and computational biology › systems biology › metabolic network analysis
metabolic pathway prediction
0.012002
Evaluation of computational metabolic-pathway predictions for Helicobacter pylori · Bioinform. 2002
Bioinformatics and computational biology › biological database
biological database curation
0.012001
Database verification studies of SWISS-PROT and GenBank · Bioinform. 2001
Bioinformatics and computational biology › knowledge representation in biology
biomedical ontology
0.012000
An ontology for biological function based on molecular interactions · Bioinform. 2000
Bioinformatics and computational biology
sequence analysis
0.011998
What we do not know about sequence analysis and sequence databases · Bioinform. 1998
Bioinformatics and computational biology › biological database
sequence database
0.011998
What we do not know about sequence analysis and sequence databases · Bioinform. 1998
Computational science and engineering
computational chemistry
0.011997
Estimation of equilibrium constants using automated group contribution methods · Comput. Appl. Biosci. 1997
Computational science and engineering › computational chemistry › molecular simulation › molecular dynamics
free energy estimation
0.011997
Estimation of equilibrium constants using automated group contribution methods · Comput. Appl. Biosci. 1997
Bioinformatics and computational biology › sequence analysis
sequence classification
0.011997
Prediction of Enzyme Classification from Protein Sequence without the Use of Sequence Similarity · ISMB 1997
Bioinformatics and computational biology › genomics › genomic data management
genome database
0.022002
The Pathway Tools software · ISMB 2002
HinCyc: A Knowledge Base of the Complete Genome and Metabolic Pathways of H. influenzae · ISMB 1996
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge management
knowledge sharing
0.011995
The Generic Frame Protocol · IJCAI (1) 1995
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.012003
Knowledge acquisition, consistency checking and concurrency control for Gene Ontology (GO) · Bioinform. 2003

Methods — techniques the papers use, named apart from their topics

taxonomic weighting · 0.4optimization-based gap-filling · 0.4ontology inference · 0.2branch-and-bound search · 0.2mixed integer linear programming · 0.1genome context analysis · 0.1clustering · 0.1name-based annotation analysis · 0.1SQL · 0.1API access · 0.1protégé-2000 · 0.0protégé axiom language · 0.0prompt · 0.0graph representation · 0.0global optimization for classification upgrade · 0.0knowledge representation · 0.0generic frame protocol · 0.0
YearPublicationVenuePosition
2021 Pathway Tools version 23.0 update: software for pathway/genome informatics and systems biology
abstract
Abstract Motivation Biological systems function through dynamic interactions among genes and their products, regulatory circuits and metabolic networks. Our development of the Pathway Tools software was motivated by the need to construct biological knowledge resources that combine these many types of data, and that enable users to find and comprehend data of interest as quickly as possible through query and visualization tools. Further, we sought to support the development of metabolic flux models from pathway databases, and to use pathway information to leverage the interpretation of high-throughput data sets. Results In the past 4 years we have enhanced the already extensive Pathway Tools software in several respects. It can now support metabolic-model execution through the Web, it provides a more accurate gap filler for metabolic models; it supports development of models for organism communities distributed across a spatial grid; and model results may be visualized graphically. Pathway Tools supports several new omics-data analysis tools including the Omics Dashboard, multi-pathway diagrams called pathway collages, a pathway-covering algorithm for metabolomics data analysis and an algorithm for generating mechanistic explanations of multi-omics data. We have also improved the core pathway/genome databases management capabilities of the software, providing new multi-organism search tools for organism communities, improved graphics rendering, faster performance and re-designed gene and metabolite pages. Availability The software is free for academic use; a fee is required for commercial use. See http://pathwaytools.com. Contact [email protected] Supplementary information Supplementary data are available at Briefings in Bioinformatics online.
Peter D. Karp, Peter E. Midford, Richard Billington, Anamika Kothari, Markus Krummenacker, Mario Latendresse, Wai Kit Ong, Pallavi Subhraveti, Ron Caspi, Carol A. Fulcher, Ingrid M. Keseler, Suzanne M. Paley
Briefings Bioinform.1
2021 The BioCyc Metabolic Network Explorer
abstract
BACKGROUND: The Metabolic Network Explorer is a new addition to the BioCyc.org website and the Pathway Tools software suite that supports the interactive exploration of metabolic networks. Any metabolic network visualization tool must by necessity show only a subset of all possible metabolite connections, or the results will be visually overwhelming. Existing tools, even those that purport to show an organism's full metabolic network, limit the set of displayed connections based on predefined pathways or other preselected criteria. We sought instead to provide a tool that would give the user dynamic control over which connections to follow. RESULTS: The Metabolic Network Explorer is an easy-to-use, web-based software tool that allows the user to specify a starting metabolite of interest and interactively explore its immediate metabolic neighborhood in either or both directions to any desired depth, letting the user select from the full set of connected reactions. Although, as for other tools, only a small portion of the metabolic network is visible at a time, that portion is selected by the user, based on the full reaction complement, and it is easy to switch among alternate paths of interest. The display is intuitive, customizable, and provides copious links to more detailed information pages. CONCLUSIONS: The Metabolic Network Explorer fills a gap in the set of metabolic network visualization tools and complements other modes of exploration. Its primary strengths are its ease of use, diagrams that are intuitive to biologists, and its integration with the broader corpus of data provided by a BioCyc Pathway/Genome Database.
Suzanne M. Paley, Peter D. Karp
BMC Bioinform.2
2020 Taxonomic weighting improves the accuracy of a gap-filling algorithm for metabolic models
abstract
MOTIVATION: The increasing availability of annotated genome sequences enables construction of genome-scale metabolic networks, which are useful tools for studying organisms of interest. However, due to incomplete genome annotations, draft metabolic models contain gaps that must be filled in a time-consuming process before they are usable. Optimization-based algorithms that fill these gaps have been developed, however, gap-filling algorithms show significant error rates and often introduce incorrect reactions. RESULTS: Here, we present a new gap-filling method that computes the costs of candidate gap-filling reactions from a universal reaction database (MetaCyc) based on taxonomic information. When gap-filling a metabolic model for an organism M (such as Escherichia coli), the cost for reaction R is based on the frequency with which R occurs in other organisms within the phylum of M (in this case, Proteobacteria). The assumption behind this method is that different taxonomic groups are biased toward using different metabolic reactions. Evaluation of the new gap-filler on randomly degraded variants of the EcoCyc metabolic model for E.coli showed an increase in the average F1-score to 99.0 (when using the variable weights by frequency method at the phylum level), compared to 91.0 using the previous MetaFlux gap-filler and 80.3 using a basic gap-filler. Evaluation on two other microbial metabolic models showed similar improvements. AVAILABILITY AND IMPLEMENTATION: The Pathway Tools software (including MetaFlux) is free for academic use and is available at http://pathwaytools.com. Additional code for reproducing the results presented here is available at www.ai.sri.com/pkarp/pubs/taxgap/supplementary.zip. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wai Kit Ong, Peter E. Midford, Peter D. Karp
Bioinform.3
2019 The BioCyc collection of microbial genomes and metabolic pathways
abstract
BioCyc.org is a microbial genome Web portal that combines thousands of genomes with additional information inferred by computer programs, imported from other databases and curated from the biomedical literature by biologist curators. BioCyc also provides an extensive range of query tools, visualization services and analysis software. Recent advances in BioCyc include an expansion in the content of BioCyc in terms of both the number of genomes and the types of information available for each genome; an expansion in the amount of curated content within BioCyc; and new developments in the BioCyc software tools including redesigned gene/protein pages and metabolite pages; new search tools; a new sequence-alignment tool; a new tool for visualizing groups of related metabolic pathways; and a facility called SmartTables, which enables biologists to perform analyses that previously would have required a programmer's assistance.
Peter D. Karp, Richard Billington, Ron Caspi, Carol A. Fulcher, Mario Latendresse, Anamika Kothari, Ingrid M. Keseler, Markus Krummenacker, Peter E. Midford, Quang Ong, Wai Kit Ong, Suzanne M. Paley, Pallavi Subhraveti
Briefings Bioinform.1
2019 The MultiOmics Explainer: explaining omics results in the context of a pathway/genome database
abstract
BACKGROUND: High-throughput experiments can bring to light associations between genes, proteins and/or metabolites, many of which will be explainable by existing knowledge. Our aim is to speed elucidation of such explanations and, in some cases, find explanations that scientists might otherwise overlook. RESULTS: We describe the MultiOmics Explainer, a new tool within the Pathway Tools software suite that leverages what is known about an organism's metabolic and regulatory network to suggest explanations for the results of omics experiments. Querying a database such as EcoCyc, the MultiOmics Explainer searches the organism's network of metabolic reactions, transporters, cofactors, enzyme substrate-level activation and inhibition relationships, and transcriptional and translational regulation relationships to identify paths of influence among input genes, proteins and metabolites. Results are presented in a combined metabolic and regulatory diagram. We present several examples of explanations generated for associations found in the Escherichia coli literature. CONCLUSIONS: The MultiOmics Explainer is a valuable tool that helps researchers understand and interpret the results of their omics experiments in the context of what is known about an organism's metabolic and regulatory network. It showcases the rich set of computational inferences that can be drawn from a database such as EcoCyc that encodes a diverse range of biological interactions.
Suzanne M. Paley, Peter D. Karp
BMC Bioinform.2
2018 How the strengths of Lisp-family languages facilitate building complex and flexible bioinformatics applications
abstract
We present a rationale for expanding the presence of the Lisp family of programming languages in bioinformatics and computational biology research. Put simply, Lisp-family languages enable programmers to more quickly write programs that run faster than in other languages. Languages such as Common Lisp, Scheme and Clojure facilitate the creation of powerful and flexible software that is required for complex and rapidly evolving domains like biology. We will point out several important key features that distinguish languages of the Lisp family from other programming languages, and we will explain how these features can aid researchers in becoming more productive and creating better code. We will also show how these features make these languages ideal tools for artificial intelligence and machine learning applications. We will specifically stress the advantages of domain-specific languages (DSLs): languages that are specialized to a particular area, and thus not only facilitate easier research problem formulation, but also aid in the establishment of standards and best programming practices as applied to the specific research field at hand. DSLs are particularly easy to build in Common Lisp, the most comprehensive Lisp dialect, which is commonly referred to as the 'programmable programming language'. We are convinced that Lisp grants programmers unprecedented power to build increasingly sophisticated artificial intelligence systems that may ultimately transform machine learning and artificial intelligence research in bioinformatics and computational biology.
Bohdan B. Khomtchouk, Edmund Weitz, Peter D. Karp, Claes Wahlestedt
Briefings Bioinform.3
2018 Evaluation of reaction gap-filling accuracy by randomization
abstract
BACKGROUND: Completion of genome-scale flux-balance models using computational reaction gap-filling is a widely used approach, but its accuracy is not well known. RESULTS: We report on computational experiments of reaction gap filling in which we generated degraded versions of the EcoCyc-20.0-GEM model by randomly removing flux-carrying reactions from a growing model. We gap-filled the degraded models and compared the resulting gap-filled models with the original model. Gap-filling was performed by the Pathway Tools MetaFlux software using its General Development Mode (GenDev) and its Fast Development Mode (FastDev). We explored 12 GenDev variants including two linear solvers (SCIP and CPLEX) for solving the Mixed Integer Linear Programming (MILP) problems for gap filling; three different sets of linear constraints were applied; and two MILP methods were implemented. We compared these 13 variants according to accuracy, speed, and amount of information returned to the user. CONCLUSIONS: We observed large variation among the performance of the 13 gap-filling variants. Although no variant was best in all dimensions, we found one variant that was fast, accurate, and returned more information to the user. Some gap-filling variants were inaccurate, producing solutions that were non-minimum or invalid (did not enable model growth). The best GenDev variant showed a best average precision of 87% and a best average recall of 61%. FastDev showed an average precision of 71% and an average recall of 59%. Thus, using the most accurate variant, approximately 13% of the gap-filled reactions were incorrect (were not the reactions removed from the model), and 39% of gap-filled reactions were not found, suggesting that curation is still an important aspect of metabolic-model development.
Mario Latendresse, Peter D. Karp
BMC Bioinform.2
2017 How the strengths of Lisp-family languages facilitate building complex and flexible bioinformatics applications
abstract
Briefings in Bioinformatics (2016) doi: 10.1093/bib/bbw130 In the above article, various author names were incorrectly abbreviated in the ‘References’ section. The corrections have been made online and in print. The publisher apologizes for this error.
Bohdan B. Khomtchouk, Edmund Weitz, Peter D. Karp, Claes Wahlestedt
Briefings Bioinform.3
2016 Pathway Tools version 19.0 update: software for pathway/genome informatics and systems biology
abstract
Pathway Tools is a bioinformatics software environment with a broad set of capabilities. The software provides genome-informatics tools such as a genome browser, sequence alignments, a genome-variant analyzer and comparative-genomics operations. It offers metabolic-informatics tools, such as metabolic reconstruction, quantitative metabolic modeling, prediction of reaction atom mappings and metabolic route search. Pathway Tools also provides regulatory-informatics tools, such as the ability to represent and visualize a wide range of regulatory interactions. This article outlines the advances in Pathway Tools in the past 5 years. Major additions include components for metabolic modeling, metabolic route search, computation of atom mappings and estimation of compound Gibbs free energies of formation; addition of editors for signaling pathways, for genome sequences and for cellular architecture; storage of gene essentiality data and phenotype data; display of multiple alignments, and of signaling and electron-transport pathways; and development of Python and web-services application programming interfaces. Scientists around the world have created more than 9800 Pathway/Genome Databases by using Pathway Tools, many of which are curated databases for important model organisms.
Peter D. Karp, Mario Latendresse, Suzanne M. Paley, Markus Krummenacker, Quang Ong, Richard Billington, Anamika Kothari, Daniel S. Weaver, Thomas J. Lee, Pallavi Subhraveti, Aaron Spaulding, Carol A. Fulcher, Ingrid M. Keseler, Ron Caspi
Briefings Bioinform.1
2016 Representation and inference of cellular architecture for metabolic reconstruction and modeling
abstract
MOTIVATION: Metabolic modeling depends on accurately representing the cellular locations of enzyme-catalyzed and transport reactions. We sought to develop a representation of cellular compartmentation that would accurately capture cellular location information. We further sought a representation that would support automated inference of the cellular compartments present in newly sequenced organisms to speed model development, and that would enable representing the cellular compartments present in multiple cell types within a multicellular organism. RESULTS: We define the cellular architecture of a unicellular organism, or of a cell type from a multicellular organism, as the collection of cellular components it contains plus the topological relationships among those components. We developed a tool for inferring cellular architectures across many domains of life and extended our Cell Component Ontology to enable representation of the inferred architectures. We provide software for visualizing cellular architectures to verify their correctness and software for editing cellular architectures to modify or correct them. We also developed a representation that records the cellular compartment assignments of reactions with minimal duplication of information. AVAILABILITY AND IMPLEMENTATION: The Cell Component Ontology is freely available. The Pathway Tools software is freely available for academic research and is available for a fee for commercial use. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Suzanne M. Paley, Markus Krummenacker, Peter D. Karp
Bioinform.3
2016 Pathway collages: personalized multi-pathway diagrams
abstract
BACKGROUND: Metabolic pathway diagrams are a classical way of visualizing a linked cascade of biochemical reactions. However, to understand some biochemical situations, viewing a single pathway is insufficient, whereas viewing the entire metabolic network results in information overload. How do we enable scientists to rapidly construct personalized multi-pathway diagrams that depict a desired collection of interacting pathways that emphasize particular pathway interactions? RESULTS: We define software for constructing personalized multi-pathway diagrams called pathway-collages using a combination of manual and automatic layouts. The user specifies a set of pathways of interest for the collage from a Pathway/Genome Database. Layouts for the individual pathways are generated by the Pathway Tools software, and are sent to a Javascript Pathway Collage application implemented using Cytoscape.js. That application allows the user to re-position pathways; define connections between pathways; change visual style parameters; and paint metabolomics, gene expression, and reaction flux data onto the collage to obtain a desired multi-pathway diagram. We demonstrate the use of pathway collages in two application areas: a metabolomics study of pathogen drug response, and an Escherichia coli metabolic model. CONCLUSIONS: Pathway collages enable facile construction of personalized multi-pathway diagrams.
Suzanne M. Paley, Paul E. O'Maille, Daniel S. Weaver, Peter D. Karp
BMC Bioinform.4
2015 Message from the ISCB: ISCB Ebola award for important future research on the computational biology of Ebola virus
abstract
UNLABELLED: Speed is of the essence in combating Ebola; thus, computational approaches should form a significant component of Ebola research. As for the development of any modern drug, computational biology is uniquely positioned to contribute through comparative analysis of the genome sequences of Ebola strains and three-dimensional protein modeling. Other computational approaches to Ebola may include large-scale docking studies of Ebola proteins with human proteins and with small-molecule libraries, computational modeling of the spread of the virus, computational mining of the Ebola literature and creation of a curated Ebola database. Taken together, such computational efforts could significantly accelerate traditional scientific approaches. In recognition of the need for important and immediate solutions from the field of computational biology against Ebola, the International Society for Computational Biology (ISCB) announces a prize for an important computational advance in fighting the Ebola virus. ISCB will confer the ISCB Fight against Ebola Award, along with a prize of US$2000, at its July 2016 annual meeting (ISCB Intelligent Systems for Molecular Biology 2016, Orlando, FL). CONTACT: [email protected] or [email protected].
Peter D. Karp, Bonnie Berger, Diane E. Kovats, Thomas Lengauer, Michal Linial, Pardis Sabeti, Winston Hide, Burkhard Rost
Bioinform.1
2015 ISCB Ebola Award for Important Future Research on the Computational Biology of Ebola Virus
abstract
Speed is of the essence in combating Ebola; thus, computational approaches should form a significant component of Ebola research.As for the development of any modern drug, computational biology is uniquely positioned to contribute through comparative analysis of the genome sequences of Ebola strains as well as 3-D protein modeling.Other computational approaches to Ebola may include large-scale docking studies of Ebola proteins with human proteins and with small-molecule libraries, computational modeling of the spread of the virus, computational mining of the Ebola literature, and creation of a curated Ebola database.Taken together, such computational efforts could significantly accelerate traditional scientific approaches.In recognition of the need for important and immediate solutions from the field of computational biology against Ebola, the International Society for Computational Biology (ISCB) announces a prize for an important computational advance in fighting the Ebola virus.ISCB will confer the ISCB Fight against Ebola Award, along with a prize of US$2,000, at its July 2016 annual meeting (ISCB Intelligent Systems for Molecular Biology [ISMB] 2016, Orlando, Florida).
Peter D. Karp, Bonnie Berger, Diane E. Kovats, Thomas Lengauer, Michal Linial, Pardis Sabeti, Winston Hide, Burkhard Rost
PLoS Comput. Biol.1
2014 Optimal metabolic route search based on atom mappings
abstract
MOTIVATION: A key computational problem in metabolic engineering is finding efficient metabolic routes from a source to a target compound in genome-scale reaction networks, potentially considering the addition of new reactions. Efficiency can be based on many factors, such as route lengths, atoms conserved and the number of new reactions, and the new enzymes to catalyze them, added to the route. Fast algorithms are needed to systematically search these large genome-scale reaction networks. RESULTS: We present the algorithm used in the new RouteSearch tool within the Pathway Tools software. This algorithm is based on a general Branch-and-Bound search and involves constructing a network of atom mappings to facilitate efficient searching. As far as we know, it is the first published algorithm that finds guaranteed optimal routes where atom conservation is part of the optimality criteria. RouteSearch includes a graphical user interface that speeds user understanding of its search results. We evaluated the algorithm on five example metabolic-engineering problems from the literature; for one problem the published solution was equivalent to the optimal route found by RouteSearch; for the remaining four problems, RouteSearch found the published solution as one of its best-scored solutions. These problems were each solved in less than 5 s of computational time. AVAILABILITY AND IMPLEMENTATION: RouteSearch is accessible at BioCyc.org by using the menu command Metabolism --> Metabolic RouteSearch and by downloading Pathway Tools. Pathway Tools software is freely available to academic users, and for a fee to commercial users. Download from: http://biocyc.org/download.shtml.
Mario Latendresse, Markus Krummenacker, Peter D. Karp
Bioinform.3
2013 A systematic comparison of the MetaCyc and KEGG pathway databases
abstract
BACKGROUND: The MetaCyc and KEGG projects have developed large metabolic pathway databases that are used for a variety of applications including genome analysis and metabolic engineering. We present a comparison of the compound, reaction, and pathway content of MetaCyc version 16.0 and a KEGG version downloaded on Feb-27-2012 to increase understanding of their relative sizes, their degree of overlap, and their scope. To assess their overlap, we must know the correspondences between compounds, reactions, and pathways in MetaCyc, and those in KEGG. We devoted significant effort to computational and manual matching of these entities, and we evaluated the accuracy of the correspondences. RESULTS: KEGG contains 179 module pathways versus 1,846 base pathways in MetaCyc; KEGG contains 237 map pathways versus 296 super pathways in MetaCyc. KEGG pathways contain 3.3 times as many reactions on average as do MetaCyc pathways, and the databases employ different conceptualizations of metabolic pathways. KEGG contains 8,692 reactions versus 10,262 for MetaCyc. 6,174 KEGG reactions are components of KEGG pathways versus 6,348 for MetaCyc. KEGG contains 16,586 compounds versus 11,991 for MetaCyc. 6,912 KEGG compounds act as substrates in KEGG reactions versus 8,891 for MetaCyc. MetaCyc contains a broader set of database attributes than does KEGG, such as relationships from a compound to enzymes that it regulates, identification of spontaneous reactions, and the expected taxonomic range of metabolic pathways. MetaCyc contains many pathways not found in KEGG, from plants, fungi, metazoa, and actinobacteria; KEGG contains pathways not found in MetaCyc, for xenobiotic degradation, glycan metabolism, and metabolism of terpenoids and polyketides. MetaCyc contains fewer unbalanced reactions, which facilitates metabolic modeling such as using flux-balance analysis. MetaCyc includes generic reactions that may be instantiated computationally. CONCLUSIONS: KEGG contains significantly more compounds than does MetaCyc, whereas MetaCyc contains significantly more reactions and pathways than does KEGG, in particular KEGG modules are quite incomplete. The number of reactions occurring in pathways in the two DBs are quite similar.
Tomer Altman, Michael Travers, Anamika Kothari, Ron Caspi, Peter D. Karp
BMC Bioinform.5
2013 Computing minimal nutrient sets from metabolic networks via linear constraint solving
abstract
BACKGROUND: As more complete genome sequences become available, bioinformatics challenges arise in how to exploit genome sequences to make phenotypic predictions. One type of phenotypic prediction is to determine sets of compounds that will support the growth of a bacterium from the metabolic network inferred from the genome sequence of that organism. RESULTS: We present a method for computationally determining alternative growth media for an organism based on its metabolic network and transporter complement. Our method predicted 787 alternative anaerobic minimal nutrient sets for Escherichia coli K-12 MG1655 from the EcoCyc database. The program automatically partitioned the nutrients within these sets into 21 equivalence classes, most of which correspond to compounds serving as sources of carbon, nitrogen, phosphorous, and sulfur, or combinations of these essential elements. The nutrient sets were predicted with 72.5% accuracy as evaluated by comparison with 91 growth experiments. Novel aspects of our approach include (a) exhaustive consideration of all combinations of nutrients rather than assuming that all element sources can substitute for one another(an assumption that can be invalid in general) (b) leveraging the notion of a machinery-duplicating constraint, namely, that all intermediate metabolites used in active reactions must be produced in increasing concentrations to prevent successive dilution from cell division, (c) the use of Satisfiability Modulo Theory solvers rather than Linear Programming solvers, because our approach cannot be formulated as linear programming, (d) the use of Binary Decision Diagrams to produce an efficient implementation. CONCLUSIONS: Our method for generating minimal nutrient sets from the metabolic network and transporters of an organism combines linear constraint solving with binary decision diagrams to efficiently produce solution sets to provided growth problems.
Steven Eker, Markus Krummenacker, Alexander Glennon Shearer, Ashish Tiwari 0001, Ingrid M. Keseler, Carolyn L. Talcott, Peter D. Karp
BMC Bioinform.7
2012 Construction and completion of flux balance models from pathway databases
abstract
MOTIVATION: Flux balance analysis (FBA) is a well-known technique for genome-scale modeling of metabolic flux. Typically, an FBA formulation requires the accurate specification of four sets: biochemical reactions, biomass metabolites, nutrients and secreted metabolites. The development of FBA models can be time consuming and tedious because of the difficulty in assembling completely accurate descriptions of these sets, and in identifying errors in the composition of these sets. For example, the presence of a single non-producible metabolite in the biomass will make the entire model infeasible. Other difficulties in FBA modeling are that model distributions, and predicted fluxes, can be cryptic and difficult to understand. RESULTS: We present a multiple gap-filling method to accelerate the development of FBA models using a new tool, called MetaFlux, based on mixed integer linear programming (MILP). The method suggests corrections to the sets of reactions, biomass metabolites, nutrients and secretions. The method generates FBA models directly from Pathway/Genome Databases. Thus, FBA models developed in this framework are easily queried and visualized using the Pathway Tools software. Predicted fluxes are more easily comprehended by visualizing them on diagrams of individual metabolic pathways or of metabolic maps. MetaFlux can also remove redundant high-flux loops, solve FBA models once they are generated and model the effects of gene knockouts. MetaFlux has been validated through construction of FBA models for Escherichia coli and Homo sapiens. AVAILABILITY: Pathway Tools with MetaFlux is freely available to academic users, and for a fee to commercial users. Download from: biocyc.org/download.shtml. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mario Latendresse, Markus Krummenacker, Miles Trupp, Peter D. Karp
Bioinform.4
2012 Regulatory network operations in the Pathway Tools software
abstract
BACKGROUND: Biologists are elucidating complex collections of genetic regulatory data for multiple organisms. Software is needed for such regulatory network data. RESULTS: The Pathway Tools software supports storage and manipulation of regulatory information through a variety of strategies. The Pathway Tools regulation ontology captures transcriptional and translational regulation, substrate-level regulation of enzyme activity, post-translational modifications, and regulatory pathways. Regulatory visualizations include a novel diagram that summarizes all regulatory influences on a gene; a transcription-unit diagram, and an interactive visualization of a full transcriptional regulatory network that can be painted with gene expression data to probe correlations between gene expression and regulatory mechanisms. We introduce a novel type of enrichment analysis that asks whether a gene-expression dataset is over-represented for known regulators. We present algorithms for ranking the degree of regulatory influence of genes, and for computing the net positive and negative regulatory influences on a gene. CONCLUSIONS: Pathway Tools provides a comprehensive environment for manipulating molecular regulatory interactions that integrates regulatory data with an organism's genome and metabolic network. Curated collections of regulatory data authored using Pathway Tools are available for Escherichia coli, Bacillus subtilis, and Shewanella oneidensis.
Suzanne M. Paley, Mario Latendresse, Peter D. Karp
BMC Bioinform.3
2011 Discovering novel subsystems using comparative genomics
abstract
MOTIVATION: Key problems for computational genomics include discovering novel pathways in genome data, and discovering functional interaction partners for genes to define new members of partially elucidated pathways. RESULTS: We propose a novel method for the discovery of subsystems from annotated genomes. For each gene pair, a score measuring the likelihood that the two genes belong to a same subsystem is computed using genome context methods. Genes are then grouped based on these scores, and the resulting groups are filtered to keep only high-confidence groups. Since the method is based on genome context analysis, it relies solely on structural annotation of the genomes. The method can be used to discover new pathways, find missing genes from a known pathway, find new protein complexes or other kinds of functional groups and assign function to genes. We tested the accuracy of our method in Escherichia coli K-12. In one configuration of the system, we find that 31.6% of the candidate groups generated by our method match a known pathway or protein complex closely, and that we rediscover 31.2% of all known pathways and protein complexes of at least 4 genes. We believe that a significant proportion of the candidates that do not match any known group in E.coli K-12 corresponds to novel subsystems that may represent promising leads for future laboratory research. We discuss in-depth examples of these findings. AVAILABILITY: Predicted subsystems are available at http://brg.ai.sri.com/pwy-discovery/journal.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Luciana Ferrer, Alexander Glennon Shearer, Peter D. Karp
Bioinform.3
2011 Web-Based Metabolic Network Visualization with a Zooming User Interface
abstract
BACKGROUND: Displaying complex metabolic-map diagrams, for Web browsers, and allowing users to interact with them for querying and overlaying expression data over them is challenging. DESCRIPTION: We present a Web-based metabolic-map diagram, which can be interactively explored by the user, called the Cellular Overview. The main characteristic of this application is the zooming user interface enabling the user to focus on appropriate granularities of the network at will. Various searching commands are available to visually highlight sets of reactions, pathways, enzymes, metabolites, and so on. Expression data from single or multiple experiments can be overlaid on the diagram, which we call the Omics Viewer capability. The application provides Web services to highlight the diagram and to invoke the Omics Viewer. This application is entirely written in JavaScript for the client browsers and connect to a Pathway Tools Web server to retrieve data and diagrams. It uses the OpenLayers library to display tiled diagrams. CONCLUSIONS: This new online tool is capable of displaying large and complex metabolic-map diagrams in a very interactive manner. This application is available as part of the Pathway Tools software that powers multiple metabolic databases including Biocyc.org: The Cellular Overview is accessible under the Tools menu.
Mario Latendresse, Peter D. Karp
BMC Bioinform.2
2010 Development of Large Scientific Knowledge Bases
Peter D. Karp
ICAART (1)1
2010 Pathway Tools version 13.0: integrated software for pathway/genome informatics and systems biology
abstract
MOTIVATION: Biological systems function through dynamic interactions among genes and their products, regulatory circuits and metabolic networks. Our development of the Pathway Tools software was motivated by the need to construct biological knowledge resources that combine these many types of data, and that enable users to find and comprehend data of interest as quickly as possible through query and visualization tools. Further, we sought to support the development of metabolic flux models from pathway databases, and to use pathway information to leverage the interpretation of high-throughput data sets. RESULTS: In the past 4 years we have enhanced the already extensive Pathway Tools software in several respects. It can now support metabolic-model execution through the Web, it provides a more accurate gap filler for metabolic models; it supports development of models for organism communities distributed across a spatial grid; and model results may be visualized graphically. Pathway Tools supports several new omics-data analysis tools including the Omics Dashboard, multi-pathway diagrams called pathway collages, a pathway-covering algorithm for metabolomics data analysis and an algorithm for generating mechanistic explanations of multi-omics data. We have also improved the core pathway/genome databases management capabilities of the software, providing new multi-organism search tools for organism communities, improved graphics rendering, faster performance and re-designed gene and metabolite pages. AVAILABILITY: The software is free for academic use; a fee is required for commercial use. See http://pathwaytools.com. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online.
Peter D. Karp, Suzanne M. Paley, Markus Krummenacker, Mario Latendresse, Joseph M. Dale, Thomas J. Lee, Pallavi Kaipa, Fred Gilham, Aaron Spaulding, Liviu Popescu, Tomer Altman, Ian T. Paulsen, Ingrid M. Keseler, Ron Caspi
Briefings Bioinform.1
2010 Machine learning methods for metabolic pathway prediction
abstract
BACKGROUND: A key challenge in systems biology is the reconstruction of an organism's metabolic network from its genome sequence. One strategy for addressing this problem is to predict which metabolic pathways, from a reference database of known pathways, are present in the organism, based on the annotated genome of the organism. RESULTS: To quantitatively validate methods for pathway prediction, we developed a large "gold standard" dataset of 5,610 pathway instances known to be present or absent in curated metabolic pathway databases for six organisms. We defined a collection of 123 pathway features, whose information content we evaluated with respect to the gold standard. Feature data were used as input to an extensive collection of machine learning (ML) methods, including naïve Bayes, decision trees, and logistic regression, together with feature selection and ensemble methods. We compared the ML methods to the previous PathoLogic algorithm for pathway prediction using the gold standard dataset. We found that ML-based prediction methods can match the performance of the PathoLogic algorithm. PathoLogic achieved an accuracy of 91% and an F-measure of 0.786. The ML-based prediction methods achieved accuracy as high as 91.2% and F-measure as high as 0.787. The ML-based methods output a probability for each predicted pathway, whereas PathoLogic does not, which provides more information to the user and facilitates filtering of predicted pathways. CONCLUSIONS: ML methods for pathway prediction perform as well as existing methods, and have qualitative advantages in terms of extensibility, tunability, and explainability. More advanced prediction methods and/or more sophisticated input features may improve the performance of ML methods. However, pathway prediction performance appears to be limited largely by the ability to correctly match enzymes to the reactions they catalyze based on genome annotations.
Joseph M. Dale, Liviu Popescu, Peter D. Karp
BMC Bioinform.3
2010 A systematic study of genome context methods: calibration, normalization and combination
abstract
BACKGROUND: Genome context methods have been introduced in the last decade as automatic methods to predict functional relatedness between genes in a target genome using the patterns of existence and relative locations of the homologs of those genes in a set of reference genomes. Much work has been done in the application of these methods to different bioinformatics tasks, but few papers present a systematic study of the methods and their combination necessary for their optimal use. RESULTS: We present a thorough study of the four main families of genome context methods found in the literature: phylogenetic profile, gene fusion, gene cluster, and gene neighbor. We find that for most organisms the gene neighbor method outperforms the phylogenetic profile method by as much as 40% in sensitivity, being competitive with the gene cluster method at low sensitivities. Gene fusion is generally the worst performing of the four methods. A thorough exploration of the parameter space for each method is performed and results across different target organisms are presented. We propose the use of normalization procedures as those used on microarray data for the genome context scores. We show that substantial gains can be achieved from the use of a simple normalization technique. In particular, the sensitivity of the phylogenetic profile method is improved by around 25% after normalization, resulting, to our knowledge, on the best-performing phylogenetic profile system in the literature. Finally, we show results from combining the various genome context methods into a single score. When using a cross-validation procedure to train the combiners, with both original and normalized scores as input, a decision tree combiner results in gains of up to 20% with respect to the gene neighbor method. Overall, this represents a gain of around 15% over what can be considered the state of the art in this area: the four original genome context methods combined using a procedure like that used in the STRING database. Unfortunately, we find that these gains disappear when the combiner is trained only with organisms that are phylogenetically distant from the target organism. CONCLUSIONS: Our experiments indicate that gene neighbor is the best individual genome context method and that gains from the combination of individual methods are very sensitive to the training data used to obtain the combiner's parameters. If adequate training data is not available, using the gene neighbor score by itself instead of a combined score might be the best choice.
Luciana Ferrer, Joseph M. Dale, Peter D. Karp
BMC Bioinform.3
2008 Annotation-based inference of transporter function
abstract
MOTIVATION: We present a method for inferring and constructing transport reactions for transporter proteins based primarily on the analysis of the names of individual proteins in the genome annotation of an organism. Transport reactions are declarative descriptions of transporter activities, and thus can be manipulated computationally, unlike free-text protein names. Once transporter activities are encoded as transport reactions, a number of computational analyses are possible including database queries by transporter activity; inclusion of transporters into an automatically generated metabolic-map diagram that can be painted with omics data to aid in their interpretation; detection of anomalies in the metabolic and transport networks, such as substrates that are transported into the cell but are not inputs to any metabolic reaction or pathway; and comparative analyses of the transport capabilities of different organisms. RESULTS: On randomly selected organisms, the method achieves precision and recall rates of 0.93 and 0.90, respectively in identifying transporter proteins by name within the complete genome. The method obtains 67.5% accuracy in predicting complete transport reactions; if allowance is made for predictions that are overly general yet not incorrect, reaction prediction accuracy is 82.5%. AVAILABILITY: The method is implemented as part of PathoLogic, the inference component of the Pathway Tools software. Pathway Tools is freely available to researchers at non-commercial institutions, including source code; a fee applies to commercial institutions. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thomas J. Lee, Ian T. Paulsen, Peter D. Karp
ISMB3
2007 A survey of orphan enzyme activities
abstract
BACKGROUND: Using computational database searches, we have demonstrated previously that no gene sequences could be found for at least 36% of enzyme activities that have been assigned an Enzyme Commission number. Here we present a follow-up literature-based survey involving a statistically significant sample of such "orphan" activities. The survey was intended to determine whether sequences for these enzyme activities are truly unknown, or whether these sequences are absent from the public sequence databases but can be found in the literature. RESULTS: We demonstrate that for ~80% of sampled orphans, the absence of sequence data is bona fide. Our analyses further substantiate the notion that many of these enzyme activities play biologically important roles. CONCLUSION: This survey points toward significant scientific cost of having such a large fraction of characterized enzyme activities disconnected from sequence data. It also suggests that a larger effort, beginning with a comprehensive survey of all putative orphan activities, would resolve nearly 300 artifactual orphans and reconnect a wealth of enzyme research with modern genomics. For these reasons, we propose that a systematic effort to identify the cognate genes of orphan enzymes be undertaken.
Yannick Pouliot, Peter D. Karp
BMC Bioinform.2
2006 BioWarehouse: a bioinformatics database warehouse toolkit
abstract
BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases. RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research. CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.
Thomas J. Lee, Yannick Pouliot, Valerie Wagner, David W. J. Stringer-Calvert, Jessica D. Tenenbaum, Peter D. Karp
BMC Bioinform.7
2006 The comprehensive updated regulatory network of Escherichia coli K-12
abstract
BACKGROUND: Escherichia coli is the model organism for which our knowledge of its regulatory network is the most extensive. Over the last few years, our project has been collecting and curating the literature concerning E. coli transcription initiation and operons, providing in both the RegulonDB and EcoCyc databases the largest electronically encoded network available. A paper published recently by Ma et al. (2004) showed several differences in the versions of the network present in these two databases. Discrepancies have been corrected, annotations from this and other groups (Shen-Orr et al., 2002) have been added, making the RegulonDB and EcoCyc databases the largest comprehensive and constantly curated regulatory network of E. coli K-12. RESULTS: Several groups have been using these curated data as part of their bioinformatics and systems biology projects, in combination with external data obtained from other sources, thus enlarging the dataset initially obtained from either RegulonDB or EcoCyc of the E. coli K12 regulatory network. We kindly obtained from the groups of Uri Alon and Hong-Wu Ma the interactions they have added to enrich their public versions of the E. coli regulatory network. These were used to search for original references and curate them with the same standards we use regularly, adding in several cases the original references (instead of reviews or missing references), as well as adding the corresponding experimental evidence codes. We also corrected all discrepancies in the two databases available as explained below. CONCLUSION: One hundred and fifty new interactions have been added to our databases as a result of this specific curation effort, in addition to those added as a result of our continuous curation work. RegulonDB gene names are now based on those of EcoCyc to avoid confusion due to gene names and synonyms, and the public releases of RegulonDB and EcoCyc are henceforth synchronized to avoid confusion due to different versions. Public flat files are available providing direct access to the regulatory network interactions thus avoiding errors due to differences in database modelling and representation. The regulatory network available in RegulonDB and EcoCyc is the most comprehensive and regularly updated electronically-encoded regulatory network of E. coli K-12.
Heladia Salgado, Alberto Santos-Zavaleta, Socorro Gama-Castro, Martín Peralta-Gil, Mónica I. Peñaloza-Spínola, Agustino Martínez-Antonio, Peter D. Karp, Julio Collado-Vides
BMC Bioinform.7
2005 Querying and computing with BioCyc databases
abstract
Summary: We describe multiple methods for accessing and querying the complex and integrated cellular data in the BioCyc family of databases: access through multiple file formats, access through Application Program Interfaces (APIs) for LISP, Perl and Java, and SQL access through the BioWarehouse relational database. Availability: The Pathway Tools software and 20 BioCyc DBs in Tiers 1 and 2 are freely available to academic users; fees apply to some types of commercial use. For download instructions see http://BioCyc.org/download.shtml Supplementary information: For more details on programmatic access to BioCyc DBs, see http://bioinformatics.ai.sri.com/ptools/ptools-resources.html Contact: [email protected]
Markus Krummenacker, Suzanne M. Paley, Lukas A. Mueller, Thomas Yan, Peter D. Karp
Bioinform.5
2004 Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases
abstract
MOTIVATION: The prediction of transcription units (TUs, which are similar to operons) is an important problem that has been tackled using many different approaches. The availability of complete microbial genomes has made genome-wide TU predictions possible. Pathway-genome databases (PGDBs) add metabolic and other organizational (i.e. protein complexes) information to the annotated genome, and are able to capture TU organization information. These characteristics of PGDBs make them a suitable framework for the development and implementation of TU predictors. RESULTS: We implemented a TU predictor that uses only intergenic distance and functional classification of genes to predict TU boundaries, and applied it to EcoCyc, our PGDB of Escherichia coli. To this original predictor, we added information on metabolic pathways, protein complexes and transporters, all readily available in EcoCyc, in order to generate an enhanced predictor. The enhanced predictor correctly predicted 80% of the known E.coli TUs (69% of the known operons), a moderate improvement over the original predictor's performance (75% of TUs and 65% of operons correctly predicted), demonstrating that the extra information available in the PGDB does indeed improve prediction performance. Performance of this E.coli-based predictor on a genome other than that of E.coli was tested on BsubCyc, our computationally generated PGDB for Bacillus subtilis, for which a set of 100 known operons is available. Prediction accuracy decreased substantially (46% of the known operons correctly predicted). This was due in part to missing information in BsubCyc, which prevented full use of the predictor's features. The augmented predictor has been implemented as part of our Pathway Tools software suite, and can be used to populate a PGDB with predicted TUs. AVAILABILITY: The TU predictor is included in version 7.0 of the Pathway Tools software suite. Pathway Tools 7.0 is available free of charge to academic institutions and for a fee to commercial enterprises. It runs on Sun Solaris 8, Linux and Windows. TUs predicted on the Caulobacter crescentus and Mycobacterium tuberculosis (H37Rv) genomes are available in our CauloCyc and MtbrvCyc databases, available at the BioCyc web site (http://biocyc.org). To obtain version 7.0 of Pathway Tools, follow the directions in our web site, http://biocyc.org/download.shtml.
Pedro Romero, Peter D. Karp
Bioinform.2
2004 A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases
abstract
BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.
Michelle L. Green, Peter D. Karp
BMC Bioinform.2
2003 Knowledge acquisition, consistency checking and concurrency control for Gene Ontology (GO)
abstract
MOTIVATION: A critical element of the computational infrastructure required for functional genomics is a shared language for communicating biological data and knowledge. The Gene Ontology (GO; http://www.geneontology.org) provides a taxonomy of concepts and their attributes for annotating gene products. As GO increases in size, its ongoing construction and maintenance becomes more challenging. In this paper, we assess the applicability of a Knowledge Base Management System (KBMS), Protégé-2000, to the maintenance and development of GO. RESULTS: We transferred GO to Protégé-2000 in order to evaluate its suitability for GO. The graphical user interface supported browsing and editing of GO. Tools for consistency checking identified minor inconsistencies in GO and opportunities to reduce redundancy in its representation. The Protégé Axiom Language proved useful for checking ontological consistency. The PROMPT tool allowed us to track changes to GO. Using Protégé-2000, we tested our ability to make changes and extensions to GO to refine the semantics of attributes and classify more concepts. AVAILABILITY: Gene Ontology in Protégé-2000 and the associated code are located at http://smi.stanford.edu/projects/helix/gokbms/. Protégé-2000 is available from http://protege.stanford.edu.
Iwei Yeh, Peter D. Karp, Natasha F. Noy, Russ B. Altman
Bioinform.2
2002 The Pathway Tools software
abstract
Abstract Motivation: Bioinformatics requires reusable software tools for creating model-organism databases (MODs). Results: The Pathway Tools is a reusable, production-quality software environment for creating a type of MOD called a Pathway/Genome Database (PGDB). A PGDB such as EcoCyc (see http://ecocyc.org) integrates our evolving understanding of the genes, proteins, metabolic network, and genetic network of an organism. This paper provides an overview of the four main components of the Pathway Tools: The PathoLogic component supports creation of new PGDBs from the annotated genome of an organism. The Pathway/Genome Navigator provides query, visualization, and Web-publishing services for PGDBs. The Pathway/Genome Editors support interactive updating of PGDBs. The Pathway Tools ontology defines the schema of PGDBs. The Pathway Tools makes use of the Ocelot object database system for data management services for PGDBs. The Pathway Tools has been used to build PGDBs for 13 organisms within SRI and by external users. Availability: The software is freely available to academics and is available for a fee to commercial institutions. Contact [email protected] for information on obtaining the software. Contact: [email protected] Keywords: Bioinformatics; model organism database; genome analyses; metabolic pathways. *To whom correspondence should be addressed.
Peter D. Karp, Suzanne M. Paley, Pedro Romero
ISMB1
2002 Evaluation of computational metabolic-pathway predictions for Helicobacter pylori
abstract
MOTIVATION: We seek to determine the accuracy of computational methods for predicting metabolic pathways in sequenced genomes, and to understand the contributions of both the prediction algorithms, and the reference pathway databases used by those algorithms, to the prediction accuracy. RESULTS: The comparisons we performed were as follows. (1) We compared two predictions of the pathway complements of Helicobacter pylori that were computed by an early version of our pathway-prediction algorithm: prediction A used the EcoCyc E. coli pathway DB as the reference database (DB) for prediction, and prediction B used the MetaCyc pathway DB (a superset of EcoCyc) as the reference pathway DB. The MetaCyc-based prediction contained 75% more pathway predictions, but we believe a significant number of those predictions were false positives. (2) We compared two predictions of the pathway complement of H. pylori that used MetaCyc as the reference pathway DB, but that used different algorithms: the original PathoLogic algorithm, and an enhanced version of the algorithm designed to eliminate false-positive pathway predictions. The improved algorithm predicted 30\% fewer metabolic pathways than the original algorithm; all of the eliminated pathways are believed to be false-positive predictions. (3) We compared the 98 pathways predicted by the enhanced algorithm with the results of a manual analysis of the pathways of H. pylori. Results: 40 of the computationally predicted pathways were consistent with the manual analysis, 13 pathways are considered false-positive predictions, and four pathways had partially overlapping topologies. Twenty-six predicted pathways were not mentioned in the manual analysis; we believe these are correct predictions by PathoLogic that were not found by the manual analysis. Five pathways from the manual analysis were not found computationally. Agreement between the computational and manual predictions was good overall, with the computational analysis inferring many pathways that the manual analysis did not identify. Ultimately the manual analysis is also partially speculative, and therefore is not an absolute measure of correctness. The algorithm is designed to err on the side of more false positives to bring more potential pathways to the user's attention. The resulting H. pylori pathway DB is freely available at http://ecocyc.org:1555/HPY/organism-summary?object=HPY. AVAILABILITY: The Pathway Tools software is freely available to academic users, and is available to commercial users for a fee. Contact [email protected] for information on obtaining the software.
Suzanne M. Paley, Peter D. Karp
Bioinform.2
2001 Database verification studies of SWISS-PROT and GenBank
abstract
PROBLEM STATEMENT: We have studied the relationships among SWISS-PROT, TrEMBL, and GenBank with two goals. First is to determine whether users can reliably identify those proteins in SWISS-PROT whose functions were determined experimentally, as opposed to proteins whose functions were predicted computationally. If this information was present in reasonable quantities, it would allow researchers to decrease the propagation of incorrect function predictions during sequence annotation, and to assemble training sets for developing the next generation of sequence-analysis algorithms. Second is to assess the consistency between translated GenBank sequences and sequences in SWISS-PROT and TrEMBL. RESULTS: (1) Contrary to claims by the SWISS-PROT authors, we conclude that SWISS-PROT does not identify a significant number of experimentally characterized proteins. (2) SWISS-PROT is more incomplete than we expected in that version 38.0 from July 1999 lacks many proteins from the full genomes of important organisms that were sequenced years earlier. (3) Even if we combine SWISS-PROT and TrEMBL, some sequences from the full genomes are missing from the combined dataset. (4) In many cases, translated GenBank genes do not exactly match the corresponding SWISS-PROT sequences, for reasons that include missing or removed methionines, differing translation start positions, individual amino-acid differences, and inclusion of sequence data from multiple sequencing projects. For example, results show that for Escherichia coli, 80.6% of the proteins in the GenBank entry for the complete genome have identical sequence matches with SWISS-PROT/TrEMBL sequences, 13.4% have exact substring matches, and matches for 4.1% can be found using BLAST search; the remaining 2.0% of E.coli protein sequences (most of which are ORFs) have no clear matches to SWISS-PROT/TrEMBL. Although many of these differences can be explained by the complexity of the DB, and by the curation processes used to create it, the scale of the differences is notable.
Peter D. Karp, Suzanne M. Paley, Jingchun Zhu
Bioinform.1
2000 An Evaluation of Ontology Exchange Languages for Bioinformatics
Robin McEntire, Peter D. Karp, Neil F. Abernethy, David Benton, Gregg Helt, Matt DeJongh, Robert Kent, Anthony Kosky, Suzanna Lewis, Dan Hodnett, Eric P. Neumann, Frank Olken, Dhiraj K. Pathak, Peter Tarczy-Hornoch, Luca Toldo, Thodoros Topaloglou
ISMB2
2000 An ontology for biological function based on molecular interactions
abstract
MOTIVATIONS: A number of important bioinformatics computations involve computing with function: executing computational operations whose inputs or outputs are descriptions of the functions of biomolecules. Examples include performing functional queries to sequence and pathway databases, and determining functional equality to evaluate algorithms that predict function from sequence. A prerequisite to computing with function is the existence of an ontology that provides a structured semantic encoding of function. Functional bioinformatics is an emerging subfield of bioinformatics that is concerned with developing ontologies and algorithms for computing with biological function. RESULTS: The article explores the notion of computing with function, and explains the importance of ontologies of function to bioinformatics. The functional ontology developed for the EcoCyc database is presented. This ontology can encode a diverse array of biochemical processes, including enzymatic reactions involving small-molecule substrates and macromolecular substrates, signal-transduction processes, transport events, and mechanisms of regulation of gene expression. The ontology is validated through its use to express complex functional queries for the EcoCyc DB. CONTACT: [email protected]
Peter D. Karp
Bioinform.1
1999 A Collaborative Environment for Authoring Large Knowledge Bases
Peter D. Karp, Vinay K. Chaudhri, Suzanne M. Paley
J. Intell. Inf. Syst.1
1998 What we do not know about sequence analysis and sequence databases
abstract
P D Karp; What we do not know about sequence analysis and sequence databases., Bioinformatics, Volume 14, Issue 9, 1 January 1998, Pages 753–754, https://doi.or
Peter D. Karp
Bioinform.1
1997 Prediction of Enzyme Classification from Protein Sequence without the Use of Sequence Similarity
Marie desJardins, Peter D. Karp, Markus Krummenacker, Thomas J. Lee, Christos A. Ouzounis
ISMB2
1997 Estimation of equilibrium constants using automated group contribution methods
abstract
MOTIVATION: Group contribution methods are frequently used for estimating physical properties of compounds from their molecular structures. An algorithm for estimating Gibbs energies of formation through group contribution methods has been automated in an object-oriented framework. The algorithm decomposes compound structures according to a basis set of groups. It permits the use of wildcards and is able to distinguish between ring groups and chain groups that use similar search structures. Past methods relied on manual decomposition of compounds into constituent groups. RESULTS: The software is written in Common LISP and requires < 2 min to estimate Gibbs energies of formation for a database of 780 species of varying size and complexity. The software allows rapid expansion to incorporate different basis sets and to estimate a variety of other physical properties.
Ronald G. Forsythe Jr., Peter D. Karp, Michael L. Mavrovouniotis
Comput. Appl. Biosci.2
1996 HinCyc: A Knowledge Base of the Complete Genome and Metabolic Pathways of H. influenzae
Peter D. Karp, Christos A. Ouzounis, Suzanne M. Paley
ISMB1
1995 The Generic Frame Protocol
Peter D. Karp, Karen L. Myers, Thomas R. Gruber
IJCAI (1)1
1995 Knowledge Representation in the Large
Peter D. Karp, Suzanne M. Paley
IJCAI (1)1
1994 A Storage System for Scalable Knowledge Representation
abstract
Twenty years of AI research in knowledge representation has produced frame knowledge representation systems (FRSs) that incorporate a number of important advances. However, FRSs lack two important capabilities that prevent them from scaling up to realistic applications: they cannot provide high-speed access to large knowledge bases (KBs), and they do not support shared, concurrent KB access by multiple users. Our research investigates the hypothesis that one can employ an existing database management system (DBMS) as a storage subsystem for an FRS, to provide high-speed access to large, shared KBs. We describe the design and implementation of a general storage system that incrementally loads referenced frames from a DBMS, and saves modified frames back to the DBMS, for two different FRSs: LOOM and THEO. We also present experimental results showing that the performance of our prototype storage subsystem exceeds that of flat files for simulated applications that reference or update up to one third of the frames from a large LOOM KB.
Peter D. Karp, Suzanne M. Paley, Ira B. Greenberg
CIKM1
1994 Representations of Metabolic Knowledge: Pathways
Peter D. Karp, Suzanne M. Paley
ISMB1
1993 Representations of Metabolic Knowledge
Peter D. Karp, Monica Riley
ISMB1
1993 Detection and elimination of inference channels in multilevel relational database systems
abstract
Multilevel relational database systems store information at different security classifications. An inference problem exists if it is possible for a user with a low-level clearance to draw conclusions about information at higher classifications. The authors are developing DISSECT, a tool for analyzing multilevel relational database schemas to assist in the detection and elimination of inference problems. A translation is defined from schemas to an equivalent graph representation, which can be presented graphically in DISSECT. The initial focus is on detection of inference problems that depend only on information all of which is stored in the database. In particular, potential inference problems are identified as different sequences of foreign key relationships that connect the same entities. Inferences can be blocked by upgrading the security classification of some of foreign key relationships. A global optimization approach to upgrading is suggested to block a set of inference problems that allows upgrade costs to be considered, and supports security categories as well as levels.>
Xiaolei Qian, Mark E. Stickel, Peter D. Karp, Teresa F. Lunt, Thomas D. Garvey
S&P3
1993 Design Methods for Scientific Hypothesis Formation and Their Application to Molecular Biology
Peter D. Karp
Mach. Learn.1
1992 A knowledge base of the chemical compounds of intermediary metabolism
abstract
This paper describes a publicly available knowledge base of the chemical compounds involved in intermediary metabolism. We consider the motivations for constructing a knowledge base of metabolic compounds, the methodology by which it was constructed, and the information that it currently contains. Currently the knowledge base describes 981 compounds, listing for each: synonyms for its name, a systematic name, CAS registry number, chemical formula, molecular weight, chemical structure and two-dimensional display coordinates for the structure. The Compound Knowledge Base (CompoundKB) illustrates several methodological principles that should guide the development of biological knowledge bases. I argue that biological datasets should be made available in multiple representations to increase their accessibility to end users, and I present multiple representations of the CompoundKB (knowledge base, relational data base and ASN. 1 representations). I also analyze the general characteristics of these representations to provide an understanding of their relative advantages and disadvantages. Another principle is that the error rate of biological data bases should be estimated and documented-this analysis is performed for the CompoundKB.
Peter D. Karp
Comput. Appl. Biosci.1
1991 Artificial intelligence methods for theory representation and hypothesis formation
abstract
This article describes artificial intelligence methods for representing theories in molecular biology, and for improving the predictive power of these theories using experimental data. A program called GENSIM provides a framework for representing theories that includes descriptions of classes of biological objects (genes, enzymes, etc.), and processes that specify potential interactions among these objects (such as enzymatic reactions). GENSIM can employ a theory specified within this framework to predict the outcomes of biological experiments. A program called HYPGENE comes into play when the observed outcome of an experiment does not match the outcome predicted by GENSIM. HYPGENE works backward from the error in GENSIMs prediction to postulate changes to both the theory embodied by GENSIM, and the presumed initial conditions of the experiment. I view HYPGENEs hypothesis generation task as a design problem, and I have adapted AI methods developed for design and planning to this task. These techniques were developed in conjunction with an in-depth study of the discovery of the gene regulation mechanism of attenuation in the E. coli tryptophan operon. Both GENSIM and HYPGENE have been tested on sample problems from the history of attenuation, and produced many of the same solutions as biologists did.
Peter D. Karp
Comput. Appl. Biosci.1