Ross A. Overbeek

dblp:o/RossAOverbeek · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-authorTheory of computation · 11 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 7Systems, architecture and hardware · 4Software engineering, systems software and programming languages · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
3 papers
Automated reasoning and model checking · 52% Algorithms and data structures · 43% Logic in computer science · 5%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › metagenomics
metagenome annotation
0.112012
Real Time Metagenomics: Using k-mers to annotate metagenomes · Bioinform. 2012
Bioinformatics and computational biology
metagenomics
0.112012
Real Time Metagenomics: Using k-mers to annotate metagenomes · Bioinform. 2012
Bioinformatics and computational biology › metagenomics
taxonomic classification
0.012012
Real Time Metagenomics: Using k-mers to annotate metagenomes · Bioinform. 2012
Bioinformatics and computational biology › phylogenetics › phylogenetic inference
maximum likelihood estimation
0.011994
fastDNAmL: a tool for construction of phylogenetic trees of DNA sequences using maximum likelihood · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology
multiple sequence alignment
0.011994
The genetic data environment an expandable GUI for multiple sequence analysis · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology
phylogenetics
0.011994
fastDNAmL: a tool for construction of phylogenetic trees of DNA sequences using maximum likelihood · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology › phylogenetics
phylogenetic inference
0.011994
fastDNAmL: a tool for construction of phylogenetic trees of DNA sequences using maximum likelihood · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology
sequence analysis
0.011994
The genetic data environment an expandable GUI for multiple sequence analysis · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA structure prediction
0.011990
Structure detection through automated covariance search · Comput. Appl. Biosci. 1990
User interface design and tools
graphical user interface
0.011994
The genetic data environment an expandable GUI for multiple sequence analysis · Comput. Appl. Biosci. 1994
Algorithms and data structures › computational biology
sequence analysis
0.011990
Structure detection through automated covariance search · Comput. Appl. Biosci. 1990
Data models and query languages
entity-relationship model
0.011980
A Practical Design Methodology for the Implementation of IMS Databases, Using the Entity-Relationship Model · SIGMOD Conference 1980
Automated reasoning and model checking
automated theorem proving
0.021976
Problems and Experiments for and with Automated Theorem-Proving Programs · IEEE Trans. Computers 1976
A New Class of Automated Theorem-Proving Algorithms · J. ACM 1974
Automated reasoning and model checking
satisfiability
0.011974
A New Class of Automated Theorem-Proving Algorithms · J. ACM 1974
Data models and query languages › data modeling
hierarchical data model
0.011980
A Practical Design Methodology for the Implementation of IMS Databases, Using the Entity-Relationship Model · SIGMOD Conference 1980

Methods — techniques the papers use, named apart from their topics

last-common ancestor · 0.1k-mer matching · 0.1x-windows · 0.0external program integration · 0.0covariation analysis · 0.0maximum likelihood · 0.0schema translation · 0.0statement sequence generation · 0.0completeness proof · 0.0
YearPublicationVenuePosition
2019 PATRIC as a unique resource for studying antimicrobial resistance
abstract
The Pathosystems Resource Integration Center (PATRIC, www.patricbrc.org) is designed to provide researchers with the tools and services that they need to perform genomic and other 'omic' data analyses. In response to mounting concern over antimicrobial resistance (AMR), the PATRIC team has been developing new tools that help researchers understand AMR and its genetic determinants. To support comparative analyses, we have added AMR phenotype data to over 15 000 genomes in the PATRIC database, often assembling genomes from reads in public archives and collecting their associated AMR panel data from the literature to augment the collection. We have also been using this collection of AMR metadata to build machine learning-based classifiers that can predict the AMR phenotypes and the genomic regions associated with resistance for genomes being submitted to the annotation service. Likewise, we have undertaken a large AMR protein annotation effort by manually curating data from the literature and public repositories. This collection of 7370 AMR reference proteins, which contains many protein annotations (functional roles) that are unique to PATRIC and RAST, has been manually curated so that it projects stably across genomes. The collection currently projects to 1 610 744 proteins in the PATRIC database. Finally, the PATRIC Web site has been expanded to enable AMR-based custom page views so that researchers can easily explore AMR data and design experiments based on whole genomes or individual genes.
Dionysios A. Antonopoulos, Rida Assaf, Ramy K. Aziz, Thomas S. Brettin, Christopher Bun, Neal Conrad, James J. Davis 0002, Emily M. Dietrich, Terry Disz, Svetlana Gerdes, Ron Kenyon, Dustin Machi, Chunhong Mao, Daniel E. Murphy-Olson, Eric K. Nordberg, Gary J. Olsen, Robert Olson, Ross A. Overbeek, Bruce D. Parrello, Gordon D. Pusch, John Santerre, Maulik Shukla, Rick L. Stevens, Margo VanOeffelen, Veronika Vonstein, Andrew S. Warren, Alice R. Wattam, Fangfang Xia, Hyun Seung Yoo
Briefings Bioinform.18
2019 A machine learning-based service for estimating quality of genomes using PATRIC
abstract
BACKGROUND: Recent advances in high-volume sequencing technology and mining of genomes from metagenomic samples call for rapid and reliable genome quality evaluation. The current release of the PATRIC database contains over 220,000 genomes, and current metagenomic technology supports assemblies of many draft-quality genomes from a single sample, most of which will be novel. DESCRIPTION: We have added two quality assessment tools to the PATRIC annotation pipeline. EvalCon uses supervised machine learning to calculate an annotation consistency score. EvalG implements a variant of the CheckM algorithm to estimate contamination and completeness of an annotated genome.We report on the performance of these tools and the potential utility of the consistency score. Additionally, we provide contamination, completeness, and consistency measures for all genomes in PATRIC and in a recent set of metagenomic assemblies. CONCLUSION: EvalG and EvalCon facilitate the rapid quality control and exploration of PATRIC-annotated draft genomes.
Bruce D. Parrello, Rory Butler, Philippe Chlenski, Robert Olson, Jamie C. Overbeek, Gordon D. Pusch, Veronika Vonstein, Ross A. Overbeek
BMC Bioinform.8
2014 Genome-scale bacterial transcriptional regulatory networks: reconstruction and integrated analysis with metabolic models
abstract
Advances in sequencing technology are resulting in the rapid emergence of large numbers of complete genome sequences. High-throughput annotation and metabolic modeling of these genomes is now a reality. The high-throughput reconstruction and analysis of genome-scale transcriptional regulatory networks represent the next frontier in microbial bioinformatics. The fruition of this next frontier will depend on the integration of numerous data sources relating to mechanisms, components and behavior of the transcriptional regulatory machinery, as well as the integration of the regulatory machinery into genome-scale cellular models. Here, we review existing repositories for different types of transcriptional regulatory data, including expression data, transcription factor data and binding site locations and we explore how these data are being used for the reconstruction of new regulatory networks. From template network-based methods to de novo reverse engineering from expression data, we discuss how regulatory networks can be reconstructed and integrated with metabolic models to improve model predictions and performance. We also explore the impact these integrated models can have in simulating phenotypes, optimizing the production of compounds of interest or paving the way to a whole-cell model.
José P. Faria, Ross A. Overbeek, Fangfang Xia, Miguel Rocha 0001, Isabel Rocha, Christopher S. Henry
Briefings Bioinform.2
2012 Real Time Metagenomics: Using k-mers to annotate metagenomes
abstract
Abstract Summary: Annotation of metagenomes involves comparing the individual sequence reads with a database of known sequences and assigning a unique function to each read. This is a time-consuming task that is computationally intensive (though not computationally complex). Here we present a novel approach to annotate metagenomes using unique k-mer oligopeptide sequences from 7 to 12 amino acids long. We demonstrate that k-mer-based annotations are faster and approach the sensitivity and precision of blastx-based annotations without loosing accuracy. A last-common ancestor approach was also developed to describe the members of the community. Availability and implementation: This open-source application was implemented in Perl and can be accessed via a user-friendly website at http://edwards.sdsu.edu/rtmg. In addition, code to access the annotation servers is available for download from http://www.theseed.org/. FIGfams and k-mers are available for download from ftp://ftp.theseed.org/FIGfams/. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Robert A. Edwards, Robert Olson, Terry Disz, Gordon D. Pusch, Veronika Vonstein, Rick L. Stevens, Ross A. Overbeek
Bioinform.7
2010 Accessing the SEED genome databases via Web services API: tools for programmers
abstract
BACKGROUND: The SEED integrates many publicly available genome sequences into a single resource. The database contains accurate and up-to-date annotations based on the subsystems concept that leverages clustering between genomes and other clues to accurately and efficiently annotate microbial genomes. The backend is used as the foundation for many genome annotation tools, such as the Rapid Annotation using Subsystems Technology (RAST) server for whole genome annotation, the metagenomics RAST server for random community genome annotations, and the annotation clearinghouse for exchanging annotations from different resources. In addition to a web user interface, the SEED also provides Web services based API for programmatic access to the data in the SEED, allowing the development of third-party tools and mash-ups. RESULTS: The currently exposed Web services encompass over forty different methods for accessing data related to microbial genome annotations. The Web services provide comprehensive access to the database back end, allowing any programmer access to the most consistent and accurate genome annotations available. The Web services are deployed using a platform independent service-oriented approach that allows the user to choose the most suitable programming platform for their application. Example code demonstrate that Web services can be used to access the SEED using common bioinformatics programming languages such as Perl, Python, and Java. CONCLUSIONS: We present a novel approach to access the SEED database. Using Web services, a robust API for access to genomics data is provided, without requiring large volume downloads all at once. The API ensures timely access to the most current datasets available, including the new genomes as soon as they come online.
Terry Disz, Sajia Akhter, Daniel Cuevas, Robert Olson, Ross A. Overbeek, Veronika Vonstein, Rick L. Stevens, Robert A. Edwards
BMC Bioinform.5
1998 Metabolic Pathway Interface to Molecular Biology Databases
abstract
We present results of providing database support to biomedicine via federation of SDB Cooperation/Integration based upon the KEGG GUI for molecular biology. The federation provides a common link to three molecular biology databases. The added value of the federation is freedom from consulting multiple references to ascertain the full set of enzymatic reactions in a metabolic pathway, and the option of selecting multiple queries to submit to the federated SDBs. Each of the SDBs is extensive, but incomplete. The union of the SDBs, implemented transparently by the federation, is more complete. Each SDB provides a different approach to the options available for data presentation and a different set of Web server tools for data analysis. Thus, an important part of the added value of the federation is the cross-fertilization available in the union of the molecular biological content, the presentation of data, and the tools available for analysis.
Barry Zeeberg, Kevin Watanabe, Susumu Goto, Ross A. Overbeek, Larry Kerschberg, George Michaels
SSDBM4
1994 Fast phylogenetic analysis on a massively parallel machine
abstract
We developed a parallel processing system for analyzing phylogenetic relationships of microorganisms based on a maximum likelihood method. Methods for inferring relationships from molecular sequence data are especially valuable, given the enormous increases in DNA sequence data. The maximum likelihood method uses concrete models of the evolutionary process and are well-motivated statistically, but the computational cost has hindered the use of this method for inferring trees with more than about 20 organisms. We parallelized the maximum likelihood method by utilizing two types of parallelism, parallel evaluation of phylogenetic trees and parallel computation of likelihood values. By combining these two parallelisms, we obtained significant speedup on the Intel Touchstone DELTA.
Hideo Matsuda, Gary J. Olsen, Ross A. Overbeek, Yukio Kaneda
International Conference on Supercomputing3
1994 fastDNAmL: a tool for construction of phylogenetic trees of DNA sequences using maximum likelihood
abstract
We have developed a new tool, called fastDNAml, for constructing phylogenetic trees from DNA sequences. The program can be run on a wide variety of computers ranging from Unix workstations to massively parallel systems, and is available from the Ribosomal Database Project (RDP) by anonymous FTP. Our program uses a maximum likelihood approach and is based on version 3.3 of Felsenstein's dnaml program. Several enhancements, including algorithmic changes, significantly improve performance and reduce memory usage, making it feasible to construct even very large trees. Trees containing 40-100 taxa have been easily generated, and phylogenetic estimates are possible even when hundreds of sequences exist. We are currently using the tool to construct a phylogenetic tree based on 473 small subunit rRNA sequences from prokaryotes.
Gary J. Olsen, Hideo Matsuda, Ray Hagstrom, Ross A. Overbeek
Comput. Appl. Biosci.4
1994 The genetic data environment an expandable GUI for multiple sequence analysis
abstract
An X-Windows-based graphic user interface is presented which allows the seamless integration of numerous existing biomolecular programs into a single analysis environment. This environment is based on a core multiple sequence editor that is linked to external programs by a user-expandable menu system and is supported on Sun and DEC workstations. There is no limitation to the number of external functions that can be linked to the interface. The length and number of sequences that can be handled are limited only by the size of virtual memory present on the workstation. The sequence data itself is used as the reference point from which analysis is done, and scalable graphic views are supported. It is suggested that future software development utilizing this expandable, user-defined menu system and the I/O linkage of external programs will allow biologists to easily integrate expertise from disparate fields into a single environment.
Steven W. Smith, Ross A. Overbeek, Carl R. Woese, W. Gilbert, P. M. Gillevet
Comput. Appl. Biosci.2
1994 Formula Databases for High-Performance Resolution/Paramodulation Systems
Ralph M. Butler, Ross A. Overbeek
J. Autom. Reason.2
1994 Maximum likelihood genetic sequence reconstruction from oligo content
abstract
Abstract One promising technique for determining long genetic sequences is sequencing by oligonucleotide content. This technique involves probing a segment of the unknown multimillion “character” genetic sequence for the presence or absence of known short subsequences. The information obtained from such hybridization experiments may be represented in network form. Network optimization methods may then be applied to identify the most likely forms of the unknown target sequence. © 1994 by John Wiley & Sons, Inc.
Jane N. Hagstrom, Ray Hagstrom, Ross A. Overbeek, Morgan Price, Linus Schrage
Networks3
1993 The CADE-11 Competitions: A Personal View
Ross A. Overbeek
J. Autom. Reason.1
1992 Exploitation of Parallel Processing for Implementing High-Performance Deduction Systems
Anita Jindal, Ross A. Overbeek, Waldo C. Kabat
J. Autom. Reason.2
1990 A High-Performance Parallel Theorem Prover
Ralph M. Butler, Ian T. Foster, Anita Jindal, Ross A. Overbeek
CADE4
1990 Automated Reasoning Contributed to Mathematics and Logic
Larry Wos, Steven K. Winker, William McCune, Ross A. Overbeek, Ewing L. Lusk, Rick L. Stevens, Ralph M. Butler
CADE4
1990 Structure detection through automated covariance search
abstract
This paper summarizes our investigations into the computational detection of secondary and tertiary structure of ribosomal RNA. We have developed a new automated procedure that not only identifies potential secondary and tertiary structural interactions, but also provides the covariation evidence that supports the proposed bondings, and any counterevidence that can be detected in the known sequences. A small number of previously unknown higher-order structural features have been detected in individual RNA molecules (16S rRNA and 7S RNA) through the use of our automated procedure. We are systematically studying mitochondrial rRNA, seeking tertiary structure within 16S rRNA and quaternary structure between 16S and 23S rRNA. To test hypotheses suggested by an examination of our program's output, our colleagues in biology are sequencing key portions of the 23S ribosomal RNA for species in which the known 16S ribosomal RNA exhibits variation (from the dominant pattern) at the site of a proposed bonding. Our ultimate hope is that automated covariation analysis will contribute significantly to a refined picture of ribosomal structure.
Steven K. Winker, Ross A. Overbeek, Carl R. Woese, Gary J. Olsen, N. Pfluger
Comput. Appl. Biosci.2
1988 Geometric specification of scheduling constraints: A simplified approach to multiprocessing
Barney Glickfeld, Ross A. Overbeek
Parallel Comput.2
1987 Experiments with OR-Parallel Logic Programs
Terry Disz, Ewing L. Lusk, Ross A. Overbeek
ICLP3
1986 Paths to High-Performance Automated Theorem Proving
Ralph M. Butler, Ewing L. Lusk, William McCune, Ross A. Overbeek
CADE4
1986 ITP at Argonne National Laboratory
Ewing L. Lusk, William McCune, Ross A. Overbeek
CADE3
1986 Parallel Logic Programming for Numeric Applications
Ralph M. Butler, Ewing L. Lusk, William McCune, Ross A. Overbeek
ICLP4
1986 Set Theory in First-Order Logic: Clauses for Gödel's Axioms
Robert S. Boyer, Ewing L. Lusk, William McCune, Ross A. Overbeek, Mark E. Stickel, Larry Wos
J. Autom. Reason.4
1986 A Foray Into Combinatory Logic
Barney Glickfeld, Ross A. Overbeek
J. Autom. Reason.2
1985 The Design of Entity-Relationship Models for General Ledger Systems
Bruce D. Parrello, Ross A. Overbeek, Ewing L. Lusk
Data Knowl. Eng.2
1985 Non-Horn Problems
Ewing L. Lusk, Ross A. Overbeek
J. Autom. Reason.2
1985 Reasoning about Equality
Ewing L. Lusk, Ross A. Overbeek
J. Autom. Reason.2
1985 A technique for achieving portability among multiprocessors: Implementation on the Lemur
J. A. Clausing, Ray Hagstrom, Ewing L. Lusk, Ross A. Overbeek
Parallel Comput.4
1984 A Portable Environment for Research in Automated Reasoning
Ewing L. Lusk, Ross A. Overbeek
CADE2
1983 Data Management: A Practical View (Panel)
Paul K. Blackwell, Dan Kapp, Ross A. Overbeek, H. J. Spencer, Gio Wiederhold, Stanley B. Zdonik
ER3
1983 Tools for the Creation of IMS Database Designs from Entity-Relationship Diagrams
G. Margrave, Ewing L. Lusk, Ross A. Overbeek
ER3
1982 Logic Machine Architecture: Kernel Funtions
Ewing L. Lusk, William McCune, Ross A. Overbeek
CADE3
1982 Logic Machine Architecture: Inference Mechanisms
Ewing L. Lusk, William McCune, Ross A. Overbeek
CADE3
1981 Item Tracking Entity-Relationship Models
Ewing L. Lusk, Gene Petrie, Ross A. Overbeek
ER3
1980 Data Structures and Control Architectures for Implementation of Theorem-Proving Programs
Ross A. Overbeek, Ewing L. Lusk
CADE1
1980 Hyperparamodulation: A Refinement of Paramodulation
Larry Wos, Ross A. Overbeek, Lawrence J. Henschen
CADE2
1980 A Practical Design Methodology for the Implementation of IMS Databases, Using the Entity-Relationship Model
abstract
Article Free Access Share on A practical design methodology for the implementation of IMS databases, using the entity-relationship model Authors: Ewing L. Lusk Northern Illinois University, DeKalb, Illinois Northern Illinois University, DeKalb, IllinoisView Profile , Ross A. Overbeek Northern Illinois University, DeKalb, Illinois Northern Illinois University, DeKalb, IllinoisView Profile , Bruce Parrello Northern Illinois University, DeKalb, Illinois Northern Illinois University, DeKalb, IllinoisView Profile Authors Info & Claims SIGMOD '80: Proceedings of the 1980 ACM SIGMOD international conference on Management of dataMay 1980Pages 9–21https://doi.org/10.1145/582250.582253Published:14 May 1980Publication History 2citation582DownloadsMetricsTotal Citations2Total Downloads582Last 12 Months50Last 6 weeks10 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Publisher SiteeReaderPDF
Ewing L. Lusk, Ross A. Overbeek, Bruce D. Parrello
SIGMOD Conference2
1979 A DML for Entity-Relationship Models
Ewing L. Lusk, Ross A. Overbeek
ER2
1976 Problems and Experiments for and with Automated Theorem-Proving Programs
abstract
The two objectives of this paper are 1) to give a large and varied problem set, complete clause sets for use in testing automated theorem-proving programs and 2) the presentation of a number of experiments with an existing program under a variety of conditions.
John D. McCharen, Ross A. Overbeek, Larry Wos
IEEE Trans. Computers2
1974 A New Class of Automated Theorem-Proving Algorithms
abstract
A procedure is defined for deriving from any statementSan infinite sequence of statementsS0,S1,S2,S3, ··· such that: (a) if there exists anisuch thatSiis unsatisfiable, thenSis unsatisfiable; (b) ifSis unsatisfiable, then there exists anisuch thatSiis unsatisfiable; (c) for allithe Herbrand universe ofSiis finite; hence, for eachithe satisfiability ofSiis decidable. The new algorithms are then based on the idea of generating successiveSiin the sequence and testing eachSifor satisfiability. Each element in the class of new algorithms is complete.
Ross A. Overbeek
J. ACM1