Edward C. Uberbacher

dblp:26/2676 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17Artificial intelligence and machine learning · 3Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
15 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genome annotation
gene prediction
0.272012
Gene and translation initiation site prediction in metagenomic sequences · Bioinform. 2012
Reference-based gene model prediction on DNA contigs (extended abstract) · RECOMB 1997
GRAIL: a multi-agent neural network system for gene identification · Proc. IEEE 1996
Bioinformatics and computational biology
knowledge base
0.112012
BESC knowledgebase public portal · Bioinform. 2012
Bioinformatics and computational biology
metagenomics
0.112012
Gene and translation initiation site prediction in metagenomic sequences · Bioinform. 2012
Bioinformatics and computational biology › genome annotation
translation initiation site prediction
0.112012
Gene and translation initiation site prediction in metagenomic sequences · Bioinform. 2012
Bioinformatics and computational biology
protein structure prediction
0.142000
Sequence-structure specificity of a knowledge based energy function at the secondary structure level · Bioinform. 2000
A new method for modeling and solving the protein fold recognition problem (extended abstract) · RECOMB 1998
A polynomial-time algorithm for a class of protein threading problems · Comput. Appl. Biosci. 1996
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition
0.132000
Sequence-structure specificity of a knowledge based energy function at the secondary structure level · Bioinform. 2000
A new method for modeling and solving the protein fold recognition problem (extended abstract) · RECOMB 1998
A polynomial-time algorithm for a class of protein threading problems · Comput. Appl. Biosci. 1996
Bioinformatics and computational biology
sequence analysis
0.012012
Gene and translation initiation site prediction in metagenomic sequences · Bioinform. 2012
Bioinformatics and computational biology › sequence analysis
motif discovery
0.012003
Background rareness-based iterative multiple sequence alignment algorithm for regulatory element detection · Bioinform. 2003
Bioinformatics and computational biology › gene regulation
regulatory genomics
0.012003
Background rareness-based iterative multiple sequence alignment algorithm for regulatory element detection · Bioinform. 2003
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.012003
Background rareness-based iterative multiple sequence alignment algorithm for regulatory element detection · Bioinform. 2003
Bioinformatics and computational biology › genome annotation › gene prediction
exon prediction
0.021997
Reference-based gene model prediction on DNA contigs (extended abstract) · RECOMB 1997
An Improved System for Exon Recognition and Gene Modeling in Human DNA Sequence · ISMB 1994
Bioinformatics and computational biology › genome annotation
gene structure prediction
0.021997
Inferring Gene Structures in Genomic Sequences Using Pattern Recognition and Expressed Sequence Tags · ISMB 1997
Constructing gene models from accurately predicted exons: an application of dynamic programming · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology › protein structure analysis
knowledge-based potential
0.012000
Sequence-structure specificity of a knowledge based energy function at the secondary structure level · Bioinform. 2000
Bioinformatics and computational biology › sequence alignment › DNA-protein alignment
frameshift-aware alignment
0.011996
Alignments of DNA and protein sequences containing frameshift errors · Comput. Appl. Biosci. 1996
Bioinformatics and computational biology › genome annotation › gene prediction
homology-based gene prediction
0.011996
Gene Prediction by Pattern Recognition and Homology Search · ISMB 1996
Bioinformatics and computational biology › protein sequence analysis
protein homology detection
0.011996
Alignments of DNA and protein sequences containing frameshift errors · Comput. Appl. Biosci. 1996
Bioinformatics and computational biology
sequence alignment
0.011996
Alignments of DNA and protein sequences containing frameshift errors · Comput. Appl. Biosci. 1996
Bioinformatics and computational biology › structural bioinformatics
sequence-structure alignment
0.011996
A polynomial-time algorithm for a class of protein threading problems · Comput. Appl. Biosci. 1996
Algorithms and data structures
polynomial-time algorithms
0.011996
A polynomial-time algorithm for a class of protein threading problems · Comput. Appl. Biosci. 1996
Bioinformatics and computational biology › genome annotation
coding region prediction
0.011995
Correcting sequencing errors in DNA coding regions using a dynamic programming approach · Comput. Appl. Biosci. 1995
Bioinformatics and computational biology › protein structure prediction › protein folding
protein fold prediction
0.011995
Predicting Protein Folding Classes without Overly Relying on Homology · ISMB 1995
Bioinformatics and computational biology › sequence analysis
sequencing error correction
0.011995
Correcting sequencing errors in DNA coding regions using a dynamic programming approach · Comput. Appl. Biosci. 1995
Bioinformatics and computational biology › transcriptomics
expressed sequence tag analysis
0.011997
Inferring Gene Structures in Genomic Sequences Using Pattern Recognition and Expressed Sequence Tags · ISMB 1997

Methods — techniques the papers use, named apart from their topics

gene prediction algorithm · 0.1data integration · 0.1dynamic programming · 0.1pattern recognition · 0.1maximum a posteriori scoring · 0.0markov chain background model · 0.0gibbs sampling · 0.0mutation analysis · 0.0energy function evaluation · 0.0approximation scheme · 0.0smith-waterman alignment · 0.0
YearPublicationVenuePosition
2016 MicroRNAs Form Triplexes with Double Stranded DNA at Sequence-Specific Binding Sites; a Eukaryotic Mechanism via which microRNAs Could Directly Alter Gene Expression
abstract
MicroRNAs are important regulators of gene expression, acting primarily by binding to sequence-specific locations on already transcribed messenger RNAs (mRNA) and typically down-regulating their stability or translation. Recent studies indicate that microRNAs may also play a role in up-regulating mRNA transcription levels, although a definitive mechanism has not been established. Double-helical DNA is capable of forming triple-helical structures through Hoogsteen and reverse Hoogsteen interactions in the major groove of the duplex, and we show physical evidence (i.e., NMR, FRET, SPR) that purine or pyrimidine-rich microRNAs of appropriate length and sequence form triple-helical structures with purine-rich sequences of duplex DNA, and identify microRNA sequences that favor triplex formation. We developed an algorithm (Trident) to search genome-wide for potential triplex-forming sites and show that several mammalian and non-mammalian genomes are enriched for strong microRNA triplex binding sites. We show that those genes containing sequences favoring microRNA triplex formation are markedly enriched (3.3 fold, p<2.2 × 10(-16)) for genes whose expression is positively correlated with expression of microRNAs targeting triplex binding sequences. This work has thus revealed a new mechanism by which microRNAs could interact with gene promoter regions to modify gene transcription.
Steven W. Paugh, David R. Coss, Ju Bao, Lucas T. Laudermilk, Christy R. Grace, Antonio M. Ferreira, M. Brett Waddell, Granger Ridout, Deanna Naeve, Michael R. Leuze, Philip F. LoCascio, John C. Panetta, Mark R. Wilkinson, Ching-Hon Pui, Clayton W. Naeve, Edward C. Uberbacher, Erik J. Bonten, William E. Evans
PLoS Comput. Biol.16
2012 Gene and translation initiation site prediction in metagenomic sequences
abstract
MOTIVATION: Gene prediction in metagenomic sequences remains a difficult problem. Current sequencing technologies do not achieve sufficient coverage to assemble the individual genomes in a typical sample; consequently, sequencing runs produce a large number of short sequences whose exact origin is unknown. Since these sequences are usually smaller than the average length of a gene, algorithms must make predictions based on very little data. RESULTS: We present MetaProdigal, a metagenomic version of the gene prediction program Prodigal, that can identify genes in short, anonymous coding sequences with a high degree of accuracy. The novel value of the method consists of enhanced translation initiation site identification, ability to identify sequences that use alternate genetic codes and confidence values for each gene call. We compare the results of MetaProdigal with other methods and conclude with a discussion of future improvements. AVAILABILITY: The Prodigal software is freely available under the General Public License from http://code.google.com/p/prodigal/.
Doug Hyatt, Philip F. LoCascio, Loren J. Hauser, Edward C. Uberbacher
Bioinform.4
2012 BESC knowledgebase public portal
abstract
UNLABELLED: The BioEnergy Science Center (BESC) is undertaking large experimental campaigns to understand the biosynthesis and biodegradation of biomass and to develop biofuel solutions. BESC is generating large volumes of diverse data, including genome sequences, omics data and assay results. The purpose of the BESC Knowledgebase is to serve as a centralized repository for experimentally generated data and to provide an integrated, interactive and user-friendly analysis framework. The Portal makes available tools for visualization, integration and analysis of data either produced by BESC or obtained from external resources. AVAILABILITY: http://besckb.ornl.gov.
Mustafa H. Syed, Tatiana V. Karpinets, Morey Parang, Michael R. Leuze, Doug Hyatt, Steven D. Brown, Steve Moulton, Michael D. Galloway, Edward C. Uberbacher
Bioinform.10
2003 Background rareness-based iterative multiple sequence alignment algorithm for regulatory element detection
abstract
MOTIVATION: Experimental methods capable of generating sets of co-regulated genes have become commonplace, however, recognizing the regulatory motifs responsible for this regulation remains difficult. As a result, computational detection of transcription factor binding sites in such data sets has been an active area of research. Most approaches have utilized either Gibbs sampling or greedy strategies to identify such elements in sets of sequences. These existing methods have varying degrees of success depending on the strength and length of the signals and the number of available sequences. We present a new deterministic iterative algorithm for regulatory element detection based on a Markov chain background. As in other methods, sequences in the entire genome and the training set are taken into account in order to discriminate against commonly occurring signals and produce patterns, which are significant in the training set. RESULTS: The results of the algorithm compare favorably with existing tools on previously known and newly compiled data sets. The iteration based search appears rather rigorous, not only finding the binding sites, but also showing how the binding site stands out from genomic background. The approach used to score the results is critical and a discussion of various scoring schemes and options is also presented. Benchmarking of several methods shows that while most tools are good at detecting strong signals, Gibbs sampling algorithms give inconsistent results when the regulatory element signal becomes weak. A Markov chain based background model alleviates the drawbacks of MAP (maximum a posteriori log likelihood) scores. AVAILABILITY: Available on request from the authors. SUPPLEMENTARY INFORMATION: Data and the results presented in this paper are available on the web at http://compbio.ornl.gov/mira/index.html
Chandrasegaran Narasimhan, Philip F. LoCascio, Edward C. Uberbacher
Bioinform.3
2000 Sequence-structure specificity of a knowledge based energy function at the secondary structure level
abstract
MOTIVATION: This paper investigates the sequence-structure specificity of a representative knowledge based energy function by applying it to threading at the level of secondary structures of proteins. Assessing the strengths and weaknesses of an energy function at this fundamental level provides more detailed and insightful information than at the tertiary structure level and the results obtained can be useful in tertiary level threading. RESULTS: We threaded each of the 293 non-redundant proteins onto the secondary structures contained in its respective native protein (host template). We also used 68 pairs of proteins with similar folds and low sequence identity. For each pair, we threaded the sequence of one protein onto the secondary structures of the other protein. The discerning power of the total energy function and its one-body, pairwise, and mutation components is studied. We then applied our energy function to a recent study which demonstrated how a designed 11-amino acid sequence can replace distinct segments (one segment is an alpha-helix, the other is a beta-sheet) of a protein without changing its fold. We conducted random mutations of the designed sequence to determine the patterns for favorable mutations. We also studied the sequence-structure specificity at the boundaries of a secondary structure. Finally, we demonstrated how to speed up tertiary level threading by filtering out alignments found to be energetically unfavorable during the secondary structure threading. AVAILABILITY: The program is available on request from the authors. CONTACT: [email protected]
Dong Xu 0002, Michael A. Unseren, Ying Xu 0001, Edward C. Uberbacher
Bioinform.4
1998 A new method for modeling and solving the protein fold recognition problem (extended abstract)
abstract
Computational recognition of native-like folds from a protein fold database is considered to be a promising alternative approach to the ab initio fold prediction.We present a new and egective method forprotein fold recognition through optimally aligning (threading) an amino acid sequence and a protein fold (template).A protein fold, in our database, is represented as a se953 of core secondary structures, and the alignment quality is cletermined by three factors.They are (I) the fitness between each amino acid and the environment of its assigned (aligned) template position; (2) pairwise interaction preferences between amino acids that are spatially close,-and (3) alignment gap penalties.Our threading algorithm cons&ucts an optimum alignment between an amino acid sequence of size n and a protein fold template of size m in O((m-l-n1+o-5cMlog(n))nc+1) time and O(nm+nc+2) space, where M is the number of core secondary structures in the fold, and C is a (small) nonnegative integer, determined by a mathematical property of the pairwise interactions in the fold.C is less than or equal to 4 for about 75% of the 296 unique folds in our .database,when pairwise interactions are restricted to amino acids 5 7A apart (measured between their beta carbon atoms).An approximation scheme is developed for fold templates with C > 4, when threading requires too much memo y and time to be practical on a typical workstation.Permission to make digitalibrud copies of all or pat ofthis material for personal orclassroomuseisgranteduithoutfeeprovidedthatthecopies are not made or diiuted for profit or commercial advantage, the wpy-riShtnotice,thetitleofthepublicationanditsdateappear, andnoticek giventhatcopylightkbype -on ofthe AChL Inc To copy othenvise, to republish, to post on servers or to rediiiuteto lists, requires specific pamission ardor fee.
Ying Xu 0001, Dong Xu 0002, Edward C. Uberbacher
RECOMB3
1998 A segmentation algorithm for noisy images: Design and evaluation
Ying Xu 0001, Victor Olman, Edward C. Uberbacher
Pattern Recognit. Lett.3
1997 Inferring Gene Structures in Genomic Sequences Using Pattern Recognition and Expressed Sequence Tags
Ying Xu 0001, Richard J. Mural, Edward C. Uberbacher
ISMB3
1997 Reference-based gene model prediction on DNA contigs (extended abstract)
abstract
This paper presents an algorithm for constructing multiple gene models on a set of contigs of a large genomic clone.The algorithm first uses pattern recognition-based methods to locute ex0n.s or partial exons in each contig, and then applies protein homology or EST information from the databases, as rejerence models, to parse the predicted exons into genr models.In the phase of gene model construction, the algorithm uses a unified framework for genes ranging from situation with homologous proteins/EST3 to no homologous pvotein/EST in the database.By exploiting protein homol-0.~1~or EST information, the algorithm is able to (1) parse exons into multiple gene models over a set of DNA contigs (possibly unoriented and unordered); (2) remove falsely predicted exons; and (3) identify and locate exons missed by the initial exon prediction.
Ying Xu 0001, Edward C. Uberbacher
RECOMB2
1997 2D image segmentation using minimum spanning trees
Ying Xu 0001, Edward C. Uberbacher
Image Vis. Comput.2
1996 Gene Prediction by Pattern Recognition and Homology Search
Ying Xu 0001, Edward C. Uberbacher
ISMB2
1996 Alignments of DNA and protein sequences containing frameshift errors
abstract
Molecular sequences, like all experimental data, are subject to error. Many current DNA sequencing protocols have very significant error rates and often generate artefactual insertions and deletions of bases (indels) which corrupt the translation of sequences and compromise the detection of protein homologies. The impact of these errors on the utility of molecular sequence data is dependent on the analytic technique used to interpret the data. In the presence of frameshift errors, standard algorithms using six-frame translation can miss important homologies because only subfragments of the correct translation are available in any given frame. We present a new algorithm which can detect and correct frameshift errors in DNA sequences during comparison of translated sequences with protein sequences in the databases. This algorithm can recognize homologous proteins sharing 30% identity even in the presence of a 7% frameshift error rate. Our algorithm uses dynamic programming, producing a guaranteed optimal alignment in the presence of frameshifts, and has a sensitivity equivalent to Smith-Waterman. The computational efficiency of the algorithm is O(nm) where n and m are the sizes of two sequences being compared. The algorithm does not rely on prior knowledge or heuristic rules and performs significantly better than any previously reported method.
Xiaojun Guan, Edward C. Uberbacher
Comput. Appl. Biosci.2
1996 A polynomial-time algorithm for a class of protein threading problems
abstract
This paper presents an algorithm for constructing an optimal alignment between a three-dimensional protein structure template and an amino acid sequence. A protein structure template is given as a sequence of amino acid residue positions in three-dimensional space, along with an array of physical properties attached to each position; these residue positions are sequentially grouped into a series of core secondary structures (central helices and beta sheets). In addition to match scores and gap penalties, as in a traditional sequence-sequence alignment problem, the quality of a structure-sequence alignment is also determined by interaction preferences among amino acids aligned with structure positions that are spatially close (we call these 'long-range interactions'). Although it is known that constructing such a structure-sequence alignment in the most general form is NP-hard, our algorithm runs in polynomial time when restricted to structures with a 'modest' number of long-range amino acid interactions. In the current work, long-range interactions are limited to interactions between amino acids from different core secondary structures. Dividing the series of core secondary structures into two subseries creates a cut set of long-range interactions. If we use N, M and C to represent the size of an amino acid sequence, the size of a structure template, and the maximum cut size of long-range interactions, respectively, the algorithm finds an optimal structure-sequence alignment in O(21C NM) time, a polynomial function of N and M when C = O(log(N + M)). When running on structure-sequence alignment problems without long-range intersections, i.e. C = 0, the algorithm achieves the same asymptotic computational complexity of the Smith-Waterman sequence-sequence alignment algorithm.
Ying Xu 0001, Edward C. Uberbacher
Comput. Appl. Biosci.2
1996 GRAIL: a multi-agent neural network system for gene identification
abstract
Identifying genes within large regions of uncharacterized DNA is a difficult undertaking and is currently the focus of many research efforts. We describe a gene localization and modeling system, called GRAIL. GRAIL is a multiple sensor-neural network-based system. It localizes genes in anonymous DNA sequence by recognizing features related to protein-coding regions and the boundaries of coding regions, and then combines the recognized features using a neural network system. Localized coding regions are then "optimally" parsed into a gene model. Through years of extensive testing GRAIL consistently achieves about 90% of coding portions of test genes with a false positive rate of about 10% A number of genes for major genetic diseases have been located through the use of GRAIL, and over 1000 research laboratories worldwide use GRAIL on regular bases for localization of genes on their newly sequenced DNA.
Ying Xu 0001, Richard J. Mural, J. Ralph Einstein, Manesh B. Shah, Edward C. Uberbacher
Proc. IEEE5
1995 Use of Neural Networks for Prediction of Graft Failure following Liver Transplantation
abstract
Liver transplantation k a well-established therapeutic option forpatients with end-stage liver dkease.However, up to 20% of transplanted livers fail to have adequate function initially, and at least harf of those will eventually fail.Accurate, early prediction of outcome may ameliorate thk situation by encouraging retransplantation before the patient's condition becomes irreversible.In thk shrdy, clinical information was gathered prospectively for 295 patients who underwent liver transplantation at the Universiy of Pittsburgh Medical Center, and was divided into sets.The feed-forward, fully connected, neural networks had 7 or 8 inputs, a single hidden layer conskting of 3 nodes and a single output node cfailure=I, success=O).The networks were trained with data from a randomly selected subset of 240 patients while the remaining 55 patients made up the test set.The preoperative (day 0) data conskted of patient demographics plus the results of standard liver function tests.The "day 1" data consisted information gathered during surgery plus the prediction of outcome from day 0. Data for days 2-5 included resultsfrom standard liver function tests plus the prediction of outcome from the previous day's network The network was trained using a standard back propagation algorithm.naining was assessed by testing the abiliy of the network to correctly predict the outcome of the 55 patients in the test set.The accuracy of prediction by the neural network improved each day and so by day 5, 98% of the grafi survivors in the test set were correctly predicted while 88% of grafi failures in the test set were correctly predicted.
Sherri Matis, Howard Doyle, Ignazio Marino, Richard J. Mural, Edward C. Uberbacher
CBMS5
1995 Predicting Protein Folding Classes without Overly Relying on Homology
Mark Craven, Richard J. Mural, Loren J. Hauser, Edward C. Uberbacher
ISMB4
1995 Correcting sequencing errors in DNA coding regions using a dynamic programming approach
abstract
This paper presents an algorithm for detecting and 'correcting' sequencing errors that occur in DNA coding regions. The types of sequencing errors addressed are insertions and deletions (indels) of DNA bases. The goal is to provide a capability which makes single-pass or low-redundancy sequence data more informative, reducing the need for high-redundancy sequencing for gene identification and characterization purposes. This would permit improved sequencing efficiency and reduce genome sequencing costs. The algorithm detects sequencing errors by discovering changes in the statistically preferred reading frame within a putative coding region and then inserts a number of 'neutral' bases at a perceived reading frame transition point to make the putative exon candidate frame consistent. We have implemented the algorithm as a front-end subsystem of the GRAIL DNA sequence analysis system to construct a version which is very error tolerant and also intend to use this as a testbed for further development of sequencing error-correction technology. Preliminary test results have shown the usefulness of this algorithm and also exhibited some of its weakness, providing possible directions for further improvement. On a test set consisting of 68 human DNA sequences with 1% randomly generated indels in coding regions, the algorithm detected and corrected 76% of the indels. The average distance between the position of an indel and the predicted one was 9.4 bases. With this subsystem in place, GRAIL correctly predicted 89% of the coding messages with 10% false message on the 'corrected' sequences, compared to 69% correctly predicted coding messages and 11% falsely predicted messages on the 'corrupted' sequences using standard GRAIL II method (version 1.2).(ABSTRACT TRUNCATED AT 250 WORDS)
Ying Xu 0001, Richard J. Mural, Edward C. Uberbacher
Comput. Appl. Biosci.3
1994 An Improved System for Exon Recognition and Gene Modeling in Human DNA Sequence
J. Ralph Einstein, Richard J. Mural, Manesh J. Shah, Edward C. Uberbacher
ISMB5
1994 Constructing gene models from accurately predicted exons: an application of dynamic programming
abstract
This paper presents a computationally efficient algorithm, the Gene Assembly Program III (GAP III), for constructing gene models from a set of accurately-predicted 'exons'. The input to the algorithm is a set of clusters of exon candidates, generated by a new version of the GRAIL coding region recognition system. The exon candidates of a cluster differ in their presumed edges and occasionally in their reading frames. Each exon candidate has a numerical score representing its 'probability' of being an actual exon. GAP III uses a dynamic programming algorithm to construct a gene model, complete or partial, by optimizing a predefined objective function. The optimal gene models constructed by GAP III correspond very well with the structures of genes which have been determined experimentally and reported in the Genome Sequence Database (GSDB). On a test set of 137 human and mouse DNA sequences consisting of 954 true exons, GAP III constructed 137 gene models using 892 exons, among which 859 (859/954 = 90%) are true exons and 33 (33/892 = 3%) are false positive. Among the 859 true positives, 635 (74%) match the actual exons exactly, and 838 (98%) have at least one edge correct. GAP III is computationally efficient. If we use E and C to represent the total number of exon candidates in all clusters and the number of clusters, respectively, the running time of GAP III is proportional to (E x C).
Ying Xu 0001, Richard J. Mural, Edward C. Uberbacher
Comput. Appl. Biosci.3