Yuri I. Wolf

dblp:65/3472 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0002-0247-8708ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
2 papers
Graph algorithms and graph theory · 94% Information theory · 6%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
comparative genomics
0.112010
A low-polynomial algorithm for assembling clusters of orthologous groups from intergenomic symmetric best matches · Bioinform. 2010
Bioinformatics and computational biology › comparative genomics › orthology analysis
ortholog identification
0.112010
A low-polynomial algorithm for assembling clusters of orthologous groups from intergenomic symmetric best matches · Bioinform. 2010
Graph algorithms and graph theory
graph algorithms
0.112010
A low-polynomial algorithm for assembling clusters of orthologous groups from intergenomic symmetric best matches · Bioinform. 2010
Bioinformatics and computational biology
birth-and-death process model
0.012003
Simple stochastic birth andz death models of genome evolution: was there enough time for us to evolve? · Bioinform. 2003
Bioinformatics and computational biology › molecular evolution › evolutionary bioinformatics › evolutionary genomics
genome evolution
0.012003
Simple stochastic birth andz death models of genome evolution: was there enough time for us to evolve? · Bioinform. 2003
Bioinformatics and computational biology › sequence alignment
local alignment
0.011999
IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices · Bioinform. 1999
Bioinformatics and computational biology › sequence alignment
protein sequence alignment
0.011999
IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices · Bioinform. 1999
Bioinformatics and computational biology › protein sequence analysis
protein sequence database search
0.011999
IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices · Bioinform. 1999
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.011999
Winnowing sequences from a database search · RECOMB 1999
Bioinformatics and computational biology › genomics
primer design
0.011992
DIROM: an experimental design interactive system for directed mutagenesis and nucleic acids engineering · Comput. Appl. Biosci. 1992
Bioinformatics and computational biology › synthetic biology
synthetic gene design
0.011992
DIROM: an experimental design interactive system for directed mutagenesis and nucleic acids engineering · Comput. Appl. Biosci. 1992
Information theory › random number generation
interval algorithm
0.011999
Winnowing sequences from a database search · RECOMB 1999

Methods — techniques the papers use, named apart from their topics

symmetric best matches · 0.2edgesearch algorithm · 0.2dominance-based filtering · 0.0stochastic birth and death process · 0.0smith-waterman · 0.0PSI-BLAST · 0.0
YearPublicationVenuePosition
2019 Microbial genome analysis: the COG approach
abstract
For the past 20 years, the Clusters of Orthologous Genes (COG) database had been a popular tool for microbial genome annotation and comparative genomics. Initially created for the purpose of evolutionary classification of protein families, the COG have been used, apart from straightforward functional annotation of sequenced genomes, for such tasks as (i) unification of genome annotation in groups of related organisms; (ii) identification of missing and/or undetected genes in complete microbial genomes; (iii) analysis of genomic neighborhoods, in many cases allowing prediction of novel functional systems; (iv) analysis of metabolic pathways and prediction of alternative forms of enzymes; (v) comparison of organisms by COG functional categories; and (vi) prioritization of targets for structural and functional characterization. Here we review the principles of the COG approach and discuss its key advantages and drawbacks in microbial genome analysis.
Michael Y. Galperin, David M. Kristensen, Kira S. Makarova, Yuri I. Wolf, Eugene V. Koonin
Briefings Bioinform.4
2013 The Vast, Conserved Mammalian lincRNome
abstract
We compare the sets of experimentally validated long intergenic non-coding (linc)RNAs from human and mouse and apply a maximum likelihood approach to estimate the total number of lincRNA genes as well as the size of the conserved part of the lincRNome. Under the assumption that the sets of experimentally validated lincRNAs are random samples of the lincRNomes of the corresponding species, we estimate the total lincRNome size at approximately 40,000 to 50,000 species, at least twice the number of protein-coding genes. We further estimate that the fraction of the human and mouse euchromatic genomes encoding lincRNAs is more than twofold greater than the fraction of protein-coding sequences. Although the sequences of most lincRNAs are much less strongly conserved than protein sequences, the extent of orthology between the lincRNomes is unexpectedly high, with 60 to 70% of the lincRNA genes shared between human and mouse. The orthologous mammalian lincRNAs can be predicted to perform equivalent functions; accordingly, it appears likely that thousands of evolutionarily conserved functional roles of lincRNAs remain to be characterized.
David Managadze, Alexander E. Lobkovsky, Yuri I. Wolf, Svetlana A. Shabalina, Igor B. Rogozin, Eugene V. Koonin
PLoS Comput. Biol.3
2012 Universal Pacemaker of Genome Evolution
abstract
A fundamental observation of comparative genomics is that the distribution of evolution rates across the complete sets of orthologous genes in pairs of related genomes remains virtually unchanged throughout the evolution of life, from bacteria to mammals. The most straightforward explanation for the conservation of this distribution appears to be that the relative evolution rates of all genes remain nearly constant, or in other words, that evolutionary rates of different genes are strongly correlated within each evolving genome. This correlation could be explained by a model that we denoted Universal PaceMaker (UPM) of genome evolution. The UPM model posits that the rate of evolution changes synchronously across genome-wide sets of genes in all evolving lineages. Alternatively, however, the correlation between the evolutionary rates of genes could be a simple consequence of molecular clock (MC). We sought to differentiate between the MC and UPM models by fitting thousands of phylogenetic trees for bacterial and archaeal genes to supertrees that reflect the dominant trend of vertical descent in the evolution of archaea and bacteria and that were constrained according to the two models. The goodness of fit for the UPM model was better than the fit for the MC model, with overwhelming statistical significance, although similarly to the MC, the UPM is strongly overdispersed. Thus, the results of this analysis reveal a universal, genome-wide pacemaker of evolution that could have been in operation throughout the history of life.
Sagi Snir, Yuri I. Wolf, Eugene V. Koonin
PLoS Comput. Biol.2
2011 Computational methods for Gene Orthology inference
abstract
Accurate inference of orthologous genes is a pre-requisite for most comparative genomics studies, and is also important for functional annotation of new genomes. Identification of orthologous gene sets typically involves phylogenetic tree analysis, heuristic algorithms based on sequence conservation, synteny analysis, or some combination of these approaches. The most direct tree-based methods typically rely on the comparison of an individual gene tree with a species tree. Once the two trees are accurately constructed, orthologs are straightforwardly identified by the definition of orthology as those homologs that are related by speciation, rather than gene duplication, at their most recent point of origin. Although ideal for the purpose of orthology identification in principle, phylogenetic trees are computationally expensive to construct for large numbers of genes and genomes, and they often contain errors, especially at large evolutionary distances. Moreover, in many organisms, in particular prokaryotes and viruses, evolution does not appear to have followed a simple 'tree-like' mode, which makes conventional tree reconciliation inapplicable. Other, heuristic methods identify probable orthologs as the closest homologous pairs or groups of genes in a set of organisms. These approaches are faster and easier to automate than tree-based methods, with efficient implementations provided by graph-theoretical algorithms enabling comparisons of thousands of genomes. Comparisons of these two approaches show that, despite conceptual differences, they produce similar sets of orthologs, especially at short evolutionary distances. Synteny also can aid in identification of orthologs. Often, tree-based, sequence similarity- and synteny-based approaches can be combined into flexible hybrid methods.
David M. Kristensen, Yuri I. Wolf, Arcady R. Mushegian, Eugene V. Koonin
Briefings Bioinform.2
2011 Predictability of Evolutionary Trajectories in Fitness Landscapes
abstract
Experimental studies on enzyme evolution show that only a small fraction of all possible mutation trajectories are accessible to evolution. However, these experiments deal with individual enzymes and explore a tiny part of the fitness landscape. We report an exhaustive analysis of fitness landscapes constructed with an off-lattice model of protein folding where fitness is equated with robustness to misfolding. This model mimics the essential features of the interactions between amino acids, is consistent with the key paradigms of protein folding and reproduces the universal distribution of evolutionary rates among orthologous proteins. We introduce mean path divergence as a quantitative measure of the degree to which the starting and ending points determine the path of evolution in fitness landscapes. Global measures of landscape roughness are good predictors of path divergence in all studied landscapes: the mean path divergence is greater in smooth landscapes than in rough ones. The model-derived and experimental landscapes are significantly smoother than random landscapes and resemble additive landscapes perturbed with moderate amounts of noise; thus, these landscapes are substantially robust to mutation. The model landscapes show a deficit of suboptimal peaks even compared with noisy additive landscapes with similar overall roughness. We suggest that smoothness and the substantial deficit of peaks in the fitness landscapes of protein evolution are fundamental consequences of the physics of protein folding.
Alexander E. Lobkovsky, Yuri I. Wolf, Eugene V. Koonin
PLoS Comput. Biol.2
2010 A low-polynomial algorithm for assembling clusters of orthologous groups from intergenomic symmetric best matches
abstract
MOTIVATION: Identifying orthologous genes in multiple genomes is a fundamental task in comparative genomics. Construction of intergenomic symmetrical best matches (SymBets) and joining them into clusters is a popular method of ortholog definition, embodied in several software programs. Despite their wide use, the computational complexity of these programs has not been thoroughly examined. RESULTS: In this work, we show that in the standard approach of iteration through all triangles of SymBets, the memory scales with at least the number of these triangles, O(g(3)) (where g = number of genomes), and construction time scales with the iteration through each pair, i.e. O(g(6)). We propose the EdgeSearch algorithm that iterates over edges in the SymBet graph rather than triangles of SymBets, and as a result has a worst-case complexity of only O(g(3)log g). Several optimizations reduce the run-time even further in realistically sparse graphs. In two real-world datasets of genomes from bacteriophages (POGs) and Mollicutes (MOGs), an implementation of the EdgeSearch algorithm runs about an order of magnitude faster than the original algorithm and scales much better with increasing number of genomes, with only minor differences in the final results, and up to 60 times faster than the popular OrthoMCL program with a 90% overlap between the identified groups of orthologs. AVAILABILITY AND IMPLEMENTATION: C++ source code freely available for download at ftp.ncbi.nih.gov/pub/wolf/COGs/COGsoft/. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.
David M. Kristensen, Lavanya Kannan, Michael K. Coleman, Yuri I. Wolf, Alexander Sorokin, Eugene V. Koonin, Arcady R. Mushegian
Bioinform.4
2004 Computational approaches for the analysis of gene neighbourhoods in prokaryotic genomes
abstract
Gene order in prokaryotes is conserved to a much lesser extent than protein sequences. Only some operons, primarily those that encode physically interacting proteins, are conserved in all or most of the bacterial and archaeal genomes. Nevertheless, even the limited conservation of operon organisation that is observed provides valuable evolutionary and functional clues through multiple genome comparisons. With the rapid growth in the number and diversity of sequenced prokaryotic genomes, functional inferences for uncharacterized genes located in the same conserved gene neighborhood with well-studied genes are becoming increasingly important. In this review, we discuss various computational approaches for identification of conserved gene strings and construction of local alignments of gene orders in prokaryotic genomes.
Igor B. Rogozin, Kira S. Makarova, Yuri I. Wolf, Eugene V. Koonin
Briefings Bioinform.3
2003 Simple stochastic birth andz death models of genome evolution: was there enough time for us to evolve?
abstract
MOTIVATION: The distributions of many genome-associated quantities, including the membership of paralogous gene families can be approximated with power laws. We are interested in developing mathematical models of genome evolution that adequately account for the shape of these distributions and describe the evolutionary dynamics of their formation. RESULTS: We show that simple stochastic models of genome evolution lead to power-law asymptotics of protein domain family size distribution. These models, called Birth, Death and Innovation Models (BDIM), represent a special class of balanced birth-and-death processes, in which domain duplication and deletion rates are asymptotically equal up to the second order. The simplest, linear BDIM shows an excellent fit to the observed distributions of domain family size in diverse prokaryotic and eukaryotic genomes. However, the stochastic version of the linear BDIM explored here predicts that the actual size of large paralogous families is reached on an unrealistically long timescale. We show that introduction of non-linearity, which might be interpreted as interaction of a particular order between individual family members, allows the model to achieve genome evolution rates that are much better compatible with the current estimates of the rates of individual duplication/loss events.
Georgy P. Karev, Yuri I. Wolf, Eugene V. Koonin
Bioinform.2
2003 The COG database: an updated version includes eukaryotes
abstract
BACKGROUND: The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. RESULTS: We describe here a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes. CONCLUSION: The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies.
Roman L. Tatusov, Natalie D. Fedorova, John D. Jackson, Aviva R. Jacobs, Boris Kiryutin, Eugene V. Koonin, Dmitri M. Krylov, Raja Mazumder, Sergei L. Mekhedov, Anastasia N. Nikolskaya, B. Sridhar Rao, Sergei Smirnov, Alexander V. Sverdlov, Sona Vasudevan, Yuri I. Wolf, Jodie J. Yin, Darren A. Natale
BMC Bioinform.15
1999 Winnowing sequences from a database search
abstract
In database searches for sequence similarity, matches to a distinct sequence region (e.g. protein domain) are frequently obscured by numerous matches to another region of the same sequence.In order to cope with this problem, algorithms are developed to discard redundant matches.One model for this problem begins with a list of intervals, each with an associated score; each interval gives the range of positions in the query sequence that align to a database sequence, and the score is that of the alignment.If interval I is contained in interval J, and I's score is less than J's, then I is said to be dominated by J.The problem is then to identify each interval that is dominated by at least K other intervals, where K is a given level of "tolerable redundancy."An algorithm is developed to solve the problem in O(N log N) time and O(N*) space, where N is the number of intervals and N' is a precisely defined value that never exceeds N and is frequently much smaller.This criterion for discarding database hits has been implemented in the Blast program, as illustrated herein with examples.Several variations and extensions of this approach are also described.
Piotr Berman, Zheng Zhang 0004, Yuri I. Wolf, Eugene V. Koonin, Webb Miller
RECOMB3
1999 IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices
abstract
Abstract Motivation: Many studies have shown that database searches using position-specific score matrices (PSSMs) or profiles as queries are more effective at identifying distant protein relationships than are searches that use simple sequences as queries. One popular program for constructing a PSSM and comparing it with a database of sequences is Position-Specific Iterated BLAST (PSI-BLAST). Results: This paper describes a new software package, IMPALA, designed for the complementary procedure of comparing a single query sequence with a database of PSI-BLAST-generated PSSMs. We illustrate the use of IMPALA to search a database of PSSMs for protein folds, and one for protein domains involved in signal transduction. IMPALA’s sensitivity to distant biological relationships is very similar to that of PSI-BLAST. However, IMPALA employs a more refined analysis of statistical significance and, unlike PSI-BLAST, guarantees the output of the optimal local alignment by using the rigorous Smith–Waterman algorithm. Also, it is considerably faster when run with a large database of PSSMs than is BLAST or PSI-BLAST when run against the complete non-redundant protein database. Availability: The IMPALA source code, the wolf1187 database, and the aravind105 database are freely available from the NCBI ftp site ncbi.nlm.nih.gov. The databases may be found in the subdirectory ftp://ncbi.nlm.nih.gov/pub/impala. The source code is in ftp://ncbi.nlm.nih.gov/toolbox/ncbi˙tools. Some IMPALA executables for different implementations of UNIX are in ftp://ncbi.nlm.nih.gov/blast/executables. IMPALA has been added as a search option on the Blocks Database Server (http://blocks.fhcrc.org/blocks/impala.html)using a library of PSSMs derived from the BLOCKS database. Contact: [email protected]
Alejandro A. Schäffer, Yuri I. Wolf, Chris P. Ponting, Eugene V. Koonin, L. Aravind, Stephen F. Altschul
Bioinform.2
1992 DIROM: an experimental design interactive system for directed mutagenesis and nucleic acids engineering
abstract
A computer system DIROM for oligonucleotide-directed mutagenesis and artificial gene design has been designed for better experimental planning and control. DIROM permits searching for optimal oligonucleotides with respect to certain important parameters, namely sufficient energy of oligonucleotide-target hybridization, the secondary structure of oligonucleotide and target DNA, the presence of alternate binding sites in the target DNA and terminal G/C pairs. It can also be used to plan polymerase chain reaction experiments, for optimal primer selection, in sequencing, etc. DIROM enables one to search for both existing and potential restriction sites, to perform vector + target sequence construction. The system consists of a set of original algorithms that formalize the empirical knowledge of oligonucleotide action as primers.
Kira S. Makarova, A. V. Mazin, Yuri I. Wolf, V. V. Soloviev
Comput. Appl. Biosci.3