VLDB 2026 Research / reviewers in the wild / expert
Golan Yona
dblp:09/4046
· DBLP profile ↗
20ranked-venue papers
4as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 4 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 50% Graph data management · 50% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics › genomic data compression
genome compression |
0.2 | 1 | 2013 | The human genome contracts again · Bioinform. 2013 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2006 | Effective similarity measures for expression profiles · Bioinform. 2006 |
Bioinformatics and computational biology › protein sequence analysis › protein family analysis
protein family modeling |
0.1 | 2 | 2001 | Variations on probabilistic suffix trees: statistical modeling and prediction of protein families · Bioinform. 2001 Modeling protein families using probabilistic suffix trees · RECOMB 1999 |
Bioinformatics and computational biology › structural bioinformatics
protein structure classification |
0.1 | 2 | 2000 | A unified sequence-structure classification of protein sequences: combining sequence and structure in a map of the protein space · RECOMB 2000 Towards a Complete Map of the Protein Space Based on a Unified Sequence and Structure Analysis of All Known Proteins · ISMB 2000 |
Bioinformatics and computational biology › bioinformatics infrastructure
genomic data storage |
0.0 | 1 | 2013 | The human genome contracts again · Bioinform. 2013 |
Bioinformatics and computational biology › protein structure analysis › protein domain identification
protein domain prediction |
0.0 | 1 | 2004 | Automatic prediction of protein domains from sequence information using a hybrid learning system · Bioinform. 2004 |
Algorithms and data structures
embedding |
0.0 | 1 | 2004 | Distributional Scaling: An Algorithm for Structure-Preserving Embedding of Metric and Nonmetric Spaces · J. Mach. Learn. Res. 2004 |
Algorithms and data structures
metric embedding |
0.0 | 1 | 2004 | Distributional Scaling: An Algorithm for Structure-Preserving Embedding of Metric and Nonmetric Spaces · J. Mach. Learn. Res. 2004 |
Bioinformatics and computational biology › protein structure analysis
protein domain identification |
0.0 | 1 | 2003 | A multi-expert system for the automatic detection of protein domains from sequence information · RECOMB 2003 |
Bioinformatics and computational biology
protein function prediction |
0.0 | 1 | 2003 | Using a mixture of probabilistic decision trees for direct prediction of protein function · RECOMB 2003 |
Bioinformatics and computational biology
protein structure prediction |
0.0 | 1 | 2003 | A multi-expert system for the automatic detection of protein domains from sequence information · RECOMB 2003 |
Bioinformatics and computational biology › protein sequence analysis
protein homology detection |
0.0 | 1 | 2000 | A unified sequence-structure classification of protein sequences: combining sequence and structure in a map of the protein space · RECOMB 2000 |
Bioinformatics and computational biology › structural bioinformatics › sequence-structure relationship
sequence-structure analysis |
0.0 | 1 | 2000 | Towards a Complete Map of the Protein Space Based on a Unified Sequence and Structure Analysis of All Known Proteins · ISMB 2000 |
Bioinformatics and computational biology › protein function prediction
protein classification |
0.0 | 1 | 1998 | A Map of the Protein Space: An Automatic Hierarchical Classification of all Protein Sequences · ISMB 1998 |
Bioinformatics and computational biology
multiple sequence alignment |
0.0 | 1 | 2003 | A multi-expert system for the automatic detection of protein domains from sequence information · RECOMB 2003 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 2003 | A multi-expert system for the automatic detection of protein domains from sequence information · RECOMB 2003 |
Bioinformatics and computational biology › protein function prediction › protein classification
protein sequence classification |
0.0 | 1 | 2001 | Variations on probabilistic suffix trees: statistical modeling and prediction of protein families · Bioinform. 2001 |
Methods — techniques the papers use, named apart from their topics
entropy coding · 0.2SNP-based compression · 0.2probabilistic model · 0.1neural network · 0.1topology detection algorithms · 0.1statistical analysis · 0.1similarity measure · 0.1multiple sequence alignment · 0.0multidimensional scaling · 0.0mixture of probabilistic decision trees · 0.0probabilistic suffix tree · 0.0sequence analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | The human genome contracts againabstractUNLABELLED: The number of human genomes that have been sequenced completely for different individuals has increased rapidly in recent years. Storing and transferring complete genomes between computers for the purpose of applying various applications and analysis tools will soon become a major hurdle, hindering the analysis phase. Therefore, there is a growing need to compress these data efficiently. Here, we describe a technique to compress human genomes based on entropy coding, using a reference genome and known Single Nucleotide Polymorphisms (SNPs). Furthermore, we explore several intrinsic features of genomes and information in other genomic databases to further improve the compression attained. Using these methods, we compress James Watson's genome to 2.5 megabytes (MB), improving on recent work by 37%. Similar compression is obtained for most genomes available from the 1000 Genomes Project. Our biologically inspired techniques promise even greater gains for genomes of lower organisms and for human genomes as more genomic data become available. AVAILABILITY: Code is available at sourceforge.net/projects/genomezip/ Dmitri S. Pavlichin, Tsachy Weissman, Golan Yona |
Bioinform. | 3 |
| 2013 | QualComp: a new lossy compressor for quality scores based on rate distortion theoryabstractBACKGROUND: Next Generation Sequencing technologies have revolutionized many fields in biology by reducing the time and cost required for sequencing. As a result, large amounts of sequencing data are being generated. A typical sequencing data file may occupy tens or even hundreds of gigabytes of disk space, prohibitively large for many users. This data consists of both the nucleotide sequences and per-base quality scores that indicate the level of confidence in the readout of these sequences. Quality scores account for about half of the required disk space in the commonly used FASTQ format (before compression), and therefore the compression of the quality scores can significantly reduce storage requirements and speed up analysis and transmission of sequencing data. RESULTS: In this paper, we present a new scheme for the lossy compression of the quality scores, to address the problem of storage. Our framework allows the user to specify the rate (bits per quality score) prior to compression, independent of the data to be compressed. Our algorithm can work at any rate, unlike other lossy compression algorithms. We envisage our algorithm as being part of a more general compression scheme that works with the entire FASTQ file. Numerical experiments show that we can achieve a better mean squared error (MSE) for small rates (bits per quality score) than other lossy compression schemes. For the organism PhiX, whose assembled genome is known and assumed to be correct, we show that it is possible to achieve a significant reduction in size with little compromise in performance on downstream applications (e.g., alignment). CONCLUSIONS: QualComp is an open source software package, written in C and freely available for download at https://sourceforge.net/projects/qualcomp. Idoia Ochoa, Himanshu Asnani, Dinesh Bharadia, Mainak Chowdhury, Tsachy Weissman, Golan Yona |
BMC Bioinform. | 6 |
| 2007 | Topology Search over Biological DatabasesabstractWe introduce the notion of a data topology and the problem of topology search over databases. A data topology summarizes the set of all possible relationships that connect a given set of entities. Topology search enables users to search for data topologies that relate entities in a large database, and to effectively summarize and rank these relationships. Using topology search over a biological database, users can ask, for example, how transcription factor proteins are related to DNAs in humans. However, detecting topologies in large databases is a difficult problem because entities can be connected in multiple ways. In this paper, we formalize the notion of data topologies, develop efficient algorithms for computing data topologies based on user queries, and evaluate our algorithms using a real biological database, the Biozon database (www.biozon.org). Jayavel Shanmugasundaram, Golan Yona |
ICDE | 3 |
| 2006 | Effective similarity measures for expression profilesabstractIt is commonly accepted that genes with similar expression profiles are functionally related. However, there are many ways one can measure the similarity of expression profiles, and it is not clear a priori what is the most effective one. Moreover, so far no clear distinction has been made as for the type of the functional link between genes as suggested by microarray data. Similarly expressed genes can be part of the same complex as interacting partners; they can participate in the same pathway without interacting directly; they can perform similar functions; or they can simply have similar regulatory sequences. Here we conduct a study of the notion of functional link as implied from expression data. We analyze different similarity measures of gene expression profiles and assess their usefulness and robustness in detecting biological relationships by comparing the similarity scores with results obtained from databases of interacting proteins, promoter signals and cellular pathways, as well as through sequence comparisons. We also introduce variations on similarity measures that are based on statistical analysis and better discriminate genes which are functionally nearby and faraway. Our tools can be used to assess other similarity measures for expression profiles, and are accessible at biozon.org/tools/expression/ Golan Yona, William Dirks, Shafquat Rahman, David M. Lin |
Bioinform. | 1 |
| 2006 | BIOZON: a system for unification, management and analysis of heterogeneous biological dataabstractBACKGROUND: Integration of heterogeneous data types is a challenging problem, especially in biology, where the number of databases and data types increase rapidly. Amongst the problems that one has to face are integrity, consistency, redundancy, connectivity, expressiveness and updatability. DESCRIPTION: Here we present a system (Biozon) that addresses these problems, and offers biologists a new knowledge resource to navigate through and explore. Biozon unifies multiple biological databases consisting of a variety of data types (such as DNA sequences, proteins, interactions and cellular pathways). It is fundamentally different from previous efforts as it uses a single extensive and tightly connected graph schema wrapped with hierarchical ontology of documents and relations. Beyond warehousing existing data, Biozon computes and stores novel derived data, such as similarity relationships and functional predictions. The integration of similarity data allows propagation of knowledge through inference and fuzzy searches. Sophisticated methods of query that span multiple data types were implemented and first-of-a-kind biological ranking systems were explored and integrated. CONCLUSION: The Biozon system is an extensive knowledge resource of heterogeneous biological data. Currently, it holds more than 100 million biological documents and 6.5 billion relations between them. The database is accessible through an advanced web interface that supports complex queries, "fuzzy" searches, data materialization and more, online at http://biozon.org. Aaron Birkland, Golan Yona |
BMC Bioinform. | 2 |
| 2006 | Hubs of knowledge: using the functional link structure in Biozon to mine for biologically significant entitiesabstractBACKGROUND: Existing biological databases support a variety of queries such as keyword or definition search. However, they do not provide any measure of relevance for the instances reported, and result sets are usually sorted arbitrarily. RESULTS: We describe a system that builds upon the complex infrastructure of the Biozon database and applies methods similar to those of Google to rank documents that match queries. We explore different prominence models and study the spectral properties of the corresponding data graphs. We evaluate the information content of principal and non-principal eigenspaces, and test various scoring functions which combine contributions from multiple eigenspaces. We also test the effect of similarity data and other variations which are unique to the biological knowledge domain on the quality of the results. Query result sets are assessed using a probabilistic approach that measures the significance of coherence between directly connected nodes in the data graph. This model allows us, for the first time, to compare different prominence models quantitatively and effectively and to observe unique trends. CONCLUSION: Our tests show that the ranked query results outperform unsorted results with respect to our significance measure and the top ranked entities are typically linked to many other biological entities. Our study resulted in a working ranking system of biological entities that was integrated into Biozon at http://biozon.org. Paul Shafer, Timothy Isganitis, Golan Yona |
BMC Bioinform. | 3 |
| 2005 | The distance-profile representation and its application to detection of distantly related protein familiesabstractBACKGROUND: Detecting homology between remotely related protein families is an important problem in computational biology since the biological properties of uncharacterized proteins can often be inferred from those of homologous proteins. Many existing approaches address this problem by measuring the similarity between proteins through sequence or structural alignment. However, these methods do not exploit collective aspects of the protein space and the computed scores are often noisy and frequently fail to recognize distantly related protein families. RESULTS: We describe an algorithm that improves over the state of the art in homology detection by utilizing global information on the proximity of entities in the protein space. Our method relies on a vectorial representation of proteins and protein families and uses structure-specific association measures between proteins and template structures to form a high-dimensional feature vector for each query protein. These vectors are then processed and transformed to sparse feature vectors that are treated as statistical fingerprints of the query proteins. The new representation induces a new metric between proteins measured by the statistical difference between their corresponding probability distributions. CONCLUSION: Using several performance measures we show that the new tool considerably improves the performance in recognizing distant homologies compared to existing approaches such as PSIBLAST and FUGUE. Chin-Jen Ku, Golan Yona |
BMC Bioinform. | 2 |
| 2005 | Automation of gene assignments to metabolic pathways using high-throughput expression dataabstractBACKGROUND: Accurate assignment of genes to pathways is essential in order to understand the functional role of genes and to map the existing pathways in a given genome. Existing algorithms predict pathways by extrapolating experimental data in one organism to other organisms for which this data is not available. However, current systems classify all genes that belong to a specific EC family to all the pathways that contain the corresponding enzymatic reaction, and thus introduce ambiguity. RESULTS: Here we describe an algorithm for assignment of genes to cellular pathways that addresses this problem by selectively assigning specific genes to pathways. Our algorithm uses the set of experimentally elucidated metabolic pathways from MetaCyc, together with statistical models of enzyme families and expression data to assign genes to enzyme families and pathways by optimizing correlated co-expression, while minimizing conflicts due to shared assignments among pathways. Our algorithm also identifies alternative ("backup") genes and addresses the multi-domain nature of proteins. We apply our model to assign genes to pathways in the Yeast genome and compare the results for genes that were assigned experimentally. Our assignments are consistent with the experimentally verified assignments and reflect characteristic properties of cellular pathways. CONCLUSION: We present an algorithm for automatic assignment of genes to metabolic pathways. The algorithm utilizes expression data and reduces the ambiguity that characterizes assignments that are based only on EC numbers. Liviu Popescu, Golan Yona |
BMC Bioinform. | 2 |
| 2004 | Automatic prediction of protein domains from sequence information using a hybrid learning systemabstractMOTIVATION: We describe a novel method for detecting the domain structure of a protein from sequence information alone. The method is based on analyzing multiple sequence alignments that are derived from a database search. Multiple measures are defined to quantify the domain information content of each position along the sequence and are combined into a single predictor using a neural network. The output is further smoothed and post-processed using a probabilistic model to predict the most likely transition positions between domains. RESULTS: The method was assessed using the domain definitions in SCOP and CATH for proteins of known structure and was compared with several other existing methods. Our method performs well both in terms of accuracy and sensitivity. It improves significantly over the best methods available, even some of the semi-manual ones, while being fully automatic. Our method can also be used to suggest and verify domain partitions based on structural data. A few examples of predicted domain definitions and alternative partitions, as suggested by our method, are also discussed. AVAILABILITY: An online domain-prediction server is available at http://biozon.org/tools/domains/ Niranjan Nagarajan, Golan Yona |
Bioinform. | 2 |
| 2004 | Protein family comparison using statistical models and predicted structural informationabstractBACKGROUND: This paper presents a simple method to increase the sensitivity of protein family comparisons by incorporating secondary structure (SS) information. We build upon the effective information theory approach towards profile-profile comparison described in [Yona & Levitt 2002]. Our method augments profile columns using PSIPRED secondary structure predictions and assesses statistical similarity using information theoretical principles. RESULTS: Our tests show that this tool detects more similarities between protein families of distant homology than the previous primary sequence-based method. A very significant improvement in performance is observed when the real secondary structure is used. CONCLUSIONS: Integration of primary and secondary structure information can substantially improve detection of relationships between remotely related protein families. Richard Chung, Golan Yona |
BMC Bioinform. | 2 |
| 2004 | On Prediction Using Variable Order Markov ModelsabstractThis paper is concerned with algorithms for prediction of discrete sequences over a finite alphabet, using variable order Markov models. The class of such algorithms is large and in principle includes any lossless compression algorithm. We focus on six prominent prediction algorithms, including Context Tree Weighting (CTW), Prediction by Partial Match (PPM) and Probabilistic Suffix Trees (PSTs). We discuss the properties of these algorithms and compare their performance using real life sequences from three domains: proteins, English text and music pieces. The comparison is made with respect to prediction quality as measured by the average log-loss. We also compare classification algorithms based on these predictors with respect to a number of large protein classification tasks. Our results indicate that a ``decomposed'' CTW (a variant of the CTW algorithm) and PPM outperform all other algorithms in sequence prediction tasks. Somewhat surprisingly, a different algorithm, which is a modification of the Lempel-Ziv compression algorithm, significantly outperforms all algorithms on the protein classification problems. Ron Begleiter, Ran El-Yaniv, Golan Yona |
J. Artif. Intell. Res. | 3 |
| 2004 | Distributional Scaling: An Algorithm for Structure-Preserving Embedding of Metric and Nonmetric Spaces
Michael Quist, Golan Yona |
J. Mach. Learn. Res. | 2 |
| 2003 | A multi-expert system for the automatic detection of protein domains from sequence informationabstractWe describe a novel method for detecting the domain structure of a protein from sequence information alone. The method is based on analyzing multiple sequence alignments that are derived from a database search. Multiple measures are defined to quantify the domain information content of each position along the sequence, and are combined into a single predictor using a neural network. The output is further smoothed and post-processed using a probabilistic model to predict the most likely transition or boundary positions between domains. The method was assessed using the domain definitions in SCOP for proteins of known structures and was compared to several other existing methods. Our method improves significantly over the best method available, the semi-manual PFam domain database, while being fully automatic. Our method can also be used to verify domain partitions based on structural data. Few examples of predicted domain definitions and alternative partitions, as suggested by our method, are also discussed. Niranjan Nagarajan, Golan Yona |
RECOMB | 2 |
| 2003 | Using a mixture of probabilistic decision trees for direct prediction of protein functionabstractWe study the direct relationship between basic protein properties and their function. Our goal is to develop a new tool for functional prediction that can be used to complement and support other techniques based on sequence or structure information. In order to define this new measure of similarity between proteins we collected a set of 453 features and properties that characterize proteins and are believed to be correlated and related to structural and functional aspects of proteins. Among these properties are the composition and fraction of different groups of amino acids, predicted secondary structure content, molecular weight, average hydrophobicity, isoelectric point and others, as well as a set of properties that are extracted from database records of known protein sequences, such as subcellular location, tissue specificity, and others.We introduce the mixture model of probabilistic decision trees to learn the set of potentially complex relationships between features and function. To study these correlations, trees are created and tested on the Pfam sequence-based classification of proteins and the EC classification of enzyme families. The model is very effective in learning highly diverged protein families or families that are not defined based on sequence. The resulting tree structure indicates the properties that are strongly correlated with structural and functional aspects of protein families, and can be used to suggest a concise definition of a protein family. Umar Syed, Golan Yona |
RECOMB | 2 |
| 2002 | A New Nonparametric Pairwise Clustering Algorithm Based on Iterative Estimation of Distance Profiles
Shlomo Dubnov, Ran El-Yaniv, Yoram Gdalyahu, Elad Schneidman, Naftali Tishby, Golan Yona |
Mach. Learn. | 6 |
| 2001 | Variations on probabilistic suffix trees: statistical modeling and prediction of protein familiesabstractAbstract Motivation: We present a method for modeling protein families by means of probabilistic suffix trees (PSTs). The method is based on identifying significant patterns in a set of related protein sequences. The patterns can be of arbitrary length, and the input sequences do not need to be aligned, nor is delineation of domain boundaries required. The method is automatic, and can be applied, without assuming any preliminary biological information, with surprising success. Basic biological considerations such as amino acid background probabilities, and amino acids substitution probabilities can be incorporated to improve performance. Results: The PST can serve as a predictive tool for protein sequence classification, and for detecting conserved patterns (possibly functionally or structurally important) within protein sequences. The method was tested on the Pfam database of protein families with more than satisfactory performance. Exhaustive evaluations show that the PST model detects much more related sequences than pairwise methods such as Gapped-BLAST, and is almost as sensitive as a hidden Markov model that is trained from a multiple alignment of the input sequences, while being much faster. Availability: The programs are available upon request from the authors. Contact: [email protected]; [email protected] * To whom correspondence should be addressed. 3 Address starting from January 2001: Department of Computer Science, Cornell University, Ithaca, NY 14853, USA. Gill Bejerano, Golan Yona |
Bioinform. | 2 |
| 2000 | Towards a Complete Map of the Protein Space Based on a Unified Sequence and Structure Analysis of All Known Proteins
Golan Yona, Michael Levitt 0001 |
ISMB | 1 |
| 2000 | A unified sequence-structure classification of protein sequences: combining sequence and structure in a map of the protein spaceabstractWe analyze all known protein sequences in search for a global map of protein space that is consistent in terms of both sequence and structure. Our goal is to define clusters of homologous protein domains, beyond those detected by sequence-based methods alone, and then to build a three-dimensional (3D) model for each of the sequences that are homologous to sequences of known 3D structure. This analysis uses both sequence and structure based metrics in the analysis of all protein sequences in a non-redundant (NR) database, comprising all major sequence databases. Golan Yona, Michael Levitt 0001 |
RECOMB | 1 |
| 1999 | Modeling protein families using probabilistic suffix treesabstractproteins which await analysis.We present a method for modeling protein families by means of probabilistic suffix trees (PSTs).The method is based on identifying significant patterns in a set of related protein sequences.The input sequences do not Gill Bejerano, Golan Yona |
RECOMB | 2 |
| 1998 | A Map of the Protein Space: An Automatic Hierarchical Classification of all Protein Sequences
Golan Yona, Nathan Linial, Naftali Tishby, Michal Linial |
ISMB | 1 |