VLDB 2026 Research / reviewers in the wild / expert
Ron Unger
dblp:75/3917
· DBLP profile ↗
21ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0003-4153-3922ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
12 papers |
Bioinformatics and computational biology · 100% Computational science and engineering · 0% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
comparative genomics |
0.8 | 1 | 2024 | ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024 |
Bioinformatics and computational biology › molecular evolution
evolutionary conservation analysis |
0.8 | 1 | 2024 | ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024 |
Bioinformatics and computational biology › gene regulation
gene regulation analysis |
0.8 | 1 | 2024 | ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024 |
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction |
0.8 | 1 | 2024 | ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024 |
Bioinformatics and computational biology
protein sequence analysis |
0.7 | 3 | 2022 | DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts · Bioinform. 2022 Low folding propensity and high translation efficiency distinguish in vivo substrates of GroEL from other Escherichia coli proteins · Bioinform. 2007 A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999 |
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model |
0.6 | 1 | 2022 | DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts · Bioinform. 2022 |
Bioinformatics and computational biology › protein function prediction › protein classification
protein sequence classification |
0.6 | 1 | 2022 | DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts · Bioinform. 2022 |
Bioinformatics and computational biology › structural bioinformatics
protein structure |
0.2 | 1 | 2016 | e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures · Bioinform. 2016 |
Bioinformatics and computational biology › transcriptomics
RNA editing |
0.2 | 1 | 2016 | e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures · Bioinform. 2016 |
Bioinformatics and computational biology
protein structure analysis |
0.2 | 2 | 2013 | Non-local residue-residue contacts in proteins are more conserved than local ones · Bioinform. 2013 A tale of two tails: why are terminal residues of proteins exposed? · Bioinform. 2007 |
Bioinformatics and computational biology
protein structure prediction |
0.1 | 2 | 2011 | Hidden conformations in protein structures · Bioinform. 2011 Genetic Algorithms for Protein Threading · ISMB 1998 |
Bioinformatics and computational biology › systems biology
gene regulatory network modeling |
0.1 | 1 | 2012 | Static network structure can be used to model the phenotypic effects of perturbations in regulatory networks · Bioinform. 2012 |
Bioinformatics and computational biology › protein structure prediction
protein folding |
0.1 | 2 | 2007 | Low folding propensity and high translation efficiency distinguish in vivo substrates of GroEL from other Escherichia coli proteins · Bioinform. 2007 A tale of two tails: why are terminal residues of proteins exposed? · Bioinform. 2007 |
Bioinformatics and computational biology › protein structure prediction
residue contact prediction |
0.1 | 1 | 2011 | Hidden conformations in protein structures · Bioinform. 2011 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2010 | Composition bias and the origin of ORFan genes · Bioinform. 2010 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.1 | 1 | 2016 | e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures · Bioinform. 2016 |
Bioinformatics and computational biology › systems bioinformatics
pathway analysis |
0.0 | 1 | 2012 | Static network structure can be used to model the phenotypic effects of perturbations in regulatory networks · Bioinform. 2012 |
Bioinformatics and computational biology › protein structure analysis › structural alignment
circular permutation detection |
0.0 | 1 | 1999 | A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition |
0.0 | 1 | 1998 | Genetic Algorithms for Protein Threading · ISMB 1998 |
Algorithms and data structures › sequence algorithms › string algorithms
edit distance |
0.0 | 1 | 1999 | A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999 |
Algorithms and data structures › sequence algorithms › string algorithms
sequence alignment |
0.0 | 1 | 1999 | A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 1986 | DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis · Comput. Appl. Biosci. 1986 |
Bioinformatics and computational biology › sequence analysis
sequence homology |
0.0 | 1 | 1986 | DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis · Comput. Appl. Biosci. 1986 |
Computational science and engineering
structural analysis |
0.0 | 1 | 1986 | DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis · Comput. Appl. Biosci. 1986 |
Methods — techniques the papers use, named apart from their topics
motif scanning · 0.8TSS dataset quality assessment · 0.8transfer learning · 0.6protein language model · 0.6knowledge distillation · 0.6structure mapping · 0.2recoding event modeling · 0.2sequence analysis · 0.2structural comparison · 0.2static network structure analysis · 0.1dynamic programming · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ElemeNT 2023: an enhanced tool for detection and curation of core promoter elementsabstractMOTIVATION: Prediction and identification of core promoter elements and transcription factor binding sites is essential for understanding the mechanism of transcription initiation and deciphering the biological activity of a specific locus. Thus, there is a need for an up-to-date tool to detect and curate core promoter elements/motifs in any provided nucleotide sequences. RESULTS: Here, we introduce ElemeNT 2023-a new and enhanced version of the Elements Navigation Tool, which provides novel capabilities for assessing evolutionary conservation and for readily evaluating the quality of high-throughput transcription start site (TSS) datasets, leveraging preferential motif positioning. ElemeNT 2023 is accessible both as a fast web-based tool and via command line (no coding skills are required to run the tool). While this tool is focused on core promoter elements, it can also be used for searching any user-defined motif, including sequence-specific DNA binding sites. Furthermore, ElemeNT's CORE database, which contains predicted core promoter elements around annotated TSSs, is now expanded to cover 10 species, ranging from worms to human. In this applications note, we describe the new workflow and demonstrate a case study using ElemeNT 2023 for core promoter composition analysis of diverse species, revealing motif prevalence and highlighting evolutionary insights. We discuss how this tool facilitates the exploration of uncharted transcriptomic data, appraises TSS quality, and aids in designing synthetic promoters for gene expression optimization. Taken together, ElemeNT 2023 empowers researchers with comprehensive tools for meticulous analysis of sequence elements and gene expression strategies. AVAILABILITY AND IMPLEMENTATION: ElemeNT 2023 is freely available at https://www.juven-gershonlab.org/resources/element-v2023/. The source code and command line version of ElemeNT 2023 are available at https://github.com/OritAdato/ElemeNT. No coding skills are required to run the tool. Orit Adato, Anna Sloutskin, Hodaya Komemi, Ian Brabb, Sascha Duttke, Philipp Bucher, Ron Unger, Tamar Juven-Gershon |
Bioinform. | 7 |
| 2022 | DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterpartsabstractSUMMARY: Recently, deep learning models, initially developed in the field of natural language processing (NLP), were applied successfully to analyze protein sequences. A major drawback of these models is their size in terms of the number of parameters needed to be fitted and the amount of computational resources they require. Recently, 'distilled' models using the concept of student and teacher networks have been widely used in NLP. Here, we adapted this concept to the problem of protein sequence analysis, by developing DistilProtBert, a distilled version of the successful ProtBert model. Implementing this approach, we reduced the size of the network and the running time by 50%, and the computational resources needed for pretraining by 98% relative to ProtBert model. Using two published tasks, we showed that the performance of the distilled model approaches that of the full model. We next tested the ability of DistilProtBert to distinguish between real and random protein sequences. The task is highly challenging if the composition is maintained on the level of singlet, doublet and triplet amino acids. Indeed, traditional machine-learning algorithms have difficulties with this task. Here, we show that DistilProtBert preforms very well on singlet, doublet and even triplet-shuffled versions of the human proteome, with AUC of 0.92, 0.91 and 0.87, respectively. Finally, we suggest that by examining the small number of false-positive classifications (i.e. shuffled sequences classified as proteins by DistilProtBert), we may be able to identify de novo potential natural-like proteins based on random shuffling of amino acid sequences. AVAILABILITY AND IMPLEMENTATION: https://github.com/yarongef/DistilProtBert. Yaron Geffen, Yanay Ofran, Ron Unger |
Bioinform. | 3 |
| 2016 | e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structuresabstractUNLABELLED: e23D, a database of A-to-I RNA editing sites from human, mouse and fly mapped to evolutionary related protein 3D structures, is presented. Genomic coordinates of A-to-I RNA editing sites are converted to protein coordinates and mapped onto 3D structures from PDB or theoretical models from ModBase. e23D allows visualization of the protein structure, modeling of recoding events and orientation of the editing with respect to nearby genomic functional sites from databases of disease causing mutations and genomic polymorphism. AVAILABILITY AND IMPLEMENTATION: http://www.sheba-cancer.org.il/e23D CONTACT: [email protected] or [email protected]. Oz Solomon, Eran Eyal, Ninette Amariglio, Ron Unger, Gideon Rechavi |
Bioinform. | 4 |
| 2013 | Non-local residue-residue contacts in proteins are more conserved than local onesabstractNon-covalent residue-residue contacts drive the folding of proteins and stabilize them. They may be local-i.e. involve residues that are close in sequence, or non-local. It has been suggested that, in most proteins, local contacts drive protein folding by providing crucial constraints of the conformational space, thus allowing proteins to fold. We compared residues that are involved in local contacts to residues that are involved in non-local contacts and found that, in most proteins, residues in non-local contacts are significantly more conserved evolutionarily than residues in local contacts. Moreover, non-local contacts are more structurally conserved: a contact between positions that are distant in sequence is more likely to exist in many structural homologues compared with a contact between positions that are close in sequence. These results provide new insights into the mechanisms of protein folding and may allow for better prediction of critical intra-chain contacts. Orly Noivirt-Brik, Gershon Hazan, Ron Unger, Yanay Ofran |
Bioinform. | 3 |
| 2012 | Static network structure can be used to model the phenotypic effects of perturbations in regulatory networksabstractMOTIVATION: Biological processes are dynamic, whereas the networks that depict them are typically static. Quantitative modeling using differential equations or logic-based functions can offer quantitative predictions of the behavior of biological systems, but they require detailed experimental characterization of interaction kinetics, which is typically unavailable. To determine to what extent complex biological processes can be modeled and analyzed using only the static structure of the network (i.e. the direction and sign of the edges), we attempt to predict the phenotypic effect of perturbations in biological networks from the static network structure. RESULTS: We analyzed three networks from different sources: The EGFR/MAPK and PI3K/AKT network from a detailed experimental study, the TNF regulatory network from the STRING database and a large network of all NCI-curated pathways from the Protein Interaction Database. Altogether, we predicted the effect of 39 perturbations (e.g. by one or two drugs) on 433 target proteins/genes. In up to 82% of the cases, an algorithm that used only the static structure of the network correctly predicted whether any given protein/gene is upregulated or downregulated as a result of perturbations of other proteins/genes. CONCLUSION: While quantitative modeling requires detailed experimental data and heavy computations, which limit its scalability for large networks, a wiring-based approach can use available data from pathway and interaction databases and may be scalable. These results lay the foundations for a large-scale approach of predicting phenotypes based on the schematic structure of networks. Ariel Feiglin, Adar Hacohen, Avital Sarusi, Jasmin Fisher, Ron Unger, Yanay Ofran |
Bioinform. | 5 |
| 2011 | Hidden conformations in protein structuresabstractMOTIVATION: Prediction of interactions between protein residues (contact map prediction) can facilitate various aspects of 3D structure modeling. However, the accuracy of ab initio contact prediction is still limited. As structural genomics initiatives move ahead, solved structures of homologous proteins can be used as multiple templates to improve contact prediction of the major conformation of an unsolved target protein. Furthermore, multiple templates may provide a wider view of the protein's conformational space. However, successful usage of multiple structural templates is not straightforward, due to their variable relevance to the target protein, and because of data redundancy issues. RESULTS: We present here an algorithm that addresses these two limitations in the use of multiple structure templates. First, the algorithm unites contact maps extracted from templates sharing high sequence similarity with each other in a fashion that acknowledges the possibility of multiple conformations. Next, it weights the resulting united maps in inverse proportion to their evolutionary distance from the target protein. Testing this algorithm against CASP8 targets resulted in high precision contact maps. Remarkably, based solely on structural data of remote homologues, our algorithm identified residue-residue interactions that account for all the known conformations of calmodulin, a multifaceted protein. Therefore, employing multiple templates, which improves prediction of contact maps, can also be used to reveal novel conformations. As multiple templates will soon be available for most proteins, our scheme suggests an effective procedure for their optimal consideration. AVAILABILITY: A Perl script implementing the WMC algorithm described in this article is freely available for academic use at http://tau.ac.il/~haimash/WMC. Haim Ashkenazy, Ron Unger, Yossef Kliger |
Bioinform. | 2 |
| 2010 | Analysis of the effects of lifetime learning on population fitness using vose modelabstractVose's dynamical systems model of the simple genetic algorithm (SGA) is an exact model that uses mathematical operations to capture the dynamical behavior of genetic algorithms. The original model was defined for a simple genetic algorithm. This paper suggests how to extend the model and incorporate two kinds of learning, Darwinian and Lamarckian, into the framework of the Vose model. The extension provides a new theoretical framework to examine the effects of lifetime learning on the fitness of a population. We analyze the asymptotic behavior of different hybrid algorithms on an infinite population vector and compare it to the behavior of the classical genetic algorithm on various population sizes. Our experiments show that Lamarckian-like inheritance - direct transfer of lifetime learning results to offsprings - allows quicker genetic adaptation. However, functions exist where the simple genetic algorithms without learning, as well as Lamarckian evolution, converge to the same local optimum, while genetic search based on Darwinian inheritance converges to the global optimum. Roi Yehoshua, Mireille Avigal, Ron Unger |
GECCO | 3 |
| 2010 | Composition bias and the origin of ORFan genesabstractMOTIVATION: Intriguingly, sequence analysis of genomes reveals that a large number of genes are unique to each organism. The origin of these genes, termed ORFans, is not known. Here, we explore the origin of ORFan genes by defining a simple measure called 'composition bias', based on the deviation of the amino acid composition of a given sequence from the average composition of all proteins of a given genome. RESULTS: For a set of 47 prokaryotic genomes, we show that the amino acid composition bias of real proteins, random 'proteins' (created by using the nucleotide frequencies of each genome) and 'proteins' translated from intergenic regions are distinct. For ORFans, we observed a correlation between their composition bias and their relative evolutionary age. Recent ORFan proteins have compositions more similar to those of random 'proteins', while the compositions of more ancient ORFan proteins are more similar to those of the set of all proteins of the organism. This observation is consistent with an evolutionary scenario wherein ORFan genes emerged and underwent a large number of random mutations and selection, eventually adapting to the composition preference of their organism over time. Inbal Yomtovian, Nuttinee Teerakulkittipong, Byungkook Lee, John Moult, Ron Unger |
Bioinform. | 5 |
| 2009 | A conflict based SAW method for Constraint Satisfaction ProblemsabstractEvolutionary algorithms have employed the SAW (stepwise adaptation of weights) method in order to solve CSPs (constraint satisfaction problems). This method originated in hill-climbing algorithms used to solve instances of 3-SAT by adapting a weight for each clause. Originally, adaptation of weights for solving CSPs was done by assigning a weight for each variable or each constraint. Here we investigate a SAW method which assigns a weight for each conflict. Two simple stochastic CSP solvers are presented. For both we show that constraint based SAW and conflict based SAW perform equally on easy CSP samples, but the conflict based SAW outperforms the constraint based SAW when applied to hard CSPs. Moreover, the best of the two suggested algorithms in its conflict based SAW version performs better than the best known evolutionary algorithm for CSPs that uses weight adaptation, and even better than the best known evolutionary algorithm for CSPs in general. Rafi Shalom, Mireille Avigal, Ron Unger |
IEEE Congress on Evolutionary Computation | 3 |
| 2009 | RNAslider: a faster engine for consecutive windows folding and its application to the analysis of genomic folding asymmetryabstractBACKGROUND: Scanning large genomes with a sliding window in search of locally stable RNA structures is a well motivated problem in bioinformatics. Given a predefined window size L and an RNA sequence S of size N (L < N), the consecutive windows folding problem is to compute the minimal free energy (MFE) for the folding of each of the L-sized substrings of S. The consecutive windows folding problem can be naively solved in O(NL3) by applying any of the classical cubic-time RNA folding algorithms to each of the N-L windows of size L. Recently an O(NL2) solution for this problem has been described. RESULTS: Here, we describe and implement an O(NLpsi(L)) engine for the consecutive windows folding problem, where psi(L) is shown to converge to O(1) under the assumption of a standard probabilistic polymer folding model, yielding an O(L) speedup which is experimentally confirmed. Using this tool, we note an intriguing directionality (5'-3' vs. 3'-5') folding bias, i.e. that the minimal free energy (MFE) of folding is higher in the native direction of the DNA than in the reverse direction of various genomic regions in several organisms including regions of the genomes that do not encode proteins or ncRNA. This bias largely emerges from the genomic dinucleotide bias which affects the MFE, however we see some variations in the folding bias in the different genomic regions when normalized to the dinucleotide bias. We also present results from calculating the MFE landscape of a mouse chromosome 1, characterizing the MFE of the long ncRNA molecules that reside in this chromosome. CONCLUSION: The efficient consecutive windows folding engine described in this paper allows for genome wide scans for ncRNA molecules as well as large-scale statistics. This is implemented here as a software tool, called RNAslider, and applied to the scanning of long chromosomes, leading to the observation of features that are visible only on a large scale. Yair Horesh, Ydo Wexler, Ilana Lebenthal, Michal Ziv-Ukelson, Ron Unger |
BMC Bioinform. | 5 |
| 2009 | Trade-off between Positive and Negative Design of Protein Stability: From Lattice Models to Real ProteinsabstractTwo different strategies for stabilizing proteins are (i) positive design in which the native state is stabilized and (ii) negative design in which competing non-native conformations are destabilized. Here, the circumstances under which one strategy might be favored over the other are explored in the case of lattice models of proteins and then generalized and discussed with regard to real proteins. The balance between positive and negative design of proteins is found to be determined by their average "contact-frequency", a property that corresponds to the fraction of states in the conformational ensemble of the sequence in which a pair of residues is in contact. Lattice model proteins with a high average contact-frequency are found to use negative design more than model proteins with a low average contact-frequency. A mathematical derivation of this result indicates that it is general and likely to hold also for real proteins. Comparison of the results of correlated mutation analysis for real proteins with typical contact-frequencies to those of proteins likely to have high contact-frequencies (such as disordered proteins and proteins that are dependent on chaperonins for their folding) indicates that the latter tend to have stronger interactions between residues that are not in contact in their native conformation. Hence, our work indicates that negative design is employed when insufficient stabilization is achieved via positive design owing to high contact-frequencies. Orly Noivirt-Brik, Amnon Horovitz, Ron Unger |
PLoS Comput. Biol. | 3 |
| 2008 | Psiscan: a computational approach to identify H/ACA-like and AGA-like non-coding RNA in trypanosomatid genomesabstractBACKGROUND: Detection of non coding RNA (ncRNA) molecules is a major bioinformatics challenge. This challenge is particularly difficult when attempting to detect H/ACA molecules which are involved in converting uridine to pseudouridine on rRNA in trypanosomes, because these organisms have unique H/ACA molecules (termed H/ACA-like) that lack several of the features that characterize H/ACA molecules in most other organisms. RESULTS: We present here a computational tool called Psiscan, which was designed to detect H/ACA-like molecules in trypanosomes. We started by analyzing known H/ACA-like molecules and characterized their crucial elements both computationally and experimentally. Next, we set up constraints based on this analysis and additional phylogenic and functional data to rapidly scan three trypanosome genomes (T. brucei, T. cruzi and L. major) for sequences that observe these constraints and are conserved among the species. In the next step, we used minimal energy calculation to select the molecules that are predicted to fold into a lowest energy structure that is consistent with the constraints. In the final computational step, we used a Support Vector Machine that was trained on known H/ACA-like molecules as positive examples and on negative examples of molecules that were identified by the computational analyses but were shown experimentally not to be H/ACA-like molecules. The leading candidate molecules predicted by the SVM model were then subjected to experimental validation. CONCLUSION: The experimental validation showed 11 molecules to be expressed (4 out of 25 in the intermediate stage and 7 out of 19 in the final validation after the machine learning stage). Five of these 11 molecules were further shown to be bona fide H/ACA-like molecules. As snoRNA in trypanosomes are organized in clusters, the new H/ACA-like molecules could be used as starting points to manually search for additional molecules in their neighbourhood. All together this study increased our repertoire by fourteen H/ACA-like and six C/D snoRNAs molecules from T. brucei and L. Major. In addition the experimental analysis revealed that six ncRNA molecules that are expressed are not downregulated in CBF5 silenced cells, suggesting that they have structural features of H/ACA-like molecules but do not have their standard function. We termed this novel class of molecules AGA-like, and we are exploring their function. This study demonstrates the power of tight collaboration between computational and experimental approaches in a combined effort to reveal the repertoire of ncRNA molecles. Inna Myslyuk, Tirza Doniger, Yair Horesh, Avraham Hury, Ran Hoffer, Yaara Ziporen, Shulamit Michaeli, Ron Unger |
BMC Bioinform. | 8 |
| 2007 | A tale of two tails: why are terminal residues of proteins exposed?abstractMOTIVATION: It is widely known that terminal residues of proteins (i.e. the N- and C-termini) are predominantly located on the surface of proteins and exposed to the solvent. However, there is no good explanation as to the forces driving this phenomenon. The common explanation that terminal residues are charged, and charged residues prefer to be on the surface, cannot explain the magnitude of the phenomenon. Here, we survey a large number of proteins from the PDB in order to explore, quantitatively, this phenomenon, and then we use a lattice model to study the mechanisms involved. RESULTS: The location of the termini was examined for 425 small monomeric proteins (50-200 amino acids) and it was found that the average solvent accessibility of termini residues is 87.1% compared with 49.2% of charged residues and 35.9% of all residues. Using a cutoff of 50% of the maximal possible exposure, 80.3% of the N-terminal and 86.1% of the C-terminal residues are exposed compared to 32% for all residues. In addition, terminal residues are much more distant from the center of mass of their proteins than other residues. Using a 2D lattice, a large population of model proteins was studied on three levels: structural selection of compact structures, thermodynamic selection of conformations with a pronounced energy gap and kinetic selection of fast folding proteins using Monte-Carlo simulations. Progressively, each selection raises the proportion of proteins with termini on the surface, resulting in similar proportions to those observed for real proteins. Etai Jacob, Ron Unger |
Bioinform. | 2 |
| 2007 | Low folding propensity and high translation efficiency distinguish in vivo substrates of GroEL from other Escherichia coli proteinsabstractMOTIVATION: Theoretical considerations have indicated that the amount of chaperonin GroEL in Escherichia coli cells is sufficient to fold only approximately 2-5% of newly synthesized proteins under normal physiological conditions, thereby suggesting that only a subset of E.coli proteins fold in vivo in a GroEL-dependent manner. Recently, members of this subset were identified in two independent studies that resulted in two partially overlapping lists of GroEL-interacting proteins. The objective of the work described here was to identify sequence-based features of GroEL-interacting proteins that distinguish them from other E.coli proteins and that may account for their dependence on the chaperonin system. RESULTS: Our analysis shows that GroEL-interacting proteins have, on average, low folding propensities and high translation efficiencies. These two properties in combination can increase the risk of aggregation of these proteins and, thus, cause their folding to be chaperonin-dependent. Strikingly, we find that these properties are absent in proteins homologous to the E.coli GroEL-interacting proteins in Ureaplasma urealyticum, an organism that lacks a chaperonin system, thereby confirming our conclusions. Orly Noivirt-Brik, Ron Unger, Amnon Horovitz |
Bioinform. | 2 |
| 2007 | RNAspa: a shortest path approach for comparative prediction of the secondary structure of ncRNA moleculesabstractBACKGROUND: In recent years, RNA molecules that are not translated into proteins (ncRNAs) have drawn a great deal of attention, as they were shown to be involved in many cellular functions. One of the most important computational problems regarding ncRNA is to predict the secondary structure of a molecule from its sequence. In particular, we attempted to predict the secondary structure for a set of unaligned ncRNA molecules that are taken from the same family, and thus presumably have a similar structure. RESULTS: We developed the RNAspa program, which comparatively predicts the secondary structure for a set of ncRNA molecules in linear time in the number of molecules. We observed that in a list of several hundred suboptimal minimal free energy (MFE) predictions, as provided by the RNAsubopt program of the Vienna package, it is likely that at least one suggested structure would be similar to the true, correct one. The suboptimal solutions of each molecule are represented as a layer of vertices in a graph. The shortest path in this graph is the basis for structural predictions for the molecule. We also show that RNA secondary structures can be compared very rapidly by a simple string Edit-Distance algorithm with a minimal loss of accuracy. We show that this approach allows us to more deeply explore the suboptimal structure space. CONCLUSION: The algorithm was tested on three datasets which include several ncRNA families taken from the Rfam database. These datasets allowed for comparison of the algorithm with other methods. In these tests, RNAspa performed better than four other programs. Yair Horesh, Tirza Doniger, Shulamit Michaeli, Ron Unger |
BMC Bioinform. | 4 |
| 2006 | Evolving High-Performance Evolutionary Computations for Space Vehicle DesignabstractThe nuclear electric vehicle optimization toolset (NEVOT) optimizes the design of all major nuclear electric propulsion (NEP) vehicle subsystems for a defined mission within constraints and optimization parameters chosen by a user. The tool currently uses a number of evolutionary computations (ECs) for designing NEP vehicles. Since evaluating candidate vehicle designs is computationally expensive, it is important that a set of robust control parameters be discovered. In order to accomplish this, a meta-genetic algorithm (meta-GA) was developed to discover control parameters for generational, steady-state, and steady-generational GAs as well as for particle swarm optimizers (PSOs) with ring, star, and random topologies. Our results show that the high-performance GAs are more efficient than the high-performance PSOs on a NASA asteroid mission problem. Gerry V. Dozier, Winard Britt, Michael P. SanSoucie, Patrick V. Hull, Michael L. Tinker, Ron Unger, Steve Bancroft, Trevor Moeller, Dan Rooney |
IEEE Congress on Evolutionary Computation | 6 |
| 2004 | The Made-In-Israel Bioinformatics PortalabstractYossi Rosenberg, Ron Unger; The Made-In-Israel Bioinformatics Portal, Briefings in Bioinformatics, Volume 5, Issue 4, 1 December 2004, Pages 389–390, https://do Yossi Rosenberg, Ron Unger |
Briefings Bioinform. | 2 |
| 1999 | A simple algorithm for detecting circular permutations in proteinsabstractMOTIVATION: Circular permutation of a protein is a genetic operation in which part of the C-terminal of the protein is moved to its N-terminal. Recently, it has been shown that proteins that undergo engineered circular permutations generally maintain their three dimensional structure and biological function. This observation raises the possibility that circular permutation has occurred in Nature during evolution. In this scenario a protein underwent circular permutation into another protein, thereafter both proteins further diverged by standard genetic operations. To study this possibility one needs an efficient algorithm that for a given pair of proteins can detect the underlying event of circular permutations. A possible formal description of the question is: given two sequences, find a circular permutation of one of them under which the edit distance between the proteins is minimal. A naive algorithm might take time proportional to N3 or even N4, which is prohibitively slow for a large-scale survey. A sophisticated algorithm that runs in asymptotic time of N2 was recently suggested, but it is not practical for a large-scale survey. RESULTS: A simple and efficient algorithm that runs in time N2 is presented. The algorithm is based on duplicating one of the two sequences, and then performing a modified version of the standard dynamic programming algorithm. While the algorithm is not guaranteed to find the optimal results, we present data that indicate that in practice the algorithm performs very well. AVAILABILITY: A Fortran program that calculates the optimal edit distance under circular permutation is available upon request from the authors. CONTACT: [email protected]. S. Uliel, A. Fliess, Amihood Amir, Ron Unger |
Bioinform. | 4 |
| 1998 | Genetic Algorithms for Protein Threading
Jacqueline Yadgari, Amihood Amir, Ron Unger |
ISMB | 3 |
| 1996 | Shuffling Biological Sequences
Denise B. Kandel, Yossi Matias, Ron Unger, Peter Winkler 0001 |
Discret. Appl. Math. | 3 |
| 1986 | DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysisabstractWe present a fast algorithm to produce a graphic matrix representation of sequence homology. The algorithm is based on lexicographical ordering of fragments. It preserves most of the options of a simple naive algorithm with a significant increase in speed. This algorithm was the bais for a program, called DNAMAT, that has been extensively tested during the last three years at the Weizmann Institute of Science and has proven to be very useful. In addition we suggest a way to extend our approach to analyse a series of related DNA or RNA sequences, in order to determine certain common structural features. The analysis is done by 'summing' a set of dot-matrices to produce an overall matrix that displays structural elements common to most of the sequences. We give an example of this procedure by analysing tRNA sequences. Ron Unger, David Harel, Joel L. Sussman |
Comput. Appl. Biosci. | 1 |