Ron Unger

dblp:75/3917 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0003-4153-3922ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
12 papers
Bioinformatics and computational biology · 100% Computational science and engineering · 0%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
comparative genomics
0.812024
ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024
Bioinformatics and computational biology › molecular evolution
evolutionary conservation analysis
0.812024
ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024
Bioinformatics and computational biology › gene regulation
gene regulation analysis
0.812024
ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.812024
ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements · Bioinform. 2024
Bioinformatics and computational biology
protein sequence analysis
0.732022
DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts · Bioinform. 2022
Low folding propensity and high translation efficiency distinguish in vivo substrates of GroEL from other Escherichia coli proteins · Bioinform. 2007
A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.612022
DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts · Bioinform. 2022
Bioinformatics and computational biology › protein function prediction › protein classification
protein sequence classification
0.612022
DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts · Bioinform. 2022
Bioinformatics and computational biology › structural bioinformatics
protein structure
0.212016
e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures · Bioinform. 2016
Bioinformatics and computational biology › transcriptomics
RNA editing
0.212016
e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures · Bioinform. 2016
Bioinformatics and computational biology
protein structure analysis
0.222013
Non-local residue-residue contacts in proteins are more conserved than local ones · Bioinform. 2013
A tale of two tails: why are terminal residues of proteins exposed? · Bioinform. 2007
Bioinformatics and computational biology
protein structure prediction
0.122011
Hidden conformations in protein structures · Bioinform. 2011
Genetic Algorithms for Protein Threading · ISMB 1998
Bioinformatics and computational biology › systems biology
gene regulatory network modeling
0.112012
Static network structure can be used to model the phenotypic effects of perturbations in regulatory networks · Bioinform. 2012
Bioinformatics and computational biology › protein structure prediction
protein folding
0.122007
Low folding propensity and high translation efficiency distinguish in vivo substrates of GroEL from other Escherichia coli proteins · Bioinform. 2007
A tale of two tails: why are terminal residues of proteins exposed? · Bioinform. 2007
Bioinformatics and computational biology › protein structure prediction
residue contact prediction
0.112011
Hidden conformations in protein structures · Bioinform. 2011
Bioinformatics and computational biology
genomics
0.112010
Composition bias and the origin of ORFan genes · Bioinform. 2010
Bioinformatics and computational biology › genome annotation
genomic variant annotation
0.112016
e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures · Bioinform. 2016
Bioinformatics and computational biology › systems bioinformatics
pathway analysis
0.012012
Static network structure can be used to model the phenotypic effects of perturbations in regulatory networks · Bioinform. 2012
Bioinformatics and computational biology › protein structure analysis › structural alignment
circular permutation detection
0.011999
A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition
0.011998
Genetic Algorithms for Protein Threading · ISMB 1998
Algorithms and data structures › sequence algorithms › string algorithms
edit distance
0.011999
A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999
Algorithms and data structures › sequence algorithms › string algorithms
sequence alignment
0.011999
A simple algorithm for detecting circular permutations in proteins · Bioinform. 1999
Bioinformatics and computational biology
sequence analysis
0.011986
DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis · Comput. Appl. Biosci. 1986
Bioinformatics and computational biology › sequence analysis
sequence homology
0.011986
DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis · Comput. Appl. Biosci. 1986
Computational science and engineering
structural analysis
0.011986
DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis · Comput. Appl. Biosci. 1986

Methods — techniques the papers use, named apart from their topics

motif scanning · 0.8TSS dataset quality assessment · 0.8transfer learning · 0.6protein language model · 0.6knowledge distillation · 0.6structure mapping · 0.2recoding event modeling · 0.2sequence analysis · 0.2structural comparison · 0.2static network structure analysis · 0.1dynamic programming · 0.0
YearPublicationVenuePosition
2024 ElemeNT 2023: an enhanced tool for detection and curation of core promoter elements
abstract
MOTIVATION: Prediction and identification of core promoter elements and transcription factor binding sites is essential for understanding the mechanism of transcription initiation and deciphering the biological activity of a specific locus. Thus, there is a need for an up-to-date tool to detect and curate core promoter elements/motifs in any provided nucleotide sequences. RESULTS: Here, we introduce ElemeNT 2023-a new and enhanced version of the Elements Navigation Tool, which provides novel capabilities for assessing evolutionary conservation and for readily evaluating the quality of high-throughput transcription start site (TSS) datasets, leveraging preferential motif positioning. ElemeNT 2023 is accessible both as a fast web-based tool and via command line (no coding skills are required to run the tool). While this tool is focused on core promoter elements, it can also be used for searching any user-defined motif, including sequence-specific DNA binding sites. Furthermore, ElemeNT's CORE database, which contains predicted core promoter elements around annotated TSSs, is now expanded to cover 10 species, ranging from worms to human. In this applications note, we describe the new workflow and demonstrate a case study using ElemeNT 2023 for core promoter composition analysis of diverse species, revealing motif prevalence and highlighting evolutionary insights. We discuss how this tool facilitates the exploration of uncharted transcriptomic data, appraises TSS quality, and aids in designing synthetic promoters for gene expression optimization. Taken together, ElemeNT 2023 empowers researchers with comprehensive tools for meticulous analysis of sequence elements and gene expression strategies. AVAILABILITY AND IMPLEMENTATION: ElemeNT 2023 is freely available at https://www.juven-gershonlab.org/resources/element-v2023/. The source code and command line version of ElemeNT 2023 are available at https://github.com/OritAdato/ElemeNT. No coding skills are required to run the tool.
Orit Adato, Anna Sloutskin, Hodaya Komemi, Ian Brabb, Sascha Duttke, Philipp Bucher, Ron Unger, Tamar Juven-Gershon
Bioinform.7
2022 DistilProtBert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts
abstract
SUMMARY: Recently, deep learning models, initially developed in the field of natural language processing (NLP), were applied successfully to analyze protein sequences. A major drawback of these models is their size in terms of the number of parameters needed to be fitted and the amount of computational resources they require. Recently, 'distilled' models using the concept of student and teacher networks have been widely used in NLP. Here, we adapted this concept to the problem of protein sequence analysis, by developing DistilProtBert, a distilled version of the successful ProtBert model. Implementing this approach, we reduced the size of the network and the running time by 50%, and the computational resources needed for pretraining by 98% relative to ProtBert model. Using two published tasks, we showed that the performance of the distilled model approaches that of the full model. We next tested the ability of DistilProtBert to distinguish between real and random protein sequences. The task is highly challenging if the composition is maintained on the level of singlet, doublet and triplet amino acids. Indeed, traditional machine-learning algorithms have difficulties with this task. Here, we show that DistilProtBert preforms very well on singlet, doublet and even triplet-shuffled versions of the human proteome, with AUC of 0.92, 0.91 and 0.87, respectively. Finally, we suggest that by examining the small number of false-positive classifications (i.e. shuffled sequences classified as proteins by DistilProtBert), we may be able to identify de novo potential natural-like proteins based on random shuffling of amino acid sequences. AVAILABILITY AND IMPLEMENTATION: https://github.com/yarongef/DistilProtBert.
Yaron Geffen, Yanay Ofran, Ron Unger
Bioinform.3
2016 e23D: database and visualization of A-to-I RNA editing sites mapped to 3D protein structures
abstract
UNLABELLED: e23D, a database of A-to-I RNA editing sites from human, mouse and fly mapped to evolutionary related protein 3D structures, is presented. Genomic coordinates of A-to-I RNA editing sites are converted to protein coordinates and mapped onto 3D structures from PDB or theoretical models from ModBase. e23D allows visualization of the protein structure, modeling of recoding events and orientation of the editing with respect to nearby genomic functional sites from databases of disease causing mutations and genomic polymorphism. AVAILABILITY AND IMPLEMENTATION: http://www.sheba-cancer.org.il/e23D CONTACT: [email protected] or [email protected].
Oz Solomon, Eran Eyal, Ninette Amariglio, Ron Unger, Gideon Rechavi
Bioinform.4
2013 Non-local residue-residue contacts in proteins are more conserved than local ones
abstract
Non-covalent residue-residue contacts drive the folding of proteins and stabilize them. They may be local-i.e. involve residues that are close in sequence, or non-local. It has been suggested that, in most proteins, local contacts drive protein folding by providing crucial constraints of the conformational space, thus allowing proteins to fold. We compared residues that are involved in local contacts to residues that are involved in non-local contacts and found that, in most proteins, residues in non-local contacts are significantly more conserved evolutionarily than residues in local contacts. Moreover, non-local contacts are more structurally conserved: a contact between positions that are distant in sequence is more likely to exist in many structural homologues compared with a contact between positions that are close in sequence. These results provide new insights into the mechanisms of protein folding and may allow for better prediction of critical intra-chain contacts.
Orly Noivirt-Brik, Gershon Hazan, Ron Unger, Yanay Ofran
Bioinform.3
2012 Static network structure can be used to model the phenotypic effects of perturbations in regulatory networks
abstract
MOTIVATION: Biological processes are dynamic, whereas the networks that depict them are typically static. Quantitative modeling using differential equations or logic-based functions can offer quantitative predictions of the behavior of biological systems, but they require detailed experimental characterization of interaction kinetics, which is typically unavailable. To determine to what extent complex biological processes can be modeled and analyzed using only the static structure of the network (i.e. the direction and sign of the edges), we attempt to predict the phenotypic effect of perturbations in biological networks from the static network structure. RESULTS: We analyzed three networks from different sources: The EGFR/MAPK and PI3K/AKT network from a detailed experimental study, the TNF regulatory network from the STRING database and a large network of all NCI-curated pathways from the Protein Interaction Database. Altogether, we predicted the effect of 39 perturbations (e.g. by one or two drugs) on 433 target proteins/genes. In up to 82% of the cases, an algorithm that used only the static structure of the network correctly predicted whether any given protein/gene is upregulated or downregulated as a result of perturbations of other proteins/genes. CONCLUSION: While quantitative modeling requires detailed experimental data and heavy computations, which limit its scalability for large networks, a wiring-based approach can use available data from pathway and interaction databases and may be scalable. These results lay the foundations for a large-scale approach of predicting phenotypes based on the schematic structure of networks.
Ariel Feiglin, Adar Hacohen, Avital Sarusi, Jasmin Fisher, Ron Unger, Yanay Ofran
Bioinform.5
2011 Hidden conformations in protein structures
abstract
MOTIVATION: Prediction of interactions between protein residues (contact map prediction) can facilitate various aspects of 3D structure modeling. However, the accuracy of ab initio contact prediction is still limited. As structural genomics initiatives move ahead, solved structures of homologous proteins can be used as multiple templates to improve contact prediction of the major conformation of an unsolved target protein. Furthermore, multiple templates may provide a wider view of the protein's conformational space. However, successful usage of multiple structural templates is not straightforward, due to their variable relevance to the target protein, and because of data redundancy issues. RESULTS: We present here an algorithm that addresses these two limitations in the use of multiple structure templates. First, the algorithm unites contact maps extracted from templates sharing high sequence similarity with each other in a fashion that acknowledges the possibility of multiple conformations. Next, it weights the resulting united maps in inverse proportion to their evolutionary distance from the target protein. Testing this algorithm against CASP8 targets resulted in high precision contact maps. Remarkably, based solely on structural data of remote homologues, our algorithm identified residue-residue interactions that account for all the known conformations of calmodulin, a multifaceted protein. Therefore, employing multiple templates, which improves prediction of contact maps, can also be used to reveal novel conformations. As multiple templates will soon be available for most proteins, our scheme suggests an effective procedure for their optimal consideration. AVAILABILITY: A Perl script implementing the WMC algorithm described in this article is freely available for academic use at http://tau.ac.il/~haimash/WMC.
Haim Ashkenazy, Ron Unger, Yossef Kliger
Bioinform.2
2010 Analysis of the effects of lifetime learning on population fitness using vose model
abstract
Vose's dynamical systems model of the simple genetic algorithm (SGA) is an exact model that uses mathematical operations to capture the dynamical behavior of genetic algorithms. The original model was defined for a simple genetic algorithm. This paper suggests how to extend the model and incorporate two kinds of learning, Darwinian and Lamarckian, into the framework of the Vose model. The extension provides a new theoretical framework to examine the effects of lifetime learning on the fitness of a population. We analyze the asymptotic behavior of different hybrid algorithms on an infinite population vector and compare it to the behavior of the classical genetic algorithm on various population sizes. Our experiments show that Lamarckian-like inheritance - direct transfer of lifetime learning results to offsprings - allows quicker genetic adaptation. However, functions exist where the simple genetic algorithms without learning, as well as Lamarckian evolution, converge to the same local optimum, while genetic search based on Darwinian inheritance converges to the global optimum.
Roi Yehoshua, Mireille Avigal, Ron Unger
GECCO3
2010 Composition bias and the origin of ORFan genes
abstract
MOTIVATION: Intriguingly, sequence analysis of genomes reveals that a large number of genes are unique to each organism. The origin of these genes, termed ORFans, is not known. Here, we explore the origin of ORFan genes by defining a simple measure called 'composition bias', based on the deviation of the amino acid composition of a given sequence from the average composition of all proteins of a given genome. RESULTS: For a set of 47 prokaryotic genomes, we show that the amino acid composition bias of real proteins, random 'proteins' (created by using the nucleotide frequencies of each genome) and 'proteins' translated from intergenic regions are distinct. For ORFans, we observed a correlation between their composition bias and their relative evolutionary age. Recent ORFan proteins have compositions more similar to those of random 'proteins', while the compositions of more ancient ORFan proteins are more similar to those of the set of all proteins of the organism. This observation is consistent with an evolutionary scenario wherein ORFan genes emerged and underwent a large number of random mutations and selection, eventually adapting to the composition preference of their organism over time.
Inbal Yomtovian, Nuttinee Teerakulkittipong, Byungkook Lee, John Moult, Ron Unger
Bioinform.5
2009 A conflict based SAW method for Constraint Satisfaction Problems
abstract
Evolutionary algorithms have employed the SAW (stepwise adaptation of weights) method in order to solve CSPs (constraint satisfaction problems). This method originated in hill-climbing algorithms used to solve instances of 3-SAT by adapting a weight for each clause. Originally, adaptation of weights for solving CSPs was done by assigning a weight for each variable or each constraint. Here we investigate a SAW method which assigns a weight for each conflict. Two simple stochastic CSP solvers are presented. For both we show that constraint based SAW and conflict based SAW perform equally on easy CSP samples, but the conflict based SAW outperforms the constraint based SAW when applied to hard CSPs. Moreover, the best of the two suggested algorithms in its conflict based SAW version performs better than the best known evolutionary algorithm for CSPs that uses weight adaptation, and even better than the best known evolutionary algorithm for CSPs in general.
Rafi Shalom, Mireille Avigal, Ron Unger
IEEE Congress on Evolutionary Computation3
2009 RNAslider: a faster engine for consecutive windows folding and its application to the analysis of genomic folding asymmetry
abstract
BACKGROUND: Scanning large genomes with a sliding window in search of locally stable RNA structures is a well motivated problem in bioinformatics. Given a predefined window size L and an RNA sequence S of size N (L < N), the consecutive windows folding problem is to compute the minimal free energy (MFE) for the folding of each of the L-sized substrings of S. The consecutive windows folding problem can be naively solved in O(NL3) by applying any of the classical cubic-time RNA folding algorithms to each of the N-L windows of size L. Recently an O(NL2) solution for this problem has been described. RESULTS: Here, we describe and implement an O(NLpsi(L)) engine for the consecutive windows folding problem, where psi(L) is shown to converge to O(1) under the assumption of a standard probabilistic polymer folding model, yielding an O(L) speedup which is experimentally confirmed. Using this tool, we note an intriguing directionality (5'-3' vs. 3'-5') folding bias, i.e. that the minimal free energy (MFE) of folding is higher in the native direction of the DNA than in the reverse direction of various genomic regions in several organisms including regions of the genomes that do not encode proteins or ncRNA. This bias largely emerges from the genomic dinucleotide bias which affects the MFE, however we see some variations in the folding bias in the different genomic regions when normalized to the dinucleotide bias. We also present results from calculating the MFE landscape of a mouse chromosome 1, characterizing the MFE of the long ncRNA molecules that reside in this chromosome. CONCLUSION: The efficient consecutive windows folding engine described in this paper allows for genome wide scans for ncRNA molecules as well as large-scale statistics. This is implemented here as a software tool, called RNAslider, and applied to the scanning of long chromosomes, leading to the observation of features that are visible only on a large scale.
Yair Horesh, Ydo Wexler, Ilana Lebenthal, Michal Ziv-Ukelson, Ron Unger
BMC Bioinform.5
2009 Trade-off between Positive and Negative Design of Protein Stability: From Lattice Models to Real Proteins
abstract
Two different strategies for stabilizing proteins are (i) positive design in which the native state is stabilized and (ii) negative design in which competing non-native conformations are destabilized. Here, the circumstances under which one strategy might be favored over the other are explored in the case of lattice models of proteins and then generalized and discussed with regard to real proteins. The balance between positive and negative design of proteins is found to be determined by their average "contact-frequency", a property that corresponds to the fraction of states in the conformational ensemble of the sequence in which a pair of residues is in contact. Lattice model proteins with a high average contact-frequency are found to use negative design more than model proteins with a low average contact-frequency. A mathematical derivation of this result indicates that it is general and likely to hold also for real proteins. Comparison of the results of correlated mutation analysis for real proteins with typical contact-frequencies to those of proteins likely to have high contact-frequencies (such as disordered proteins and proteins that are dependent on chaperonins for their folding) indicates that the latter tend to have stronger interactions between residues that are not in contact in their native conformation. Hence, our work indicates that negative design is employed when insufficient stabilization is achieved via positive design owing to high contact-frequencies.
Orly Noivirt-Brik, Amnon Horovitz, Ron Unger
PLoS Comput. Biol.3
2008 Psiscan: a computational approach to identify H/ACA-like and AGA-like non-coding RNA in trypanosomatid genomes
abstract
BACKGROUND: Detection of non coding RNA (ncRNA) molecules is a major bioinformatics challenge. This challenge is particularly difficult when attempting to detect H/ACA molecules which are involved in converting uridine to pseudouridine on rRNA in trypanosomes, because these organisms have unique H/ACA molecules (termed H/ACA-like) that lack several of the features that characterize H/ACA molecules in most other organisms. RESULTS: We present here a computational tool called Psiscan, which was designed to detect H/ACA-like molecules in trypanosomes. We started by analyzing known H/ACA-like molecules and characterized their crucial elements both computationally and experimentally. Next, we set up constraints based on this analysis and additional phylogenic and functional data to rapidly scan three trypanosome genomes (T. brucei, T. cruzi and L. major) for sequences that observe these constraints and are conserved among the species. In the next step, we used minimal energy calculation to select the molecules that are predicted to fold into a lowest energy structure that is consistent with the constraints. In the final computational step, we used a Support Vector Machine that was trained on known H/ACA-like molecules as positive examples and on negative examples of molecules that were identified by the computational analyses but were shown experimentally not to be H/ACA-like molecules. The leading candidate molecules predicted by the SVM model were then subjected to experimental validation. CONCLUSION: The experimental validation showed 11 molecules to be expressed (4 out of 25 in the intermediate stage and 7 out of 19 in the final validation after the machine learning stage). Five of these 11 molecules were further shown to be bona fide H/ACA-like molecules. As snoRNA in trypanosomes are organized in clusters, the new H/ACA-like molecules could be used as starting points to manually search for additional molecules in their neighbourhood. All together this study increased our repertoire by fourteen H/ACA-like and six C/D snoRNAs molecules from T. brucei and L. Major. In addition the experimental analysis revealed that six ncRNA molecules that are expressed are not downregulated in CBF5 silenced cells, suggesting that they have structural features of H/ACA-like molecules but do not have their standard function. We termed this novel class of molecules AGA-like, and we are exploring their function. This study demonstrates the power of tight collaboration between computational and experimental approaches in a combined effort to reveal the repertoire of ncRNA molecles.
Inna Myslyuk, Tirza Doniger, Yair Horesh, Avraham Hury, Ran Hoffer, Yaara Ziporen, Shulamit Michaeli, Ron Unger
BMC Bioinform.8
2007 A tale of two tails: why are terminal residues of proteins exposed?
abstract
MOTIVATION: It is widely known that terminal residues of proteins (i.e. the N- and C-termini) are predominantly located on the surface of proteins and exposed to the solvent. However, there is no good explanation as to the forces driving this phenomenon. The common explanation that terminal residues are charged, and charged residues prefer to be on the surface, cannot explain the magnitude of the phenomenon. Here, we survey a large number of proteins from the PDB in order to explore, quantitatively, this phenomenon, and then we use a lattice model to study the mechanisms involved. RESULTS: The location of the termini was examined for 425 small monomeric proteins (50-200 amino acids) and it was found that the average solvent accessibility of termini residues is 87.1% compared with 49.2% of charged residues and 35.9% of all residues. Using a cutoff of 50% of the maximal possible exposure, 80.3% of the N-terminal and 86.1% of the C-terminal residues are exposed compared to 32% for all residues. In addition, terminal residues are much more distant from the center of mass of their proteins than other residues. Using a 2D lattice, a large population of model proteins was studied on three levels: structural selection of compact structures, thermodynamic selection of conformations with a pronounced energy gap and kinetic selection of fast folding proteins using Monte-Carlo simulations. Progressively, each selection raises the proportion of proteins with termini on the surface, resulting in similar proportions to those observed for real proteins.
Etai Jacob, Ron Unger
Bioinform.2
2007 Low folding propensity and high translation efficiency distinguish in vivo substrates of GroEL from other Escherichia coli proteins
abstract
MOTIVATION: Theoretical considerations have indicated that the amount of chaperonin GroEL in Escherichia coli cells is sufficient to fold only approximately 2-5% of newly synthesized proteins under normal physiological conditions, thereby suggesting that only a subset of E.coli proteins fold in vivo in a GroEL-dependent manner. Recently, members of this subset were identified in two independent studies that resulted in two partially overlapping lists of GroEL-interacting proteins. The objective of the work described here was to identify sequence-based features of GroEL-interacting proteins that distinguish them from other E.coli proteins and that may account for their dependence on the chaperonin system. RESULTS: Our analysis shows that GroEL-interacting proteins have, on average, low folding propensities and high translation efficiencies. These two properties in combination can increase the risk of aggregation of these proteins and, thus, cause their folding to be chaperonin-dependent. Strikingly, we find that these properties are absent in proteins homologous to the E.coli GroEL-interacting proteins in Ureaplasma urealyticum, an organism that lacks a chaperonin system, thereby confirming our conclusions.
Orly Noivirt-Brik, Ron Unger, Amnon Horovitz
Bioinform.2
2007 RNAspa: a shortest path approach for comparative prediction of the secondary structure of ncRNA molecules
abstract
BACKGROUND: In recent years, RNA molecules that are not translated into proteins (ncRNAs) have drawn a great deal of attention, as they were shown to be involved in many cellular functions. One of the most important computational problems regarding ncRNA is to predict the secondary structure of a molecule from its sequence. In particular, we attempted to predict the secondary structure for a set of unaligned ncRNA molecules that are taken from the same family, and thus presumably have a similar structure. RESULTS: We developed the RNAspa program, which comparatively predicts the secondary structure for a set of ncRNA molecules in linear time in the number of molecules. We observed that in a list of several hundred suboptimal minimal free energy (MFE) predictions, as provided by the RNAsubopt program of the Vienna package, it is likely that at least one suggested structure would be similar to the true, correct one. The suboptimal solutions of each molecule are represented as a layer of vertices in a graph. The shortest path in this graph is the basis for structural predictions for the molecule. We also show that RNA secondary structures can be compared very rapidly by a simple string Edit-Distance algorithm with a minimal loss of accuracy. We show that this approach allows us to more deeply explore the suboptimal structure space. CONCLUSION: The algorithm was tested on three datasets which include several ncRNA families taken from the Rfam database. These datasets allowed for comparison of the algorithm with other methods. In these tests, RNAspa performed better than four other programs.
Yair Horesh, Tirza Doniger, Shulamit Michaeli, Ron Unger
BMC Bioinform.4
2006 Evolving High-Performance Evolutionary Computations for Space Vehicle Design
abstract
The nuclear electric vehicle optimization toolset (NEVOT) optimizes the design of all major nuclear electric propulsion (NEP) vehicle subsystems for a defined mission within constraints and optimization parameters chosen by a user. The tool currently uses a number of evolutionary computations (ECs) for designing NEP vehicles. Since evaluating candidate vehicle designs is computationally expensive, it is important that a set of robust control parameters be discovered. In order to accomplish this, a meta-genetic algorithm (meta-GA) was developed to discover control parameters for generational, steady-state, and steady-generational GAs as well as for particle swarm optimizers (PSOs) with ring, star, and random topologies. Our results show that the high-performance GAs are more efficient than the high-performance PSOs on a NASA asteroid mission problem.
Gerry V. Dozier, Winard Britt, Michael P. SanSoucie, Patrick V. Hull, Michael L. Tinker, Ron Unger, Steve Bancroft, Trevor Moeller, Dan Rooney
IEEE Congress on Evolutionary Computation6
2004 The Made-In-Israel Bioinformatics Portal
abstract
Yossi Rosenberg, Ron Unger; The Made-In-Israel Bioinformatics Portal, Briefings in Bioinformatics, Volume 5, Issue 4, 1 December 2004, Pages 389–390, https://do
Yossi Rosenberg, Ron Unger
Briefings Bioinform.2
1999 A simple algorithm for detecting circular permutations in proteins
abstract
MOTIVATION: Circular permutation of a protein is a genetic operation in which part of the C-terminal of the protein is moved to its N-terminal. Recently, it has been shown that proteins that undergo engineered circular permutations generally maintain their three dimensional structure and biological function. This observation raises the possibility that circular permutation has occurred in Nature during evolution. In this scenario a protein underwent circular permutation into another protein, thereafter both proteins further diverged by standard genetic operations. To study this possibility one needs an efficient algorithm that for a given pair of proteins can detect the underlying event of circular permutations. A possible formal description of the question is: given two sequences, find a circular permutation of one of them under which the edit distance between the proteins is minimal. A naive algorithm might take time proportional to N3 or even N4, which is prohibitively slow for a large-scale survey. A sophisticated algorithm that runs in asymptotic time of N2 was recently suggested, but it is not practical for a large-scale survey. RESULTS: A simple and efficient algorithm that runs in time N2 is presented. The algorithm is based on duplicating one of the two sequences, and then performing a modified version of the standard dynamic programming algorithm. While the algorithm is not guaranteed to find the optimal results, we present data that indicate that in practice the algorithm performs very well. AVAILABILITY: A Fortran program that calculates the optimal edit distance under circular permutation is available upon request from the authors. CONTACT: [email protected].
S. Uliel, A. Fliess, Amihood Amir, Ron Unger
Bioinform.4
1998 Genetic Algorithms for Protein Threading
Jacqueline Yadgari, Amihood Amir, Ron Unger
ISMB3
1996 Shuffling Biological Sequences
Denise B. Kandel, Yossi Matias, Ron Unger, Peter Winkler 0001
Discret. Appl. Math.3
1986 DNAMAT: an efficient graphic matrix sequence homology algorithm and its application to structural analysis
abstract
We present a fast algorithm to produce a graphic matrix representation of sequence homology. The algorithm is based on lexicographical ordering of fragments. It preserves most of the options of a simple naive algorithm with a significant increase in speed. This algorithm was the bais for a program, called DNAMAT, that has been extensively tested during the last three years at the Weizmann Institute of Science and has proven to be very useful. In addition we suggest a way to extend our approach to analyse a series of related DNA or RNA sequences, in order to determine certain common structural features. The analysis is done by 'summing' a set of dot-matrices to produce an overall matrix that displays structural elements common to most of the sequences. We give an example of this procedure by analysing tRNA sequences.
Ron Unger, David Harel, Joel L. Sussman
Comput. Appl. Biosci.1