VLDB 2026 Research / reviewers in the wild / expert
Michal Linial
dblp:13/1697
· DBLP profile ↗
44ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-9357-4526ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 44 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Advancing causal inference in medicine using biobank data
Hadasa Kaufman, Nadav Rappoport, Amir Gilad, Michal Linial |
J. Biomed. Informatics | 4 |
| 2024 | Discovering predisposing genes for hereditary breast cancer using deep learningabstractBreast cancer (BC) is the most common malignancy affecting Western women today. It is estimated that as many as 10% of BC cases can be attributed to germline variants. However, the genetic basis of the majority of familial BC cases has yet to be identified. Discovering predisposing genes contributing to familial BC is challenging due to their presumed rarity, low penetrance, and complex biological mechanisms. Here, we focused on an analysis of rare missense variants in a cohort of 12 families of Middle Eastern origins characterized by a high incidence of BC cases. We devised a novel, high-throughput, variant analysis pipeline adapted for family studies, which aims to analyze variants at the protein level by employing state-of-the-art machine learning models and three-dimensional protein structural analysis. Using our pipeline, we analyzed 1218 rare missense variants that are shared between affected family members and classified 80 genes as candidate pathogenic. Among these genes, we found significant functional enrichment in peroxisomal and mitochondrial biological pathways which segregated across seven families in the study and covered diverse ethnic groups. We present multiple evidence that peroxisomal and mitochondrial pathways play an important, yet underappreciated, role in both germline BC predisposition and BC survival. Gal Passi, Sari Lieberman, Fouad Zahdeh, Omer Murik, Paul Renbaum, Rachel Beeri, Michal Linial, Dalit May, Ephrat Levy-Lahad, Dina Schneidman |
Briefings Bioinform. | 7 |
| 2024 | Automated annotation of disease subtypes
Dan Ofer, Michal Linial |
J. Biomed. Informatics | 2 |
| 2022 | ProteinBERT: a universal deep-learning model of protein sequence and functionabstractSUMMARY: Self-supervised deep language modeling has shown unprecedented success across natural language tasks, and has recently been repurposed to biological sequences. However, existing models and pretraining methods are designed and optimized for text analysis. We introduce ProteinBERT, a deep language model specifically designed for proteins. Our pretraining scheme combines language modeling with a novel task of Gene Ontology (GO) annotation prediction. We introduce novel architectural elements that make the model highly efficient and flexible to long sequences. The architecture of ProteinBERT consists of both local and global representations, allowing end-to-end processing of these types of inputs and outputs. ProteinBERT obtains near state-of-the-art performance, and sometimes exceeds it, on multiple benchmarks covering diverse protein properties (including protein structure, post-translational modifications and biophysical attributes), despite using a far smaller and faster model than competing deep-learning methods. Overall, ProteinBERT provides an efficient framework for rapidly training protein predictors, even with limited labeled data. AVAILABILITY AND IMPLEMENTATION: Code and pretrained model weights are available at https://github.com/nadavbra/protein_bert. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nadav Brandes, Dan Ofer, Yam Peleg, Nadav Rappoport, Michal Linial |
Bioinform. | 5 |
| 2020 | Functional Evolutionary Modeling Exposes Overlooked Protein-Coding Genes Involved in Cancer
Nadav Brandes, Nathan Linial, Michal Linial |
ISBRA | 3 |
| 2020 | PWAS: Proteome-Wide Association Study
Nadav Brandes, Nathan Linial, Michal Linial |
RECOMB | 3 |
| 2020 | BIRD: identifying cell doublets via biallelic expression from single cellsabstractSUMMARY: Current technologies for single-cell transcriptomics allow thousands of cells to be analyzed in a single experiment. The increased scale of these methods raises the risk of cell doublets contamination. Available tools and algorithms for identifying doublets and estimating their occurrence in single-cell experimental data focus on doublets of different species, cell types or individuals. In this study, we analyze transcriptomic data from single cells having an identical genetic background. We claim that the ratio of monoallelic to biallelic expression provides a discriminating power toward doublets' identification. We present a pipeline called BIallelic Ratio for Doublets (BIRD) that relies on heterologous genetic variations, from single-cell RNA sequencing. For each dataset, doublets were artificially created from the actual data and used to train a predictive model. BIRD was applied on Smart-seq data from 163 primary fibroblast single cells. The model achieved 100% accuracy in annotating the randomly simulated doublets. Bonafide doublets were verified based on a biallelic expression signal amongst X-chromosome of female fibroblasts. Data from 10X Genomics microfluidics of human peripheral blood cells achieved in average 83% (±3.7%) accuracy, and an area under the curve of 0.88 (±0.04) for a collection of ∼13 300 single cells. BIRD addresses instances of doublets, which were formed from cell mixtures of identical genetic background and cell identity. Maximal performance is achieved for high-coverage data from Smart-seq. Success in identifying doublets is data specific which varies according to the experimental methodology, genomic diversity between haplotypes, sequence coverage and depth. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kerem Wainer-Katsir, Michal Linial |
Bioinform. | 2 |
| 2019 | A cell-based probabilistic approach unveils the concerted action of miRNAsabstractMature microRNAs (miRNAs) regulate most human genes through direct base-pairing with mRNAs. We investigate the underlying principles of miRNA regulation in living cells. To this end, we overexpressed miRNAs in different cell types and measured the mRNA decay rate under a paradigm of a transcriptional arrest. Based on an exhaustive matrix of mRNA-miRNA binding probabilities, and parameters extracted from our experiments, we developed a computational framework that captures the cooperative action of miRNAs in living cells. The framework, called COMICS, simulates the stochastic binding events between miRNAs and mRNAs in cells. The input of COMICS is cell-specific profiles of mRNAs and miRNAs, and the outcome is the retention level of each mRNA at the end of 100,000 iterations. The results of COMICS from thousands of miRNA manipulations reveal gene sets that exhibit coordinated behavior with respect to all miRNAs (total of 248 families). We identified a small set of genes that are highly responsive to changes in the expression of almost any of the miRNAs. In contrast, about 20% of the tested genes remain insensitive to a broad range of miRNA manipulations. The set of insensitive genes is strongly enriched with genes that belong to the translation machinery. These trends are shared by different cell types. We conclude that the stochastic nature of miRNAs reveals unexpected robustness of gene expression in living cells. By applying a systematic probabilistic approach some key design principles of cell states are revealed, emphasizing in particular, the immunity of the translational machinery vis-a-vis miRNA manipulations across cell types. We propose COMICS as a valuable platform for assessing the outcome of miRNA regulation of cells in health and disease. Shelly Mahlab-Aviv, Nathan Linial, Michal Linial |
PLoS Comput. Biol. | 3 |
| 2015 | Message from the ISCB: ISCB Ebola award for important future research on the computational biology of Ebola virusabstractUNLABELLED: Speed is of the essence in combating Ebola; thus, computational approaches should form a significant component of Ebola research. As for the development of any modern drug, computational biology is uniquely positioned to contribute through comparative analysis of the genome sequences of Ebola strains and three-dimensional protein modeling. Other computational approaches to Ebola may include large-scale docking studies of Ebola proteins with human proteins and with small-molecule libraries, computational modeling of the spread of the virus, computational mining of the Ebola literature and creation of a curated Ebola database. Taken together, such computational efforts could significantly accelerate traditional scientific approaches. In recognition of the need for important and immediate solutions from the field of computational biology against Ebola, the International Society for Computational Biology (ISCB) announces a prize for an important computational advance in fighting the Ebola virus. ISCB will confer the ISCB Fight against Ebola Award, along with a prize of US$2000, at its July 2016 annual meeting (ISCB Intelligent Systems for Molecular Biology 2016, Orlando, FL). CONTACT: [email protected] or [email protected]. Peter D. Karp, Bonnie Berger, Diane E. Kovats, Thomas Lengauer, Michal Linial, Pardis Sabeti, Winston Hide, Burkhard Rost |
Bioinform. | 5 |
| 2015 | ProFET: Feature engineering captures high-level protein functionsabstractMOTIVATION: The amount of sequenced genomes and proteins is growing at an unprecedented pace. Unfortunately, manual curation and functional knowledge lag behind. Homologous inference often fails at labeling proteins with diverse functions and broad classes. Thus, identifying high-level protein functionality remains challenging. We hypothesize that a universal feature engineering approach can yield classification of high-level functions and unified properties when combined with machine learning approaches, without requiring external databases or alignment. RESULTS: In this study, we present a novel bioinformatics toolkit called ProFET (Protein Feature Engineering Toolkit). ProFET extracts hundreds of features covering the elementary biophysical and sequence derived attributes. Most features capture statistically informative patterns. In addition, different representations of sequences and the amino acids alphabet provide a compact, compressed set of features. The results from ProFET were incorporated in data analysis pipelines, implemented in python and adapted for multi-genome scale analysis. ProFET was applied on 17 established and novel protein benchmark datasets involving classification for a variety of binary and multi-class tasks. The results show state of the art performance. The extracted features' show excellent biological interpretability. The success of ProFET applies to a wide range of high-level functions such as subcellular localization, structural classes and proteins with unique functional properties (e.g. neuropeptide precursors, thermophilic and nucleic acid binding). ProFET allows easy, universal discovery of new target proteins, as well as understanding the features underlying different high-level protein functions. AVAILABILITY AND IMPLEMENTATION: ProFET source code and the datasets used are freely available at https://github.com/ddofer/ProFET. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dan Ofer, Michal Linial |
Bioinform. | 2 |
| 2015 | ISCB Ebola Award for Important Future Research on the Computational Biology of Ebola VirusabstractSpeed is of the essence in combating Ebola; thus, computational approaches should form a significant component of Ebola research.As for the development of any modern drug, computational biology is uniquely positioned to contribute through comparative analysis of the genome sequences of Ebola strains as well as 3-D protein modeling.Other computational approaches to Ebola may include large-scale docking studies of Ebola proteins with human proteins and with small-molecule libraries, computational modeling of the spread of the virus, computational mining of the Ebola literature, and creation of a curated Ebola database.Taken together, such computational efforts could significantly accelerate traditional scientific approaches.In recognition of the need for important and immediate solutions from the field of computational biology against Ebola, the International Society for Computational Biology (ISCB) announces a prize for an important computational advance in fighting the Ebola virus.ISCB will confer the ISCB Fight against Ebola Award, along with a prize of US$2,000, at its July 2016 annual meeting (ISCB Intelligent Systems for Molecular Biology [ISMB] 2016, Orlando, Florida). Peter D. Karp, Bonnie Berger, Diane E. Kovats, Thomas Lengauer, Michal Linial, Pardis Sabeti, Winston Hide, Burkhard Rost |
PLoS Comput. Biol. | 5 |
| 2014 | NeuroPID: a predictor for identifying neuropeptide precursors from metazoan proteomesabstractMOTIVATION: The evolution of multicellular organisms is associated with increasing variability of molecules governing behavioral and physiological states. This is often achieved by neuropeptides (NPs) that are produced in neurons from a longer protein, named neuropeptide precursor (NPP). The maturation of NPs occurs through a sequence of proteolytic cleavages. The difficulty in identifying NPPs is a consequence of their diversity and the lack of applicable sequence similarity among the short functionally related NPs. RESULTS: Herein, we describe Neuropeptide Precursor Identifier (NeuroPID), a machine learning scheme that predicts metazoan NPPs. NeuroPID was trained on hundreds of identified NPPs from the UniProtKB database. Some 600 features were extracted from the primary sequences and processed using support vector machines (SVM) and ensemble decision tree classifiers. These features combined biophysical, chemical and informational-statistical properties of NPs and NPPs. Other features were guided by the defining characteristics of the dibasic cleavage sites motif. NeuroPID reached 89-94% accuracy and 90-93% precision in cross-validation blind tests against known NPPs (with an emphasis on Chordata and Arthropoda). NeuroPID also identified NPP-like proteins from extensively studied model organisms as well as from poorly annotated proteomes. We then focused on the most significant sets of features that contribute to the success of the classifiers. We propose that NPPs are attractive targets for investigating and modulating behavior, metabolism and homeostasis and that a rich repertoire of NPs remains to be identified. AVAILABILITY: NeuroPID source code is freely available at http://www.protonet.cs.huji.ac.il/neuropid Dan Ofer, Michal Linial |
Bioinform. | 2 |
| 2014 | Entropy-driven partitioning of the hierarchical protein spaceabstractMOTIVATION: Modern protein sequencing techniques have led to the determination of >50 million protein sequences. ProtoNet is a clustering system that provides a continuous hierarchical agglomerative clustering tree for all proteins. While ProtoNet performs unsupervised classification of all included proteins, finding an optimal level of granularity for the purpose of focusing on protein functional groups remain elusive. Here, we ask whether knowledge-based annotations on protein families can support the automatic unsupervised methods for identifying high-quality protein families. We present a method that yields within the ProtoNet hierarchy an optimal partition of clusters, relative to manual annotation schemes. The method's principle is to minimize the entropy-derived distance between annotation-based partitions and all available hierarchical partitions. We describe the best front (BF) partition of 2 478 328 proteins from UniRef50. Of 4,929,553 ProtoNet tree clusters, BF based on Pfam annotations contain 26,891 clusters. The high quality of the partition is validated by the close correspondence with the set of clusters that best describe thousands of keywords of Pfam. The BF is shown to be superior to naïve cut in the ProtoNet tree that yields a similar number of clusters. Finally, we used parameters intrinsic to the clustering process to enrich a priori the BF's clusters. We present the entropy-based method's benefit in overcoming the unavoidable limitations of nested clusters in ProtoNet. We suggest that this automatic information-based cluster selection can be useful for other large-scale annotation schemes, as well as for systematically testing and comparing putative families derived from alternative clustering methods. AVAILABILITY AND IMPLEMENTATION: A catalog of BF clusters for thousands of Pfam keywords is provided at http://protonet.cs.huji.ac.il/bestFront/. Nadav Rappoport, Amos Stern, Nathan Linial, Michal Linial |
Bioinform. | 4 |
| 2014 | The automated function prediction SIG looks back at 2013 and prepares for 2014abstractAbstract Contact: [email protected] or [email protected] Mark N. Wass, Sean D. Mooney, Michal Linial, Predrag Radivojac, Iddo Friedberg |
Bioinform. | 3 |
| 2014 | Speed Controls in Translating Secretory Proteins in Eukaryotes - an Evolutionary PerspectiveabstractProtein translation is the most expensive operation in dividing cells from bacteria to humans. Therefore, managing the speed and allocation of resources is subject to tight control. From bacteria to humans, clusters of relatively rare tRNA codons at the N'-terminal of mRNAs have been implicated in attenuating the process of ribosome allocation, and consequently the translation rate in a broad range of organisms. The current interpretation of "slow" tRNA codons does not distinguish between protein translations mediated by free- or endoplasmic reticulum (ER)-bound ribosomes. We demonstrate that proteins translated by free- or ER-bound ribosomes exhibit different overall properties in terms of their translation efficiency and speed in yeast, fly, plant, worm, bovine and human. We note that only secreted or membranous proteins with a Signal peptide (SP) are specified by segments of "slow" tRNA at the N'-terminal, followed by abundant codons that are considered "fast." Such profiles apply to 3100 proteins of the human proteome that are composed of secreted and signal peptide (SP)-assisted membranous proteins. Remarkably, the bulks of the proteins (12,000), or membranous proteins lacking SP (3400), do not have such a pattern. Alternation of "fast" and "slow" codons was found also in proteins that translocate to mitochondria through transit peptides (TP). The differential clusters of tRNA adapted codons is not restricted to the N'-terminal of transcripts. Specifically, Glycosylphosphatidylinositol (GPI)-anchored proteins are unified by clusters of low adapted tRNAs codons at the C'-termini. Furthermore, selection of amino acids types and specific codons was shown as the driving force which establishes the translation demands for the secretory proteome. We postulate that "hard-coded" signals within the secretory proteome assist the steps of protein maturation and folding. Specifically, "speed control" signals for delaying the translation of a nascent protein fulfill the co- and post-translational stages such as membrane translocation, proteins processing and folding. Shelly Mahlab-Aviv, Michal Linial |
PLoS Comput. Biol. | 2 |
| 2013 | Functional inference by ProtoNet family tree: the uncharacterized proteome of Daphnia pulexabstractBACKGROUND: Daphnia pulex (Water flea) is the first fully sequenced crustacean genome. The crustaceans and insects have diverged from a common ancestor. It is a model organism for studying the molecular makeup for coping with the environmental challenges. In the complete proteome, there are 30,550 putative proteins. However, about 10,000 of them have no known homologues. Currently, the UniProtoKB reports on 95% of the Daphnia's proteins as putative and uncharacterized proteins. RESULTS: We have applied ProtoNet, an unsupervised hierarchical protein clustering method that covers about 10 million sequences, for automatic annotation of the Daphnia's proteome. 98.7% (26,625) of the Daphnia full-length proteins were successfully mapped to 13,880 ProtoNet stable clusters, and only 1.3% remained unmapped. We compared the properties of the Daphnia's protein families with those of the mouse and the fruitfly proteomes. Functional annotations were successfully assigned for 86% of the proteins. Most proteins (61%) were mapped to only 2953 clusters that contain Daphnia's duplicated genes. We focused on the functionality of maximally amplified paralogs. Cuticle structure components and a variety of ion channels protein families were associated with a maximal level of gene amplification. We focused on gene amplification as a leading strategy of the Daphnia in coping with environmental toxicity. CONCLUSIONS: Automatic inference is achieved through mapping of sequences to the protein family tree of ProtoNet 6.0. Applying a careful inference protocol resulted in functional assignments for over 86% of the complete proteome. We conclude that the scaffold of ProtoNet can be used as an alignment-free protocol for large-scale annotation task of uncharacterized proteomes. Nadav Rappoport, Michal Linial |
BMC Bioinform. | 2 |
| 2012 | Susceptibility of the human pathways graphs to fragmentation by small sets of microRNAsabstractMOTIVATION: MicroRNAs (miRNAs) are short sequences that negatively regulate gene expression. The current understanding of miRNA and their corresponding mRNA targets is primarily based on prediction programs. This study addresses the potential of a coordinated action of miRNAs to manipulate the human pathways. Specifically, we investigate the effectiveness of disrupting the topology of human pathway graphs through a regulation by miRNAs. RESULTS: From a set of miRNA candidates that is associated with a pathway, an exhaustive search for all possible doubles and triplets (coined miR-Duo, miR-Trios) is performed. The impact of each miR-combination on the connectivity of the pathway graph was quantified. About 170 human pathways were tested, and the miR-Duos and miR-Trios were scored for their ability to disrupt these pathway graphs. We show that 75% of all pathways are effectively disconnected by a small number of pathway-specific miR-Trios. Only 15% of the human pathways are resistant to fragmentation by miR-Duos or miR-Trios. Significantly, the combination of the most effective miR-Trios is unique. Thus, a specific regulation of a pathway within the cell is guaranteed. The impact of the selected miR-Duo/Trios on various diseases is discussed. CONCLUSIONS: The methodology presented shows that the synthesis of the topology of a network with a detailed understanding of the miRNAs' regulation is useful in exposing critical nodes of the network. We propose the miR-Trio approach as a basis for rationally designed perturbation experiments. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guy Naamati, Yitzhak Friedman, Ohad Balaga, Michal Linial |
Bioinform. | 4 |
| 2012 | Paving the future: finding suitable ISMB venuesabstractThe International Society for Computational Biology, ISCB, organizes the largest event in the field of computational biology and bioinformatics, namely the annual international conference on Intelligent Systems for Molecular Biology, the ISMB. This year at ISMB 2012 in Long Beach, ISCB celebrated the 20th anniversary of its flagship meeting. ISCB is a young, lean and efficient society that aspires to make a significant impact with only limited resources. Many constraints make the choice of venues for ISMB a tough challenge. Here, we describe those challenges and invite the contribution of ideas for solutions. Burkhard Rost, Terry Gaasterland, Thomas Lengauer, Michal Linial, Scott Markel, B. J. Morrison McKay, Reinhard Schneider 0002, Paul Horton, Janet Kelso |
Bioinform. | 4 |
| 2012 | Viral Proteins Acquired from a Host Converge to Simplified Domain ArchitecturesabstractThe infection cycle of viruses creates many opportunities for the exchange of genetic material with the host. Many viruses integrate their sequences into the genome of their host for replication. These processes may lead to the virus acquisition of host sequences. Such sequences are prone to accumulation of mutations and deletions. However, in rare instances, sequences acquired from a host become beneficial for the virus. We searched for unexpected sequence similarity among the 900,000 viral proteins and all proteins from cellular organisms. Here, we focus on viruses that infect metazoa. The high-conservation analysis yielded 187 instances of highly similar viral-host sequences. Only a small number of them represent viruses that hijacked host sequences. The low-conservation sequence analysis utilizes the Pfam family collection. About 5% of the 12,000 statistical models archived in Pfam are composed of viral-metazoan proteins. In about half of Pfam families, we provide indirect support for the directionality from the host to the virus. The other families are either wrongly annotated or reflect an extensive sequence exchange between the viruses and their hosts. In about 75% of cross-taxa Pfam families, the viral proteins are significantly shorter than their metazoan counterparts. The tendency for shorter viral proteins relative to their related host proteins accounts for the acquisition of only a fragment of the host gene, the elimination of an internal domain and shortening of the linkers between domains. We conclude that, along viral evolution, the host-originated sequences accommodate simplified domain compositions. We postulate that the trimmed proteins act by interfering with the fundamental function of the host including intracellular signaling, post-translational modification, protein-protein interaction networks and cellular trafficking. We compiled a collection of hijacked protein sequences. These sequences are attractive targets for manipulation of viral infection. Nadav Rappoport, Michal Linial |
PLoS Comput. Biol. | 2 |
| 2011 | Geometric Interpretation of Gene Expression by Sparse Reconstruction of Transcript Profiles
Yosef Prat, Menachem Fromer, Michal Linial, Nathan Linial |
RECOMB | 3 |
| 2011 | Recovering key biological constituents through sparse representation of gene expressionabstractMOTIVATION: Large-scale RNA expression measurements are generating enormous quantities of data. During the last two decades, many methods were developed for extracting insights regarding the interrelationships between genes from such data. The mathematical and computational perspectives that underlie these methods are usually algebraic or probabilistic. RESULTS: Here, we introduce an unexplored geometric view point where expression levels of genes in multiple experiments are interpreted as vectors in a high-dimensional space. Specifically, we find, for the expression profile of each particular gene, its approximation as a linear combination of profiles of a few other genes. This method is inspired by recent developments in the realm of compressed sensing in the machine learning domain. To demonstrate the power of our approach in extracting valuable information from the expression data, we independently applied it to large-scale experiments carried out on the yeast and malaria parasite whole transcriptomes. The parameters extracted from the sparse reconstruction of the expression profiles, when fed to a supervised learning platform, were used to successfully predict the relationships between genes throughout the Gene Ontology hierarchy and protein-protein interaction map. Extensive assessment of the biological results shows high accuracy in both recovering known predictions and in yielding accurate predictions missing from the current databases. We suggest that the geometrical approach presented here is suitable for a broad range of high-dimensional experimental data. Yosef Prat, Menachem Fromer, Nathan Linial, Michal Linial |
Bioinform. | 4 |
| 2011 | Generative probabilistic models for protein-protein interaction networks - the biclique perspectiveabstractMOTIVATION: Much of the large-scale molecular data from living cells can be represented in terms of networks. Such networks occupy a central position in cellular systems biology. In the protein-protein interaction (PPI) network, nodes represent proteins and edges represent connections between them, based on experimental evidence. As PPI networks are rich and complex, a mathematical model is sought to capture their properties and shed light on PPI evolution. The mathematical literature contains various generative models of random graphs. It is a major, still largely open question, which of these models (if any) can properly reproduce various biologically interesting networks. Here, we consider this problem where the graph at hand is the PPI network of Saccharomyces cerevisiae. We are trying to distinguishing between a model family which performs a process of copying neighbors, represented by the duplication-divergence (DD) model, and models which do not copy neighbors, with the Barabási-Albert (BA) preferential attachment model as a leading example. RESULTS: The observed property of the network is the distribution of maximal bicliques in the graph. This is a novel criterion to distinguish between models in this area. It is particularly appropriate for this purpose, since it reflects the graph's growth pattern under either model. This test clearly favors the DD model. In particular, for the BA model, the vast majority (92.9%) of the bicliques with both sides ≥4 must be already embedded in the model's seed graph, whereas the corresponding figure for the DD model is only 5.1%. Our results, based on the biclique perspective, conclusively show that a naïve unmodified DD model can capture a key aspect of PPI networks. Regev Schweiger, Michal Linial, Nathan Linial |
Bioinform. | 2 |
| 2010 | MiRror: a combinatorial analysis web tool for ensembles of microRNAs and their targetsabstractSUMMARY: The miRror application provides insights on microRNA (miRNA) regulation. It is based on the notion of a combinatorial regulation by an ensemble of miRNAs or genes. miRror integrates predictions from a dozen of miRNA resources that are based on complementary algorithms into a unified statistical framework. For miRNAs set as input, the online tool provides a ranked list of targets, based on set of resources selected by the user, according to their significance of being coordinately regulated. Symmetrically, a set of genes can be used as input to suggest a set of miRNAs. The user can restrict the analysis for the preferred tissue or cell line. miRror is suitable for analyzing results from miRNAs profiling, proteomics and gene expression arrays. AVAILABILITY: http://www.proto.cs.huji.ac.il/mirror Yitzhak Friedman, Guy Naamati, Michal Linial |
Bioinform. | 3 |
| 2010 | Exposing the co-adaptive potential of protein-protein interfaces through computational sequence designabstractMOTIVATION: In nature, protein-protein interactions are constantly evolving under various selective pressures. Nonetheless, it is expected that crucial interactions are maintained through compensatory mutations between interacting proteins. Thus, many studies have used evolutionary sequence data to extract such occurrences of correlated mutation. However, this research is confounded by other evolutionary pressures that contribute to sequence covariance, such as common ancestry. RESULTS: Here, we focus exclusively on the compensatory mutations deriving from physical protein interactions, by performing large-scale computational mutagenesis experiments for >260 protein-protein interfaces. We investigate the potential for co-adaptability present in protein pairs that are always found together in nature (obligate) and those that are occasionally in complex (transient). By modeling each complex both in bound and unbound forms, we find that naturally transient complexes possess greater relative capacity for correlated mutation than obligate complexes, even when differences in interface size are taken into account. Menachem Fromer, Michal Linial |
Bioinform. | 2 |
| 2010 | SPRINT: side-chain prediction inference toolbox for multistate protein designabstractUNLABELLED: SPRINT is a software package that performs computational multistate protein design using state-of-the-art inference on probabilistic graphical models. The input to SPRINT is a list of protein structures, the rotamers modeled for each structure and the pre-calculated rotamer energies. Probabilistic inference is performed using the belief propagation or A* algorithms, and dead-end elimination can be applied as pre-processing. The output can either be a list of amino acid sequences simultaneously compatible with these structures, or probabilistic amino acid profiles compatible with the structures. In addition, higher order (e.g. pairwise) amino acid probabilities can also be predicted. Finally, SPRINT also has a module for protein side-chain prediction and single-state design. AVAILABILITY: The full C++ source code for SPRINT can be freely downloaded from http://www.protonet.cs.huji.ac.il/sprint. Menachem Fromer, Chen Yanover, Amir Harel, Ori Shachar, Yair Weiss, Michal Linial |
Bioinform. | 6 |
| 2010 | A predictor for toxin-like proteins exposes cell modulator candidates within viral genomesabstractMOTIVATION: Animal toxins operate by binding to receptors and ion channels. These proteins are short and vary in sequence, structure and function. Sporadic discoveries have also revealed endogenous toxin-like proteins in non-venomous organisms. Viral proteins are the largest group of quickly evolving proteomes. We tested the hypothesis that toxin-like proteins exist in viruses and that they act to modulate functions of their hosts. RESULTS: We updated and improved a classifier for compact proteins resembling short animal toxins that is based on a machine-learning method. We applied it in a large-scale setting to identify toxin-like proteins among short viral proteins. Among the approximately 26 000 representatives of such short proteins, 510 sequences were positively identified. We focused on the 19 highest scoring proteins. Among them, we identified conotoxin-like proteins, growth factors receptor-like proteins and anti-bacterial peptides. Our predictor was shown to enhance annotation inference for many 'uncharacterized' proteins. We conclude that our protocol can expose toxin-like proteins in unexplored niches including metagenomics data and enhance the systematic discovery of novel cell modulators for drug development. AVAILABILITY: ClanTox is available at http://www.clantox.cs.huji.ac.il. Guy Naamati, Manor Askenazi, Michal Linial |
Bioinform. | 3 |
| 2010 | UFFizi: a generic platform for ranking informative featuresabstractBACKGROUND: Feature selection is an important pre-processing task in the analysis of complex data. Selecting an appropriate subset of features can improve classification or clustering and lead to better understanding of the data. An important example is that of finding an informative group of genes out of thousands that appear in gene-expression analysis. Numerous supervised methods have been suggested but only a few unsupervised ones exist. Unsupervised Feature Filtering (UFF) is such a method, based on an entropy measure of Singular Value Decomposition (SVD), ranking features and selecting a group of preferred ones. RESULTS: We analyze the statistical properties of UFF and present an efficient approximation for the calculation of its entropy measure. This allows us to develop a web-tool that implements the UFF algorithm. We propose novel criteria to indicate whether a considered dataset is amenable to feature selection by UFF. Relying on formalism similar to UFF we propose also an Unsupervised Detection of Outliers (UDO) method, providing a novel definition of outliers and producing a measure to rank the "outlier-degree" of an instance.Our methods are demonstrated on gene and microRNA expression datasets, covering viral infection disease and cancer. We apply UFFizi to select genes from these datasets and discuss their biological and medical relevance. CONCLUSIONS: Statistical properties extracted from the UFF algorithm can distinguish selected features from others. UFFizi is a framework that is based on the UFF algorithm and it is applicable for a wide range of diseases. The framework is also implemented as a web-tool.The web-tool is available at: http://adios.tau.ac.il/UFFizi. Assaf Gottlieb, Roy Varshavsky, Michal Linial, David Horn 0001 |
BMC Bioinform. | 3 |
| 2008 | Efficient algorithms for accurate hierarchical clustering of huge datasets: tackling the entire protein spaceabstractMOTIVATION: UPGMA (average linking) is probably the most popular algorithm for hierarchical data clustering, especially in computational biology. However, UPGMA requires the entire dissimilarity matrix in memory. Due to this prohibitive requirement, UPGMA is not scalable to very large datasets. APPLICATION: We present a novel class of memory-constrained UPGMA (MC-UPGMA) algorithms. Given any practical memory size constraint, this framework guarantees the correct clustering solution without explicitly requiring all dissimilarities in memory. The algorithms are general and are applicable to any dataset. We present a data-dependent characterization of hardness and clustering efficiency. The presented concepts are applicable to any agglomerative clustering formulation. RESULTS: We apply our algorithm to the entire collection of protein sequences, to automatically build a comprehensive evolutionary-driven hierarchy of proteins from sequence alone. The newly created tree captures protein families better than state-of-the-art large-scale methods such as CluSTr, ProtoNet4 or single-linkage clustering. We demonstrate that leveraging the entire mass embodied in all sequence similarities allows to significantly improve on current protein family clusterings which are unable to directly tackle the sheer mass of this data. Furthermore, we argue that non-metric constraints are an inherent complexity of the sequence space and should not be overlooked. The robustness of UPGMA allows significant improvement, especially for multidomain proteins, and for large or divergent families. AVAILABILITY: A comprehensive tree built from all UniProt sequence similarities, together with navigation and classification tools will be made available as part of the ProtoNet service. A C++ implementation of the algorithm is available on request. Yaniv Loewenstein, Elon Portugaly, Menachem Fromer, Michal Linial |
ISMB | 4 |
| 2008 | ISMB 2008 TorontoabstractISCB) presents the Sixteenth International Conference on Intelligent Systems for Molecular Biology (ISMB 2008), to be held in Toronto, Canada, July 19-23, 2008.Now in the final phases of scheduling selected presentations, demonstrations, and posters, the organizers are preparing what will likely be recognized as the premier conference on computational biology in 2008.ISMB 2008 (http://www.iscb.org/ismb2008/)will follow the road paved by the ISMB/ ECCB 2007 (http://www.iscb.org/ismbeccb2007/) in Vienna in the attempt to specifically encourage increased participation from previously under-represented disciplines of computational biology.This conference will feature the best of the computer and life sciences through a variety of core sessions running in multiple parallel tracks, along with single-tracked Keynote Presentations, posters on display throughout the duration of the conference, and an extensive commercial exposition.The first day (July 18) of the meeting is reserved for two-day Special Interest Group (SIG) and Satellite meetings, the second day (July 19) runs SIGs for the first time in parallel with Tutorials and the Student Council Symposium, and for the first time two SIGs are running in parallel with the main ISMB meeting (July 20-23). Michal Linial, Jill P. Mesirov, B. J. Morrison McKay, Burkhard Rost |
PLoS Comput. Biol. | 1 |
| 2007 | Clustering Algorithms Optimizer: A Framework for Large Datasets
Roy Varshavsky, David Horn 0001, Michal Linial |
ISBRA | 3 |
| 2007 | When Less Is More: Improving Classification of Protein Families with a Minimal Set of Global Features
Roy Varshavsky, Menachem Fromer, Amit Man, Michal Linial |
WABI | 4 |
| 2007 | Unsupervised feature selection under perturbations: meeting the challenges of biological dataabstractMOTIVATION: Feature selection methods aim to reduce the complexity of data and to uncover the most relevant biological variables. In reality, information in biological datasets is often incomplete as a result of untrustworthy samples and missing values. The reliability of selection methods may therefore be questioned. METHOD: Information loss is incorporated into a perturbation scheme, testing which features are stable under it. This method is applied to data analysis by unsupervised feature filtering (UFF). The latter has been shown to be a very successful method in analysis of gene-expression data. RESULTS: We find that the UFF quality degrades smoothly with information loss. It remains successful even under substantial damage. Our method allows for selection of a best imputation method on a dataset treated by UFF. More importantly, scoring features according to their stability under information loss is shown to be correlated with biological importance in cancer studies. This scoring may lead to novel biological insights. Roy Varshavsky, Assaf Gottlieb, David Horn 0001, Michal Linial |
Bioinform. | 4 |
| 2006 | The Secrets of a Functional Synapse - From a Computational and Experimental ViewpointabstractBACKGROUND: Neuronal communication is tightly regulated in time and in space. The neuronal transmission takes place in the nerve terminal, at a specialized structure called the synapse. Following neuronal activation, an electrical signal triggers neurotransmitter (NT) release at the active zone. The process starts by the signal reaching the synapse followed by a fusion of the synaptic vesicle and diffusion of the released NT in the synaptic cleft; the NT then binds to the appropriate receptor, and as a result, a potential change at the target cell membrane is induced. The entire process lasts for only a fraction of a millisecond. An essential property of the synapse is its capacity to undergo biochemical and morphological changes, a phenomenon that is referred to as synaptic plasticity. RESULTS: In this survey, we consider the mammalian brain synapse as our model. We take a cell biological and a molecular perspective to present fundamental properties of the synapse:(i) the accurate and efficient delivery of organelles and material to and from the synapse; (ii) the coordination of gene expression that underlies a particular NT phenotype; (iii) the induction of local protein expression in a subset of stimulated synapses. We describe the computational facet and the formulation of the problem for each of these topics. CONCLUSION: Predicting the behavior of a synapse under changing conditions must incorporate genomics and proteomics information with new approaches in computational biology. Michal Linial |
BMC Bioinform. | 1 |
| 2006 | EVEREST: automatic identification and classification of protein domains in all protein sequencesabstractBACKGROUND: Proteins are comprised of one or several building blocks, known as domains. Such domains can be classified into families according to their evolutionary origin. Whereas sequencing technologies have advanced immensely in recent years, there are no matching computational methodologies for large-scale determination of protein domains and their boundaries. We provide and rigorously evaluate a novel set of domain families that is automatically generated from sequence data. Our domain family identification process, called EVEREST (EVolutionary Ensembles of REcurrent SegmenTs), begins by constructing a library of protein segments that emerge in an all vs. all pairwise sequence comparison. It then proceeds to cluster these segments into putative domain families. The selection of the best putative families is done using machine learning techniques. A statistical model is then created for each of the chosen families. This procedure is then iterated: the aforementioned statistical models are used to scan all protein sequences, to recreate a library of segments and to cluster them again. RESULTS: Processing the Swiss-Prot section of the UniProt Knoledgebase, release 7.2, EVEREST defines 20,230 domains, covering 85% of the amino acids of the Swiss-Prot database. EVEREST annotates 11,852 proteins (6% of the database) that are not annotated by Pfam A. In addition, in 43,086 proteins (20% of the database), EVEREST annotates a part of the protein that is not annotated by Pfam A. Performance tests show that EVEREST recovers 56% of Pfam A families and 63% of SCOP families with high accuracy, and suggests previously unknown domain families with at least 51% fidelity. EVEREST domains are often a combination of domains as defined by Pfam or SCOP and are frequently sub-domains of such domains. CONCLUSION: The EVEREST process and its output domain families provide an exhaustive and validated view of the protein domain world that is automatically generated from sequence data. The EVEREST library of domain families, accessible for browsing and download at 1, provides a complementary view to that provided by other existing libraries. Furthermore, since it is automatic, the EVEREST process is scalable and we will run it in the future on larger databases as well. The EVEREST source files are available for download from the EVEREST web site. Elon Portugaly, Amir Harel, Nathan Linial, Michal Linial |
BMC Bioinform. | 4 |
| 2005 | Predicting fold novelty based on ProtoNet hierarchical classificationabstractAbstract Motivation: Structural genomics projects aim to solve a large number of protein structures with the ultimate objective of representing the entire protein space. The computational challenge is to identify and prioritize a small set of proteins with new, currently unknown, superfamilies or folds. Results: We develop a method that assigns each protein a likelihood of it belonging to a new, yet undetermined, structural superfamily. The method relies on a variant of ProtoNet, an automatic hierarchical classification scheme of all protein sequences from SwissProt. Our results show that proteins that are remote from solved structures in the ProtoNet hierarchy are more likely to belong to new superfamilies. The results are validated against SCOP releases from recent years that account for about half of the solved structures known to date. We show that our new method and the representation of ProtoNet are superior in detecting new targets, compared to our previous method using ProtoMap classification. Furthermore, our method outperforms PSI-BLAST search in detecting potential new superfamilies. Availability: An interactive tool implementing this method, named ProTarget, is available at http://www.protarget.cs.huji.ac.il. It can be used interactively to retrieve a list of candidate proteins for Structural genomics projects. Supplementary material is available at http://www.protarget.cs.huji.ac.il/supplement Contact: [email protected] Ilona Kifer, Ori Sasson, Michal Linial |
Bioinform. | 3 |
| 2005 | Automatic detection of false annotations via binary property clusteringabstractBACKGROUND: Computational protein annotation methods occasionally introduce errors. False-positive (FP) errors are annotations that are mistakenly associated with a protein. Such false annotations introduce errors that may spread into databases through similarity with other proteins. Generally, methods used to minimize the chance for FPs result in decreased sensitivity or low throughput. We present a novel protein-clustering method that enables automatic separation of FP from true hits. The method quantifies the biological similarity between pairs of proteins by examining each protein's annotations, and then proceeds by clustering sets of proteins that received similar annotation into biological groups. RESULTS: Using a test set of all PROSITE signatures that are marked as FPs, we show that the method successfully separates FPs in 69% of the 327 test cases supplied by PROSITE. Furthermore, we constructed an extensive random FP simulation test and show a high degree of success in detecting FP, indicating that the method is not specifically tuned for PROSITE and performs well on larger scales. We also suggest some means of predicting in which cases this approach would be successful. CONCLUSION: Automatic detection of FPs may greatly facilitate the manual validation process and increase annotation sensitivity. With the increasing number of automatic annotations, the tendency of biological properties to be clustered, once a biological similarity measure is introduced, may become exceedingly helpful in the development of such automatic methods. Noam Kaplan, Michal Linial |
BMC Bioinform. | 2 |
| 2004 | A functional hierarchical organization of the protein sequence spaceabstractBACKGROUND: It is a major challenge of computational biology to provide a comprehensive functional classification of all known proteins. Most existing methods seek recurrent patterns in known proteins based on manually-validated alignments of known protein families. Such methods can achieve high sensitivity, but are limited by the necessary manual labor. This makes our current view of the protein world incomplete and biased. This paper concerns ProtoNet, a automatic unsupervised global clustering system that generates a hierarchical tree of over 1,000,000 proteins, based solely on sequence similarity. RESULTS: In this paper we show that ProtoNet correctly captures functional and structural aspects of the protein world. Furthermore, a novel feature is an automatic procedure that reduces the tree to 12% its original size. This procedure utilizes only parameters intrinsic to the clustering process. Despite the substantial reduction in size, the system's predictive power concerning biological functions is hardly affected. We then carry out an automatic comparison with existing functional protein annotations. Consequently, 78% of the clusters in the compressed tree (5,300 clusters) get assigned a biological function with a high confidence. The clustering and compression processes are unsupervised, and robust. CONCLUSIONS: We present an automatically generated unbiased method that provides a hierarchical classification of all currently known proteins. Noam Kaplan, Moriah Friedlich, Menachem Fromer, Michal Linial |
BMC Bioinform. | 4 |
| 2002 | The metric space of proteins-comparative study of clustering algorithmsabstractAbstract Motivation: A large fraction of biological research concentrates on individual proteins and on small families of proteins. One of the current major challenges in bioinformatics is to extend our knowledge to very large sets of proteins. Several major projects have tackled this problem. Such undertakings usually start with a process that clusters all known proteins or large subsets of this space. Some work in this area is carried out automatically, while other attempts incorporate expert advice and annotation. Results: We propose a novel technique that automatically clusters protein sequences. We consider all proteins in SWISSPROT, and carry out an all-against-all BLAST similarity test among them. With this similarity measure in hand we proceed to perform a continuous bottom-up clustering process by applying alternative rules for merging clusters. The outcome of this clustering process is a classification of the input proteins into a hierarchy of clusters of varying degrees of granularity. Here we compare the clusters that result from alternative merging rules, and validate the results against InterPro. Our preliminary results show that clusters that are consistent with several rather than a single merging rule tend to comply with InterPro annotation. This is an affirmation of the view that the protein space consists of families that differ markedly in their evolutionary conservation. Availability: The outcome of these investigations can be viewed in an interactive Web site at http://www.protonet.cs.huji.ac.il Supplementary information: Biological examples for comparing the performance of the different algorithms used for classification are presented in http://www.protonet.cs.huji.ac.il/examples.html Contact: [email protected] Keywords: protein families; protein classification; sequence alignment; clustering. Ori Sasson, Nathan Linial, Michal Linial |
ISMB | 3 |
| 2002 | Functional Consequences in Metabolic Pathways from Phylogenetic Profiles
Yonatan Bilu, Michal Linial |
WABI | 2 |
| 2002 | Selecting targets for structural determination by navigating in a graph of protein familiesabstractMOTIVATION: A major goal in structural genomics is to enrich the catalogue of proteins whose 3D structures are known. In an attempt to address this problem we mapped over 10 000 proteins with solved structures onto a graph of all Swissprot protein sequences (release 36, approximately 73 000 proteins) provided by ProtoMap, with the goal of sorting proteins according to their likelihood of belonging to new superfamilies. We hypothesized that proteins within neighbouring clusters tend to share common structural superfamilies or folds. If true, the likelihood of finding new superfamilies increases in clusters that are distal from other solved structures within the graph. RESULTS: We defined an order relation between unsolved proteins according to their 'distance' from solved structures in the graph, and sorted approximately 48 000 proteins. Our list can be partitioned into three groups: approximately 35 000 proteins sharing a cluster with at least one known structure; approximately 6500 proteins in clusters with no solved structure but with neighbouring clusters containing known structures; and a third group contains the rest of the proteins, approximately 6100 (in 1274 clusters). We tested the quality of the order relation using thousands of recently solved structures that were not included when the order was defined. The tests show that our order is significantly better (P-value approximately 10(5)) than a random order. More interestingly, the order within the union of the second and third groups, and the order within the third group alone, perform better than random (P-values: 0.0008 and 0.15, respectively) and are better than alternative orders created using PSI-BLAST. Herein, we present a method for selecting targets to be used in structural genomics projects. AVAILABILITY: List of proteins to be used for targets selection combined with a set of biological filters for narrowing down potential targets is in http://www.protarget.cs.huji.ac.il. Elon Portugaly, Ilona Kifer, Michal Linial |
Bioinform. | 3 |
| 2001 | On the predictive power of sequence similarity in yeastabstractPerhaps the most direct way to infer functional linkage of proteins is through structural similarity. However, structure determination lags behind DNA sequencing. Here we show that sequence similarity based on nucleotide sequences alone between ORFs in yeast is indicative of the corresponding genes being of the same functional group, having a similar gene expression pattern or of being involved in a protein-protein interaction. In particular, we compare the nucleotide sequences corresponding to the 6280 yeast ORFs using BLAST, and then cluster them together using a simple neighbor-joining algorithm. This, in effect, gives us hierarchical clustering of 53 levels, where higher levels have bigger clusters. We compare the clustering to large databases that are not based a-priori on sequence information to get a notion of how well our clustering is correlated with this data. For functional annotation we use the SGD database that gives one of 540 annotations for about half the yeast genes. For all pairs that appear within a cluster, we test the hypothesis that almost all genes within the same cluster have the same function. We get very high percentage rates of correct annotation at the lower levels of the hierarchy, which decreases gradually at higher ones. From the results of the large scale gene expression experiments we generate a list of pairs of genes, whose expression is highly correlated (≱0.9). We then go over this list of pairs, and count how many of them are contained within a cluster as opposed to the expected number by a random model. We estimate the significance of our results using simulation, and get for all levels p-value≰≰0.001. The third type of data obtained from the protein-protein interaction database that is given as a list of airs of proteins involved in interaction. Yonatan Bilu, Michal Linial |
RECOMB | 2 |
| 2000 | Using Bayesian networks to analyze expression dataabstractDNA hybridization arrays simultaneously measure the expression level for thousands of genes. These measurements provide a “snapshot” of transcription levels within the cell. A major challenge in computational biology is to uncover, from such measurements, gene/protein interactions and key biological features of cellular systems. Nir Friedman, Michal Linial, Iftach Nachman, Dana Pe'er |
RECOMB | 2 |
| 2000 | Probabilities for having a new fold on the basis of a map of all protein sequencesabstractIt is a major problem in the study of protein structure to predict which proteins have new, currently unknown structural folds. In an attempt to address this problem we studied the location of all proteins with solved structures within the map of all known protein sequences provided by ProtoMap. The mutual distances in this map among solved structures are used to derive a probabilistic model from which we infer an estimate for the probability of an unsolved protein to have a new fold. The probabilities were based on data from SCOP release 1.37. The results were evaluated against the more recent SCOP pre-release 1.41. Our predicted probabilities for unsolved proteins to have a new fold are very well correlated with the proportion of new folds among recently released structures. Thus, information about the structure of proteins can be inferred from a global relational view of protein sequences. Finally, the same procedure was applied to estimate probabilities on the basis of SCOP 1.41. A list of the highest scoring proteins is provided: These are about 80 non-membranous proteins that belong to clusters with more than 5 proteins and achieve the highest probability to have a new fold. A rational selection for 3D determination of those targets is expected to accelerate the pace of new fold discovery. Elon Portugaly, Michal Linial |
RECOMB | 2 |
| 1998 | A Map of the Protein Space: An Automatic Hierarchical Classification of all Protein Sequences
Golan Yona, Nathan Linial, Naftali Tishby, Michal Linial |
ISMB | 4 |