Michael Heinzinger

dblp:255/3206 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0002-9601-3580ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein function prediction
1.732023
CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Clustering FunFams using sequence embeddings improves EC purity · Bioinform. 2021
Bioinformatics and computational biology › protein structure analysis › protein domain identification
protein domain classification
0.712023
CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023
Bioinformatics and computational biology › structural bioinformatics
protein structure classification
0.712023
CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection
0.712023
CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023
Natural language and speech › Language models and text generation › neural language model
protein language model
0.612022
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Bioinformatics and computational biology
protein structure prediction
0.612022
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Bioinformatics and computational biology › protein function prediction
protein subcellular localization prediction
0.612022
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction
0.612022
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Bioinformatics and computational biology › protein function prediction
protein classification
0.512021
Clustering FunFams using sequence embeddings improves EC purity · Bioinform. 2021

Methods — techniques the papers use, named apart from their topics

t5 · 1.1self-supervised learning · 1.1albert · 1.1XLNet · 1.1Transformer-XL · 1.1ELECTRA · 1.1BERT · 1.1protein language model embeddings · 0.7neural network · 0.7hidden markov model · 0.7
YearPublicationVenuePosition
2023 CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models
abstract
MOTIVATION: CATH is a protein domain classification resource that exploits an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues missed by state-of-the-art hidden Markov model (HMM)-based approaches. The method developed (CATHe) combines a neural network with sequence representations obtained from protein language models. It was assessed using a dataset of remote homologues having less than 20% sequence identity to any domain in the training set. RESULTS: The CATHe models trained on 1773 largest and 50 largest CATH superfamilies had an accuracy of 85.6 ± 0.4% and 98.2 ± 0.3%, respectively. As a further test of the power of CATHe to detect more remote homologues missed by HMMs derived from CATH domains, we used a dataset consisting of protein domains that had annotations in Pfam, but not in CATH. By using highly reliable CATHe predictions (expected error rate <0.5%), we were able to provide CATH annotations for 4.62 million Pfam domains. For a subset of these domains from Homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold2 structures with structures from the CATH superfamilies to which they were assigned. AVAILABILITY AND IMPLEMENTATION: The code for the developed models is available on https://github.com/vam-sin/CATHe, and the datasets developed in this study can be accessed on https://zenodo.org/record/6327572. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vamsi Nallapareddy, Nicola Bordin, Ian Sillitoe, Michael Heinzinger, Maria Littmann, Vaishali P. Waman, Neeladri Sen, Burkhard Rost, Christine A. Orengo
Bioinform.4
2022 ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
abstract
Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The protein LMs (pLMs) were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw pLM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks: (1) a per-residue (per-token) prediction of protein secondary structure (3-state accuracy Q3=81%-87%); (2) per-protein (pooling) predictions of protein sub-cellular location (ten-state accuracy: Q10=81%) and membrane versus water-soluble (2-state accuracy Q2=91%). For secondary structure, the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without multiple sequence alignments (MSAs) or evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that pLMs learned some of the grammar of the language of life. All our models are available through https://github.com/agemagician/ProtTrans.
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang 0008, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, Burkhard Rost
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Mutations in transmembrane proteins: diseases, evolutionary insights, prediction and comparison with globular proteins
abstract
Membrane proteins are unique in that they interact with lipid bilayers, making them indispensable for transporting molecules and relaying signals between and across cells. Due to the significance of the protein's functions, mutations often have profound effects on the fitness of the host. This is apparent both from experimental studies, which implicated numerous missense variants in diseases, as well as from evolutionary signals that allow elucidating the physicochemical constraints that intermembrane and aqueous environments bring. In this review, we report on the current state of knowledge acquired on missense variants (referred to as to single amino acid variants) affecting membrane proteins as well as the insights that can be extrapolated from data already available. This includes an overview of the annotations for membrane protein variants that have been collated within databases dedicated to the topic, bioinformatics approaches that leverage evolutionary information in order to shed light on previously uncharacterized membrane protein structures or interaction interfaces, tools for predicting the effects of mutations tailored specifically towards the characteristics of membrane proteins as well as two clinically relevant case studies explaining the implications of mutated membrane proteins in cancer and cardiomyopathy.
Jan Zaucha, Michael Heinzinger, A. Kulandaisamy, Evans Kataka, Óscar Llorian Salvádor, Petr Popov, Burkhard Rost, M. Michael Gromiha, Boris S. Zhorov, Dmitrij Frishman
Briefings Bioinform.2
2021 Clustering FunFams using sequence embeddings improves EC purity
abstract
MOTIVATION: Classifying proteins into functional families can improve our understanding of protein function and can allow transferring annotations within one family. For this, functional families need to be 'pure', i.e., contain only proteins with identical function. Functional Families (FunFams) cluster proteins within CATH superfamilies into such groups of proteins sharing function. 11% of all FunFams (22 830 of 203 639) contain EC annotations and of those, 7% (1526 of 22 830) have inconsistent functional annotations. RESULTS: We propose an approach to further cluster FunFams into functionally more consistent sub-families by encoding their sequences through embeddings. These embeddings originate from language models transferring knowledge gained from predicting missing amino acids in a sequence (ProtBERT) and have been further optimized to distinguish between proteins belonging to the same or a different CATH superfamily (PB-Tucker). Using distances between embeddings and DBSCAN to cluster FunFams and identify outliers, doubled the number of pure clusters per FunFam compared to random clustering. Our approach was not limited to FunFams but also succeeded on families created using sequence similarity alone. Complementing EC annotations, we observed similar results for binding annotations. Thus, we expect an increased purity also for other aspects of function. Our results can help generating FunFams; the resulting clusters with improved functional consistency allow more reliable inference of annotations. We expect this approach to succeed equally for any other grouping of proteins by their phenotypes. AVAILABILITY AND IMPLEMENTATION: Code and embeddings are available via GitHub: https://github.com/Rostlab/FunFamsClustering. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maria Littmann, Nicola Bordin, Michael Heinzinger, Konstantin Schütze, Christian Dallago, Christine A. Orengo, Burkhard Rost
Bioinform.3
2019 Modeling aspects of the language of life through transfer-learning protein sequences
abstract
BACKGROUND: Predicting protein function and structure from sequence is one important challenge for computational biology. For 26 years, most state-of-the-art approaches combined machine learning and evolutionary information. However, for some applications retrieving related proteins is becoming too time-consuming. Additionally, evolutionary information is less powerful for small families, e.g. for proteins from the Dark Proteome. Both these problems are addressed by the new methodology introduced here. RESULTS: We introduced a novel way to represent protein sequences as continuous vectors (embeddings) by using the language model ELMo taken from natural language processing. By modeling protein sequences, ELMo effectively captured the biophysical properties of the language of life from unlabeled big data (UniRef50). We refer to these new embeddings as SeqVec (Sequence-to-Vector) and demonstrate their effectiveness by training simple neural networks for two different tasks. At the per-residue level, secondary structure (Q3 = 79% ± 1, Q8 = 68% ± 1) and regions with intrinsic disorder (MCC = 0.59 ± 0.03) were predicted significantly better than through one-hot encoding or through Word2vec-like approaches. At the per-protein level, subcellular localization was predicted in ten classes (Q10 = 68% ± 1) and membrane-bound were distinguished from water-soluble proteins (Q2 = 87% ± 1). Although SeqVec embeddings generated the best predictions from single sequences, no solution improved over the best existing method using evolutionary information. Nevertheless, our approach improved over some popular methods using evolutionary information and for some proteins even did beat the best. Thus, they prove to condense the underlying principles of protein sequences. Overall, the important novelty is speed: where the lightning-fast HHblits needed on average about two minutes to generate the evolutionary information for a target protein, SeqVec created embeddings on average in 0.03 s. As this speed-up is independent of the size of growing sequence databases, SeqVec provides a highly scalable approach for the analysis of big data in proteomics, i.e. microbiome or metaproteome analysis. CONCLUSION: Transfer-learning succeeded to extract information from unlabeled sequence databases relevant for various protein prediction tasks. SeqVec modeled the language of life, namely the principles underlying protein sequences better than any features suggested by textbooks and prediction methods. The exception is evolutionary information, however, that information is not available on the level of a single sequence.
Michael Heinzinger, Ahmed Elnaggar, Yu Wang 0008, Christian Dallago, Dmitrii Nechaev, Florian Matthes, Burkhard Rost
BMC Bioinform.1