VLDB 2026 Research / reviewers in the wild / expert
Vamsi Nallapareddy
dblp:340/5336
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0003-4750-038XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › protein structure analysis › protein domain identification
protein domain classification |
0.7 | 1 | 2023 | CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023 |
Bioinformatics and computational biology
protein function prediction |
0.7 | 1 | 2023 | CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023 |
Bioinformatics and computational biology › structural bioinformatics
protein structure classification |
0.7 | 1 | 2023 | CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023 |
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection |
0.7 | 1 | 2023 | CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models · Bioinform. 2023 |
Methods — techniques the papers use, named apart from their topics
protein language model embeddings · 0.7neural network · 0.7hidden markov model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language modelsabstractMOTIVATION: CATH is a protein domain classification resource that exploits an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues missed by state-of-the-art hidden Markov model (HMM)-based approaches. The method developed (CATHe) combines a neural network with sequence representations obtained from protein language models. It was assessed using a dataset of remote homologues having less than 20% sequence identity to any domain in the training set. RESULTS: The CATHe models trained on 1773 largest and 50 largest CATH superfamilies had an accuracy of 85.6 ± 0.4% and 98.2 ± 0.3%, respectively. As a further test of the power of CATHe to detect more remote homologues missed by HMMs derived from CATH domains, we used a dataset consisting of protein domains that had annotations in Pfam, but not in CATH. By using highly reliable CATHe predictions (expected error rate <0.5%), we were able to provide CATH annotations for 4.62 million Pfam domains. For a subset of these domains from Homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold2 structures with structures from the CATH superfamilies to which they were assigned. AVAILABILITY AND IMPLEMENTATION: The code for the developed models is available on https://github.com/vam-sin/CATHe, and the datasets developed in this study can be accessed on https://zenodo.org/record/6327572. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vamsi Nallapareddy, Nicola Bordin, Ian Sillitoe, Michael Heinzinger, Maria Littmann, Vaishali P. Waman, Neeladri Sen, Burkhard Rost, Christine A. Orengo |
Bioinform. | 1 |