EDBT 2026 Demo / reviewers in the wild / expert
Andrew Dickson
dblp:370/5300
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2024
0000-0002-1146-6346ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein function prediction |
1.4 | 2 | 2024 | Fine-tuning protein embeddings for functional similarity evaluation · Bioinform. 2024 GO Bench: shared hub for universal benchmarking of machine learning-based protein functional annotations · Bioinform. 2023 |
Bioinformatics and computational biology › protein function prediction
gene ontology annotation |
0.8 | 1 | 2024 | Fine-tuning protein embeddings for functional similarity evaluation · Bioinform. 2024 |
Bioinformatics and computational biology › protein sequence analysis › protein family analysis
protein family clustering |
0.2 | 1 | 2024 | Fine-tuning protein embeddings for functional similarity evaluation · Bioinform. 2024 |
Methods — techniques the papers use, named apart from their topics
language model fine-tuning · 0.8k-nearest neighbor · 0.8machine learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fine-tuning protein embeddings for functional similarity evaluationabstractMOTIVATION: Proteins with unknown function are frequently compared to better characterized relatives, either using sequence similarity, or recently through similarity in a learned embedding space. Through comparison, protein sequence embeddings allow for interpretable and accurate annotation of proteins, as well as for downstream tasks such as clustering for unsupervised discovery of protein families. However, it is unclear whether embeddings can be deliberately designed to improve their use in these downstream tasks. RESULTS: We find that for functional annotation of proteins, as represented by Gene Ontology (GO) terms, direct fine-tuning of language models on a simple classification loss has an immediate positive impact on protein embedding quality. Fine-tuned embeddings show stronger performance as representations for K-nearest neighbor classifiers, reaching stronger performance for GO annotation than even directly comparable fine-tuned classifiers, while maintaining interpretability through protein similarity comparisons. They also maintain their quality in related tasks, such as rediscovering protein families with clustering. AVAILABILITY AND IMPLEMENTATION: github.com/mofradlab/go_metric. Andrew Dickson, Mohammad R. K. Mofrad |
Bioinform. | 1 |
| 2023 | GO Bench: shared hub for universal benchmarking of machine learning-based protein functional annotationsabstractMOTIVATION: Gene annotation is the problem of mapping proteins to their functions represented as Gene Ontology (GO) terms, typically inferred based on the primary sequences. Gene annotation is a multi-label multi-class classification problem, which has generated growing interest for its uses in the characterization of millions of proteins with unknown functions. However, there is no standard GO dataset used for benchmarking the newly developed new machine learning models within the bioinformatics community. Thus, the significance of improvements for these models remains unclear. RESULTS: The Gene Benchmarking database is the first effort to provide an easy-to-use and configurable hub for the learning and evaluation of gene annotation models. It provides easy access to pre-specified datasets and takes the non-trivial steps of preprocessing and filtering all data according to custom presets using a web interface. The GO bench web application can also be used to evaluate and display any trained model on leaderboards for annotation tasks. AVAILABILITY AND IMPLEMENTATION: The GO Benchmarking dataset is freely available at www.gobench.org. Code is hosted at github.com/mofradlab, with repositories for website code, core utilities and examples of usage (Supplementary Section S.7). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andrew Dickson, Ehsaneddin Asgari, Alice C. McHardy, Mohammad R. K. Mofrad |
Bioinform. | 1 |