VLDB 2026 Research / reviewers in the wild / expert
Alex Warwick Vesztrocy
dblp:213/7041
· DBLP profile ↗
4ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-4074-4261ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
comparative genomics |
0.9 | 1 | 2025 | Annotation matters: the effect of structural gene annotation on orthology inference · Bioinform. 2025 |
Bioinformatics and computational biology › genome annotation
gene annotation |
0.9 | 1 | 2025 | Annotation matters: the effect of structural gene annotation on orthology inference · Bioinform. 2025 |
Bioinformatics and computational biology › comparative genomics › orthology analysis
orthology inference |
0.9 | 1 | 2025 | Annotation matters: the effect of structural gene annotation on orthology inference · Bioinform. 2025 |
Bioinformatics and computational biology › functional genomics
gene function prediction |
0.8 | 2 | 2020 | Benchmarking gene ontology function predictions using negative annotations · Bioinform. 2020 Prioritising candidate genes causing QTL using hierarchical orthologous groups · Bioinform. 2018 |
Bioinformatics and computational biology
phylogenetics |
0.5 | 1 | 2021 | OMAmer: tree-driven and alignment-free protein assignment to subfamilies outperforms closest sequence approaches · Bioinform. 2021 |
Bioinformatics and computational biology › protein function prediction
gene ontology annotation |
0.4 | 1 | 2020 | Benchmarking gene ontology function predictions using negative annotations · Bioinform. 2020 |
Bioinformatics and computational biology › statistical genetics
quantitative trait locus analysis |
0.3 | 1 | 2018 | Prioritising candidate genes causing QTL using hierarchical orthologous groups · Bioinform. 2018 |
Methods — techniques the papers use, named apart from their topics
smith-waterman · 0.5evolutionarily informed k-mers · 0.5DIAMOND · 0.5orthology-based prediction · 0.4BLAST · 0.4hierarchical orthologous groups · 0.3gene ontology annotation propagation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Annotation matters: the effect of structural gene annotation on orthology inferenceabstractMOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark. Silvia Prieto-Baños, Yannis Nevers, Adrian M. Altenhoff, Alex Warwick Vesztrocy, Christophe Dessimoz, Natasha M. Glover |
Bioinform. | 4 |
| 2021 | OMAmer: tree-driven and alignment-free protein assignment to subfamilies outperforms closest sequence approachesabstractMOTIVATION: Assigning new sequences to known protein families and subfamilies is a prerequisite for many functional, comparative and evolutionary genomics analyses. Such assignment is commonly achieved by looking for the closest sequence in a reference database, using a method such as BLAST. However, ignoring the gene phylogeny can be misleading because a query sequence does not necessarily belong to the same subfamily as its closest sequence. For example, a hemoglobin which branched out prior to the hemoglobin alpha/beta duplication could be closest to a hemoglobin alpha or beta sequence, whereas it is neither. To overcome this problem, phylogeny-driven tools have emerged but rely on gene trees, whose inference is computationally expensive. RESULTS: Here, we first show that in multiple animal and plant datasets, 18-62% of assignments by closest sequence are misassigned, typically to an over-specific subfamily. Then, we introduce OMAmer, a novel alignment-free protein subfamily assignment method, which limits over-specific subfamily assignments and is suited to phylogenomic databases with thousands of genomes. OMAmer is based on an innovative method using evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. Whilst able to reject non-homologous family-level assignments, we show that OMAmer provides better and quicker subfamily-level assignments than approaches relying on the closest sequence, whether inferred exactly by Smith-Waterman or by the fast heuristic DIAMOND. AVAILABILITYAND IMPLEMENTATION: OMAmer is available from the Python Package Index (as omamer), with the source code and a precomputed database available at https://github.com/DessimozLab/omamer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victor Rossier, Alex Warwick Vesztrocy, Marc Robinson-Rechavi, Christophe Dessimoz |
Bioinform. | 2 |
| 2020 | Benchmarking gene ontology function predictions using negative annotationsabstractMOTIVATION: With the ever-increasing number and diversity of sequenced species, the challenge to characterize genes with functional information is even more important. In most species, this characterization almost entirely relies on automated electronic methods. As such, it is critical to benchmark the various methods. The Critical Assessment of protein Function Annotation algorithms (CAFA) series of community experiments provide the most comprehensive benchmark, with a time-delayed analysis leveraging newly curated experimentally supported annotations. However, the definition of a false positive in CAFA has not fully accounted for the open world assumption (OWA), leading to a systematic underestimation of precision. The main reason for this limitation is the relative paucity of negative experimental annotations. RESULTS: This article introduces a new, OWA-compliant, benchmark based on a balanced test set of positive and negative annotations. The negative annotations are derived from expert-curated annotations of protein families on phylogenetic trees. This approach results in a large increase in the average information content of negative annotations. The benchmark has been tested using the naïve and BLAST baseline methods, as well as two orthology-based methods. This new benchmark could complement existing ones in future CAFA experiments. AVAILABILITY AND IMPLEMENTATION: All data, as well as code used for analysis, is available from https://lab.dessimoz.org/20_not. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alex Warwick Vesztrocy, Christophe Dessimoz |
Bioinform. | 1 |
| 2018 | Prioritising candidate genes causing QTL using hierarchical orthologous groupsabstractMotivation: A key goal in plant biotechnology applications is the identification of genes associated to particular phenotypic traits (for example: yield, fruit size, root length). Quantitative Trait Loci (QTL) studies identify genomic regions associated with a trait of interest. However, to infer potential causal genes in these regions, each of which can contain hundreds of genes, these data are usually intersected with prior functional knowledge of the genes. This process is however laborious, particularly if the experiment is performed in a non-model species, and the statistical significance of the inferred candidates is typically unknown. Results: This paper introduces QTLSearch, a method and software tool to search for candidate causal genes in QTL studies by combining Gene Ontology annotations across many species, leveraging hierarchical orthologous groups. The usefulness of this approach is demonstrated by re-analysing two metabolic QTL studies: one in Arabidopsis thaliana, the other in Oryza sativa subsp. indica. Even after controlling for statistical significance, QTLSearch inferred potential causal genes for more QTL than BLAST-based functional propagation against UniProtKB/Swiss-Prot, and for more QTL than in the original studies. Availability and implementation: QTLSearch is distributed under the LGPLv3 license. It is available to install from the Python Package Index (as qtlsearch), with the source available from https://bitbucket.org/alex-warwickvesztrocy/qtlsearch. Supplementary information: Supplementary data are available at Bioinformatics online. Alex Warwick Vesztrocy, Christophe Dessimoz, Henning Redestig |
Bioinform. | 1 |