VLDB 2026 Research / reviewers in the wild / expert
Tomás Pluskal
dblp:03/8709
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-6940-3006ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
2 papers |
Representation and self-supervised learning · 62% Generative modeling · 19% Deep learning architectures and training · 19% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning |
0.8 | 1 | 2024 | Learning to design protein-protein interactions with enhanced generalization · ICLR 2024 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.8 | 1 | 2024 | MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024 |
Bioinformatics and computational biology
metabolomics |
0.8 | 1 | 2024 | MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024 |
Bioinformatics and computational biology › structural biology
molecular structure determination |
0.8 | 1 | 2024 | MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024 |
Bioinformatics and computational biology
protein engineering |
0.8 | 1 | 2024 | Learning to design protein-protein interactions with enhanced generalization · ICLR 2024 |
Bioinformatics and computational biology › protein design
protein-protein interaction design |
0.8 | 1 | 2024 | Learning to design protein-protein interactions with enhanced generalization · ICLR 2024 |
Machine learning › Generative modeling
protein design |
0.2 | 1 | 2024 | Learning to design protein-protein interactions with enhanced generalization · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
thermodynamically motivated loss · 1.5spectrum simulation · 1.5pre-training · 1.5molecule retrieval · 1.5fine-tuning · 1.5de novo molecular structure generation · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MARTS-DB: a database of mechanisms and reactions of terpene synthasesabstractBACKGROUND: Terpene synthases (TPSs) are enzymes that catalyze some of the most complex reactions in nature-the cyclizations of terpenes, which form the carbon backbones to the largest group of natural products, the terpenoids. On average, more than half of the carbon atoms in a terpene scaffold undergo a change in connectivity or configuration during these enzymatic cascades. Understanding TPS reaction mechanisms remains challenging, often requiring intricate computational modeling and isotopic labelling studies. Moreover, the relationship between TPS sequence and catalytic function is difficult to decipher, while data-driven approaches remain limited due to a lack of comprehensive, high-quality data sources. MAIN: We introduce the Mechanisms And Reactions of Terpene Synthases DataBase (MARTS-DB)-a manually curated, structured, and searchable database that integrates TPS enzymes, the terpenes they produce, and their detailed reaction mechanisms. MARTS-DB includes over 2850 reactions catalyzed by 1432 annotated enzymes from across all domains of life, with reaction mechanisms mapped as stepwise cascades for more than 500 terpenes. Accessible at https://www.marts-db.org , the database provides advanced search functionality and supports full dataset downloads in machine-readable formats. It also encourages community contributions to promote continuous growth. CONCLUSION: User-friendly and comprehensive, MARTS-DB enables the systematic exploration of TPS catalysis, opening new avenues for computational analysis and machine learning, as recently demonstrated in the prediction of novel TPSs. Martin Engst, Martin Brokes, Tereza Calounová, Raman Samusevich, Roman Bushuiev, Anton Bushuiev, Ratthachat Chatpatanasiri, Adéla Tajovská, Safa Mert Akmese, Milana Perkovic, Matous Soldát, Josef Sivic, Tomás Pluskal |
BMC Bioinform. | 13 |
| 2025 | An evaluation methodology for machine learning-based tandem mass spectra similarity predictionabstractBACKGROUND: Untargeted tandem mass spectrometry serves as a scalable solution for the organization of small molecules. One of the most prevalent techniques for analyzing the acquired tandem mass spectrometry data (MS/MS) - called molecular networking - organizes and visualizes putatively structurally related compounds. However, a key bottleneck of this approach is the comparison of MS/MS spectra used to identify nearby structural neighbors. Machine learning (ML) approaches have emerged as a promising technique to predict structural similarity from MS/MS that may surpass the current state-of-the-art algorithmic methods. However, the comparison between these different ML methods remains a challenge because there is a lack of standardization to benchmark, evaluate, and compare MS/MS similarity methods, and there are no methods that address data leakage between training and test data in order to analyze model generalizability. RESULT: In this work, we present the creation of a new evaluation methodology using a train/test split that allows for the evaluation of machine learning models at varying degrees of structural similarity between training and test sets. We also introduce a training and evaluation framework that measures prediction accuracy on domain-inspired annotation and retrieval metrics designed to mirror real-world applications. We further show how two alternative training methods that leverage MS specific insights (e.g., similar instrumentation, collision energy, adduct) affect method performance and demonstrate the orthogonality of the proposed metrics. We especially highlight the role that collision energy plays in prediction errors. Finally, we release a continually updated version of our dataset online along with our data cleaning and splitting pipelines for community use. CONCLUSION: It is our hope that this benchmark will serve as the basis of development for future machine learning approaches in MS/MS similarity and facilitate comparison between models. We anticipate that the introduced set of evaluation metrics allows for a better reflection of practical performance. Michael Strobel 0002, Alberto Gil-de-la-Fuente, Mohammad Reza Shahneh, Yasin El Abiead, Roman Bushuiev, Anton Bushuiev, Tomás Pluskal, Mingxun Wang 0001 |
BMC Bioinform. | 7 |
| 2024 | Learning to design protein-protein interactions with enhanced generalizationabstractDiscovering mutations enhancing protein-protein interactions (PPIs) is critical for advancing biomedical research and developing improved therapeutics. While machine learning approaches have substantially advanced the field, they often struggle to generalize beyond training data in practical scenarios. The contributions of this work are three-fold. First, we construct PPIRef, the largest and non-redundant dataset of 3D protein-protein interactions, enabling effective large-scale learning. Second, we leverage the PPIRef dataset to pre-train PPIformer, a new SE(3)-equivariant model generalizing across diverse protein-binder variants. We fine-tune PPIformer to predict effects of mutations on protein-protein interactions via a thermodynamically motivated adjustment of the pre-training loss function. Finally, we demonstrate the enhanced generalization of our new PPIformer approach by outperforming other state-of-the-art methods on new, non-leaking splits of standard labeled PPI mutational data and independent case studies optimizing a human antibody against SARS-CoV-2 and increasing the thrombolytic activity of staphylokinase. Anton Bushuiev, Roman Bushuiev, Petr Kouba, Anatolii Filkin, Marketa Gabrielova, Michal Gabriel, Jirí Sedlár, Tomás Pluskal, Jirí Damborský, Stanislav Mazurenko, Josef Sivic |
ICLR | 8 |
| 2024 | MassSpecGym: A benchmark for the discovery and identification of moleculesabstractThe discovery and identification of molecules in biological and environmental samples is crucial for advancing biomedical and chemical sciences. Tandem mass spectrometry (MS/MS) is the leading technique for high-throughput elucidation of molecular structures. However, decoding a molecular structure from its mass spectrum is exceptionally challenging, even when performed by human experts. As a result, the vast majority of acquired MS/MS spectra remain uninterpreted, thereby limiting our understanding of the underlying (bio)chemical processes. Despite decades of progress in machine learning applications for predicting molecular structures from MS/MS spectra, the development of new methods is severely hindered by the lack of standard datasets and evaluation protocols. To address this problem, we propose MassSpecGym -- the first comprehensive benchmark for the discovery and identification of molecules from MS/MS data. Our benchmark comprises the largest publicly available collection of high-quality MS/MS spectra and defines three MS/MS annotation challenges: \textit{de novo} molecular structure generation, molecule retrieval, and spectrum simulation. It includes new evaluation metrics and a generalization-demanding data split, therefore standardizing the MS/MS annotation tasks and rendering the problem accessible to the broad machine learning community. MassSpecGym is publicly available at \url{https://github.com/pluskal-lab/MassSpecGym}. Roman Bushuiev, Anton Bushuiev, Niek F. de Jonge, Adamo Young, Fleming Kretschmer, Raman Samusevich, Janne Heirman, Fei Wang 0062, Luke Zhang, Kai Dührkop, Marcus Ludwig, Nils A. Haupt, Apurva Kalia, Corinna Brungs, Robin Schmid, Russell Greiner, Bo Wang 0044, David S. Wishart, Liping Liu 0001, Juho Rousu, Wout Bittremieux, Hannes L. Röst, Tytus D. Mak, Soha Hassoun, Florian Huber 0001, Justin J. J. van der Hooft, Michael A. Stravs, Sebastian Böcker, Josef Sivic, Tomás Pluskal |
NeurIPS | 30 |
| 2023 | transXpress: a Snakemake pipeline for streamlined de novo transcriptome assembly and annotationabstractBACKGROUND: RNA-seq followed by de novo transcriptome assembly has been a transformative technique in biological research of non-model organisms, but the computational processing of RNA-seq data entails many different software tools. The complexity of these de novo transcriptomics workflows therefore presents a major barrier for researchers to adopt best-practice methods and up-to-date versions of software. RESULTS: Here we present a streamlined and universal de novo transcriptome assembly and annotation pipeline, transXpress, implemented in Snakemake. transXpress supports two popular assembly programs, Trinity and rnaSPAdes, and allows parallel execution on heterogeneous cluster computing hardware. CONCLUSIONS: transXpress simplifies the use of best-practice methods and up-to-date software for de novo transcriptome assembly, and produces standardized output files that can be mined using SequenceServer to facilitate rapid discovery of new genes and proteins in non-model organisms. Timothy R. Fallon, Tereza Calounová, Martin Mokrejs, Jing-Ke Weng, Tomás Pluskal |
BMC Bioinform. | 5 |
| 2010 | MZmine 2: Modular framework for processing, visualizing, and analyzing mass spectrometry-based molecular profile dataabstractBACKGROUND: Mass spectrometry (MS) coupled with online separation methods is commonly applied for differential and quantitative profiling of biological samples in metabolomic as well as proteomic research. Such approaches are used for systems biology, functional genomics, and biomarker discovery, among others. An ongoing challenge of these molecular profiling approaches, however, is the development of better data processing methods. Here we introduce a new generation of a popular open-source data processing toolbox, MZmine 2. RESULTS: A key concept of the MZmine 2 software design is the strict separation of core functionality and data processing modules, with emphasis on easy usability and support for high-resolution spectra processing. Data processing modules take advantage of embedded visualization tools, allowing for immediate previews of parameter settings. Newly introduced functionality includes the identification of peaks using online databases, MSn data support, improved isotope pattern support, scatter plot visualization, and a new method for peak list alignment based on the random sample consensus (RANSAC) algorithm. The performance of the RANSAC alignment was evaluated using synthetic datasets as well as actual experimental data, and the results were compared to those obtained using other alignment algorithms. CONCLUSIONS: MZmine 2 is freely available under a GNU GPL license and can be obtained from the project website at: http://mzmine.sourceforge.net/. The current version of MZmine 2 is suitable for processing large batches of data and has been applied to both targeted and non-targeted metabolomic analyses. Tomás Pluskal, Sandra Castillo, Alejandro Villar-Briones, Matej Oresic |
BMC Bioinform. | 1 |