Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Pavel Sulimov

dblp:220/2447 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-2885-2646ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 96% Machine learning and data management · 4%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query optimization › learned query optimization
learned query optimizer
1.622025
GenJoin: Conditional Generative Plan-to-Plan Query Optimizer that Learns from Subplan Hints · Proc. ACM Manag. Data 2025
Is Your Learned Query Optimizer Behaving As You Expect? A Machine Learning Perspective · Proc. VLDB Endow. 2024
Query processing and optimization
query optimization
1.622025
GenJoin: Conditional Generative Plan-to-Plan Query Optimizer that Learns from Subplan Hints · Proc. ACM Manag. Data 2025
Is Your Learned Query Optimizer Behaving As You Expect? A Machine Learning Perspective · Proc. VLDB Endow. 2024
Query processing and optimization
query execution
0.912025
GenJoin: Conditional Generative Plan-to-Plan Query Optimizer that Learns from Subplan Hints · Proc. ACM Manag. Data 2025
Query processing and optimization › query planning
query plan generation
0.912025
GenJoin: Conditional Generative Plan-to-Plan Query Optimizer that Learns from Subplan Hints · Proc. ACM Manag. Data 2025
Bioinformatics and computational biology › proteomics › peptide identification
peptide-spectrum match scoring
0.412020
Annotation of tandem mass spectrometry data using stochastic neural networks in shotgun proteomics · Bioinform. 2020
Bioinformatics and computational biology
proteomics
0.412020
Annotation of tandem mass spectrometry data using stochastic neural networks in shotgun proteomics · Bioinform. 2020
Machine learning and data management
learned database components
0.212024
Is Your Learned Query Optimizer Behaving As You Expect? A Machine Learning Perspective · Proc. VLDB Endow. 2024
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.112020
Annotation of tandem mass spectrometry data using stochastic neural networks in shotgun proteomics · Bioinform. 2020

Methods — techniques the papers use, named apart from their topics

machine learning · 0.9generative modeling · 0.9reinforcement learning · 0.8deep learning · 0.8stochastic neural network · 0.4restricted boltzmann machine · 0.4
YearPublicationVenuePosition
2025 GenJoin: Conditional Generative Plan-to-Plan Query Optimizer that Learns from Subplan Hints
abstract
Query optimization has become a research area where classical algorithms are being challenged by machine learning algorithms. At the same time, recent trends in learned query optimizers have shown that it is prudent to take advantage of decades of database research and augment classical query optimizers by shrinking the plan search space through different types of hints (e.g. by specifying the join type, scan type or the order of joins) rather than completely replacing the classical query optimizer with machine learning models. It is especially relevant for cases when classical optimizers cannot fully enumerate all logical and physical plans and, as an alternative, need to rely on less robust approaches like genetic algorithms. However, even symbiotically learned query optimizers are hampered by the need for vast amounts of training data, slow plan generation during inference and unstable results across various workload conditions. In this paper, we present GenJoin - a novel learned query optimizer that considers the query optimization problem as a generative task and is capable of learning from a random set of subplan hints to produce query plans that outperform classical optimizers. GenJoin is the first learned query optimizer that significantly and consistently outperforms PostgreSQL as well as state-of-the-art methods on two well-known real-world benchmarks across a variety of workloads using rigorous machine learning evaluations.
Pavel Sulimov, Claude Lehmann, Kurt Stockinger
Proc. ACM Manag. Data1
2024 Is Your Learned Query Optimizer Behaving As You Expect? A Machine Learning Perspective
abstract
The current boom of learned query optimizers (LQO) can be explained not only by the general continuous improvement of deep learning (DL) methods but also by the straightforward formulation of a query optimization problem (QOP) as a machine learning (ML) one. The idea is often to replace dynamic programming approaches, widespread for solving QOP, with more powerful methods such as reinforcement learning. However, such a rapid "game change" in the field of QOP could not pass without consequences - other parts of the ML pipeline, except for predictive model development, have large improvement potential. For instance, different LQOs introduce their own restrictions on training data generation from queries, use an arbitrary train/validation approach, and evaluate on a voluntary split of benchmark queries. In this paper, we attempt to standardize the ML pipeline for evaluating LQOs by introducing a new end-to-end benchmarking framework. Additionally, we guide the reader through each data science stage in the ML pipeline and provide novel insights from the machine learning perspective, considering the specifics of QOP. Finally, we perform a rigorous evaluation of existing LQOs, showing that PostgreSQL outperforms these LQOs in almost all experiments depending on the train/test splits.
Claude Lehmann, Pavel Sulimov, Kurt Stockinger
Proc. VLDB Endow.2
2020 Annotation of tandem mass spectrometry data using stochastic neural networks in shotgun proteomics
abstract
MOTIVATION: The discrimination ability of score functions to separate correct from incorrect peptide-spectrum-matches in database-searching-based spectrum identification is hindered by many superfluous peaks belonging to unexpected fragmentation ions or by the lacking peaks of anticipated fragmentation ions. RESULTS: Here, we present a new method, called BoltzMatch, to learn score functions using a particular stochastic neural networks, called restricted Boltzmann machines, in order to enhance their discrimination ability. BoltzMatch learns chemically explainable patterns among peak pairs in the spectrum data, and it can augment peaks depending on their semantic context or even reconstruct lacking peaks of expected ions during its internal scoring mechanism. As a result, BoltzMatch achieved 50% and 33% more annotations on high- and low-resolution MS2 data than XCorr at a 0.1% false discovery rate in our benchmark; conversely, XCorr yielded the same number of spectrum annotations as BoltzMatch, albeit with 4-6 times more errors. In addition, BoltzMatch alone does yield 14% more annotations than Prosit (which runs with Percolator), and BoltzMatch with Percolator yields 32% more annotations than Prosit at 0.1% FDR level in our benchmark. AVAILABILITY AND IMPLEMENTATION: BoltzMatch is freely available at: https://github.com/kfattila/BoltzMatch. CONTACT: [email protected]. SUPPORTING INFORMATION: Supplementary data are available at Bioinformatics online.
Pavel Sulimov, Anastasia Voronkova, Attila Kertész-Farkas
Bioinform.1