VLDB 2026 Research / reviewers in the wild / expert
Zakhar Ostrovsky
dblp:411/5578 · also Zakhar Ostrovskyi
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0003-4644-3587ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › structural bioinformatics › ligand binding site analysis
binding pocket prediction |
0.9 | 1 | 2025 | Leveraging large language models for literature-driven prioritization of protein binding pockets · Bioinform. 2025 |
Bioinformatics and computational biology
drug discovery |
0.3 | 1 | 2025 | Leveraging large language models for literature-driven prioritization of protein binding pockets · Bioinform. 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 0.9geometric pocket detection · 0.9fpocket · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging large language models for literature-driven prioritization of protein binding pocketsabstractMOTIVATION: Accurately identifying and prioritizing protein binding pockets is a foundational element of small-molecule drug discovery. Defining these known pockets currently relies on a laborious manual process of extracting key residue data from selected publications, reconciling inconsistent terminology, and independently computing volumetric representations. This manual curation to ensure biological relevance is time-consuming, error-prone, and represents a major bottleneck for efficient, high-throughput drug discovery. RESULTS: We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with large language models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLM's performance in paper relevance assessment and pocket extraction. AVAILABILITY AND IMPLEMENTATION: The developed benchmark dataset and methodology are freely available at the GitHub repository (https://github.com/receptor-ai/LLM-benchmark-dataset) and Zenodo (DOI: 10.5281/zenodo.15798647). Roman Stratiichuk, Mykola Melnychenko, Ihor Koleiev, Taras Voitsitskyi, Vladyslav Husak, Nazar Shevchuk, Zakhar Ostrovsky, Volodymyr G. Bdzhola, Semen O. Yesylevskyy, Serhii Starosyla, Alan Nafiiev |
Bioinform. | 7 |