VLDB 2026 Research / reviewers in the wild / expert
Andrey Sakhovskiy
dblp:262/6042
· DBLP profile ↗
6ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0000-0003-2762-2910ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (2 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Millennium of Arabic Manuscripts in Three Styles: A Line-Level OCR Benchmark for Naskh, Taliq, and Nastaliq
Maxim Novopoltsev, Ruslan Murtazin, Andrey Sakhovskiy, Emilia Bojarskaja, Vladimir Kokh, Ivan Ulitin, Botirjon Abdullayev, Khamidulla Aminov, Masudkhon Ismoilov, Semen A. Budennyy |
ICDAR (2) | 3 |
| 2026 | LLM Agents Factory: Retrieval of Domain-Specific LLM AgentsabstractLarge language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent design for each user request. To address this, we present LLM Agents Factory, a retrieval-based framework that constructs domain-specific and Wikipedia-grounded agents on demand using a base of over 20K predetermined agent profiles. Our framework supports two modes: (1) agent profile retrieval via semantic search and (2) distillation into a compact model fine-tuned for direct agent generation. Experiments on MMLU, BIG-bench, and BIG-bench Hard in a single-agent scenario demonstrate that our retrieval-based agent construction surpasses non-agent baselines in accuracy while matching AutoGen generation quality with a 120B backbone at a substantially lower inference cost. Our work reveals that retrieval from a structured agent repository provides a cost-efficient, accurate, and controllable alternative to dynamic agent generation, responding to the strict demands of industrial applications. We provide the implementation code and the agent base in https://huggingface.co/frontier-ai/llm-agent-factory. Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen A. Budennyy |
SIGIR | 3 |
| 2026 | FactOWL: A Cost-Efficient Tool for Long-Form Factuality Evaluation
Andrey Sakhovskiy, Nikita Sushko, Maria Marina, Vasily Konovalov, Elena Tutubalina, Alexander Panchenko, Pavel Braslavski 0001 |
SIGIR | 1 |
| 2025 | BioASQ at CLEF2025: The Thirteenth Edition of the Large-Scale Biomedical Semantic Indexing and Question Answering Challenge
Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodríguez-Ortega, Natalia V. Loukachevitch, Andrey Sakhovskiy, Elena Tutubalina, Grigorios Tsoumakas, George Giannakoulas, Alexandra Bekiaridou, Athanasios Samaras, Giorgio Maria Di Nunzio, Nicola Ferro 0001, Stefano Marchesin 0001, Laura Menotti, Gianmaria Silvello, Georgios Paliouras |
ECIR (5) | 7 |
| 2025 | ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs
Mikhail Salnikov, Andrey Sakhovskiy, Irina Nikishina, Aida Usmanova, Angelie Kraft, Cedric Möller, Debayan Banerjee, Junbo Huang, Longquan Jiang 0001, Rana Abdullah, Xi Yan 0001, Elena Tutubalina, Ricardo Usbeck, Alexander Panchenko |
NLDB (1) | 2 |
| 2025 | BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model AlignmentabstractIn recent years, there has been substantial progress in using pretrained Language Models (LMs) on a range of tasks aimed at improving the understanding of biomedical texts. Nonetheless, existing biomedical LLMs show limited comprehension of complex, domain-specific concept structures and the factual information encoded in biomedical Knowledge Graphs (KGs). In this work, we propose BALI (Biomedical Knowledge Graph and Language Model Ali gnment), a novel joint LM and KG pre-training method that augments an LM with external knowledge by the simultaneous learning of a dedicated KG encoder and aligning the representations of both the LM and the graph. For a given textual sequence, we link biomedical concept mentions to the Unified Medical Language System (UMLS) KG and utilize local KG subgraphs as cross-modal positive samples for these mentions. Our empirical findings indicate that implementing our method on several leading biomedical LMs, such as PubMedBERT and BioLinkBERT, improves their performance on a range of language understanding tasks and the quality of entity representations, even with minimal pre-training on a small alignment dataset sourced from PubMed scientific abstracts. Andrey Sakhovskiy, Elena Tutubalina |
SIGIR | 1 |