EDBT 2026 Demo / reviewers in the wild / expert
Nikita Sorokin
dblp:306/8664
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0002-2437-953XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAGE: Solver-Aligned Guided ExplorationabstractModern code agents achieve strong results on challenging software engineering benchmarks such as SWE-bench, but solving each issue remains expensive: most inference cost is spent not on patch generation, but on repository exploration and search, accounting for up to 56% of tokens in our experiments. We propose SAGE, a modular approach that trains a small searcher model to handle codebase exploration as a tool for a frozen large solver model. We first distill the search trajectories from a strong agent to obtain a compact searcher with comparable retrieval quality. We then apply reinforcement learning to optimize the searcher for usefulness under the solver's fixed interface, using step-level feedback that directly evaluates whether retrieved context is actionable for downstream patching. On SWE-bench Verified, SAGE improves resolve rate while reducing search overhead by 67% and overall cost by 21%, demonstrating a practical path to cheaper, modular software engineering agents. Nikita Sorokin, Ivan Sedykh, Timur Ionov, Valentin Malykh |
SIGIR | 1 |
| 2026 | Hierarchical Embedding Fusion for Retrieval-Augmented Code GenerationabstractRetrieval-augmented code generation commonly conditions a decoder on large retrieved snippets, which couples online cost to repository size and introduces long-context noise. We present Hierarchical Embedding Fusion (HEF), a two-stage repository representation for code completion: (i) an offline cache that compresses repository chunks into a reusable hierarchy of dense vectors using a small fuser model, and (ii) an online interface that maps a small number of retrieved vectors into learned pseudo-tokens consumed by a code generator. This replaces thousands of retrieved tokens with a fixed pseudo-token budget while retaining access to repository-level information. Nikita Sorokin, Ivan Sedykh, Valentin Malykh |
SIGIR | 1 |
| 2025 | Iterative Self-training for Code Generation via Reinforced Re-ranking
Nikita Sorokin, Ivan Sedykh, Valentin Malykh |
ECIR (3) | 1 |
| 2024 | Searching by Code: A New SearchBySnippet Dataset and SnippeR Retrieval Model for Searching by Code SnippetsabstractCode search is an important and well-studied task, but it usually means searching for code by a text query. We argue that using a code snippet (and possibly an error traceback) as a query while looking for bugfixing instructions and code samples is a natural use case not covered by prior art. Moreover, existing datasets use code comments rather than full-text descriptions as text, making them unsuitable for this use case. We present a new SearchBySnippet dataset implementing the search-by-code use case based on StackOverflow data; we show that on SearchBySnippet, existing architectures fall short of a simple BM25 baseline even after fine-tuning. We present a new single encoder model SnippeR that outperforms several strong baselines on SearchBySnippet with a result of 0.451 Recall@10; we propose the SearchBySnippet dataset and SnippeR as a new important benchmark for code search evaluation. Ivan Sedykh, Nikita Sorokin, Dmitry Abulkhanov, Sergey I. Nikolenko, Valentin Malykh |
LREC/COLING | 2 |
| 2023 | LAPCA: Language-Agnostic Pretraining with Cross-Lingual AlignmentabstractData collection and mining is a crucial bottleneck for cross-lingual information retrieval (CLIR). While previous works used machine translation and iterative training, we present a novel approach to cross-lingual pretraining called LAPCA (language-agnostic pretraining with cross-lingual alignment). We train the LAPCA-LM model based on XLM-RoBERTa and łexa that significantly improves cross-lingual knowledge transfer for question answering and sentence retrieval on, e.g., XOR-TyDi and Mr. TyDi datasets, and in the zero-shot cross-lingual scenario performs on par with supervised methods, outperforming many of them on MKQA. Dmitry Abulkhanov, Nikita Sorokin, Sergey I. Nikolenko, Valentin Malykh |
SIGIR | 2 |
| 2022 | Ask Me Anything in Your Native LanguageabstractNikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh |
NAACL-HLT | 1 |