Nikita Sorokin

dblp:306/8664 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0002-2437-953XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SAGE: Solver-Aligned Guided Exploration
abstract
Modern code agents achieve strong results on challenging software engineering benchmarks such as SWE-bench, but solving each issue remains expensive: most inference cost is spent not on patch generation, but on repository exploration and search, accounting for up to 56% of tokens in our experiments. We propose SAGE, a modular approach that trains a small searcher model to handle codebase exploration as a tool for a frozen large solver model. We first distill the search trajectories from a strong agent to obtain a compact searcher with comparable retrieval quality. We then apply reinforcement learning to optimize the searcher for usefulness under the solver's fixed interface, using step-level feedback that directly evaluates whether retrieved context is actionable for downstream patching. On SWE-bench Verified, SAGE improves resolve rate while reducing search overhead by 67% and overall cost by 21%, demonstrating a practical path to cheaper, modular software engineering agents.
Nikita Sorokin, Ivan Sedykh, Timur Ionov, Valentin Malykh
SIGIR1
2026 Hierarchical Embedding Fusion for Retrieval-Augmented Code Generation
abstract
Retrieval-augmented code generation commonly conditions a decoder on large retrieved snippets, which couples online cost to repository size and introduces long-context noise. We present Hierarchical Embedding Fusion (HEF), a two-stage repository representation for code completion: (i) an offline cache that compresses repository chunks into a reusable hierarchy of dense vectors using a small fuser model, and (ii) an online interface that maps a small number of retrieved vectors into learned pseudo-tokens consumed by a code generator. This replaces thousands of retrieved tokens with a fixed pseudo-token budget while retaining access to repository-level information.
Nikita Sorokin, Ivan Sedykh, Valentin Malykh
SIGIR1
2025 Iterative Self-training for Code Generation via Reinforced Re-ranking
Nikita Sorokin, Ivan Sedykh, Valentin Malykh
ECIR (3)1
2024 Searching by Code: A New SearchBySnippet Dataset and SnippeR Retrieval Model for Searching by Code Snippets
abstract
Code search is an important and well-studied task, but it usually means searching for code by a text query. We argue that using a code snippet (and possibly an error traceback) as a query while looking for bugfixing instructions and code samples is a natural use case not covered by prior art. Moreover, existing datasets use code comments rather than full-text descriptions as text, making them unsuitable for this use case. We present a new SearchBySnippet dataset implementing the search-by-code use case based on StackOverflow data; we show that on SearchBySnippet, existing architectures fall short of a simple BM25 baseline even after fine-tuning. We present a new single encoder model SnippeR that outperforms several strong baselines on SearchBySnippet with a result of 0.451 Recall@10; we propose the SearchBySnippet dataset and SnippeR as a new important benchmark for code search evaluation.
Ivan Sedykh, Nikita Sorokin, Dmitry Abulkhanov, Sergey I. Nikolenko, Valentin Malykh
LREC/COLING2
2023 LAPCA: Language-Agnostic Pretraining with Cross-Lingual Alignment
abstract
Data collection and mining is a crucial bottleneck for cross-lingual information retrieval (CLIR). While previous works used machine translation and iterative training, we present a novel approach to cross-lingual pretraining called LAPCA (language-agnostic pretraining with cross-lingual alignment). We train the LAPCA-LM model based on XLM-RoBERTa and łexa that significantly improves cross-lingual knowledge transfer for question answering and sentence retrieval on, e.g., XOR-TyDi and Mr. TyDi datasets, and in the zero-shot cross-lingual scenario performs on par with supervised methods, outperforming many of them on MKQA.
Dmitry Abulkhanov, Nikita Sorokin, Sergey I. Nikolenko, Valentin Malykh
SIGIR2
2022 Ask Me Anything in Your Native Language
abstract
Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh
NAACL-HLT1