Fedor Nikolaev

dblp:166/3090 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-6343-1623ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 84% Knowledge graphs · 16%
Artificial intelligence
2 papers
Generative modeling · 92% Question answering and dialogue systems · 8%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search engines › semantic search
entity retrieval
1.542024
Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024
DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017
Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge Graph · SIGIR 2016
Machine learning › Generative modeling
diffusion model
0.912025
Diffusion on Language Model Encodings for Protein Sequence Generation · ICML 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.912025
Diffusion on Language Model Encodings for Protein Sequence Generation · ICML 2025
Machine learning › Generative modeling › protein design
protein sequence generation
0.912025
Diffusion on Language Model Encodings for Protein Sequence Generation · ICML 2025
Bioinformatics and computational biology
protein design
0.912025
Diffusion on Language Model Encodings for Protein Sequence Generation · ICML 2025
Bioinformatics and computational biology › protein design
protein sequence generation
0.912025
Diffusion on Language Model Encodings for Protein Sequence Generation · ICML 2025
Information retrieval › search engines › semantic search › entity retrieval
entity ranking
0.812024
Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024
Knowledge graphs
knowledge graph querying
0.812024
Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024
Bioinformatics and computational biology › structural biology › protein structure and function
protein stability prediction
0.712023
PROSTATA: a framework for protein stability assessment using transformers · Bioinform. 2023
Information retrieval
retrieval models
0.522016
Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge Graph · SIGIR 2016
Fielded Sequential Dependence Model for Ad-Hoc Entity Retrieval in the Web of Data · SIGIR 2015
Information retrieval
evaluation
0.312017
DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017
Information retrieval › evaluation
relevance judgment
0.312017
DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017
Information retrieval › evaluation
test collection
0.312017
DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017
Information retrieval › search engines › semantic search › entity retrieval
knowledge graph entity retrieval
0.212016
Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge Graph · SIGIR 2016
Natural language and speech › Question answering and dialogue systems
conversational search
0.212024
Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024
Information retrieval › retrieval models › language model
term dependency models
0.212015
Fielded Sequential Dependence Model for Ad-Hoc Entity Retrieval in the Web of Data · SIGIR 2015
Information retrieval › document retrieval
structured document retrieval
0.112015
Fielded Sequential Dependence Model for Ad-Hoc Entity Retrieval in the Web of Data · SIGIR 2015

Methods — techniques the papers use, named apart from their topics

transformer · 2.2latent diffusion · 1.7continuous diffusion · 1.7conditional generation · 1.7neural architecture · 1.5LSTM · 1.5knowledge transfer · 0.7crowdsourcing · 0.3parameterized fielded sequential dependence model · 0.2parameterized fielded full dependence model · 0.2learning to rank · 0.2
YearPublicationVenuePosition
2025 Diffusion on Language Model Encodings for Protein Sequence Generation
abstract
Protein *sequence* design has seen significant advances through discrete diffusion and autoregressive approaches, yet the potential of continuous diffusion remains underexplored. Here, we present *DiMA*, a latent diffusion framework that operates on protein language model representations. Through systematic exploration of architectural choices and diffusion components, we develop a robust methodology that generalizes across multiple protein encoders ranging from 8M to 3B parameters. We demonstrate that our framework achieves consistently high performance across sequence-only (ESM-2, ESMc), dual-decodable (CHEAP), and multimodal (SaProt) representations using the same architecture and training approach. We conduct extensive evaluation of existing methods alongside *DiMA* using multiple metrics across two protein modalities, covering quality, diversity, novelty, and distribution matching of generated proteins. *DiMA* consistently produces novel, high-quality and diverse protein sequences and achieves strong results compared to baselines such as autoregressive, discrete diffusion and flow matching language models. The model demonstrates versatile functionality, supporting conditional generation tasks including protein family-generation, motif scaffolding and infilling, and fold-specific sequence design, despite being trained solely on sequence data. This work provides a universal continuous diffusion framework for protein sequence generation, offering both architectural insights and practical applicability across various protein design scenarios. Code is released at [GitHub](https://github.com/MeshchaninovViacheslav/DiMA).
Viacheslav Meshchaninov, Pavel V. Strashnov, Andrey Shevtsov, Fedor Nikolaev, Nikita Ivanisenko, Olga L. Kardymon, Dmitry P. Vetrov
ICML4
2024 Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph
abstract
This paper introduces a novel information retrieval (IR) task of Conversational Entity Retrieval from a Knowledge Graph (CER-KG), which extends non-conversational entity retrieval from a knowledge graph (KG) to the conversational scenario. The user queries in CER-KG dialog turns may rely on the results of the preceding turns, which are KG entities. Similar to the conversational document IR, CER-KG can be viewed as a sequence of interrelated ranking tasks. To enable future research on CER-KG, we created QBLink-KG, a publicly available benchmark that was adapted from QBLink, a benchmark for text-based conversational reading comprehension of Wikipedia. As an initial approach to CER-KG, we experimented with Transformer- and LSTM-based query encoders in combination with the Neural Architecture for Conversational Entity Retrieval (NACER), our proposed feature-based neural architecture for entity ranking in CER-KG. NACER computes the ranking score of a candidate KG entity by taking into account diverse lexical and semantic matching signals between various KG components in its neighborhood, such as entities, categories, and literals, as well as entities in the results of the preceding turns in dialog history. The reported experimental results reveal the key challenges of CER-KG along with the possible directions for new approaches to this task.
Mona Zamiri, Yao Qiang, Fedor Nikolaev, Dongxiao Zhu, Alexander Kotov 0001
WWW3
2023 PROSTATA: a framework for protein stability assessment using transformers
abstract
MOTIVATION: Accurate prediction of change in protein stability due to point mutations is an attractive goal that remains unachieved. Despite the high interest in this area, little consideration has been given to the transformer architecture, which is dominant in many fields of machine learning. RESULTS: In this work, we introduce PROSTATA, a predictive model built in a knowledge-transfer fashion on a new curated dataset. PROSTATA demonstrates advantage over existing solutions based on neural networks. We show that the large improvement margin is due to both the architecture of the model and the quality of the new training dataset. This work opens up opportunities to develop new lightweight and accurate models for protein stability assessment. AVAILABILITY AND IMPLEMENTATION: PROSTATA is available at https://github.com/AIRI-Institute/PROSTATA and https://prostata.airi.net.
Dmitry Umerenkov, Fedor Nikolaev, Tatiana Shashkova, Pavel V. Strashnov, Maria Sindeeva, Andrey Shevtsov, Nikita Ivanisenko, Olga L. Kardymon
Bioinform.2
2020 Joint Word and Entity Embeddings for Entity Retrieval from a Knowledge Graph
Fedor Nikolaev, Alexander Kotov 0001
ECIR (1)1
2018 Attentive Neural Architecture for Ad-hoc Structured Document Retrieval
abstract
The problem of ad-hoc structured document retrieval arises in many information access scenarios, from Web to product search. Yet neither deep neural networks, which have been successfully applied to ad-hoc information retrieval and Web search, nor the attention mechanism, which has been shown to significantly improve the performance of deep neural networks on natural language processing tasks, have been explored in the context of this problem. In this paper, we propose a deep neural architecture for ad-hoc structured document retrieval, which utilizes attention mechanism to determine important phrases in keyword queries as well as the relative importance of matching those phrases in different fields of structured documents. Experimental evaluation on publicly available collections for Web document, product and entity retrieval from knowledge graphs indicates superior retrieval accuracy of the proposed neural architecture relative to both state-of-the-art neural architectures for ad-hoc document retrieval and probabilistic models for ad-hoc structured document retrieval.
Saeid Balaneshinkordan, Alexander Kotov 0001, Fedor Nikolaev
CIKM3
2017 DBpedia-Entity v2: A Test Collection for Entity Search
abstract
The DBpedia-entity collection has been used as a standard test collection for entity search in recent years. We develop and release a new version of this test collection, DBpedia-Entity v2, which uses a more recent DBpedia dump and a unified candidate result pool from the same set of retrieval models. Relevance judgments are also collected in a uniform way, using the same group of crowdsourcing workers, following the same assessment guidelines. The result is an up-to-date and consistent test collection.To facilitate further research, we also provide details about the pre-processing and indexing steps, and include baseline results from both classical and recently developed entity search methods.
Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov 0001, Jamie Callan
SIGIR2
2016 Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge Graph
abstract
Accurate projection of terms in free-text queries onto structured entity representations is one of the fundamental problems in entity retrieval from knowledge graphs. In this paper, we demonstrate that existing retrieval models for ad-hoc structured and unstructured document retrieval fall short of addressing this problem, due to their rigid assumptions. According to these assumptions, either all query concepts of the same type (unigrams and bigrams) are projected onto the fields of entity representations with identical weights or such projection is determined based only on one simple statistic, which makes it sensitive to data sparsity. To address this issue, we propose the Parametrized Fielded Sequential Dependence Model (PFSDM) and the Parametrized Fielded Full Dependence Model (PFFDM), two novel models for entity retrieval from knowledge graphs, which infer the user's intent behind each individual query concept by dynamically estimating its projection onto the fields of structured entity representations based on a small number of statistical and linguistic features. Experimental results obtained on several publicly available benchmarks indicate that PFSDM and PFFDM consistently outperform state-of-the-art retrieval models for the task of entity retrieval from knowledge graph.
Fedor Nikolaev, Alexander Kotov 0001, Nikita Zhiltsov
SIGIR1
2015 Fielded Sequential Dependence Model for Ad-Hoc Entity Retrieval in the Web of Data
abstract
Previously proposed approaches to ad-hoc entity retrieval in the Web of Data (ERWD) used multi-fielded representation of entities and relied on standard unigram bag-of-words retrieval models. Although retrieval models incorporating term dependencies have been shown to be significantly more effective than the unigram bag-of-words ones for ad hoc document retrieval, it is not known whether accounting for term dependencies can improve retrieval from the Web of Data. In this work, we propose a novel retrieval model that incorporates term dependencies into structured document retrieval and apply it to the task of ERWD. In the proposed model, the document field weights and the relative importance of unigrams and bigrams are optimized with respect to the target retrieval metric using a learning-to-rank method. Experiments on a publicly available benchmark indicate significant improvement of the accuracy of retrieval results by the proposed model over state-of-the-art retrieval models for ERWD.
Nikita Zhiltsov, Alexander Kotov 0001, Fedor Nikolaev
SIGIR3