EDBT 2026 Demo / reviewers in the wild / expert
Mikhail Burtsev 0001
dblp:85/5425 · also Mikhail S. Burtsev
· DBLP profile ↗
20ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-1614-1695ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Wikontic: A Tool for Building Knowledge Graphs from Text Aligned with the Wikidata OntologyabstractKnowledge Graphs (KGs) provide structured, verifiable representations that ground facts and supply large language models (LLMs) with reliable real-world information. Building high-quality KGs from open-domain text remains difficult due to redundancy, inconsistency, and lack of ontology grounding. We present Wikontic, a pipeline that extracts triples from text with LLMs and refines them through ontology-based typing, schema validation, and entity deduplication, yielding compact and coherent graphs. Unlike prior frameworks that lack ontology grounding or perform only partial deduplication, Wikontic uniquely integrates entity canonicalization, alias tracking, and automatic enforcement of Wikidata’s ontology, enabling robust schema-aware construction without manual schema design. Its web interface lets users upload text, visualize graphs, and perform multi-hop question answering. By combining LLM flexibility with Wikidata’s ontological rigor, Wikontic transforms ambiguous text into structured, interpretable, and actionable knowledge. Alla Chepurova, Aydar Bulatov, Mikhail Burtsev 0001, Yuri Kuratov |
AAAI | 3 |
| 2025 | Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space CapacityabstractA range of recent works addresses the problem of compression of sequence of tokens into a shorter sequence of real-valued vectors to be used as inputs instead of token embeddings or key-value cache. These approaches are focused on reduction of the amount of compute in existing language models rather than minimization of number of bits needed to store text. Despite relying on powerful models as encoders, the maximum attainable lossless compression ratio is typically not higher than x10. This fact is highly intriguing because, in theory, the maximum information capacity of large real-valued vectors is far beyond the presented rates even for 16-bit precision and a modest vector size. In this work, we explore the limits of compression by replacing the encoder with a per-sample optimization procedure. We show that vectors with compression ratios up to x1500 exist, which highlights two orders of magnitude gap between existing and practically attainable solutions. Furthermore, we empirically show that the compression limits are determined not by the length of the input but by the amount of uncertainty to be reduced, namely, the cross-entropy loss on this sequence without any conditioning. The obtained limits highlight the substantial gap between the theoretical capacity of input embeddings and their practical utilization, suggesting significant room for optimization in model design. Yuri Kuratov, Mikhail Arkhipov, Aydar Bulatov, Mikhail Burtsev 0001 |
ACL (1) | 4 |
| 2025 | AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM AgentsabstractAdvancements in the capabilities of Large Language Models (LLMs) have created a promising foundation for developing autonomous agents. With the right tools, these agents could learn to solve tasks in new environments by accumulating and updating their knowledge. Current LLM-based agents process past experiences using a full history of observations, summarization, retrieval augmentation. However, these unstructured memory representations do not facilitate the reasoning and planning essential for complex decision-making. In our study, we introduce AriGraph, a novel method wherein the agent constructs and updates a memory graph that integrates semantic and episodic memories while exploring the environment. We demonstrate that our Ariadne LLM agent, consisting of the proposed memory architecture augmented with planning and decision-making, effectively handles complex tasks within interactive text game environments difficult even for human players. Results show that our approach markedly outperforms other established memory methods and strong RL baselines in a range of problems of varying complexity. Additionally, AriGraph demonstrates competitive performance compared to dedicated knowledge graph-based methods in static multi-hop question-answering. Petr Anokhin, Nikita Semenov, Artyom Y. Sorokin, Dmitry Evseev, Andrey Kravchenko, Mikhail Burtsev 0001, Evgeny Burnaev |
IJCAI | 6 |
| 2024 | Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent MemoryabstractA major limitation for the broader scope of problems solvable by transformers is the quadratic scaling of computational complexity with input size. In this study, we investigate the recurrent memory augmentation of pre-trained transformer models to extend input context length while linearly scaling compute. Our approach demonstrates the capability to store information in memory for sequences of up to an unprecedented two million tokens while maintaining high retrieval accuracy. Experiments with language modeling tasks show perplexity improvement as the number of processed input segments increases. These results underscore the effectiveness of our method, which has significant potential to enhance long-term dependency handling in natural language understanding and generation tasks, as well as enable large-scale context processing for memory-intensive applications. Aydar Bulatov, Yuri Kuratov, Yermek Kapushev, Mikhail Burtsev 0001 |
AAAI | 4 |
| 2024 | BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-HaystackabstractIn recent years, the input context sizes of large language models (LLMs) have increased dramatically. However, existing evaluation methods have not kept pace, failing to comprehensively assess the efficiency of models in handling long contexts. To bridge this gap, we introduce the BABILong benchmark, designed to test language models' ability to reason across facts distributed in extremely long documents. BABILong includes a diverse set of 20 reasoning tasks, including fact chaining, simple induction, deduction, counting, and handling lists/sets. These tasks are challenging on their own, and even more demanding when the required facts are scattered across long natural text. Our evaluations show that popular LLMs effectively utilize only 10-20% of the context and their performance declines sharply with increased reasoning complexity. Among alternatives to in-context reasoning, Retrieval-Augmented Generation methods achieve a modest 60% accuracy on single-fact question answering, independent of context length. Among context extension methods, the highest performance is demonstrated by recurrent memory transformers after fine-tuning, enabling the processing of lengths up to 50 million tokens. The BABILong benchmark is extendable to any length to support the evaluation of new upcoming models with increased capabilities, and we provide splits up to 10 million token lengths. Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Y. Sorokin, Mikhail Burtsev 0001 |
NeurIPS | 7 |
| 2024 | Associative Learning and Active InferenceabstractAssociative learning is a behavioral phenomenon in which individuals develop connections between stimuli or events based on their co-occurrence. Initially studied by Pavlov in his conditioning experiments, the fundamental principles of learning have been expanded on through the discovery of a wide range of learning phenomena. Computational models have been developed based on the concept of minimizing reward prediction errors. The Rescorla-Wagner model, in particular, is a well-known model that has greatly influenced the field of reinforcement learning. However, the simplicity of these models restricts their ability to fully explain the diverse range of behavioral phenomena associated with learning. In this study, we adopt the free energy principle, which suggests that living systems strive to minimize surprise or uncertainty under their internal models of the world. We consider the learning process as the minimization of free energy and investigate its relationship with the Rescorla-Wagner model, focusing on the informational aspects of learning, different types of surprise, and prediction errors based on beliefs and values. Furthermore, we explore how well-known behavioral phenomena such as blocking, overshadowing, and latent inhibition can be modeled within the active inference framework. We accomplish this by using the informational and novelty aspects of attention, which share similar ideas proposed by seemingly contradictory models such as Mackintosh and Pearce-Hall models. Thus, we demonstrate that the free energy principle, as a theoretical framework derived from first principles, can integrate the ideas and models of associative learning proposed based on empirical experiments and serve as a framework for a better understanding of the computational processes behind associative learning in the brain. Petr Anokhin, Artyom Y. Sorokin, Mikhail Burtsev 0001, Karl J. Friston |
Neural Comput. | 3 |
| 2023 | Hybrid Uncertainty Quantification for Selective Text Classification in Ambiguous TasksabstractArtem Vazhentsev, Gleb Kuzmin, Akim Tsvigun, Alexander Panchenko, Maxim Panov, Mikhail Burtsev, Artem Shelmanov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Artem Vazhentsev, Gleb Kuzmin, Akim Tsvigun, Alexander Panchenko, Maxim Panov, Mikhail Burtsev 0001, Artem Shelmanov |
ACL (1) | 6 |
| 2023 | Uncertainty Guided Global Memory Improves Multi-Hop Question AnsweringabstractTransformers have become the gold standard for many natural language processing tasks and, in particular, for multi-hop question answering (MHQA).This task includes processing a long document and reasoning over the multiple parts of it.The landscape of MHQA approaches can be classified into two primary categories.The first group focuses on extracting supporting evidence, thereby constraining the QA model's context to predicted facts.Conversely, the second group relies on the attention mechanism of the long input encoding model to facilitate multi-hop reasoning.However, attention-based token representations lack explicit global contextual information to connect reasoning steps.To address these issues, we propose GEMFormer, a two-stage method that first collects relevant information over the entire document to the memory and then combines it with local context to solve the task 1 .Our experimental results show that fine-tuning a pre-trained model with memory-augmented input, including the most certain global elements, improves the model's performance on three MHQA datasets compared to the baseline.We also found that the global explicit memory contains information from supporting facts required for the correct answer. Alsu Sagirova, Mikhail Burtsev 0001 |
EMNLP | 2 |
| 2022 | Uncertainty Estimation of Transformer Predictions for Misclassification DetectionabstractArtem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev, Manvel Avetisian, Leonid Zhukov. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Artem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev 0001, Manvel Avetisian, Leonid Zhukov |
ACL (1) | 10 |
| 2022 | Attention Understands Semantic RelationsabstractToday, natural language processing heavily relies on pre-trained large language models. Even though such models are criticized for the poor interpretability, they still yield state-of-the-art solutions for a wide set of very different tasks. While lots of probing studies have been conducted to measure the models’ awareness of grammatical knowledge, semantic probing is less popular. In this work, we introduce the probing pipeline to study the representedness of semantic relations in transformer language models. We show that in this task, attention scores are nearly as expressive as the layers’ output activations, despite their lesser ability to represent surface cues. This supports the hypothesis that attention mechanisms are focusing not only on the syntactic relational information but also on the semantic one. Anastasia Chizhikova, Sanzhar Murzakhmetov, Oleg Serikov, Tatiana Shavrina, Mikhail Burtsev 0001 |
LREC | 5 |
| 2022 | Recurrent Memory TransformerabstractTransformer-based models show their effectiveness across multiple domains and tasks. The self-attention allows to combine information from all sequence elements into context-aware representations. However, global and local information has to be stored mostly in the same element-wise representations. Moreover, the length of an input sequence is limited by quadratic computational complexity of self-attention. In this work, we propose and study a memory-augmented segment-level recurrent Transformer (RMT). Memory allows to store and process local and global information as well as to pass information between segments of the long sequence with the help of recurrence. We implement a memory mechanism with no changes to Transformer model by adding special memory tokens to the input or output sequence. Then the model is trained to control both memory operations and sequence representations processing. Results of experiments show that RMT performs on par with the Transformer-XL on language modeling for smaller memory sizes and outperforms it for tasks that require longer sequence processing. We show that adding memory tokens to Tr-XL is able to improve its performance. This makes Recurrent Memory Transformer a promising architecture for applications that require learning of long-term dependencies and general purpose in memory processing, such as algorithmic tasks and reasoning. Aydar Bulatov, Yuri Kuratov, Mikhail Burtsev 0001 |
NeurIPS | 3 |
| 2022 | Explain My Surprise: Learning Efficient Long-Term Memory by predicting uncertain outcomesabstractIn many sequential tasks, a model needs to remember relevant events from the distant past to make correct predictions. Unfortunately, a straightforward application of gradient based training requires intermediate computations to be stored for every element of a sequence. This requires to store prohibitively large intermediate data if a sequence consists of thousands or even millions elements, and as a result, makes learning of very long-term dependencies infeasible. However, the majority of sequence elements can usually be predicted by taking into account only temporally local information. On the other hand, predictions affected by long-term dependencies are sparse and characterized by high uncertainty given only local information. We propose \texttt{MemUP}, a new training method that allows to learn long-term dependencies without backpropagating gradients through the whole sequence at a time. This method can potentially be applied to any recurrent architecture. LSTM network trained with \texttt{MemUP} performs better or comparable to baselines while requiring to store less intermediate data. Artyom Y. Sorokin, Nazar Buzun, Leonid Pugachev, Mikhail Burtsev 0001 |
NeurIPS | 4 |
| 2022 | A review of neural architecture search
Dilyara Baymurzina, Eugene A. Golikov, Mikhail Burtsev 0001 |
Neurocomputing | 3 |
| 2021 | Building and Evaluating Open-Domain Dialogue Corpora with Clarifying QuestionsabstractEnabling open-domain dialogue systems to ask clarifying questions when appropriate is an important direction for improving the quality of the system response.Namely, for cases when a user request is not specific enough for a conversation system to provide an answer right away, it is desirable to ask a clarifying question to increase the chances of retrieving a satisfying answer.To address the problem of 'asking clarifying questions in opendomain dialogues': (1) we collect and release a new dataset focused on open-domain singleand multi-turn conversations, (2) we benchmark several state-of-the-art neural baselines, and (3) we propose a pipeline consisting of offline and online steps for evaluating the quality of clarifying questions in various dialogues.These contributions are suitable as a foundation for further research. Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton 0001, Mikhail Burtsev 0001 |
EMNLP (1) | 5 |
| 2014 | Evolution, Development and Learning with Predictor Neural NetworksabstractStudy of mechanisms, that can make possible effective learning of artificial systems in complex environments, is one of the key issues in the adaptive systems research. In this paper we make an attempt to implement and test a number of ideas motivated by brain theory. Proposed model integrates evolutionary, developmental and learning phases. The main concept of this paper is the notion of predictor neural network which provide distributed evaluation of the effectiveness of goal-directed behavior on the neuronal level. We also propose learning mechanism based on gradually inclusion of new neuronal functional groups in case when the existing behavior fails to deliver adaptive result. We performed basic computational study of the model to investigate some of its’ core properties such as evolution of innate and learned behavior and dynamics of the learning process. Konstantin Lakhman, Mikhail Burtsev 0001 |
ALIFE | 2 |
| 2013 | Neuroevolution results in emergence of short-term memory in multi-goal environmentabstractAnimals behave adaptively in environments with multiple competing goals. Understanding of mechanisms underlying such goal-directed behavior remains a challenge for neuroscience as well as for adaptive system and machine learning research. To address this problem we developed an evolutionary model of adaptive behavior in a multi-goal stochastic environment. The proposed neuroevolutionary algorithm is based on neuron's duplication as a basic mechanism of agent's recurrent neural network development. Results of simulations demonstrate that in the course of evolution agents acquire the ability to store the short-term memory and use it in behavior with alternative actions. We found that evolution discovered two mechanisms for short-term memory. The first mechanism is integration of sensory signals and ongoing internal neural activity, resulting in emergence of cell groups specialized on alternative actions. And the second mechanism is slow neurodynamical process that makes possible to encode the previous behavioral choice. Konstantin Lakhman, Mikhail Burtsev 0001 |
GECCO | 2 |
| 2008 | Basic Principles of Adaptive Learning through Variation and Selection
Mikhail Burtsev 0001 |
ALIFE | 1 |
| 2008 | Learning drives the accumulation of adaptive complexity in simulated evolution
Mikhail Burtsev 0001, Konstantin V. Anokhin, Patrick Bateson |
ALIFE | 1 |
| 2004 | Theory of functional systems, adaptive critics and neural networksabstractWe propose a general scheme of intelligent adaptive control system based on the Petr K. Anokhin's theory of functional systems. This scheme is aimed at controlling adaptive purposeful behavior of an animat (a simulated animal) that has several natural needs (e.g., energy replenishment, reproduction). The control system consists of a set of hierarchically linked functional systems and enables predictive and goal-directed behavior. Each functional system includes a neural network based adaptive critic design. We also discuss schemes of prognosis, decision making, action selection and learning that occur in the functional systems and in the whole control system of the animat. Vladimir G. Red'ko, Danil V. Prokhorov, Mikhail Burtsev 0001 |
IJCNN | 3 |
| 2004 | Tracking the Trajectories of EvolutionabstractThis article proposes a method of visualizing and measuring evolution in artificial life simulations. The evolving population of agents is treated as a dynamical system. The proposed method is inspired by the notion of trajectory. The article provides examples of tracking of trajectories of evolutionary systems in the spaces of genotypes, strategies, and some global characteristics. Visualization similar to a bifurcation diagram is used to represent results of a series of simulations. Mikhail Burtsev 0001 |
Artif. Life | 1 |