EDBT 2026 Demo / reviewers in the wild / expert
Aleksandr I. Panov
dblp:177/9975 · also Aleksandr Panov 0001, Alexander I. Panov
· DBLP profile ↗
34ranked-venue papers
0as first author
33since 2021 · last 2026
0000-0002-9747-3837ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Systems, architecture and hardware · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAMAR: Continuous Actions Multi-Agent RoutingabstractMulti-agent reinforcement learning (MARL) is a powerful paradigm for solving cooperative and competitive decision-making problems. While many MARL benchmarks have been proposed, few combine continuous state and action spaces with challenging coordination and planning tasks. We introduce CAMAR, a new MARL benchmark designed explicitly for multi-agent pathfinding in environments with continuous actions. CAMAR supports cooperative and competitive interactions between agents and runs efficiently at up to 100,000 environment steps per second. We also propose a three-tier evaluation protocol to better track algorithmic progress and enable deeper analysis of performance. In addition, CAMAR allows the integration of classical planning methods such as RRT and RRT* into MARL pipelines. We use them as standalone baselines and combine RRT* with popular MARL algorithms to create hybrid approaches. We provide a suite of test scenarios and benchmarking tools to ensure reproducibility and fair comparison. Experiments show that CAMAR presents a challenging and realistic testbed for the MARL community. Artem Pshenitsyn, Aleksandr I. Panov, Aleksey Skrynnik |
AAAI | 2 |
| 2026 | Scene graph-driven reasoning for action planning of humanoid robot
Dmitry A. Yudin, Alexander Lazarev, Eva Bakaeva, Angelika Kochetkova, Alexey K. Kovalev, Aleksandr I. Panov |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Say it better: RL-based prompt tuning for enhancing open-vocabulary recognition
Mikhail Avshalumov, Zoya Volovikova, Dmitry A. Yudin, Aleksandr I. Panov |
Neurocomputing | 4 |
| 2025 | MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at ScaleabstractMulti-agent pathfinding (MAPF) is a problem that generally requires finding collision-free paths for multiple agents in a shared environment. Solving MAPF optimally, even under restrictive assumptions, is NP-hard, yet efficient solutions for this problem are critical for numerous applications, such as automated warehouses and transportation systems. Recently, learning-based approaches to MAPF have gained attention, particularly those leveraging deep reinforcement learning. Typically, such learning-based MAPF solvers are augmented with additional components like single-agent planning or communication. Orthogonally, in this work we rely solely on imitation learning that leverages a large dataset of expert MAPF solutions and transformer-based neural network to create a foundation model for MAPF called MAPF-GPT. The latter is capable of generating actions without additional heuristics or communication. MAPF-GPT demonstrates zero-shot learning abilities when solving the MAPF problems that are not present in the training dataset. We show that MAPF-GPT notably outperforms the current best-performing learnable MAPF solvers on a diverse range of problem instances and is computationally efficient during inference. Anton Andreychuk, Konstantin S. Yakovlev, Aleksandr I. Panov, Aleksey Skrynnik |
AAAI | 3 |
| 2025 | AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentabstractAs a part of an embodied agent, Large Language Models (LLMs) are typically used for behavior planning given natural language instructions from the user.However, dealing with ambiguous instructions in real-world environments remains a challenge for LLMs.Various methods for task ambiguity detection have been proposed.However, it is difficult to compare them because they are tested on different datasets and there is no universal benchmark.For this reason, we propose AmbiK (Ambiguous Tasks in Kitchen Environment), the fully textual dataset of ambiguous instructions addressed to a robot in a kitchen environment.AmbiK was collected with the assistance of LLMs and is human-validated.It comprises 1000 pairs of ambiguous tasks and their unambiguous counterparts, categorized by ambiguity type (Human Preferences, Common Sense Knowledge, Safety), with environment descriptions, clarifying questions and answers, user intents, and task plans, for a total of 2000 tasks.We hope that AmbiK will enable researchers to perform a unified comparison of ambiguity detection methods.AmbiK is available at https: //github.com/cog-model/AmbiK-dataset. Anastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev, Aleksandr I. Panov |
ACL (1) | 5 |
| 2025 | CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended WorldabstractFollowing instructions in real-world conditions requires a capability to adapt to the world's volatility and entanglement: the environment is dynamic and unpredictable, instructions can be linguistically complex with diverse vocabulary, and the number of possible goals an agent may encounter is vast.Despite extensive research in this area, most studies are conducted in static environments with simple instructions and a limited vocabulary, making it difficult to assess agent performance in more diverse and challenging settings.To address this gap, we introduce CrafText, a benchmark for evaluating instruction following in a multimodal environment with diverse instructions and dynamic interactions.CrafText includes 3,924 instructions with 3,423 unique words, covering Localization, Conditional, Building, and Achievement tasks.Additionally, we propose an evaluation protocol that measures an agent's ability to generalize to novel instruction formulations and dynamically evolving task configurations, providing a rigorous test of both linguistic understanding and adaptive decisionmaking. Zoya Volovikova, Gregory Gorbov, Petr Kuderov, Aleksandr I. Panov, Aleksey Skrynnik |
ACL (1) | 4 |
| 2025 | Safe Planning and Policy Optimization via World Model LearningabstractReinforcement Learning (RL) applications in real-world scenarios must prioritize safety and reliability, which impose strict constraints on agent behavior. Model-based RL leverages predictive world models for action planning and policy optimization, but inherent model inaccuracies can lead to catastrophic failures in safety-critical settings. We propose a novel model-based RL framework that jointly optimizes task performance and safety. To address world model errors, our method incorporates an adaptive mechanism that dynamically switches between model-based planning and direct policy execution. We resolve the objective mismatch problem of traditional model-based approaches using an implicit world model. Furthermore, our framework employs dynamic safety thresholds that adapt to the agent’s evolving capabilities, consistently selecting actions that surpass safe policy suggestions in both performance and safety. Experiments demonstrate significant improvements over non-adaptive methods, showing that our approach optimizes safety and performance simultaneously rather than merely meeting minimum safety requirements. The proposed framework achieves robust performance on diverse safety-critical continuous control tasks, outperforming existing methods. Artem Latyshev, Gregory Gorbov, Aleksandr I. Panov |
ECAI | 3 |
| 2025 | Accelerating Transformers in Online RLabstractThe appearance of transformer-based models in Reinforcement Learning (RL) has expanded the horizons of possibilities in robotics tasks, but it has simultaneously brought a wide range of challenges during its implementation, especially in model-free online RL. Some of the existing learning algorithms cannot be easily implemented with transformer-based models due to the instability of the latter. In this paper, we propose a method that uses the Accelerator policy as a transformer’s trainer. The Accelerator, a simpler and more stable model, interacts with the environment independently while simultaneously training the transformer through behavior cloning during the first stage of the proposed algorithm. In the second stage, the pretrained transformer starts to interact with the environment in a fully online setting. As a result, this model-free algorithm accelerates the transformer in terms of its performance and helps it to train online in a more stable and faster way. By conducting experiments on both state-based and image-based ManiSkill environments, as well as on MuJoCo tasks in MDP and POMDP settings, we show that applying our algorithm not only enables stable training of transformers but also reduces training time on image-based environments by up to a factor of two. Moreover, it decreases the required replay buffer size in off-policy methods to 10–20 thousand, which significantly lowers the overall computational demands. Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov |
ECAI | 3 |
| 2025 | Object-Centric Dreamer
Leonid Ugadiarov, Vitaliy Vorobyov, Aleksandr I. Panov |
ICANN (1) | 3 |
| 2025 | Learning Successor Features with Distributed Hebbian Temporal MemoryabstractThis paper presents a novel approach to address the challenge of online sequence learning for decision making under uncertainty in non-stationary, partially observable environments. The proposed algorithm, Distributed Hebbian Temporal Memory (DHTM), is based on the factor graph formalism and a multi-component neuron model. DHTM aims to capture sequential data relationships and make cumulative predictions about future observations, forming Successor Features (SFs). Inspired by neurophysiological models of the neocortex, the algorithm uses distributed representations, sparse transition matrices, and local Hebbian-like learning rules to overcome the instability and slow learning of traditional temporal memory algorithms such as RNN and HMM. Experimental results show that DHTM outperforms LSTM, RWKV and a biologically inspired HMM-like algorithm, CSCG, on non-stationary data sets. Our results suggest that DHTM is a promising approach to address the challenges of online sequence learning and planning in dynamic environments. Evgenii Dzhivelikian, Petr Kuderov, Aleksandr I. Panov |
ICLR | 3 |
| 2025 | POGEMA: A Benchmark Platform for Cooperative Multi-Agent PathfindingabstractMulti-agent reinforcement learning (MARL) has recently excelled in solving challenging cooperative and competitive multi-agent problems in various environments, typically involving a small number of agents and full observability. Moreover, a range of crucial robotics-related tasks, such as multi-robot pathfinding, which have traditionally been approached with classical non-learnable methods (e.g., heuristic search), are now being suggested for solution using learning-based or hybrid methods. However, in this domain, it remains difficult, if not impossible, to conduct a fair comparison between classical, learning-based, and hybrid approaches due to the lack of a unified framework that supports both learning and evaluation. To address this, we introduce POGEMA, a comprehensive set of tools that includes a fast environment for learning, a problem instance generator, a collection of predefined problem instances, a visualization toolkit, and a benchmarking tool for automated evaluation. We also introduce and define an evaluation protocol that specifies a range of domain-related metrics, computed based on primary evaluation indicators (such as success rate and path length), enabling a fair multi-fold comparison. The results of this comparison, which involves a variety of state-of-the-art MARL, search-based, and hybrid methods, are presented. Aleksey Skrynnik, Anton Andreychuk, Anatolii Borzilov, Alexander Chernyavskiy, Konstantin S. Yakovlev, Aleksandr I. Panov |
ICLR | 6 |
| 2025 | Advancing Learnable Multi-Agent Pathfinding Solvers with Active Fine-TuningabstractMulti-agent pathfinding (MAPF) is a common abstraction of multi-robot trajectory planning problems, where multiple homogeneous robots simultaneously move in the shared environment. While solving MAPF optimally has been proven to be NP-hard, scalable, and efficient, solvers are vital for real-world applications like logistics, search-and-rescue, etc. To this end, decentralized suboptimal MAPF solvers that leverage machine learning have come on stage. Building on the success of the recently introduced MAPF-GPT, a pure imitation learning solver, we introduce MAPF-GPT-DDG. This novel approach effectively fine-tunes the pre-trained MAPF model using centralized expert data. Leveraging a novel delta-data generation mechanism, MAPF-GPT-DDG accelerates training while significantly improving performance at test time. Our experiments demonstrate that MAPF-GPT-DDG surpasses all existing learning-based MAPF solvers, including the original MAPF-GPT, regarding solution quality across many testing scenarios. Remarkably, it can work with MAPF instances involving up to 1 million agents in a single environment, setting a new milestone for scalability in MAPF domains. Anton Andreychuk, Konstantin S. Yakovlev, Aleksandr I. Panov, Aleksey Skrynnik |
IROS | 3 |
| 2025 | VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for RobotsabstractIn the field of robotics, researchers face a critical challenge in ensuring reliable and efficient task planning. Verifying high-level task plans before execution significantly reduces errors and enhance the overall performance of these systems. In this paper, we propose an architecture for automatically verifying high-level task plans before their execution in simulator or real-world environments. Leveraging Large Language Models (LLMs), our approach consists of two key steps: first, the conversion of natural language instructions into Linear Temporal Logic (LTL), followed by a comprehensive analysis of action sequences. The module uses the reasoning capabilities of the LLM to evaluate logical coherence and identify potential gaps in the plan. Rigorous testing on datasets of varying complexity demonstrates the broad applicability of the module to household tasks. We contribute to improving the reliability and efficiency of task planning and addresses the critical need for robust pre-execution verification in autonomous systems. The project page is available at https://verifyllm.github.io. Danil S. Grigorev, Alexey K. Kovalev, Aleksandr I. Panov |
IROS | 3 |
| 2025 | M3PO: Massively Multi-Task Model-Based Policy OptimizationabstractWe introduce Massively Multi-Task Model-Based Policy Optimization (M3PO), a scalable model-based reinforcement learning (MBRL) framework designed to address the challenges of sample efficiency in single-task settings and generalization in multi-task domains. Existing model-based approaches like DreamerV3 rely on generative world models that prioritize pixel-level reconstruction, often at the cost of control-centric representations, while model-free methods such as PPO suffer from high sample complexity and limited exploration. M3PO integrates an implicit world model, trained to predict task outcomes without reconstructing observations, with a hybrid exploration strategy that combines model-based planning and model-free uncertainty-driven bonuses. This approach eliminates the bias-variance trade-off inherent in prior methods (e.g., POME’s exploration bonuses) by using the discrepancy between model-based and model-free value estimates to guide exploration while maintaining stable policy updates via a trust-region optimizer. M3PO is introduced as an advanced alternative to existing model-based policy optimization methods. Aditya Narendra, Dmitry Makarov, Aleksandr I. Panov |
IROS | 3 |
| 2025 | LERa: Replanning with Visual Feedback in Instruction FollowingabstractLarge Language Models are increasingly used in robotics for task planning, but their reliance on textual inputs limits their adaptability to real-world changes and failures. To address these challenges, we propose LERa — Look, Explain, Replan — a Visual Language Model-based replanning approach that utilizes visual feedback. Unlike existing methods, LERa requires only a raw RGB image, a natural language instruction, an initial task plan, and failure detection — without additional information such as object detection or predefined conditions that may be unavailable in a given scenario. The replanning process consists of three steps: (i) Look — where LERa generates a scene description and identifies errors; (ii) Explain — where it provides corrective guidance; and (iii) Replan — where it modifies the plan accordingly. LERa is adaptable to various agent architectures and can handle errors from both dynamic scene changes and task execution failures. We evaluate LERa on the newly introduced ALFRED-ChaOS and VirtualHome-ChaOS datasets, achieving a 40% improvement over baselines in dynamic environments. In tabletop manipulation tasks with a predefined probability of task failure within the PyBullet simulator, LERa improves success rates by up to 67%. Further experiments, including real-world trials with a tabletop manipulator robot, confirm LERa’s effectiveness in replanning. We demonstrate that LERa is a robust and adaptable solution for error-aware task execution in robotics. The project page is available at https://lera-robo.github.io. Svyatoslav Pchelintsev, Maxim Patratskiy, Anatoly Onishchenko, Alexandr Korchemnyi, Aleksandr Medvedev, Uliana Vinogradova, Ilya Galuzinsky, Aleksey Postnikov, Alexey K. Kovalev, Aleksandr I. Panov |
IROS | 10 |
| 2025 | Generative models for grid-based and image-based pathfinding
Daniil E. Kirilenko, Anton Andreychuk, Aleksandr I. Panov, Konstantin S. Yakovlev |
Artif. Intell. | 3 |
| 2025 | Polygon decomposition for obstacle representation in motion planning with Model Predictive Control
Aleksey Logunov, Muhammad Alhaddad, Konstantin Mironov, Konstantin S. Yakovlev, Aleksandr I. Panov |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | SegmATRon: Embodied adaptive semantic segmentation for indoor environment
Tatiana Zemskova, Margarita Kichik, Dmitry A. Yudin, Aleksey Staroverov, Aleksandr I. Panov |
Neurocomputing | 5 |
| 2024 | Learn to Follow: Decentralized Lifelong Multi-Agent Pathfinding via Planning and LearningabstractMulti-agent Pathfinding (MAPF) problem generally asks to find a set of conflict-free paths for a set of agents confined to a graph and is typically solved in a centralized fashion. Conversely, in this work, we investigate the decentralized MAPF setting, when the central controller that possesses all the information on the agents' locations and goals is absent and the agents have to sequentially decide the actions on their own without having access to the full state of the environment. We focus on the practically important lifelong variant of MAPF, which involves continuously assigning new goals to the agents upon arrival to the previous ones. To address this complex problem, we propose a method that integrates two complementary approaches: planning with heuristic search and reinforcement learning through policy optimization. Planning is utilized to construct and re-plan individual paths. We enhance our planning algorithm with a dedicated technique tailored to avoid congestion and increase the throughput of the system. We employ reinforcement learning to discover the collision avoidance policies that effectively guide the agents along the paths. The policy is implemented as a neural network and is effectively trained without any reward-shaping or external guidance. We evaluate our method on a wide range of setups comparing it to the state-of-the-art solvers. The results show that our method consistently outperforms the learnable competitors, showing higher throughput and better ability to generalize to the maps that were unseen at the training stage. Moreover our solver outperforms a rule-based one in terms of throughput and is an order of magnitude faster than a state-of-the-art search-based solver. The code is available at https://github.com/AIRI-Institute/learn-to-follow. Aleksey Skrynnik, Anton Andreychuk, Maria Nesterova, Konstantin S. Yakovlev, Aleksandr I. Panov |
AAAI | 5 |
| 2024 | Decentralized Monte Carlo Tree Search for Partially Observable Multi-Agent PathfindingabstractThe Multi-Agent Pathfinding (MAPF) problem involves finding a set of conflict-free paths for a group of agents confined to a graph. In typical MAPF scenarios, the graph and the agents' starting and ending vertices are known beforehand, allowing the use of centralized planning algorithms. However, in this study, we focus on the decentralized MAPF setting, where the agents may observe the other agents only locally and are restricted in communications with each other. Specifically, we investigate the lifelong variant of MAPF, where new goals are continually assigned to the agents upon completion of previous ones. Drawing inspiration from the successful AlphaZero approach, we propose a decentralized multi-agent Monte Carlo Tree Search (MCTS) method for MAPF tasks. Our approach utilizes the agent's observations to recreate the intrinsic Markov decision process, which is then used for planning with a tailored for multi-agent tasks version of neural MCTS. The experimental results show that our approach outperforms state-of-the-art learnable MAPF solvers. The source code is available at https://github.com/AIRI-Institute/mats-lp. Aleksey Skrynnik, Anton Andreychuk, Konstantin S. Yakovlev, Aleksandr I. Panov |
AAAI | 4 |
| 2024 | Instruction Following with Goal-Conditioned Reinforcement Learning in Virtual EnvironmentsabstractIn this study, we address the issue of enabling an artificial intelligence agent to execute complex language instructions within virtual environments. In our framework, we assume that these instructions involve intricate linguistic structures and multiple interdependent tasks that must be navigated successfully to achieve the desired outcomes. To effectively manage these complexities, we propose a hierarchical framework that combines the deep language comprehension of large language models with the adaptive action-execution capabilities of reinforcement learning agents: the language module (based on LLM) translates the language instruction into a high-level action plan, which is then executed by a pre-trained reinforcement learning agent.We have demonstrated the effectiveness of our approach in two different environments: in IGLU, where agents are instructed to build structures, and in Crafter, where agents perform tasks and interact with objects in the surrounding environment according to language commands. Zoya Volovikova, Aleksey Skrynnik, Petr Kuderov, Aleksandr I. Panov |
ECAI | 4 |
| 2024 | Object-Centric Learning with Slot Mixture ModuleabstractObject-centric architectures usually apply a differentiable module to the entire feature map to decompose it into sets of entity representations called slots. Some of these methods structurally resemble clustering algorithms, where the cluster's center in latent space serves as a slot representation. Slot Attention is an example of such a method, acting as a learnable analog of the soft k-means algorithm. Our work employs a learnable clustering method based on the Gaussian Mixture Model. Unlike other approaches, we represent slots not only as centers of clusters but also incorporate information about the distance between clusters and assigned vectors, leading to more expressive slot representations. Our experiments demonstrate that using this approach instead of Slot Attention improves performance in object-centric scenarios, achieving state-of-the-art results in the set property prediction task. Daniil E. Kirilenko, Vitaliy Vorobyov, Alexey K. Kovalev, Aleksandr I. Panov |
ICLR | 4 |
| 2024 | Gradual Optimization Learning for Conformational Energy MinimizationabstractMolecular conformation optimization is crucial to computer-aided drug discovery and materials design.
Traditional energy minimization techniques rely on iterative optimization methods that use molecular forces calculated by a physical simulator (oracle) as anti-gradients.
However, this is a computationally expensive approach that requires many interactions with a physical simulator.
One way to accelerate this procedure is to replace the physical simulator with a neural network.
Despite recent progress in neural networks for molecular conformation energy prediction, such models are prone to errors due to distribution shift, leading to inaccurate energy minimization.
We find that the quality of energy minimization with neural networks can be improved by providing optimization trajectories as additional training data.
Still, obtaining complete optimization trajectories demands a lot of additional computations.
To reduce the required additional data, we present the Gradual Optimization Learning Framework (GOLF) for energy minimization with neural networks.
The framework consists of an efficient data-collecting scheme and an external optimizer.
The external optimizer utilizes gradients from the energy prediction model to generate optimization trajectories, and the data-collecting scheme selects additional training data to be processed by the physical simulator.
Our results demonstrate that the neural network trained with GOLF performs \textit{on par} with the oracle on a benchmark of diverse drug-like molecules using significantly less additional data. Artem Tsypin, Leonid Ugadiarov, Kuzma Khrabrov, Alexander Telepov, Egor Rumiantsev, Aleksey Skrynnik, Aleksandr I. Panov, Dmitry P. Vetrov, Elena Tutubalina, Artur Kadurin |
ICLR | 7 |
| 2024 | Neural Potential Field for Obstacle-Aware Local Motion PlanningabstractModel predictive control (MPC) may provide local motion planning for mobile robotic platforms. The challenging aspect is the analytic representation of collision cost for the case when both the obstacle map and robot footprint are arbitrary. We propose a Neural Potential Field: a neural network model that returns a differentiable collision cost based on robot pose, obstacle map, and robot footprint. The differentiability of our model allows its usage within the MPC solver. It is computationally hard to solve problems with a very high number of parameters. Therefore, our architecture includes neural image encoders, which transform obstacle maps and robot footprints into embeddings, which reduce problem dimensionality by two orders of magnitude. The reference data for network training are generated based on algorithmic calculation of a signed distance function. Comparative experiments showed that the proposed approach is comparable with existing local planners: it provides trajectories with outperforming smoothness, comparable path length, and safe distance from obstacles. Muhammad Alhaddad, Konstantin Mironov, Aleksey Staroverov, Aleksandr I. Panov |
ICRA | 4 |
| 2024 | Model-based Policy Optimization using Symbolic World ModelabstractThe application of learning-based control methods in robotics presents significant challenges. One is that model-free reinforcement learning algorithms use observation data with low sample efficiency. To address this challenge, a prevalent approach is model-based reinforcement learning, which involves employing an environment dynamics model. We suggest approximating transition dynamics with symbolic expressions, which are generated via symbolic regression. Approximation of a mechanical system with a symbolic model has fewer parameters than approximation with neural networks, which can potentially lead to higher accuracy and quality of extrapolation. We use a symbolic dynamics model to generate trajectories in model-based policy optimization to improve the sample efficiency of the learning algorithm. We evaluate our approach across various tasks within simulated environments. Our method demonstrates superior sample efficiency in these tasks compared to model-free and model-based baseline methods. Andrey Gorodetskiy, Konstantin Mironov, Aleksandr I. Panov |
IROS | 3 |
| 2024 | Hierarchical waste detection with weakly supervised segmentation in images from recycling plants
Dmitry A. Yudin, Nikita Zakharenko, Artem Smetanin, Roman Filonov, Margarita Kichik, Vladislav Kuznetsov, Dmitry Larichev, Evgeny Gudov, Semen A. Budennyy, Aleksandr I. Panov |
Eng. Appl. Artif. Intell. | 10 |
| 2024 | When to Switch: Planning and Learning for Partially Observable Multi-Agent PathfindingabstractMulti-agent pathfinding (MAPF) is a problem that involves finding a set of non-conflicting paths for a set of agents confined to a graph. In this work, we study a MAPF setting, where the environment is only partially observable for each agent, i.e., an agent observes the obstacles and other agents only within a limited field-of-view. Moreover, we assume that the agents do not communicate and do not share knowledge on their goals, intended actions, etc. The task is to construct a policy that maps the agent's observations to actions. Our contribution is multifold. First, we propose two novel policies for solving partially observable MAPF (PO-MAPF): one based on heuristic search and another one based on reinforcement learning (RL). Next, we introduce a mixed policy that is based on switching between the two. We suggest three different switch scenarios: the heuristic, the deterministic, and the learnable one. A thorough empirical evaluation of all the proposed policies in a variety of setups shows that the mixing policy demonstrates the best performance is able to generalize well to the unseen maps and problem instances, and, additionally, outperforms the state-of-the-art counterparts (PRIMAL2 and PICO). The source-code is available at https://github.com/AIRI-Institute/when-to-switch. Aleksey Skrynnik, Anton Andreychuk, Konstantin S. Yakovlev, Aleksandr I. Panov |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | TransPath: Learning Heuristics for Grid-Based Pathfinding via TransformersabstractHeuristic search algorithms, e.g. A*, are the commonly used tools for pathfinding on grids, i.e. graphs of regular structure that are widely employed to represent environments in robotics, video games, etc. Instance-independent heuristics for grid graphs, e.g. Manhattan distance, do not take the obstacles into account, and thus the search led by such heuristics performs poorly in obstacle-rich environments. To this end, we suggest learning the instance-dependent heuristic proxies that are supposed to notably increase the efficiency of the search. The first heuristic proxy we suggest to learn is the correction factor, i.e. the ratio between the instance-independent cost-to-go estimate and the perfect one (computed offline at the training phase). Unlike learning the absolute values of the cost-to-go heuristic function, which was known before, learning the correction factor utilizes the knowledge of the instance-independent heuristic. The second heuristic proxy is the path probability, which indicates how likely the grid cell is lying on the shortest path. This heuristic can be employed in the Focal Search framework as the secondary heuristic, allowing us to preserve the guarantees on the bounded sub-optimality of the solution. We learn both suggested heuristics in a supervised fashion with the state-of-the-art neural networks containing attention blocks (transformers). We conduct a thorough empirical evaluation on a comprehensive dataset of planning tasks, showing that the suggested techniques i) reduce the computational effort of the A* up to a factor of 4x while producing the solutions, whose costs exceed those of the optimal solutions by less than 0.3% on average; ii) outperform the competitors, which include the conventional techniques from the heuristic search, i.e. weighted A*, as well as the state-of-the-art learnable planners. The project web-page is: https://airi-institute.github.io/TransPath/. Daniil E. Kirilenko, Anton Andreychuk, Aleksandr I. Panov, Konstantin S. Yakovlev |
AAAI | 3 |
| 2023 | Interpreting Decision Process in Offline Reinforcement Learning for Interactive Recommendation Systems
Zoya Volovikova, Petr Kuderov, Aleksandr I. Panov |
ICONIP (9) | 3 |
| 2022 | HPointLoc: Point-Based Indoor Place Recognition Using Synthetic RGB-D Images
Dmitry A. Yudin, Yaroslav K. Solomentsev, Ruslan Musaev, Aleksey Staroverov, Aleksandr I. Panov |
ICONIP (3) | 5 |
| 2022 | Vector Symbolic Scene Representation for Semantic Place RecognitionabstractMost state-of-the-art methods do not explicitly use scene semantics for place recognition by the images. We address this problem and propose a new two-stage approach referred to as TSVLoc. It solves the place recognition task as the image retrieval problem and enriches any well-known method. In the first model-agnostic stage, any modern neural network model that does not directly use semantics, e.g., HF-Net, NetVLAD, or Patch-NetVLAD, can be used. In the second stage, we apply the Vector Symbolic Architectures (VSA) framework to construct semantic scene representation. Our method uses semantic segmentation of an image to extract objects and their relations and applies VSA operations to form semantic scene representation. For this, an optional usage of the depth map was considered, which showed promising results. The effectiveness of our approach is demonstrated through extensive experiments on the open large-scale datasets: the indoor HPointLoc dataset built in the Habitat simulation environment and the outdoor Oxford RobotCar dataset. The proposed solution significantly improves the quality of the place recognition. Daniil E. Kirilenko, Alexey K. Kovalev, Yaroslav K. Solomentsev, Alexander Melekhin, Dmitry A. Yudin, Aleksandr I. Panov |
IJCNN | 6 |
| 2021 | Planning with Hierarchical Temporal Memory for Deterministic Markov Decision Problem
Petr Kuderov, Aleksandr I. Panov |
ICAART (2) | 2 |
| 2021 | Forgetful experience replay in hierarchical reinforcement learning from expert demonstrations
Aleksey Skrynnik, Aleksey Staroverov, Ermek Aitygulov, Kirill Aksenov, Vasilii Davydov, Aleksandr I. Panov |
Knowl. Based Syst. | 6 |
| 2015 | Assessment of Dendritic Cell Therapy Effectiveness Based on the Feature Extraction from Scientific Publications
Alexey Yu. Lupatov, Aleksandr I. Panov, Roman E. Suvorov, Alexander V. Shvets, Konstantin N. Yarygin, Galina D. Volkova |
ICPRAM (2) | 2 |