EDBT 2026 Demo / reviewers in the wild / expert
Dmitry A. Yudin
dblp:278/1638
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-1407-2633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scene graph-driven reasoning for action planning of humanoid robot
Dmitry A. Yudin, Alexander Lazarev, Eva Bakaeva, Angelika Kochetkova, Alexey K. Kovalev, Aleksandr I. Panov |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Say it better: RL-based prompt tuning for enhancing open-vocabulary recognition
Mikhail Avshalumov, Zoya Volovikova, Dmitry A. Yudin, Aleksandr I. Panov |
Neurocomputing | 3 |
| 2025 | 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene UnderstandingabstractA 3D scene graph represents a compact scene model by capturing both the objects present and the semantic relationships between them, making it a promising structure for robotic applications. To effectively interact with users, an embodied intelligent agent should be able to answer a wide range of natural language queries about the surrounding 3D environment. Large Language Models (LLMs) are beneficial solutions for user-robot interaction due to their natural language understanding and reasoning abilities. Recent methods for learning scene representations have shown that adapting these representations to the 3D world can significantly improve the quality of LLM responses. However, existing methods typically rely only on geometric information, such as object coordinates, and overlook the rich semantic relationships between objects. In this work, we propose 3DGraphLLM, a method for constructing a learnable representation of a 3D scene graph that explicitly incorporates semantic relationships. This representation is used as input to LLMs for performing 3D vision-language tasks. In our experiments on popular ScanRefer, Multi3DRefer, ScanQA, Sqa3D, and Scan2cap datasets, we demonstrate that our approach outperforms baselines that do not leverage semantic relationships between objects. The code is publicly available at https://github.com/CognitiveAISystems/3DGraphLLM. Tatiana Zemskova, Dmitry A. Yudin |
ICCV | 2 |
| 2025 | Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene GraphabstractLocating objects described in natural language presents a significant challenge for autonomous agents. Existing CLIP-based open-vocabulary methods successfully perform 3D object grounding with simple (bare) queries, but cannot cope with ambiguous descriptions that demand an understanding of object relations. To tackle this problem, we propose a modular approach called BBQ (Beyond Bare Queries), which constructs 3D scene graph representation with metric and semantic spatial edges and utilizes a large language model as a human-toagent interface through our deductive scene reasoning algorithm. BBQ employs robust DINO-powered associations to construct 3D object-centric map and an advanced raycasting algorithm with a 2D vision-language model to describe them as graph nodes. On the Replica and ScanNet datasets, we have demonstrated that BBQ takes a leading place in openvocabulary 3D semantic segmentation compared to other zeroshot methods. Also, we show that leveraging spatial relations is especially effective for scenes containing multiple entities of the same semantic class. On challenging Sr3D+, Nr3D and ScanRefer benchmarks, our deductive approach demonstrates a significant improvement, enabling objects grounding by complex queries compared to other state-of-the-art methods. The combination of our design choices and software implementation has resulted in significant data processing speed in experiments on the robot on-board computer. This promising performance enables the application of our approach in intelligent robotics projects. We made the code publicly available at linukc.github.io/BeyondBareQueries. Sergey Linok, Tatiana Zemskova, Svetlana Ladanova, Roman Titkov, Dmitry A. Yudin, Maxim Monastyrny, Aleksei Valenkov |
ICRA | 5 |
| 2025 | LaMDEN: Addressing Elevator-Based Navigation with Large Language Models and 3D Scene GraphsabstractMulti-Floor navigation has become an increasingly important topic in the robotics research community. Among various approaches, 3D scene graphs have emerged as an effective tool for addressing multi-floor navigation challenges. However, existing methods typically construct scene graphs using pre-collected datasets that mainly focus on stair-based navigation, largely overlooking the more common elevator-based navigation in multi-floor buildings. In this paper, we bridge this gap by introducing LaMDEN (Large language Model-Driven Elevator Navigation), a novel framework designed for elevator-based multi-floor navigation. LaMDEN operates in two stages: it first constructs a 3D scene graph from RGB-D sequences paired with camera poses, and then leverages Large Language Models (LLMs) to decompose high-level, long-horizon instructions into actionable primitives, such as pressing elevator buttons and entering or exiting elevators, enabling smooth cross-floor navigation. We validate the effectiveness of LaMDEN in complex elevator-based environments built within the Isaac Sim simulator. Furthermore, we release a newly collected dataset for 3D scene graph construction, providing a valuable resource for advancing research in this area. Code and dataset are publicly available at: https://github.com/zhanghuzhenyu/mul-floor-navigation. Huzhenyu Zhang, Dmitry A. Yudin |
IJCNN | 2 |
| 2025 | SegmATRon: Embodied adaptive semantic segmentation for indoor environment
Tatiana Zemskova, Margarita Kichik, Dmitry A. Yudin, Aleksey Staroverov, Aleksandr I. Panov |
Neurocomputing | 3 |
| 2024 | Hierarchical waste detection with weakly supervised segmentation in images from recycling plants
Dmitry A. Yudin, Nikita Zakharenko, Artem Smetanin, Roman Filonov, Margarita Kichik, Vladislav Kuznetsov, Dmitry Larichev, Evgeny Gudov, Semen A. Budennyy, Aleksandr I. Panov |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | OFMPNet: Deep end-to-end model for occupancy and flow prediction in urban environment
Youshaa Murhij, Dmitry A. Yudin |
Neurocomputing | 2 |
| 2023 | TASFormer: Task-Aware Image Segmentation Transformer
Dmitry A. Yudin, Aleksandr Khorin, Tatiana Zemskova, Darya Ovchinnikova |
ICONIP (5) | 1 |
| 2022 | Rethinking Voxelization and Classification for 3D Object Detection
Youshaa Murhij, Alexander Golodkov, Dmitry A. Yudin |
ICONIP (6) | 3 |
| 2022 | HPointLoc: Point-Based Indoor Place Recognition Using Synthetic RGB-D Images
Dmitry A. Yudin, Yaroslav K. Solomentsev, Ruslan Musaev, Aleksey Staroverov, Aleksandr I. Panov |
ICONIP (3) | 1 |
| 2022 | Vector Symbolic Scene Representation for Semantic Place RecognitionabstractMost state-of-the-art methods do not explicitly use scene semantics for place recognition by the images. We address this problem and propose a new two-stage approach referred to as TSVLoc. It solves the place recognition task as the image retrieval problem and enriches any well-known method. In the first model-agnostic stage, any modern neural network model that does not directly use semantics, e.g., HF-Net, NetVLAD, or Patch-NetVLAD, can be used. In the second stage, we apply the Vector Symbolic Architectures (VSA) framework to construct semantic scene representation. Our method uses semantic segmentation of an image to extract objects and their relations and applies VSA operations to form semantic scene representation. For this, an optional usage of the depth map was considered, which showed promising results. The effectiveness of our approach is demonstrated through extensive experiments on the open large-scale datasets: the indoor HPointLoc dataset built in the Habitat simulation environment and the outdoor Oxford RobotCar dataset. The proposed solution significantly improves the quality of the place recognition. Daniil E. Kirilenko, Alexey K. Kovalev, Yaroslav K. Solomentsev, Alexander Melekhin, Dmitry A. Yudin, Aleksandr I. Panov |
IJCNN | 5 |
| 2022 | FMFNet: Improve the 3D Object Detection and Tracking via Feature Map FlowabstractThree-dimensional object detection and tracking from point clouds are important computer vision tasks for robots and vehicles where objects can be represented as 3D boxes. Improving the accuracy of understanding the environment is critical for successful autonomous driving. This paper presents a simple yet efficient method called “Feature Map Flow, FMF” for 3D object detection and tracking, considering time-spatial feature map aggregation from different timesteps of deep neural model inference. Several versions of the FMF are proposed: from common concatenation to context-based feature map fusion and odometry usage for previous feature map affine transform. The proposed approach significantly improves the quality of 3D detection and tracking baseline on the nuScenes and Waymo benchmarks. The software implementation of the proposed method has been carried out for the server platform and the NVidia Jetson AGX Xavier single-board computer. Its results have demonstrated high application prospects both for automated 3D point cloud labeling and for embedded on-board vehicle systems. The developed code is publicly available at: https://github.com/YoushaaMurhij/FMFNet. Youshaa Murhij, Dmitry A. Yudin |
IJCNN | 2 |
| 2022 | Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo ImagesabstractThe detection of dynamic and static obstacles is a key task for the navigation of autonomous ground vehicles. The article presents a new algorithm for generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras. The camera images are segmented by the proposed deep neural network FCN-ResNet-M-OC, which combines the speed of the FCN-ResNet method and improves the quality of the model using the concept of object context representation. The paper investigates supervised approaches to network training on unbalanced samples with road scenes such as the weighted cross entropy and the Focal Loss. The occupancy map is built from point clouds with semantic labels, in which static environment and potentially dynamic obstacles are highlighted. Our solution is operational in real time and applicable on platforms with limited computing resources. The approach was tested on autonomous vehicle datasets: Semantic KITTI, KITTI-360, Mapillary Vistas and custom OpenTaganrog. The usage of semantically labeled point clouds increased the precision of obstacle detection by an average of 17%. The performance of the entire approach on various computing platforms with Jetson Xavier, RTX3070, GPUs NVidia Tesla V100 is respectively from 10 to 15 FPS for input image resolution$1920\times 1080$pixels. Ilya Shepel, Vasily Adeshkin, Ilya Belkin, Dmitry A. Yudin |
IEEE Trans. Intell. Transp. Syst. | 4 |