EDBT 2026 Demo / reviewers in the wild / expert
Artem Lykov
dblp:332/1788
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0001-6119-2366ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Legged, aerial and field robots · 16% Planning, search and constraint satisfaction · 16% Autonomous driving · 16% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Legged, aerial and field robots
aerial robots |
0.9 | 1 | 2025 | UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation · HRI 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
cognitive robotics |
0.9 | 1 | 2025 | CognitiveOS: Large Multimodal Model Based System to Endow Any Type of Robot with Generative AI · ICRA 2025 |
Robotics › Autonomous driving
end-to-end driving |
0.9 | 1 | 2025 | METDrive: Multimodal End-to-End Autonomous Driving with Temporal Guidance · ICRA 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
mission planning |
0.9 | 1 | 2025 | UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation · HRI 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | CognitiveOS: Large Multimodal Model Based System to Endow Any Type of Robot with Generative AI · ICRA 2025 |
Robotics › Robot navigation and mapping › mobile robot navigation › navigation planning
waypoint prediction |
0.9 | 1 | 2025 | METDrive: Multimodal End-to-End Autonomous Driving with Temporal Guidance · ICRA 2025 |
Human-robot interaction › teleoperation
gesture-based robot control |
0.9 | 1 | 2025 | GestLLM: Advanced Hand Gesture Interpretation via Large Language Models for Human-Robot Interaction · HRI 2025 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation · HRI 2025 |
Human-robot interaction
assistive robotics |
0.3 | 1 | 2025 | GestLLM: Advanced Hand Gesture Interpretation via Large Language Models for Human-Robot Interaction · HRI 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9transformer · 0.9temporal feature embedding · 0.9multimodal fusion · 0.9multi-agent system · 0.9mediapipe · 0.9large multimodal model · 0.9large language model · 0.9k-nearest neighbors · 0.9feature extraction · 0.9GPT · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GestLLM: Advanced Hand Gesture Interpretation via Large Language Models for Human-Robot InteractionabstractThis paper introduces GestLLM, an advanced system for human-robot interaction that enables intuitive robot control through hand gestures. Unlike conventional systems, which rely on a limited set of predefined gestures, GestLLM leverages large language models and feature extraction via MediaPipe [1] to interpret a diverse range of gestures. This integration addresses key limitations in existing systems, such as restricted gesture flexibility and the inability to recognize complex or unconventional gestures commonly used in human communication. By combining state-of-the-art feature extraction and language model capabilities, GestLLM achieves performance comparable to leading vision-language models while supporting gestures underrepresented in traditional datasets. For example, this includes gestures from popular culture, such as the “Vulcan salute” from Star Trek, without any additional pretraining, prompt engineering, etc. This flexibility enhances the naturalness and inclusivity of robot control, making interactions more intuitive and user-friendly. GestLLM provides a significant step forward in gesture-based interaction, enabling robots to understand and respond to a wide variety of hand gestures effectively. This paper outlines its design, implementation, and evaluation, demonstrating its potential applications in advanced human-robot collaboration, assistive robotics, and interactive entertainment. Oleg Kobzarev, Artem Lykov, Dzmitry Tsetserukou |
HRI | 2 |
| 2025 | UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission GenerationabstractThe UAV-VLA (Visual-Language-Action) system is a tool designed to facilitate communication with aerial robots. By integrating satellite imagery processing with the Visual Language Model (VLM) and the powerful capabilities of GPT, UAV-VLA enables users to generate general flight paths-and-action plans through simple text requests. This system leverages the rich contextual information provided by satellite images, allowing for enhanced decision-making and mission planning. The combination of visual analysis by VLM and natural language processing by GPT can provide the user with the path-and-action set, making aerial operations more efficient and accessible. The newly developed method showed the difference in the length of the created trajectory in 22% and the mean error in finding the objects of interest on a map in 34.22 m by Euclidean distance in the K-Nearest Neighbors (KNN) approach. Additionally, the UAV-VLA system generates all flight plans in just 5 minutes and 24 seconds, making it 6.5 times faster than an experienced human operator. The code is available here: https://github.com/sautenich/uav-vla Oleg Sautenkov, Yasheerah Yaqoot, Artem Lykov, Muhammad Ahsan Mustafa, Grik Tadevosyan, Aibek Akhmetkazy, Miguel Altamirano, Mikhail Martynov, Sausar Karaf, Dzmitry Tsetserukou |
HRI | 3 |
| 2025 | METDrive: Multimodal End-to-End Autonomous Driving with Temporal GuidanceabstractMultimodal end-to-end autonomous driving has shown promising advancements in recent work. By embedding more modalities into end-to-end networks, the system's understanding of both static and dynamic aspects of the driving environment is enhanced, thereby improving the safety of autonomous driving. In this paper, we introduce METDrive, an end-to-end system that leverages temporal guidance from the embedded time series features of ego states, including rotation angles, steering, throttle signals, and waypoint vectors. The geometric features derived from the perception sensor data and the time series features of ego state data jointly guide the waypoint prediction with the proposed temporal guidance loss function. We evaluated METDrive on the CARLA leaderboard benchmarks, achieving a driving score of 70%, a route completion score of 94%, and an infraction score of 0.78. Ziang Guo, Xinhao Lin, Zakhar Yagudin, Artem Lykov, Yanqiang Li, Dzmitry Tsetserukou |
ICRA | 4 |
| 2025 | CognitiveOS: Large Multimodal Model Based System to Endow Any Type of Robot with Generative AIabstractThis paper introduces CognitiveOS, the first operating system designed for cognitive robots capable of functioning across diverse robotic platforms. CognitiveOS is structured as a multi-agent system comprising modules built upon a transformer architecture, facilitating communication through an internal monologue format. These modules collectively empower the robot to tackle intricate real-world tasks. The paper delineates the operational principles of the system along with the descriptions of its nine distinct modules. The modular design endows the system with distinctive advantages over traditional end-to-end methodologies, notably in terms of adaptability and scalability. The system's modules are configurable, modifiable, or deactivatable depending on the task requirements, while new modules can be seamlessly integrated. This system serves as a foundational resource for researchers and developers in the Cognitive Robotics domain, alleviating the burden of constructing a cognitive robot system from scratch. Experimental findings demonstrate the system's advanced task comprehension and adaptability across varied tasks, robotic platforms, and module configurations, underscoring its potential for realworld applications. Moreover, in the category of Reasoning it outperformed CognitiveDog (by 15%) and RT2 (by 31%), achieving the highest to date rate of 77 %. We provide a code repository and dataset for the replication of CognitiveOS: https://github.com/Arcwy0/cognitiveos Artem Lykov, Mikhail Konenkov, Koffivi Fidèle Gbagbe, Mikhail Litvinov, Denis Davletshin, Aleksey Fedoseev, Miguel Altamirano, Robinroy Peter, Dzmitry Tsetserukou |
ICRA | 1 |
| 2025 | Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous RobotsabstractThis paper presents the concept of Industry 6.0, which introduces the world’s first fully automated production system that autonomously handles the entire product design and manufacturing process based on user-provided natural language descriptions. By leveraging generative AI, the system automates critical aspects of production, including product blueprint design, component manufacturing, logistics, and assembly. A heterogeneous swarm of robots, each equipped with individual AI through integration with Large Language Models (LLMs), orchestrates the production process. The robotic system includes manipulator arms, delivery drones, and 3D printers capable of generating assembly blueprints. The system was evaluated using commercial and open source LLMs, operating via APIs and local deployment. A user study demonstrated that the system reduced the average production time to 119.10 minutes, significantly outperforming a team of expert human developers, who averaged 528.64 minutes (an improvement factor of 4.4). Furthermore, in the product blueprinting stage, the system outperformed human CAD operators by an unprecedented factor of 47, completing the task in 0.5 minutes compared to 23.5 minutes. This breakthrough represents a major leap towards fully autonomous manufacturing. Artem Lykov, Miguel Altamirano, Mikhail Konenkov, Valerii Serpiva, Koffivi Fidèle Gbagbe, Ali Alabbas, Aleksey Fedoseev, Luis Moreno 0007, Muhammad Haris Khan, Ziang Guo, Dzmitry Tsetserukou |
IROS | 1 |
| 2025 | FADet: A Multi-sensor 3D Object Detection Network based on Local Featured AttentionabstractCamera, LiDAR, and radar are common perception sensors for autonomous driving tasks. Robust prediction of 3D object detection is optimally based on the fusion of these sensors. Taking advantage of their abilities remains a challenge, because each of these sensors has its own characteristics. Specifically, different sensors present different scales in their corresponding extracted features. To address this problem, considering the feature alignment in different scales, in this paper, we propose FADet, a multi-sensor 3D detection network, which specifically studies the characteristics of different sensors across the dimensions of their data input based on our local featured attention modules. For camera images, we propose a dual-attention-based submodule. For LiDAR point clouds, the triple-attention-based submodule is utilized, while the mixed-attention-based submodule is applied for features of radar points. With local featured attention submodules, our FADet has effective detection results in long-tail and complex scenes from camera, LiDAR and radar input. In the NuScenes validation dataset, FADet achieves state-of-the-art performance on LiDAR-camera object detection tasks with 71.8% NDS and 69.0% mAP, at the same time, on radar-camera object detection tasks with 51.7% NDS and 40.3% mAP. Ziang Guo, Zakhar Yagudin, Selamawit Asfaw, Artem Lykov, Dzmitry Tsetserukou |
IV | 4 |
| 2025 | UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and RescueabstractEmergency search and rescue (SAR) operations often require rapid and precise target identification in complex environments where traditional manual drone control is inefficient. In order to address these scenarios, a rapid SAR system, UAV-VLRR (Vision-Language-Rapid-Response), is developed in this research. This system consists of two aspects: 1) A multimodal system which harnesses the power of Visual Language Model (VLM) and the natural language processing capabilities of ChatGPT-4o (LLM) for scene interpretation. 2) A non-linear model predictive control (NMPC) with built-in obstacle avoidance for rapid response by a drone to fly according to the output of the multimodal system. This work aims at improving response times in emergency SAR operations by providing a more intuitive and natural approach to the operator to plan the SAR mission while allowing the drone to carry out that mission in a rapid and safe manner. When tested, our approach was faster on an average by 33.75% when compared with an off-the-shelf autopilot and 54.6% when compared with a human pilot. Github: https://github.com/ahsan-mustafa/uav-vlrr Video of UAV-VLRR: https://youtu.be/KJqQGKKt1xY Yasheerah Yaqoot, Muhammad Ahsan Mustafa, Oleg Sautenkov, Artem Lykov, Valerii Serpiva, Dzmitry Tsetserukou |
IV | 4 |
| 2024 | Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous ManipulationsabstractThis research introduces the Bi-VLA (Vision-Language-Action) model, a novel system designed for bimanual robotic dexterous manipulation that seamlessly integrates vision for scene understanding, language comprehension for translating human instructions into executable code, and physical action generation. We evaluated the system's functionality through a series of household tasks, including the preparation of a desired salad upon human request. Bi-VLA demonstrates the ability to interpret complex human instructions, perceive and understand the visual context of ingredients, and execute precise bimanual actions to prepare the requested salad. We assessed the system's performance in terms of accuracy, efficiency, and adaptability to different salad recipes and human preferences through a series of experiments. Our results show a 100 % success rate in generating the correct executable code by the Language Module, a 96.06 % success rate in detecting specific ingredients by the Vision Module, and an overall success rate of 83.4 % in correctly executing user-requested tasks. Koffivi Fidèle Gbagbe, Miguel Altamirano, Ali Alabbas, Oussama Alyounes, Artem Lykov, Dzmitry Tsetserukou |
SMC | 5 |