Dzmitry Tsetserukou

dblp:62/6503 · DBLP profile ↗
← Back
79ranked-venue papers
4as first author
59since 2021 · last 2025
0000-0001-8055-5345ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 41 · 3 first-author · 36 since 2021Artificial intelligence and machine learning · 33 · 2 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 1 first-author · 26 since 2021Systems, architecture and hardware · 20 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 ViewVR: Visual Feedback Modes to Achieve Quality of VR-Based Telemanipulation
abstract
The paper focuses on an immersive teleoperation system that enhances operator's ability to actively perceive the robot's surroundings. A consumer-grade HTC Vive VR system was used to synchronize the operator's hand and head movements with a UR3 robot and a custom-built robotic head with two degrees of freedom (2-DoF). The system's usability, manipulation efficiency, and intuitiveness of control were evaluated in comparison with static head camera positioning across three distinct tasks. Code and other supplementary materials can be accessed by link: https://github.com/ErkhovArtemNiewVR.
Artem Erkhov, Artem Bazhenov, Sergei Satsevich, Danil Belov, Farit Khabibullin, Sergei Egorov, Maxim Gromakov, Miguel Altamirano, Dzmitry Tsetserukou
HRI9
2025 Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing
abstract
This paper introduces Shake-VLA, a Vision-Language-Action (VLA) model-based system designed to enable bimanual robotic manipulation for automated cocktail preparation. The system integrates a vision module for detecting ingredient bottles and reading labels, a speech-to-text module for interpreting user commands, and a language model to generate task-specific robotic instructions. Force Torque (FT) sensors are employed to precisely measure the quantity of liquid poured, ensuring accuracy in ingredient proportions during the mixing process. The system architecture includes a Retrieval-Augmented Generation (RAG) module for accessing and adapting recipes, an anomaly detection mechanism to address ingredient availability issues, and bimanual robotic arms for dexterous manipulation. Experimental evaluations demonstrated a high success rate across system components, with the speech-to-text module achieving a 93% success rate in noisy environments, the vision module attaining a 91% success rate in object and label detection in cluttered environment, the anomaly module successfully identified 95% of discrepancies between detected ingredients and recipe requirements, and the system achieved an overall success rate of 100% in preparing cocktails, from recipe formulation to action generation.
Muhamamd Haris Khan, Selamawit Asfaw, Dmitrii Iarchuk, Miguel Altamirano, Luis Moreno 0007, Issatay Tokmurziyev, Dzmitry Tsetserukou
HRI7
2025 GestLLM: Advanced Hand Gesture Interpretation via Large Language Models for Human-Robot Interaction
abstract
This paper introduces GestLLM, an advanced system for human-robot interaction that enables intuitive robot control through hand gestures. Unlike conventional systems, which rely on a limited set of predefined gestures, GestLLM leverages large language models and feature extraction via MediaPipe [1] to interpret a diverse range of gestures. This integration addresses key limitations in existing systems, such as restricted gesture flexibility and the inability to recognize complex or unconventional gestures commonly used in human communication. By combining state-of-the-art feature extraction and language model capabilities, GestLLM achieves performance comparable to leading vision-language models while supporting gestures underrepresented in traditional datasets. For example, this includes gestures from popular culture, such as the “Vulcan salute” from Star Trek, without any additional pretraining, prompt engineering, etc. This flexibility enhances the naturalness and inclusivity of robot control, making interactions more intuitive and user-friendly. GestLLM provides a significant step forward in gesture-based interaction, enabling robots to understand and respond to a wide variety of hand gestures effectively. This paper outlines its design, implementation, and evaluation, demonstrating its potential applications in advanced human-robot collaboration, assistive robotics, and interactive entertainment.
Oleg Kobzarev, Artem Lykov, Dzmitry Tsetserukou
HRI3
2025 UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
abstract
The UAV-VLA (Visual-Language-Action) system is a tool designed to facilitate communication with aerial robots. By integrating satellite imagery processing with the Visual Language Model (VLM) and the powerful capabilities of GPT, UAV-VLA enables users to generate general flight paths-and-action plans through simple text requests. This system leverages the rich contextual information provided by satellite images, allowing for enhanced decision-making and mission planning. The combination of visual analysis by VLM and natural language processing by GPT can provide the user with the path-and-action set, making aerial operations more efficient and accessible. The newly developed method showed the difference in the length of the created trajectory in 22% and the mean error in finding the objects of interest on a map in 34.22 m by Euclidean distance in the K-Nearest Neighbors (KNN) approach. Additionally, the UAV-VLA system generates all flight plans in just 5 minutes and 24 seconds, making it 6.5 times faster than an experienced human operator. The code is available here: https://github.com/sautenich/uav-vla
Oleg Sautenkov, Yasheerah Yaqoot, Artem Lykov, Muhammad Ahsan Mustafa, Grik Tadevosyan, Aibek Akhmetkazy, Miguel Altamirano, Mikhail Martynov, Sausar Karaf, Dzmitry Tsetserukou
HRI10
2025 SafeSwarm: Decentralized Safe RL for the Swarm of Drones Landing in Dense Crowds
abstract
This paper introduces a safe swarm of drones capable of performing landings in crowded environments robustly by relying on Reinforcement Learning techniques combined with Safe Learning. The developed system allows us to teach the swarm of drones with different dynamics to land on moving landing pads in an environment while avoiding collisions with obstacles and between agents. The safe barrier net algorithm was developed and evaluated using a swarm of Crazyflie 2.1 micro quadrotors, which were tested indoors with the Vicon motion capture system to ensure precise localization and control. Experimental results show that our system achieves landing accuracy of 2.25 cm with a mean time of 17 s and collision-free landings, underscoring its effectiveness and robustness in real-world scenarios. This work offers a promising foundation for applications in environments where safety and precision are paramount.
Grik Tadevosyan, Maksim Osipenko, Demetros Aschu, Aleksey Fedoseev, Valerii Serpiva, Oleg Sautenkov, Sausar Karaf, Dzmitry Tsetserukou
HRI8
2025 GazeGrasp: DNN-Driven Robotic Grasping with Wearable Eye-Gaze Interface
abstract
We present GazeGrasp, a gaze-based manipulation system enabling individuals with motor impairments to control collaborative robots using eye-gaze. The system employs an ESP32 CAM for eye tracking, MediaPipe for gaze detection, and YOLOv8 for object localization, integrated with a Uni-versal Robot UR10 for manipulation tasks. After user-specific calibration, the system allows intuitive object selection with a magnetic snapping effect and robot control via eye gestures. Experimental evaluation involving 13 participants demonstrated that the magnetic snapping effect significantly reduced gaze alignment time, improving task efficiency by 31%. GazeGrasp provides a robust, hands-free interface for assistive robotics, enhancing accessibility and autonomy for users.
Issatay Tokmurziyev, Miguel Altamirano, Luis Moreno 0007, Muhammad Haris Khan, Dzmitry Tsetserukou
HRI5
2025 Towards Intuitive Drone Operation Using a Handheld Motion Controller
abstract
We present an intuitive human-drone interaction system that utilizes a gesture-based motion controller to enhance the drone operation experience in real and simulated environments. The handheld motion controller enables natural control of the drone through the movements of the operator's hand, thumb, and index finger: the trigger press manages the throttle, the tilt of the hand adjusts pitch and roll, and the thumbstick controls yaw rotation. Communication with drones is facilitated via the ExpressLRS radio protocol, ensuring robust connectivity across various frequencies. The user evaluation of the flight experience with the designed drone controller using the UEQ-S survey showed high scores for both Pragmatic (mean=2.2, SD = 0.8) and Hedonic (mean=2.3, SD = 0.9) Qualities. This versatile control interface supports applications such as research, drone racing, and training programs in real and simulated environments, thereby contributing to advances in the field of human-drone interaction.
Daria Trinitatova, Sofia Shevelo, Dzmitry Tsetserukou
HRI3
2025 METDrive: Multimodal End-to-End Autonomous Driving with Temporal Guidance
abstract
Multimodal end-to-end autonomous driving has shown promising advancements in recent work. By embedding more modalities into end-to-end networks, the system's understanding of both static and dynamic aspects of the driving environment is enhanced, thereby improving the safety of autonomous driving. In this paper, we introduce METDrive, an end-to-end system that leverages temporal guidance from the embedded time series features of ego states, including rotation angles, steering, throttle signals, and waypoint vectors. The geometric features derived from the perception sensor data and the time series features of ego state data jointly guide the waypoint prediction with the proposed temporal guidance loss function. We evaluated METDrive on the CARLA leaderboard benchmarks, achieving a driving score of 70%, a route completion score of 94%, and an infraction score of 0.78.
Ziang Guo, Xinhao Lin, Zakhar Yagudin, Artem Lykov, Yanqiang Li, Dzmitry Tsetserukou
ICRA7
2025 CognitiveOS: Large Multimodal Model Based System to Endow Any Type of Robot with Generative AI
abstract
This paper introduces CognitiveOS, the first operating system designed for cognitive robots capable of functioning across diverse robotic platforms. CognitiveOS is structured as a multi-agent system comprising modules built upon a transformer architecture, facilitating communication through an internal monologue format. These modules collectively empower the robot to tackle intricate real-world tasks. The paper delineates the operational principles of the system along with the descriptions of its nine distinct modules. The modular design endows the system with distinctive advantages over traditional end-to-end methodologies, notably in terms of adaptability and scalability. The system's modules are configurable, modifiable, or deactivatable depending on the task requirements, while new modules can be seamlessly integrated. This system serves as a foundational resource for researchers and developers in the Cognitive Robotics domain, alleviating the burden of constructing a cognitive robot system from scratch. Experimental findings demonstrate the system's advanced task comprehension and adaptability across varied tasks, robotic platforms, and module configurations, underscoring its potential for realworld applications. Moreover, in the category of Reasoning it outperformed CognitiveDog (by 15%) and RT2 (by 31%), achieving the highest to date rate of 77 %. We provide a code repository and dataset for the replication of CognitiveOS: https://github.com/Arcwy0/cognitiveos
Artem Lykov, Mikhail Konenkov, Koffivi Fidèle Gbagbe, Mikhail Litvinov, Denis Davletshin, Aleksey Fedoseev, Miguel Altamirano, Robinroy Peter, Dzmitry Tsetserukou
ICRA9
2025 ImpedanceGPT: VLM-driven Impedance Control of Swarm of Mini-drones for Intelligent Navigation in Dynamic Environment
abstract
Swarm robotics plays a crucial role in enabling autonomous operations in dynamic and unpredictable environments. However, a major challenge remains ensuring safe and efficient navigation in environments shared by both dynamic alive (e.g., humans) and dynamic inanimate (e.g., non-living objects) obstacles. In this paper, we propose ImpedanceGPT, a novel system that leverages a Vision-Language Model (VLM) with Retrieval-Augmented Generation (RAG) framework to enable real-time reasoning for adaptive navigation of mini-drone swarm in complex environments. The key innovation of ImpedanceGPT lies in the merging VLM-RAG system with impedance control method, which is an active compliance strategy. This system provides drones with an enhanced semantic understanding of their surroundings and allows them to dynamically adjust impedance control parameters in response to obstacle types and environmental conditions. Our approach not only ensures safe and precise navigation but also improves coordination between drones in the swarm. Experimental evaluations demonstrate the effectiveness of our system. The VLM-RAG framework achieved an obstacle detection and retrieval accuracy of 80% under optimal lighting. In static environments, drones navigated dynamic inanimate obstacles at 1.4 m/s but slowed to 0.7 m/s with increased safety margin around humans. In dynamic environments, speed adjusted to 1.0 m/s near hard obstacles, while reducing to 0.6 m/s with higher deflection region to safely avoid moving humans.Video of ImpedanceGPT: https://youtu.be/JTdeg9bAzL4 Github: https://github.com/Faryal-Batool/ImpedanceGPT
Faryal Batool, Yasheerah Yaqoot, Malaika Zafar, Roohan Ahmed Khan, Muhammad Haris Khan, Aleksey Fedoseev, Dzmitry Tsetserukou
IROS7
2025 Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots
abstract
This paper presents the concept of Industry 6.0, which introduces the world’s first fully automated production system that autonomously handles the entire product design and manufacturing process based on user-provided natural language descriptions. By leveraging generative AI, the system automates critical aspects of production, including product blueprint design, component manufacturing, logistics, and assembly. A heterogeneous swarm of robots, each equipped with individual AI through integration with Large Language Models (LLMs), orchestrates the production process. The robotic system includes manipulator arms, delivery drones, and 3D printers capable of generating assembly blueprints. The system was evaluated using commercial and open source LLMs, operating via APIs and local deployment. A user study demonstrated that the system reduced the average production time to 119.10 minutes, significantly outperforming a team of expert human developers, who averaged 528.64 minutes (an improvement factor of 4.4). Furthermore, in the product blueprinting stage, the system outperformed human CAD operators by an unprecedented factor of 47, completing the task in 0.5 minutes compared to 23.5 minutes. This breakthrough represents a major leap towards fully autonomous manufacturing.
Artem Lykov, Miguel Altamirano, Mikhail Konenkov, Valerii Serpiva, Koffivi Fidèle Gbagbe, Ali Alabbas, Aleksey Fedoseev, Luis Moreno 0007, Muhammad Haris Khan, Ziang Guo, Dzmitry Tsetserukou
IROS11
2025 FADet: A Multi-sensor 3D Object Detection Network based on Local Featured Attention
abstract
Camera, LiDAR, and radar are common perception sensors for autonomous driving tasks. Robust prediction of 3D object detection is optimally based on the fusion of these sensors. Taking advantage of their abilities remains a challenge, because each of these sensors has its own characteristics. Specifically, different sensors present different scales in their corresponding extracted features. To address this problem, considering the feature alignment in different scales, in this paper, we propose FADet, a multi-sensor 3D detection network, which specifically studies the characteristics of different sensors across the dimensions of their data input based on our local featured attention modules. For camera images, we propose a dual-attention-based submodule. For LiDAR point clouds, the triple-attention-based submodule is utilized, while the mixed-attention-based submodule is applied for features of radar points. With local featured attention submodules, our FADet has effective detection results in long-tail and complex scenes from camera, LiDAR and radar input. In the NuScenes validation dataset, FADet achieves state-of-the-art performance on LiDAR-camera object detection tasks with 71.8% NDS and 69.0% mAP, at the same time, on radar-camera object detection tasks with 51.7% NDS and 40.3% mAP.
Ziang Guo, Zakhar Yagudin, Selamawit Asfaw, Artem Lykov, Dzmitry Tsetserukou
IV5
2025 MAGNNET: Multi-Agent Graph Neural Network-Based Efficient Task Allocation for Autonomous Vehicles with Deep Reinforcement Learning
abstract
This paper addresses the challenge of decentralized task allocation within heterogeneous multiagent systems operating under communication constraints. We introduce a novel framework that integrates Graph Neural Networks (GNNs) with a centralized training and decentralized execution (CTDE) paradigm, further enhanced by a tailored Proximal Policy Optimization (PPO) algorithm for multi-agent deep reinforcement learning (MARL). Our approach enables unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) to dynamically allocate tasks efficiently without necessitating central coordination in a 3D grid environment. The framework minimizes total travel time while simultaneously avoiding conflicts in task assignments. For the cost calculation and routing, we employ reservation-based$A^{*}$and$R^{*}$path planners. Experimental results revealed that our method achieves a high 92.5% conflict-free success rate, with only a 7.49% performance gap compared to the centralized Hungarian method, while outperforming the heuristic decentralized baseline based on a greedy approach. Additionally, the framework exhibits scalability with up to 20 agents with allocation processing of 2.8 s and robustness in responding to dynamically generated tasks, underscoring its potential for real-world applications in complex multi-agent scenarios.
Lavanya Ratnabala, Aleksey Fedoseev, Robinroy Peter, Dzmitry Tsetserukou
IV4
2025 UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and Rescue
abstract
Emergency search and rescue (SAR) operations often require rapid and precise target identification in complex environments where traditional manual drone control is inefficient. In order to address these scenarios, a rapid SAR system, UAV-VLRR (Vision-Language-Rapid-Response), is developed in this research. This system consists of two aspects: 1) A multimodal system which harnesses the power of Visual Language Model (VLM) and the natural language processing capabilities of ChatGPT-4o (LLM) for scene interpretation. 2) A non-linear model predictive control (NMPC) with built-in obstacle avoidance for rapid response by a drone to fly according to the output of the multimodal system. This work aims at improving response times in emergency SAR operations by providing a more intuitive and natural approach to the operator to plan the SAR mission while allowing the drone to carry out that mission in a rapid and safe manner. When tested, our approach was faster on an average by 33.75% when compared with an off-the-shelf autopilot and 54.6% when compared with a human pilot. Github: https://github.com/ahsan-mustafa/uav-vlrr Video of UAV-VLRR: https://youtu.be/KJqQGKKt1xY
Yasheerah Yaqoot, Muhammad Ahsan Mustafa, Oleg Sautenkov, Artem Lykov, Valerii Serpiva, Dzmitry Tsetserukou
IV6
2025 NMPC-Lander: Nonlinear MPC with Control Barrier Function for UAV Landing on a Moving Platform
abstract
Quadcopters are versatile aerial robots gaining popularity in numerous critical applications. However, their operational effectiveness is constrained by limited battery life and restricted flight range. To address these challenges, autonomous drone landing on stationary or mobile charging and battery-swapping stations has become an essential capability. In this study, we present NMPC-Lander, a novel control architecture that integrates Nonlinear Model Predictive Control (NMPC) with Control Barrier Functions (CBF) to achieve precise and safe autonomous landing on both static and dynamic platforms. Our approach employs NMPC for accurate trajectory tracking and landing, while simultaneously incorporating CBF to ensure collision avoidance with static obstacles. Experimental evaluations on the real hardware demonstrate high precision in landing scenarios, with an average final position error of 9.0 cm and 11 cm for stationary and mobile platforms, respectively. Notably, NMPC-Lander outperforms the B-spline combined with the A* planning method by nearly threefold in terms of position tracking, underscoring its superior robustness and practical effectiveness. Video of NMPC-Lander: https://youtu.be/JAmYipRCiWo
Amber Batool, Faryal Batool, Roohan Ahmed Khan, Muhammad Ahsan Mustafa, Aleksey Fedoseev, Dzmitry Tsetserukou
SMC6
2025 EEG Study of the Influence of Imagined Temperature Sensations on Neuronal Activity in the Sensorimotor Cortex
abstract
Understanding the neural correlates of sensory imagery is crucial for advancing cognitive neuroscience and developing novel Brain-Computer Interface (BCI) paradigms. This study investigated the influence of imagined temperature sensations (ITS) on neural activity within the sensorimotor cortex. The experimental study involved the evaluation of neural activity using electroencephalography (EEG) during both real thermal stimulation (TS: 40 °C Hot, 20 °C Cold) applied to the participants’ hand, and the mental temperature imagination (ITS) of the corresponding hot and cold sensations. The analysis focused on quantifying the event-related desynchronization (ERD) of the sensorimotor μ-rhythm (8-13 Hz). The experimental results revealed a characteristic μ-ERD localized over central scalp regions (e.g., C3) during both TS and ITS conditions. Although the magnitude of μ-ERD during ITS was slightly lower than during TS, this difference was not statistically significant (p > .05). However, ERD during both ITS and TS was statistically significantly different from the resting baseline (p < .001). These findings demonstrate that imagining temperature sensations engages sensorimotor cortical mechanisms in a manner comparable to actual thermal perception. This insight expands our understanding of the neurophysiological basis of sensory imagery and suggests the potential utility of ITS for non-motor BCI control and neurorehabilitation technologies.
Anton Belichenko, Daria Trinitatova, Aigul Nasibullina, Lev Yakovlev, Dzmitry Tsetserukou
SMC5
2025 HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
abstract
This paper introduces HapticVLM, a novel multimodal system that integrates vision-language reasoning and deep convolutional networks to enable real-time haptic feedback. HapticVLM leverages a ConvNeXt-based material recognition module to generate robust visual embeddings for accurate identification of object materials. A state-of-the-art Vision-Language Model (Qwen2-VL-2B-Instruct) infers ambient temperature from environmental cues. The system synthesizes tactile sensations by delivering vibrotactile feedback through speakers and thermal cues with a Peltier module, thereby bridging the gap between visual perception and tactile experience. Experimental evaluations demonstrate an average recognition accuracy of 84.7% across five distinct auditory-tactile patterns and a temperature estimation accuracy of 86.7% using an 8 °C margin error across 15 scenarios. Although promising, the current study is limited by the use of a small set of patterns and participants. Future work will focus on expanding the range of tactile patterns and increasing user studies to further refine and validate the system’s performance. Overall, HapticVLM presents a significant step toward intelligent, context-aware, multimodal haptic interaction for Virtual Reality (VR) and assistive technologies.
Muhammad Haris Khan, Miguel Altamirano, Dmitrii Iarchuk, Yara Mahmoud, Daria Trinitatova, Issatay Tokmurziyev, Dzmitry Tsetserukou
SMC7
2024 GrainGrasp: Dexterous Grasp Generation with Fine-grained Contact Guidance
abstract
One goal of dexterous robotic grasping is to allow robots to handle objects with the same level of flexibility and adaptability as humans. However, it remains a challenging task to generate an optimal grasping strategy for dexterous hands, especially when it comes to delicate manipulation and accurate adjustment the desired grasping poses for objects of varying shapes and sizes. In this paper, we propose a novel dexterous grasp generation scheme called GrainGrasp that provides fine-grained contact guidance for each fingertip. In particular, we employ a generative model to predict separate contact maps for each fingertip on the object point cloud, effectively capturing the specifics of finger-object interactions. In addition, we develop a new dexterous grasping optimization algorithm that solely relies on the point cloud as input, eliminating the necessity for complete mesh information of the object. By leveraging the contact maps of different fingertips, the proposed optimization algorithm can generate precise and determinable strategies for human-like object grasping. Experimental results confirm the efficiency of the proposed scheme. Our code is available at https://github.com/wmtlab/GrainGrasp.
Fuqiang Zhao, Dzmitry Tsetserukou, Qian Liu 0001
ICRA2
2024 OmniRace: 6D Hand Pose Estimation for Intuitive Guidance of Racing Drone
abstract
This paper presents the OmniRace approach to controlling a racing drone with 6-degree of freedom (DoF) hand pose estimation and gesture recognition. To our knowledge, this is the first technology enabling low-level control of high-speed drones through gestures. OmniRace employs a gesture interface based on computer vision and a deep neural network to estimate 6-DoF hand pose. The advanced machine learning algorithm robustly interprets human gestures, allowing users to control drone motion intuitively. Real-time control tests validate the system’s effectiveness and its potential to revolutionize drone racing and other applications. Experimental results conducted in simulation environment revealed that OmniRace allows the users to complite the UAV race track significantly (by 25.1%) faster and to decrease the length of the test drone path (from 102.9 to 83.7 m). Users preferred the gesture interface for attractiveness (1.57 UEQ score), hedonic quality (1.56 UEQ score), and lower perceived temporal demand (32.0 score in NASA-TLX), while noting the high efficiency (0.75 UEQ score) and low physical demand (19.0 score in NASA-TLX) of the baseline remote controller. The deep neural network attains an average accuracy of 99.75% when applied to both normalized datasets and raw datasets. OmniRace can potentially change the way humans interact with and navigate racing drones in dynamic and complex environments. The source code is available at https://github.com/SerValera/OmniRace.git.
Valerii Serpiva, Aleksey Fedoseev, Sausar Karaf, Ali Alridha Abdulkarim, Dzmitry Tsetserukou
IROS5
2024 HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene
abstract
Many established vision perception systems for autonomous driving scenarios ignore the influence of light conditions, one of the key elements for driving safety. To address this problem, we present HawkDrive, a novel perception system with hardware and software solutions. Hardware that utilizes stereo vision perception, which has been demonstrated to be a more reliable way of estimating depth information than monocular vision, is partnered with the edge computing device Nvidia Jetson Xavier AGX. Our software for low light enhancement, depth estimation, and semantic segmentation tasks, is a transformer-based neural network. Our software stack, which enables fast inference and noise reduction, is packaged into system modules in Robot Operating System 2 (ROS2). Our experimental results have shown that the proposed end-to-end system is effective in improving the depth estimation and semantic segmentation performance. Our dataset and codes will be released at https://github.com/ZionGo6/HawkDrive.
Ziang Guo, Stepan Perminov, Mikhail Konenkov, Dzmitry Tsetserukou
IV4
2024 MARLander: A Local Path Planning for Drone Swarms using Multiagent Deep Reinforcement Learning
abstract
Achieving safe and precise landings for a swarm of drones poses a significant challenge, primarily attributed to conventional control and planning methods. This paper presents the implementation of multi-agent deep reinforcement learning (MADRL) techniques for the precise landing of a drone swarm at relocated target locations. The system is trained in a realistic simulated environment with a maximum linear velocity of 3 m/s in training spaces of 4 m3and deployed utilizing Crazyflie drones with a Vicon indoor localization system. The experimental results revealed that the proposed approach achieved a landing accuracy of 2.26 cm on stationary and 3.93 cm on moving platforms surpassing a baseline method used with a Proportional-integral-derivative (PID) controller with an Artificial Potential Field (APF). This research high-lights drone landing technologies that eliminate the need for analytical centralized systems, potentially offering scalability and revolutionizing applications in logistics, safety, and rescue missions.
Demetros Aschu, Robinroy Peter, Sausar Karaf, Aleksey Fedoseev, Dzmitry Tsetserukou
SMC5
2024 Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous Manipulations
abstract
This research introduces the Bi-VLA (Vision-Language-Action) model, a novel system designed for bimanual robotic dexterous manipulation that seamlessly integrates vision for scene understanding, language comprehension for translating human instructions into executable code, and physical action generation. We evaluated the system's functionality through a series of household tasks, including the preparation of a desired salad upon human request. Bi-VLA demonstrates the ability to interpret complex human instructions, perceive and understand the visual context of ingredients, and execute precise bimanual actions to prepare the requested salad. We assessed the system's performance in terms of accuracy, efficiency, and adaptability to different salad recipes and human preferences through a series of experiments. Our results show a 100 % success rate in generating the correct executable code by the Language Module, a 96.06 % success rate in detecting specific ingredients by the Vision Module, and an overall success rate of 83.4 % in correctly executing user-requested tasks.
Koffivi Fidèle Gbagbe, Miguel Altamirano, Ali Alabbas, Oussama Alyounes, Artem Lykov, Dzmitry Tsetserukou
SMC6
2024 MorphoMove: Bi-Modal Path Planner with MPC-based Path Follower for Multi-Limb Morphogenetic UAV
abstract
This paper discusses developments for a multi-limb morphogenetic UAV, MorphoGear, that is capable of both aerial flight and ground locomotion. A hybrid path planning algorithm based on the A* strategy has been developed, enabling seamless transition between air-to-ground navigation modes, thereby enhancing robot's mobility in complex environments. Moreover, precise path following is achieved during ground locomotion with a Model Predictive Control (MPC) architecture for its novel walking behaviour. Experimental validation was conducted in the Unity simulation environment utilizing Python scripts to compute control values. The algorithm's performance is validated by the Root Mean Squared Error (RMSE) of 0.91 cm and a maximum error of 1.85 cm, as demonstrated by the results. These developments highlight the adaptability of MorphoGear in navigation through cluttered environments, establishing it as a usable tool in autonomous exploration, both aerial and ground-based.
Muhammad Ahsan Mustafa, Yasheerah Yaqoot, Mikhail Martynov, Sausar Karaf, Dzmitry Tsetserukou
SMC5
2024 Dynamic Subgoal Based Path Formation and Task Allocation: A NeuroFleets Approach to Scalable Swarm Robotics
abstract
This paper addresses the challenges of exploration and navigation in unknown environments from the perspective of evolutionary swarm robotics. A key focus is on path formation, which is essential for enabling cooperative swarm robots to navigate effectively. We designed the task allocation and path formation process based on a finite state machine, ensuring systematic decision-making and efficient state transitions. The approach is decentralized, allowing each robot to make decisions independently based on local information, which enhances scalability and robustness. We present a novel subgoal-based path formation method that establishes paths between locations by leveraging visually connected subgoals. Simulation experiments conducted in the Argos simulator show that this method successfully forms paths in the majority of trials. However, inter-collision (traffic) among numerous robots during path formation can negatively impact performance. To address this issue, we propose a task allocation strategy that uses local communication protocols and light signal-based communication to manage robot deployment. This strategy assesses the distance between points and determines the optimal number of robots needed for the path formation task, thereby reducing unnecessary exploration and traffic congestion. The performance of both the subgoal-based path formation method and the task allocation strategy is evaluated by comparing the path length, time, and resource usage against the A * algorithm. Simulation results demonstrate the effectiveness of our approach, highlighting its scalability, robustness, and fault tolerance.
Robinroy Peter, Lavanya Ratnabala, Eugene Yugarajah Andrew Charles, Dzmitry Tsetserukou
SMC4
2024 HyperSurf: Quadruped Robot Leg Capable of Surface Recognition with GRU and Real-to-Sim Transferring
abstract
This paper introduces a system of data collection acceleration and real-to-sim transferring for surface recognition on a quadruped robot. The system features a mechanical single-leg setup capable of stepping on various easily interchangeable surfaces. Additionally, it incorporates a GRU-based Surface Recognition System, inspired by the system detailed in the Dog-Surf paper [1]. This setup facilitates the expansion of dataset collection for model training, enabling data acquisition from hard-to-reach surfaces in laboratory conditions. Furthermore, it opens avenues for transferring surface properties from reality to simulation, thereby allowing the training of optimal gaits for legged robots in simulation environments using a pre-prepared library of digital twins of surfaces. Moreover, enhancements have been made to the GRU-based Surface Recognition System, allowing for the integration of data from both the quadruped robot and the single-leg setup. The dataset and code have been made publicly available.
Sergei Satsevich, Yaroslav Savotin, Danil Belov, Elisaveta Pestova, Artem Erhov, Batyr Khabibullin, Artem Bazhenov, Vyacheslav Kovalev, Aleksey Fedoseev, Dzmitry Tsetserukou
SMC10
2024 GazeRace: Revolutionizing Remote Piloting with Eye-Gaze Control
abstract
This paper presents GazeRace, a novel system that leverages eye-tracking technology for intuitive drone control. Using the MediaPipe library, the system translates eye movements into precise drone commands, enabling effective remote piloting. In testing, GazeRace demonstrated an 18% reduction in drone trajectory length while maintaining competitive speed with traditional controls. The results suggest that this approach enhances control accuracy and reduces user frustration, offering a significant advancement in the field of human-computer interaction and drone navigation.
Issatay Tokmurziyev, Valerii Serpiva, Aleksey Fedoseev, Miguel Altamirano, Dzmitry Tsetserukou
SMC5
2024 Pose estimation in robotic electric vehicle plug-in charging tasks using auto-annotation and deep learning-based keypoint detector
Viktor Rakhmatulin, Miguel Altamirano, Andrei Puchkov, Evgeny Burnaev, Dzmitry Tsetserukou
Eng. Appl. Artif. Intell.5
2024 OmniCharger: CNN-Based Hand Gesture Interface to Operate an Electric Car Charging Robot through Teleconference
abstract
The automation of the car charging process is motivated by the rapid development of technologies for self-driving cars and the increasing importance of ecological transportation units. Automation of this process requires the implementation of Computer Vision (CV) techniques. However, it remains challenging to precisely position the charger plug autonomously due to the sensitivity of CV algorithms to lighting and weather conditions. We introduce a novel robotic operation system based on hand gesture recognition through teleconferencing software. The users, connected by teleconference, use their hand gestures to teleoperate the electric plug located on the collaborative robot end-effector. We conducted a user study to evaluate the system performance and suitability using OmniCharger and two baseline interfaces (a UR10 Teach Pendant and a Logitech F710 Wireless Gamepad). Except for two trials, all the users were able to locate the plug inside of a 5 cm target using the interfaces. The distance to the target and the orientation error did not present statistically significant differences ( \(p=0.1099 \gt 0.05\) and \(p=0.0903 \gt 0.05\) , respectively) in the use of the three interfaces. The NASA-TLX questionnaire results showed low values in all the sub-classes, the SUS results rated the usability of the proposed interface above average (68%), and the UEQ showed excellent performance of the OmniCharger interface in the attractiveness, stimulation, and novelty attributes.
Miguel Altamirano, Viktor Rakhmatulin, Aleksey Fedoseev, Oleg Sautenkov, Oussama Alyounes, Andrei Puchkov, Dzmitry Tsetserukou
ACM Trans. Hum. Robot Interact.7
2023 Multi-Sensor Large-Scale Dataset for Multi-View 3D Reconstruction
abstract
We present a new multi-sensor dataset for multi-view 3D surface reconstruction. It includes registered RGB and depth data from sensors of different resolutions and modalities: smartphones, Intel RealSense, Microsoft Kinect, industrial cameras, and structured-light scanner. The scenes are selected to emphasize a diverse set of material properties challenging for existing algorithms. We provide around 1.4 million images of 107 different scenes acquired from 100 viewing directions under 14 lighting conditions. We expect our dataset will be useful for evaluation and training of 3D reconstruction algorithms and for related tasks. The dataset is available at skol tech3d. appliedai. tech.
Oleg Voynov, Gleb Bobrovskikh, Pavel A. Karpyshev, Saveliy Galochkin, Andrei-Timotei Ardelean, Arseniy Bozhenko, Ekaterina Karmanova, Pavel Kopanev, Yaroslav Labutin-Rymsho, Ruslan Rakhimov, Aleksandr Safin, Valerii Serpiva, Alexey Artemov, Evgeny Burnaev, Dzmitry Tsetserukou, Denis Zorin
CVPR15
2023 ArUcoGlide: a Novel Wearable Robot for Position Tracking and Haptic Feedback to Increase Safety During Human-Robot Interaction
abstract
The current capabilities of robotic systems make human collaboration necessary to accomplish complex tasks effectively. In this work, we are introducing a framework to ensure safety in a human-robot collaborative environment. The system is composed of a wearable 2-DoFs robot, a low-cost and easy-to-install tracking system, and a collision avoidance algorithm based on the Artificial Potential Field (APF). The wearable robot is designed to hold a fiducial marker and maintain its visibility to the tracking system, which, in turn, localizes the user’s hand with good accuracy and low latency and provides haptic feedback to the user. The system is designed to enhance the performance of collaborative tasks while ensuring user safety. Three experiments were carried out to evaluate the performance of the proposed system. The first one evaluated the accuracy of the tracking system. The second experiment analyzed human-robot behavior during an imminent collision. The third experiment evaluated the system in a collaborative activity in a shared working environment. The results show that the implementation of the introduced system reduces the operation time by 16% and increases the average distance between the user’s hand and the robot by 5 cm.
Ali Alabbas, Miguel Altamirano, Oussama Alyounes, Dzmitry Tsetserukou
ETFA4
2023 Hierarchical Whole-body Control of the cable-Suspended Aerial Manipulator endowed with Winch-based Actuation
abstract
During operation, aerial manipulation systems are affected by various disturbances. Among them is a gravitational torque caused by the weight of the robotic arm. Common propeller-based actuation is ineffective against such disturbances because of possible overheating and high power consumption. To overcome this issue, in this paper we propose a winch-based actuation for the crane-stationed cable-suspended aerial manipulator. Three winch-controlled suspension rigging cables produce a desired cable tension distribution to generate a wrench that reduces the effect of gravitational torque. In order to coordinate the robotic arm and the winch-based actuation, a model-based hierarchical whole-body controller is adapted. It resolves two tasks: keeping the robotic arm end-effector at the desired pose and shifting the system center of mass in the location with zero gravitational torque. The performance of the introduced actuation system as well as control strategy is validated through experimental studies.
Yuri S. Sarkisov, Andre Coelho, Maihara Santos, Minjun Kim 0003, Dzmitry Tsetserukou, Christian Ott 0001, Konstantin Kondak
ICRA5
2023 MorphoLander: Reinforcement Learning Based Landing of a Group of Drones on the Adaptive Morphogenetic UAV
abstract
This paper focuses on a novel robotic system MorphoLander representing heterogeneous swarm of drones for exploring rough terrain environments. The morphogenetic leader drone is capable of landing on uneven terrain, traversing it, and maintaining horizontal position to deploy smaller drones for extensive area exploration. After completing their tasks, these drones return and land back on the landing pads of MorphoGear. The reinforcement learning algorithm was developed for a precise landing of drones on the leader robot that either remains static during their mission or relocates to the new position. Several experiments were conducted to evaluate the performance of the developed landing algorithm under both even and uneven terrain conditions. The experiments revealed that the proposed system results in high landing accuracy of 0.5 cm when landing on the leader drone under even terrain conditions and 2.35 cm under uneven terrain conditions. MorphoLander has the potential to significantly enhance the efficiency of the industrial inspections, seismic surveys, and rescue missions in highly cluttered and unstructured environments.
Sausar Karaf, Aleksey Fedoseev, Mikhail Martynov, Zhanibek Darush, Aleksei Shcherbak, Dzmitry Tsetserukou
SMC6
2023 DNFOMP: Dynamic Neural Field Optimal Motion Planner for Navigation of Autonomous Robots in Cluttered Environment
abstract
Motion planning in dynamically changing environments is one of the most complex challenges in autonomous driving. Safety is a crucial requirement, along with driving comfort and speed limits. While classical sampling-based, lattice-based, and optimization-based planning methods can generate smooth and short paths, they often do not consider the dynamics of the environment. Some techniques do consider it, but they rely on updating the environment on-the-go rather than explicitly accounting for the dynamics, which is not suitable for self-driving. To address this, we propose a novel method based on the Neural Field Optimal Motion Planner (NFOMP), which outperforms state-of-the-art approaches in terms of normalized curvature and the number of cusps. Our approach embeds previously known moving obstacles into the neural field collision model to account for the dynamics of the environment. We also introduce time profiling of the trajectory and non-linear velocity constraints by adding Lagrange multipliers to the trajectory loss function. We applied our method to solve the optimal motion planning problem in an urban environment using the BeamNG.tech driving simulator. An autonomous car drove the generated trajectories in three city scenarios while sharing the road with the obstacle vehicle. Our evaluation shows that the maximum acceleration the passenger can experience instantly is −7.5 m/s2and that 89.6% of the driving time is devoted to normal driving with accelerations below 3.5 m/s2. The driving style is characterized by 46.0% and 31.4% of the driving time being devoted to the light rail transit style and the moderate driving style, respectively.
Maksim Katerishich, Mikhail Kurenkov, Sausar Karaf, Artem Nenashev, Dzmitry Tsetserukou
SMC5
2023 LocoNeRF: A NeRF-Based Approach for Local Structure from Motion for Precise Localization
abstract
Visual localization is a critical task in mobile robotics, and researchers are continuously developing new approaches to enhance its efficiency. In this article, we propose a novel approach to improve the accuracy of visual localization using Structure from Motion (SfM) techniques. We highlight the limitations of global SfM, which suffers from high latency, and the challenges of local SfM, which requires large image databases for accurate reconstruction. To address these issues, we propose utilizing Neural Radiance Fields (NeRF), as opposed to image databases, to cut down on the space required for storage. We suggest that sampling reference images around the prior query position can lead to further improvements. We evaluate the accuracy of our proposed method against ground truth obtained using LIDAR and Advanced Lidar Odometry and Mapping in Real-time (A-LOAM), and compare its storage usage against local SfM with COLMAP in the conducted experiments. Our proposed method achieves an accuracy of 0.068 meters compared to the ground truth, which is slightly lower than the most advanced method COLMAP, which has an accuracy of 0.022 meters. However, the size of the database required for COLMAP is 400 megabytes, whereas the size of our NeRF model is only 160 megabytes. Finally, we perform an ablation study to assess the impact of using reference images from the NeRF reconstruction.
Artem Nenashev, Mikhail Kurenkov, Andrei Potapov, Iana Zhura, Maksim Katerishich, Dzmitry Tsetserukou
SMC6
2023 GHACPP: Genetic-Based Human-Aware Coverage Path Planning Algorithm for Autonomous Disinfection Robot
abstract
Numerous mobile robots with mounted Ultraviolet-C (UV-C) lamps were developed recently, yet they cannot work in the same space as humans without irradiating them by UV-C. This paper proposes a novel modular and scalable Human-Aware Genetic-based Coverage Path Planning algorithm (GHACPP), that aims to solve the problem of disinfecting of unknown environments by UV-C irradiation and preventing human eyes and skin from being harmed. The proposed genetic-based algorithm alternates between the stages of exploring a new area, generating parts of the resulting disinfection trajectory, called mini-trajectories, and updating the current state around the robot. The system performance in effectiveness and human safety is validated and compared with one of the latest state-of-the-art online coverage path planning algorithms called SimExCoverage-STC. The experimental results confirmed both the high level of safety for humans and the efficiency of the developed algorithm in terms of decrease of path length (by 37.1%), number (39.5%) and size (35.2%) of turns, and time (7.6%) to complete the disinfection task, with a small loss in the percentage of area covered (0.6 %), in comparison with the state-of-the-art approach.
Stepan Perminov, Ivan Kalinov, Dzmitry Tsetserukou
SMC3
2023 POA: Passable Obstacles Aware Path-Planning Algorithm for Navigation of a Two-Wheeled Robot in Highly Cluttered Environments
abstract
This paper focuses on Passable Obstacles Aware (POA) planner - a novel navigation method for two-wheeled robots in a highly cluttered environment. The navigation algorithm detects and classifies objects to distinguish two types of obstacles - passable and unpassable. Our algorithm allows two-wheeled robots to find a path through passable obstacles. Such a solution helps the robot working in areas inaccessible to standard path planners and find optimal trajectories in scenarios with a high number of objects in the robot's vicinity. The POA planner can be embedded into other planning algorithms and enables them to build a path through obstacles. Our method decreases path length and the total travel time to the final destination up to 43 % and 39 %, respectively, comparing to standard path planners such as GVD, A*, and RRT*.
Alexander A. Petrovsky, Yomna Youssef, Kirill Myasoedov, Artem Timoshenko, Vladimir Guneavoi, Ivan Kalinov, Dzmitry Tsetserukou
SMC7
2023 AirTouch: Towards Safe Human-Robot Interaction Using Air Pressure Feedback and IR Mocap System
abstract
The growing use of robots in urban environments has raised concerns about potential safety hazards, especially in public spaces where humans and robots may interact. In this paper, we present a system for safe human-robot interaction that combines an infrared (IR) camera with a wearable marker and airflow potential field. IR cameras enable real-time detection and tracking of humans in challenging environments, while controlled airflow creates a physical barrier that guides humans away from dangerous proximity to robots without the need for wearable devices. A preliminary experiment was conducted to measure the accuracy of the perception of safety barriers rendered by controlled air pressure. In a second experiment, we evaluated our approach in an imitation scenario of an interaction between an inattentive person and an autonomous robotic system. Experimental results show that the proposed system significantly improves a participant's ability to maintain a safe distance from the operating robot compared to trials without the system.
Viktor Rakhmatulin, Denis Grankin, Mikhail Konenkov, Sergei Davidenko, Daria Trinitatova, Oleg Sautenkov, Dzmitry Tsetserukou
SMC7
2023 NeuroSwarm: Multi-Agent Neural 3D Scene Reconstruction and Segmentation with UAV for Optimal Navigation of Quadruped Robot
abstract
Quadruped robots have the distinct ability to adapt their body and step height to navigate through cluttered environments. Nonetheless, for these robots to utilize their full potential in real-world scenarios, they require awareness of their environment and obstacle geometry. We propose a novel multi-agent robotic system that incorporates cutting-edge technologies. The proposed solution features a 3D neural reconstruction algorithm that enables navigation of a quadruped robot in both static and semi-static environments. The prior areas of the environment are also segmented according to the quadruped robots' abilities to pass them. Moreover, we have developed an adaptive neural field optimal motion planner (AN-FOMP) that considers both collision probability and obstacle height in 2D space. Our new navigation and mapping approach enables quadruped robots to adjust their height and behavior to navigate under arches and push through obstacles with smaller dimensions. The multi-agent mapping operation has proven to be highly accurate, with an obstacle reconstruction precision of 82%. Moreover, the quadruped robot can navigate with 3D obstacle information and the ANFOMP system, resulting in a 33.3% reduction in path length and a 70% reduction in navigation time. The developed project with code is available on GitHub11https://github.com/Iana-Zhura/NeuroSwarm.
Iana Zhura, Denis Davletshin, Nipun Dhananjaya Weerakkodi Mudalige, Aleksey Fedoseev, Robinroy Peter, Dzmitry Tsetserukou
SMC6
2023 Hierarchical Visual Localization Based on Sparse Feature Pyramid for Adaptive Reduction of Keypoint Map Size
abstract
Visual localization is a fundamental task for a wide range of applications in the field of robotics. Yet, it is still a complex problem with no universal solution, and the existing approaches are difficult to scale: most state-of-the-art solutions are unable to provide accurate localization without a significant amount of storage space. We propose a hierarchical, low-memory approach to localization based on keypoints with different descriptor lengths. It becomes possible with the use of the developed unsupervised neural network, which predicts a feature pyramid with different descriptor lengths for images. This structure allows applying coarse-to-fine paradigms for localization based on keypoint map, and varying the accuracy of localization by changing the type of the descriptors used in the pipeline. Our approach achieves comparable results in localization accuracy and a significant reduction in memory consumption (up to 16 times) among state-of-the-art methods.
Andrei Potapov, Mikhail Kurenkov, Pavel A. Karpyshev, Evgeny Yudin, Alena Savinykh, Evgeny Kruzhkov, Dzmitry Tsetserukou
VTC2023-Spring7
2023 CloudVision: DNN-based Visual Localization of Autonomous Robots using Prebuilt LiDAR Point Cloud
abstract
In this study, we propose a novel visual localization approach to accurately estimate six degrees of freedom (6-DoF) poses of the robot within the 3D LiDAR map based on visual data from an RGB camera. The 3D map is obtained utilizing an advanced LiDAR-based simultaneous localization and mapping (SLAM) algorithm capable of collecting a precise sparse map. The features extracted from the camera images are compared with the points of the 3D map, and then the geometric optimization problem is being solved to achieve precise visual localization. Our approach allows employing a scout robot equipped with an expensive LiDAR only once — for mapping of the environment, and multiple operational robots with only RGB cameras onboard — for performing mission tasks, with the localization accuracy higher than common camera-based solutions. The proposed method was tested on the custom dataset collected in the Skolkovo Institute of Science and Technology (Skoltech). During the process of assessing the localization accuracy, we managed to achieve centimeter-level accuracy; the median translation error was as low as 1.3 cm. The precise positioning achieved with only cameras makes possible the usage of autonomous mobile robots to solve the most complex tasks that require high localization accuracy.
Evgeny Yudin, Pavel A. Karpyshev, Mikhail Kurenkov, Alena Savinykh, Andrei Potapov, Evgeny Kruzhkov, Dzmitry Tsetserukou
VTC2023-Spring7
2023 SwipeBot: DNN-based Autonomous Robot Navigation among Movable Obstacles in Cluttered Environments
abstract
In this paper, we propose a novel approach to wheeled robot navigation through an environment with movable obstacles. A robot exploits knowledge about different obstacle classes and selects the minimally invasive action to perform to clear the path. We trained a convolutional neural network (CNN), so the robot can classify an RGB-D image and decide whether to push a blocking object and which force to apply. After known objects are segmented, they are being projected to a cost-map, and a robot calculates an optimal path to the goal. If the blocking objects are allowed to be moved, a robot drives through them while pushing them away. We implemented our algorithm in ROS, and an extensive set of simulations showed that the robot successfully overcomes the blocked regions. Our approach allows a robot to successfully build a path through regions, where it would have stuck with traditional path-planning techniques.
Nikolay Zherdev, Mikhail Kurenkov, Kristina Belikova, Dzmitry Tsetserukou
VTC2023-Spring4
2022 DroneARchery: Human-Drone Interaction through Augmented Reality with Haptic Feedback and Multi-UAV Collision Avoidance Driven by Deep Reinforcement Learning
abstract
We propose a novel concept of augmented reality (AR) human-drone interaction driven by RL-based swarm behavior to achieve intuitive and immersive control of a swarm formation of unmanned aerial vehicles. The DroneARchery system developed by us allows the user to quickly deploy a swarm of drones, generating flight paths simulating archery. The haptic interface LinkGlide delivers a tactile stimulus of the bowstring tension to the forearm to increase the precision of aiming. The swarm of released drones dynamically avoids collisions between each other, the drone following the user, and external obstacles with behavior control based on deep reinforcement learning. The developed concept was tested in the scenario with a human, where the user shoots from a virtual bow with a real drone to hit the target. The human operator observes the ballistic trajectory of the drone in an AR and achieves a realistic and highly recognizable experience of the bowstring tension through the haptic display. The experimental results revealed that the system improves trajectory prediction accuracy by 63.3% through applying AR technology and conveying haptic feedback of pulling force. DroneARchery users highlighted the naturalness (4.3 out of 5 point Likert scale) and increased confidence (4.7 out of 5) when controlling the drone. We have designed the tactile patterns to present four sliding distances (tension) and three applied force levels (stiffness) of the haptic display. Users demonstrated the ability to distinguish tactile patterns produced by the haptic display representing varying bowstring tension(average recognition rate is of 72.8%) and stiffness (average recognition rate is of 94.2%). The novelty of the research is the development of an AR-based approach for drone control that does not require special skills and training from the operator. In the future, the proposed interaction can be applied in various fields, for example, for fast swarm deployment in search and rescue missions, crop monitoring, inspection and maintenance.
Ekaterina Dorzhieva, Ahmed Baza, Ayush Gupta 0003, Aleksey Fedoseev, Miguel Altamirano, Ekaterina Karmanova, Dzmitry Tsetserukou
ISMAR7
2022 HyperGuider: Virtual Reality Framework for Interactive Path Planning of Quadruped Robot in Cluttered and Multi-Terrain Environments
abstract
Quadruped platforms have become an active topic of research due to their high mobility and traversability in rough terrain. However, it is highly challenging to determine whether the clattered environment could be passed by the robot and how exactly its path should be calculated. Moreover, the calculated path may pass through areas with dynamic objects or environments that are dangerous for the robot or people around. Therefore, we propose a novel conceptual approach of teaching quadruped robots navigation through user-guided path planning in virtual reality (VR). Our system contains both global and local path planners, allowing robot to generate path through iterations of learning. The VR interface allows user to interact with environment and to assist quadruped robot in challenging scenarios. The results of comparison experiments show that cooperation between human and path planning algorithms can increase the computational speed of the algorithm by 35.58% in average, and non-critically increasing of the path length (average of 6.66%) in test scenario. Additionally, users described VR interface as not requiring physical demand (2.3 out of 10) and highly evaluated their performance (7.1 out of 10). The ability to find a less optimal but safer path remains in demand for the task of navigating in a cluttered and unstructured environment.
Ildar Babataev, Aleksey Fedoseev, Nipun Weerakkodi, Elena Nazarova, Dzmitry Tsetserukou
SMC5
2022 SwarMan: Anthropomorphic Swarm of Drones Avatar with Body Tracking and Deep Learning-Based Gesture Recognition
abstract
Anthropomorphic robot avatars present a conceptually novel approach to remote affective communication, allowing people across the world a wider specter of emotional and social exchanges over traditional 2D and 3D image data. However, there are several limitations of current telepresence robots, such as the high weight, complexity of the system that prevents its fast deployment, and the limited workspace of the avatars mounted on either static or wheeled mobile platforms. In this paper, we present a novel concept of telecommunication through a robot avatar based on an anthropomorphic swarm of drones; SwarMan. The developed system consists of nine nanocopters controlled remotely by the operator through a gesture recognition interface. SwarMan allows operators to communicate by directly following their motions and by recognizing one of the prerecorded emotional patterns, thus rendering the captured emotion as illumination on the drones. (a) The LSTM MediaPipe network was trained on a collected dataset of 600 short videos with five emotional gestures. The accuracy of achieved emotion recognition was 97% on the test dataset.As communication through the swarm avatar significantly changes the visual appearance of the operator, we investigated the ability of the users to recognize and respond to emotions performed by the swarm of drones. The experimental results revealed a high consistency between the users in rating emotions. Additionally, users indicated low physical demand (2.25 on the Likert scale) and were satisfied with their performance (1.38 on the Likert scale) when communicating by the SwarMan interface.
Ahmed Baza, Ayush Gupta 0003, Ekaterina Dorzhieva, Aleksey Fedoseev, Dzmitry Tsetserukou
SMC5
2022 DandelionTouch: High Fidelity Haptic Rendering of Soft Objects in VR by a Swarm of Drones
abstract
To achieve high fidelity haptic rendering of soft objects in a high mobility virtual environment, we propose a novel haptic display DandelionTouch. The tactile actuators are delivered to the fingertips of the user by a swarm of drones. Users of DandelionTouch are capable of experiencing tactile feedback in a large space that is not limited by the device’s working area. Importantly, they will not experience muscle fatigue during long interactions with virtual objects. Hand tracking and swarm control algorithm allow guiding the swarm with hand motions and avoid collisions inside the formation.Several topologies of impedance connection between swarm units were investigated in this research. The experiment, in which drones performed a point following task on a square trajectory in real-time, revealed that drones connected in a Star topology performed the trajectory with low mean positional error (RMSE decreased by 20.6% in comparison with other impedance topologies and by 40.9% in comparison with potential field-based swarm control). The achieved velocities of the drones in all formations with impedance behavior were 28% higher than for the swarm controlled with the potential field algorithm.Additionally, the perception of several vibrotactile patterns was evaluated in a user study with 7 participants. The study has shown that the proposed combination of temporal delay and frequency modulation allows users to successfully recognize the surface property and motion direction in VR simultaneously (mean recognition rate of 70%, maximum of 93%). Dandelion-Touch suggests a new type of haptic feedback in VR systems where no hand-held or wearable interface is required.
Aleksey Fedoseev, Ahmed Baza, Ayush Gupta 0003, Ekaterina Dorzhieva, Riya Neelesh Gujarathi, Dzmitry Tsetserukou
SMC6
2022 MeSLAM: Memory Efficient SLAM based on Neural Fields
abstract
Existing Simultaneous Localization and Mapping (SLAM) approaches are limited in their scalability due to growing map size in long-term robot operation. Moreover, processing such maps for localization and planning tasks leads to the increased computational resources required onboard. To address the problem of memory consumption in long-term operation, we develop a novel real-time SLAM algorithm, MeSLAM, that is based on neural field implicit map representation. It combines the proposed global mapping strategy, including neural networks distribution and region tracking, with an external odometry system. As a result, the algorithm is able to efficiently train multiple networks representing different map regions and track poses accurately in large-scale environments. Experimental results show that the accuracy of the proposed approach is comparable to the state-of-the-art methods (on average, 6.6 cm on TUM RGB-D sequences) and outperforms the baseline, iMAP*. Moreover, the proposed SLAM approach provides the most compact-sized maps without details distortion (1.9 MB to store 57 m3) among the state-of-the-art SLAM approaches.
Evgeny Kruzhkov, Alena Savinykh, Pavel A. Karpyshev, Mikhail Kurenkov, Evgeny Yudin, Andrei Potapov, Dzmitry Tsetserukou
SMC7
2022 HyperDog: An Open-Source Quadruped Robot Platform Based on ROS2 and micro-ROS
abstract
Nowadays, design and development of legged quadruped robots is a quite active area of scientific research. In fact, the legged robots have become popular due to their capabilities to adapt to harsh terrains and diverse environmental conditions in comparison to other mobile robots. With the higher demand for legged robot experiments, more researches and engineers need an affordable and quick way of locomotion algorithm development. In this paper, we present a new open source quadruped robot HyperDog platform, which features 12 RC servo motors, onboard NVIDIA Jetson nano computer and STM32F4 Discovery board. HyperDog is an open-source platform for quadruped robotic software development, which is based on Robot Operating System 2 (ROS2) and micro-ROS. Moreover, the HyperDog is a quadrupedal robotic dog entirely built from 3D printed parts and carbon fiber, which allows the robot to have light weight and good strength. The idea of this work is to demonstrate an affordable and customizable way of robot development and provide researches and engineers with the legged robot platform, where different algorithms can be tested and validated in simulation and real environment. The developed project with code is available on GitHub.1
Nipun Dhananjaya Weerakkodi Mudalige, Iana Zhura, Ildar Babataev, Elena Nazarova, Aleksey Fedoseev, Dzmitry Tsetserukou
SMC6
2022 HyperPalm: DNN-based hand gesture recognition interface for intelligent communication with quadruped robot in 3D space
abstract
Nowadays, autonomous mobile robots support people in many areas where human presence either redundant or too dangerous. They have successfully proven themselves in expeditions, gas industry, mines, warehouses, etc. However, even legged robots may stuck in rough terrain conditions requiring human cognitive abilities to navigate the system. While gamepads and keyboards are convenient for wheeled robot control, the quadruped robot in 3D space can move along all linear coordinates and Euler angles, requiring at least 12 buttons for independent control of their DoF. Therefore, more convenient interfaces of control are required.In this paper we present HyperPalm: a novel gesture interface for intuitive human-robot interaction with quadruped robots. Without additional devices, the operator has full position and orientation control of the quadruped robot in 3D space through hand gesture recognition with only 5 gestures and 6 DoF hand motion.The experimental results revealed to classify 5 static gestures with high accuracy (96.5%), accurately predict the position of the 6D position of the hand in three-dimensional space. The absolute linear deviation Root mean square deviation (RMSD) of the proposed approach is 11.7 mm, which is almost 50% lower than for the second tested approach, the absolute angular deviation RMSD of the proposed approach is 2.6 degrees, which is almost 27% lower than for the second tested approach. Moreover, the user study was conducted to explore user’s subjective experience from human-robot interaction through the proposed gesture interface. The participants evaluated their interaction with HyperPalm as intuitive (2.0), not causing frustration (2.63), and requiring low physical demand (2.0).
Elena Nazarova, Ildar Babataev, Nipun Weerakkodi, Aleksey Fedoseev, Dzmitry Tsetserukou
SMC5
2022 SwarmHive: Heterogeneous Swarm of Drones for Robust Autonomous Landing on Moving Robot
abstract
The paper focuses on a heterogeneous swarm of drones to achieve a dynamic landing of formation on a moving robot. This challenging task was not yet achieved by scientists. The key technology is that instead of facilitating each agent of the swarm of drones with computer vision that considerably increases the payload and shortens the flight time, we propose to install only one camera on the leader drone. The follower drones receive the commands from the leader UAV and maintain a collision-free trajectory with the artificial potential field. The experimental results revealed a high accuracy of the swarm landing on a static mobile platform (RMSE of 4.48 cm). RMSE of swarm landing on the mobile platform moving with the maximum velocities of 1.0 m/s and 1.5 m/s equals 8.76 cm and 8.98 cm, respectively. The proposed SwarmHive technology will allow the time-saving landing of the swarm for further drone recharging. This will make it possible to achieve self-sustainable operation of a multi-agent robotic system for such scenarios as rescue operations, inspection and maintenance, autonomous warehouse inventory, cargo delivery, and etc.
Ayush Gupta 0003, Ahmed Baza, Ekaterina Dorzhieva, Mert Alper, Mariia Makarova, Stepan Perminov, Aleksey Fedoseev, Dzmitry Tsetserukou
VTC Spring8
2022 DogTouch: CNN-based Recognition of Surface Textures by Quadruped Robot with High Density Tactile Sensors
abstract
The ability to perform locomotion in various terrains is critical for legged robots. However, the robot has to have a better understanding of the surface it is walking on to perform robust locomotion on different terrains. Animals and humans are able to recognize the surface with the help of the tactile sensation on their feet. Although, the foot tactile sensation for legged robots has not been much explored. This paper presents research on a novel quadruped robot DogTouch with tactile sensing feet (TSF). TSF allows the recognition of different surface textures utilizing a tactile sensor and a convolutional neural network (CNN). The experimental results show a sufficient validation accuracy of 74.37% for our trained CNN-based model, with the highest recognition for line patterns of 90%. In the future, we plan to improve the prediction model by presenting surface samples with the various depths of patterns and applying advanced Deep Learning and Shallow learning models for surface recognition.Additionally, we propose a novel approach to navigation of quadruped and legged robots. We can arrange the tactile paving textured surface (similar that used for blind or visually impaired people). Thus, DogTouch will be capable of locomotion in unknown environment by just recognizing the specific tactile patterns which will indicate the straight path, left or right turn, pedestrian crossing, road, and etc. That will allow robust navigation regardless of lighting condition. Future quadruped robots equipped with visual and tactile perception system will be able to safely and intelligently navigate and interact in the unstructured indoor and outdoor environment.
Nipun Dhananjaya Weerakkodi Mudalige, Elena Nazarova, Ildar Babataev, Pavel Kopanev, Aleksey Fedoseev, Miguel Altamirano, Dzmitry Tsetserukou
VTC Spring7
2022 DarkSLAM: GAN-assisted Visual SLAM for Reliable Operation in Low-light Conditions
abstract
Existing visual SLAM approaches are sensitive to illumination, with their precision drastically falling in dark conditions due to feature extractor limitations. The algorithms currently used to overcome this issue are not able to provide reliable results due to poor performance and noisiness, and the localization quality in dark conditions is still insufficient for practical use. In this paper, we present a novel SLAM method capable of working in low light using Generative Adversarial Network (GAN) preprocessing module to enhance the light conditions on input images, thus improving the localization robustness. The proposed algorithm was evaluated on a custom indoor dataset consisting of 14 sequences with varying illumination levels and ground truth data collected using a motion capture system. According to the experimental results, the reliability of the proposed approach remains high even in extremely low light conditions, providing 25.1% tracking time on darkest sequences, whereas existing approaches achieve tracking only 0.6% of the sequence time.
Alena Savinykh, Mikhail Kurenkov, Evgeny Kruzhkov, Evgeny Yudin, Andrei Potapov, Pavel A. Karpyshev, Dzmitry Tsetserukou
VTC Spring7
2021 ChromoUpdate: Fast Design Iteration of Photochromic Color Textures Using Grayscale Previews and Local Color Updates
abstract
ChromoUpdate is a texture transfer system for fast design iteration. For the early stages of design, ChromoUpdate provides a fast grayscale preview that enables a texture to be transferred in under one minute. Once designers are satisfied with the grayscale texture, ChromoUpdate supports designers in coloring the texture by transitioning individual pixels directly to a desired target color. Finally, if designers need to make a change to the color texture already transferred, ChromoUpdate can quickly transition individual pixels from one color to a new target color. ChromoUpdate accomplishes this by (1) using a UV projector rather than a UV LED, which enables pixels to be saturated individually rather than resetting the entire texture to black, and (2) providing two new texture transfer algorithms that allow for fast grayscale previews and color-to-color transitions. Our evaluation shows a significant increase in texture transfer speed for both the grayscale preview (89%) and color-to-color updates (11%).
Michael Wessely, Yuhua Jin, Cattalyya Nuengsigkapian, Aleksei Kashapov, Isabel P. S. Qamar, Dzmitry Tsetserukou, Stefanie Mueller 0001
CHI6
2021 DeepScanner: a Robotic System for Automated 2D Object Dataset Collection with Annotations
abstract
In the proposed study, we describe the possibility of automated dataset collection using an articulated robot. The proposed technology reduces the number of pixel errors on a polygonal dataset and the time spent on manual labeling of 2D objects. The paper describes a novel automatic dataset collection and annotation system, and compares the results of automated and manual dataset labeling. Our approach increases the speed of data labeling 240-fold, and improves the accuracy compared to manual labeling 13-fold. We also present a comparison of metrics for training a neural network on a manually annotated and an automatically collected dataset.
Valeriy Ilin, Ivan Kalinov, Pavel A. Karpyshev, Dzmitry Tsetserukou
ETFA4
2021 UltraBot: Autonomous Mobile Robot for Indoor UV-C Disinfection with Non-trivial Shape of Disinfection Zone
abstract
The paper focuses on the development of an autonomous disinfection robot UltraBot to reduce COVID-19 transmission along with other harmful bacteria and viruses. The motivation behind the research is to develop such a robot that is capable of performing disinfection tasks without the use of harmful sprays and chemicals that can leave residues and require airing the room afterward for a long time. UltraBot technology has the potential to offer the most optimal autonomous disinfection performance along with taking care of people, keeping them from getting under the UV-C radiation. The paper highlights UltraBot's mechanical and electrical design as well as disinfection performance. The conducted experiments demonstrate the effectiveness of robot disinfection ability and actual disinfection area per each side with UV-C lamp array. The disinfection effectiveness results show actual performance for the multi-pass technique that provides 1-log reduction with combined direct UV-C exposure and ozone-based air purification after two robot passes at a speed of 0.14 m/s. This technique has the same performance as ten minutes static disinfection. Finally, we have calculated the nontrivial form of the robot disinfection zone by two consecutive experiment to produce optimal path planning and to provide full disinfection in selected areas.
Nikita Mikhailovskiy, Alexander Sedunin, Stepan Perminov, Ivan Kalinov, Dzmitry Tsetserukou
ETFA5
2021 MobileCharger: an Autonomous Mobile Robot with Inverted Delta Actuator for Robust and Safe Robot Charging
abstract
MobileCharger is a novel mobile charging robot with an Inverted Delta actuator for safe and robust energy transfer between two mobile robots. The RGB-D camera-based computer vision system allows to detect the electrodes on the target mobile robot using a convolutional neural network (CNN). The embedded high-fidelity tactile sensors are applied to estimate the misalignment between the electrodes on the charger mechanism and the electrodes on the main robot using CNN based on pressure data on the contact surfaces. Thus, the developed vision-tactile perception system allows precise positioning of the end effector of the actuator and ensures a reliable connection between the electrodes of the two robots. The experimental results showed high average precision (84.2%) for electrode detection using CNN. The percentage of successful trials of the CNN-based electrode search algorithm reached 83% and the average execution time accounted for 60 s. MobileCharger could introduce a new level of charging systems and increase the prevalence of autonomous mobile robots.
Iaroslav Okunevich, Daria Trinitatova, Pavel Kopanev, Dzmitry Tsetserukou
ETFA4
2021 SwarmPlay: Interactive Tic-tac-toe Board Game with Swarm of Nano-UAVs driven by Reinforcement Learning
abstract
Reinforcement learning (RL) methods have been actively applied in the field of robotics, allowing the system itself to find a solution for a task otherwise requiring a complex decision-making algorithm. In this paper, we present a novel RL-based Tic-tac-toe scenario, i.e. SwarmPlay, where each playing component is presented by an individual drone that has its own mobility and swarm intelligence to win against a human player. Thus, the combination of challenging swarm strategy and human-drone collaboration aims to make the games with machines tangible and interactive. Although some research on AI for board games already exists, e.g., chess, the SwarmPlay technology has the potential to offer much more engagement and interaction with the user as it proposes a multi-agent swarm instead of a single interactive robot. We explore user’s evaluation of RL-based swarm behavior in comparison with the game theory-based behavior. The preliminary user study revealed that participants were highly engaged in the game with drones (70% put a maximum score on the Likert scale) and found it less artificial compared to the regular computer-based systems (80%). The affection of the user’s game perception from its outcome was analyzed and put under discussion. User study revealed that SwarmPlay has the potential to be implemented in a wider range of games, significantly improving human-drone interactivity.
Ekaterina Karmanova, Valerii Serpiva, Stepan Perminov, Aleksey Fedoseev, Dzmitry Tsetserukou
RO-MAN5
2021 WareVR: Virtual Reality Interface for Supervision of Autonomous Robotic System Aimed at Warehouse Stocktaking
abstract
WareVR is a novel human-robot interface based on a virtual reality (VR) application to interact with a heterogeneous robotic system for automated inventory management. We have created an interface to supervise an autonomous robot remotely from a secluded workstation in a warehouse that could benefit during the current pandemic COVID-19 since the stocktaking is a necessary and regular process in warehouses, which involves a group of people. The proposed interface allows regular warehouse workers without experience in robotics to control the heterogeneous robotic system consisting of an unmanned ground vehicle (UGV) and unmanned aerial vehicle (UAV). WareVR provides visualization of the robotic system in a digital twin of the warehouse, which is accompanied by a real-time video stream from the real environment through an onboard UAV camera. Using the WareVR interface, the operator can conduct different levels of stocktaking, monitor the inventory process remotely, and teleoperate the drone for a more detailed inspection. Besides, the developed interface includes remote control of the UAV for intuitive and straightforward human interaction with the autonomous robot for stocktaking. The effectiveness of the VR-based interface was evaluated through the user study in a “visual inspection” scenario.
Ivan Kalinov, Daria Trinitatova, Dzmitry Tsetserukou
SMC3
2021 CobotAR: Interaction with Robots using Omnidirectionally Projected Image and DNN-based Gesture Recognition
abstract
Several technological solutions supported the creation of interfaces for Augmented Reality (AR) multi-user collaboration in the last years. However, these technologies require the use of wearable devices. We present CobotAR -a new AR technology to achieve the Human-Robot Interaction (HRI) by gesture recognition based on Deep Neural Network (DNN) - without an extra wearable device for the user. The system allows users to have a more intuitive experience with robotic applications using just their hands. The CobotAR system assumes the AR spatial display created by a mobile projector mounted on a 6 DoF robot. The proposed technology suggests a novel way of interaction with machines to achieve safe, intuitive, and immersive control mediated by a robotic projection system and DNN-based algorithm. We conducted the experiment with several parameters assessment during this research, which allows the users to define the positives and negatives of the new approach. The mental demand of CobotAR system is twice less than Wireless Gamepad and by 16% less than Teach Pendant.
Elena Nazarova, Oleg Sautenkov, Miguel Altamirano, Jonathan Tirado, Valerii Serpiva, Viktor Rakhmatulin, Dzmitry Tsetserukou
SMC7
2021 CoboGuider: Haptic Potential Fields for Safe Human-Robot Interaction
abstract
Modern industry still relies on manual manufacturing operations and safe human-robot interaction is of great interest nowadays. Speed and Separation Monitoring (SSM) allows close and efficient collaborative scenarios by maintaining a protective separation distance during robot operation. The paper focuses on a novel approach to strengthen the SSM safety requirements by introducing haptic feedback to a robotic cell worker. Tactile stimuli provide early warning of dangerous movements and proximity to the robot, based on the human reaction time and instantaneous velocities of robot and op-erator. A preliminary experiment was performed to identify the reaction time of participants when they are exposed to tactile stimuli in a collaborative environment with controlled conditions. In a second experiment, we evaluated our approach into a study case where human worker and cobot performed collaborative planetary gear assembly. Results show that the applied approach increased the average minimum distance between the robot’s end-effector and hand by 44% compared to the operator relying only on the visual feedback. Moreover, the participants without the haptic support have failed several times to maintain the protective separation distance.
Viktor Rakhmatulin, Miguel Altamirano, Fikre Hagos, Oleg Sautenkov, Jonathan Tirado, Ighor Uzhinsky, Dzmitry Tsetserukou
SMC7
2020 Realizing Body-Machine Interface for Quadrotor Control Through Kalman Filters and Recurrent Neural Network
abstract
Unmanned Aerial Vehicles (UAV) have been recently applied in several various civilian applications. Based on this, there is a growing need for intuitive UAV control interfaces. In this work, we report on the Body-Machine Interface (BMI), helping a human operator to control a quadrotor through the gesture commands. We perform the human motion capture through wearable sensors and Kalman filter to reduce the noise. For the gesture command recognition, we designed the Recurrent Neural Network recognizing gestures within 65 ms. For the quadrotor orientation estimation, we designed the Extended Kalman Filter (EKF). We assess the proposed BMI via the simulations and experiments: the standard deviation of the trajectories varies for up to 10 cm.
Alexander Menshchikov, Daniil Lopatkin, Evgeny V. Tsykunov, Dzmitry Tsetserukou, Andrey Somov
ETFA4
2020 Customer behavior analytics using an autonomous robotics-based system
abstract
This paper suggests a novel method for customer behavior analytics and demand distribution based on Radio Frequency Identification (RFID) stocktaking. Existing solutions lack applicability to real-life situations in retailing, which may result in unobservable loss of sales. The proposed solution provides new parameters of demand distribution to the retailer using a mobile robot for autonomous stocktaking of RFID-equipped shopping rooms. Built models depict location-related demand dependencies, the most and the least purchasable areas in a store, and precise localization of lost and moved items. Our research differs from the related works by the sheer size of the underlying data set collected in a real-world environment for more than ten months.
Alexander A. Petrovsky, Ivan Kalinov, Pavel A. Karpyshev, Mikhail Kurenkov, Vladimir Ramzhaev, Valeriy Ilin, Dzmitry Tsetserukou
ICARCV7
2020 Optimal Oscillation Damping Control of cable-Suspended Aerial Manipulator with a Single IMU Sensor
abstract
This paper presents a design of oscillation damping control for the cable-Suspended Aerial Manipulator (SAM). The SAM is modeled as a double pendulum, and it can generate a body wrench as a control action. The main challenge is the fact that there is only one onboard IMU sensor which does not provide full information on the system state. To overcome this difficulty, we design a controller motivated by a simplified SAM model. The proposed controller is very simple yet robust to model uncertainties. Moreover, we propose a gain tuning rule by formulating the proposed controller in the form of output feedback linear quadratic regulation problem. Consequently, it is possible to quickly dampen oscillations with minimal energy consumption. The proposed approach is validated through simulations and experiments.
Yuri S. Sarkisov, Minjun Kim 0003, Andre Coelho, Dzmitry Tsetserukou, Christian Ott 0001, Konstantin Kondak
ICRA4
2020 TT-TSDF: Memory-Efficient TSDF with Low-Rank Tensor Train Decomposition
abstract
In this paper we apply the low-rank Tensor Train decomposition for compression and operations on 3D objects and scenes represented by volumetric distance functions. Our study shows that not only it allows for a very efficient compression of the high-resolution TSDF maps (up to three orders of magnitude of the original memory footprint at resolution of 5123), but also allows to perform TSDF-Fusion directly in the low-rank form. This can potentially enable much more efficient 3D mapping on low-power mobile and consumer robot platforms.
Alexey I. Boyko, Mikhail Matrosov, Ivan V. Oseledets, Dzmitry Tsetserukou, Gonzalo Ferrer 0001
IROS4
2020 DroneLight: Drone Draws in the Air using Long Exposure Light Painting and ML
abstract
We propose a novel human-drone interaction paradigm where a user directly interacts with a drone to light-paint predefined patterns or letters through hand gestures. The user wears a glove which is equipped with an IMU sensor to draw letters or patterns in the midair. The developed ML algorithm detects the drawn pattern and the drone light-paints each pattern in midair in the real time. The proposed classification model correctly predicts all of the input gestures. The DroneLight system can be applied in drone shows, advertisements, distant communication through text or pattern, rescue, and etc. To our knowledge, it would be the world's first human-centric robotic system that people can use to send messages based on light-painting over distant locations (drone-based instant messaging). Another unique application of the system would be the development of vision-driven rescue system that reads light-painting by person who is in distress and triggers rescue alarm.
Roman Ibrahimov, Nikolay Zherdev, Dzmitry Tsetserukou
RO-MAN3
2019 IceVisionSet: lossless video dataset collected on Russian winter roads with traffic sign annotations
abstract
Ability of autonomous vehicles to operate in complex dynamic environments requires, among other things, fast and accurate perception of surroundings, which includes recognition and tracking of traffic signs.For development and testing of modern sophisticated computer vision systems large and diverse datasets are of the major importance. To test the robustness of algorithms, image data with different moving speeds, camera settings, lighting and weather conditions are especially important.In this work we present a comprehensive, lifelike dataset of traffic sign images collected on the Russian winter roads in varying conditions, which include different weather, camera exposure, illumination and moving speeds. The dataset was annotated in accordance with the Russian traffic code. Annotation results and images are published under open CC BY 4.0 license and can be downloaded from the project website: http://oscar.skoltech.ru/.
Artem L. Pavlov, Pavel A. Karpyshev, G. V. Ovchinnikov, Ivan V. Oseledets, Dzmitry Tsetserukou
ICRA5
2019 Development of SAM: cable-Suspended Aerial Manipulator*
abstract
High risk of a collision between rotor blades and the obstacles in a complex environment imposes restrictions on the aerial manipulators. To solve this issue, a novel system cable-Suspended Aerial Manipulator (SAM) is presented in this paper. Instead of attaching a robotic manipulator directly to an aerial carrier, it is mounted on an active platform which is suspended on the carrier by means of a cable. As a result, higher safety can be achieved because the aerial carrier can keep a distance from the obstacles. For self-stabilization, the SAM is equipped with two actuation systems: winches and propulsion units. This paper presents an overview of the SAM including the concept behind, hardware realization, control strategy, and the first experimental results.
Yuri S. Sarkisov, Minjun Kim 0003, Davide Bicego, Dzmitry Tsetserukou, Christian Ott 0001, Antonio Franchi, Konstantin Kondak
ICRA4
2019 Asynchronous Behavior Trees with Memory aimed at Aerial Vehicles with Redundancy in Flight Controller
abstract
Complex aircraft systems are becoming a target for automation. For successful operation, they require both efficient and readable mission execution system (MES). Flight control computer (FCC) units, as well as all important subsystems, are often duplicated. Discrete nature of MES does not allow small differences in data flow among redundant FCCs which are acceptable for continuous control algorithms. Therefore, mission state consistency has to be specifically maintained. We present a novel MES which includes FCC state synchronization. To achieve this result we developed the new concept of Asynchronous Behavior Tree with Memory (ABTM) and proposed a state synchronization algorithm. The implemented system was tested and proven to work in a real-time simulation of High Altitude Pseudo Satellite (HAPS) mission.
Evgenii Safronov, Michael Vilzmann, Dzmitry Tsetserukou, Konstantin Kondak
IROS3
2019 DronePick: Object Picking and Delivery Teleoperation with the Drone Controlled by a Wearable Tactile Display
abstract
We report on the teleoperation system DronePick which provides remote object picking and delivery by a human- controlled quadcopter. The main novelty of the proposed system is that the human user continuously gets the visual and haptic feedback for accurate teleoperation. DronePick consists of a quadcopter equipped with a magnetic grabber, a tactile glove with finger motion tracking sensor, hand tracking system, and the Virtual Reality (VR) application. The human operator teleoperates the quadcopter by changing the position of the hand. The proposed vibrotactile patterns representing the location of the remote object relative to the quadcopter are delivered to the glove. It helps the operator to determine when the quadcopter is right above the object. When the “pick” command is sent by clasping the hand in the glove, the quadcopter decreases its altitude and the magnetic grabber attaches the target object. The whole scenario is in parallel simulated in VR. The air flow from the quadcopter and the relative positions of VR objects help the operator to determine the exact position of the delivered object to be picked. The experiments showed that the vibrotactile patterns were recognized by the users at the high recognition rates: the average 99% recognition rate and the average 2.36s recognition time. The real-life implementation of DronePick featuring object picking and delivering to the human was developed and tested.
Roman Ibrahimov, Evgeny V. Tsykunov, Vladimir Shirokun, Andrey Somov, Dzmitry Tsetserukou
RO-MAN5
2019 Development of MirrorShape: High Fidelity Large-Scale Shape Rendering Framework for Virtual Reality
abstract
Today there is a high variety of haptic devices capable of providing tactile feedback. Although most of existing designs are aimed at realistic simulation of the surface properties, their capabilities are limited in attempts of displaying shape and position of virtual objects.
Aleksey Fedoseev, Nikita Chernyadev, Dzmitry Tsetserukou
VRST3
2019 RecyGlide : A Forearm-worn Multi-modal Haptic Display aimed to Improve User VR Immersion Submission
abstract
Haptic devices have been employed to immerse users in VR environments. In particular, hand and finger haptic devices have been deeply developed. However, this type of devices occludes hand detection for some tracking systems, or, for some other tracking systems, it is uncomfortable for the users to wear two different devices (haptic and tracking device) on both hands. We introduce RecyGlide, a novel wearable multimodal display located at the forearm. The RecyGlide is composed of inverted five-bar linkages with 2 degrees of freedom (DoF) and vibration motors (see Fig. 1.(a). The device provides multimodal tactile feedback such as slippage, force vector, pressure, and vibration. We tested the discrimination ability of monomodal and multimodal stimuli patterns on the forearm and confirmed that the multimodal patterns have higher recognition rate. This haptic device was used in VR applications, and we proved that it enhances VR experience and makes it more interactive.
Juan Heredia 0001, Jonathan Tirado, Vladislav Panov, Miguel Altamirano, Kamal Youcef-Toumi, Dzmitry Tsetserukou
VRST6
2019 SlingDrone: Mixed Reality System for Pointing and Interaction Using a Single Drone
abstract
We propose SlingDrone, a novel Mixed Reality interaction paradigm that utilizes a micro-quadrotor as both pointing controller and interactive robot with a slingshot motion type. The drone attempts to hover at a given position while the human pulls it in desired direction using a hand grip and a leash. Based on the displacement, a virtual trajectory is defined. To allow for intuitive and simple control, we use virtual reality (VR) technology to trace the path of the drone based on the displacement input. The user receives force feedback propagated through the leash. Force feedback from SlingDrone coupled with visualized trajectory in VR creates an intuitive and user friendly pointing device. When the drone is released, it follows the trajectory that was shown in VR. Onboard payload (e.g. magnetic gripper) can perform various scenarios for real interaction with the surroundings, e.g. manipulation or sensing. Unlike HTC Vive controller, SlingDrone does not require handheld devices, thus it can be used as a standalone pointing technology in VR.
Evgeny V. Tsykunov, Roman Ibrahimov, Derek Vasquez, Dzmitry Tsetserukou
VRST4
2019 WiredSwarm: High Resolution Haptic Feedback Provided by a Swarm of Drones to the User's Fingers for VR interaction
abstract
We propose a concept of a novel interaction strategy for providing rich haptic feedback in Virtual Reality (VR), when each user’s finger is connected to micro-quadrotor with a wire. Described technology represents the first flying wearable haptic interface. The solution potentially is able to deliver high resolution force feedback to each finger during fine motor interaction in VR. The tips of tethers are connected to the centers of quadcopters under their bottom. Therefore, flight stability is increasing and the interaction forces are becoming stronger which allows to use smaller drones.
Evgeny V. Tsykunov, Dzmitry Tsetserukou
VRST2
2019 High-Precision UAV Localization System for Landing on a Mobile Collaborative Robot Based on an IR Marker Pattern Recognition
abstract
We present a novel high-precision UAV localization system for interconnection between two collaborative robots, i.e., unmanned ground robot (UGR) and unmanned aerial vehicle (UAV) capable of autonomous navigation and precise localization in an indoor environment. Based on our localization system we have achieved robust UAV landing on the moving robot using a fusion of 2D LIDAR sensors, camera, and ultrasonic system for localization. In addition, UAV is capable of accurate high-altitude indoor flights (up to 15 m) relative to the ground robot. Localization of UAV is based on the developed adaptive active IR marker system to achieve reliable flight on different altitudes and light conditions. In this paper, we describe the operating principle of the system and present the results of UAV flight experiments. One of promising applications of the developed system is automated inventory management of warehouses.
Ivan Kalinov, Evgenii Safronov, Ruslan Agishev, Mikhail Kurenkov, Dzmitry Tsetserukou
VTC Spring5
2018 AA-ICP: Iterative Closest Point with Anderson Acceleration
abstract
Iterative Closest Point (ICP) is a widely used method for performing scan-matching and registration. Being simple and robust, this method is still computationally expensive and may be challenging to use in real-time applications with limited resources on mobile platforms. In this paper we propose a novel effective method for acceleration of ICP which does not require substantial modifications to the existing code. This method is based on an idea of Anderson acceleration which is an iterative procedure for finding a fixed point of contractive mapping. The latter is often faster than a standard Picard iteration, usually used in ICP implementations. We show that ICP, being a fixed point problem, can be significantly accelerated by this method enhanced by heuristics to improve overall robustness. We implement proposed approach into Point Cloud Library (PCL) and make it available online. Benchmarking on the real-world data fully supports our claims.
Artem L. Pavlov, G. V. Ovchinnikov, Dmitry Yu. Derbyshev, Dzmitry Tsetserukou, Ivan V. Oseledets
ICRA4
2018 SwarmTouch: Tactile Interaction of Human with Impedance Controlled Swarm of Nano-Quadrotors
abstract
We propose a novel interaction strategy for a human-swarm communication when a human operator guides a formation of quadrotors with impedance control and receives vibrotactile feedback. The presented approach takes into account the human hand velocity and changes the formation shape and dynamics accordingly using impedance interlinks simulated between quadrotors, which helps to achieve a life-like swarm behavior. Experimental results with Crazyflie 2.0 quadrotor platform validate the proposed control algorithm. The tactile patterns representing dynamics of the swarm (extension or contraction) are proposed. The user feels the state of the swarm at his fingertips and receives valuable information to improve the controllability of the complex life-like formation. The user study revealed the patterns with high recognition rates. Subjects stated that tactile sensation improves the ability to guide the drone formation and makes the human-swarm communication much more interactive. The proposed technology can potentially have a strong impact on the human-swarm interaction, providing a new level of intuitiveness and immersion into the swarm navigation.
Evgeny V. Tsykunov, Luiza Labazanova, Akerke Tleugazy, Dzmitry Tsetserukou
IROS4
2011 Belt tactile interface for communication with mobile robot allowing intelligent obstacle detection
abstract
This paper focuses on the construction of a novel belt tactile interface and telepresence system intended for mobile robot control. The robotic system consists of a mobile robot and a wearable master robot. The elaborated algorithms allow the robot to precisely recognize the shape, boundaries, movement direction, speed, and distance to the obstacle by means of the laser range finders. The designed tactile belt interface receives the detected information and maps it through the vibrotactile patterns. We designed the patterns in such a way that they convey the obstacle parameters in a very intuitive, robust, and unobtrusive manner. The robot movement direction and speed are governed by the tilt of the user's torso. The sensors embedded into the belt interface measure the user orientation and gestures precisely. Such an interface lets to deeply engage the user into the teleoperation process and to deliver them the tactile perception of the remote environment at the same time. The key point is that the user gets the opportunity to use own arms, hands, fingers for operation of the robotic manipulators and another devices installed on the mobile robot platform. The experimental results of user study revealed the effectiveness of the designed vibration patterns for obstacle parameter presentation. The accuracy in 100% for detection of the moving object by participants was achieved. We believe that the developed robotic system has significant potential in facilitating the navigation of mobile robot while providing a high degree of immersion into remote space.
Dzmitry Tsetserukou, Junichi Sugiyama, Jun Miura
World Haptics1
2009 FlexTorque: innovative haptic interface for realistic physical interaction in virtual reality
abstract
In order to realize haptic interaction (e.g., holding, pushing, and contacting the object) in virtual environment and mediated haptic communication with human beings (e.g., handshaking), the force feedback is required. Recently there has been a substantial need and interest in haptic displays, which can provide realistic and high fidelity physical interaction in virtual environment. The aim of our research is to implement a wearable haptic display for presentation of realistic feedback (kinesthetic stimulus) to the human arm. We developed a wearable device FlexTorque that induces forces to the human arm and does not require holding any additional haptic interfaces in the human hand. It is completely new technology for Virtual Reality that allows user to explore surroundings freely. The concept of Karate (empty hand) Haptics proposed by us is opposite to conventional interfaces (e.g., Wii Remote, SensAble's PHANTOM, SPIDAR [Murayama et al. 2004]) that require holding haptic interface in the hand, restricting thus the motion of the fingers in midair.
Dzmitry Tsetserukou, Katsunari Sato, Alena Neviarouskaya, Naoki Kawakami, Susumu Tachi
SIGGRAPH ASIA Sketches1
2007 Towards Safe Human-Robot Interaction: Joint Impedance Control of a New Teleoperated Robot Arm
abstract
The paper focuses on design and joint impedance control of a new teleoperated robot arm enabling torque measurement in each joint by means of incorporation of devised optical torque sensors. When the contact of arm with an object occurs, joint impedance algorithm provides active compliance of corresponding robot arm joint. Thus, the whole structure of the manipulator can safely interact with unstructured environment. In the paper, we describe detailed design procedure of the 4-DOF robot arm and optical torque sensors. To extract force signal from measured data, the gravity compensation algorithm was elaborated and verified. The experimental results of joint impedance control show that proposed strategy provides safe interaction of entire structure of robot arm with human beings and ensures the collision avoidance.
Dzmitry Tsetserukou, Riichiro Tadakuma, Hiroyuki Kajimoto, Naoki Kawakami, Susumu Tachi
RO-MAN1
2006 Pervasive Sensor System for Evidence-based Nursing Care Support
abstract
This paper introduces a pervasive sensor system for nursing homes, where daily activities of elderly persons are monitored by pervasive sensors all the time. Deterioration in the quality of nursing care for the elderly has become one of the biggest problems in the aging society and the authors have been challenging the problem by sensors embedded in a nursing room. The sensors accumulate position information of a subject person and his wheelchair, then it is utilized for prompt assists for the subject and also used for obtaining his life log. In our experiments, we obtained position data of an elderly person for a month and a half and analyzed his sleeping hours and the number of times of going to a restroom. This paper presents the concept of the system, overview of the current system and experimental results obtained
Dzmitry Tsetserukou, Riichiro Tadakuma, Hiroyuki Kajimoto, Susumu Tachi
ICRA1