EDBT 2026 Demo / reviewers in the wild / expert
Miguel Altamirano
dblp:213/7883 · also Miguel Altamirano Cabrera
· DBLP profile ↗
17ranked-venue papers
1as first author
16since 2021 · last 2025
0000-0002-5974-9257ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ViewVR: Visual Feedback Modes to Achieve Quality of VR-Based TelemanipulationabstractThe paper focuses on an immersive teleoperation system that enhances operator's ability to actively perceive the robot's surroundings. A consumer-grade HTC Vive VR system was used to synchronize the operator's hand and head movements with a UR3 robot and a custom-built robotic head with two degrees of freedom (2-DoF). The system's usability, manipulation efficiency, and intuitiveness of control were evaluated in comparison with static head camera positioning across three distinct tasks. Code and other supplementary materials can be accessed by link: https://github.com/ErkhovArtemNiewVR. Artem Erkhov, Artem Bazhenov, Sergei Satsevich, Danil Belov, Farit Khabibullin, Sergei Egorov, Maxim Gromakov, Miguel Altamirano, Dzmitry Tsetserukou |
HRI | 8 |
| 2025 | Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid MixingabstractThis paper introduces Shake-VLA, a Vision-Language-Action (VLA) model-based system designed to enable bimanual robotic manipulation for automated cocktail preparation. The system integrates a vision module for detecting ingredient bottles and reading labels, a speech-to-text module for interpreting user commands, and a language model to generate task-specific robotic instructions. Force Torque (FT) sensors are employed to precisely measure the quantity of liquid poured, ensuring accuracy in ingredient proportions during the mixing process. The system architecture includes a Retrieval-Augmented Generation (RAG) module for accessing and adapting recipes, an anomaly detection mechanism to address ingredient availability issues, and bimanual robotic arms for dexterous manipulation. Experimental evaluations demonstrated a high success rate across system components, with the speech-to-text module achieving a 93% success rate in noisy environments, the vision module attaining a 91% success rate in object and label detection in cluttered environment, the anomaly module successfully identified 95% of discrepancies between detected ingredients and recipe requirements, and the system achieved an overall success rate of 100% in preparing cocktails, from recipe formulation to action generation. Muhamamd Haris Khan, Selamawit Asfaw, Dmitrii Iarchuk, Miguel Altamirano, Luis Moreno 0007, Issatay Tokmurziyev, Dzmitry Tsetserukou |
HRI | 4 |
| 2025 | UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission GenerationabstractThe UAV-VLA (Visual-Language-Action) system is a tool designed to facilitate communication with aerial robots. By integrating satellite imagery processing with the Visual Language Model (VLM) and the powerful capabilities of GPT, UAV-VLA enables users to generate general flight paths-and-action plans through simple text requests. This system leverages the rich contextual information provided by satellite images, allowing for enhanced decision-making and mission planning. The combination of visual analysis by VLM and natural language processing by GPT can provide the user with the path-and-action set, making aerial operations more efficient and accessible. The newly developed method showed the difference in the length of the created trajectory in 22% and the mean error in finding the objects of interest on a map in 34.22 m by Euclidean distance in the K-Nearest Neighbors (KNN) approach. Additionally, the UAV-VLA system generates all flight plans in just 5 minutes and 24 seconds, making it 6.5 times faster than an experienced human operator. The code is available here: https://github.com/sautenich/uav-vla Oleg Sautenkov, Yasheerah Yaqoot, Artem Lykov, Muhammad Ahsan Mustafa, Grik Tadevosyan, Aibek Akhmetkazy, Miguel Altamirano, Mikhail Martynov, Sausar Karaf, Dzmitry Tsetserukou |
HRI | 7 |
| 2025 | GazeGrasp: DNN-Driven Robotic Grasping with Wearable Eye-Gaze InterfaceabstractWe present GazeGrasp, a gaze-based manipulation system enabling individuals with motor impairments to control collaborative robots using eye-gaze. The system employs an ESP32 CAM for eye tracking, MediaPipe for gaze detection, and YOLOv8 for object localization, integrated with a Uni-versal Robot UR10 for manipulation tasks. After user-specific calibration, the system allows intuitive object selection with a magnetic snapping effect and robot control via eye gestures. Experimental evaluation involving 13 participants demonstrated that the magnetic snapping effect significantly reduced gaze alignment time, improving task efficiency by 31%. GazeGrasp provides a robust, hands-free interface for assistive robotics, enhancing accessibility and autonomy for users. Issatay Tokmurziyev, Miguel Altamirano, Luis Moreno 0007, Muhammad Haris Khan, Dzmitry Tsetserukou |
HRI | 2 |
| 2025 | CognitiveOS: Large Multimodal Model Based System to Endow Any Type of Robot with Generative AIabstractThis paper introduces CognitiveOS, the first operating system designed for cognitive robots capable of functioning across diverse robotic platforms. CognitiveOS is structured as a multi-agent system comprising modules built upon a transformer architecture, facilitating communication through an internal monologue format. These modules collectively empower the robot to tackle intricate real-world tasks. The paper delineates the operational principles of the system along with the descriptions of its nine distinct modules. The modular design endows the system with distinctive advantages over traditional end-to-end methodologies, notably in terms of adaptability and scalability. The system's modules are configurable, modifiable, or deactivatable depending on the task requirements, while new modules can be seamlessly integrated. This system serves as a foundational resource for researchers and developers in the Cognitive Robotics domain, alleviating the burden of constructing a cognitive robot system from scratch. Experimental findings demonstrate the system's advanced task comprehension and adaptability across varied tasks, robotic platforms, and module configurations, underscoring its potential for realworld applications. Moreover, in the category of Reasoning it outperformed CognitiveDog (by 15%) and RT2 (by 31%), achieving the highest to date rate of 77 %. We provide a code repository and dataset for the replication of CognitiveOS: https://github.com/Arcwy0/cognitiveos Artem Lykov, Mikhail Konenkov, Koffivi Fidèle Gbagbe, Mikhail Litvinov, Denis Davletshin, Aleksey Fedoseev, Miguel Altamirano, Robinroy Peter, Dzmitry Tsetserukou |
ICRA | 7 |
| 2025 | Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous RobotsabstractThis paper presents the concept of Industry 6.0, which introduces the world’s first fully automated production system that autonomously handles the entire product design and manufacturing process based on user-provided natural language descriptions. By leveraging generative AI, the system automates critical aspects of production, including product blueprint design, component manufacturing, logistics, and assembly. A heterogeneous swarm of robots, each equipped with individual AI through integration with Large Language Models (LLMs), orchestrates the production process. The robotic system includes manipulator arms, delivery drones, and 3D printers capable of generating assembly blueprints. The system was evaluated using commercial and open source LLMs, operating via APIs and local deployment. A user study demonstrated that the system reduced the average production time to 119.10 minutes, significantly outperforming a team of expert human developers, who averaged 528.64 minutes (an improvement factor of 4.4). Furthermore, in the product blueprinting stage, the system outperformed human CAD operators by an unprecedented factor of 47, completing the task in 0.5 minutes compared to 23.5 minutes. This breakthrough represents a major leap towards fully autonomous manufacturing. Artem Lykov, Miguel Altamirano, Mikhail Konenkov, Valerii Serpiva, Koffivi Fidèle Gbagbe, Ali Alabbas, Aleksey Fedoseev, Luis Moreno 0007, Muhammad Haris Khan, Ziang Guo, Dzmitry Tsetserukou |
IROS | 2 |
| 2025 | HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic InteractionabstractThis paper introduces HapticVLM, a novel multimodal system that integrates vision-language reasoning and deep convolutional networks to enable real-time haptic feedback. HapticVLM leverages a ConvNeXt-based material recognition module to generate robust visual embeddings for accurate identification of object materials. A state-of-the-art Vision-Language Model (Qwen2-VL-2B-Instruct) infers ambient temperature from environmental cues. The system synthesizes tactile sensations by delivering vibrotactile feedback through speakers and thermal cues with a Peltier module, thereby bridging the gap between visual perception and tactile experience. Experimental evaluations demonstrate an average recognition accuracy of 84.7% across five distinct auditory-tactile patterns and a temperature estimation accuracy of 86.7% using an 8 °C margin error across 15 scenarios. Although promising, the current study is limited by the use of a small set of patterns and participants. Future work will focus on expanding the range of tactile patterns and increasing user studies to further refine and validate the system’s performance. Overall, HapticVLM presents a significant step toward intelligent, context-aware, multimodal haptic interaction for Virtual Reality (VR) and assistive technologies. Muhammad Haris Khan, Miguel Altamirano, Dmitrii Iarchuk, Yara Mahmoud, Daria Trinitatova, Issatay Tokmurziyev, Dzmitry Tsetserukou |
SMC | 2 |
| 2024 | Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous ManipulationsabstractThis research introduces the Bi-VLA (Vision-Language-Action) model, a novel system designed for bimanual robotic dexterous manipulation that seamlessly integrates vision for scene understanding, language comprehension for translating human instructions into executable code, and physical action generation. We evaluated the system's functionality through a series of household tasks, including the preparation of a desired salad upon human request. Bi-VLA demonstrates the ability to interpret complex human instructions, perceive and understand the visual context of ingredients, and execute precise bimanual actions to prepare the requested salad. We assessed the system's performance in terms of accuracy, efficiency, and adaptability to different salad recipes and human preferences through a series of experiments. Our results show a 100 % success rate in generating the correct executable code by the Language Module, a 96.06 % success rate in detecting specific ingredients by the Vision Module, and an overall success rate of 83.4 % in correctly executing user-requested tasks. Koffivi Fidèle Gbagbe, Miguel Altamirano, Ali Alabbas, Oussama Alyounes, Artem Lykov, Dzmitry Tsetserukou |
SMC | 2 |
| 2024 | GazeRace: Revolutionizing Remote Piloting with Eye-Gaze ControlabstractThis paper presents GazeRace, a novel system that leverages eye-tracking technology for intuitive drone control. Using the MediaPipe library, the system translates eye movements into precise drone commands, enabling effective remote piloting. In testing, GazeRace demonstrated an 18% reduction in drone trajectory length while maintaining competitive speed with traditional controls. The results suggest that this approach enhances control accuracy and reduces user frustration, offering a significant advancement in the field of human-computer interaction and drone navigation. Issatay Tokmurziyev, Valerii Serpiva, Aleksey Fedoseev, Miguel Altamirano, Dzmitry Tsetserukou |
SMC | 4 |
| 2024 | Pose estimation in robotic electric vehicle plug-in charging tasks using auto-annotation and deep learning-based keypoint detector
Viktor Rakhmatulin, Miguel Altamirano, Andrei Puchkov, Evgeny Burnaev, Dzmitry Tsetserukou |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | OmniCharger: CNN-Based Hand Gesture Interface to Operate an Electric Car Charging Robot through TeleconferenceabstractThe automation of the car charging process is motivated by the rapid development of technologies for self-driving cars and the increasing importance of ecological transportation units. Automation of this process requires the implementation of Computer Vision (CV) techniques. However, it remains challenging to precisely position the charger plug autonomously due to the sensitivity of CV algorithms to lighting and weather conditions. We introduce a novel robotic operation system based on hand gesture recognition through teleconferencing software. The users, connected by teleconference, use their hand gestures to teleoperate the electric plug located on the collaborative robot end-effector. We conducted a user study to evaluate the system performance and suitability using OmniCharger and two baseline interfaces (a UR10 Teach Pendant and a Logitech F710 Wireless Gamepad). Except for two trials, all the users were able to locate the plug inside of a 5 cm target using the interfaces. The distance to the target and the orientation error did not present statistically significant differences ( \(p=0.1099 \gt 0.05\) and \(p=0.0903 \gt 0.05\) , respectively) in the use of the three interfaces. The NASA-TLX questionnaire results showed low values in all the sub-classes, the SUS results rated the usability of the proposed interface above average (68%), and the UEQ showed excellent performance of the OmniCharger interface in the attractiveness, stimulation, and novelty attributes. Miguel Altamirano, Viktor Rakhmatulin, Aleksey Fedoseev, Oleg Sautenkov, Oussama Alyounes, Andrei Puchkov, Dzmitry Tsetserukou |
ACM Trans. Hum. Robot Interact. | 1 |
| 2023 | ArUcoGlide: a Novel Wearable Robot for Position Tracking and Haptic Feedback to Increase Safety During Human-Robot InteractionabstractThe current capabilities of robotic systems make human collaboration necessary to accomplish complex tasks effectively. In this work, we are introducing a framework to ensure safety in a human-robot collaborative environment. The system is composed of a wearable 2-DoFs robot, a low-cost and easy-to-install tracking system, and a collision avoidance algorithm based on the Artificial Potential Field (APF). The wearable robot is designed to hold a fiducial marker and maintain its visibility to the tracking system, which, in turn, localizes the user’s hand with good accuracy and low latency and provides haptic feedback to the user. The system is designed to enhance the performance of collaborative tasks while ensuring user safety. Three experiments were carried out to evaluate the performance of the proposed system. The first one evaluated the accuracy of the tracking system. The second experiment analyzed human-robot behavior during an imminent collision. The third experiment evaluated the system in a collaborative activity in a shared working environment. The results show that the implementation of the introduced system reduces the operation time by 16% and increases the average distance between the user’s hand and the robot by 5 cm. Ali Alabbas, Miguel Altamirano, Oussama Alyounes, Dzmitry Tsetserukou |
ETFA | 2 |
| 2022 | DroneARchery: Human-Drone Interaction through Augmented Reality with Haptic Feedback and Multi-UAV Collision Avoidance Driven by Deep Reinforcement LearningabstractWe propose a novel concept of augmented reality (AR) human-drone interaction driven by RL-based swarm behavior to achieve intuitive and immersive control of a swarm formation of unmanned aerial vehicles. The DroneARchery system developed by us allows the user to quickly deploy a swarm of drones, generating flight paths simulating archery. The haptic interface LinkGlide delivers a tactile stimulus of the bowstring tension to the forearm to increase the precision of aiming. The swarm of released drones dynamically avoids collisions between each other, the drone following the user, and external obstacles with behavior control based on deep reinforcement learning. The developed concept was tested in the scenario with a human, where the user shoots from a virtual bow with a real drone to hit the target. The human operator observes the ballistic trajectory of the drone in an AR and achieves a realistic and highly recognizable experience of the bowstring tension through the haptic display. The experimental results revealed that the system improves trajectory prediction accuracy by 63.3% through applying AR technology and conveying haptic feedback of pulling force. DroneARchery users highlighted the naturalness (4.3 out of 5 point Likert scale) and increased confidence (4.7 out of 5) when controlling the drone. We have designed the tactile patterns to present four sliding distances (tension) and three applied force levels (stiffness) of the haptic display. Users demonstrated the ability to distinguish tactile patterns produced by the haptic display representing varying bowstring tension(average recognition rate is of 72.8%) and stiffness (average recognition rate is of 94.2%). The novelty of the research is the development of an AR-based approach for drone control that does not require special skills and training from the operator. In the future, the proposed interaction can be applied in various fields, for example, for fast swarm deployment in search and rescue missions, crop monitoring, inspection and maintenance. Ekaterina Dorzhieva, Ahmed Baza, Ayush Gupta 0003, Aleksey Fedoseev, Miguel Altamirano, Ekaterina Karmanova, Dzmitry Tsetserukou |
ISMAR | 5 |
| 2022 | DogTouch: CNN-based Recognition of Surface Textures by Quadruped Robot with High Density Tactile SensorsabstractThe ability to perform locomotion in various terrains is critical for legged robots. However, the robot has to have a better understanding of the surface it is walking on to perform robust locomotion on different terrains. Animals and humans are able to recognize the surface with the help of the tactile sensation on their feet. Although, the foot tactile sensation for legged robots has not been much explored. This paper presents research on a novel quadruped robot DogTouch with tactile sensing feet (TSF). TSF allows the recognition of different surface textures utilizing a tactile sensor and a convolutional neural network (CNN). The experimental results show a sufficient validation accuracy of 74.37% for our trained CNN-based model, with the highest recognition for line patterns of 90%. In the future, we plan to improve the prediction model by presenting surface samples with the various depths of patterns and applying advanced Deep Learning and Shallow learning models for surface recognition.Additionally, we propose a novel approach to navigation of quadruped and legged robots. We can arrange the tactile paving textured surface (similar that used for blind or visually impaired people). Thus, DogTouch will be capable of locomotion in unknown environment by just recognizing the specific tactile patterns which will indicate the straight path, left or right turn, pedestrian crossing, road, and etc. That will allow robust navigation regardless of lighting condition. Future quadruped robots equipped with visual and tactile perception system will be able to safely and intelligently navigate and interact in the unstructured indoor and outdoor environment. Nipun Dhananjaya Weerakkodi Mudalige, Elena Nazarova, Ildar Babataev, Pavel Kopanev, Aleksey Fedoseev, Miguel Altamirano, Dzmitry Tsetserukou |
VTC Spring | 6 |
| 2021 | CobotAR: Interaction with Robots using Omnidirectionally Projected Image and DNN-based Gesture RecognitionabstractSeveral technological solutions supported the creation of interfaces for Augmented Reality (AR) multi-user collaboration in the last years. However, these technologies require the use of wearable devices. We present CobotAR -a new AR technology to achieve the Human-Robot Interaction (HRI) by gesture recognition based on Deep Neural Network (DNN) - without an extra wearable device for the user. The system allows users to have a more intuitive experience with robotic applications using just their hands. The CobotAR system assumes the AR spatial display created by a mobile projector mounted on a 6 DoF robot. The proposed technology suggests a novel way of interaction with machines to achieve safe, intuitive, and immersive control mediated by a robotic projection system and DNN-based algorithm. We conducted the experiment with several parameters assessment during this research, which allows the users to define the positives and negatives of the new approach. The mental demand of CobotAR system is twice less than Wireless Gamepad and by 16% less than Teach Pendant. Elena Nazarova, Oleg Sautenkov, Miguel Altamirano, Jonathan Tirado, Valerii Serpiva, Viktor Rakhmatulin, Dzmitry Tsetserukou |
SMC | 3 |
| 2021 | CoboGuider: Haptic Potential Fields for Safe Human-Robot InteractionabstractModern industry still relies on manual manufacturing operations and safe human-robot interaction is of great interest nowadays. Speed and Separation Monitoring (SSM) allows close and efficient collaborative scenarios by maintaining a protective separation distance during robot operation. The paper focuses on a novel approach to strengthen the SSM safety requirements by introducing haptic feedback to a robotic cell worker. Tactile stimuli provide early warning of dangerous movements and proximity to the robot, based on the human reaction time and instantaneous velocities of robot and op-erator. A preliminary experiment was performed to identify the reaction time of participants when they are exposed to tactile stimuli in a collaborative environment with controlled conditions. In a second experiment, we evaluated our approach into a study case where human worker and cobot performed collaborative planetary gear assembly. Results show that the applied approach increased the average minimum distance between the robot’s end-effector and hand by 44% compared to the operator relying only on the visual feedback. Moreover, the participants without the haptic support have failed several times to maintain the protective separation distance. Viktor Rakhmatulin, Miguel Altamirano, Fikre Hagos, Oleg Sautenkov, Jonathan Tirado, Ighor Uzhinsky, Dzmitry Tsetserukou |
SMC | 2 |
| 2019 | RecyGlide : A Forearm-worn Multi-modal Haptic Display aimed to Improve User VR Immersion SubmissionabstractHaptic devices have been employed to immerse users in VR environments. In particular, hand and finger haptic devices have been deeply developed. However, this type of devices occludes hand detection for some tracking systems, or, for some other tracking systems, it is uncomfortable for the users to wear two different devices (haptic and tracking device) on both hands. We introduce RecyGlide, a novel wearable multimodal display located at the forearm. The RecyGlide is composed of inverted five-bar linkages with 2 degrees of freedom (DoF) and vibration motors (see Fig. 1.(a). The device provides multimodal tactile feedback such as slippage, force vector, pressure, and vibration. We tested the discrimination ability of monomodal and multimodal stimuli patterns on the forearm and confirmed that the multimodal patterns have higher recognition rate. This haptic device was used in VR applications, and we proved that it enhances VR experience and makes it more interactive. Juan Heredia 0001, Jonathan Tirado, Vladislav Panov, Miguel Altamirano, Kamal Youcef-Toumi, Dzmitry Tsetserukou |
VRST | 4 |