Lisa Gutzeit

dblp:185/4997 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-7447-0609ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Improving Human-Robot Communication in Noisy Environments with Visual Voice Activity Detection
abstract
This paper investigates the integration of Visual Voice Activity Detection (VVAD) into human-robot dialogue systems to enhance communication in noisy environments. Usually, speech recognition systems often falter under acoustic interference, limiting their effectiveness in real-world human-robot interactions. By leveraging visual cues, especially lip movements, VVAD supports more accurate speech detection and turn-taking. In this paper, we present a multimodal dialogue system combining VVAD with speech recognition, dialogue management, and context-aware intention recognition. The system is evaluated through a user study involving 30 participants in a noise-rich, simulated restaurant scenario. Results show a substantial improvement in task completion rates and user satisfaction when VVAD is enabled, alongside with a significant reduction in false activations. These findings underscore the value of visual input in robust, socially aware human-robot interaction and suggest VVAD as a critical component in the design of next-generation interactive robots.
Arunima Gopikrishnan, Adrian Auer, Lisa Gutzeit
CHIRA (2)3
2025 Bayesian Inverse Physics for Neuro-Symbolic Robot Learning
abstract
Real-world robotic applications, from autonomous exploration to assistive technologies, require adaptive, interpretable, and data-efficient learning paradigms. While deep learning architectures and foundation models have driven significant advances in diverse robotic applications, they remain limited in their ability to operate efficiently and reliably in unknown and dynamic environments. In this position paper, we critically assess these limitations and introduce a conceptual framework for combining data-driven learning with deliberate, structured reasoning. Specifically, we propose leveraging differentiable physics for efficient world modeling, Bayesian inference for uncertainty-aware decision-making, and meta-learning for rapid adaptation to new tasks. By embedding physical symbolic reasoning within neural models, robots could generalize beyond their training data, reason about novel situations, and continuously expand their knowledge. We argue that such hybrid neuro-symbolic architectures are essential for the next generation of autonomous systems, and to this end, we provide a research roadmap to guide and accelerate their development.
Octavio Arriaga, Rebecca Adam, Melvin Laux, Lisa Gutzeit, Marco Ragni, Jan Peters 0001, Frank Kirchner
NeSy4
2024 Fusion of Inertial Sensor Suit and Monocular Camera for 3D Human Pelvis Pose Estimation
abstract
In real-world scenarios, robots come closer to humans in many applications, sharing the same workspace or even manipulating the same objects. To ensure safe and intuitive collaboration, it is crucial to have an accurate knowledge of the human’s 3D position in space, which should be estimated with high precision, high frequency and low latency. However, individual sensors such as inertial measurement units (IMUs) or cameras cannot meet all requirements for reliable human pose estimation under conditions such as long operating times, large distances and occlusions. In this study, we highlight the limitations of different visual pose methods and present a fused approach for real-time estimation of the 3D position of the human pelvis using machine learning-based visual pose from a monocular camera and an IMU sensor suit. The multimodal fusion is based on the Invariant Extended Kalman filter (InEKF) on Lie Groups, which fuses drift-free visual poses with high-frequency inertial measurements in a loosely-coupled manner. The evaluation is performed on a recorded dataset of multiple subjects performing various experimental scenarios. The results show that the fused approach can increase the accuracy and robustness of the estimates, taking a step closer towards smooth human-robot collaboration.
Mihaela Popescu, Kashmira Shinde, Proneet Sharma, Lisa Gutzeit, Frank Kirchner
RO-MAN4
2023 CoBaIR: A Python Library for Context-Based Intention Recognition in Human-Robot-Interaction
abstract
Human-Robot Interaction (HRI) becomes more and more important in a world where robots integrate fast in all aspects of our lives but HRI applications depend massively on the utilized robotic system as well as the deployment environment and cultural differences. Because of these variable dependencies it is often not feasible to use a data-driven approach to train a model for human intent recognition. Expert systems have been proven to close this gap very efficiently. Furthermore, it is important to support understandability in HRI systems to establish trust in the system. To address the above-mentioned challenges in HRI we present an adaptable python library in which current state-of-the-art Models for context recognition can be integrated. For Context-Based Intention Recognition a two-layer Bayesian Network (BN) is used. The bayesian approach offers explainability and clarity in the creation of scenarios and is easily extendable with more modalities. Additionally, it can be used as an expert system if no data is available but can as well be fine-tuned when data becomes available
Adrian Lubitz, Lisa Gutzeit, Frank Kirchner
RO-MAN2
2022 Hierarchical Segmentation of Human Manipulation Movements
abstract
This paper introduces a segmentation algorithm which splits complex human manipulation movements into movement segments and automatically groups these into labeled actions. With this hierarchical algorithm the basic movement identities, which we call building blocks, as well as their concatenation to more complex actions can be identified in one handed as well as dual arm human manipulation movements. The algorithm can be used, e.g., in robotic applications such as imitation learning, in which human movement examples are directly used to generate robotic behavior. In this paper, we present two variants of the hierarchical segmentation algorithm, one supervised approach which requires a small number of pre-labeled movements as training data, as well as an approach which uses unsupervised algorithms to group building block segments which belong to the same movement. In both variants, the building block movements are detected based on the velocity of the hand(s), using the velocity-based multiple change-point inference algorithm. We evaluate both methods on human manipulation movements recorded from several participants with a marker-based motion tracking system. The first evaluations are done on simple one-handed point-to-point movements, followed by an evaluation on a complex dual arm manipulation task. The results show, that the presented approaches are able to identify basic movements as well as their concatenations into more complex, labeled actions.
Lisa Gutzeit
ICPR1
2022 The Influence of Labeling Techniques in Classifying Human Manipulation Movement of Different Speed
abstract
In this work, we investigate the influence of labeling methods on the classification of human movements on data recorded using a marker-based motion capture system. The dataset is labeled using two different approaches, one based on video data of the movements, the other based on the movement trajectories recorded using the motion capture system. The dataset is labeled using two different approaches, one based on video data of the movements, the other based on the movement trajectories recorded using the motion capture system. The data was recorded from one participant performing a stacking scenario comprising simple arm movements at three different speeds (slow, normal, fast). Machine learning algorithms that include k-Nearest Neighbor, Random Forest, Extreme Gradient Boosting classifier, Convolutional Neural networks (CNN), Long Short-Term Memory networks (LSTM), and a combination of CNN-LSTM networks are compared on their performance in recognition of these arm movements. The models were trained on actions performed on slow and normal speed movements segments and generalized on actions consisting of fast-paced human movement. It was observed that all the models trained on normal-paced data labeled using trajectories have almost 20% improvement in accuracy on test data in comparison to the models trained on data labeled using videos of the performed experiments.
Sadique Adnan Siddiqui, Lisa Gutzeit, Frank Kirchner
ICPRAM2
2021 A Comparison of Few-shot Classification of Human Movement Trajectories
Lisa Gutzeit
ICPRAM1