Matthias Kerzel

dblp:70/11043 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0002-1378-0435ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 7 first-author · 14 since 2021Systems, architecture and hardware · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Workshop Proposal: Dungeons, Neurons, and Dialogues 2 Edition: Social Interaction Dynamics in Contextual Games (DnD-SIDC)
abstract
Join us for the second edition of the “Dungeons, Neurons, and Dialogues: Social Interaction Dynamics in Contextual Games” (DnD-SIDC) workshop at HRI 2025! This engaging event will delve into the exciting intersection of human-robot interaction (HRI) and contextual games, exploring how embodied agents can enhance social dynamics within rich, narrative-driven environments. Participants will engage in thought-provoking discussions centered on empathy, trust, and cooperation as they navigate complex emotional landscapes in gaming scenarios. The workshop features a unique “Quest for Solutions” activity, where attendees will collaborate in small groups to tackle specific challenges and design innovative pre-registration studies. This format fosters interdisciplinary collaboration among researchers, practitioners, and students from diverse fields, including robotics, psvchology, game design, and AI, By leveraging the advancements in technology and the rising popularity of digital platforms, this workshop aims to inspire new ideas and partnerships that contribute to the sustainable development of social interactions in virtual spaces. Don't miss this opportunity to explore cutting-edge research and contribute to the future of social computing in gaming! Join us and embark on a collaborative adventure in the realm of social interaction dynamics!
Pablo V. A. Barros, Laura Triglia, Nikhil Churamani, Matthias Kerzel
HRI4
2025 Shaken, Not Stirred: A Novel Dataset for Visual Understanding of Glasses in Human-Robot Bartending Tasks
abstract
Datasets for object detection often do not account for enough variety of glasses, due to their transparent and reflective properties. Specifically, open-vocabulary object detectors, widely used in embodied robotic agents, fail to distinguish subclasses of glasses. This scientific gap poses an issue for robotic applications that suffer from accumulating errors between detection, planning, and action execution. This paper introduces a novel method for acquiring real-world data from RGB-D sensors that minimizes human effort. We propose an auto-labeling pipeline that generates labels for all the acquired frames based on the depth measurements. We provide a novel real-world glass object dataset3that was collected on the Neuro-Inspired COLlaborator (NICOL), a humanoid robot platform. The dataset consists of 7850 images recorded from five different cameras. We show that our trained baseline model outperforms state-of-the-art open-vocabulary approaches. In addition, we deploy our baseline model in an embodied agent approach to the NICOL platform, on which it achieves a success rate of 81% in a human-robot bartending scenario.
Lukás Gajdosech, Hassan Ali 0005, Jan-Gerrit Habekost, Martin Madaras, Matthias Kerzel, Stefan Wermter
IROS5
2024 Learning Low-Level Causal Relations Using a Simulated Robotic Arm
Miroslav Cibula, Matthias Kerzel, Igor Farkas
ICANN (10)2
2024 Details Make a Difference: Object State-Sensitive Neurorobotic Task Planning
Xiaowen Sun, Xufeng Zhao 0002, Jae Hee Lee 0001, Wenhao Lu, Matthias Kerzel, Stefan Wermter
ICANN (4)5
2023 Acquisition and Formalization of Tacit Knowledge for Value Chain Generation in Local Production Networks
abstract
Interest in producing goods locally again has risen. This leads to new challenges for companies producing locally, especially since they are mostly small enterprises and do not always have resources to adapt Industry 4.0 technologies. Therefore, collaborating in networks can strengthen local production. We propose an online system with an underlying planning component that is supported by a large-scale language model to coordinate value chains within a network by utilizing the tacit production knowledge within the companies. Before any type of information processing can happen, however, the data – in this case, the tacit knowledge – needs to be acquired and formalized in such a way that is easy and quick, but also sufficient enough in detail and quality for the computer system. To this end, we conducted a study with 16 participants to simulate the collection of knowledge regarding the production of four pieces of furniture by having them describe simplified production steps. We analyze the results and show that the use of the collaborative system has a positive effect on the soundness of resulting production plans. In a second step, we utilize artificial intelligence methods to fill incomplete plans. Results and implications for future research are presented as well.
Matthias Kerzel, Julia Markert, Emad Aghajanzadeh, Stephanie von Riegen, Lothar Hotz, Pascal Krenz
ECAI1
2023 CycleIK: Neuro-inspired Inverse Kinematics
abstract
Abstract The paper introduces CycleIK, a neuro-robotic approach that wraps two novel neuro-inspired methods for the inverse kinematics (IK) task—a Generative Adversarial Network (GAN), and a Multi-Layer Perceptron architecture. These methods can be used in a standalone fashion, but we also show how embedding these into a hybrid neuro-genetic IK pipeline allows for further optimization via sequential least-squares programming (SLSQP) or a genetic algorithm (GA). The models are trained and tested on dense datasets that were collected from random robot configurations of the new Neuro-Inspired COLlaborator (NICOL), a semi-humanoid robot with two redundant 8-DoF manipulators. We utilize the weighted multi-objective function from the state-of-the-art BioIK method to support the training process and our hybrid neuro-genetic architecture. We show that the neural models can compete with state-of-the-art IK approaches, which allows for deployment directly to robotic hardware. Additionally, it is shown that the incorporation of the genetic algorithm improves the precision while simultaneously reducing the overall runtime.
Jan-Gerrit Habekost, Erik Strahl, Philipp Allgeuer, Matthias Kerzel, Stefan Wermter
ICANN (1)4
2023 Clarifying the Half Full or Half Empty Question: Multimodal Container Classification
abstract
Abstract Multimodal integration is a key component of allowing robots to perceive the world. Multimodality comes with multiple challenges that have to be considered, such as how to integrate and fuse the data. In this paper, we compare different possibilities of fusing visual, tactile and proprioceptive data. The data is directly recorded on the NICOL robot in an experimental setup in which the robot has to classify containers and their content. Due to the different nature of the containers, the use of the modalities can wildly differ between the classes. We demonstrate the superiority of multimodal solutions in this use case and evaluate three fusion strategies that integrate the data at different time steps. We find that the accuracy of the best fusion strategy is 15% higher than the best strategy using only one singular sense.
Josua Spisak, Matthias Kerzel, Stefan Wermter
ICANN (1)2
2022 Sim-to-Real Neural Learning with Domain Randomisation for Humanoid Robot Grasping
Connor Gaede, Matthias Kerzel, Erik Strahl, Stefan Wermter
ICANN (1)2
2022 Learning Flexible Translation Between Robot Actions and Language Descriptions
Ozan Özdemir, Matthias Kerzel, Cornelius Weber, Jae Hee Lee 0001, Stefan Wermter
ICANN (2)2
2022 Learning Visually Grounded Human-Robot Dialog in a Hybrid Neural Architecture
Xiaowen Sun, Cornelius Weber, Matthias Kerzel, Tom Weber, Mengdi Li 0006, Stefan Wermter
ICANN (2)3
2022 What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning
abstract
Understanding spatial relations is essential for intelligent agents to act and communicate in the physical world. Relative directions are spatial relations that describe the relative positions of target objects with regard to the intrinsic orientation of reference objects. Grounding relative directions is more difficult than grounding absolute directions because it not only requires a model to detect objects in the image and to identify spatial relation based on this information, but it also needs to recognize the orientation of objects and integrate this information into the reasoning process. We investigate the challenging problem of grounding relative directions with end-to-end neural networks. To this end, we provide GRiD-3D, a novel dataset that features relative directions and complements existing visual question answering (VQA) datasets, such as CLEVR, that involve only absolute directions. We also provide baselines for the dataset with two established end-to-end VQA models. Experimental evaluations show that answering questions on relative directions is feasible when questions in the dataset simulate the necessary subtasks for grounding relative directions. We discover that those subtasks are learned in an order that reflects the steps of an intuitive pipeline for processing relative directions.
Jae Hee Lee 0001, Matthias Kerzel, Kyra Ahrens, Cornelius Weber, Stefan Wermter
IJCAI2
2021 Pruning Neural Networks with Supermasks
abstract
The Lottery Ticket hypothesis by Frankle and Carbin states that a randomly initialized dense network contains a smaller subnetwork that, when trained in isolation, will match the performance of the original network.However, identifying this pruned subnetwork usually requires repeated training to determine optimal pruning thresholds.We present a novel approach to accelerate the pruning: By methodically evaluating different Supermasks, the threshold for selecting neurons as part of a pruned Lottery Ticket network can be determined without additional training.We evaluate the method on the MNIST dataset and achieve a size reduction of over 60% without a drop in performance.
Vincent Rolfs, Matthias Kerzel, Stefan Wermter
ESANN2
2021 Robotic Occlusion Reasoning for Efficient Object Existence Prediction
abstract
Reasoning about potential occlusions is essential for robots to efficiently predict whether an object exists in an environment. Though existing work shows that a robot with active perception can achieve various tasks, it is still unclear if occlusion reasoning can be achieved. To answer this question, we introduce the task of robotic object existence prediction: when being asked about an object, a robot needs to move as few steps as possible around a table with randomly placed objects to predict whether the queried object exists. To address this problem, we propose a novel recurrent neural network model that can be jointly trained with supervised and reinforcement learning methods using a curriculum training strategy. Experimental results show that 1) both active perception and occlusion reasoning are necessary to successfully achieve the task; 2) the proposed model demonstrates a good occlusion reasoning ability by achieving a similar prediction accuracy to an exhaustive exploration baseline while requiring only about 10% of the baseline’s number of movement steps on average; and 3) the model generalizes to novel object combinations with a moderate loss of accuracy.
Mengdi Li 0006, Cornelius Weber, Matthias Kerzel, Jae Hee Lee 0001, Zheni Zeng, Zhiyuan Liu 0001, Stefan Wermter
IROS3
2021 Solving visual object ambiguities when pointing: an unsupervised learning approach
abstract
Abstract Whenever we are addressing a specific object or refer to a certain spatial location, we are using referential or deictic gestures usually accompanied by some verbal description. Particularly, pointing gestures are necessary to dissolve ambiguities in a scene and they are of crucial importance when verbal communication may fail due to environmental conditions or when two persons simply do not speak the same language. With the currently increasing advances of humanoid robots and their future integration in domestic domains, the development of gesture interfaces complementing human–robot interaction scenarios is of substantial interest. The implementation of an intuitive gesture scenario is still challenging because both the pointing intention and the corresponding object have to be correctly recognized in real time. The demand increases when considering pointing gestures in a cluttered environment, as is the case in households. Also, humans perform pointing in many different ways and those variations have to be captured. Research in this field often proposes a set of geometrical computations which do not scale well with the number of gestures and objects and use specific markers or a predefined set of pointing directions. In this paper, we propose an unsupervised learning approach to model the distribution of pointing gestures using a growing-when-required (GWR) network. We introduce an interaction scenario with a humanoid robot and define the so-called ambiguity classes. Our implementation for the hand and object detection is independent of any markers or skeleton models; thus, it can be easily reproduced. Our evaluation comparing a baseline computer vision approach with our GWR model shows that the pointing-object association is well learned even in cases of ambiguities resulting from close object proximity.
Doreen Jirak, David Biertimpel, Matthias Kerzel, Stefan Wermter
Neural Comput. Appl.3
2020 Neuro-Genetic Visuomotor Architecture for Robotic Grasping
Matthias Kerzel, Josua Spisak, Erik Strahl, Stefan Wermter
ICANN (2)1
2020 Model Mediated Teleoperation with a Hand-Arm Exoskeleton in Long Time Delays Using Reinforcement Learning
abstract
Telerobotic systems must adapt to new environmental conditions and deal with high uncertainty caused by long-time delays. As one of the best alternatives to human-level intelligence, Reinforcement Learning (RL) may offer a solution to cope with these issues. This paper proposes to integrate RL with the Model Mediated Teleoperation (MMT) concept. The teleoperator interacts with a simulated virtual environment, which provides instant feedback. Whereas feedback from the real environment is delayed, feedback from the model is instantaneous, leading to high transparency. The MMT is realized in combination with an intelligent system with two layers. The first layer utilizes Dynamic Movement Primitives (DMP) which accounts for certain changes in the avatar environment. And, the second layer addresses the problems caused by uncertainty in the model using RL methods. Augmented reality was also provided to fuse the avatar device and virtual environment models for the teleoperator. Implemented on DLR's Exodex Adam hand-arm haptic exoskeleton, the results show RL methods are able to find different solutions when changes are applied to the object position after the demonstration. The results also show DMPs to be effective at adapting to new conditions where there is no uncertainty involved.
Hadi Beik-Mohammadi, Matthias Kerzel, Benedikt Pleintinger, Thomas Hulin, Philipp Reisich, Annika Schmidt, Aaron Pereira, Stefan Wermter, Neal Y. Lii
RO-MAN2
2019 Mixed-Reality Deep Reinforcement Learning for a Reach-to-grasp Task
Hadi Beik-Mohammadi, Mohammad-Ali Zamani, Matthias Kerzel, Stefan Wermter
ICANN (1)3
2019 Curious Meta-Controller: Adaptive Alternation between Model-Based and Model-Free Control in Deep Reinforcement Learning
abstract
Recent success in deep reinforcement learning for continuous control has been dominated by model-free approaches which, unlike model-based approaches, do not suffer from representational limitations in making assumptions about the world dynamics and model errors inevitable in complex domains. However, they require a lot of experiences compared to model-based approaches that are typically more sample-efficient. We propose to combine the benefits of the two approaches by presenting an integrated approach called Curious Meta-Controller. Our approach alternates adaptively between model-based and model-free control using a curiosity feedback based on the learning progress of a neural model of the dynamics in a learned latent space. We demonstrate that our approach can significantly improve the sample efficiency and achieve near-optimal performance on learning robotic reaching and grasping tasks from raw-pixel input in both dense and sparse reward settings.
Muhammad Burhan Hafez, Cornelius Weber, Matthias Kerzel, Stefan Wermter
IJCNN3
2019 Neuro-Robotic Haptic Object Classification by Active Exploration on a Novel Dataset
abstract
We present an embodied neural model for haptic object classification by active haptic exploration with the humanoid robot NICO. When NICO's newly developed robotic hand closes around an object, multiple sensory readings from a tactile fingertip sensor, motor positions, and motor currents are recorded. We created a haptic dataset with 83200 haptic measurements, based on 100 samples of each of 16 different objects, every sample containing 52 measurements. First, we provide an analysis of neural classification models with regard to isolated haptic sensory channels for object classification. Based on this, we develop a series of neural models (MLP, CNN, LSTM) that integrate the haptic sensory channels to classify explored objects. As an initial baseline, our best model achieves a 66.6% classification accuracy over 16 objects. We show that this result is due to the ability of the network to integrate the haptic data both over time domain and over different haptic sensory channels. Furthermore, we make the dataset publically available to address the issue of sparse haptic datasets for machine learning research.
Matthias Kerzel, Erik Strahl, Connor Gaede, Emil Gasanov, Stefan Wermter
IJCNN1
2019 Exploring Low-level and High-level Transfer Learning for Multi-task Facial Recognition with a Semi-supervised Neural Network
abstract
Facial recognition tasks like identity, age, gender, and emotion recognition received substantial attention in recent years. Their deployment in robotic platforms became necessary for the characterization of most of the non-verbal Human-Robot Interaction (HRI) scenarios. In this regard, deep convolution neural networks have shown to be effective on processing different facial representations but with a high cost: to achieve maximum generalization, they require an enormous amount of task-specific labeled data. This paper proposes a unified semi-supervised deep neural model to address this problem. Our hybrid model is composed of an unsupervised deep generative adversarial network which learns fundamental characteristics of facial representations, and a set of convolution channels that fine-tunes the high-level facial concepts for the recognition of identity, age group, gender, and facial expressions. Our network employs progressive lateral connections between the convolution channels so that they share the high-abstraction particularities of each of these tasks in order to reduce the necessity of a large amount of strongly labeled training data. We propose a series of experiments to evaluate each individual mechanism of our hybrid model, in particular, the impact of the progressive connections on learning the specific facial recognition tasks and we observe that our model achieves a better performance when compared to task-specific models.
Pablo V. A. Barros, Erik Fließwasser, Matthias Kerzel, Stefan Wermter
IROS3
2019 Continuous convolutional object tracking in developmental robot scenarios
abstract
Tracking arbitrary objects in natural environments is a challenging task in visual computing. A central problem is the need to adapt to changing appearances under strong transformation and occlusion. We propose a tracking framework that utilises the strength of Convolutional Neural Networks to create a robust and adaptive model of the object from training data produced during tracking. An incremental update mechanism provides increased performance and reduces the computational costs for training during tracking, allowing for robust real-time tracking with state-of-the-art performance. Together with optimisations for deploying the framework on humanoid robots and distributed devices, this shows its viability for research in developmental robotics on questions around infant cognition or active exploration.
Stefan Heinrich, Peer Springstübe, Tobias Knöppler, Matthias Kerzel, Stefan Wermter
Neurocomputing4
2018 Slowness-based neural visuomotor control with an Intrinsically motivated Continuous Actor-Critic
Muhammad Burhan Hafez, Matthias Kerzel, Cornelius Weber, Stefan Wermter
ESANN2
2018 Classification of MRI Migraine Medical Data Using 3D Convolutional Neural Network
Hwei Geok Ng, Matthias Kerzel, Jan Mehnert, Arne May, Stefan Wermter
ICANN (3)2
2018 Accelerating Deep Continuous Reinforcement Learning through Task Simplification
abstract
Robotic motor policies can, in theory, be learned via deep continuous reinforcement learning. In practice, however, collecting the enormous amount of required training samples in realistic time, surpasses the possibilities of many robotic platforms. To address this problem, we propose a novel method for accelerating the learning process by task simplification inspired by the Goldilocks effect known from developmental psychology. We present results on a reach-for-grasp task that is learned with the Deep Deterministic Policy Gradients (DDPG) algorithm. Task simplification is realized by initially training the system with “larger-than-life” training objects that adapt their reachability dynamically during training. We achieve a significant acceleration compared to the unaltered training setup. We describe modifications to the DDPG algorithm with regard to the replay buffer to prevent artifacts during the learning process from the simplified learning instances while maintaining the speed of learning. With this result, we contribute towards the realistic application of deep reinforcement learning on robotic platforms.
Matthias Kerzel, Hadi Beik-Mohammadi, Mohammad-Ali Zamani, Stefan Wermter
IJCNN1
2018 Deep Neural Object Analysis by Interactive Auditory Exploration with a Humanoid Robot
abstract
We present a novel approach for interactive auditory object analysis with a humanoid robot. The robot elicits sensory information by physically shaking visually indistinguishable plastic capsules. It gathers the resulting audio signals from microphones that are embedded into the robotic ears. A neural network architecture learns from these signals to analyze properties of the contents of the containers. Specifically, we evaluate the material classification and weight prediction accuracy and demonstrate that the framework is fairly robust to acoustic real-world noise.
Manfred Eppe, Matthias Kerzel, Erik Strahl, Stefan Wermter
IROS2
2018 Object Detection and Pose Estimation Based on Convolutional Neural Networks Trained with Synthetic Data
abstract
Instance-based object detection and fine pose estimation is an active research problem in computer vision. While the traditional interest-point-based approaches for pose estimation are precise, their applicability in robotic tasks relies on controlled environments and rigid objects with detailed textures. CNN-based approaches, on the other hand, have shown impressive results in uncontrolled environments for more general object recognition tasks like category-based coarse pose estimation, but the need of large datasets of fully-annotated training images makes them unfavourable for tasks like instance-based pose estimation. We present a novel approach that combines the robustness of CNNs with a fine-resolution instance-based 3D pose estimation, where the model is trained with fully-annotated synthetic training data, generated automatically from the 3D models of the objects. We propose an experimental setup in which we can carefully examine how the model trained with synthetic data performs on real images of the objects. Results show that the proposed model can be trained only with synthetic renderings of the objects' 3D models and still be successfully applied on images of the real objects, with precision suitable for robotic tasks like object grasping. Based on the results, we present more general insights about training neural models with synthetic images for application on real-world images.
Josip Josifovski, Matthias Kerzel, Christoph Pregizer, Lukas Posniak, Stefan Wermter
IROS2
2018 Hear the Egg - Demonstrating Robotic Interactive Auditory Perception
abstract
We present an illustrative example of an interactive auditory perception approach performed by a humanoid robot called NICO, the Neuro Inspired COmpanion [1]. The video demonstrates a material classification task in the style of a classic TV game show. NICO and another candidate are supposed to determine the content of small plastic capsules that are visually indistinguishable. Shaking the capsules produces audio signals that range from rattling stones, over tinkling coins to swooshing sand. NICO can perceive and analyze these sounds to determine the material of the capsules content.
Erik Strahl, Matthias Kerzel, Manfred Eppe, Sascha S. Griffiths, Stefan Wermter
IROS2
2017 Neural End-to-End Self-learning of Visuomotor Skills by Environment Interaction
Matthias Kerzel, Stefan Wermter
ICANN (1)1
2017 Teaching emotion expressions to a human companion robot using deep neural architectures
abstract
Human companion robots need to be sociable and responsive towards emotions to better interact with the human environment they are expected to operate in. This paper is based on the Neuro-Inspired COmpanion robot (NICO) and investigates a hybrid, deep neural network model to teach the NICO to associate perceived emotions with expression representations using its on-board capabilities. The proposed model consists of a Convolutional Neural Network (CNN) and a Self-organising Map (SOM) to perceive the emotions expressed by a human user towards NICO and trains two parallel Multilayer Perceptron (MLP) networks to learn general as well as person-specific associations between perceived emotions and the robot's facial expressions.
Nikhil Churamani, Matthias Kerzel, Erik Strahl, Pablo V. A. Barros, Stefan Wermter
IJCNN2
2017 Haptic material classification with a multi-channel neural network
abstract
We present a novel approach for haptic material classification based on an adaptation of human haptic exploratory procedures executed by a robot arm with an optical force sensor. A multi-channel neural architecture informed by findings from human haptic perception performs a spectral analysis on vibration and texture data gathered during material exploration and integrates this analysis with information gathered on material compliance. Experimental results show a high classification accuracy on a test set of 32 common household materials. Furthermore, we show that haptic material properties, relevant for robot grasping, can be classified with a simple haptic exploration while actual material classification requires more complex exploration and computation.
Matthias Kerzel, Moaaz Ali, Hwei Geok Ng, Stefan Wermter
IJCNN1
2017 NICO - Neuro-inspired companion: A developmental humanoid robot platform for multimodal interaction
abstract
Interdisciplinary research, drawing from robotics, artificial intelligence, neuroscience, psychology, and cognitive science, is a cornerstone to advance the state-of-the-art in multimodal human-robot interaction and neuro-cognitive modeling. Research on neuro-cognitive models benefits from the embodiment of these models into physical, humanoid agents that possess complex, human-like sensorimotor capabilities for multimodal interaction with the real world. For this purpose, we develop and introduce NICO (Neuro-Inspired COmpanion), a humanoid developmental robot that fills a gap between necessary sensing and interaction capabilities and flexible design. This combination makes it a novel neuro-cognitive research platform for embodied sensorimotor computational and cognitive models in the context of multimodal interaction as shown in our results.
Matthias Kerzel, Erik Strahl, Sven Magg, Nicolás Navarro-Guerrero, Stefan Heinrich, Stefan Wermter
RO-MAN1
2013 Event Recognition during the Exploration of Line-Based Graphics in Virtual Haptic Environments
Matthias Kerzel, Christopher Habel
COSIT1