Markus Eisenbach 0001

dblp:86/6537-1 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0003-4951-622XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 9 since 2021Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
YearPublicationVenuePosition
2025 Including Semantic Information via Word Embeddings for Skeleton-based Action Recognition
abstract
Effective human action recognition is widely used for cobots in Industry 4.0 to assist in assembly tasks. However, conventional skeleton-based methods often lose keypoint semantics, limiting their effectiveness in complex interactions. In this work, we introduce a novel approach to skeleton-based action recognition that enriches input representations by leveraging word embeddings to encode semantic information. Our method replaces one-hot encodings with semantic volumes, enabling the model to capture meaningful relationships between joints and objects. Through extensive experiments on multiple assembly datasets, we demonstrate that our approach significantly improves classification performance, and enhances generalization capabilities by simultaneously supporting different skeleton types and object classes. Our findings highlight the potential of incorporating semantic information to enhance skeleton-based action recognition in dynamic and diverse environments.
Dustin Aganian, Erik Franze, Markus Eisenbach 0001, Horst-Michael Groß
IJCNN3
2024 Few-Shot Object Detection: A Comprehensive Survey
abstract
Humans are able to learn to recognize new objects even from a few examples. In contrast, training deep-learning-based object detectors requires huge amounts of annotated data. To avoid the need to acquire and annotate these huge amounts of data, few-shot object detection (FSOD) aims to learn from few object instances of new categories in the target domain. In this survey, we provide an overview of the state of the art in FSOD. We categorize approaches according to their training scheme and architectural layout. For each type of approach, we describe the general realization as well as concepts to improve the performance on novel categories. Whenever appropriate, we give short takeaways regarding these concepts in order to highlight the best ideas. Eventually, we introduce commonly used datasets and their evaluation protocols and analyze the reported benchmark results. As a result, we emphasize common challenges in evaluation and identify the most promising current trends in this emerging field of FSOD.
Mona Köhler, Markus Eisenbach 0001, Horst-Michael Groß
IEEE Trans. Neural Networks Learn. Syst.2
2023 Fusing Hand and Body Skeletons for Human Action Recognition in Assembly
Dustin Aganian, Mona Köhler, Benedict Stephan, Markus Eisenbach 0001, Horst-Michael Groß
ICANN (1)4
2023 ATTACH Dataset: Annotated Two-Handed Assembly Actions for Human Action Understanding
abstract
With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions during assembly. To effectively train models for this task, a dataset containing suitable assembly actions in a realistic setting is cru-cial. For this purpose, we present the ATTACH dataset, which contains 51.6 hours of assembly with 95.2k annotated fine-grained actions monitored by three cameras, which represent potential viewpoints of a cobot. Since in an assembly context workers tend to perform different actions simultaneously with their two hands, we annotated the performed actions for each hand separately. Therefore, in the ATTACH dataset, more than 68% of annotations overlap with other annotations, which is many times more than in related datasets, typically featuring more simplistic assembly tasks. For better generalization with respect to the background of the working area, we did not only record color and depth images, but also used the Azure Kinect body tracking SDK for estimating 3D skeletons of the worker. To create a first baseline, we report the performance of state-of-the-art methods for action recognition as well as action detection on video and skeleton-sequence inputs. The dataset is available at https://www.tu-ilmenau.de/neurob/data-sets-code/attach-dataset.
Dustin Aganian, Benedict Stephan, Markus Eisenbach 0001, Corinna Stretz, Horst-Michael Groß
ICRA3
2023 A Little Bit Attention Is All You Need for Person Re-Identification
abstract
Person re-identification plays a key role in applications where a mobile robot needs to track its users over a long period of time, even if they are partially unobserved for some time, in order to follow them or be available on demand. In this context, deep-learning-based real-time feature extraction on a mobile robot is often performed on special-purpose devices whose computational resources are shared for multiple tasks. Therefore, the inference speed has to be taken into account. In contrast, person re-identification is often improved by architectural changes that come at the cost of significantly slowing down inference. Attention blocks are one such example. We will show that some well-performing attention blocks used in the state of the art are subject to inference costs that are far too high to justify their use for mobile robotic applications. As a consequence, we propose an attention block that only slightly affects the inference speed while keeping up with much deeper networks or more complex attention blocks in terms of re-identification accuracy. We perform extensive neural architecture search to derive rules at which locations this attention block should be integrated into the architecture in order to achieve the best trade-off between speed and accuracy. Finally, we confirm that the best performing configuration on a re-identification benchmark also performs well on an indoor robotic dataset.
Markus Eisenbach 0001, Jannik Lübberstedt, Dustin Aganian, Horst-Michael Groß
ICRA1
2023 How Object Information Improves Skeleton-based Human Action Recognition in Assembly Tasks
abstract
As the use of collaborative robots (cobots) in industrial manufacturing continues to grow, human action recognition for effective human-robot collaboration becomes increasingly important. This ability is crucial for cobots to act autonomously and assist in assembly tasks. Recently, skeleton-based approaches are often used as they tend to generalize better to different people and environments. However, when processing skeletons alone, information about the objects a human interacts with is lost. Therefore, we present a novel approach of integrating object information into skeleton-based action recognition. We enhance two state-of-the-art methods by treating object centers as further skeleton joints. Our experiments on the assembly dataset IKEA ASM show that our approach improves the performance of these state-of-the-art methods to a large extent when combining skeleton joints with objects predicted by a state-of-the-art instance segmentation model. Our research sheds light on the benefits of combining skeleton joints with object information for human action recognition in assembly tasks. We analyze the effect of the object detector on the combination for action classification and discuss the important factors that must be taken into account.
Dustin Aganian, Mona Köhler, Sebastian Baake, Markus Eisenbach 0001, Horst-Michael Groß
IJCNN4
2022 On the Importance of Label Encoding and Uncertainty Estimation for Robotic Grasp Detection
abstract
Automated grasping of arbitrary objects is an essential skill for many applications such as smart manufacturing and human robot interaction. This makes grasp detection a vital skill for automated robotic systems. Recent work in model-free grasp detection uses point cloud data as input and typically outperforms the earlier work on RGB(D)-based methods. We show that RGB(D)-based methods are being underestimated due to suboptimal label encodings used for training. Using the evaluation pipeline of the GraspNet-1Billion dataset, we investigate different encodings and propose a novel encoding that significantly improves grasp detection on depth images. Additionally, we show shortcomings of the 2D rectangle grasps supplied by the GraspNet-1Billion dataset and propose a filtering scheme by which the ground truth labels can be improved significantly. Furthermore, we apply established methods for uncertainty estimation on our trained models since knowing when we can trust the model's decisions provides an advantage for real-world application. By doing so, we are the first to directly estimate uncertainties of detected grasps. We also investigate the applicability of the estimated aleatoric and epistemic uncertainties based on their theoretical properties. Additionally, we demonstrate the correlation between estimated uncertainties and grasp quality, thus improving selection of high quality grasp detections. By all these modifications, our approach using only depth images can compete with point-cloud-based approaches for grasp detection despite the lower degree of freedom for grasp poses in 2D image space.
Benedict Stephan, Dustin Aganian, Lars Hinneburg, Markus Eisenbach 0001, Steffen Müller 0001, Horst-Michael Groß
IROS4
2021 Revisiting Loss Functions for Person Re-identification
Dustin Aganian, Markus Eisenbach 0001, Joachim Wagner 0005, Daniel Seichter, Horst-Michael Groß
ICANN (5)2
2021 Evaluation of Transfer Learning for Visual Road Condition Assessment
Christoph Peter Balada, Markus Eisenbach 0001, Horst-Michael Groß
ICANN (5)2
2019 Improving Visual Road Condition Assessment by Extensive Experiments on the Extended GAPs Dataset
abstract
Aging public roads need frequent inspections in order to guarantee their permanent availability. In many countries, this includes the standardized visual assessment of millions of images. Due to the lack of sophisticated approaches, often, the evaluation is done manually and therefore requires excessive manual labor. GAPs is the most extensive publicly available dataset that provides standardized, high-quality images for training deep neural networks for pavement distress detection. We further enlarge this dataset and provide refined annotations. By conducting extensive experiments on the GAPs dataset, we improve the performance of automated visual road condition assessment. We evaluate the performance gain of several modern neural network architectures and advanced training techniques.
Ronny Stricker, Markus Eisenbach 0001, Maximilian Sesselmann, Klaus Debes, Horst-Michael Groß
IJCNN2
2017 Mobile robot companion for walking training of stroke patients in clinical post-stroke rehabilitation
abstract
This paper introduces a novel robot-based approach to the stroke rehabilitation scenario, in which a mobile robot companion accompanies stroke patients during their walking self-training. This assistance enables them to move freely in the clinic practicing both their mobility and spatial orientation skills. Based on a set of questions for systematic evaluating the autonomy and practicability of assistive robots and a three-stage approach in conducting function and user tests in the clinical setting, we present the results of user trials performed with N=30 stroke patients in a stroke rehabilitation center between 4/2015 and 3/2016. This allowed us to make an honest inventory of the strengths and weaknesses of the developed robot companion and its already achieved practicability for clinical use. The results of the user studies show that patients and fellow patients were very open-minded and accepted the robotic coach. The robot motivated them for independent training and leaving their room, despite severe consequences of stroke (lower limbs paralysis, speech/language problems, loss of orientation, depression), provided a very self-determined training regime, and encouraged them to expand the radius of their training in the clinic.
Horst-Michael Groß, Sibylle Meyer, Andrea Scheidig, Markus Eisenbach 0001, Steffen Müller 0001, Thanh Quang Trinh, Tim Wengefeld, Andreas Bley, Christian Martin 0001, Christa Fricke
ICRA4
2017 How to get pavement distress detection ready for deep learning? A systematic approach
abstract
Road condition acquisition and assessment are the key to guarantee their permanent availability. In order to maintain a country's whole road network, millions of high-resolution images have to be analyzed annually. Currently, this requires cost and time excessive manual labor. We aim to automate this process to a high degree by applying deep neural networks. Such networks need a lot of data to be trained successfully, which are not publicly available at the moment. In this paper, we present the GAPs dataset, which is the first freely available pavement distress dataset of a size, large enough to train high-performing deep neural networks. It provides high quality images, recorded by a standardized process fulfilling German federal regulations, and detailed distress annotations. For the first time, this enables a fair comparison of research in this field. Furthermore, we present a first evaluation of the state of the art in pavement distress detection and an analysis of the effectiveness of state of the art regularization techniques on this dataset.
Markus Eisenbach 0001, Ronny Stricker, Daniel Seichter, Karl Amende, Klaus Debes, Maximilian Sesselmann, Dirk Ebersbach, Ulrike Stoeckert, Horst-Michael Groß
IJCNN1
2016 Cooperative multi-scale Convolutional Neural Networks for person detection
abstract
Robust person detection is required by many computer vision applications. We present a deep learning approach, that combines three Convolutional Neural Networks to detect people at different scales, which is the first time that a multi-resolution model is combined with deep learning techniques in the pedestrian detection domain. The networks learn features from raw pixel information, which is also rare for pedestrian detection. Due to the use of multiple Convolutional Neural Networks at different scales, the learned features are specific for far, medium, and near scales respectively, and thus, the overall performance is improved. Furthermore, we show, that neural approaches can also be applied successfully for the remaining processing steps of classification and non-maximum suppression. The evaluation on the most popular Caltech pedestrian detection benchmark shows that the proposed method can compete with state of the art methods without using Caltech training data and without fine tuning. Therefore, it is shown that our method generalizes well on domains it is not trained on.
Markus Eisenbach 0001, Daniel Seichter, Tim Wengefeld, Horst-Michael Groß
IJCNN1
2015 Evaluation of multi feature fusion at score-level for appearance-based person re-identification
abstract
Robust appearance-based person re-identification can only be achieved by combining multiple diverse features describing the subject. Since individual features perform different, it is not trivial to combine them. Often this problem is bypassed by concatenating all feature vectors and learning a distance metric for the combined feature vector. However, to perform well, metric learning approaches need many training samples which are not available in most real-world applications. In contrast, in our approach we perform score-level fusion to combine the matching scores of different features. To evaluate which score-level fusion techniques perform best for appearance-based person re-identification, we examine several score normalization and feature weighting approaches employing the the widely used and very challenging VIPeR dataset. Experiments show that in fusing a large ensemble of features, the proposed score-level fusion approach outperforms linear metric learning approaches which fuse at feature-level. Furthermore, a combination of linear metric learning and score-level fusion even outperforms the currently best non-linear kernel-based metric learning approaches, regarding both accuracy and computation time.
Markus Eisenbach 0001, Alexander Kolarow, Alexander Vorndran, Julia Niebling, Horst-Michael Groß
IJCNN1
2015 User recognition for guiding and following people with a mobile robot in a clinical environment
abstract
Rehabilitative follow-up care is important for stroke patients to regain their motor and cognitive skills. We aim to develop a robotic rehabilitation assistant for walking exercises in late stages of rehabilitation. The robotic rehab assistant is to accompany inpatients during their self-training, practicing both mobility and spatial orientation skills. To hold contact to the patient, even after temporally full occlusions, robust user re-identification is essential. Therefore, we implemented a person re-identification module that continuously re-identifies the patient, using only few amount of the robot's processing resources. It is robust to varying illumination and occlusions. State-of-the-art performance is confirmed on a standard benchmark dataset, as well as on a recorded scenario-specific dataset. Additionally, the benefit of using a visual re-identification component is verified by live-tests with the robot in a stroke rehab clinic.
Markus Eisenbach 0001, Alexander Vorndran, Sven Sorge, Horst-Michael Groß
IROS1
2013 APFel: The intelligent video analysis and surveillance system for assisting human operators
abstract
The rising need for security in the last years has led to an increased use of surveillance cameras in both public and private areas. The increasing amount of footage makes it necessary to assist human operators with automated systems to monitor and analyze the video data in reasonable time. In this paper we summarize our work of the past three years in the field of intelligent and automated surveillance. Our proposed system extends the common active monitoring of camera footage into an intelligent automated investigative person-search and walk path reconstruction of a selected person within hours of image data. Our system is evaluated and tested under life-like conditions in real-world surveillance scenarios. Our experiments show that with our system an operator can reconstruct a case in a fraction of time, compared to manually searching the recorded data.
Alexander Kolarow, Konrad Schenk, Markus Eisenbach 0001, Michael Dose, Michael Brauckmann, Klaus Debes, Horst-Michael Groß
AVSS3
2012 View Invariant Appearance-Based Person Reidentification Using Fast Online Feature Selection and Score Level Fusion
abstract
Fast and robust person reidentification is an important task in multi-camera surveillance and automated access control. We present an efficient appearance-based algorithm, able to reidentify a person regardless of occlusions, distance to the camera, and changes in view and lighting. The use of fast online feature selection techniques enables us to perform reidentification in hyper-real-time for a multi-camera system, by taking only 10 seconds for evaluating 100 minutes of HD-video data. We demonstrate, that our approach surpasses current appearance-based state-of-the-art in reidentification quality and computational speed and sets a new reference in non-biometric reidentification.
Markus Eisenbach 0001, Alexander Kolarow, Konrad Schenk, Klaus Debes, Horst-Michael Groß
AVSS1
2012 Automatic Calibration of Multiple Stationary Laser Range Finders Using Trajectories
abstract
Laser based detection and tracking of persons can be used for numerous tasks, like statistical measurements for determining bottlenecks in public buildings, optimizing passenger flow, or planning camera placement. Only a network of multiple LRF is sufficient to fulfill these tasks in larger spaces. Calibrating multiple LRF into a global coordinate system is usually done by hand in a time consuming procedure. In this paper, we address the problem of automatically calibrating such a sensor network. We introduce an automatic calibration mechanism, which is able to obtain the positions and orientations of all LRF in a global coordinate system, without any prior knowledge of the scene. Our approach is based on comparing person tracks, determined by each individual LRF unit and matching them in order to obtain constraints between the LRF units. By resolving these constraints, we are able to estimate the poses of all LRF. We evaluate and compare our method to the current state of the art approach methodically and experimentally. Experiments show that our calibration approach outperforms this approach.
Konrad Schenk, Alexander Kolarow, Markus Eisenbach 0001, Klaus Debes, Horst-Michael Groß
AVSS3
2012 Vision-based hyper-real-time object tracker for robotic applications
abstract
Fast vision-based object and person tracking is important for various applications in mobile robotics and Human-Robot Interaction. While current state-of-the-art methods use descriptive features for visual tracking, we propose a novel approach using a sparse template based feature set, which is drawn from homogeneous regions on the object to be tracked. Using only a small number of simple features, without complex descriptors in combination with logarithmic-search, the tracker performs at hyper-real-time on HD-images without the use of parallelized hardware. Detailed benchmark experiments show that it outperforms most other state-of-the-art approaches for real-time object and person tracking in quality and runtime. In the experiments we also show the robustness of the tracker and evaluate the effects of different initialization methods, feature sets, and parameters on the tracker. Although we focus on the scenario of person and object tracking in robot applications, the proposed tracker can be used for a variety of other tracking tasks.
Alexander Kolarow, Michael Brauckmann, Markus Eisenbach 0001, Konrad Schenk, Erik Einhorn, Klaus Debes, Horst-Michael Groß
IROS3
2012 Automatic calibration of a stationary network of laser range finders by matching movement trajectories
abstract
Laser based detection and tracking of persons can be used for numerous tasks. While a single laser range finder (LRF) is sufficient for detecting and tracking persons on a mobile robot platform, a network of multiple LRF is required to observe persons in larger spaces. Calibrating multiple LRF into a global coordinate system is usually done by hand in a time consuming procedure. An automatic calibration mechanism for such a sensor network is introduced in this paper. Without the need of prior knowledge about the environment, this mechanism is able to obtain the positions and orientations of all LRF in a global coordinate system. By comparing person tracks, determined for each individual LRF unit and matching them, constrains between the LRF units can be calculated. We are able to estimate the poses of all LRF by resolving these constrains. We evaluate and compare our method to the current state of the art approach methodically and experimentally. Experiments show that our calibration approach outperforms this approach.
Konrad Schenk, Alexander Kolarow, Markus Eisenbach 0001, Klaus Debes, Horst-Michael Groß
IROS3