VLDB 2026 Research / reviewers in the wild / expert
Markus Hofbauer
dblp:50/5798
· DBLP profile ↗
16ranked-venue papers
6as first author
10since 2021 · last 2022
0000-0002-8167-5485ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Evaluation of Video Coding for Machines without Ground TruthabstractIn the emerging field of video coding for machines, video datasets with pristine video quality and high-quality annotations are required for a comprehensive evaluation. However, existing video datasets with detailed annotations are severely limited in size and video quality. Thus, current methods have to either evaluate their codecs on still images or on already compressed data. To mitigate this problem, we propose an evaluation method based on pseudo ground-truth data from the field of semantic segmentation to the evaluation of video coding for machines. Through extensive evaluation, this paper shows that the proposed ground-truth-agnostic evaluation method results in an acceptable absolute measurement error below 0.7 percentage points on the Bjøntegaard Delta Rate compared to using the true ground truth for mid-range bitrates. We evaluate on the three tasks of semantic segmentation, instance segmentation, and object detection. Lastly, we utilize the ground-truth-agnostic method to measure the coding performances of the VVC compared against HEVC on the Cityscapes sequences. This reveals that the coding position has a significant influence on the task performance. Kristian Fischer 0001, Markus Hofbauer, Christopher B. Kuhn, Eckehard G. Steinbach, André Kaup |
ICASSP | 2 |
| 2022 | Reverse Error Modeling for Improved Semantic SegmentationabstractWe propose the concept of error-reversing autoencoders (ERA) for correcting pixel-wise errors made by an arbitrary semantic segmentation model. For this, we reframe the segmentation model as an error function applied to the ground truth labels. Then, we train an autoencoder to reverse this error function. During testing, the autoencoder reverses the approximated error function to correct the classification errors. We consider two sources of errors. First, we target the errors made by a model despite having being trained with clean, accurately labeled images. In this case, our proposed approach achieves an improvement of around 1% on the Cityscapes data set with the state-of-the-art DeepLabV3+ model. Second, we target errors introduced by compromised images. With JPEG-compressed images as input, our approach improves the segmentation performance by over 70% for high levels of compression. The proposed architecture is simple to implement, fast to train and can be applied to any semantic segmentation model as a post-processing step. Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach |
ICIP | 2 |
| 2022 | Measuring the Influence of Image Preprocessing on the Rate-Distortion Performance of Video EncodingabstractIn this paper, we conduct an extensive analysis of the rate-distortion (RD) performance achieved by using different preprocessing steps before encoding the video. We propose a novel evaluation method called the Mean Saving-Cost Ratio (MSCR) to compare the RD performance for different preprocessing algorithms. We define MSCR as the logarithmic mean ratio of maximum bitrate savings over maximum quality cost for all parameters of a preprocessing algorithm. Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach |
ISM | 1 |
| 2022 | Improving Multimodal Object Detection with Individual Sensor MonitoringabstractMultimodal object detection fuses different sensors such as camera or LIDAR to improve the detection performance. However, individual sensor inputs can also be detrimental to a system, for example when sun glare hits a camera. In this work, we propose to monitor each sensor individually to predict when an input would lead to incorrect detections. We first train one detection network for each sensor separately, using only that sensor as input. Then, we record the performance for each single-sensor network and train an introspective performance prediction network for each sensor. Finally, we train a multimodal fusion network where we weight the impact of each sensor with its predicted performance. This allows us to dynamically adapt the fusion to reduce the influence of harmful sensor readings based only on the current data. We apply the proposed concept to the state-of-the-art AVOD architecture and evaluate on the KITTI data set. The proposed sensor monitoring system improves the mean intersection-over-union performance by 4.6%. For inputs with a low predicted performance, the proposed approach outperforms the state of the art by over 10%, demonstrating the potential of using individual sensor monitoring to react to problematic input. The proposed approach can be applied to any fusion network with two or more sensors and could also be used for classification or segmentation tasks. Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach |
ISM | 2 |
| 2022 | S2CMAF: Multi-Method Assessment Fusion for Scan-to-CAD MethodsabstractScan-to-CAD-based 3D reconstruction of indoor environments has become increasingly more popular in recent years. The inherent structure of Scan-to-CAD consists of object detection, model retrieval, and alignment. Therefore, a variety of metrics are required to assess these three aspects. This can lead to ambiguous evaluation results and incorrect quality assumptions. To impede the problem of incorrect evaluation, we introduce S2CMAF, a multi-method assessment fusion approach for Scan-to-CAD pipelines. S2CMAF merges several metrics used in evaluating these pipelines into one unique quality score. We show that S2CMAF significantly improves the correlation between Scan-to-CAD results and the ground truth, compared to the conventionally used Scan2CAD benchmark. Additionally, we train S2CMAF using different optimization techniques and demonstrate the advantages of our approach on real-world data. Driton Salihu, Adam Misik, Markus Hofbauer, Eckehard G. Steinbach |
ISM | 3 |
| 2022 | Traffic-Aware Multi-View Video Stream Adaptation for Teleoperated DrivingabstractRemote control of an autonomous vehicle by a human operator requires low delay video transmission to resolve complex situations and ensure safety. The remote operator perceives the current traffic scenario via video streams from multiple cameras. To provide the operator with the best possible scene understanding while matching the available network resources, the video streams need to be automatically adapted. In this paper, we propose a traffic-aware multi-view video stream adaptation scheme. We estimate the importance of each camera view based on the vehicle’s real-time movement in traffic. The resulting prioritization together with the total available transmission rate determines a specific bit-budget for each camera view. We optimize the video quality of each individual video stream for the given bit-budget using a quality-of-experience-driven multi-dimensional adaptation scheme. Additionally, we apply a region-of-interest mask to the rear-facing camera views. The mask removes less important areas from the image which reduces the required bitrate. All modules are implemented to extend the existing TELECARLA framework. We evaluate the proposed traffic-aware adaptation scheme in a user study. We observe a high correlation between the proposed view prioritization module and the subjective ratings obtained in the user study. The region-of-interest masking achieves Bjøntegaard Delta Rate savings of at least 19.8% compared to streaming the full camera view. The overall system improves the VMAF score by 1.86 per camera when considering the importance of the individual camera views as rated by the users. This demonstrates the potential of an individual adaptation for each camera view optimized for the current traffic situation. Markus Hofbauer, Christopher B. Kuhn, Mariem Khlifi, Goran Petrovic, Eckehard G. Steinbach |
VTC Spring | 1 |
| 2022 | Preprocessor Rate Control for Adaptive Multi-View Live Video Streaming Using a Single EncoderabstractCurrently, an increasing number of technical systems are equipped with multiple cameras. Limited by cost and size, they are often restricted to a single hardware encoder. The combination of all views into a single superframe allows for streaming all camera views at the same time, but it prevents individual rate/quality adaptations on those camera views. We propose a preprocessing filter concept that allows for individual rate/quality adaptation while using a single encoder. Additionally, we create a preprocessor model that estimates the required preprocessing filter parameters from the specified encoding parameters. This means our approach can be used with any existing multi-view adaptation scheme designed for controlling multiple encoders. We design both an analytical and a Machine Learning-based bitrate model. Because both models perform equally well, we suggest using either one as the core part of our preprocessor model. Both models are specifically designed for estimating the influence of the quantization parameter, frame rate, frame size, group of pictures length, and a Gaussian low-pass filter on the video bitrate. Furthermore, the rate models outperform state-of-the-art bitrate models by at least 22% regarding the overall root mean square error. Our bitrate models are the first of their kind to consider the influence of a Gaussian low-pass filter. We evaluate the preprocessing approach by streaming six camera views in a teledriving scenario with a single encoder and compare it to using six individual encoders. The experimental results demonstrate that the preprocessing approach achieves bitrates similar to the individual encoders for all views. While achieving a comparable rate and quality for the most important views, our approach requires a total bitrate that is 50% smaller than when using a single encoder approach without preprocessing. Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Introspective Failure Prediction for Autonomous Driving Using Late Fusion of State and Camera InformationabstractWe present an introspective failure prediction approach for autonomous vehicles. In autonomous driving, complex or unknown scenarios can cause a disengagement of the self-driving system. Disengagements can be triggered either by automatic safety measures or by human intervention. We propose to use recorded disengagement sequences from test drives as training data to learn to predict future failures. The system then learns introspectively from its own previous mistakes. In order to predict failures as early as possible, we propose a machine learning approach where sequences of sensor data are classified as either failure or success. The car itself is treated as a black box. Our method combines two sensor modalities that contain different types of information. An image-based model learns to detect generally challenging situations such as crowded intersections accurately multiple seconds in advance. A state data based model allows to detect fast changes immediately before a failure, such as sudden braking or swerving. The outcome of the individual models is fused by averaging the individual failure probabilities. We evaluate our approach on a data set provided by the BMW Group containing 14 hours of autonomous driving. The proposed late fusion approach allows for predicting failures at an accuracy of more than 85% seven seconds in advance, at a false positive rate of 20%. The proposed method outperforms state-of-the-art failure prediction by more than 15% while being a flexible framework that allows for straightforward addition of further sensor modalities. Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Pixel-Wise Failure Prediction For Semantic Video SegmentationabstractWe propose a pixel-accurate failure prediction approach for semantic video segmentation. The proposed scheme improves previously proposed failure prediction methods which so far disregarded the temporal information in videos. Our approach consists of two main steps: First, we train an LSTM-based model to detect spatio-temporal patterns that indicate pixel-wise misclassifications in the current video frame. Second, we use sequences of failure predictions to train a denoising autoencoder that both refines the current failure prediction and predicts future misclassifications. Since public data sets for this scenario are limited, we introduce the large-scale densely annotated video driving (DAVID) data set generated using the CARLA simulator. We evaluate our approach on the real-world Cityscapes data set and the simulator-based DAVID data set. Our experimental results show that spatiotemporal failure prediction outperforms single-image failure prediction by up to 8.8%. Refining the prediction using a sequence of previous failure predictions further improves the performance by a significant 15.2% and allows to accurately predict misclassifications for future frames. While we focus our study on driving videos, the proposed approach is general and can be easily used in other scenarios as well. Christopher B. Kuhn, Markus Hofbauer, Ziqin Xu, Goran Petrovic, Eckehard G. Steinbach |
ICIP | 2 |
| 2021 | Trajectory-Based Failure Prediction for Autonomous DrivingabstractIn autonomous driving, complex traffic scenarios can cause situations that require human supervision to resolve safely. Instead of only reacting to such events, it is desirable to predict them early in advance. While predicting the future is challenging, there is a source of information about the future readily available in autonomous driving: the planned trajectory the car intends to drive. In this paper, we propose to analyze the trajectories planned by the vehicle to predict failures early on. We consider sequences of trajectories and use machine learning to detect patterns that indicate impending failures. Since no public data of disengagements of autonomous vehicles is available, we use data provided by development vehicles of the BMW Group. From over six months of test drives, we obtain more than 2600 disengagements of the automated system. We train a Long Short-Term Memory classifier with sequences of planned trajectories that either resulted in successful driving or disengagements. The proposed approach outperforms existing state-of-the-art failure prediction with low-dimensional data by more than 3 % in a Receiver Operating Characteristic analysis. Since our approach makes no assumptions on the underlying system, it can be applied to predict failures in other safety-critical areas of robotics as well. Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach |
IV | 2 |
| 2020 | Measuring Driver Situation Awareness Using Region-of-Interest Prediction and Eye TrackingabstractWith increasing progress in autonomous driving, the human does not have to be in control of the vehicle for the entire drive. A human driver obtains the control of the vehicle in case of an autonomous system failure or when the vehicle encounters an unknown traffic situation it cannot handle on its own. A critical part of this transition to human control is to ensure a sufficient driver situation awareness. Currently, no direct method to explicitly estimate driver awareness exists. In this paper, we propose a novel system to explicitly measure the situation awareness of the driver. Our approach is inspired by methods used in aviation. However, in contrast to aviation, the situation awareness in driving is determined by the detection and understanding of dynamically changing and previously unknown situation elements. Our approach uses machine learning to define the best possible situation awareness. We also propose to measure the actual situation awareness of the driver using eye tracking. Comparing the actual awareness to the target awareness allows us to accurately assess the awareness the driver has of the current traffic situation. To test our approach, we conducted a user study. We measured the situation awareness score of our model for 8 unique traffic scenarios. The results experimentally validate the accuracy of the proposed driver awareness model. Markus Hofbauer, Christopher B. Kuhn, Lukas Püttner, Goran Petrovic, Eckehard G. Steinbach |
ISM | 1 |
| 2020 | Adaptive Multi-View Live Video Streaming for Teledriving Using a Single Hardware EncoderabstractTeleoperated driving (TOD) is a possible solution to cope with failures of autonomous vehicles. In TOD, the human operator perceives the traffic situation via video streams of multiple cameras from a remote location. Adaptation mechanisms are needed in order to match the available transmission resources and provide the operator with the best possible situation awareness. This includes the adjustment of individual camera video streams according to the current traffic situation. The limited video encoding hardware in vehicles requires the combination of individual camera frames into a larger superframe video. While this enables the encoding of multiple camera views with a single encoder, it does not allow for rate/quality adaptation of the individual views. To this end, we propose a novel concept that uses preprocessing filters to enable individual rate/quality adaptations in the superframe video. The proposed preprocessing filters allow for the usage of existing multidimensional adaptation models in the same way as for individual video streams using multiple encoders. Our experiments confirm that the proposed concept is able to control the spatial, temporal and quality resolution of individual segments in the superframe video. Additionally, we demonstrate the usability of the proposed method by applying it in a multi-view teledriving scenario. We compare our approach to individually encoded video streams and a multiplexing solution without preprocessing. The results show that the proposed approach produces bitrates for the individual video streams which are comparable to the bitrates achieved with separate encoders. While achieving a similar bitrate for the most important views, our approach requires a total bitrate that is 40% smaller compared to the multiplexing approach without preprocessing. Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach |
ISM | 1 |
| 2020 | Better Look Twice - Improving Visual Scene Perception Using a Two-Stage ApproachabstractAccurate visual scene perception plays an important role in fields such as medical imaging or autonomous driving. Recent advances in computer vision allow for accurate image classification, object detection and even pixel-wise semantic segmentation. Human vision has repeatedly been used as an inspiration for developing new machine vision approaches. In this work, we propose to adapt the “zoom lens model” from psychology for semantic scene segmentation. According to this model, humans first distribute their attention evenly across the entire field of view at low processing power. Then, they follow visual cues to look at a few smaller areas with increased attention. By looking twice, it is possible to refine the initial scene understanding without requiring additional input. We propose to perform semantic segmentation the same way. To obtain visual cues for deciding where to look twice, we use a failure region prediction approach based on a state-of-the-art failure prediction method. Then, the second, focused look is performed by a dedicated classifier that reclassifies the most challenging patches. Finally, pixels predicted to be errors are updated in the original semantic prediction. While focusing only on areas with the highest predicted failure probability, we achieve a classification accuracy of over 63% for the predicted failure regions. After updating the initial semantic prediction of 4000 test images from a large-scale driving data set, we reduce the absolute pixel-wise error of 232 road participants by 10% or more. Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach |
ISM | 2 |
| 2020 | TELECARLA: An Open Source Extension of the CARLA Simulator for Teleoperated Driving Research Using Off-the-Shelf ComponentsabstractTeledriving is a possible fallback mode to cope with failures of fully autonomous vehicles. One important requirement for teleoperated vehicles is a reliable low delay data transmission solution, which adapts to the current network conditions to provide the operator with the best possible situation awareness. Currently, there is no easily accessible solution for the evaluation of such systems and algorithms in a fully controllable environment available. To this end we propose an open source framework for teleoperated driving research using low-cost off-the-shelf components. The proposed system is an extension of the open source simulator CARLA, which is responsible for rendering the driving environment and providing reproducible scenario evaluation. As a proof of concept, we evaluated our teledriving solution against CARLA in remote and local driving scenarios. The proposed teledriving system leads to almost identical performance measurements for local and remote driving. In contrast, remote driving using CARLA's client server communication results in drastically reduced operator performance. Further, the framework provides an interface for the adaptation of the temporal resolution and target bitrate of the compressed video streams. The proposed framework reduces the required setup effort for teleoperated driving research in academia and industry. Markus Hofbauer, Christopher B. Kuhn, Goran Petrovic, Eckehard G. Steinbach |
IV | 1 |
| 2020 | Introspective Black Box Failure Prediction for Autonomous DrivingabstractFailures in autonomous driving caused by complex traffic situations or model inaccuracies remain inevitable in the near future. While much research is focused on how to prevent such failures, comparatively little research has been done on predicting them. An early failure prediction would allow for more time to take actions to resolve challenging situations. In this work, we propose an introspective approach to predict future disengagements of the car by learning from previous disengagement sequences. Our method is designed to detect failures as early as possible by using sensor data from up to ten seconds before each disengagement. The car itself is treated as a black box, with only its state data and the number of detected objects being required. Since no model-specific knowledge is needed, our method is applicable to any self-driving system. Currently, no public data of real-life disengagements is available. To test our approach, we therefore use autonomous driving data provided by BMW that was collected with BMW research vehicles over multiple months. We show that an LSTM classifier trained with sequences of state data can predict failures up to seven seconds in advance with an accuracy of more than 80%. This is two seconds earlier than comparable approaches from the literature. Christopher B. Kuhn, Markus Hofbauer, Goran Petrovic, Eckehard G. Steinbach |
IV | 2 |
| 2007 | A Robust, Responsive, Distributed Tree-Based Routing Algorithm Guaranteeing N Valid Links per Node in Wireless Ad-Hoc NetworksabstractIn this paper, we present two new tree growing algorithms. The pairing algorithm allows for the local approximate implementation of global algorithms such as Prim's or Dijkstra 's algorithm. A node requires only information from its neighborhood acquired by local message exchange. Any global cost function that can be locally calculated can be used with this algorithm. The N-SafeLinks Algorithm establishes look-up tables with N possible links per node. Implementing an additional constraint it is guaranteed that each link leads to the sink, ruling out the possibility of loops. Therefore, if maximal N-l links per node are broken there is still a guaranteed connection to the destination node (sink) for every node. A proof of this property is presented. As demonstrated in simulations the trees are close to optimum. The algorithms can be utilized for static routing in wireless single-sink Ad-Hoc networks with safety critical applications where timeliness, robustness and energy efficiency is crucial. Tunc Ikikardes, Markus Hofbauer, August Kaelin, Martin May |
ISCC | 2 |