Johann Marius Zöllner

dblp:97/3452 · also J. Marius Zoellner, J. Marius Zöllner, Marius Zöllner 0001 · DBLP profile ↗
← Back
121ranked-venue papers
0as first author
60since 2021 · last 2026
0000-0001-6190-7202ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 109 · 55 since 2021Systems, architecture and hardware · 20 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Efficient Cross-Country Data Acquisition Strategy for ADAS via Street-View Imagery
Daniel Slieter, Carl Esselborn, Ahmed Abouelazm, Tsung Yuan Tseng, Johann Marius Zöllner
IV6
2026 DenseBEV: Transforming BEV Grid Cells into 3D Objects
abstract
In current research, Bird’s-Eye-View (BEV)-based transformers are increasingly utilized for multi-camera 3D object detection. Traditional models often employ random queries as anchors, optimizing them successively. Recent advancements complement or replace these random queries with detections from auxiliary networks. We propose a more intuitive and efficient approach by using BEV feature cells directly as anchors. This end-to-end approach leverages the dense grid of BEV queries, considering each cell as a potential object for the final detection task. As a result, we introduce a novel two-stage anchor generation method specifically designed for multi-camera 3D object detection. To address the scaling issues of attention with a large number of queries, we apply BEV-based Non-Maximum Suppression, allowing gradients to flow only through non-suppressed objects. This ensures efficient training without the need for post-processing. By using BEV features from encoders such as BEVFormer directly as object queries, temporal BEV information is inherently embedded. Building on the temporal BEV information already embedded in our object queries, we introduce a hybrid temporal modeling approach by integrating prior detections to further enhance detection performance. Evaluating our method on the nuScenes dataset shows consistent and significant improvements in NDS and mAP over the baseline, even with sparser BEV grids and therefore fewer initial anchors. It is particularly effective for small objects, enhancing pedestrian detection with a 3.8% mAP increase on nuScenes and an 8% increase in LET-mAP on Waymo. Applying our method, named DenseBEV, to the challenging Waymo Open dataset yields state-of-the-art performance, achieving a LET-mAP of 60.7%, surpassing the previous best by 5.4%. Code is available at https://github.com/mdaehl/DenseBEV.
Marius Dähling, Sebastian Krebs, Johann Marius Zöllner
WACV3
2025 Diverse and Adaptive Behavior Curriculum for Autonomous Driving: A Student-Teacher Framework with Multi-Agent RL
abstract
Autonomous driving faces challenges in navigating complex real-world traffic, requiring safe handling of both common and critical scenarios. Reinforcement learning (RL), a prominent method in end-to-end driving, enables agents to learn through trial and error in simulation. However, RL training often relies on rule-based traffic scenarios, limiting generalization. Additionally, current scenario generation methods focus heavily on critical scenarios, neglecting a balance with routine driving behaviors. Curriculum learning, which progressively trains agents on increasingly complex tasks, is a promising approach to improving the robustness and coverage of RL driving policies. However, existing research mainly emphasizes manually designed curricula, focusing on scenery and actor placement rather than traffic behavior dynamics. This work introduces a novel student-teacher framework for automatic curriculum learning. The teacher, a graph-based multi-agent RL component, adaptively generates traffic behaviors across diverse difficulty levels. An adaptive mechanism adjusts task difficulty based on student performance, ensuring exposure to behaviors ranging from common to critical. The student, though exchangeable, is realized as a deep RL agent with partial observability, reflecting real-world perception constraints. Results demonstrate the teacher’s ability to generate diverse traffic behaviors. The student, trained with automatic curricula, outperformed agents trained on rule-based traffic, achieving higher rewards and exhibiting balanced, assertive driving.
Ahmed Abouelazm, Johannes Ratz, Philip Schörner, Johann Marius Zöllner
IROS4
2025 CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded Scenes
abstract
This paper introduces a novel method for end-to-end crowd detection that leverages object density information to enhance existing transformer-based detectors. We present CrowdQuery (CQ), whose core component is our CQ module that predicts and subsequently embeds an object density map. The embedded density information is then systematically integrated into the decoder. Existing density map definitions typically depend on head positions or object-based spatial statistics. Our method extends these definitions to include individual bounding box dimensions. By incorporating density information into object queries, our method utilizes density-guided queries to improve detection in crowded scenes. CQ is universally applicable to both 2D and 3D detection without requiring additional data. Consequently, we are the first to design a method that effectively bridges 2D and 3D detection in crowded environments. We demonstrate the integration of CQ into both a general 2D and 3D transformer-based object detector, introducing the architectures CQ2D and CQ3D. CQ is not limited to the specific transformer models we selected. Experiments on the STCrowd dataset for both 2D and 3D domains show significant performance improvements compared to the base models, outperforming most state-of-the-art methods. When integrated into a state-of-the-art crowd detector, CQ can further improve performance on the challenging CrowdHuman dataset, demonstrating its generalizability. The code is released at https://github.com/mdaehl/CrowdQuery.
Marius Dähling, Sebastian Krebs, Johann Marius Zöllner
IROS3
2025 Boundary-Guided Trajectory Prediction for Road Aware and Physically Feasible Autonomous Driving
abstract
Accurate prediction of surrounding road users' trajectories is essential for safe and efficient autonomous driving. While deep learning models have improved performance, challenges remain in preventing off-road predictions and ensuring kinematic feasibility. Existing methods incorporate road-awareness modules and enforce kinematic constraints but lack plausibility guarantees and often introduce trade-offs in complexity and flexibility. This paper proposes a novel framework that formulates trajectory prediction as a constrained regression guided by permissible driving directions and their boundaries. Using the agent's current state and an HD map, our approach defines the valid boundaries and ensures on-road predictions by training the network to learn superimposed paths between left and right boundary polylines. To guarantee feasibility, the model predicts acceleration profiles that determine the vehicle's travel distance along these paths while adhering to kinematic constraints. We evaluate our approach on the Argoverse-2 dataset against the HPTR baseline. Our approach shows a slight decrease in benchmark metrics compared to HPTR but notably improves final displacement error and eliminates infeasible trajectories. Moreover, the proposed approach has a superior generalization to less prevalent maneuvers and unseen out-of-distribution scenarios, reducing the off-road rate under adversarial attacks from 66 % to just 1 %. These results highlight the effectiveness of our approach in generating feasible and robust predictions.
Ahmed Abouelazm, Mianzhi Liu, Christian Hubschneider, Daniel Slieter, Johann Marius Zöllner
IV6
2025 Balancing Progress and Safety: A Novel Risk-Aware Objective for RL in Autonomous Driving
abstract
Reinforcement Learning (RL) is a promising approach for achieving autonomous driving due to robust decision-making capabilities. RL learns a driving policy through trial and error in traffic scenarios, guided by a reward function that combines the driving objectives. The design of such reward function has received insufficient attention, yielding ill-defined rewards with various pitfalls. Safety, in particular, has long been regarded only as a penalty for collisions. This leaves the risks associated with actions leading up to a collision unaddressed, limiting the applicability of RL in real-world scenarios. To address these shortcomings, our work focuses on enhancing the reward formulation by defining a set of driving objectives and structuring them hierarchically. Furthermore, we discuss the formulation of these objectives in a normalized manner to transparently determine their contribution to the overall reward. Additionally, we introduce a novel risk-aware objective for various driving interactions based on a two-dimensional ellipsoid function and an extension of Responsibility-Sensitive Safety (RSS) concepts. We evaluate the efficacy of our proposed reward in unsignalized intersection scenarios with varying traffic densities. The approach decreases collision rates by 21% on average compared to baseline rewards and consistently surpasses them in route progress and cumulative reward, demonstrating its capability to promote safer driving behaviors while maintaining high-performance levels.
Ahmed Abouelazm, Jonas Michel, Helen Gremmelmaier, Tim Joseph, Philip Schörner, Johann Marius Zöllner
IV6
2025 Automatic Curriculum Learning for Driving Scenarios: Towards Robust and Efficient Reinforcement Learning
abstract
This paper addresses the challenges of training end-to-end autonomous driving agents using Reinforcement Learning (RL). RL agents are typically trained in a fixed set of scenarios and nominal behavior of surrounding road users in simulations, limiting their generalization and real-life deployment. While Domain Randomization offers a potential solution by randomly sampling driving scenarios, it frequently results in inefficient training and sub-optimal policies due to the high variance among training scenarios. To address these limitations, we propose an automatic curriculum learning framework that dynamically generates driving scenarios with adaptive complexity based on the agent's evolving capabilities. Unlike manually designed curricula that introduce expert bias and lack scalability, our framework incorporates a “teacher” that automatically generates and mutates driving scenarios based on their learning potential-an agent-centric metric derived from the agent's current policy, eliminating the need for expert design. The framework enhances training efficiency by excluding scenarios the agent has mastered or finds too challenging. We evaluate our framework in a reinforcement learning setting where the agent learns a driving policy from camera images. Comparative results against baseline methods, including fixed scenario training and domain randomization, demonstrate that our approach leads to enhanced generalization, achieving higher success rates, +9 % in low traffic density, +21 % in high traffic density, and faster convergence with fewer training steps. Our findings highlight the potential of ACL in improving the robustness and efficiency of RL-based autonomous driving agents.
Ahmed Abouelazm, Tim Weinstein, Tim Joseph, Philip Schörner, Johann Marius Zöllner
IV5
2025 TPK: Trustworthy Trajectory Prediction Integrating Prior Knowledge for Interpretability and Kinematic Feasibility
abstract
Trajectory prediction is crucial for autonomous driving, enabling vehicles to navigate safely by anticipating the movements of surrounding road users. However, current deep learning models often lack trustworthiness as their predictions can be physically infeasible and illogical to humans. To make predictions more trustworthy, recent research has incorporated prior knowledge, like the social force model for modeling interactions and kinematic models for physical realism. However, these approaches focus on priors that suit either vehicles or pedestrians and do not generalize to traffic with mixed agent classes. We propose incorporating interaction and kinematic priors of all agent classes-vehicles, pedestrians, and cyclists with class-specific interaction layers to capture agent behavioral differences. To improve the interpretability of the agent interactions, we introduce DG-SFM, a rule-based interaction importance score that guides the interaction layer. To ensure physically feasible predictions, we proposed suitable kinematic models for all agent classes with a novel pedestrian kinematic model. We benchmark our approach on the Argoverse 2 dataset, using the state-of-the-art transformer HPTR as our baseline. Experiments demonstrate that our method improves interaction interpretability, revealing a correlation between incorrect predictions and divergence from our interaction prior. Even though incorporating the kinematic models causes a slight decrease in accuracy, they eliminate infeasible trajectories found in the dataset and the baseline model. Thus, our approach fosters trust in trajectory prediction as its interaction reasoning is interpretable, and its predictions adhere to physics.
Marius Baden, Ahmed Abouelazm, Christian Hubschneider, Daniel Slieter, Johann Marius Zöllner
IV6
2025 Label-Free Model Failure Detection for Lidar-based Point Cloud Segmentation
abstract
Autonomous vehicles drive millions of miles on the road each year. Under such circumstances, deployed machine learning models are prone to failure both in seemingly normal situations and in the presence of outliers. However, in the training phase, they are only evaluated on small validation and test sets, which are unable to reveal model failures due to their limited scenario coverage. While it is difficult and expensive to acquire large and representative labeled datasets for evaluation, large-scale unlabeled datasets are typically available. In this work, we introduce label-free model failure detection for lidar-based point cloud segmentation, taking advantage of the abundance of unlabeled data available. We leverage different data characteristics by training a supervised and self-supervised stream for the same task to detect failure modes. We perform a large-scale qualitative analysis and present LidarCODA, the first publicly available dataset with labeled anomalies in real-world lidar data, for an extensive quantitative analysis.
Daniel Bogdoll, Finn Sartoris, Vincent Geppert, Svetlana Pavlitska, Johann Marius Zöllner
IV5
2025 MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations
abstract
World models for autonomous driving have the potential to dramatically improve the reasoning capabilities of today's systems. However, most works focus on camera data, with only a few that leverage lidar data or combine both to better represent autonomous vehicle sensor setups. In addition, raw sensor predictions are less actionable than 3D occupancy predictions, but there are no works examining the effects of combining both multimodal sensor data and 3D occupancy prediction. In this work, we perform a set of experiments with a MUltimodal World Model with Geometric VOxel represen-tations (MUVO) to evaluate different sensor fusion strategies to better understand the effects on sensor data prediction. We also analyze potential weaknesses of current sensor fusion approaches and examine the benefits of additionally predicting 3D occupancy.
Daniel Bogdoll, Yitian Yang, Tim Joseph, Melih Yazgan, Johann Marius Zöllner
IV5
2025 Towards Intelligent Control Centers: Case-Based Reasoning for Waypoint Assistance
abstract
Autonomous-driving research has traditionally focused on refining on-board intelligence for individual vehicles. Recent studies highlight the potential of intelligent control centers-equipped with machine learning capabilities-to complement on-board systems by sharing insights across an entire fleet and assisting with corner cases in real time. In this paper, we advance the concept of intelligent control centers through a case-based reasoning approach for remote waypoint assistance. Our system captures and reuses operator interventions, converting human-generated solutions for unforeseen obstacles into transferable “cases”. The remote operator remains in the loop to validate every suggested waypoint, ensuring safety before execution. Preliminary field tests with two autonomous shuttles demonstrate the feasibility of retrieving previous waypoint interventions under realistic conditions. Our initial small-scale results indicate that (i) the prototype can achieve performance comparable to fully manual interventions and (ii) solutions devised for one vehicle can effectively be transferred to another. Together, these outcomes lay a solid foundation for more sophisticated, learning-based control-center architectures.
Martin Gontscharow, Stefan Orf, Albert Schotschneider, Tobias Fleck, Johann Marius Zöllner
IV5
2025 A Chef's KISS - Utilizing Semantic Information in Both ICP and SLAM Framework
abstract
For utilizing autonomous vehicle in urban areas a reliable localization is needed. Especially when HD maps are used, a precise and repeatable method has to be chosen. Therefore accurate map generation but also re-localization against these maps is necessary. Due to best 3D reconstruction of the surrounding, LiDAR has become a reliable modality for localization. The latest LiDAR odometry estimation are based on iterative closest point (ICP) approaches, namely KISS-ICP [1] and SAGE-ICP [2]. We extend the capabilities of KISS-ICP by incorporating semantic information into the point alignment process using a generalizable approach with minimal parameter tuning. This enhancement allows us to surpass KISS-ICP in terms of absolute trajectory error (ATE), the primary metric for map accuracy. Additionally, we improve the Cartographer mapping framework to handle semantic information. Cartographer facilitates loop closure detection over larger areas, mitigating odometry drift and further enhancing ATE accuracy. By integrating semantic information into the mapping process, we enable the filtering of specific classes, such as parked vehicles, from the resulting map. This filtering improves relocalization quality by addressing temporal changes, such as vehicles being moved.
Sven Ochs, Marc Heinrich, Philip Schömer, Marc Rene Zofka, Johann Marius Zöllner
IV5
2025 Fool the Stoplight: Realistic Adversarial Patch Attacks on Traffic Light Detectors
abstract
Realistic adversarial attacks on various camera-based perception tasks of autonomous vehicles have been successfully demonstrated so far. However, only a few works considered attacks on traffic light detectors. This work shows how CNNs for traffic light detection can be attacked with printed patches. We propose a threat model, where each instance of a traffic light is attacked with a patch placed under it, and describe a training strategy. We demonstrate successful adversarial patch attacks in universal settings. Our experiments show realistic targeted red - to-green label-flipping attacks and attacks on pictogram classification. Finally, we perform a real-world evaluation with printed patches and demonstrate attacks in the lab settings with a mobile traffic light for construction sites and in a test area with stationary traffic lights. Our code is available at https://github.com/KASTEL-MobilityLab/attacks-on-traffic-light-detection.
Svetlana Pavlitska, Jamie Robb, Nikolai Polley, Melih Yazgan, Johann Marius Zöllner
IV5
2025 Fool the Stoplight: Realistic Adversarial Patch Attacks on Traffic Light Detectors
abstract
Realistic adversarial attacks on various camera-based perception tasks of autonomous vehicles have been suc-cessfully demonstrated so far. However, only a few works con-sidered attacks on traffic light detectors. This work shows how CNNs for traffic light detection can be attacked with printed patches. We propose a threat model, where each instance of a traffic light is attacked with a patch placed under it, and describe a training strategy. We demonstrate successful adver-sarial patch attacks in universal settings. Our experiments show realistic targeted red-to-green label-flipping attacks and attacks on pictogram classification. Finally, we perform a real-world evaluation with printed patches and demonstrate attacks in the lab settings with a mobile traffic light for construction sites and in a test area with stationary traffic lights. Our code is avail-able at https://github.com/KASTEL-MobilityLab/attacks-on-traffic-light-detection.
Svetlana Pavlitska, Jamie Robb, Nikolai Polley, Melih Yazgan, Johann Marius Zöllner
IV5
2025 Self-Supervised Pretraining for Aerial Road Extraction
abstract
Deep neural networks for aerial image segmentation require large amounts of labeled data, but high-quality aerial datasets with precise annotations are scarce and costly to produce. To address this limitation, we propose a self-supervised pretraining method that improves segmentation performance while reducing reliance on labeled data. Our approach uses inpainting-based pretraining, where the model learns to reconstruct missing regions in aerial images, capturing their inherent structure before being fine-tuned for road extraction. This method improves generalization, enhances robustness to domain shifts, and is invariant to model architecture and dataset choice. Experiments show that our pretraining significantly boosts segmentation accuracy, especially in low-data regimes, making it a scalable solution for aerial image analysis.
Rupert Polley, Sai Vignesh Abishek Deenadayalan, Johann Marius Zöllner
IV3
2025 The ATLAS of Traffic Lights: A Reliable Perception Framework for Autonomous Driving
abstract
Traffic light perception is an essential component of the camera-based perception system for autonomous vehicles, enabling accurate detection and interpretation of traffic lights to ensure safe navigation through complex urban environments. In this work, we propose a modularized perception framework that integrates state-of-the-art detection models with a novel real-time association and decision framework, enabling seamless deployment into an autonomous driving stack. To address the limitations of existing public datasets, we introduce the ATLAS dataset, which provides comprehensive annotations of traffic light states and pictograms across diverse environmental conditions and camera setups. This dataset is publicly available at https://url.fzi.de/ATLAS. We train and evaluate several state-of-the-art traffic light detection architectures on ATLAS, demonstrating significant performance improvements in both accuracy and robustness. Finally, we evaluate the framework in real-world scenarios by deploying it in an autonomous vehicle to make decisions at traffic light-controlled intersections, highlighting its reliability and effectiveness for real-time operation.
Rupert Polley, Nikolai Polley, Dominik Heid, Marc Heinrich, Sven Ochs, Johann Marius Zöllner
IV6
2025 Efficient Data Representation for Motion Forecasting: A Scene-Specific Trajectory Set Approach
abstract
Representing diverse and plausible future trajectories is critical for motion forecasting in autonomous driving. However, efficiently capturing these trajectories in a compact set remains challenging. This study introduces a novel approach for generating scene-specific trajectory sets tailored to different contexts, such as intersections and straight roads, by leveraging map information and actor dynamics. A deterministic goal sampling algorithm identifies relevant map regions, while our Recursive In-Distribution Subsampling (RIDS) method enhances trajectory plausibility by condensing redundant representations. Experiments on the Argoverse 2 dataset demonstrate that our method achieves up to a 45% improvement in Driving Area Compliance (DAC) compared to baseline methods while maintaining competitive displacement errors. Our work highlights the benefits of mining such scene-aware trajectory sets and how they could capture the complex and heterogeneous nature of actor behavior in real-world driving scenarios.
Abhishek Vivekanandan, Johann Marius Zöllner
IV2
2025 Centralized Decision-Making for Platooning By Using SPaT-Driven Reference Speeds
abstract
This paper introduces a centralized approach for fuel-efficient urban platooning by leveraging real-time Vehicle-to-Everything (V2X) communication and Signal Phase and Timing (SPaT) data. A nonlinear Model Predictive Control (MPC) algorithm optimizes the trajectories of platoon leader vehicles, employing an asymmetric cost function to minimize fuel-intensive acceleration. Following vehicles utilize a gap-and velocity-based control strategy, complemented by dynamic platoon splitting logic communicated through Platoon Control Messages (PCM) and Platoon Awareness Messages (PAM). Simulation results obtained from the CARLA environment demonstrate substantial fuel savings of up to 41.2 %, along with smoother traffic flows, fewer vehicle stops, and improved intersection throughput.
Melih Yazgan, Süleyman Tatar, Johann Marius Zöllner
IV3
2025 Runtime Safety Monitoring of Deep Neural Networks for Perception: A Survey
abstract
Deep neural networks (DNNs) are widely used in perception systems for safety-critical applications, such as autonomous driving and robotics. However, DNNs remain vulnerable to various safety concerns, including generalization errors, out-of-distribution (OOD) inputs, and adversarial attacks, which can lead to hazardous failures. This survey provides a comprehensive overview of runtime safety monitoring approaches, which operate in parallel to DNNs during inference to detect these safety concerns without modifying the DNN itself. We categorize existing methods into three main groups: Monitoring inputs, internal representations, and outputs. We analyze the state-of-the-art for each category, identify strengths and limitations, and map methods to the safety concerns they address. In addition, we highlight open challenges and future research directions.
Albert Schotschneider, Svetlana Pavlitska, Johann Marius Zöllner
SMC3
2024 Towards Adversarial Robustness of Model-Level Mixture-of-Experts Architectures for Semantic Segmentation
abstract
Vulnerability to adversarial attacks is a well-known deficiency of deep neural networks. Larger networks are generally more robust, and ensembling is one method to increase adversarial robustness: each model's weaknesses are compensated by the strengths of others. While an ensemble uses a deterministic rule to combine model outputs, a mixture of experts (MoE) includes an additional learnable gating component that predicts weights for the outputs of the expert models, thus determining their contributions to the final prediction. MoEs have been shown to outperform ensembles on specific tasks, yet their susceptibility to adversarial attacks has not been studied yet. In this work, we evaluate the adversarial vulnerability of MoEs for semantic segmentation of urban and highway traffic scenes. We show that MoEs are, in most cases, more robust to per-instance and universal white-box adversarial attacks and can better withstand transfer attacks. Our code is available at https://2ithub.com/KASTEL-MobilitvLab/mixtures-of-exuerts/.
Svetlana Pavlitska, Enrico Eisen, Johann Marius Zöllner
ICMLA3
2024 Evaluating Adversarial Attacks on Traffic Sign Classifiers Beyond Standard Baselines
abstract
Adversarial attacks on traffic sign classification models were among the first successfully tried in the real world. Since then, the research in this area has been mainly restricted to repeating baseline models, such as LISA-CNN or GTSRB-CNN, and similar experiment settings, including white and black patches on traffic signs. In this work, we decouple model architectures from the datasets and evaluate on further generic models to make a fair comparison. Furthermore, we compare two attack settings, inconspicuous and visible, which are usually regarded without direct comparison. Our results show that standard baselines like LISA-CNN or GTSRB-CNN are significantly more susceptible than the generic ones. We, therefore, suggest evaluating new attacks on a broader spectrum of baselines in the future. Our code is available at https://github.com/KASTEL-MobilityLab/attacks-on-traffic-sign-recognition/.
Svetlana Pavlitska, Leopold Müller, Johann Marius Zöllner
ICMLA3
2024 Informed Reinforcement Learning for Situation-Aware Traffic Rule Exceptions
abstract
Reinforcement Learning is a highly active research field with promising advancements. In the field of autonomous driving, however, often very simple scenarios are being examined. Common approaches use non-interpretable control commands as the action space and unstructured reward designs, which are unsuitable for complex scenarios. In this work, we introduce Informed Reinforcement Learning, where a structured rulebook is integrated as a knowledge source. We learn trajectories and asses them with a situation-aware reward design, leading to a dynamic reward that allows the agent to learn situations that require controlled traffic rule exceptions. Our method is applicable to arbitrary RL models. We successfully demonstrate high completion rates of complex scenarios with recent model-based agents.
Daniel Bogdoll, Moritz Nekolla, Ahmed Abouelazm, Tim Joseph, Johann Marius Zöllner
ICRA6
2024 Iterative Filter Pruning for Concatenation-based CNN Architectures
abstract
Model compression and hardware acceleration are essential for the resource-efficient deployment of deep neural networks. Modern object detectors have highly interconnected convolutional layers with concatenations. In this work, we study how pruning can be applied to such architectures, exemplary for YOLOv7. We propose a method to handle concatenation layers, based on the connectivity graph of convolutional layers. By automating iterative sensitivity analysis, pruning, and subsequent model fine-tuning, we can significantly reduce model size both in terms of the number of parameters and FLOPs, while keeping comparable model accuracy. Finally, we deploy pruned models to FPGA and NVIDIA Jetson Xavier AGX. Pruned models demonstrate a 2x speedup for the convolutional layers in comparison to the unpruned counterparts and reach real-time capability with 14 FPS on FPGA. Our code is available at https://github.com/fzi-forschungszentrum-informatik/iterative-yolo-pruning.
Svetlana Pavlitska, Oliver Bagge, Federico Nicolás Peccia, Toghrul Mammadov, Johann Marius Zöllner
IJCNN5
2024 A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
abstract
Reinforcement learning has emerged as an important approach for autonomous driving. A reward function is used in reinforcement learning to establish the learned skill objectives and guide the agent toward the optimal policy. Since autonomous driving is a complex domain with partly conflicting objectives with varying degrees of priority, developing a suitable reward function represents a fundamental challenge. This paper aims to highlight the gap in such function design by assessing different proposed formulations in the literature and dividing individual objectives into Safety, Comfort, Progress, and Traffic Rules compliance categories. Additionally, the limitations of the reviewed reward functions are discussed, such as objectives aggregation and indifference to driving context. Furthermore, the reward categories are frequently inadequately formulated and lack standardization. This paper concludes by proposing future research that potentially addresses the observed shortcomings in rewards, including a reward validation framework and structured rewards that are context-aware and able to resolve conflicts.
Ahmed Abouelazm, Jonas Michel, Johann Marius Zöllner
IV3
2024 Scalable Remote Operation for Autonomous Vehicles: Integration of Cooperative Perception and Open Source Communication
abstract
As autonomous vehicles become increasingly preva-lent, robust remote operation systems are imperative to ensure safety and reliability in unpredictable scenarios. Current remote operation systems in research often lack scalability and adaptability, hindering their integration into diverse autonomous driving platforms. This paper addresses these challenges by introducing a scalable remote operation system that leverages cooperative perception and an open-source communication module. Field tests conducted with an SAE Level 3 autonomous shuttle have validated the effectiveness of our system in real-world scenarios. The code for a key component of this system, the communication module, is available online1.
Martin Gontscharow, Jens Doll, Albert Schotschneider, Daniel Bogdoll, Stefan Orf, Johannes Jestram, Marc Rene Zofka, Johann Marius Zöllner
IV8
2024 Last Mile Delivery with Autonomous Shuttles: ROS-based Integration of Smart Cargo Cages
abstract
With consistently increasing amounts of transported goods, autonomous cargo transport has gained increasing interest as a potential solution. In addition to reliable autonomous driving functions, Autonomous cargo transport requires a wide range of additional software and hardware components to ensure a safe and efficient transport of cargo as well as a pleasant user experience for the customers. This work presents a general concept for an autonomous and flexible cargo transport system, targeting point-to-point transports in the range of the typical last mile. The proposed concept provides flexibility for demand-responsive passenger transport as a mixed cargo-passenger transport solution. Furthermore, the proposed concept is realized through designing a removable cargo hold with electronic locks, and software modules such as a Booking App, a Scanner App, and a central backend. The implementation was developed, deployed in an autonomous shuttle, and extensively tested in a peri-urban quarter of the Test Area Autonomous Driving Baden-Württemberg.
Sven Ochs, Nico Lambing, Ahmed Abouelazm, Marc Rene Zofka, Johann Marius Zöllner
IV5
2024 Can You See Me Now? Blind Spot Estimation for Autonomous Vehicles using Scenario-Based Simulation with Random Reference Sensors
abstract
In this paper, we introduce a method for estimating blind spots for sensor setups of autonomous or automated vehicles and/or robotics applications. In comparison to previous methods that rely on geometric approximations, our presented approach provides more realistic coverage estimates by utilizing accurate and detailed 3D simulation environments. Our method leverages point clouds from LiDAR sensors or camera depth images from high-fidelity simulations of target scenarios to provide accurate and actionable visibility estimates. A Monte Carlo-based reference sensor simulation enables us to accurately estimate blind spot size as a metric of coverage, as well as detection probabilities of objects at arbitrary positions.
Marc Uecker, Johann Marius Zöllner
IV2
2024 KI-PMF: Knowledge Integrated Plausible Motion Forecasting
abstract
The accurate prediction of surrounding traffic actors’ movements is vital for the large-scale safe deployment of autonomous vehicles. Existing motion forecasting methods primarily aim to minimize prediction error by optimizing a loss function, which can sometimes lead to physically infeasible predictions or states that violate external constraints. This paper proposes a method that integrates explicit knowledge priors, allowing a network to forecast future trajectories that comply with both the vehicle’s kinematic constraints and the driving environment’s geometry. This is achieved by introducing a non-parametric pruning layer, and learnable attention layers to incorporate the defined knowledge priors. The proposed method aims to ensure reachability guarantees for traffic actors in both complex and dynamic situations. By conditioning the network to adhere to physical laws, we can achieve accurate and safe predictions, which are crucial for maintaining the safety and efficiency of autonomous vehicles in real-world settings.
Abhishek Vivekanandan, Ahmed Abouelazm, Philip Schörner, Johann Marius Zöllner
IV4
2024 Shuttle2X -Overcoming Operational Borders of Autonomous Shuttles by Infrastructure Support
abstract
Automated shuttles are currently limited to very simple and clear scenarios, e.g. with dedicated lanes, low traffic density and little flexibility in interaction with other road users. Furthermore, the vehicles are always monitored by a safety driver and the systems are highly adapted to a certain route and only suitable for one specific application. This whitepaper presents the general test site architecture, algorithmic challenges and a general overview of the German research project "Shuttle2X". It has the aim to go significantly beyond these dedicated solutions currently found in very isolated environments. This is achieved by expanding the operating area through research and development of new technologies, in particular through a connection to an intelligent infrastructure and selective expansion of automated driving (AD) capabilities as well as a highly reliable and functional safe communication with the infrastructure at dedicated and essential route points. For this purpose, functional algorithms are also being further developed using artificial intelligence (AI) processes by considering fault tolerance in order to enable robust self-driving operation and thus ultimately contribute to a replacement of the safety driver.
Melih Yazgan, Jennifer Amritzer, Tobias Fleck, Marc Rene Zofka, Johann Marius Zöllner, Florian Alexander Schiegg, Keno Garlichs, Mihai Kocsis, Johannes Buyer, Raoul Daniel Zöllner
IV5
2024 Collaborative Perception Datasets in Autonomous Driving: A Survey
abstract
This survey offers a comprehensive examination of collaborative perception datasets in the context of Vehicle-to-Infrastructure (V2I), Vehicle-to-Vehicle (V2V), and Vehicle-to-Everything (V2X). It highlights the latest developments in large-scale benchmarks that accelerate advancements in perception tasks for autonomous vehicles. The paper systematically analyzes a variety of datasets, comparing them based on aspects such as diversity, sensor setup, quality, public availability, and their applicability to downstream tasks. It also highlights the key challenges such as domain shift, sensor setup limitations, and gaps in dataset diversity and availability. The importance of addressing privacy and security concerns in the development of datasets is emphasized, regarding data sharing and dataset creation. The conclusion underscores the necessity for comprehensive, globally accessible datasets and collaborative efforts from both technological and research communities to overcome these challenges and fully harness the potential of autonomous driving.
Melih Yazgan, Mythra Varun Akkanapragada, Johann Marius Zöllner
IV3
2024 A Survey on Intermediate Fusion Methods for Collaborative Perception Categorized by Real World Challenges
abstract
This survey analyzes intermediate fusion methods in collaborative perception for autonomous driving, categorized by real-world challenges. We examine various methods, detailing their features and the evaluation metrics they employ. The focus is on addressing challenges like transmission efficiency, localization errors, communication disruptions, and heterogeneity. Moreover, we explore strategies to counter adversarial attacks and defenses, as well as approaches to adapt to domain shifts. The objective is to present an overview of how intermediate fusion methods effectively meet these diverse challenges, highlighting their role in advancing the field of collaborative perception in autonomous driving.
Melih Yazgan, Thomas Graf, Tobias Fleck, Johann Marius Zöllner
IV5
2024 Scene-Extrapolation: Generating Interactive Traffic Scenarios
abstract
Verifying highly automated driving functions can be challenging, requiring identifying relevant test scenarios. Scenario-based testing will likely play a significant role in verifying these systems, predominantly occurring within simulation. In our approach, we use traffic scenes as a starting point (seedscene) to address the individuality of various highly automated driving functions and to avoid the problems associated with a predefined test traffic scenario. Different highly autonomous driving functions, or their distinct iterations, may display different behaviors under the same operating conditions. To make a generalizable statement about a seed-scene, we simulate possible outcomes based on various behavior profiles. We utilize our lightweight simulation environment and populate it with rule-based and machine learning behavior models for individual actors in the scenario. We analyze resulting scenarios using a variety of criticality metrics. The density distributions of the resulting criticality values enable us to make a profound statement about the significance of a particular scene, considering various eventualities.
Maximilian Zipfl, Barbara Schütt, Johann Marius Zöllner
IV3
2023 Unscented Autoencoder
abstract
The Variational Autoencoder (VAE) is a seminal approach in deep generative modeling with latent variables. Interpreting its reconstruction process as a nonlinear transformation of samples from the latent posterior distribution, we apply the Unscented Transform (UT) – a well-known distribution approximation used in the Unscented Kalman Filter (UKF) from the field of filtering. A finite set of statistics called sigma points, sampled deterministically, provides a more informative and lower-variance posterior representation than the ubiquitous noise-scaling of the reparameterization trick, while ensuring higher-quality reconstruction. We further boost the performance by replacing the Kullback-Leibler (KL) divergence with the Wasserstein distribution metric that allows for a sharper posterior. Inspired by the two components, we derive a novel, deterministic-sampling flavor of the VAE, the Unscented Autoencoder (UAE), trained purely with regularization-like terms on the per-sample posterior. We empirically show competitive performance in Fréchet Inception Distance scores over closely-related models, in addition to a lower training variance than the VAE.
Faris Janjos, Lars Rosenbaum, Maxim Dolgov, Johann Marius Zöllner
ICML4
2023 Robotic Control Using Model Based Meta Adaption
abstract
In machine learning, meta-learning methods aim for fast adaptability to unknown tasks using prior knowledge. Model-based meta-reinforcement learning combines reinforcement learning via world models with Meta Reinforcement Learning (MRL) for increased sample efficiency. However, adaption to unknown tasks does not always result in preferable agent behavior. This paper introduces a new Meta Adaptation Controller (MAC) that employs MRL to apply a preferred robot behavior from one task to many similar tasks. To do this, MAC aims to find actions an agent has to take in a new task to reach a similar outcome as in a learned task. As a result, the agent will adapt quickly to the change in the dynamic and behave appropriately without the need to construct a reward function that enforces the preferred behavior.
Karam Daaboul, Joel Ikels, Johann Marius Zöllner
ICRA3
2023 Holistic Graph-based Motion Prediction
abstract
Motion prediction for automated vehicles in complex environments is a difficult task that is to be mastered when automated vehicles are to be used in arbitrary situations. Many factors influence the future motion of traffic participants starting with traffic rules and reaching from the interaction between each other to personal habits of human drivers. Therefore, we present a novel approach for a graph-based prediction based on a heterogeneous holistic graph representation that combines temporal information, properties and relations between traffic participants as well as relations with static elements such as the road network. The information is encoded through different types of nodes and edges that both are enriched with arbitrary features. We evaluated the approach on the INTERACTION and the Argoverse dataset and conducted an informative ablation study to demonstrate the benefit of different types of information for the motion prediction quality.
Daniel Grimm, Philip Schörner, Moritz Dreßler, Johann Marius Zöllner
ICRA4
2023 Sparsely-gated Mixture-of-Expert Layers for CNN Interpretability
abstract
Sparsely-gated Mixture of Expert (MoE) layers have been recently successfully applied for scaling large transformers, especially for language modeling tasks. An intriguing side effect of sparse MoE layers is that they convey inherent interpretability to a model via natural expert specialization. In this work, we apply sparse MoE layers to CNNs for computer vision tasks and analyze the resulting effect on model interpretability. To stabilize MoE training, we present both soft and hard constraint-based approaches. With hard constraints, the weights of certain experts are allowed to become zero, while soft constraints balance the contribution of experts with an additional auxiliary loss. As a result, soft constraints handle expert utilization better and support the expert specialization process, while hard constraints maintain more generalized experts and increase overall model performance. Our findings demonstrate that experts can implicitly focus on individual sub-domains of the input space. For example, experts trained for CIFAR-100 image classification specialize in recognizing different domains such as flowers or animals without previous data clustering. Experiments with RetinaNet and the COCO dataset further indicate that object detection experts can also specialize in detecting objects of distinct sizes.
Svetlana Pavlitska, Christian Hubschneider, Lukas Struppek, Johann Marius Zöllner
IJCNN4
2023 Perception Datasets for Anomaly Detection in Autonomous Driving: A Survey
abstract
Deep neural networks (DNN) which are employed in perception systems for autonomous driving require a huge amount of data to train on, as they must reliably achieve high performance in all kinds of situations. However, these DNN are usually restricted to a closed set of semantic classes available in their training data, and are therefore unreliable when confronted with previously unseen instances. Thus, multiple perception datasets have been created for the evaluation of anomaly detection methods, which can be categorized into three groups: real anomalies in real-world, synthetic anomalies augmented into real-world and completely synthetic scenes. This survey provides a structured and, to the best of our knowledge, complete overview and comparison of perception datasets for anomaly detection in autonomous driving. Each chapter provides information about tasks and ground truth, context information, and licenses. Additionally, we discuss current weaknesses and gaps in existing datasets to underline the importance of developing further data.
Daniel Bogdoll, Svenja Uhlemeyer, Kamil Kowol, Johann Marius Zöllner
IV4
2023 Bridging the Gap Between Multi-Step and One-Shot Trajectory Prediction via Self-Supervision
abstract
Accurate vehicle trajectory prediction is an unsolved problem in autonomous driving with various open research questions. State-of-the-art approaches regress trajectories either in a one-shot or step-wise manner. Although one-shot approaches are usually preferred for their simplicity, they relinquish powerful self-supervision schemes that can be constructed by chaining multiple time-steps. We address this issue by proposing a middle-ground where multiple trajectory segments are chained together. Our proposed Multi-Branch Self-Supervised Predictor receives additional training on new predictions starting at intermediate future segments. In addition, the model ’imagines’ the latent context and ’predicts the past’ while combining multi-modal trajectories in a tree-like manner. We deliberately keep aspects such as interaction and environment modeling simplistic and nevertheless achieve competitive results on the INTERACTION dataset. Furthermore, we investigate the sparsely explored uncertainty estimation of deterministic predictors. We find positive correlations between the prediction error and two proposed metrics, which might pave way for determining prediction confidence.
Faris Janjos, Max Keller, Maxim Dolgov, Johann Marius Zöllner
IV4
2023 Inverse Universal Traffic Quality - a Criticality Metric for Crowded Urban Traffic Scenes
abstract
An essential requirement for scenario-based testing the identification of critical scenes and their associated scenarios. However, critical scenes, such as collisions, occur comparatively rarely. Accordingly, large amounts of data must be examined. A further issue is that recorded real-world traffic often consists of scenes with a high number of vehicles, and it can be challenging to determine which are the most critical vehicles regarding the safety of an ego vehicle. Therefore, we present the inverse universal traffic quality, a criticality metric for urban traffic independent of predefined adversary vehicles and vehicle constellations such as intersection trajectories or car-following scenarios. Our metric is universally applicable for different urban traffic situations, e.g., intersections or roundabouts, and can be adjusted to certain situations if needed. Additionally, in this paper, we evaluate the proposed metric and compares its result to other well-known criticality metrics of this field, such as time-to-collision or post-encroachment time.
Barbara Schütt, Maximilian Zipfl, Johann Marius Zöllner, Eric Sax
IV3
2023 A Comprehensive Review on Ontologies for Scenario-based Testing in the Context of Autonomous Driving
abstract
The verification and validation of autonomous driving vehicles remains a major challenge due to the high complexity of autonomous driving functions. Scenario-based testing is a promising method for validating such a complex system. Ontologies can be utilized to produce test scenarios that are both meaningful and relevant. One crucial aspect of this process is selecting the appropriate method for describing the entities involved. The level of detail and specific entity classes required will vary depending on the system being tested. It is important to choose an ontology that properly reflects these needs.This paper summarizes key representative ontologies for scenario-based testing and related use cases in the field of autonomous driving. The considered ontologies are classified according to their level of detail for both static facts and dynamic aspects. Furthermore, the ontologies are evaluated based on the presence of important entity classes and the relations between them.
Maximilian Zipfl, Nina Koch, Johann Marius Zöllner
IV3
2022 Point Cloud Generation with Continuous Conditioning
abstract
Generative models can be used to synthesize 3D objects of high quality and diversity. However, there is typically no control over the properties of the generated object.This paper proposes a novel generative adversarial network (GAN) setup that generates 3D point cloud shapes conditioned on a continuous parameter. In an exemplary application, we use this to guide the generative process to create a 3D object with a custom-fit shape. We formulate this generation process in a multi-task setting by using the concept of auxiliary classifier GANs. Further, we propose to sample the generator label input for training from a kernel density estimation (KDE) of the dataset. Our ablations show that this leads to significant performance increase in regions with few samples. Extensive quantitative and qualitative experiments show that we gain explicit control over the object dimensions while maintaining good generation quality and diversity.
Larissa T. Triess, Andre Bühler, David Peter, Fabian Flohr, Johann Marius Zöllner
AISTATS5
2022 Quantification of Actual Road User Behavior on the Basis of Given Traffic Rules
abstract
Driving on roads is restricted by various traffic rules, aiming to ensure safety for all traffic participants. However, human road users usually do not adhere to these rules strictly, resulting in varying degrees of rule conformity. Such deviations from given rules are key components of today’s road traffic. In autonomous driving, robotic agents can disturb traffic flow, when rule deviations are not taken into account. In this paper, we present an approach to derive the distribution of degrees of rule conformity from human driving data. We demonstrate our method with the Waymo Open Motion dataset and Safety Distance and Speed Limit rules.
Daniel Bogdoll, Moritz Nekolla, Tim Joseph, Johann Marius Zöllner
IV4
2022 SAN: Scene Anchor Networks for Joint Action-Space Prediction
abstract
In this work, we present a novel multi-modal trajectory prediction architecture. We decompose the uncertainty of future trajectories along higher-level scene characteristics and lower-level motion characteristics, and model multi-modality along both dimensions separately. The scene uncertainty is captured in a joint manner, where diversity of scene modes is ensured by training multiple separate anchor networks which specialize to different scene realizations. At the same time, each network outputs multiple trajectories that cover smaller deviations given a scene mode, thus capturing motion modes. In addition, we train our architectures with an outlier-robust regression loss function, which offers a trade-off between the outlier-sensitive L2and outlier-insensitive L1losses. Our scene anchor model achieves improvements over the state of the art on the INTERACTION dataset, outperforming the StarNet architecture from our previous work.
Faris Janjos, Maxim Dolgov, Muhamed Kuric, Yinzhe Shen, Johann Marius Zöllner
IV5
2022 StarNet: Joint Action-Space Prediction with Star Graphs and Implicit Global-Frame Self-Attention
abstract
In this work, we present a novel multi-modal multi-agent trajectory prediction architecture, focusing on map and interaction modeling using graph representation. For the purposes of map modeling, we capture rich topological structure into vector-based star graphs, which enable an agent to directly attend to relevant regions along polylines that are used to represent the map. We denote this architecture StarNet, and integrate it into a single-agent prediction setting. As the main result, we extend this architecture to joint scene-level prediction, which produces multiple agents’ predictions simultaneously. The key idea in joint-StarNet is integrating the awareness of one agent in its own reference frame with how it is perceived from the points of view of other agents. We achieve this via masked self-attention. Both proposed architectures are built on top of the action-space prediction framework introduced in our previous work, which ensures kinematically feasible trajectory predictions. We evaluate the methods on the interaction-rich inD and INTERACTION datasets, with both StarNet and joint-StarNet achieving improvements over state of the art.
Faris Janjos, Maxim Dolgov, Johann Marius Zöllner
IV3
2022 Learning Reward Models for Cooperative Trajectory Planning with Inverse Reinforcement Learning and Monte Carlo Tree Search
abstract
Cooperative trajectory planning methods for automated vehicles can solve traffic scenarios that require a high degree of cooperation between traffic participants. However, for cooperative systems to integrate into human-centered traffic, the automated systems must behave human-like so that humans can anticipate the system’s decisions. While Reinforcement Learning has made remarkable progress in solving the decision-making part, it is non-trivial to parameterize a reward model that yields predictable actions. This work employs feature-based Maximum Entropy Inverse Reinforcement Learning combined with Monte Carlo Tree Search to learn reward models that maximize the likelihood of recorded multi-agent cooperative expert trajectories. The evaluation demonstrates that the approach can recover a reasonable reward model that mimics the expert and performs similarly to a manually tuned baseline reward model.
Karl Kurzer, Matthias Bitzer 0001, Johann Marius Zöllner
IV3
2022 A Unified Description of Proving Grounds and Test Areas for Automated and Connected Vehicles
abstract
Highly automated and connected vehicles are being tested more and more on proving grounds as well as in test areas in public traffic. Between simulation based testing in laboratories and real-world testing, such testing environments offer the opportunity to evaluate scenarios while providing partially controllable and partially observable environments. Although the number of public test areas is increasing, their capabilities and opportunities have not yet been analyzed and certainly have not been brought to a formal description.To overcome this shortage, this paper presents a unified taxonomy that describes proving grounds and test areas for connected autonomous driving in a uniform manner. We introduce a machine-readable and processable representation, that makes it possible to analyze test areas regarding their abilities and benefits to facilitate testing of assisted, automated and connected vehicles. So, necessary technological bricks are classified, a corresponding ontology is presented, common algorithms are discussed and evaluated at the example of smart and connected infrastructure in the Test Area Autonomous Driving Baden-Württemberg. We conclude by giving an overview of future research questions to motivate researchers to use the proposed model as a description baseline for further V&V approaches.
Marc Rene Zofka, Tobias Fleck, Johann Marius Zöllner
IV3
2022 Multimodal Detection of Unknown Objects on Roads for Autonomous Driving
abstract
Tremendous progress in deep learning over the last years has led towards a future with autonomous vehicles on our roads. Nevertheless, the performance of their perception systems is strongly dependent on the quality of the utilized training data. As these usually only cover a fraction of all object classes an autonomous driving system will face, such systems struggle with handling the unexpected. In order to safely operate on public roads, the identification of objects from unknown classes remains a crucial task. In this paper, we propose a novel pipeline to detect unknown objects. Instead of focusing on a single sensor modality, we make use of lidar and camera data by combining state-of-the art detection models in a sequential manner. We evaluate our approach on the Waymo Open Perception Dataset and point out current research gaps in anomaly detection.
Daniel Bogdoll, Enrico Eisen, Maximilian Nitsche, Christin Scheib, Johann Marius Zöllner
SMC5
2022 Ad-datasets: A Meta-collection of Data Sets for Autonomous Driving
abstract
Autonomous driving is among the largest domains in which deep learning has been fundamental for progress within the last years. The rise of datasets went hand in hand with this development. All the more striking is the fact that researchers do not have a tool available that provides a quick, comprehensive and up-to-date overview of data sets and their features in the domain of autonomous driving. In this paper, we present ad-datasets, an online tool that provides such an overview for more than 150 data sets. The tool enables users to sort and filter the data sets according to currently 16 different categories. ad-datasets is an open-source project with community contributions. It is in constant development, ensuring that the content stays up-to-date.
Daniel Bogdoll, Felix Schreyer, Johann Marius Zöllner
VEHITS3
2022 An Application of Scenario Exploration to Find New Scenarios for the Development and Testing of Automated Driving Systems in Urban Scenarios
abstract
Verification and validation are major challenges for developing automated driving systems. A concept that gets more and more recognized for testing in automated driving is scenario-based testing. However, it introduces the problem of what scenarios are relevant for testing and which are not. This work aims to find relevant, interesting, or critical parameter sets within logical scenarios by utilizing Bayes optimization and Gaussian processes. The parameter optimization is done by comparing and evaluating six different metrics in two urban intersection scenarios. Finally, a list of ideas this work leads to and should be investigated further is presented.
Barbara Schütt, Marc Heinrich, Sonja Marahrens, Johann Marius Zöllner, Eric Sax
VEHITS4
2022 A Realism Metric for Generated LiDAR Point Clouds
abstract
Abstract A considerable amount of research is concerned with the generation of realistic sensor data. LiDAR point clouds are generated by complex simulations or learned generative models. The generated data is usually exploited to enable or improve downstream perception algorithms. Two major questions arise from these procedures: First, how to evaluate the realism of the generated data? Second, does more realistic data also lead to better perception performance? This paper addresses both questions and presents a novel metric to quantify the realism of LiDAR point clouds. Relevant features are learned from real-world and synthetic point clouds by training on a proxy classification task. In a series of experiments, we demonstrate the application of our metric to determine the realism of generated LiDAR data and compare the realism estimation of our metric to the performance of a segmentation model. We confirm that our metric provides an indication for the downstream segmentation performance.
Larissa T. Triess, Christoph Rist, David Peter, Johann Marius Zöllner
Int. J. Comput. Vis.4
2021 Sample-Specific Output Constraints for Neural Networks
abstract
It is common practice to constrain the output space of a neural network with the final layer to a problem-specific value range. However, for many tasks it is desired to restrict the output space for each input independently to a different subdomain with a non-trivial geometry, e.g. in safety-critical applications, to exclude hazardous outputs sample-wise. We propose ConstraintNet—a scalable neural network architecture which constrains the output space in each forward pass independently. Contrary to prior approaches, which perform a projection in the final layer, ConstraintNet applies an input-dependent parametrization of the constrained output space. Thereby, the complete interior of the constrained region is covered and computational costs are reduced significantly. For constraints in form of convex polytopes, we leverage the vertex representation to specify the parametrization. The second modification consists of adding an auxiliary input in form of a tensor description of the constraint to enable the handling of multiple constraints for the same sample. Finally, ConstraintNet is end-to-end trainable with almost no overhead in the forward and backward pass. We demonstrate ConstraintNet on two regression tasks: First, we modify a CNN and construct several constraints for facial landmark detection tasks. Second, we demonstrate the application to a follow object controller for vehicles and accomplish safe reinforcement learning in this case. In both experiments, ConstraintNet improves performance and we conclude that our approach is promising for applying neural networks in safety-critical environments.
Mathis Brosowsky, Florian Keck, Olaf Dünkel, Johann Marius Zöllner
AAAI4
2021 Data-Driven Merging of Car-Following Models for Interaction-Aware Vehicle Speed Prediction
Johannes Buyer, Dominic Waldenmayer, Raoul Daniel Zöllner, Johann Marius Zöllner
FUSION4
2021 Tuning Multi Object Tracking Systems using Bayesian Optimization
Tobias Fleck, Johann Marius Zöllner
FUSION2
2021 Safe Continuous Control with Constrained Model-Based Policy Optimization
abstract
The applicability of reinforcement learning (RL) algorithms in real-world domains often requires adherence to safety constraints, a need difficult to address given the asymptotic nature of the classic RL optimization objective. In contrast to the traditional RL objective, safe exploration considers the maximization of expected returns under safety constraints expressed in expected cost returns. We introduce a model-based safe exploration algorithm for constrained high-dimensional control to address the often prohibitively high sample complexity of model-free safe exploration algorithms. Further, we provide theoretical and empirical analyses regarding the implications of model-usage on constrained policy optimization problems and introduce a practical algorithm that accelerates policy search with model-generated data. The need for accurate estimates of a policy’s constraint satisfaction is in conflict with accumulating model-errors. We address this issue by quantifying model-uncertainty as the expected Kullback-Leibler divergence between predictions of an ensemble of probabilistic dynamics models and constrain this error-measure, resulting in an adaptive resampling scheme and dynamically limited rollout horizons. We evaluate this approach on several simulated constrained robot locomotion tasks with high-dimensional action- and state-spaces. Our empirical studies find that our algorithm reaches model-free performances with a 10-20 fold reduction of training samples while maintaining approximate constraint satisfaction levels of model-free methods.
Moritz A. Zanger, Karam Daaboul, Johann Marius Zöllner
IROS3
2021 Safe Deep Reinforcement Learning for Adaptive Cruise Control by Imposing State-Specific Safe Sets
abstract
Deep reinforcement learning has been increasingly discussed for solving continuous control tasks in the field of autonomous driving and driver assistance systems. However, trial-and-error learning and the black-box character of neural networks make it prone to accidental damage in safety-critical environments. We propose to learn a safe vehicle following controller with deep reinforcement learning by imposing state-specific safe sets as output constraints on the policy and call the approach ACC 4S. The main safety goal is the avoidance of rear-end collisions with the front vehicle. To achieve this, we build on the Responsibility-Sensitive Safety model and derive an upper bound for the demanded acceleration. Further limitations emerge from regulatory standards and system limits. We end up with state-specific intervals of safe actions, the safe sets. To impose these safe sets as hard output constraints on the policy, we leverage the recently proposed neural network architecture ConstraintNet. We compare ConstraintNet with an unconstrained neural network, additional clipping as postprocessing, and clipping as part of the neural network. The results show, that the proposed safe sets ensure collision avoidance and ConstraintNet shows superior performance compared to the other approaches.
Mathis Brosowsky, Florian Keck, Jakob Ketterer, Simon T. Isele, Daniel Slieter, Johann Marius Zöllner
IV6
2021 Learning Semantics on Radar Point-Clouds
abstract
Localization and perception research for Autonomous Driving is mainly focused on camera and LiDAR data, rarely on radar data. We apply an automated labeling pipeline to semantically annotate real world radar measurements, manually correct point-wise labels to obtain ground-truth, and apply supervised learning models on this data. To assign an attribute, called class label, to every point of an input cloud is hereby referred to as semantic segmentation. Transferring approaches of LiDAR segmentation into the similar data structure, we research deep-learning semantic segmentation on radar point clouds. Compared to classical Cartesian coordinates, a polar coordinate input discretization benefits the dynamically changing number of radar detections per sensing cycle and simplifies to model the quasi-radial sensor resolution. Moreover, we evaluate different network architectures, examine radar feature channels and also temporal consistency by attention map concatenation. Our contribution is twofold. First, featuring a semantically labeled real world radar dataset for ground truth. Second, our supervised learning approach to solve semantic segmentation on radar point-cloud data. Our classification benchmark network yields 56.1 % weighted Intersection of a Union of relevant classes for radar, while reaching a real-time framerate of 12.4ms.
Simon T. Isele, Fabian E. Klein, Mathis Brosowsky, Johann Marius Zöllner
IV4
2021 Self-Supervised Action-Space Prediction for Automated Driving
abstract
Making informed driving decisions requires reliable prediction of other vehicles' trajectories. In this paper, we present a novel learned multi-modal trajectory prediction architecture for automated driving. It achieves kinematically feasible predictions by casting the learning problem into the space of accelerations and steering angles - by performing action-space prediction, we can leverage valuable model knowledge. Additionally, the dimensionality of the action manifold is lower than that of the state manifold, whose intrinsically correlated states are more difficult to capture in a learned manner. For the purpose of action-space prediction, we present the simple Feed-Forward Action-Space Prediction (FFW-ASP) architecture. Then, we build on this notion and introduce the novel Self-Supervised Action-Space Prediction (SSP-ASP) architecture that outputs future environment context features in addition to trajectories. A key element in the self-supervised architecture is that, based on an observed action history and past context features, future context features are predicted prior to future trajectories. The proposed methods are evaluated on real-world datasets containing urban intersections and roundabouts, and show accurate predictions, outperforming state-of-the-art for kinematically feasible predictions in several prediction metrics.
Faris Janjos, Maxim Dolgov, Johann Marius Zöllner
IV3
2021 Generalizing Decision Making for Automated Driving with an Invariant Environment Representation using Deep Reinforcement Learning
abstract
Data driven approaches for decision making applied to automated driving require appropriate generalization strategies, to ensure applicability to the world's variability. Current approaches either do not generalize well beyond the training data or are not capable to consider a variable number of traffic participants. Therefore we propose an invariant environment representation from the perspective of the ego vehicle. The representation encodes all necessary information for safe decision making. To assess the generalization capabilities of the novel environment representation, we train our agents on a small subset of scenarios and evaluate on the entire diverse set of scenarios. Here we show that the agents are capable to generalize successfully to unseen scenarios, due to the abstraction. In addition we present a simple occlusion model that enables our agents to navigate intersections with occlusions without a significant change in performance.
Karl Kurzer, Philip Schörner, Alexander Albers, Hauke Thomsen, Karam Daaboul, Johann Marius Zöllner
IV6
2021 Temporal Feature Networks for CNN based Object Detection
abstract
For reliable environment perception, the use of temporal information is essential in some situations. Especially for object detection, sometimes a situation can only be understood in the right perspective through temporal information. Since image-based object detectors are currently based almost exclusively on CNN architectures, an extension of their feature extraction with temporal features seems promising. Within this work we investigate different architectural components for a CNN-based temporal information extraction. We present a Temporal Feature Network which is based on the insights gained from our architectural investigations. This network is trained from scratch without any ImageNet information based pre-training as these images are not available with temporal information. The object detector based on this network is evaluated against the non-temporal counterpart as baseline and achieves competitive results in an evaluation on the KITTI object detection dataset.
Michael Weber 0009, Tassilo Wald, Johann Marius Zöllner
IV3
2021 Radar Artifact Labeling Framework (RALF): Method for Plausible Radar Detections in Datasets
abstract
Research on localization and perception for Autonomous Driving is mainly focused on camera and LiDAR datasets, rarely on radar data. Manually labeling sparse radar point clouds is challenging. For a dataset generation, we propose the cross sensor Radar Artifact Labeling Framework (RALF). Automatically generated labels for automotive radar data help to cure radar shortcomings like artifacts for the application of artificial intelligence. RALF provides plausibility labels for radar raw detections, distinguishing between artifacts and targets. The optical evaluation backbone consists of a generalized monocular depth image estimation of surround view cameras plus LiDAR scans. Modern car sensor sets of cameras and LiDAR allow to calibrate image-based relative depth information in overlapping sensing areas. K-Nearest Neighbors matching relates the optical perception point cloud with raw radar detections. In parallel, a temporal tracking evaluation part considers the radar detections' transient behavior. Based on the distance between matches, respecting both sensor and model uncertainties, we propose a plausibility rating of every radar detection. We validate the results by evaluating error metrics on semi-manually labeled ground truth dataset of $3.28\cdot10^6$ points. Besides generating plausible radar detections, the framework enables further labeled low-level radar signal datasets for applications of perception and Autonomous Driving learning tasks.
Simon T. Isele, Marcel P. Schilling, Fabian E. Klein, Sascha Saralajew, Johann Marius Zöllner
VEHITS5
2020 Automated Focal Loss for Image based Object Detection
abstract
Current state-of-the-art object detection algorithms still suffer the problem of imbalanced distribution of training data over object classes and background. Recent work introduced a new loss function called focal loss to mitigate this problem, but at the cost of an additional hyperparameter. Manually tuning this hyperparameter for each training task is highly time-consuming. With automated focal loss we introduce a new loss function which substitutes this hyperparameter by a parameter that is automatically adapted during the training progress and controls the amount of focusing on hard training examples. We show on the COCO benchmark that this leads to an up to 30% faster training convergence. We further introduced a focal regression loss which on the more challenging task of 3D vehicle detection outperforms other loss functions by up to 1.8 AOS and can be used as a value range independent metric for regression.
Michael Weber 0009, Michael Fürst, Johann Marius Zöllner
IV3
2020 Runtime Optimization of a CNN Model for Environment Perception
abstract
For self driving cars one of the current key technologies are deep neural networks. Especially in camera based environment perception they are absolutely irreplaceable. The currently developed network models are usually executed on high end consumer or server GPUs. Also the verification of the real-time properties is mostly based on these GPUs. However, if these models are to be used in near-series applications, the question arises whether they can also be used on significantly reduced hardware. To address this question, we conduct a case study with a camera based traffic light detection system. Promising optimization techniques are adapted and applied to the model to investigate potential performance gains achievable with these techniques in the context of self driving car environment perception. In particular, the trade-off between quality and speed is to be examined in detail.
Michael Weber 0009, Christof Wendenius, Johann Marius Zöllner
IV3
2020 Robust Tracking of Reference Trajectories for Autonomous Driving in Intelligent Roadside Infrastructure
abstract
High quality reference data is crucial for the development of autonomous driving applications. Unfortunately, datasets including fixed, reproducible static environments that contain manifold interactions between traffic participants are not widely available. In this paper we propose a camera based trajectory estimation framework that enables the generation of reference trajectory data in stationary roadside infrastructure. We develop a Simple Online Realtime Tracking (SORT) algorithm that tracks objects in image space utilizing the tracking-by-detection paradigm with a deep neural network detector. By projecting tracks to a ground model, we are able to gather cartesian and georeferenced trajectories for manually driven and autonomous vehicles in the field. We evaluate the framework in stationary roadside infrastructure in the Test Area Autonomous Driving Baden-Württemberg, Germany. A vehicle equipped with inertial measurement unit and differential GPS is used to generate ground truth positions that are compared with our framework.
Tobias Fleck, Sven Ochs, Marc Rene Zofka, Johann Marius Zöllner
IV4
2020 Accelerating Cooperative Planning for Automated Vehicles with Learned Heuristics and Monte Carlo Tree Search
abstract
Efficient driving in urban traffic scenarios requires foresight. The observation of other traffic participants and the inference of their possible next actions depending on the own action is considered cooperative prediction and planning. Humans are well equipped with the capability to predict the actions of multiple interacting traffic participants and plan accordingly, without the need to directly communicate with others. Prior work has shown that it is possible to achieve effective cooperative planning without the need for explicit communication. However, the search space for cooperative plans is so large that most of the computational budget is spent on exploring the search space in unpromising regions that are far away from the solution. To accelerate the planning process, we combined learned heuristics with a cooperative planning method to guide the search towards regions with promising actions, yielding better solutions at lower computational costs.
Karl Kurzer, Marcus Fechner, Johann Marius Zöllner
IV3
2020 Optimization of Sampling-Based Motion Planning in Dynamic Environments Using Neural Networks
abstract
Motion planning for autonomous vehicles is a challenging task, especially in dynamic environments. The motion of the vehicle itself needs to be considered while the vehicle needs to react to its surroundings at the same time. Sampling-based algorithms proved to be suitable to cope with these challenges. However, the performance of these algorithms is highly dependent on the sampling heuristics, which in turn are often hand crafted and thus need a large amount of tuning. Therefore, we developed two approaches based on deep learning to learn these heuristics for sampling-based motion planning in dynamic environments. The first approach predicts a discrete probability distribution for each point in time of the future trajectory, whereas the second approach directly predicts a variety of trajectories by using dropout sampling. Both approaches are based on an environment representation encoded as a grid-based tensor. The learned heuristics are integrated into an existing planning framework based on particle swarm optimization and are evaluated in several situations. This shows how to combine the strengths of machine learning based approaches and the traceability of rule- or model-based approaches. The evaluation demonstrates that, in total, we were able to improve on our current heuristics. However, none of the approaches performed consistently better in all scenarios evaluated.
Philip Schörner, Mark Timon Hüneberg, Johann Marius Zöllner
IV3
2020 Scan-based Semantic Segmentation of LiDAR Point Clouds: An Experimental Study
abstract
Autonomous vehicles need to have a semantic understanding of the three-dimensional world around them in order to reason about their environment. State of the art methods use deep neural networks to predict semantic classes for each point in a LiDAR scan. A powerful and efficient way to process LiDAR measurements is to use two-dimensional, image-like projections. In this work, we perform a comprehensive experimental study of image-based semantic segmentation architectures for LiDAR point clouds. We demonstrate various techniques to boost the performance and to improve runtime as well as memory constraints. First, we examine the effect of network size and suggest that much faster inference times can be achieved at a very low cost to accuracy. Next, we introduce an improved point cloud projection technique that does not suffer from systematic occlusions. We use a cyclic padding mechanism that provides context at the horizontal field-of-view boundaries. In a third part, we perform experiments with a soft Dice loss function that directly optimizes for the intersection-over-union metric. Finally, we propose a new kind of convolution layer with a reduced amount of weight-sharing along one of the two spatial dimensions, addressing the large difference in appearance along the vertical axis of a LiDAR scan. We propose a final set of the above methods with which the model achieves an increase of 3.2% in mIoU segmentation performance over the baseline while requiring only 42% of the original inference time.
Larissa T. Triess, David Peter, Christoph Rist, Johann Marius Zöllner
IV4
2019 Beyond Bounding Boxes: Using Bounding Shapes for Real-Time 3D Vehicle Detection from Monocular RGB Images
abstract
The representation of objects as 2D bounding boxes in monocular RGB images limits the faculty of current computer vision systems to 2D object detection. It fails to provide crucial information such as the orientation of other vehicles, which is vital for autonomous driving. At the same time, real-time performance is essential to qualify an approach for deployment in a productive environment. In order to tackle this problem, we present an approach that predicts several key points selected from a virtual 3D bounding box around a vehicle instead of a pure 2D bounding box. These key points can be interpreted as a bounding shape. With this novel representation we can calculate the actual 3D bounding box of the corresponding object. Thanks to the straightforward implementation of bounding shape in any current state-of-the-art 2D object detector both for singleshot frameworks like YOLO or SSD as well as for two-stage detectors like Faster-RCNN with a minimum of computational overhead, it is able to be run in real-time while providing additional useful information for vehicle detection. We exemplify the extension of SSD to Bounding Shape SSD ( BS3D) and evaluate our approach using the challenging KITTI as well as the novel VIPER dataset.
Nils Gählert, Jun-Jun Wan, Michael Weber 0009, Johann Marius Zöllner, Uwe Franke, Joachim Denzler
IV4
2019 Generation of Scenes in Intersections for the Validation of Highly Automated Driving Functions
abstract
The simulation of traffic scenes in the environment of an automated vehicle promises to make significant contribution to the validation of automated driving. The construction of models which describe traffic scenes in a generic manner is complicated, since the parameter space of the scenes is infinite. This paper introduces a statistical approach to generate traffic scenes in intersections. A generic model which allows us to represent the scenes is proposed. The concept of Bayesian networks is used to fit the model onto a publicly accessible dataset and to infer traffic scenes from the model. A quantitative evaluation of the results is achieved by the calculation of the total variation distance (TVD) between the distributions of several physical properties.
Stefan Jesenski, Jan Erik Stellet, Florian Alexander Schiegg, Johann Marius Zöllner
IV4
2019 Benchmarking and Functional Decomposition of Automotive Lidar Sensor Models
abstract
Simulation-based testing is seen as a major requirement for the safety validation of highly automated driving. One crucial part of such test architectures are models of environment perception sensors such as camera, lidar and radar sensors. Currently, an objective evaluation and the comparison of different modeling approaches for automotive lidar sensors are still a challenge. In this work, a real lidar sensor system used for object recognition is first functionally decomposed. The resulting sequence of processing blocks and interfaces is then mapped onto simulation methods. Subsequently, metrics applied to the aforementioned interfaces are derived, enabling a quantitative comparison between simulated and real sensor data at different steps of the processing pipeline. Benchmarks for several existing sensor models at a concrete selected interface are performed using those metrics by comparing them to measurements gained from the real sensor. Finally, we outline how metrics on low-level interfaces can correlate with results on more abstract ones. A major achievement of this work lies within the commonly accepted interfaces and a common understanding of real and virtual lidar sensor systems and, even more important, an initial guideline for the quantitative comparison of sensor models with the ambition to support future validation of virtual sensor models.
Philipp Rosenberger, Martin Holder, Sebastian Hueh, Hermann Winner, Tobias Fleck, Marc Rene Zofka, Johann Marius Zöllner, Thomas D'hondt, Benjamin Wassermann
IV7
2019 Predictive Trajectory Planning in Situations with Hidden Road Users Using Partially Observable Markov Decision Processes
abstract
State of the art emergency brake assistant systems solely based on sensor measurements reduced the number of traffic accidents and casualties drastically in recent years. In order to be able to react on road users who elude a vehicle's field of view because of sensor limits or occlusions, this paper presents an approach to anticipate potential hidden traffic participants in occluded areas in the decision making process of an autonomous vehicle. A Partially Observable Markov Decision Process is used to determine the vehicle's longitudinal motion. Observations are made using the vehicle's field of view. Therefore the field of view is calculated with a generic model of a sensor setup in dependence of the current or the predicted environment. In this way, the vehicle can either observe that it detects a previously hidden road user or receives information that the road is clear. In total, that allows the vehicle to better anticipate future developments. Therefore, assumptions about vehicles that may be located in hidden areas need to be made. We demonstrate the approach in two scenarios. Firstly in a scenario, where the vehicle has to move cautiously into the intersection with a minimum number of actions and secondly in a typical scenario for urban traffic. Evaluation shows, that the approach is able to anticipate hidden road users correctly and act accordingly.
Philip Schörner, Lars Töttel, Jens Doll, Johann Marius Zöllner
IV4
2019 CNN-based synthesis of realistic high-resolution LiDAR data
abstract
This paper presents a novel CNN-based approach for synthesizing high-resolution LiDAR point cloud data. Our approach generates semantically and perceptually realistic results with guidance from specialized loss-functions. First, we utilize a modified per-point loss that addresses missing LiDAR point measurements. Second, we align the quality of our generated output with real-world sensor data by applying a perceptual loss. In large-scale experiments on real-world datasets, we evaluate both the geometric accuracy and semantic segmentation performance using our generated data vs. ground truth. In a mean opinion score testing we further assess the perceptual quality of our generated point clouds. Our results demonstrate a significant quantitative and qualitative improvement in both geometry and semantics over traditional non CNN-based upsampling methods.
Larissa T. Triess, David Peter, Christoph Rist, Markus Enzweiler, Johann Marius Zöllner
IV5
2019 Direct 3D Detection of Vehicles in Monocular Images with a CNN based 3D Decoder
abstract
In autonomous driving, the detection of objects like surrounding vehicles based on monocular RGB images is usually performed by 2D bounding box detectors. The resulting 2D objects can be used for a first coarse 3D position estimate but for a precise location, additional sensor data has to be taken into account. For further use in sensor fusion systems and environment maps it is preferable to detect objects, their orientation and dimensions directly in 3D coordinates. To address this 3D object detection task, we propose a direct 3D bounding box estimator which is realized as CNN decoder module and can be connected to most 2D object detectors like SSD[1], OverFeat[2], YOLO[3] and RetinaNet[4] or directly to CNN feature extractors like VGG [2] and ResNet [5]. The 3D parameters of the objects such as dimension and orientation are directly predicted by the CNN module. To successfully train this complex MultiNet architecture, a combination and modification of current loss functions is proposed. The fastest of the proposed network module combinations is capable of detecting objects in 3D camera coordinates at a frame rate of 28 fps.
Michael Weber 0009, Michael Fürst, Johann Marius Zöllner
IV3
2018 Predicting Ego-Vehicle Paths from Environmental Observations with a Deep Neural Network
abstract
Advanced driver assistance systems allow for increasing user comfort and safety by sensing the environment and anticipating upcoming hazards. Often, this requires to accurately predict how situations will change. Recent approaches make simplifying assumptions on the predictive model of the Ego-Vehicle motion or assume prior knowledge, such as road topologies, to be available. However, in many urban areas this assumption is not satisfied. Furthermore, temporary changes (e.g. construction areas, vehicles parked on the street) are not considered by such models. Since many cars observe the environment with several different sensors, predictive models can benefit from them by considering environmental properties. In this work, we present an approach for an Ego-Vehicle path prediction from such sensor measurements of the static vehicle environment. Besides proposing a learned model for predicting the driver's multi-modal future path as a grid-based prediction, we derive an approach for extracting paths from it. In driver assistance systems both can be used to solve varying assistance tasks. The proposed approach is evaluated on real driving data and outperforms several baseline approaches.
Ulrich Baumann, Claudius Guiser, Michael Herman, Johann Marius Zöllner
ICRA4
2018 From G2 to G3 Continuity: Continuous Curvature Rate Steering Functions for Sampling-Based Nonholonomic Motion Planning
abstract
Motion planning for car-like robots is one of the major challenges in automated driving. It requires to solve a two-point boundary value problem (BVP) in real time while taking into account the nonholonomic constraints of the vehicle and the obstacles in the non-convex environment. This paper introduces Hybrid Curvature Rate (HCR) and Continuous Curvature Rate (CCR) Steer: Two novel steering functions for car-like robots that compute a curvature rate continuous solution of the two-point BVP. Hard constraints on the maximum curvature, maximum curvature rate, and maximum curvature acceleration are satisfied resulting in directly driveable G3continuous paths. The presented steering functions are benchmarked in terms of computation time and path length against its G1and G2continuous counterparts, namely Dubins, Reeds-Shepp, Hybrid Curvature, and Continuous Curvature Steer. It is shown that curvature rate continuity can be achieved with only small computational overhead. The generic motion planner Bidirectional RRT* is finally used to present the effectiveness of HCR and CCR Steer in three challenging automated driving scenarios.
Holger Banzhaf, Nijanthan Berinpanathan, Dennis Nienhüser, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2018 Decentralized Cooperative Planning for Automated Vehicles with Hierarchical Monte Carlo Tree Search
abstract
Today's automated vehicles lack the ability to cooperate implicitly with others. This work presents a Monte Carlo Tree Search (MCTS) based approach for decentralized cooperative planning using macro-actions for automated vehicles in heterogeneous environments. Based on cooperative modeling of other agents and Decoupled-UCT (a variant of MCTS), the algorithm evaluates the state-action-values of each agent in a cooperative and decentralized manner, explicitly modeling the interdependence of actions between traffic participants. Macro-actions allow for temporal extension over multiple time steps and increase the effective search depth requiring fewer iterations to plan over longer horizons. Without predefined policies for macro-actions, the algorithm simultaneously learns policies over and within macro-actions. The proposed method is evaluated under several conflict scenarios, showing that the algorithm can achieve effective cooperative planning with learned macro-actions in heterogeneous environments.
Karl Kurzer, Chenyang Zhou 0002, Johann Marius Zöllner
Intelligent Vehicles Symposium3
2018 MultiNet: Real-time Joint Semantic Reasoning for Autonomous Driving
abstract
While most approaches to semantic reasoning have focused on improving performance, in this paper we argue that computational times are very important in order to enable real time applications such as autonomous driving. Towards this goal, we present an approach to joint classification, detection and semantic segmentation using a unified architecture where the encoder is shared amongst the three tasks. Our approach is very simple, can be trained end-to-end and performs extremely well in the challenging KITTI dataset. Our approach is also very efficient, allowing us to perform inference at more then 23 frames per second. Training scripts and trained weights to reproduce our results can be found here: https://github.com/MarvinTeichmann/MultiNet
Marvin Teichmann, Michael Weber 0009, Johann Marius Zöllner, Roberto Cipolla, Raquel Urtasun
Intelligent Vehicles Symposium3
2018 Adaptive Behavior Generation for Autonomous Driving using Deep Reinforcement Learning with Compact Semantic States
abstract
Making the right decision in traffic is a challenging task that is highly dependent on individual preferences as well as the surrounding environment. Therefore it is hard to model solely based on expert knowledge. In this work we use Deep Reinforcement Learning to learn maneuver decisions based on a compact semantic state representation. This ensures a consistent model of the environment across scenarios as well as a behavior adaptation function, enabling on-line changes of desired behaviors without re-training. The input for the neural network is a simulated object list similar to that of Radar or Lidar sensors, superimposed by a relational semantic scene description. The state as well as the reward are extended by a behavior adaptation function and a parameterization respectively. With little expert knowledge and a set of mid-level actions, it can be seen that the agent is capable to adhere to traffic rules and learns to drive safely in a variety of situations.
Karl Kurzer, Tobias Wingert, Florian Kuhnt, Johann Marius Zöllner
Intelligent Vehicles Symposium5
2018 Making Bertha Cooperate-Team AnnieWAY's Entry to the 2016 Grand Cooperative Driving Challenge
abstract
This paper presents the concepts and methods utilized by Team AnnieWAY for the 2016 Grand Cooperative Driving Challenge. The paper introduces the automated vehicle BerthaOne. The vehicle, even though being based on the Bertha platform, distinguishes itself from its siblings by its software modules and algorithms. We, therefore, describe its system architecture and algorithms for perception, cooperation and motion planning. In Particular, we present a motion planner that plans different maneuvers flexibly by augmenting the cost function with situation specific cost terms. We subsequently describe the requirements of the 2016 GCDC and evaluate our performance during the competition.
Ömer Sahin Tas, Niels Ole Salscheider, Fabian Poggenhans, Sascha Wirges, Claudio Bandera, Marc Rene Zofka, Tobias Strauß, Johann Marius Zöllner, Christoph Stiller
IEEE Trans. Intell. Transp. Syst.8
2017 Towards Grasping with Spiking Neural Networks for Anthropomorphic Robot Hands
Juan Camilo Vasquez Tieck, Heiko Donat, Jacques Kaiser, Igor Peric, Stefan Ulbrich, Arne Roennau, Johann Marius Zöllner, Rüdiger Dillmann
ICANN (1)7
2017 The future of parking: A survey on automated valet parking with an outlook on high density parking
abstract
In the near future, humans will be relieved from parking. Major improvements in autonomous driving allow the realization of automated valet parking (AVP). It enables the vehicle to drive to a parking spot and park itself. This paper presents a review of the intelligent vehicles literature on AVP. An overview and analysis of the core components of AVP such as the platforms, sensor setups, maps, localization, perception, environment model, and motion planning is provided. Leveraging the potential of AVP, high density parking (HDP) is reviewed as a future research direction with the capability to either reduce the necessary space for parking by up to 50 % or increase the capacity of future parking facilities. Finally, a synthesized view discussing the remaining challenges in automated valet parking and the technological requirements for high density parking is given.
Holger Banzhaf, Dennis Nienhüser, Steffen Knoop, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2017 High density valet parking using k-deques in driveways
abstract
Advances in autonomous driving and the introduction of automated valet parking allow the optimization of parking space. A future concept is high density valet parking with the potential to either reduce the extensive land use for parking or increase the capacity of existing parking facilities. This paper presents a novel approach that integrates high density parking into an existing parking lot, by explicitly making use of parking in the driving lane and reducing the shunting operations per vehicle. The proposed parking scheme allows vehicles to park either perpendicular to the driveway or in double-ended queues with k parking spots (k-deque) on the side of the driving lane. Leveraging the potential of such a layout increases the capacity of a parking lot by up to 25 %, while keeping the maximum number of shunts per car below [k/2] +1 between entry and exit. A dynamic simulation verifies the theoretical analysis and compares the performance of different deque lengths with respect to the distances traveled and the number of shunts per vehicle.
Holger Banzhaf, Frank-M. Quedenfeld, Dennis Nienhüser, Steffen Knoop, Johann Marius Zöllner
Intelligent Vehicles Symposium5
2017 Fully convolutional neural networks for dynamic object detection in grid maps
abstract
Grid maps are widely used in robotics to represent obstacles in the environment and differentiating dynamic objects from static infrastructure is essential for many practical applications. In this work, we present a methods that uses a deep convolutional neural network (CNN) to infer whether grid cells are covering a moving object or not. Compared to tracking approaches, that use e.g. a particle filter to estimate grid cell velocities and then make a decision for individual grid cells based on this estimate, our approach uses the entire grid map as input image for a CNN that inspects a larger area around each cell and thus takes the structural appearance in the grid map into account to make a decision. Compared to our reference method, our concept yields a performance increase from 83.9% to 97.2%. A runtime optimized version of our approach yields similar improvements with an execution time of just 10 milliseconds.
Florian Piewak, Timo Rehfeld, Michael Weber 0009, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2017 Learning how to drive in a real world simulation with deep Q-Networks
abstract
We present a reinforcement learning approach using Deep Q-Networks to steer a vehicle in a 3D physics simulation. Relying solely on camera image input the approach directly learns steering the vehicle in an end-to-end manner. The system is able to learn human driving behavior without the need of any labeled training data. An action-based reward function is proposed, which is motivated by a potential use in real world reinforcement learning scenarios. Compared to a naive distance-based reward function, it improves the overall driving behavior of the vehicle agent. The agent is even able to reach comparable to human driving performance on a previously unseen track in our simulation environment.
Christian Hubschneider, Michael Weber 0009, André Bauer 0003, Jonathan Härtl, Fabian Duerr, Johann Marius Zöllner
Intelligent Vehicles Symposium7
2017 Estimating high definition map parameters with convolutional neural networks
abstract
In this paper, we present a method to estimate abstract parameters of high definition (HD) maps from sensor data. Parameters we estimate include the distance from ego-vehicle to road boundary, orientation of the ego-vehicle with respect to lanes, number of lanes, and street type. Our method is realized as a Convolutional Neural Network (CNN) that takes pre-processed sensor information in the form of grid map images as input. The estimated parameters of the network can then either be used for localization or to validate existing map data. To generate ground truth training samples, we use a semi-automatic procedure based on a good localization method to align the HD map with the sensor information from the vehicle. Our experiments yield some first promising results of the concept.
Sebastian Bittel, Timo Rehfeld, Michael Weber 0009, Johann Marius Zöllner
SMC4
2016 Vehicle pose estimation in cluttered urban environments using multilayer adaptive Monte Carlo localization
Jan Rohde, Inga Jatzkowski, Holger Mielenz, Johann Marius Zöllner
FUSION4
2016 Post processing of laser scanner measurements for testing advanced driver assistance systems
Jan Erik Stellet, Leopold Walkling, Johann Marius Zöllner
FUSION3
2016 Unsupervised Contextual Task Learning and Recognition for Sharing Autonomy to Assist Mobile Robot Teleoperation
abstract
We focus on the problem of learning and recognizing contextual tasks from human demonstrations, aiming to efficiently assist mobile robot teleoperation through sharing autonomy. We present in this study a novel unsupervised contextual task learning and recognition approach, consisting of two phases. Firstly, we use Dirichlet Process Gaussian Mixture Model (DPGMM) to cluster the human motion patterns of task executions from unannotated demonstrations, where the number of possible motion components is inferred from the data itself instead of being manually specified a priori or determined through model selection. Post clustering, we employ Sparse Online Gaussian Process (SOGP) to classify the query point with the learned motion patterns, due to its superior introspective capability and scalability to large datasets. The effectiveness of the proposed approach is confirmed with the extensive evaluations on real data.
Ming Gao 0005, Ralf Kohlhaas, Johann Marius Zöllner
ICINCO (2)3
2016 Localization accuracy estimation with application to perception design
abstract
Landmark-based localization in dynamic environments poses high demands on the perception system of a mobile robot. The pose estimate generally has to fulfill specific accuracy requirements which might be necessitated by dependent systems, such as behavior planning. Thus, in this contribution we focus on the model-based derivation of perception requirements, i.e. detectable landmark types and minimum detection rates, to enable global localization with a specified upper bound on uncertainty. To this end, we utilize stochastic geometry to accurately capture and explicitly consider characteristics of the dynamic environment (e.g. occlusions), and the perception system (e.g. missed detections). From this point our contributions are twofold: i) We propose an analytical model of upper bounds on localization uncertainty. For continuous pose tracking, the Kalman filter equations for intermittent observations are considered and ii) perception requirements, i.e. minimum detection rates, based on specified upper bounds on pose estimation uncertainty are derived. Monte Carlo simulations are used to demonstrate the performance of the proposed methods.
Jan Rohde, Jan Erik Stellet, Holger Mielenz, Johann Marius Zöllner
ICRA4
2016 Understanding interactions between traffic participants based on learned behaviors
abstract
Predicting vehicles' behaviors in a traffic scene can be very challenging due to many influences. Especially interactions with other traffic participants like vehicles or pedestrians are very crucial for the future movement while they are hard to model even with expert knowledge. In this paper we propose an object-oriented probabilistic approach that detects interactions between vehicles and is able to infer possible routes of traffic participants. Using the Object-Oriented Probabilistic Relational Modelling Language (OPRML), the interactions between vehicles can be modeled in an intuitive direct way. The probabilistic component allows Bayesian Inference on noisy sensor data and uncertain dependencies, while the object-orientation makes the model flexible to a varying number of traffic participants. Street-dependent as well as interaction-dependent motion models are learned from simulated situations and recordings of real traffic scenes. Finally, route prediction is evaluated at an exemplary intersection showing how the awareness of interactions reduces route prediction uncertainty and wrong predictions.
Florian Kuhnt, Jens Schulz, Thomas Schamm, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2016 Analytical derivation of performance bounds of autonomous emergency brake systems
abstract
Autonomous emergency brake (AEB) systems have to decide on brake interventions based on an uncertain and incomplete perception of the environment. This paper analyses theoretical limitations in AEB systems caused by noisy sensor measurements and uncertain prediction models. Such performance bounds can be used to derive sensor accuracy constraints, to identify challenging scenarios or to develop objective metrics. In contrast to most previous studies, this work focusses on analytical derivations. To this end, the Cramér-Rao bound of the best attainable state estimation covariance is derived from a model of sensor measurement errors. This state- and time-dependent covariance is then propagated to an AEB decision making logic that is based on a criticality measure. Additional inherent prediction uncertainty in this risk assessment is taken into account. The effectiveness of the AEB subject to uncertainties is compared to the deterministic baseline case in terms of the brake activation time and the collision energy reduction.
Jan Erik Stellet, Patrick Vogt, Jan Schumacher, Wolfgang Branz, Johann Marius Zöllner
Intelligent Vehicles Symposium5
2016 Functional system architectures towards fully automated driving
abstract
The functional system architecture of an automated vehicle plays a crucial role in the performance of the vehicle. When considered as a backbone, it does not only transmit information between distinct layers, but rather serves as a feedback mechanism coordinating the degradation between them and thereby regulates the behavior of the system against failures. Hence, the design of robust functional architectures is essential to cope with the uncertainties of the world. This paper summarizes existing system architectures and investigates them regarding their robustness against measurement inaccuracies, failures, and unexpected evolution of traffic situations. After illustrating their strengths and deficiencies, we derive the requirements and propose a structure for future, robust system architectures.
Ömer Sahin Tas, Florian Kuhnt, Johann Marius Zöllner, Christoph Stiller
Intelligent Vehicles Symposium3
2016 DeepTLR: A single deep convolutional network for detection and classification of traffic lights
abstract
Reliable real-time detection of traffic lights is a major concern for the task of autonomous driving. As deep convolutional networks have proven to be a powerful tool in visual object detection, we propose DeepTLR, a camera-based system for real-time detection and classification of traffic lights. Detection and state classification are realized using a single deep convolutional network. DeepTLR does not use any prior knowledge about traffic light locations. Also the detection is executed frame by frame without using temporal information. It is able to detect traffic lights on the whole camera image without any presegmentation. This is achieved by classifying each fine-grained pixel region of the input image and performing a bounding box regression on regions of each class. We show that our algorithm is able to run on frame-rates required for real-time applications while reaching notable results.
Michael Weber 0009, Johann Marius Zöllner
Intelligent Vehicles Symposium3
2016 Testing and validating high level components for automated driving: simulation framework for traffic scenarios
abstract
Current advances in the research field of autonomous driving demand advanced simulation methods for testing and validation. By combining versatile foci of different simulations, we can provide an increased amount and diversity of realistic traffic scenarios, which are relevant to the development and verification of high level automated driving functions. The focus of the present paper is to propose a concept for realistic simulation scenarios, which is capable of running in different integration levels, from software- to vehicle-in-the-loop. Its application is demonstrated, exposing an experimental vehicle, which is used for autonomous driving development, to a traffic scenario with virtual vehicles on a real road network.
Marc Rene Zofka, Sebastian Klemm, Florian Kuhnt, Thomas Schamm, Johann Marius Zöllner
Intelligent Vehicles Symposium5
2015 Towards a unified traffic situation estimation model - Street-dependent behaviour and motion models
Florian Kuhnt, Ralf Kohlhaas, Thomas Schamm, Johann Marius Zöllner
FUSION4
2015 Data-driven simulation and parametrization of traffic scenarios for the development of advanced driver assistance systems
Marc Rene Zofka, Florian Kuhnt, Ralf Kohlhaas, Christoph Rist, Thomas Schamm, Johann Marius Zöllner
FUSION6
2015 DSRC and radar object matching for cooperative driver assistance systems
abstract
Dedicated Short Range Communication (DSRC) systems will become ubiquitous among vehicles in the near future. Because this technology enables communication between any set of DSRC-equipped vehicles, precise knowledge of these other vehicles is available to the host car. In addition to the DSRC system, onboard radars are able to provide high fidelity dynamics measurements of other objects within the sensing range. Given these two methods of measurement, environmental perception for driver assistance systems can be greatly improved, especially if the measurements are fused together. However, this is not a trivial task because of an inherent data association problem: Given the objects detected by the radar sensor, which one is truly the DSRC message sender? In this paper, we propose a system architecture to fuse DSRC and radar data. This architecture uses a reliable statistical track-to-track association algorithm in a novel way to solve this data matching problem. We present experimental results of this architecture on a system running in real traffic situations in the U. S.
Qi Chen 0011, Jörg Hillenbrand, Axel Gern, Tobias Roth, Florian Kuhnt, Johann Marius Zöllner, Jakob Breu, Miro Bogdanovic, Christian Weiss
Intelligent Vehicles Symposium7
2015 Performance bounds on change detection with application to manoeuvre recognition for advanced driver assistance systems
abstract
Recognising the intended manoeuvres of other traffic participants is a crucial task for situation interpretation in driver assistance and autonomous driving. While many works propose algorithms for (computationally feasible) inference, much less attention is paid to finding analytic upper performance bounds for these problems. This work studies the statistical properties of the optimal detector in a binary change detection problem, i.e. the Generalised Likelihood Ratio test. With analytic models of the best attainable receiver operating characteristic, the influence of system design parameters can be investigated without the need for empirical evaluation. Moreover, these bounds can be used to derive objective performance metrics.
Jan Erik Stellet, Jan Schumacher, Wolfgang Branz, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2015 Uncertainty propagation in criticality measures for driver assistance
abstract
Active safety systems employ surround environment perception in order to detect critical driving situations. Assessing the threat level, e.g. the risk of an imminent collision, is usually based on criticality measures which are calculated from the sensor measurements. However, these metrics are subject to uncertainty. Probabilistic modelling of the uncertainty allows for more informed decision making and the derivation of sensor requirements. This work derives closed-form expressions for probability distributions of criticality measures under both state estimation and prediction uncertainty. The analysis is founded on uncertainty propagation in non-linear motion models. Finding the distribution of model-based criticality metrics is then performed using closed-form expressions for the collision probability and error propagation in implicit functions. All results are illustrated and verified in Monte-Carlo simulations.
Jan Erik Stellet, Jan Schumacher, Wolfgang Branz, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2014 Contextual task-aware shared autonomy for assistive mobile robot teleoperation
abstract
For robot applications in unknown or even hazardous environments, such as search and rescue, it is difficult and stressful for human beings to merely simply teleoperate a mobile robot without its assistance. Consequently, means to facilitate an efficient shared autonomy between human and robot are the subject of much research work in the field of robotics. This paper proposes a novel shared autonomy system, which recognizes the user intention by estimating the task the user is executing based on the context information, and provides motion assistance according to the inferences. To incorporate the uncertainty of contextual task recognition, a Gaussian Mixture Regression model combined with a recursive Bayesian filter is adopted, which is adaptive to the implicit user model for task execution during operation. The proposed method is applied to the problem of controlling a flying robot in the context of two task types: doorway crossing and object inspection. Its benefits are demonstrated by the simulation results.
Ming Gao 0005, Jan Oberländer, Thomas Schamm, Johann Marius Zöllner
IROS4
2014 Reactive posture behaviors for stable legged locomotion over steep inclines and large obstacles
abstract
Multi-legged walking robots often make use of sophisticated control architectures to play their strengths in rough and unknown environments. The adaptability of these robots is an essential skill to achieve the maneuverability and autonomy needed in their application fields. In this work we present a reactive control approach for the hexapod LAURONV, which enables it to overcome large obstacles and steep slopes without any knowledge about the environment. A key to this success can also be seen in the increased kinematic adaptability due to the fourth rotational joint in the bio-inspired leg kinematics. An extended experimental evaluation shows that the reactive posture behaviors are able to create an effective and efficient locomotion in challenging environments.
Arne Roennau, Georg Heppner, Michal R. Nowicki, Johann Marius Zöllner, Rüdiger Dillmann
IROS4
2014 A semantic approach to sensor-independent vehicle localization
abstract
As intelligent vehicles become more and more capable, they must learn to navigate and localize themselves in a wide variety of environments, including GPS-denied and only crudely mapped areas. We argue that since autonomous vehicles must be able to perceive, and semantically interpret, their immediate environment, they should be able to use abstract semantic information as their sole means of localization. This simplifies the level of detail and precision required from environment maps so that, for example, a rough floor plan of a parking garage will suffice to autonomously navigate it. We propose a concept for semantic localization which only requires a conceptual semantic map of the environment, and can be made to work with any kind of sensor data from which the required semantic information can be extracted. We present a localization algorithm which may be used as a base for semantic navigation, e.g. in context of automated driving, and some initial results of its application in a parking garage scenario.
Jan Oberländer, Sebastian Klemm, Marc Essinger, Arne Roennau, Thomas Schamm, Johann Marius Zöllner, Rüdiger Dillmann
Intelligent Vehicles Symposium6
2014 Semivirtual simulations for the evaluation of vision-based ADAS
abstract
The design and development process of advanced driver assistance systems (ADAS) is divided into different phases, where the algorithms are implemented as a model, then as software and finally as hardware. Since it is unfeasable to simulate all possible driving situations for environmental perception and interpretation algorithms, there is still a need for expensive and time-consuming real test drives of thousands of kilometers. Therefore we present a novel approach for testing and evaluation of vision-based ADAS, where reliable simulations are fused with recorded data from test drives to provide a task-specific reference model. This approach provides ground truth with much higher reliability and reproducability than real test drives and authenticity than using pure simulations and can be applied already in early steps of the design process. We illustrate the effectiveness of our approach by testing a vision-based collision mitigation system on recordings of a german highway.
Marc Rene Zofka, Ralf Kohlhaas, Thomas Schamm, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2013 Context aware shared autonomy for robotic manipulation tasks
abstract
This paper describes a collaborative human-robot system that provides context information to enable more effective robotic manipulation. We take advantage of the semantic knowledge of a human co-worker who provides additional context information and interacts with the robot through a user interface. A Bayesian Network encodes the dependencies between this information provided by the user. The output of this model generates a ranked list of grasp poses best suitable for a given task which is then passed to the motion planner. Our system was implemented in ROS and tested on a PR2 robot. We compared the system to state-of-the-art implementations using quantitative (e.g. success rate, execution times) as well as qualitative (e.g. user convenience, cognitive load) metrics. We conducted a user study in which eight subjects were asked to perform a generic manipulation task, for instance to pour a bottle or move a cereal box, with a set of state-of-the-art shared autonomy interfaces. Our results indicate that an interface which is aware of the context provides benefits not currently provided by other state-of-the-art implementations.
Thomas Witzig, Johann Marius Zöllner, Dejan Pangercic, Sarah Osentoski, Rainer Jäkel, Rüdiger Dillmann
IROS2
2013 Seen and missed traffic objects: A traffic object-specific awareness estimation
abstract
Handing-over vehicle control from a human driver to an intelligent vehicle and vice versa needs elaborate and safe hand-over strategies. Before passing control it must be ensured that the driver is aware of all objects which are important in a particular traffic situation. In this work a decision tree is used to learn which objects attract the driver's gaze in a particular situation. The decision tree classifies on object features as the object's type, velocity, size, color, and brightness. This information is fused from laser-scanners, front camera, and the vehicle's CAN-bus data. Whilst driving, an awareness confidence is built for each object perceived by the laser-scanners. Unexpected gaze behavior is detected by comparing the awareness confidence of each object to the expected gaze behavior, learned by means of the decision tree. Objects overlooked by the driver are further classified as critical or uncritical. This provides valuable information for following human-car interaction, augmented-reality, or safety applications.
Tobias Bär, Denys Linke, Dennis Nienhüser, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2013 Towards driving autonomously: Autonomous cruise control in urban environments
abstract
For automatic driving, vehicles must be able to recognize their environment and take control of the vehicle. The vehicle must perceive relevant objects, which includes other traffic participants as well as infrastructure information, assess the situation and generate appropriate actions. This work is a first step of integrating previous works on environment perception and situation analysis toward automatic driving strategies. We present a method for automatic cruise control of vehicles in urban environments. The longitudinal velocity is influenced by the speed limit, the curvature of the lane, the state of the next traffic light and the most relevant target on the current lane. The necessary acceleration is computed in respect to the information which is estimated by an instrumented vehicle.
Ralf Kohlhaas, Thomas Schamm, Dominik Lenk, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2013 Batch-mode active learning for traffic sign recognition
abstract
Cognitive vehicles face a huge variety of different objects when perceiving their environment using video cameras. Their capability of recognizing certain types of objects in a robust way is tightly coupled to successfully discriminating positive samples from negative ones (background objects). Machine learning methods trained on huge data sets can accomplish this, but to reduce labeling costs and classification runtime it is desirable to minimize the number of training samples needed. We examine the feasibility of active learning paradigms to achieve this goal for traffic sign recognition and propose a batch-mode multi-class active learning query strategy for support vector machines: Angular diversity ranking is a weighted combination of relevance of unlabeled samples with a diversity measure. A quantitative evaluation on a large data set with more than 30,000 samples shows clear benefits of our query strategy for traffic sign recognition, track-pruned class-biased angular diversity ranking, compared to uncertainty sampling based active learning as well as passive learning.
Dennis Nienhiiser, Johann Marius Zöllner
Intelligent Vehicles Symposium2
2011 Lane confidence fusion for visual occupancy estimation
abstract
Detection and tracking of road lanes is vital to a wide range of driver assistance systems. Confidence measures on tracked lanes are necessary to create reliable assistance systems. On curvy, multi-lane roads, adaptive cruise control systems could greatly benefit from information of lane location and curvature to identify the leading vehicle.We present a novel approach to estimating unobstructed lane length using visual lane confidence measures. The gradient between individually calculated lane confidence segments of a lane model provided by a lane tracker is used to estimate the distance to leading cars or obstacles on lanes. To achieve this goal, various confidence measures are defined, and a confidence fusion method based on Dempster Shafer is presented.
Thomas Gumpp, Dennis Nienhüser, Johann Marius Zöllner
Intelligent Vehicles Symposium3
2011 Traffic intersection situation description ontology for advanced driver assistance
abstract
This work provides an approach to create a generic situation description for advanced driver assistance systems using logic reasoning on a traffic situation knowledge base. It contains multiple objects of different type such as vehicles and infrastructure elements like roads, lanes, intersections, traffic signs, traffic lights and relations among them. Logic inference is performed to check and extend the situation description and interpret the situation e. g. by reasoning about traffic rules. The capabilities of our ontological situation description approach are shown at the example of complex intersections with several roads, lanes, vehicles and different combinations of traffic signs and traffic lights. Real-time issues are discussed thereon.
Michael Hülsen, Johann Marius Zöllner, Christian Weiss
Intelligent Vehicles Symposium2
2010 Design of an automotive traffic sign recognition system targeting a multi-core SoC implementation
abstract
This paper describes the design of an automotive traffic sign recognition application. All stages of the design process, starting on system-level with an abstract, pure functional model down to final hardware/software implementations on an FPGA, are shown. The proposed design flow tackles existing bottlenecks of today's system-level design processes, following an early model-based performance evaluation and analysis strategy, which takes into account hardware, software and real-time operating system aspects. The experiments with the traffic sign recognition application show, that the developed mechanisms are able to identify appropriate system configurations and to provide a seamless link into the underlying implementation flows.
Matthias Müller 0004, Axel G. Braun, Joachim Gerlach, Wolfgang Rosenstiel, Dennis Nienhüser, Johann Marius Zöllner, Oliver Bringmann 0001
DATE6
2010 Physical road marker property estimation using monoscopic vision
abstract
In this paper algorithms are presented to extract lane markers and their properties from monoscopic camera images. A filter approach that takes into account the visual appearance of the markers is presented. Another contribution constitutes the measurement of physical marker width and length from perspectively distorted images using only the calibrated camera images. Also distances between markers can be estimated. All algorithms are independent from a specific lane model.
Thomas Gumpp, Dennis Nienhüser, Johann Marius Zöllner
ICRA3
2010 Proactive avoidance of moving obstacles for a service robot utilizing a behavior-based control
abstract
A main challenge in the application of service robotics is safe and reliable navigation of robots in human everyday environments. Supermarkets, which are chosen here as an example, pose a challenging scenario because they usually have a cluttered and nested character. The robot has to avoid collisions with static and even with moving obstacles while interacting with nearby humans or a dedicated user respectively. This paper presents a hierarchical approach for the proactive avoidance of moving objects as it is used on the robot shopping trolley InBOT. The behavior-based control (bbc) of InBOT is extended by a reflex and a reactive behavior to ensure adequate reaction times when confronted with a possible collision. On top of the bbc a spatio-temporal planner is situated which is able to predict environmental changes and therefore can generate a safe movement sequence accordingly.
Michael Göller, Florian Steinhardt, Thilo Kerscher, Johann Marius Zöllner, Rüdiger Dillmann
IROS4
2010 Fast and reliable recognition of supplementary traffic signs
abstract
Supplementary traffic signs are used to alter the meaning of other traffic signs. Assistance systems that recognize traffic signs therefore must also recognize supplementary signs to evaluate their influence on the meaning of detected traffic signs. We propose an algorithm which is able to detect supplementary signs in the vicinity of other signs using a novel rectangle segmentation algorithm. Support vector machines are used for the classification and rejection of other objects. The combination of both components permits to recognize a supplementary sign in less than 40 ms. First quantitative results for a test set with four different supplementary sign types show a very good classification accuracy of more than 96%.
Dennis Nienhüser, Thomas Gumpp, Johann Marius Zöllner, Koba Natroshvili
Intelligent Vehicles Symposium3
2010 On-road vehicle detection during dusk and at night
abstract
The video-based on-road detection of vehicles at daytime allows driver assistance systems to avoid collisions and thereby improve safety, and realize comfort functions, like the well known adaptive cruise control. However, at nighttime, common video sensor based vehicle detection algorithms can't be used, because most state-of-the-art features, like shadows, symmetry and others, cannot be measured. The on-road detection of vehicles at night is an obligatory feature for modern driver assistance systems, because those systems have to provide assistance functionality at day-time and at night-time, either. In this work, vehicles in front of the own car are recognized by detection of their front or rear lights, using a perspective blob filter and subsequently searching for corresponding light pairs. For preceding vehicles, the activity of the third break light is estimated, to distinguish the maneuver state of the vehicle. Experiments show the robustness of the approach during dusk and at night sequences.
Thomas Schamm, Christoph von Carlowitz, Johann Marius Zöllner
Intelligent Vehicles Symposium3
2010 3D-segmentation of traffic environments with u/v-disparity supported by radar-given masterpoints
abstract
3D-segmentation of a traffic scene with two-dimensional row- and column-disparity-histograms, namely u/v-disparities, has become more and more popular for modern stereo-camera-based driver assistance systems due to its fast computation in real-time, few memory requirements and robustness against noisy or intermittent data. In this paper, we present a novel approach to support this pure vision-based method by projecting preprocessed radar-signals directly to u-disparity-space. We called the projection result “masterpoints”. This data fusion on low feature-level improved the segmentation process and increased the obstacle detection rate significantly. No assumptions about obstacle-type or -size are needed. Furthermore, the algorithms can be parallelized easily and run in real-time.
Michael Teutsch, Thomas Heger, Thomas Schamm, Johann Marius Zöllner
Intelligent Vehicles Symposium4
2008 A region-based SLAM algorithm capturing metric, topological, and semantic properties
abstract
This paper proposes a SLAM algorithm based on FastSLAM 2.0 that maps features representing regions with a semantic type, topological properties, and an approximative geometric extent. The resulting maps enable spatial reasoning on a semantic level and provide abstract information allowing efficient semantic planning and a convenient interface for human-machine interaction. We present novel region features and an algorithm for estimating the feature parameters from uncertain measurements. In particular, we provide a means of estimating parameters even if the region feature is considerably larger than the robot's sensor range. Finally, we adapt the FastSLAM 2.0 algorithm to map the proposed features and show simulation-based results illustrating the capabilities of the proposed algorithm.
Jan Oberländer, Klaus Uhl, Johann Marius Zöllner, Rüdiger Dillmann
ICRA3
2008 Dexterous manipulation planning of objects with surface of revolution
abstract
In this paper, we propose a novel method for dexterous manipulation planning problem of rotating object with surface of revolution using a robotic multi-fingered hand. This method finds contact point trajectories from contact points between the robotic hand and the object with task-orientated manipulation quality measurement. Based on the defined manipulation quality, the pose for robotic hand relative to object can also be optimized by random sample. Experiments using Schunk anthropomorphic hand with 13 degrees of freedom screwing a light bulb into holder with screw thread demonstrates the feasibility and efficiency of the introduced method.
Zhixing Xue, Johann Marius Zöllner, Rüdiger Dillmann
IROS2
2007 Using case-based reasoning for autonomous vehicle guidance
abstract
Vehicle guidance in complex scenarios such as inner-city traffic requires an in-depth understanding of the current situation. In order to select the appropriate behavior for an autonomous vehicle, an analysis of the situation is needed. The analysis consists of an estimation of the situation's development with respect to the selected behavior. This can only be done using higher-level reasoning techniques. In this paper, an approach for situation interpretation for autonomous vehicles is presented. The approach relies on case-based reasoning in order to predict the evolvement of the current situation and to select the appropriate behavior. Case-based reasoning allows to utilize prior experiences in the task of situation assessment.
Stefan Vacek, Tobias Gindele, Johann Marius Zöllner, Rüdiger Dillmann
IROS3
2005 Localization of Walking Robots
abstract
Proper navigation of walking machines in unstructured terrain requires the knowledge of the spatial position and orientation of the robot. There are many approaches for localization of mobile robots in outdoor environment, but their application to walking robots is rather rare. In particular, middle sized robots like LAURON III don’t provide the possibility to carry large or heavy sensors. Due to many degrees of freedom of walking robots the localization task becomes a even more complex challenge. This paper discusses the problem, presents a method of resolution and describes the first steps towards a localization system for the six-legged walking robot LAURON III.
Bernd Gaßmann, Franziska Zacharias, Johann Marius Zöllner, Rüdiger Dillmann
ICRA3
2005 Compliant motion of a multi-segmented inspection robot
abstract
This paper presents a method to potentate the multi-segmented inspection robot Kairo-II to navigate in unstructured and dynamic environment. Previous methods for motion planning for such robots come from driving scenarios in highly structured areas. The virtual tube algorithm is introduced which enables a multi-segmented robot to range in such complex environment. Precise force feedback is required. Therefore, we present a sensor system which is based on strain-gauges technology. Information extracted by this sensor enables the trajectory planning algorithm to adapt its curve. Thus, the proposed system provides and evaluates key functions for compliant motion of a multi-segmented robot within unstructured environment.
Clemens Birkenhofer, Michael Hoffmeister, Johann Marius Zöllner, Rüdiger Dillmann
IROS3
2002 Understanding users intention: programming fine manipulation tasks by demonstration
abstract
The Programming by Demonstration (PbD) paradigm enable programming of service robots by inexperienced human users. The main goal of these systems is to allow the inexperienced human user to easily integrate motion and perception skills or complex problem solving strategies. Unfortunately, actual PbD systems deal only with manipulation based on Pick & Place operations. For complex service tasks these are insufficient. Therefore, this paper describes how fine manipulations like detecting screw movements can be recognized by a PbD system. In order to do this, finger movements and forces on the fingertips are gathered and analyzed while an object is grasped. This assumes sensory employment like a data glove and integrated tactile sensors. An overview of the used tactile sensors and the gathered signals is given. Furthermore the segmentation of users demonstration and the classification of the recognized dynamic grasp is pointed out. For classifying dynamic grasps a time delay method based on a Support Vector Machine (SVM) is used. Finally the symbolic representation of service tasks is briefly illustrated.
Raoul Daniel Zöllner, Oliver Rogalla, Rüdiger Dillmann, Johann Marius Zöllner
IROS4
2000 Learning methods for online-process diagnosis
abstract
Because of the very high workpiece costs in manufacturing processes, production errors should be detected online in order to avoid a series of defective workpieces. This article describes a qualitative evaluation method for time series that is applied to the diagnosis of a procedure for spraying car body parts. The determination of the parameters for the procedure is gained through learning data, which simplifies the industrial use enormously. A prototype that is already employed in production confirms the expected functionality of the procedure.
Patrick Feucht, Johann Marius Zöllner, Karsten Berns, Torsten Zirzlaff, Oskar Leisin
ICTAI2